Battery pack equalization control method and system based on deep reinforcement learning for vehicle
Patent Information
- Application Number
- CN202610894292.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2046-06-22
AI Technical Summary
[0004]本发明的主要目的在于提供了一种基于深度强化学习的车用电池组均衡控制方法及系统,旨在解决现有技术难以在复杂车载工况下,低延迟、自适应地实现电池组荷电状态主动均衡控制,无法满足车载系统低延迟和高实时性的要求的技术问题
[0015]本发明通过结合相邻单体电池连接关系、单体电池运行数据和荷电状态序列构建数字孪生均衡训练环境,使智能体训练同时受到电池动态特性、相邻均衡通道约束和荷电状态演化规律的限定;通过在数字孪生均衡训练环境中训练近端策略优化智能体,并将荷电状态序列作为状态观测数据、将相邻均衡通道的均衡电流控制量作为动作数据,使智能体的均衡策略网络能够学习不同荷电状态分布下的电流调节规律;通过将目标电池组实时荷电状态序列输入训练后的近端策略优化智能体并生成目标动作序列,使均衡控制量能够随目标电池组当前状态自适应变化;通过嵌入式均衡执行单元驱动相邻主动均衡电路并回传执行反馈数据,使智能决策结果能够转化为硬件侧闭环执行,从而面向车载动态工况生成可执行的均衡电流控制量,实现车载电池组荷电状态主动均衡控制,满足车载系统低延迟、高实时性的部署要求,有效地降低了车载单体电池过充、过放及热失控风险。
Smart Images

Figure CN122402316B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of lithium battery equalization control technology and management system, and in particular to a method and system for equalization control of automotive battery packs based on deep reinforcement learning. Background Technology
[0002] Against the backdrop of the rapid development and continuous advancement of the new energy vehicle industry, lithium-ion batteries have become an important energy storage unit for electric vehicle power systems due to their advantages such as high energy density and long cycle life. Actual automotive power battery packs are usually composed of multiple cells connected in series. However, due to factors such as differences in manufacturing processes, capacity deviations, inconsistent internal resistance, uneven temperature distribution, and long-term aging, inconsistencies in the SOC (State of Charge) of individual cells can easily occur.
[0003] Inconsistent SOC (State of Charge) reduces the usable capacity of the battery pack and increases the risk of overcharging, over-discharging, and thermal runaway of individual cells. Currently, there are still many unresolved issues in the field of automotive battery balancing: traditional balancing control strategies are highly dependent on battery models and fixed thresholds, and lack robustness in the face of dynamic operating conditions and battery aging; most reinforcement learning-based balancing schemes remain only at the simulation level, lacking integration and verification with actual hardware systems; at the same time, existing reinforcement learning algorithms are generally highly complex, making it difficult to meet the deployment requirements of low latency and high real-time performance in automotive systems. Summary of the Invention
[0004] The main objective of this invention is to provide a method and system for balancing battery packs in vehicles based on deep reinforcement learning. This aims to solve the technical problem that existing technologies are unable to achieve low-latency and adaptive active balancing control of battery pack state of charge under complex vehicle operating conditions, thus failing to meet the requirements of low latency and high real-time performance of vehicle systems.
[0005] To achieve the above objectives, this invention provides a method for equalization control of automotive battery packs based on deep reinforcement learning, the method comprising the following steps: The state of charge sequence of each battery pack is generated based on the individual cell operating data of multiple battery packs in the vehicle battery system. Each battery pack includes multiple individual cells connected in series. The individual cell operating data includes voltage data, current data and temperature data. A digital twin equalization training environment is constructed based on the connection relationship between adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence. The proximal policy optimization agent is trained based on the digital twin equalization training environment. The proximal policy optimization agent uses the charge state sequence as state observation data and the equalization current control quantity of adjacent equalization channels as action data. The real-time state of charge sequence of the target battery pack is input into the trained proximal policy optimization agent to obtain the target action sequence and generate adjacent equalization control instructions. The adjacent equalization control instructions are sent to the embedded equalization execution unit so that the embedded equalization execution unit drives the adjacent active equalization circuit to perform adjacent equalization operation and transmits execution feedback data back in real time to form closed-loop control.
[0006] Optionally, the step of constructing a digital twin equalization training environment based on the connection relationship of adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence includes: Based on the single cell operating data, the equivalent circuit model parameters of each series-connected single cell are identified, and the second-order resistance-capacitance equivalent circuit model of each series-connected single cell is configured based on the equivalent circuit model parameters. Based on the connection relationship of multiple series-connected individual cells in each battery pack, the second-order resistive-capacitive equivalent circuit model is connected to obtain the digital twin model of the battery pack. Adjacent equalization channel mapping data is generated based on the connection relationship between adjacent individual cells in each battery pack. The adjacent equalization channel mapping data includes the mapping relationship between adjacent equalization channels and adjacent individual cell pairs. An adjacent equilibrium model is constructed based on the adjacent equilibrium channel mapping data. A digital twin equilibrium training environment is constructed based on the battery pack digital twin model, the neighbor equilibrium model, and the state of charge sequence.
[0007] Optionally, training the proximal policy optimization agent based on the digital twin balanced training environment includes: The state of charge sequence is obtained by arranging the series-connected individual cells in each battery pack according to their series order; The equalization current control quantities of adjacent equalization channels are arranged based on the mapping data of adjacent equalization channels to obtain a continuous action vector; Based on the difference between the maximum and minimum state of charge values in the state observation vector, the state of charge difference data is obtained. Charge state difference penalty data is generated based on the charge state difference data and the first preset weight; Action amplitude penalty data is generated based on the sum of squares of the duty cycle parameters in the continuous action vector and the second preset weight; When the difference in state of charge is less than a preset convergence threshold, equal convergence state reward data is generated. A comprehensive reward function is constructed based on the state of charge difference penalty data, the action amplitude penalty data, and the equilibrium convergence state reward data. The comprehensive reward function is based on the following formula: in, This represents the comprehensive reward data. This data represents the difference in state of charge (SOC), used to measure the degree of inconsistency in the SOC of a battery pack. This represents the reward data for the equilibrium convergence state. Indicates the first Duty cycle parameters of adjacent equalization channels Indicates the number of adjacent equalization channels; The proximal policy optimization agent is trained based on the state observation vector, the continuous action vector, the comprehensive reward function, and the digital twin balanced training environment.
[0008] Optionally, training the proximal policy optimization agent based on the state observation vector, the continuous action vector, the comprehensive reward function, and the digital twin balanced training environment includes: Based on the near-end policy optimization algorithm, a policy network and a value network are constructed for the near-end policy optimization agent. The policy network is used to output the distribution data of the equalization current control quantity of adjacent equalization channels, and the value network is used to output the state value estimation data. Based on the comprehensive reward function, comprehensive reward data is generated. Then, based on the comprehensive reward data, the preset discount factor, and the temporal relationship of each training moment within the training round, cumulative reward target data is generated, as shown in the following formula: in, This represents the cumulative reward target data. The network parameters represent the policy network. This represents the expectation of the sample data during the training process. This represents the maximum training time within a single training round. This indicates the current training moment within a training round. Indicates the preset discount factor. This represents the instantaneous reward data at the t-th training time. The state observation vector is input into the policy network to obtain the current equalization current control quantity distribution data; The state observation vector is input into the value network to obtain state value estimation data; Advantage assessment data is generated based on the comprehensive reward data and the state value estimation data; Based on the current and historical balanced current control distribution data, strategy ratio data is generated, referring to the following formula: in, This represents the policy ratio data at training time t. This indicates that the current policy network is in the state observation data. Select action data The probability, This indicates that the historical policy network is based on state observation data. Select action data The probability, This represents the policy network parameters from the previous training round. This represents the state observation data at the t-th training time. This represents the equalization current control action at the t-th training moment; The strategy ratio data is subjected to amplitude limiting update processing based on a preset clipping factor to obtain amplitude limiting strategy update data. The amplitude limiting update processing refers to the following formula: in, This represents the value of the objective function for pruning. This represents the expectation of the sample data at each training time step. Let represent the advantage evaluation data at training time t, used to characterize the superiority or inferiority of the balanced current control action relative to the current state value estimation data. Represents the amplitude limiting function. Indicates the preset clipping factor; Based on the cumulative reward target data, the limit policy update data, and the advantage evaluation data, the policy network and value network are updated to obtain the trained proximal policy optimization agent.
[0009] Optionally, the step of inputting the real-time state of charge sequence of the target battery pack into the trained proximal policy optimization agent to obtain the target action sequence includes: The target battery pack is determined from the plurality of battery packs based on preset equalization task information; Read the real-time state of charge sequence of the target battery pack from the state of charge sequence; The real-time state of charge sequence is vectorized according to the series connection order of multiple series-connected individual cells in the target battery pack to obtain the real-time state observation vector. The real-time state observation vector is subjected to state constraint processing based on the preset upper limit value of the charged state and the preset upper limit value of the charged state to obtain the constrained state observation vector. The constrained state observation vector is input into the trained proximal policy optimization agent to obtain the initial action sequence; The initial action sequence is subjected to action constraint processing based on the preset lower limit value and the preset upper limit value of the balanced current control quantity to obtain the target action sequence.
[0010] Optionally, the step of generating adjacent equalization control instructions, sending the adjacent equalization control instructions to the embedded equalization execution unit, so that the embedded equalization execution unit drives the adjacent active equalization circuit to perform adjacent equalization operations, and transmits execution feedback data back in real time to form closed-loop control, includes: The target action sequence is broken down to obtain multiple channel equalization current control quantities; Based on the adjacent equalization channel mapping data, the equalization current control quantities of the multiple channels are mapped to the multiple adjacent equalization channels to obtain channel control data; Equalization drive control data is generated based on the channel control data; Based on the channel control data and the equalization drive control data, generate adjacent equalization control commands; The adjacent equalization control command is sent to the embedded equalization execution unit; The system receives execution feedback data from the embedded equalization execution unit based on the adjacent equalization control command, and forms closed-loop control based on the execution feedback data.
[0011] Optionally, receiving the execution feedback data returned by the embedded equalization execution unit based on the adjacent equalization control command, and forming closed-loop control based on the execution feedback data, includes: Receive channel execution status data returned by the embedded equalization execution unit; Receive voltage sampling feedback data, current sampling feedback data, and temperature sampling feedback data transmitted back by the embedded equalization execution unit; The action generation process of the trained proximal policy optimization agent and the transmission process of adjacent equalization control commands are monitored to obtain total control delay data. When the total control delay data is not greater than the preset total delay threshold, the channel execution status data, the voltage sampling feedback data, the current sampling feedback data and the temperature sampling feedback data are time aligned to obtain aligned feedback data. The real-time state of charge sequence of the target battery pack is updated based on the alignment feedback data to obtain the updated real-time state of charge sequence. The target action sequence is regenerated based on the updated real-time state of charge sequence, and closed-loop control is formed based on the regenerated target action sequence.
[0012] Optionally, the monitoring of the action generation process and the transmission process of the neighboring equalization control command of the trained proximal policy optimization agent to obtain total control delay data includes: The monitoring and training of the proximal policy optimization agent generates inference delay data corresponding to the target action sequence; Monitor the instruction transmission delay data sent from the adjacent equalization control command to the embedded equalization execution unit; Receive hardware execution latency data returned by the embedded equalization execution unit; The total control latency data is generated based on the inference latency data, the instruction transmission latency data, and the hardware execution latency data, referring to the following formula: in, Indicates total control delay. This represents the agent inference delay caused by the near-end policy optimization agent generating the target action sequence based on the real-time charge state sequence. Indicates instruction transmission delay. This indicates hardware execution latency.
[0013] Furthermore, to achieve the above objectives, this invention also proposes a deep reinforcement learning-based vehicle battery pack equalization control system that applies the deep reinforcement learning-based vehicle battery pack equalization control method described above. The deep reinforcement learning-based vehicle battery pack equalization control system includes: The data acquisition module is used to generate a state of charge sequence for each battery pack based on the operating data of individual cells in multiple battery packs in the vehicle battery system. Each battery pack includes multiple individual cells connected in series. The operating data of the individual cells includes voltage data, current data, and temperature data. The environment construction module is used to construct a digital twin equalization training environment based on the connection relationship between adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence. The agent training module is used to train a proximal policy optimization agent based on the digital twin equalization training environment. The proximal policy optimization agent uses the charge state sequence as state observation data and the equalization current control quantity of adjacent equalization channels as action data. The equalization control module is used to input the real-time state of charge sequence of the target battery pack into the trained proximal policy optimization agent to obtain the target action sequence, generate adjacent equalization control instructions, and send the adjacent equalization control instructions to the embedded equalization execution unit so that the embedded equalization execution unit drives the adjacent active equalization circuit to perform adjacent equalization operation and transmits execution feedback data back in real time to form closed-loop control.
[0014] Furthermore, to achieve the above objectives, this application also proposes a vehicle battery pack balancing control device based on deep reinforcement learning. The device includes: a memory, a processor, and a vehicle battery pack balancing control program stored in the memory. The processor is used to run the vehicle battery pack balancing control program, and the computer program is configured to implement the steps of the vehicle battery pack balancing control method based on deep reinforcement learning as described above.
[0015] This invention constructs a digital twin balancing training environment by combining the connection relationships of adjacent individual batteries, individual battery operating data, and state of charge (SOC) sequences. This allows the agent's training to be simultaneously constrained by battery dynamic characteristics, adjacent balancing channels, and the evolution law of SOC. By training a proximal policy optimization agent in the digital twin balancing training environment, using the SOC sequence as state observation data and the balancing current control quantities of adjacent balancing channels as action data, the agent's balancing policy network can learn the current regulation law under different SOC distributions. By inputting the real-time SOC sequence of the target battery pack into the trained proximal policy optimization agent and generating a target action sequence, the balancing control quantity can adaptively change with the current state of the target battery pack. By driving adjacent active balancing circuits through an embedded balancing execution unit and transmitting execution feedback data, the intelligent decision-making result can be transformed into hardware-side closed-loop execution, thereby generating executable balancing current control quantities for vehicle dynamic conditions. This achieves active SOC balancing control of the vehicle battery pack, meeting the deployment requirements of low latency and high real-time performance for vehicle systems, and effectively reducing the risks of overcharging, over-discharging, and thermal runaway of individual vehicle batteries. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the structure of a vehicle battery pack equalization control device based on deep reinforcement learning, which is part of the hardware operating environment of the embodiment of the present invention. Figure 2 This is a flowchart illustrating the first embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning according to the present invention. Figure 3 This is a flowchart illustrating the second embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning of the present invention. Figure 4 This is a structural block diagram of the first embodiment of the vehicle battery pack balancing control system based on deep reinforcement learning of the present invention; Figure 5The graph shows the variation curve of the individual cell voltage difference under the PPO equalization control strategy designed using the method of this invention. Figure 6 The curve showing the maximum SOC difference of the battery pack under the PPO equalization control strategy designed using the method of this invention; Figure 7 The simulation curve of SOC equalization of a six-cell battery under the PPO equalization control strategy designed using the method of this invention; Figure 8 This is a schematic diagram comparing the convergence times of the PPO equilibrium control strategy designed based on the method of this invention. Figure 9 This is a schematic diagram comparing the equalization accuracy of the PPO equalization control strategy designed based on the method of this invention. Figure 10 This is a schematic diagram comparing the energy loss of the PPO equalization control strategy designed based on the method of this invention. Figure 11 This is a comparative diagram showing the battery capacity improvement achieved by the PPO equalization control strategy designed based on the method of this invention.
[0018] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0019] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0020] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a vehicle battery pack equalization control device based on deep reinforcement learning, which is part of the hardware operating environment of the embodiment of the present invention.
[0021] like Figure 1As shown, the deep reinforcement learning-based vehicle battery pack balancing control device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen and an input unit such as a keyboard; the user interface 1003 may also include standard wired and wireless interfaces. The network interface 1004 may optionally include standard wired and wireless interfaces (such as Wireless-Fidelity (Wi-Fi) interfaces). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.
[0022] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the deep reinforcement learning-based vehicle battery pack equalization control device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0023] like Figure 1 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a vehicle battery pack balancing control program.
[0024] exist Figure 1 In the deep reinforcement learning-based vehicle battery pack balancing control device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the deep reinforcement learning-based vehicle battery pack balancing control device of the present invention can be set in the deep reinforcement learning-based vehicle battery pack balancing control device. The deep reinforcement learning-based vehicle battery pack balancing control device calls the vehicle battery pack balancing control program stored in the memory 1005 through the processor 1001 and executes the deep reinforcement learning-based vehicle battery pack balancing control method provided in the embodiment of the present invention.
[0025] This invention provides a method for equalization control of automotive battery packs based on deep reinforcement learning, referring to... Figure 2 , Figure 2This is a flowchart illustrating the first embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning according to the present invention.
[0026] In this embodiment, the deep reinforcement learning-based vehicle battery pack equalization control method includes the following steps: Step S10: Generate the state of charge sequence of each battery pack based on the single cell operation data of multiple battery packs in the vehicle battery system.
[0027] It should be understood that the executing entity of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a battery management controller, vehicle edge controller, embedded controller, or vehicle control device with data processing capabilities, or a terminal electronic device capable of realizing the above functions. The following description uses a deep reinforcement learning-based vehicle battery pack balancing control device (hereinafter referred to as the control device) as an example to illustrate this embodiment and the following embodiments.
[0028] It should be noted that automotive batteries can include multiple battery packs, and each battery pack can include multiple individual cells connected in series. These battery packs can correspond to different battery modules within a power battery pack, or to series-connected battery cells in different areas of the vehicle. Individual cell operating data includes voltage, current, and temperature data.
[0029] It should be noted that single-cell operating data refers to sampled data reflecting the electrical and thermal states of a single cell, including voltage, current, and temperature data. State of charge (SCC) sequence refers to a set of SCC data arranged according to the connection order of individual cells within the battery pack. Series-connected cells refer to individual cells connected in series within the battery pack to form the target output voltage.
[0030] In practical implementation, the control device can acquire the terminal voltage of each series-connected individual battery cell through a voltage sampling circuit, acquire the current in the battery pack's charging and discharging circuit through a current detection circuit, and acquire the surface temperature of the individual battery cell or the internal temperature of the battery module through a temperature sensor. It then processes all types of sampling data synchronously according to the sampling time. The control device can estimate the state of charge (SOC) of each individual battery cell at continuous sampling times by combining the initial SOC, rated capacity, and cumulative current change. It then arranges the SOCs according to the series connection order of the individual cells in the battery pack, generating a SOC sequence corresponding to each battery pack. For example, under conditions such as vehicle acceleration discharge, regenerative braking, end-of-charge charging, and resting recovery, it can continuously generate SOC sequences reflecting individual cell differences.
[0031] In one embodiment, voltage, current, and temperature data of each series-connected cell in multiple battery packs within a vehicle battery are collected; the voltage, current, and temperature data are time-aligned according to the sampling time to obtain time-series operational data; coulomb counting is performed on the current data and a preset rated capacity value in the time-series operational data to obtain initial state of charge (SOC) estimation data; the initial SOC estimation data is corrected based on the voltage and temperature data in the time-series operational data to obtain corrected SOC estimation data; and the corrected SOC estimation data is limited based on a preset SOC lower limit and a preset SOC upper limit to obtain the SOC sequence of each battery pack.
[0032] Step S20: Construct a digital twin equalization training environment based on the connection relationship between adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence.
[0033] It should be noted that the connection relationship of adjacent individual cells refers to the electrical connection sequence and topological relationship between two adjacent series-connected individual cells within a battery pack. The digital twin equalization training environment refers to a virtual training environment used to simulate changes in battery pack state, adjacent equalization responses, agent action execution, and training feedback. The adjacent equalization channel refers to the control channel connecting adjacent individual cells and used to perform equalization current regulation.
[0034] In practical implementation, the control device can construct a virtual battery model based on the operating data of individual cells, which can reflect the voltage response, state of charge change and temperature change of individual cells, and combine them according to the connection order of multiple series-connected individual cells in the battery pack to form a battery pack-level virtual model.
[0035] The control device can also establish an adjacent equalization response model based on the connection relationship between adjacent individual cells, enabling the training environment to simulate energy transfer or charge regulation between adjacent individual cells according to the equalization current control quantity of adjacent equalization channels. The digital twin equalization training environment can be configured with state input interfaces, action input interfaces, state update interfaces, and reward feedback interfaces. For example, under conditions such as low-temperature start-up, high-rate discharge, end of charging, or battery aging, the training environment can generate training samples based on different initial state of charge distributions and simulate the impact of adjacent equalization actions on the state of charge sequence.
[0036] Step S30: Train the proximal policy optimization agent based on the digital twin balanced training environment.
[0037] It should be noted that the proximal policy optimization agent uses the aforementioned state-of-charge sequence as state observation data and the equalization current control quantities of adjacent equalization channels as action data. The proximal policy optimization agent refers to a reinforcement learning control unit constructed and trained based on the Proximal Policy Optimization (PPO) algorithm. State observation data refers to the input data used by the proximal policy optimization agent to determine the current equalization state of the battery pack. Equalization current control quantities refer to the action data used to characterize the magnitude of the equalization current or normalized equalization current of adjacent equalization channels. Action data refers to the control data output by the proximal policy optimization agent that can be mapped to adjacent active equalization circuits.
[0038] It should be understood that this step uses the state of charge sequence as state observation data and the equalization current control quantity of adjacent equalization channels as action data, so that the learning object of the agent directly corresponds to the battery pack consistency state and hardware equalization execution quantity, reducing the dependence on fixed threshold control and manual parameter tuning.
[0039] In a practical implementation, the control device can construct an agent including a policy network and a value network based on a near-end policy optimization algorithm. The policy network is used to output the distribution of equalization current control quantities corresponding to each adjacent equalization channel according to the charge state sequence, and the value network is used to evaluate the state value corresponding to the current charge state sequence.
[0040] During training, the digital twin equalization training environment inputs the state of charge sequence to the proximal policy optimization agent, which outputs the equalization current control quantity of adjacent equalization channels. The digital twin equalization training environment updates the battery pack state according to the equalization current control quantity and generates training feedback based on the state of charge difference, equalization current amplitude, and equalization convergence state.
[0041] After multiple training rounds, the trained proximal policy optimization agent is obtained. As an example, the equalization current control quantity can be normalized data, which is then mapped to the actual current value; the equalization current control quantity can also be directly the target equalization current magnitude.
[0042] Step S40: Input the real-time state of charge sequence of the target battery pack into the trained proximal policy optimization agent to obtain the target action sequence, and generate the neighbor equalization control instruction. Send the neighbor equalization control instruction to the embedded equalization execution unit so that the embedded equalization execution unit drives the neighbor active equalization circuit to perform the neighbor equalization operation and transmits the execution feedback data back in real time to form a closed-loop control.
[0043] It should be noted that the target battery pack refers to the battery pack that currently requires equalization control. The real-time state of charge sequence refers to the state of charge sequence of the target battery pack at the current control moment. The target action sequence refers to the set of action data output by the trained proximal policy optimization agent based on the real-time state of charge sequence, which may include the target equalization current control quantities of each adjacent equalization channel.
[0044] Adjacent equalization control commands refer to the control information used to drive adjacent equalization channels to perform equalization operations. These commands may include equalization channel identifiers, target equalization current values, drive control parameters, and control cycle parameters. The embedded equalization execution unit can be a hardware control unit configured with a microcontroller, communication interface, and drive circuitry. Execution feedback data refers to feedback data such as channel status, voltage, current, and temperature transmitted back during the equalization operation.
[0045] It should be understood that this step converts the output of the trained agent into neighboring equalization control instructions and drives the neighboring active equalization circuit through the embedded equalization execution unit, so that the equalization strategy obtained by offline training can enter the hardware execution link; at the same time, it updates the subsequent control input by executing feedback data, thereby improving the real-time performance and execution reliability of equalization control.
[0046] In practice, the control device can determine the target battery pack based on the balancing task information, the difference in state of charge, or the operating status of the battery pack, and read the real-time state of charge sequence of the target battery pack at the current control moment.
[0047] The control device inputs the real-time state of charge sequence into the trained proximal policy optimization agent to obtain the target action sequence, and generates adjacent equalization control commands based on the target equalization current control quantities of each adjacent equalization channel in the target action sequence.
[0048] After receiving the adjacent equalization control command, the embedded equalization execution unit drives the adjacent active equalization circuit to perform equalization operations between adjacent individual cells, and sends back the channel execution status, voltage sampling feedback, current sampling feedback and temperature sampling feedback in real time.
[0049] The control device updates the real-time state of charge sequence of the target battery pack based on the execution feedback data, and continues to input the trained proximal policy optimization agent to form a closed-loop control of state perception, intelligent decision-making, hardware execution and feedback update.
[0050] This embodiment constructs a digital twin balancing training environment by combining the connection relationships of adjacent individual cells, individual cell operating data, and state of charge (SOC) sequences. This allows the agent's training to be simultaneously constrained by battery dynamic characteristics, adjacent balancing channels, and the evolution law of SOC. By training the proximal policy optimization agent in the digital twin balancing training environment, and using the SOC sequence as state observation data and the balancing current control quantities of adjacent balancing channels as action data, the agent's balancing policy network can learn the current regulation law under different SOC distributions. By inputting the real-time SOC sequence of the target battery pack into the trained proximal policy optimization agent and generating the target action sequence, the balancing control quantity can adaptively change with the current state of the target battery pack. By driving adjacent active balancing circuits through an embedded balancing execution unit and transmitting execution feedback data, the intelligent decision result can be transformed into hardware-side closed-loop execution, thereby generating executable balancing current control quantities for vehicle dynamic conditions. This achieves active SOC balancing control of the vehicle battery pack, meets the deployment requirements of low latency and high real-time performance of vehicle systems, and effectively reduces the risks of overcharging, over-discharging, and thermal runaway of individual vehicle batteries.
[0051] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning of the present invention.
[0052] Based on the first embodiment described above, in this embodiment, step S20 further includes: Step S201: Identify the equivalent circuit model parameters of each series-connected individual cell based on the single cell operation data, and configure the second-order resistance-capacitance equivalent circuit model of each series-connected individual cell based on the equivalent circuit model parameters.
[0053] It should be noted that equivalent circuit model parameters refer to model parameters used to characterize the voltage response characteristics of a single battery cell, and may include open-circuit voltage parameters, ohmic resistance parameters, polarization resistance parameters, and polarization capacitance parameters. A second-order resistive-capacitive equivalent circuit model refers to an equivalent model that describes the dynamic response of a single battery cell using an open-circuit voltage source, an ohmic resistor, and two sets of resistive-capacitive branches.
[0054] In practice, the control equipment samples and synchronizes the operating data of individual batteries, removes outliers, and classifies operating conditions. It also identifies parameters based on voltage, current, and temperature data under charging, discharging, resting, and temperature change conditions.
[0055] For example, open-circuit voltage parameters can be calibrated based on voltage changes during the resting phase, ohmic resistance parameters can be identified based on voltage jumps before and after current steps, and polarization resistance and polarization capacitance parameters can be identified based on the slow voltage recovery process. After parameter identification, the equivalent circuit model parameters are written into the second-order resistive-capacitive equivalent structure to form the second-order resistive-capacitive equivalent circuit model corresponding to each series-connected single cell.
[0056] Step S202: Based on the connection relationship of multiple series-connected individual cells in each battery pack, connect the second-order resistive-capacitive equivalent circuit model to obtain the digital twin model of the battery pack.
[0057] It should be noted that the connection relationship refers to the electrical connection sequence between multiple series-connected individual cells in a battery pack, which may include the adjacent cell numbers, the positive and negative electrode connection directions, and the battery pack port relationships. A battery pack digital twin model is a virtual model corresponding to the actual battery pack topology and operating state, used to simulate the battery pack terminal voltage, individual cell state of charge, and state changes during the equalization process.
[0058] It should be understood that this embodiment combines the second-order resistor-capacitor equivalent circuit models at the individual cell level according to the series connection relationship, so that the digital twin model can reflect the terminal voltage superposition, individual cell difference transmission and state of charge evolution relationship at the battery pack level, and provide a battery pack simulation object for equalization strategy training.
[0059] In practical implementation, the control device can sequentially connect multiple second-order resistive-capacitive equivalent circuit models according to the numbering order of the individual cells in the battery pack, forming a virtual battery pack with the same series topology as the actual battery pack. For multiple battery packs, the control device can establish corresponding virtual battery pack models for each, and configure independent ports, sampling nodes, and status output interfaces for each virtual battery pack.
[0060] For example, when there are multiple modules in a vehicle's power battery pack, multiple battery pack digital twin models can be established on a module-by-module basis. Each battery pack digital twin model outputs the terminal voltage, state of charge, and temperature state of the corresponding individual battery cells, thereby obtaining the battery pack digital twin model.
[0061] Step S203: Generate adjacent equalization channel mapping data based on the connection relationship between adjacent individual cells in each battery pack.
[0062] It should be noted that adjacent cell pairs refer to two individual cells in a battery pack that are connected in adjacent positions. Adjacent equalization channel mapping data refers to the mapping relationship between adjacent equalization channels and adjacent cell pairs, which may include the equalization channel number, the cell number corresponding to the input end, the cell number corresponding to the output end, and the channel control interface number.
[0063] It should be understood that this step generates adjacent equalization channel mapping data so that the actions output by the subsequent agent can be accurately assigned to the corresponding adjacent equalization channels, avoiding the problem of inconsistency between the action output and the hardware channels.
[0064] In practice, the control device can read the connection order of each individual cell in the battery pack and generate a channel mapping relationship with two adjacent individual cells as a single cell pair.
[0065] Step S204: Construct an adjacent equalization model based on the adjacent equalization channel mapping data.
[0066] It should be noted that the adjacent equalization model refers to a model used to describe the relationship between equalization actions, equalization channel control quantities, and state of charge changes among adjacent individual cells.
[0067] In practical implementation, the control device can establish a corresponding equalization branch model for each adjacent cell pair based on the adjacent equalization channel mapping data, and use the duty cycle parameter as the control input for the equalization branch. The adjacent equalization model can calculate the corresponding equalization strength based on the duty cycle, the voltage difference between adjacent cells, the channel inductance parameter, or the equivalent transmission parameter, and update the state of charge of adjacent cells accordingly.
[0068] Step S205: Construct a digital twin equilibrium training environment based on the battery pack digital twin model, the neighbor equilibrium model, and the state of charge sequence.
[0069] In practical implementation, the control device can use the battery pack digital twin model as the controlled object, the neighbor equilibrium model as the action execution object, the charge state sequence as the intelligent agent observation input, and configure the action interface, state feedback interface and reward calculation interface for the training environment.
[0070] Within a training cycle, the digital twin equilibrium training environment outputs the current state of charge sequence to the proximal policy optimization agent, receives the actions of the neighboring equilibrium channels output by the proximal policy optimization agent, and changes the state of charge of individual cells in the digital twin model of the battery pack through the neighboring equilibrium model. Subsequently, it outputs the updated state of charge sequence and the corresponding reward value.
[0071] This embodiment establishes the basis for the dynamic response of a single battery cell by identifying the parameters of the equivalent circuit model; it forms a battery pack-level digital twin object by connecting the second-order resistor-capacitor equivalent circuit model in series; it establishes the topological correspondence between the action output and adjacent single battery cell pairs by generating adjacent equalization channel mapping data; it limits the equalization response corresponding to the equalization current control quantity by constructing an adjacent equalization model; and it forms an interactive training environment by integrating the battery pack digital twin model, the adjacent equalization model, and the state of charge sequence. This enables the training of the near-end policy optimization agent to be simultaneously constrained by the battery dynamic characteristics, the adjacent equalization topology, and the state of charge change law, thereby improving the effectiveness and engineering adaptability of equalization policy training.
[0072] Based on the above embodiments, in the third embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning of the present invention, step S30 further includes: Step S301: Arrange the state of charge sequence based on the series connection order of multiple series-connected individual cells in each battery pack to obtain the state observation vector.
[0073] It should be noted that the state observation vector refers to the vectorized state of charge data formed by arranging multiple series-connected individual cells in the battery pack according to their physical connection order, and is used as the state input of the near-end policy optimization agent.
[0074] Step S302: Arrange the equalization current control quantities of adjacent equalization channels based on the adjacent equalization channel mapping data to obtain the continuous action vector.
[0075] It should be noted that the continuous action vector refers to continuous action data composed of equalization current control quantities corresponding to multiple adjacent equalization channels. The equalization current control quantity refers to the action quantity used to characterize the target equalization current magnitude or normalized equalization current magnitude of adjacent equalization channels. The adjacent equalization channel mapping data refers to the correspondence data between adjacent equalization channels and adjacent individual cell pairs.
[0076] In practice, the vehicle battery pack balancing control system reads the mapping data of adjacent balancing channels and arranges each balancing current control quantity according to the adjacent balancing channel number or the connection order of adjacent single cell pairs to form a continuous action vector.
[0077] Step S303: Based on the difference between the maximum and minimum state of charge values in the state observation vector, obtain the state of charge difference data.
[0078] It should be noted that state of charge difference data refers to data used to characterize the dispersion of the state of charge of each series-connected cell in the same battery pack, and can be represented by the difference between the maximum and minimum state of charge values.
[0079] In a practical implementation, the control device can traverse each state of charge value in the state observation vector, obtain the maximum state of charge value and the minimum state of charge value, and calculate the difference between the two as the state of charge difference data. For example, when there is a significant difference between the highest state of charge cell and the lowest state of charge cell in a battery pack, this difference can be used as the basis for subsequent penalty calculations to obtain the state of charge difference data.
[0080] Step S304: Generate charge state difference penalty data based on the charge state difference data and the first preset weight.
[0081] It should be noted that the first preset weight refers to a pre-set weight parameter used to adjust the influence of state-of-charge (POC) difference data on the overall reward function. For example, it can be set to a larger weight value based on the balance priority. The POC difference penalty data refers to negative evaluation data generated based on the POC difference data.
[0082] In a specific implementation, the control device can weight the state of charge difference data with a first preset weight to obtain state of charge difference penalty data. The first preset weight can be preset according to the balanced control objective. For example, a higher weight can be set in the training phase where the inconsistency of the battery pack is prioritized, and a moderate weight can be set in the training phase where the smoothness of the action is taken into account, so as to obtain state of charge difference penalty data.
[0083] Step S305: Generate motion amplitude penalty data based on the sum of squares of the duty cycle parameters in the continuous motion vector and the second preset weight.
[0084] It should be noted that the second preset weight refers to the pre-set weight parameters used to adjust the strength of the motion amplitude constraint. The motion amplitude penalty data refers to the negative evaluation data used to constrain excessive equalization current control quantity. If the equalization current control quantity is expressed in a normalized form, the motion amplitude penalty data can be calculated based on the normalized equalization current control quantity; if the equalization current control quantity is expressed in the form of actual current, the motion amplitude penalty data can be calculated based on the actual target equalization current value.
[0085] Step S306: When the state of charge difference data is less than the preset convergence threshold, generate equilibrium convergence state reward data.
[0086] It should be noted that the preset convergence threshold refers to the pre-set threshold for the difference in state of charge used to determine whether the battery pack has entered the equilibrium convergence state. The equilibrium convergence state reward data refers to the positive evaluation data generated when the difference in state of charge meets the convergence condition.
[0087] In a specific implementation, the control device can compare the state of charge difference data with a preset convergence threshold; when the state of charge difference data is less than the preset convergence threshold, it generates equilibrium convergence state reward data; when the state of charge difference data is greater than or equal to the preset convergence threshold, it may not generate a positive reward or generate a preset low reward value, thereby obtaining equilibrium convergence state reward data.
[0088] Step S307: Construct a comprehensive reward function based on the charge state difference penalty data, the action amplitude penalty data, and the equilibrium convergence state reward data.
[0089] It should be noted that the comprehensive reward function is based on the following formula: in, This represents the comprehensive reward data. This data represents the difference in state of charge (SOC), used to measure the degree of inconsistency in the SOC of a battery pack. This represents the reward data for the equilibrium convergence state. Indicates the first Duty cycle parameters of adjacent equalization channels This indicates the number of adjacent equalization channels.
[0090] In practical implementation, the control device can use the charge state difference penalty data and action amplitude penalty data as negative reward items, and the equilibrium convergence state reward data as positive reward items, and combine the reward items to obtain the comprehensive reward function. For example, at a certain training moment, if the agent's output of the equilibrium current control quantity can reduce the charge state difference and the action amplitude is small, the comprehensive reward function outputs a higher reward value; if the equilibrium current control quantity has a large amplitude but the charge state difference is not reduced sufficiently, the comprehensive reward function outputs a lower reward value.
[0091] Step S308: Train the proximal policy optimization agent based on the state observation vector, the continuous action vector, the comprehensive reward function, and the digital twin balanced training environment.
[0092] In the specific implementation, the control device initializes different state-of-charge distributions in a digital twin equilibrium training environment and inputs the state observation vectors into the proximal policy optimization agent. The proximal policy optimization agent outputs continuous action vectors based on the state observation vectors. The digital twin equilibrium training environment simulates the equilibrium response of adjacent equilibrium channels based on the continuous action vectors and updates the battery pack's state-of-charge sequence. The system generates reward data for the current training moment based on a comprehensive reward function and updates the proximal policy optimization agent using the state observation vectors, continuous action vectors, and reward data corresponding to multiple training moments until a preset number of training rounds, a preset reward convergence condition, or a preset training duration is reached.
[0093] This embodiment constructs state observation vectors in a cascaded order, preserving the topological relationships of individual cells within the battery pack; it establishes a correspondence between the equalization current control quantity and adjacent equalization channels by arranging continuous action vectors through channel mapping; it quantifies the degree of inconsistency through state-of-charge difference data, providing an equalization target for reward evaluation; it constrains the equalization effect and equalization current intensity through state-of-charge difference penalties and action amplitude penalties, respectively; it guides the strategy to maintain a low-difference state through equalization convergence state rewards; it unifies the training objective through a comprehensive reward function, and utilizes a digital twin equalization training environment to train the near-end policy optimization agent, thereby improving the training sample coverage and policy training stability.
[0094] Based on the third embodiment described above, in the fourth embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning of the present invention, step S308 further includes: Step S3081: Construct the policy network and value network of the proximal policy optimization agent based on the proximal policy optimization algorithm.
[0095] It should be noted that the policy network refers to the neural network in the near-end policy optimization agent used to generate the distribution data of the equilibrium current control quantity based on the state observation vector. The value network refers to the neural network in the near-end policy optimization agent used to estimate the long-term reward value corresponding to the state observation vector. The distribution data of the equilibrium current control quantity can be the mean parameter, variance parameter, or action probability density parameter corresponding to each adjacent equilibrium channel. The state value estimation data refers to the value network's estimation result of the expected cumulative reward corresponding to the current state observation vector.
[0096] In practical implementation, the control device can construct a proximal policy optimization agent based on a proximal policy optimization algorithm, and configure a policy network and a value network for the proximal policy optimization agent. The policy network can adopt a lightweight multilayer perceptron structure, with the input layer dimension consistent with the state observation vector dimension and the output layer dimension consistent with the number of adjacent equalization channels. The output results are used to represent the distribution data of the equalization current control quantity of each adjacent equalization channel. The value network can adopt the same or similar input structure as the policy network, and output state value estimation data through a single-value output layer. The policy network and value network can use randomly initialized parameters, or they can be pre-trained based on historical equalization samples and then enter the digital twin equalization training environment for reinforcement learning training.
[0097] Step S3082: Generate comprehensive reward data based on the comprehensive reward function, and generate cumulative reward target data based on the comprehensive reward data, the preset discount factor, and the temporal relationship of each training moment within the training round.
[0098] It should be noted that the comprehensive reward data refers to the reward value output by the comprehensive reward function at a single training time. The preset discount factor is a parameter set in advance to adjust the degree of influence of future rewards on the current training objective.
[0099] It should be noted that a training round refers to a training process in which the proximal policy optimization agent continuously interacts with the digital twin equilibrium training environment from the initial state of charge distribution until the termination condition is met. The cumulative reward target data refers to the training target data formed by discounting and accumulating the comprehensive reward data according to the training time sequence. The cumulative reward target data is calculated using the following formula: in, This represents the cumulative reward target data. The network parameters represent the policy network. This represents the expectation of the sample data during the training process. This represents the maximum training time within a single training round. This indicates the current training moment within a training round. Indicates the preset discount factor. This represents the instantaneous reward data at the t-th training time. In its implementation, the control device can call the comprehensive reward function at each training moment to generate comprehensive reward data based on the differences in state of charge, the amplitude of the equilibrium current, and the equilibrium convergence state, and record the reward sequence according to the training moment order. Subsequently, the reward sequence is accumulated in reverse or rolled over based on a preset discount factor to generate cumulative reward target data. The end condition for a training round can be reaching a preset number of training steps, satisfying a preset equilibrium convergence condition, or reaching a preset simulation duration. For example, the current training round ends after running for a preset number of control cycles in the digital twin equilibrium training environment.
[0100] Step S3083: Input the state observation vector into the policy network to obtain the current equalization current control quantity distribution data.
[0101] It should be noted that the current equalization current control quantity distribution data refers to the action distribution data output by the policy network based on the state observation vector at the current training time, which is used to determine the equalization current control quantity of each adjacent equalization channel.
[0102] In practical implementation, the control device can input the state observation vector at the current training moment into the policy network. After passing through multiple layers of nonlinear mapping, the policy network outputs the distribution data of the equalization current control quantity corresponding to each adjacent equalization channel. For the normalized equalization current control quantity, the output result of the policy network can correspond to the normalized action distribution; for the actual equalization current control quantity, the output result of the policy network can correspond to the target current action distribution. During the training phase, continuous action vectors can be sampled from the current equalization current control quantity distribution data. During the verification phase, the distribution mean or the action value after limiting can be selected as the equalization current control quantity.
[0103] Step S3084: Input the state observation vector into the value network to obtain state value estimation data.
[0104] It should be noted that the state value estimation data is used to characterize the long-term return estimate corresponding to the current state observation vector under the current policy.
[0105] In its implementation, the control device can input the state observation vectors at the same training time into the value network, which then outputs a state value estimate. This state value estimate can be compared with the comprehensive reward data or the cumulative reward target data for subsequent advantage evaluation. The value network can be updated by minimizing the error between the value estimate and the cumulative reward target data.
[0106] Step S3085: Generate advantage assessment data based on the comprehensive reward data and the state value estimation data.
[0107] It should be noted that advantage assessment data refers to data used to characterize the superiority or inferiority of the current action relative to the average strategy or state value benchmark of the current state.
[0108] In practice, the control device can calculate the advantage assessment data based on the comprehensive reward data at the current training moment, the state value estimation data at the next training moment, and the state value estimation data at the current training moment. Alternatively, it can calculate the advantage assessment data based on the difference between the cumulative reward target data and the current state value estimation data.
[0109] In some embodiments, a generalized advantage estimation method can be used to smooth the reward and value estimation results at multiple training times in order to reduce the volatility of advantage assessment data.
[0110] Step S3086: Generate strategy ratio data based on the current equalization current control quantity distribution data and the historical equalization current control quantity distribution data.
[0111] It should be noted that the historical equilibrium current control quantity distribution data refers to the equilibrium current control quantity distribution data output by the policy network in the previous training round based on the same or corresponding state observation vector. The policy ratio data refers to the ratio between the action probability under the current policy and the action probability under the historical policy, used to characterize the magnitude of change between the old and new policies. The policy ratio data is calculated using the following formula: in, This represents the policy ratio data at training time t. This indicates that the current policy network is in the state observation data. Select action data The probability, This indicates that the historical policy network is based on state observation data. Select action data The probability, This represents the policy network parameters from the previous training round. This represents the state observation data at the t-th training time. This represents the equalization current control action at the t-th training moment.
[0112] In practical implementation, the control device can read the historical equalization current control quantity distribution data saved in the previous training round, and calculate the probability ratio under the same state observation vector and the same continuous action vector based on the current equalization current control quantity distribution data. For the continuous action space, the vehicle battery pack equalization control system can calculate the probability density corresponding to the equalization current control action according to the distribution parameters output by the strategy network, and divide the current probability density by the historical probability density to generate strategy ratio data.
[0113] Step S3087: Perform amplitude limiting update processing on the strategy ratio data based on the preset clipping coefficient to obtain amplitude limiting strategy update data.
[0114] It should be noted that the preset clipping factor refers to a pre-set coefficient used to limit the range of change in the strategy ratio data. The clipping strategy update data refers to the target data for strategy updates after clipping constraints. The clipping update processing follows the formula below: in, This represents the value of the objective function for pruning. This represents the expectation of the sample data at each training time step. Let represent the advantage evaluation data at training time t, used to characterize the superiority or inferiority of the balanced current control action relative to the current state value estimation data. Represents the amplitude limiting function. This indicates the preset clipping factor.
[0115] In practical implementation, the control device can prune the strategy ratio data to upper and lower limits based on preset pruning coefficients, ensuring the strategy ratio data remains within the preset pruning range. Subsequently, the vehicle battery pack balancing control system, combining the advantage evaluation data, calculates both the unpruned and pruned strategy update terms, and selects the update result that satisfies the near-end strategy optimization objective as the amplitude-limited strategy update data. The preset pruning coefficients can be set according to the action space dimension, training stability requirements, and the complexity of the digital twin balancing training environment.
[0116] Step S3088: Update the policy network and value network based on the cumulative reward target data, the limit policy update data, and the advantage evaluation data to obtain the trained proximal policy optimization agent.
[0117] In practical implementation, the trained proximal policy optimization agent refers to a reinforcement learning control unit obtained after multiple rounds of interactive training in a digital twin equalization training environment. This control unit is used to output a target action sequence based on the real-time state-of-charge sequence. The target action sequence may include the target equalization current control quantities of each adjacent equalization channel.
[0118] In its implementation, the control device updates the value network parameters using the error between the cumulative reward target data and the state value estimation data, and updates the policy network parameters using the policy target composed of the limit policy update data and the advantage evaluation data.
[0119] During training, state observation vectors, continuous action vectors, comprehensive reward data, cumulative reward target data, and state value estimation data are repeatedly collected over multiple training rounds, and the policy network and value network are updated using a batch training method. When the number of training rounds reaches the preset number of training rounds, the cumulative reward target data tends to stabilize, or the difference in charged states meets the preset training convergence condition, the trained proximal policy optimization agent is output.
[0120] Based on the above embodiments, in the fifth embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning of the present invention, step S40 further includes: Step S401: Determine the target battery pack from the multiple battery packs based on the preset equalization task information.
[0121] It should be noted that the preset balancing task information refers to the information that is pre-set or generated by the vehicle battery management system to determine the balancing object, balancing triggering conditions and balancing execution cycle. It may include battery pack number, balancing start flag, state of charge difference threshold, temperature safety range, balancing priority and control cycle.
[0122] For example, the preset balancing task information could be: when the difference between the maximum and minimum state of charge values in a battery pack is greater than the preset balancing start threshold, the battery pack is identified as the object to be balanced.
[0123] Step S402: Read the real-time state of charge sequence of the target battery pack from the state of charge sequence.
[0124] In practical implementation, the control device can read the real-time state of charge (POC) sequence of the target battery pack from the POC sequence buffer, state estimation module, or real-time data queue corresponding to multiple battery packs, based on the target battery pack number. If there is data missing in the target battery pack within the current control cycle, a usable real-time POC sequence can be generated by interpolating data from adjacent sampling times, maintaining the state from the previous control cycle, or removing abnormal data. If there are cells with abnormal temperatures or sampling in the target battery pack, the POC data corresponding to the abnormal cells can be marked for subsequent state constraint processing.
[0125] Step S403: Vectorize the real-time state of charge sequence according to the series connection order of multiple series-connected individual cells in the target battery pack to obtain the real-time state observation vector.
[0126] In practical implementation, the control device can arrange the real-time state of charge (SOC) sequence according to the connection order of multiple series-connected individual cells in the target battery pack from the negative terminal to the positive terminal, and convert the arranged data into a one-dimensional vector, column vector, or data tensor adapted to the input interface of the intelligent agent. For example, if the target battery pack includes several series-connected individual cells, the corresponding SOC values can be arranged in ascending order according to the individual cell numbers to form a real-time state observation vector; if the near-end policy optimization agent needs to normalize the input, the real-time state observation vector can also be scaled.
[0127] Step S404: Perform state constraint processing on the real-time state observation vector based on the preset upper limit value of the charged state and the preset upper limit value of the charged state to obtain the constrained state observation vector.
[0128] It should be noted that the preset minimum charge state value refers to the minimum allowable value of the charge state set in advance; the preset maximum charge state value refers to the maximum allowable value of the charge state set in advance. State constraint processing refers to the process of restricting the charge state values in the real-time state observation vector to between the preset minimum charge state value and the preset maximum charge state value. The constrained state observation vector refers to the agent input vector obtained after restricting the range of state values.
[0129] In the specific implementation, the control device reads the state-of-charge (SOC) values from the real-time state observation vector one by one, corrects SOC values that are less than the preset SOC lower limit to the preset SOC lower limit, and corrects SOC values that are greater than the preset SOC upper limit to the preset SOC upper limit. SOC values between the preset SOC lower limit and the preset SOC upper limit are kept unchanged. After completing the state constraint processing, a constraint state observation vector is generated, and this constraint state observation vector is used as the input data for the trained proximal policy optimization agent.
[0130] Step S405: Input the constrained state observation vector into the trained proximal policy optimization agent to obtain the initial action sequence.
[0131] It should be noted that the initial action sequence refers to the set of original equalization current control actions output by the trained proximal policy optimization agent based on the constraint state observation vector.
[0132] In the specific implementation, the control device inputs the constraint state observation vector into the trained proximal policy optimization agent. The trained proximal policy optimization agent can be obtained through training on the state observation vector, continuous action vector, and comprehensive reward data in the previous digital twin equilibrium training environment. The policy network can adopt a lightweight multilayer perceptron structure, with the input layer dimension consistent with the constraint state observation vector dimension and the output layer dimension consistent with the number of adjacent equilibrium channels. During the online control phase, if the policy network outputs equilibrium current control quantity distribution data, the distribution mean or sampled value can be selected as the initial action sequence; if the policy network directly outputs action values, the action values can be arranged according to the order of adjacent equilibrium channels to obtain the initial action sequence.
[0133] Step S406: Perform action constraint processing on the initial action sequence based on the preset lower limit value and the preset upper limit value of the balanced current control quantity to obtain the target action sequence.
[0134] It should be noted that the preset lower limit of the equalization current control quantity refers to the minimum equalization current control quantity allowed to be output by adjacent equalization channels, which can be a normalized lower limit or an actual current lower limit; the preset upper limit of the equalization current control quantity refers to the maximum equalization current control quantity allowed to be output by adjacent equalization channels, which can be a normalized upper limit or an actual current upper limit.
[0135] For example, in the normalized action expression method, the preset lower limit of the equalization current control quantity can be 0, and the preset upper limit of the equalization current control quantity can be 1; in the actual current expression method, the preset upper limit of the equalization current control quantity can be set according to the rated current capability of adjacent active equalization circuits. Action constraint processing refers to the process of restricting the initial action sequence to between the preset lower limit and the preset upper limit of the equalization current control quantity. The target action sequence refers to the action dataset that, after action constraint processing, can be used to generate adjacent equalization control commands.
[0136] In the specific implementation, the control device reads the initial equalization current control values from the initial action sequence one by one. Initial equalization current control values lower than the preset lower limit are corrected to the preset lower limit, and initial equalization current control values higher than the preset upper limit are corrected to the preset upper limit. Initial equalization current control values between the preset lower and upper limits are kept unchanged. After completing the action constraint processing, the target action sequence is output according to the order of adjacent equalization channels. The target action sequence can serve as the data basis for subsequently generating adjacent equalization control commands.
[0137] This embodiment filters target battery packs by pre-setting equalization task information, giving the equalization control object a clear task orientation and avoiding indiscriminate equalization actions on multiple battery packs, thus improving the targeting of equalization scheduling. By reading the real-time state of charge (SOC) sequence of the target battery pack at the current control moment, subsequent agent inputs can reflect the current equalization requirements, reducing control lag caused by actions generated based on historical states. By organizing the real-time SOC sequence in a cascaded order, the real-time state observation vector is consistent with the physical topology of the target battery pack, preserving the SOC differences between adjacent individual cells. State constraint processing limits the impact of abnormal estimates or sampling disturbances on agent inputs, ensuring that the SOC data received by the near-end policy optimization agent meets the physical value range of the battery, improving the stability of action output. The trained near-end policy optimization agent maps the current SOC distribution of the target battery pack to the equalization current control quantity of adjacent equalization channels, enabling equalization actions to be adaptively generated according to the current inconsistent state of the target battery pack. Action constraint processing ensures that the agent output meets the current execution range of the equalization hardware, reducing the impact of out-of-range equalization current on adjacent active equalization circuits and improving the executability of the target action sequence.
[0138] Based on the above embodiments, in the sixth embodiment of the vehicle battery pack equalization control method based on deep reinforcement learning of the present invention, step S40 further includes: Step S411: Decompose the target action sequence to obtain multiple channel equalization current control quantities.
[0139] It should be noted that the channel equalization current control quantity refers to the target value of equalization current or the normalized target value of equalization current allocated to a single adjacent equalization channel, which is used to characterize the equalization intensity of the corresponding adjacent equalization channels.
[0140] In the specific implementation, the control device parses the target action sequence according to the element order, extracts each element in the target action sequence as a channel equalization current control quantity, and configures a corresponding action sequence number for each channel equalization current control quantity.
[0141] Step S412: Based on the adjacent equalization channel mapping data, map the multiple channel equalization current control quantities to multiple adjacent equalization channels to obtain channel control data.
[0142] It should be noted that channel control data refers to the control data formed by binding the channel equalization current control quantity with the corresponding adjacent equalization channel, which may include the adjacent equalization channel number, the adjacent single cell pair number, the channel equalization current control quantity, and the control interface number.
[0143] In the specific implementation, the control device reads the mapping data of adjacent equalization channels and binds the equalization current control quantities of multiple channels to the corresponding adjacent equalization channels according to the channel number.
[0144] Step S413: Generate equalization drive control data based on the channel control data.
[0145] It should be noted that the equalization drive control data refers to the data used to control the turn-on and turn-off times of power switching devices in adjacent active equalization circuits, which may include control cycle, turn-on duration, turn-off duration, dead time, and channel triggering sequence.
[0146] In practice, the control device generates equalization drive control data for the corresponding channel based on the adjacent equalization channel number and channel equalization current control quantity in the channel control data, combined with the differences in state of charge, terminal voltage, and drive requirements of adjacent active equalization circuits of adjacent individual cells.
[0147] Step S414: Generate adjacent equalization control instructions based on the channel control data and the equalization drive control data.
[0148] It should be noted that the adjacent equalization control command refers to the control information sent to the embedded equalization execution unit and used to drive the adjacent active equalization circuit to perform equalization operations. It may include channel identifier, duty cycle parameter, switching timing parameter, control cycle parameter and verification parameter.
[0149] It should be understood that this step encapsulates channel control data and equalization drive control data into adjacent equalization control instructions, thereby forming a unified instruction format for channel allocation results and switch drive logic, which facilitates parsing and execution by the embedded equalization execution unit.
[0150] In practical implementation, the control device can encapsulate the channel number and duty cycle parameters in the channel control data with the on-time, off-time, and control cycle in the equalization drive control data according to the communication protocol of the embedded equalization execution unit, generating adjacent equalization control instructions containing frame headers, channel fields, timing fields, and check fields. For multiple adjacent equalization channels, a single frame of control instructions containing all channel control information can be generated, or multiple frames of control instructions can be generated according to the channel order.
[0151] Step S415: Send the adjacent equalization control command to the embedded equalization execution unit.
[0152] It should be noted that the embedded equalization execution unit can be a hardware unit configured with a microcontroller, driver circuit, communication interface and sampling interface, used to parse adjacent equalization control instructions and drive adjacent active equalization circuits to perform equalization actions.
[0153] In practical implementation, the control device can send adjacent equalization control commands through a serial communication interface, a controller area network interface, an Ethernet interface, or a board-level communication interface. After receiving the adjacent equalization control commands, the embedded equalization execution unit can perform command verification, channel parsing, and timing buffering, and output the corresponding drive signals based on the parsing results. For example, in a hardware-in-the-loop testing scenario, the host computer or battery management controller can send adjacent equalization control commands to the microcontroller development board, which then outputs channel drive signals based on the control commands.
[0154] Step S416: Receive the execution feedback data returned by the embedded equalization execution unit based on the adjacent equalization control command, and form closed-loop control based on the execution feedback data.
[0155] It should be noted that execution feedback data refers to the status data returned by the embedded equalization execution unit after executing adjacent equalization control commands. This data may include channel execution status, voltage sampling feedback, current sampling feedback, temperature sampling feedback, and anomaly flags. Closed-loop control refers to a control method that updates the battery pack status based on the execution feedback data and continues to generate subsequent control actions based on the updated status.
[0156] It should be understood that this step, by receiving execution feedback data and using it for status updates, enables the equalization control to move beyond simply issuing open-loop commands. Instead, it allows for continuous correction of subsequent control actions based on the actual execution status and sampled feedback, thereby improving the real-time performance and execution reliability of the vehicle equalization control.
[0157] In its implementation, the control device receives execution feedback data from the embedded equalization execution unit and performs time alignment, anomaly marking, and state update processing on the feedback data. If the execution feedback data indicates that adjacent equalization channels are conducting normally and the sampled data is within a safe range, the updated voltage, current, temperature, and state of charge data can be used as the agent input for the next control cycle. If the execution feedback data indicates that a certain channel has over-temperature, over-current, or communication anomalies, the duty cycle of the corresponding channel can be reduced, equalization of the corresponding channel can be paused, or adjacent equalization control commands can be regenerated. Through the above feedback processing, a closed-loop control is formed.
[0158] Furthermore, to improve the accuracy of closed-loop control state updates, step S416 may include: Step S4161: Receive the channel execution status data returned by the embedded equalization execution unit.
[0159] It should be noted that the channel execution status data refers to the operating status data of adjacent equalization channels after executing adjacent equalization control commands, which may include channel on status, channel off status, drive confirmation status, abnormal flags and execution completion flags.
[0160] In the specific implementation, the control device receives the channel execution status data returned by the embedded equalization execution unit, and judges the execution status of each adjacent equalization channel according to the channel number, execution flag and exception flag, thereby confirming whether the adjacent equalization channels corresponding to the target action sequence have been correctly triggered.
[0161] Step S4162: Receive voltage sampling feedback data, current sampling feedback data, and temperature sampling feedback data transmitted back by the embedded equalization execution unit.
[0162] In its implementation, the control device receives voltage, current, and temperature sampling feedback data from the embedded equalization execution unit and caches them according to the sampling channel number and sampling time. For example, after adjacent equalization operations are performed, the terminal voltage changes of each series-connected cell in the target battery pack, the current changes of the equalization branch, and the temperature changes of the battery module can be read, and the three types of sampling feedback data can be used as the data source for subsequent state updates.
[0163] Step S4163: Monitor the action generation process of the trained proximal policy optimization agent and the transmission process of the neighboring equalization control command to obtain total control delay data.
[0164] It should be noted that the action generation process refers to the process by which the trained proximal policy optimization agent outputs the target action sequence based on the real-time state of charge sequence. The transmission process refers to the communication process in which adjacent equalization control commands are sent from the control device to the embedded equalization execution unit. The total control delay data refers to the sum of the delays corresponding to the action generation process and the transmission process. In some embodiments, it may also include the feedback data return delay or the embedded execution confirmation delay. The preset total delay threshold can be preset according to the vehicle equalization control cycle, for example, it can be set to 5 milliseconds, 10 milliseconds, or a delay threshold that matches the battery management control cycle.
[0165] It should be understood that this step, by monitoring the delays in the action generation process and the command transmission process, can determine whether the closed-loop balance control meets the on-board real-time requirements and reduce the risk of outdated state data participating in the next round of balance decision-making.
[0166] In a specific implementation, the control device can record the start time of action generation when the real-time state-of-charge sequence is input into the trained proximal policy optimization agent, and the end time of action generation when the target action sequence is output, and calculate the action generation delay based on the difference between the two. It can also record the transmission time when sending adjacent equalization control commands, and the acknowledgment time when the embedded equalization execution unit returns a receipt acknowledgment, and calculate the command transmission delay based on the difference between the two. Subsequently, the action generation delay and the command transmission delay are summed to obtain the total control delay data; in embodiments requiring further constraints on feedback real-time performance, the feedback data return delay can also be included in the total control delay data.
[0167] Furthermore, in order to reduce latency, improve response speed, and enhance robustness and control accuracy, step S4163 may include: Step S41631: Monitor the training of the near-end policy optimization agent to generate inference delay data corresponding to the target action sequence.
[0168] It should be noted that inference latency data refers to the time consumption data generated between the training proximal policy optimization agent receiving the real-time charged state sequence and outputting the target action sequence, which may include input preprocessing time, policy network forward computation time, and action output preparation time.
[0169] It should be understood that this step, by monitoring inference latency data, can quantify the time cost of the agent's decision-making stage, providing a basis for determining whether the deep reinforcement learning equilibrium strategy meets the vehicle's real-time control cycle.
[0170] Step S41632: Monitor the instruction transmission delay data corresponding to the adjacent equalization control instruction sent to the embedded equalization execution unit.
[0171] It should be noted that instruction transmission delay data refers to the communication time consumed between the sending end of the adjacent equalization control instruction and the confirmation of receipt by the embedded equalization execution unit. It can be related to the controller area network, serial communication, Ethernet communication or board-level communication link.
[0172] It should be understood that this step, by monitoring the instruction transmission delay data, can quantify the transmission overhead of control instructions in the communication link, and avoid excessive communication delays that could lead to a mismatch between the equalization control instructions and the real-time state of charge.
[0173] Step S41633: Receive the hardware execution delay data returned by the embedded equalization execution unit.
[0174] It should be noted that hardware execution latency data refers to the time consumption data generated between the time the embedded equalization execution unit parses the adjacent equalization control instructions and the time it drives the adjacent active equalization circuit to enter the corresponding execution state. It can include instruction parsing time, drive signal generation time, power switch response time, and execution confirmation time.
[0175] It should be understood that by receiving hardware execution latency data, this step can obtain the execution response information of the balanced hardware side, so that the total control latency evaluation includes not only algorithm reasoning and communication transmission, but also the time overhead of the hardware execution stage.
[0176] Step S41634: Generate total control delay data based on the inference delay data, the instruction transmission delay data, and the hardware execution delay data.
[0177] Understandably, the intelligent agent and hardware employ a collaborative mechanism of instruction issuance and status feedback to ensure synchronized and efficient decision-making and execution. After completing inference, the intelligent agent packages the balancing target, intensity, and timing into control commands and issues them to the master controller. The hardware performs the balancing operation and transmits voltage, temperature, and execution status back in real time. The decision-making layer dynamically adjusts its actions based on the feedback, forming a closed loop of perception-decision-execution-feedback. This mechanism can reduce latency, improve response speed, and enhance robustness and control accuracy.
[0178] The core advantage of this collaborative mechanism is the reduction of control latency, calculated using the following formula: in, Indicates total control delay (unit: ms). The unit is ms, which represents the agent inference delay (in which the agent generates the target action sequence based on the real-time charge state sequence) of the proximal policy optimization agent. Indicates command transmission delay (unit: ms). This indicates the hardware execution latency (unit: ms). Actual testing shows that the total latency under this mechanism is ≤5ms, meeting the low latency requirements for automotive applications and effectively improving robustness and control accuracy.
[0179] Step S4164: When the total control delay data is not greater than the preset total delay threshold, perform time alignment processing on the channel execution status data, the voltage sampling feedback data, the current sampling feedback data and the temperature sampling feedback data to obtain aligned feedback data.
[0180] It should be understood that this step uses the total control delay to filter feedback data that meets the real-time requirements and performs time alignment on multi-source feedback data so that subsequent state updates are based on data within the same time window, thereby improving the consistency of closed-loop control data.
[0181] In practical implementation, the control device can first compare the total control delay data with a preset total delay threshold. If the total control delay data is not greater than the preset total delay threshold, the channel execution status data, voltage sampling feedback data, current sampling feedback data, and temperature sampling feedback data are aligned according to a unified timestamp, control cycle number, or feedback frame sequence number. For data with different sampling frequencies, time alignment can be achieved using nearest neighbor matching, interpolation, or control cycle merging. If the total control delay data is greater than the preset total delay threshold, the current feedback data can be discarded, the control frequency reduced, or feedback data re-requested.
[0182] Step S4165: Update the real-time state of charge sequence of the target battery pack based on the alignment feedback data to obtain the updated real-time state of charge sequence.
[0183] It should be understood that this step updates the real-time state of charge sequence of the target battery pack by utilizing alignment feedback data, so that the next round of agent input can reflect the actual state after the adjacent equalization operation, thereby improving the accuracy of closed-loop control state update.
[0184] In practical implementation, the control device can re-estimate or correct the state of charge (SOC) of each series-connected cell in the target battery pack based on the voltage, current, and temperature sampling feedback data in the alignment feedback data. For example, the SOC of a single cell can be updated by combining changes in the equalization branch current and changes in the battery pack loop current, and abnormal states can be corrected using the voltage and temperature sampling feedback data. If the channel execution status data indicates that an adjacent equalization channel has not executed normally, the participation of the corresponding feedback data in the status update can be reduced, thereby obtaining an updated real-time SOC sequence.
[0185] Step S4166: Regenerate the target action sequence based on the updated real-time state of charge sequence, and form closed-loop control based on the regenerated target action sequence.
[0186] In practical implementation, the control device can use the updated real-time state of charge sequence as input data for the next control cycle, inputting it into the trained proximal policy optimization agent to obtain a regenerated target action sequence. Subsequently, it can continue to drive the embedded equalization execution unit to perform adjacent equalization operations according to the channel mapping, timing generation, and command sending process, and continuously receive subsequent execution feedback data. Through multiple rounds of execution, the actions of each adjacent equalization channel can be continuously adjusted according to the changes in the target battery pack state, forming a closed-loop control.
[0187] Furthermore, this embodiment of the invention also proposes a computer-readable storage medium storing a vehicle battery pack balancing control program. When the vehicle battery pack balancing control program is executed by a processor, it implements the steps of the vehicle battery pack balancing control method based on deep reinforcement learning as described above.
[0188] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0189] The aforementioned computer-readable storage medium may be included in a deep reinforcement learning-based vehicle battery pack balancing control device; or it may exist independently and not be assembled into the deep reinforcement learning-based vehicle battery pack balancing control device.
[0190] Furthermore, this invention also proposes a computer program product, including a vehicle battery pack balancing control program, which, when executed by a processor, implements the steps of the deep reinforcement learning-based vehicle battery pack balancing control method described above.
[0191] The specific implementation of the computer program product of the present invention is basically the same as the embodiments of the above-mentioned vehicle battery pack equalization control method based on deep reinforcement learning, and will not be repeated here.
[0192] Reference Figure 4 , Figure 4 This is a structural block diagram of the first embodiment of the vehicle battery pack balancing control system based on deep reinforcement learning of the present invention.
[0193] like Figure 4 As shown, the vehicle battery pack equalization control system based on deep reinforcement learning proposed in this embodiment of the invention includes: Data acquisition module 10 is used to generate a state of charge sequence of each battery pack based on the single cell operation data of multiple battery packs in the vehicle battery system. Each battery pack includes multiple single cells connected in series. The single cell operation data includes voltage data, current data and temperature data. Environment construction module 20 is used to construct a digital twin equalization training environment based on the connection relationship of adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence; The agent training module 30 is used to train a proximal policy optimization agent based on the digital twin equalization training environment. The proximal policy optimization agent uses the charge state sequence as state observation data and the equalization current control quantity of adjacent equalization channels as action data. The action generation module 40 is used to input the real-time state of charge sequence of the target battery pack into the trained proximal policy optimization agent to obtain the target action sequence. The equalization control module 50 is used to generate adjacent equalization control instructions based on the target action sequence, and send the adjacent equalization control instructions to the embedded equalization execution unit so that the embedded equalization execution unit drives the adjacent active equalization circuit to perform adjacent equalization operation, and transmits execution feedback data back in real time to form closed-loop control.
[0194] This embodiment collects individual battery cell operating data and generates a state of charge (SOC) sequence, transforming the raw electrical and thermal sampling data into a state representation recognizable by the agent. A training environment is constructed using a battery pack digital twin model and a neighboring equalization model, ensuring the training process is simultaneously constrained by battery dynamics and equalization channel structure. A proximal policy optimization agent is trained using a comprehensive reward function, ensuring the equalization strategy considers SOC consistency, action amplitude constraints, and convergence state preservation. By inputting the real-time SOC sequence of the target battery pack into the trained proximal policy optimization agent, a target action sequence matching the current state is generated. An embedded equalization execution unit drives adjacent active equalization circuits and transmits execution feedback data, enabling closed-loop execution of equalization actions on the hardware side, thereby improving the adaptability, real-time performance, and engineering deployment suitability of automotive battery pack equalization control.
[0195] In one embodiment, the deep reinforcement learning-based vehicle battery pack equalization control system adopts a hierarchical collaborative architecture, organically combining digital twin training, single-agent decision-making, and rapid hardware execution to form a real-time equalization control system. This embodiment takes a 6-cell lithium battery pack for vehicles as the research object, designing a lightweight PPO equalization strategy with SOC as the observed variable and equalization control signal as the action space to achieve adaptive optimization control of the equalization process. Simultaneously, an active equalization hardware system is built based on the STM32H7 embedded platform and GaN power devices, achieving collaborative work between the algorithm and hardware while ensuring low power consumption and high response speed.
[0196] The overall system architecture is designed as follows: (1) Digital twin layer: 6-string second-order RC battery model + inductive active balancing topology; (2) Intelligent decision-making layer: single-agent PPO equalization controller; (3) Hardware execution layer: STM32H7 main controller + GaN driver circuit; The system is divided into a digital twin layer, an intelligent decision-making layer, and a hardware execution layer. It forms a closed-loop control through data interaction. The three-layer architecture performs its own functions and works together to balance the policy adaptability with the requirements of low latency and high stability in vehicles.
[0197] The intelligent agent and hardware employ a collaborative mechanism of command issuance and status feedback to ensure synchronized and efficient decision-making and execution. After completing inference, the intelligent agent packages the balancing target, intensity, and timing into control commands and issues them to the master controller. The hardware performs the balancing operation and transmits voltage, temperature, and execution status back in real time. The decision-making layer dynamically adjusts its actions based on the feedback, forming a closed loop of perception-decision-execution-feedback. This mechanism reduces latency, improves response speed, and enhances system robustness and control accuracy.
[0198] In reinforcement learning, the agent learns the optimal control strategy by continuously interacting with the environment. At time t, the agent, based on its current state... Select Action After the environment performs an action, it transitions to the next state and returns a reward. The agent's goal is to maximize the long-term cumulative reward by continuously adjusting its policy. Its optimization objective can be expressed as: in, This represents the cumulative reward target data. The network parameters represent the policy network. This represents the expectation of the sample data during the training process. This represents the maximum training time within a single training round. This indicates the current training moment within a training round. Indicates the preset discount factor. This represents the instantaneous reward data at the t-th training time.
[0199] In this embodiment, the proximal policy optimization algorithm design agent (PPO agent) adopts an Actor-Critic structure. The Actor network outputs action policies based on the current state, while the Critic network estimates the value function of the current state. The Actor is responsible for control decisions, and the Critic evaluates the merits of the current policy; together, they complete policy updates. For the battery equalization control task in this embodiment, the Actor network outputs control actions for five adjacent equalization channels based on the six individual cell SOC states, while the Critic network evaluates the value of the policy under the current equalization state.
[0200] The core idea of an intelligent agent is to introduce probability ratios during policy updates: in, This represents the policy ratio data at training time t. This indicates that the current policy network is in the state observation data. Select action data The probability, This indicates that the historical policy network is based on state observation data. Select action data The probability, This represents the policy network parameters from the previous training round. This represents the state observation data at the t-th training time. This represents the equalization current control action at the t-th training moment.
[0201] If the differences between the old and new strategies are too large, it may lead to instability in the training process. Therefore, this embodiment uses a pruning objective function to limit the policy update magnitude, and the objective function can be expressed as: in, This represents the value of the objective function for pruning. This represents the expectation of the sample data at each training time step. Let represent the advantage evaluation data at training time t, used to characterize the superiority or inferiority of the balanced current control action relative to the current state value estimation data. Represents the amplitude limiting function. This represents the preset pruning coefficient; through this objective function, when the changes between the old and new strategies are too large, the update range will be limited to a certain range, thereby improving training stability.
[0202] In this embodiment, battery pack SOC balancing control is a continuous control problem, where balancing actions are output in the form of duty cycles, and the action values change continuously. Therefore, the PPO algorithm is suitable for the balancing control task of a six-cell battery pack in this embodiment. By training the PPO agent in the Simulink simulation environment, it can gradually learn to output reasonable balancing actions based on the degree of SOC inconsistency, thereby reducing the maximum SOC difference of the battery pack and achieving SOC balancing control.
[0203] Design of state, action, and reward functions: To transform the SOC balancing problem of a six-cell battery pack into a reinforcement learning control task, it is necessary to explicitly define the agent's state space, action space, and reward function. This embodiment takes a six-cell lithium battery pack as the research object and adopts a single-agent PPO control structure. The agent uniformly observes the SOC state of the six individual cells and outputs control actions for five adjacent balancing channels.
[0204] State-space design: The state space is the basis for the agent's decision-making. In this embodiment, the goal of the equalization control is to reduce the SOC difference among the six individual cells; therefore, the SOC of the six individual cells is selected as the agent's observed state. The state vector is defined as follows: in, Let represent the state of charge (SOC) of the i-th individual cell. Since the physical range of SOC is 0 to 1, the state-space constraint is: In the MATLAB reinforcement learning environment, the state space is defined as a six-dimensional continuous vector with the following dimensions: This state design can directly reflect the degree of inconsistency within the battery pack. When there are large differences in the state of charge (SOC) of individual cells, the agent can determine the objects and intensity of balancing that need to be performed based on the state vector.
[0205] State space and motion space design: The state space is constructed around consistency, health status, and environmental conditions, including individual unit voltages, voltage variance, temperature difference, and approximate health status values, comprehensively reflecting the degree of imbalance and safety risks. The action space corresponds to the balanced execution intensity and direction, outputting continuously adjustable parameters that directly map to hardware-executable control logic, ensuring that decisions are implementable and achievable. The design is as follows: 1. State Space: Constructed around voltage consistency, battery health status, and environmental conditions, providing comprehensive observational information for the algorithm. Core state parameters include: individual cell voltages. (i=1,2,...,n, where n is the number of individual cells), voltage variance reflects the degree of imbalance), and battery pack temperature difference. Approximate health status State-space expression: in, This represents the voltage variance, which is the average voltage of individual cells. The larger the variance, the greater the degree of imbalance.
[0206] 2. Motion Space: Corresponds to balanced execution intensity and direction, outputs continuously adjustable parameters, directly maps to hardware-executable control logic, motion space expression: in, The equalization intensity coefficient of the i-th unit is 0, which means no equalization is performed and 1 means maximum intensity equalization. The continuously adjustable action output can realize fine equalization control, ensuring that the decision can be implemented and realized.
[0207] Reward function design and optimization objectives: The reward function determines the learning direction of the agent and is crucial for designing reinforcement learning control strategies. In this embodiment, the equilibrium control objective is to reduce the maximum SOC difference of the battery pack and avoid excessive control actions. Therefore, the reward function mainly consists of an SOC difference penalty term and an action penalty term.
[0208] First, the maximum SOC difference of the battery pack is defined as: In the formula, This represents the difference in state of charge. Used to measure the degree of inconsistency in the state of charge (SOC) of a battery pack. The larger the value, the more severe the inconsistency in the battery pack. The smaller the value, the better the balancing effect.
[0209] The comprehensive reward function in this embodiment is designed as follows: Among them, the first item Used to penalize the largest difference in SOC, prompting the agent to reduce battery pack SOC inconsistency; the second term Used to suppress excessive motion output and avoid motion oscillation during the equalization control process; the third item The convergent reward term is defined as follows: When the maximum SOC difference of the battery pack is less than 0.02, it indicates that the battery pack has reached a good equilibrium state. At this time, additional rewards are given to guide the agent to learn a more stable equilibrium strategy.
[0210] This reward function balances the effects of equilibrium and control stability, enabling the agent to reduce the SOC difference during training while avoiding frequent output of excessively large control actions.
[0211] To achieve SOC equalization control of a six-cell battery pack, this embodiment models the battery pack equalization process as a sequential decision problem in reinforcement learning. At each control moment, the PPO agent outputs control actions for adjacent equalization channels based on the current SOC state of the battery pack, and evaluates the current control effect through a reward function, thereby learning the optimal equalization strategy through continuous interactive training. The PPO equalization control flow designed in this embodiment mainly includes five stages: state observation, action decision-making, equalization execution, reward calculation, and policy update.
[0212] First, within each simulation step, the environment model outputs the SOC state of the six individual batteries in real time, which constitutes the agent's observation vector. This state vector can reflect the degree of inconsistency of SOC within the current battery pack and is the basis for the agent to make equilibrium decisions.
[0213] Secondly, the agent outputs an action vector based on the current state. To prevent excessive control input from causing simulation instability, this embodiment performs amplitude limiting on the action output: Then, the action signal is input to the adjacent equalization module. The equalization module generates a corresponding equalization current based on the duty cycle, causing the cells with higher SOC to adjust their energy towards those with lower SOC. The battery pack model updates the SOC of each cell based on the equalization current to obtain the state at the next moment. .
[0214] After completing a state update, the reward function calculates an immediate reward based on the current maximum SOC difference and the action amplitude. In this embodiment, the maximum SOC difference of the battery pack is defined as: in, This represents the maximum difference in SOC of the battery pack at the t-th control or training time. This represents the state of charge of the i-th individual cell at the t-th control time. This represents the maximum state of charge (SOC) of each individual cell in the battery pack at the t-th control time. represents the minimum state of charge of each individual cell in the battery pack at the t-th control time; i represents the cell number.
[0215] The reward function is designed as follows: The first item is used to penalize inconsistencies in SOC (State of Occupation), and the second item is used to constrain excessively large control actions. To converge the reward terms. When When the SOC value is less than 0.24, it indicates that the battery pack's SOC has reached a good equilibrium state, at which point an additional reward is given. Through this reward function, the agent can gradually learn control strategies to reduce the SOC difference.
[0216] Finally, the PPO agent updates the policy network using the collected state, action, and reward data. Unlike ordinary policy gradient algorithms, PPO limits the update magnitude between the old and new policies by pruning the objective function, preventing excessive policy changes from causing training divergence and thus improving training stability. After multiple rounds of training, the agent can output reasonable balanced actions based on different SOC distributions, gradually reducing the SOC difference of the battery pack.
[0217] Termination conditions and constraint design: In this Simulink reinforcement learning environment, to ensure stability during training, the SOC signal is limited to the following range: The amplitude limit range for the motion signal is set as follows: During training, each episode ends after a set number of simulation steps. If the SOC difference becomes small, the agent is guided to maintain an equilibrium state through a convergence reward in the reward function. This method avoids premature termination of the training process while ensuring that the agent learns control laws during the complete equilibrium process.
[0218] Algorithm training and simulation in a digital twin environment: First, an equivalent model of the SOC of six battery cells and a five-channel adjacent equalization model are established in Simulink. The SOC signals of the six individual cells are combined into an observation vector by the Mux module and input to the RLAgent module. The RLAgent module outputs a five-dimensional action signal, which is distributed to the five adjacent equalization modules by the Demux module to control the equalization channels 1-2, 2-3, 3-4, 4-5 and 5-6 of the battery cells, respectively.
[0219] In the environment, the reward function is implemented by the MATLAB Function module, which takes the SOC state and action signal as input and outputs the reward value at the current moment. During the simulation, the agent outputs an balancing action based on the current SOC state, the balancing module changes the SOC of each individual agent based on the action signal, and the environment then returns the new SOC state and reward value, thus forming a closed-loop training process.
[0220] The initial SOC during the training phase is set as follows: SOC0=[0.95,0.75,0.60,0.45,0.30,0.10]. This initial state has obvious SOC inconsistency, which can effectively examine the control learning ability of the PPO agent in an unbalanced state.
[0221] Based on the established digital twin simulation environment, multiple rounds of simulation tests were conducted to verify the convergence and effectiveness of the algorithm. The simulation test conditions covered typical automotive scenarios (25℃, 1C charge / discharge, initial voltage difference of 20mV). The data on the variation of individual cell voltage difference with equalization time during the simulation tests are as follows: Figure 5 As shown, Figure 5 The graph shows the voltage difference variation curve of the individual cells under the PPO equalization control strategy designed using the method of this invention. It can be seen that the algorithm can quickly reduce the voltage difference and control the voltage difference within 3mV in 120s, which meets the equalization accuracy requirements.
[0222] In one embodiment, the simulation test platform is built based on MATLAB R2024b and Simulink, and the PPO agent is trained using the Reinforcement Learning Toolbox. The battery model uses a six-string SOC equivalent battery model, and the balancing structure uses a five-channel adjacent balancing model. In the reinforcement learning environment, the state space consists of the SOC of six individual cells, and the action space consists of the duty cycles of five adjacent balancing channels. The simulation time during the training phase is set to 500s to improve training efficiency; the simulation time during the verification phase is set to 5000s or 10000s to observe the long-term convergence trend of the SOC curve; the initial state is used to simulate a situation where there is significant SOC inconsistency in the battery pack. By observing the changing trend of the SOC of the six individual cells under the PPO balancing control, it can be determined whether the proposed balancing strategy can effectively reduce the SOC difference. During hardware-in-the-loop testing, the STM32H7 is not directly connected to the real high-voltage battery pack, but rather participates in the closed-loop control of the Simulink model as an external controller, completing SOC state reception, control decision-making, and action output.
[0223] To verify the effectiveness of the PPO equalization control strategy designed based on the method described in this invention, SOC equalization simulation tests were first conducted in the MATLAB / Simulink environment. During the test, the influence of external charging and discharging was not considered; only a state of initial SOC inconsistency was set in the six-cell battery pack, and the changes in the SOC of each cell under the equalization control were observed.
[0224] The initial SOC was set to: [0.95, 0.75, 0.60, 0.45, 0.30, 0.10]; Reference Figure 6 , Figure 6This figure shows the variation curve of the maximum SOC difference of the battery pack under the PPO equalization control strategy of the deep reinforcement learning-based vehicle battery pack equalization control method of this invention. Under the action of the PPO equalization control strategy, the SOC curves of each cell generally show a convergence trend. Cells with higher initial SOC gradually decrease, cells with lower initial SOC gradually increase, and cells with intermediate SOC show relatively gentle changes. Simulation results show that the designed control strategy can output corresponding equalization actions according to the SOC difference of the battery pack, so that the SOC difference of each cell in the battery pack gradually decreases.
[0225] like Figure 7 As shown, after 5000s of equalization control, the maximum SOC difference was reduced to approximately 0.24. This demonstrates that the PPO equalization control strategy based on deep reinforcement learning in the automotive battery pack equalization control method designed in this invention can effectively improve the battery pack SOC inconsistency problem.
[0226] Reference Figure 8 , Figure 9 , Figure 10 and Figure 11 , Figure 8 , Figure 9 , Figure 10 and Figure 11 This diagram illustrates a comparison between the PPO equilibrium control strategy designed based on the method described in this invention and the traditional fixed threshold strategy under different evaluation dimensions. Figure 8 This diagram illustrates a comparison of the convergence times of the PPO equilibrium control strategy designed based on the method described in this invention. Figure 9 This diagram illustrates a comparison of the equalization accuracy of the PPO equalization control strategy designed based on the method described in this invention. Figure 10 This is a schematic diagram comparing the energy loss of the PPO equalization control strategy designed based on the method described in this invention. Figure 11 This is a comparative schematic diagram showing the battery capacity improvement achieved by the PPO equalization control strategy designed based on the method described in this invention.
[0227] The deep reinforcement learning-based vehicle battery pack balancing control system provided in this application employs the deep reinforcement learning-based vehicle battery pack balancing control method described in the above embodiments, and can solve the technical problems of deep reinforcement learning-based vehicle battery pack balancing control. Compared with the prior art, the beneficial effects of the deep reinforcement learning-based vehicle battery pack balancing control system provided in this application are the same as those of the deep reinforcement learning-based vehicle battery pack balancing control method provided in the above embodiments, and other technical features in the deep reinforcement learning-based vehicle battery pack balancing control system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0228] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.
[0229] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.
[0230] In addition, for technical details not described in detail in this embodiment, please refer to the deep reinforcement learning-based vehicle battery pack equalization control method provided in any embodiment of the present invention, which will not be repeated here.
[0231] It should be noted that, in this embodiment, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0232] It should be noted that the user information (including but not limited to user device information, user personal information, user location information, user behavior information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0233] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0234] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0235] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A method for equalization control of automotive battery packs based on deep reinforcement learning, characterized in that, The method includes: The state of charge sequence of each battery pack is generated based on the individual cell operating data of multiple battery packs in the vehicle battery system. Each battery pack includes multiple individual cells connected in series. The individual cell operating data includes voltage data, current data and temperature data. A digital twin equalization training environment is constructed based on the connection relationship between adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence. The state of charge sequence is obtained by arranging the series-connected individual cells in each battery pack according to their series order; The equalization current control quantities of adjacent equalization channels are arranged based on the mapping data of adjacent equalization channels to obtain a continuous action vector; Based on the difference between the maximum and minimum state of charge values in the state observation vector, the state of charge difference data is obtained. Charge state difference penalty data is generated based on the charge state difference data and the first preset weight; Action amplitude penalty data is generated based on the sum of squares of the duty cycle parameters in the continuous action vector and the second preset weight; When the state of charge difference data is less than a preset convergence threshold, equilibrium convergence state reward data is generated. A comprehensive reward function is constructed based on the state of charge difference penalty data, the action amplitude penalty data, and the equilibrium convergence state reward data. The comprehensive reward function is based on the following formula: in, This represents the comprehensive reward data. This data represents the difference in state of charge (SOC), used to measure the degree of inconsistency in the SOC of a battery pack. This represents the reward data for the equilibrium convergence state. Indicates the first Duty cycle parameters of adjacent equalization channels Indicates the number of adjacent equalization channels; The proximal policy optimization agent is trained based on the state observation vector, the continuous action vector, the comprehensive reward function, and the digital twin equalization training environment. The proximal policy optimization agent uses the charged state sequence as state observation data and the equalization current control quantity of adjacent equalization channels as action data. The real-time state of charge sequence of the target battery pack is input into the trained proximal policy optimization agent to obtain the target action sequence and generate adjacent equalization control instructions. The adjacent equalization control instructions are sent to the embedded equalization execution unit so that the embedded equalization execution unit drives the adjacent active equalization circuit to perform adjacent equalization operation and transmits execution feedback data back in real time to form closed-loop control.
2. The vehicle battery pack equalization control method based on deep reinforcement learning as described in claim 1, characterized in that, The construction of a digital twin equalization training environment based on the connection relationship of adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence includes: Based on the single cell operating data, the equivalent circuit model parameters of each series-connected single cell are identified, and the second-order resistance-capacitance equivalent circuit model of each series-connected single cell is configured based on the equivalent circuit model parameters. Based on the connection relationship of multiple series-connected individual cells in each battery pack, the second-order resistive-capacitive equivalent circuit model is connected to obtain the digital twin model of the battery pack. Adjacent equalization channel mapping data is generated based on the connection relationship between adjacent individual cells in each battery pack. The adjacent equalization channel mapping data includes the mapping relationship between adjacent equalization channels and adjacent individual cell pairs. An adjacent equilibrium model is constructed based on the adjacent equilibrium channel mapping data. A digital twin equilibrium training environment is constructed based on the battery pack digital twin model, the neighbor equilibrium model, and the state of charge sequence.
3. The vehicle battery pack equalization control method based on deep reinforcement learning as described in claim 2, characterized in that, The process of training a proximal policy optimization agent based on the state observation vector, the continuous action vector, the comprehensive reward function, and the digital twin balanced training environment includes: Based on the near-end policy optimization algorithm, a policy network and a value network are constructed for the near-end policy optimization agent. The policy network is used to output the distribution data of the equalization current control quantity of adjacent equalization channels, and the value network is used to output the state value estimation data. Based on the comprehensive reward function, comprehensive reward data is generated. Then, based on the comprehensive reward data, the preset discount factor, and the temporal relationship of each training moment within the training round, cumulative reward target data is generated, as shown in the following formula: in, This represents the cumulative reward target data. The network parameters represent the policy network. This represents the expectation of the sample data during the training process. This represents the maximum training time within a single training round. This indicates the current training moment within a training round. Indicates the preset discount factor. This represents the instantaneous reward data at the t-th training time. The state observation vector is input into the policy network to obtain the current equalization current control quantity distribution data; The state observation vector is input into the value network to obtain state value estimation data; Advantage assessment data is generated based on the comprehensive reward data and the state value estimation data; Based on the current and historical balanced current control distribution data, strategy ratio data is generated, referring to the following formula: in, This represents the policy ratio data at training time t. This indicates that the current policy network is in the state observation data. Select action data The probability, This indicates that the historical policy network is based on state observation data. Select action data The probability, This represents the policy network parameters from the previous training round. This represents the state observation data at the t-th training time. This represents the equalization current control action at the t-th training moment; The strategy ratio data is subjected to amplitude limiting update processing based on a preset clipping factor to obtain amplitude limiting strategy update data. The amplitude limiting update processing refers to the following formula: in, This represents the value of the objective function for pruning. This represents the expectation of the sample data at each training time step. Let represent the advantage evaluation data at training time t, used to characterize the superiority or inferiority of the balanced current control action relative to the current state value estimation data. Represents the amplitude limiting function. Indicates the preset clipping factor; Based on the cumulative reward target data, the limit policy update data, and the advantage evaluation data, the policy network and value network are updated to obtain the trained proximal policy optimization agent.
4. The vehicle battery pack equalization control method based on deep reinforcement learning as described in claim 1, characterized in that, The step of inputting the real-time state of charge sequence of the target battery pack into the trained proximal policy optimization agent to obtain the target action sequence includes: The target battery pack is determined from the plurality of battery packs based on preset equalization task information; Read the real-time state of charge sequence of the target battery pack from the state of charge sequence; The real-time state of charge sequence is vectorized according to the series connection order of multiple series-connected individual cells in the target battery pack to obtain the real-time state observation vector. The real-time state observation vector is subjected to state constraint processing based on the preset upper limit value of the charged state and the preset upper limit value of the charged state to obtain the constrained state observation vector. The constrained state observation vector is input into the trained proximal policy optimization agent to obtain the initial action sequence; The initial action sequence is subjected to action constraint processing based on the preset lower limit value and the preset upper limit value of the balanced current control quantity to obtain the target action sequence.
5. The vehicle battery pack equalization control method based on deep reinforcement learning as described in claim 2, characterized in that, The process of generating adjacent equalization control commands and sending these commands to the embedded equalization execution unit, causing the embedded equalization execution unit to drive the adjacent active equalization circuit to perform adjacent equalization operations and to transmit execution feedback data in real time to form closed-loop control, includes: The target action sequence is broken down to obtain multiple channel equalization current control quantities; Based on the adjacent equalization channel mapping data, the equalization current control quantities of the multiple channels are mapped to the multiple adjacent equalization channels to obtain channel control data; Equalization drive control data is generated based on the channel control data; Based on the channel control data and the equalization drive control data, generate adjacent equalization control commands; The adjacent equalization control command is sent to the embedded equalization execution unit; The system receives execution feedback data from the embedded equalization execution unit based on the adjacent equalization control command, and forms closed-loop control based on the execution feedback data.
6. The vehicle battery pack equalization control method based on deep reinforcement learning as described in claim 5, characterized in that, The step of receiving execution feedback data from the embedded equalization execution unit based on the adjacent equalization control command, and forming closed-loop control based on the execution feedback data, includes: Receive channel execution status data returned by the embedded equalization execution unit; Receive voltage sampling feedback data, current sampling feedback data, and temperature sampling feedback data transmitted back by the embedded equalization execution unit; The action generation process of the trained proximal policy optimization agent and the transmission process of adjacent equalization control commands are monitored to obtain total control delay data. When the total control delay data is not greater than the preset total delay threshold, the channel execution status data, the voltage sampling feedback data, the current sampling feedback data and the temperature sampling feedback data are time aligned to obtain aligned feedback data. The real-time state of charge sequence of the target battery pack is updated based on the alignment feedback data to obtain the updated real-time state of charge sequence. The target action sequence is regenerated based on the updated real-time state of charge sequence, and closed-loop control is formed based on the regenerated target action sequence.
7. The vehicle battery pack equalization control method based on deep reinforcement learning as described in claim 6, characterized in that, The process of action generation and transmission of neighboring equalization control commands for the trained proximal policy optimization agent is monitored to obtain total control delay data, including: The monitoring and training of the proximal policy optimization agent generates inference delay data corresponding to the target action sequence; Monitor the instruction transmission delay data sent from the adjacent equalization control command to the embedded equalization execution unit; Receive hardware execution latency data returned by the embedded equalization execution unit; The total control latency data is generated based on the inference latency data, the instruction transmission latency data, and the hardware execution latency data, referring to the following formula: in, Indicates total control delay. This represents the agent inference delay caused by the near-end policy optimization agent generating the target action sequence based on the real-time charge state sequence. Indicates instruction transmission delay. This indicates hardware execution latency.
8. A vehicle battery pack balancing control system based on deep reinforcement learning, applying the vehicle battery pack balancing control method based on deep reinforcement learning as described in any one of claims 1 to 7, characterized in that, The system includes: The data acquisition module is used to generate a state of charge sequence for each battery pack based on the operating data of individual cells in multiple battery packs in the vehicle battery system. Each battery pack includes multiple individual cells connected in series. The operating data of the individual cells includes voltage data, current data, and temperature data. The environment construction module is used to construct a digital twin equalization training environment based on the connection relationship between adjacent individual cells in each battery pack, the operating data of the individual cells, and the state of charge sequence. The agent training module is used to train a proximal policy optimization agent based on the digital twin equalization training environment. The proximal policy optimization agent uses the charge state sequence as state observation data and the equalization current control quantity of adjacent equalization channels as action data. The equalization control module is used to input the real-time state of charge sequence of the target battery pack into the trained proximal policy optimization agent to obtain the target action sequence, generate adjacent equalization control instructions, and send the adjacent equalization control instructions to the embedded equalization execution unit so that the embedded equalization execution unit drives the adjacent active equalization circuit to perform adjacent equalization operation and transmits execution feedback data back in real time to form closed-loop control.
9. A vehicle battery pack balancing control device based on deep reinforcement learning, characterized in that, The deep reinforcement learning-based vehicle battery pack balancing control device includes: a memory, a processor, and a vehicle battery pack balancing control program stored in the memory. The processor is used to run the vehicle battery pack balancing control program, which is configured to implement the deep reinforcement learning-based vehicle battery pack balancing control method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Battery pack equalization method based on reinforcement learning
CN116674431A
SOC active equalization control method for nesting reconfigurable battery
CN121710470A
Circuit and Method for Cell Balancing
US20130200850A1