Lithium iron phosphate battery fast charging strategy optimization method and system based on reinforcement learning
By extracting battery aging status characteristics and comprehensive reward functions, the fast charging strategy of lithium-ion batteries is optimized, which solves the problem of balancing charging speed and safety in aging batteries and achieves a balance between safe fast charging and battery life.
Patent Information
- Application Number
- CN202511249688.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing lithium-ion battery fast-charging strategies cannot be dynamically adjusted according to the battery aging status, resulting in difficulty in balancing charging speed and safety, especially the risk of lithium plating in aging batteries.
By extracting voltage, current, and temperature time series data, combining them with capacity decay rate, designing a negative electrode potential estimation mechanism and a comprehensive reward function, and constructing a reinforcement learning model to optimize the charging strategy, adaptive adjustment and safety control of the battery aging state are achieved.
The safety and life protection of aged lithium iron phosphate batteries during fast charging are achieved, the risk of lithium plating is avoided, and the charging efficiency and battery life are improved.
Smart Images

Figure CN120749264A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of battery management, and specifically relates to a method and system for optimizing a fast-charging strategy for a lithium iron phosphate battery based on reinforcement learning. Background Art
[0002] With the widespread use of lithium-ion batteries, especially lithium iron phosphate batteries, fast charging technology has become a key link in improving user experience and device operating efficiency. However, as batteries undergo irreversible aging during long-term use, their internal electrochemical properties gradually degrade, resulting in severe challenges for traditional fast-charging strategies based on fixed current-voltage curves. On the one hand, fixed charging curves are difficult to adapt to the dynamic response differences of batteries with different degrees of aging. They often limit the charging speed of new batteries and make it difficult to fully utilize their fast-charging capabilities. On the other hand, excessive current in aged batteries can induce lithium plating at the negative electrode, seriously threatening battery safety and lifespan. On the other hand, existing methods often rely on empirical rules or simplified electrochemical models for charging control, lacking the ability to accurately perceive the real-time status of the battery and adaptively adjust the adjustment.
[0003] In recent years, deep learning technology has been introduced to the field of battery charging optimization. In particular, intelligent decision-making methods, such as reinforcement learning, have shown great potential. These methods typically construct a Markov decision process model, taking observables such as battery voltage, current, and temperature as state inputs, and training a policy network to output the optimal charging current. However, existing deep learning methods generally suffer from key flaws: First, the characterization of battery aging often remains at the level of static indicators, such as using only cycle count or capacity decay rate as additional features. This fails to fully consider aging status and real-time charging parameters, resulting in an inability to dynamically adjust the policy based on the actual battery health. Second, the negative electrode potential, a core variable for determining lithium plating risk, is difficult to measure directly. Existing estimation methods often rely on simplified models or fixed parameter mappings, resulting in inaccurate negative electrode point measurements, which in turn affects the assessment of lithium plating risk. Third, the reward function design often focuses on a single objective, such as maximizing charging speed or minimizing temperature rise, without modeling charging efficiency and lithium plating risk, leading to an uncontrolled safety margin.
[0004] The above-mentioned drawbacks jointly restrict the reliability and universality of smart charging strategies under real complex working conditions, making it difficult to achieve the coordinated optimization of safety and efficiency during the charging process of aging batteries. Summary of the Invention
[0005] In response to the problems existing in the background technology, the purpose of the present invention is to provide a method and system for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning, aiming to overcome the problems in the existing technology that it is difficult to balance battery life and charging efficiency, and the risk of lithium plating during the charging process.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A method for optimizing a fast-charging strategy for a lithium iron phosphate battery based on reinforcement learning comprises the following steps:
[0008] S1: Using a voltage sensor, a current sensor, and a temperature sensor, real-time voltage data, real-time current data, and real-time temperature data of the lithium iron phosphate battery during charging are collected, and preprocessed to obtain preprocessed voltage time series data, current time series data, and temperature time series data;
[0009] At the same time, the capacity attenuation rate is calculated by the capacity calibration method;
[0010] S2: Extract three-factor features based on the pre-processed voltage time series data, current time series data, and temperature time series data, and extract capacity decay features based on the capacity decay rate;
[0011] Calculate the potential sensitivity characteristics based on the pre-processed voltage time series data, pre-processed temperature time series data and capacity attenuation rate;
[0012] Calculate the gated fusion feature based on the three-factor feature, capacity decay feature, and potential sensitivity feature;
[0013] According to the gated fusion features, a potential boundary constraint activation function is designed to calculate the estimated value of the negative electrode potential;
[0014] The potential boundary constraint activation function performs a nonlinear transformation on the gated fusion feature based on a sigmoid function and a fully connected layer, and scales the gated fusion feature according to the difference between the capacity decay rate and the capacity decay health state threshold, so that the output range of the negative electrode potential estimate is dynamically tightened;
[0015] S3: Based on the estimated negative electrode potential, calculate the negative electrode potential estimation sequence and extract the negative electrode potential characteristics. Then, combined with the capacity decay rate, calculate the aging modulation characteristics, risk sensitivity characteristics, and aging state characteristics in sequence.
[0016] S4: Construct the state space vector and action space of the reinforcement learning model, design the aging policy network and comprehensive reward function;
[0017] S5: Use the aging strategy network and the comprehensive reward function to train the reinforcement learning model until it converges. Based on the converged reinforcement learning model, record the real-time charging current actions in the action space and arrange them in chronological order to obtain the charging current curve. Then, based on the charging current curve, execute the current regulation command through the charging device to complete the charging control of the battery.
[0018] Furthermore, the specific process of step S1 is as follows:
[0019] S11: Real-time voltage data, real-time current data, and real-time temperature data of the lithium iron phosphate battery during the charging process are collected through a voltage sensor, a current sensor, and a temperature sensor, and digitally converted through an analog-to-digital converter to obtain voltage time series data, current time series data, and temperature time series data;
[0020] S12: removing high-frequency noise from the voltage time series data, the current time series data, and the temperature time series data by moving average filtering and normalizing them using a minimum-maximum normalization method to obtain preprocessed voltage time series data, preprocessed current time series data, and preprocessed temperature time series data;
[0021] S13: Calculate the capacity attenuation rate of the aged lithium iron phosphate battery using a capacity calibration method based on the historical charge and discharge cycle data stored in the battery management system.
[0022] Furthermore, the lithium iron phosphate battery is a lithium iron phosphate battery whose aging degree reaches a threshold; the threshold is that the number of charge and discharge cycles is greater than 500 times, or its available capacity is less than 80% of the initial nominal capacity, or its DC internal resistance is greater than 120% of the initial internal resistance.
[0023] Furthermore, the specific process of step S2 is:
[0024] S21: Based on the preprocessed voltage time series data, the preprocessed current time series data, and the preprocessed temperature time series data, a bidirectional long short-term memory network is used to extract three-factor features. The calculation method is:
[0025] ,
[0026] in, is the three-factor feature at time t, t is the time index, is a bidirectional long short-term memory network, For splicing operation, is the pre-processed voltage time series data at time t, is the pre-processed current time series data at time t, is the pre-processed temperature time series data at time t;
[0027] S22: Based on the capacity decay rate, the capacity decay characteristics are extracted through the embedding layer. The calculation method is:
[0028] ,
[0029] in, is the capacity fading characteristic, is the embedding layer, is the capacity attenuation rate;
[0030] S23: Calculate the potential sensitivity characteristics based on the pre-processed voltage time series data, the pre-processed temperature time series data, and the capacity attenuation rate. The calculation method is:
[0031] ,
[0032] ,
[0033] ,
[0034] in, is the voltage change rate series, is the central difference operator, is the preprocessed voltage time series data, is the temperature change rate series, is the preprocessed temperature time series data, It is a potential sensitive feature. is the Sigmoid function;
[0035] S24: Based on the three-factor characteristics, capacity decay characteristics, and potential-sensitive characteristics, a gated fusion feature is generated through a potential-sensitive gating mechanism. The calculation method is:
[0036] ,
[0037] ,
[0038] in, is the intermediate modulation characteristic at time t, ⊙ is the Hadamard product, is the dynamic characteristic gain coefficient, is the aging modulation amplitude coefficient, is the hyperbolic tangent function, is the fully connected layer, is the gated fusion feature at time t;
[0039] S25: Based on the gated fusion features, a potential boundary constraint activation function is designed to calculate the estimated value of the negative electrode potential. The calculation method is:
[0040] ,
[0041] ,
[0042] in, is the estimated value of the negative electrode potential at time t, is the safe lower limit of negative electrode potential, is the upper limit of the negative electrode potential, is the potential boundary constraint activation function, is the capacity decay health status threshold, is the aging sensitive gain coefficient, To take the maximum value.
[0043] It should be further explained that, in order to address the technical difficulties of inaccurate estimation of the negative electrode potential of aged lithium iron phosphate batteries during fast charging and the increased sensitivity of aged batteries to charging conditions, the present invention extracts three-factor features and capacity decay features, combines them with potential-sensitive features for gated fusion, and can effectively correlate the battery's dynamic response with its aging state. This allows the gated fusion features to incorporate both the electrochemical behavior of the real-time charging process and battery aging information.
[0044] Specifically, the three-factor feature extracts the dynamic parameter change characteristics of the battery during the charging process from the voltage, current and temperature time series data, such as the linear voltage rise in the constant current stage and the current decay in the constant voltage stage; the capacity decay feature extracts the battery aging information from the capacity decay rate, reflecting aging phenomena such as the increase in battery internal resistance; the potential-sensitive feature models the interaction between the voltage change rate, temperature change rate and capacity decay rate; in the gated fusion process, the potential-sensitive feature, as part of the dynamic adjustment weight, intelligently balances the contribution ratio of the three-factor feature and the capacity decay feature. When the system detects a high-risk operating condition, it automatically increases the weight of the capacity decay feature, making the fused feature more inclined to reflect the safety constraints of the aging battery;
[0045] The potential boundary constraint activation function designed on this basis can dynamically adjust the negative electrode potential estimation range according to the difference between the capacity decay rate and the capacity decay health state threshold, so that batteries with higher degree of aging can automatically tighten the estimation range to avoid the risk of lithium plating;
[0046] This activation function adds a dynamic scaling term related to the aging state to the standard Sigmoid function. , when the battery is in good health (e.g. ), the dynamic scaling term remains at the baseline value, the negative electrode potential estimation range is normal, and the model estimates the negative electrode potential in a conventional way; however, when the battery aging degree increases (e.g. ), the dynamic scaling term automatically increases, tightening the upper limit of the negative electrode potential estimate to the safety margin. This mechanism achieves the technical effect of "the older the battery, the more conservative the negative electrode potential estimate";
[0047] This design embeds the physical safety prior between battery aging and lithium plating risk, and can automatically adjust the estimation strategy according to the actual degree of battery aging, effectively preventing the negative electrode potential from approaching the critical point of lithium plating. Compared with the existing technology, the present invention realizes accurate estimation of the negative electrode potential through multi-source information fusion, and considers the impact of the battery aging state on charging safety, so that the estimated value of the negative electrode potential can be adaptively adjusted according to the actual degree of battery aging. Since the core design of the present invention revolves around the lithium plating sensitivity unique to aging batteries, the potential boundary constraint mechanism ensures that a more conservative estimation strategy is automatically adopted when the battery aging intensifies, which not only avoids the problem of limited charging speed of new batteries, but also effectively prevents the risk of lithium plating caused by overcharging of aging batteries. Therefore, it is more adaptable to the scenario of aging lithium iron phosphate batteries, and achieves a balance between safe and fast charging and battery life protection.
[0048] Furthermore, in the S3 step, the negative electrode potential estimation values at each moment are concatenated in chronological order to obtain a negative electrode potential estimation sequence; the negative electrode potential features are extracted based on the negative electrode potential estimation sequence; the aging modulation features are generated based on the capacity decay rate; the negative electrode potential features and the aging modulation features are fused to generate risk-sensitive features; and the aging state features are generated through the attention mechanism based on the risk-sensitive features and the negative electrode potential estimation sequence.
[0049] Furthermore, the specific process of step S3 is as follows:
[0050] S31: Concatenate the estimated negative electrode potential values at each moment in chronological order to obtain a negative electrode potential estimation sequence ;
[0051] S32: Based on the negative electrode potential estimation sequence, the negative electrode potential features are extracted through a one-dimensional convolution layer and a one-dimensional maximum pooling layer. The calculation method is:
[0052] ,
[0053] in, is the negative electrode potential characteristic, is a one-dimensional convolutional layer, is a one-dimensional maximum pooling layer, is the ReLU function;
[0054] S32: Generate an aging modulation feature based on the capacity decay rate through an aging-sensitive gating mechanism. The calculation method is:
[0055] ,
[0056] in, is the aging modulation characteristic, is a multi-layer perceptron, To obtain the maximum value;
[0057] S33: Generate risk-sensitive features based on the negative electrode potential characteristics and aging modulation characteristics. The calculation method is:
[0058] ,
[0059] ,
[0060] in, is the feature fusion gating weight, is a risk-sensitive characteristic;
[0061] S34: Based on the risk-sensitive features and the negative electrode potential estimation sequence, the aging state features are generated through the attention mechanism. The calculation method is:
[0062] ,
[0063] in, Characteristic of aging state. For the attention mechanism.
[0064] It should be further explained that during the charging process, aged lithium iron phosphate batteries face the technical difficulty of using static aging states to drive dynamic charging decisions. This is because aging indicators such as capacity decay rate change slowly and cannot directly reflect real-time risk changes during the charging process. Traditional methods usually use aging states as fixed parameter inputs or simply combine them with real-time states. This results in charging strategies that are insufficiently adaptable to aged batteries, making it impossible to maximize charging speed while ensuring safety, and it is difficult to effectively prevent lithium plating risks.
[0065] The present invention splices the estimated values of the negative electrode potential at each moment into a sequence in chronological order, which represents the dynamic evolution law of the negative electrode potential during the charging process and provides complete time series information for subsequent feature extraction. On this basis, the negative electrode potential features are extracted through a one-dimensional convolution layer and a one-dimensional maximum pooling layer, which can effectively identify the key feature points in the charging curve, such as the width and slope of the voltage plateau period. These features are closely related to the risk of lithium plating. At the same time, the aging modulation feature is generated according to the capacity decay rate to quantify the impact of the battery aging degree on charging safety. When the battery is charged, the aging modulation feature is automatically enhanced, reflecting the physical property that aging batteries are more sensitive to the same charging conditions. Furthermore, the negative electrode potential feature and the aging modulation feature are fused through feature fusion gate weights to generate risk-sensitive features, realizing dynamic assessment of charging process risks. Under safe conditions, the system tends to focus on the negative electrode potential feature to maintain charging efficiency. Under high-risk conditions, the weight of the aging modulation feature is automatically increased to strengthen safety constraints. Finally, the risk-sensitive feature is combined with the negative electrode potential estimation sequence through the attention mechanism to generate an aging state feature. This feature not only contains information about the battery's aging degree, but also integrates the real-time risk level of the current charging state, solving the fundamental problem that "static aging state cannot drive real-time decision-making."
[0066] Compared with the existing technology, the present invention no longer simply uses the capacity decay rate as a fixed parameter. Instead, it converts static aging information into a dynamic risk assessment indicator through a multi-level feature extraction and fusion mechanism, so that the charging strategy can be adaptively adjusted according to the actual aging degree of the battery and the current charging status.
[0067] Furthermore, in step S4, a state space vector including the current temperature, the current charge amount, the current voltage change rate, and the estimated value of the negative electrode potential, and an action space including the minimum charging current to the maximum charging current range are constructed; according to the aging state characteristics, the hidden state of the policy network is modulated through the aging gate weight matrix, the aging gate bias vector, the offset weight matrix, and the offset bias vector, so that the aging state characteristics serve as the conditional input of the policy network;
[0068] A comprehensive reward function is designed that includes a charging efficiency reward function and a lithium plating risk penalty function. The charging efficiency reward function is calculated based on the current charge capacity and its first-order derivative, and the lithium plating risk penalty function is calculated based on the estimated negative electrode potential, the current charge capacity, and the current temperature. The comprehensive reward function is further calculated by combining the current voltage change rate and the current temperature.
[0069] Furthermore, the specific process of step S4 is as follows:
[0070] S41: Construct the state space vector and action space of the reinforcement learning model. The calculation method is:
[0071] ,
[0072] ,
[0073] in, is the state space vector, is the current temperature, is the current charge level, is the current voltage change rate, A is the action space, a is the current charging current action, is the minimum charging current, is the maximum charging current;
[0074] S42: Based on the aging state characteristics, an aging strategy network of the reinforcement learning model is designed to calculate the current charging current. The calculation method is:
[0075] ,
[0076] ,
[0077] ,
[0078] in, is the hidden state of the kth layer of the aging strategy network, k is the layer index, , K is the total number of layers in the aging strategy network, is the optimized weight matrix of the kth layer of the aging strategy network, For aging strategy network The hidden state of the layer, is the optimized bias vector of the kth layer of the aging strategy network, is the hidden state of the final layer of the aging strategy network, For aging strategy network The hidden state of the layer, is the aging gating weight matrix, is the aging gate bias vector, is the bias weight matrix, is the offset bias vector, is the final optimized weight matrix of the Kth layer of the aging strategy network, is the final optimized bias vector of the Kth layer of the aging strategy network;
[0079] S43: Design a comprehensive reward function for the reinforcement learning model based on the state space vector.
[0080] It should be further explained that existing reinforcement learning techniques for optimizing battery fast-charging strategies generally treat the aging state as an external static parameter or simply add it as an additional dimension to the state space, failing to effectively integrate it into the decision-making mechanism of the policy network. As a result, the policy function cannot dynamically adjust its behavior according to the degree of battery aging, making it difficult to achieve adaptive control of batteries at different aging stages.
[0081] This paper uses the aging state characteristics as the conditional input of the policy network and deeply embeds them into the generation process of the network's hidden state, thus changing the limitation of the traditional reinforcement learning policy function's "passive response" to the aging state.
[0082] Specifically, in the aging policy network, a dynamic gating mechanism is constructed through the weight matrix and bias vector of each layer, so that the aging state characteristics can regulate the activation strength of the hidden state of each layer of the policy network. When the aging state characteristics reflect that the battery is in a mild aging state, the aging gating weight remains at a high level, and the policy network is dominated by the pursuit of charging efficiency and outputs a higher charging current action. When the aging state characteristics reflect that the battery has entered a severe aging state, the aging gating weight is automatically reduced. At the same time, a negative offset is introduced by offsetting the weight matrix and offsetting the bias vector, so that the overall output of the policy network shifts in a conservative direction. This allows the policy behavior to continuously evolve with the degree of aging without changing the network structure. This design makes the policy function no longer a single fixed mapping relationship, but has the ability to "adaptively deform" according to the battery aging state.
[0083] Compared to existing technologies, this method no longer treats aging status as an isolated input feature and simply splices it together. Instead, it transforms it into a "control signal" that influences the policy network's result generation process. This enables the policy network to actively adjust its decision preferences based on the aging status, avoiding the problem of policy rigidity caused by the slow change of aging status. Furthermore, this method eliminates the need to train multiple independent policy models for different aging stages, significantly reducing training complexity and deployment costs.
[0084] This invention breaks through the inherent limitations of traditional reinforcement learning in dealing with the slow-changing characteristics of battery aging. It elevates the aging state from an "observed state variable" to a "conditional variable for regulating strategy behavior", allowing the charging strategy to respond to real-time electrochemical states and adapt to long-term aging evolution, achieving coordinated optimization of charging speed and battery life. Therefore, it is more adaptable to charging scenarios of aging lithium iron phosphate batteries.
[0085] Furthermore, the specific process of step S43 is as follows:
[0086] S431: Calculate the charging efficiency reward function based on the current charging amount. The calculation method is:
[0087] ,
[0088] in, is the charging efficiency reward function, is the first-order derivative of the current charge, is a natural constant;
[0089] S432: Calculate the lithium plating risk penalty function based on the estimated negative electrode potential, the current charge capacity, and the current temperature. The calculation method is:
[0090] ,
[0091] ,
[0092] in, is the lithium plating risk penalty function, is the natural logarithm function, is the risk penalty item;
[0093] S433: Calculate the comprehensive reward function based on the charging efficiency reward function, the lithium plating risk penalty function, the current voltage change rate, and the current temperature. The calculation method is:
[0094] ,
[0095] in, is the comprehensive reward function, To take the absolute value.
[0096] It should be further explained that to address the difficulty in balancing charging speed and battery life during fast charging of aged lithium iron phosphate batteries, traditional methods typically use fixed charging curves or simple linearly weighted reward functions, which cannot accurately reflect the risk of lithium plating under multi-physics coupling. This can easily lead to overcharging or insufficient charging speed, especially in high SOC areas and low temperature conditions. The comprehensive reward function designed in this invention achieves precise control of the charging process through the synergistic effects of the charging efficiency reward function and the lithium plating risk penalty function.
[0097] Specifically, the charging efficiency reward function not only considers the current charging speed but also matches the electrochemical characteristics of lithium iron phosphate batteries through nonlinear design. In the low SOC region, fast charging is fully encouraged to improve efficiency, while in the high SOC region, the natural attenuation of the reward is achieved through the combination of cubic and exponential terms.
[0098] The lithium plating risk penalty function uses the estimated negative electrode potential as a direct assessment indicator of lithium plating risk. When the negative electrode potential approaches the critical point of lithium plating at 0.05 volts, the risk penalty term increases exponentially. Taking into account the impact of the increased lithium concentration gradient in the high SOC region (over 75%) and the inhibitory effect of temperature on lithium plating, the Sigmoid function triggers an enhanced penalty mechanism, doubling the risk penalty and effectively preventing battery damage from lithium plating when the negative electrode potential drops to 0.03 volts, the critical point of accelerated lithium plating.
[0099] The comprehensive reward function achieves a dynamic balance among multiple objectives by subtracting a lithium plating risk penalty from the charging efficiency reward and incorporating safety constraints for voltage change rate and temperature. The voltage change rate constraint initiates a nonlinear penalty when the rate of change exceeds 0.005 volts per second to ensure system stability, and the temperature constraint initiates a linear penalty when the temperature exceeds 45 degrees Celsius to prevent thermal runaway risks.
[0100] This design enables the reward function to automatically adjust the reward orientation according to the real-time status, prioritizing charging speed under safe conditions and automatically shifting to safety priority under high-risk conditions. For example, when the battery is in a low-temperature environment and the negative electrode potential is close to the critical value, the comprehensive reward decreases rapidly, prompting the strategy network to automatically reduce the charging current, while maintaining a higher charging speed when the temperature is suitable and the negative electrode potential is safe. Compared with the existing technology, the present invention no longer relies on indirect indicators such as terminal voltage to assess risk, but directly constructs a risk penalty function based on the estimated value of the negative electrode potential, which is more in line with the physical mechanism of lithium plating.
[0101] The present invention also discloses a lithium iron phosphate battery fast charging strategy optimization system based on reinforcement learning, comprising:
[0102] Data acquisition module: collects the real-time voltage, current, and temperature data of the aging lithium iron phosphate battery during the charging process through voltage, current, and temperature sensors, and pre-processes them to obtain the pre-processed voltage, current, and temperature time series data; calculates the capacity decay rate;
[0103] Negative electrode potential estimation module: Based on the preprocessed voltage, current, and temperature time series data and capacity decay rate, the three-factor features and capacity decay features are extracted respectively. The potential-sensitive features and gated fusion features are then calculated in sequence. Based on the gated fusion features, a potential boundary-constrained activation function is designed to calculate the estimated negative electrode potential.
[0104] Aging state feature extraction module: Based on the estimated negative electrode potential value, the negative electrode potential estimation sequence is calculated to extract the negative electrode potential features. Then, combined with the capacity decay rate, the aging modulation features, risk sensitivity features, and aging state features are calculated in sequence.
[0105] Reinforcement Learning Module: Constructs the state space vector and action space of the reinforcement learning model, designs the aging policy network and comprehensive reward function; then uses the aging policy network and comprehensive reward function to train the reinforcement learning model until convergence;
[0106] Charging control module: Based on the converged reinforcement learning model, the real-time charging current actions in the action space are recorded and arranged in chronological order to obtain the charging current curve. Based on the charging current curve, the current regulation instructions are executed through the charging equipment to complete the charging control of the battery.
[0107] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0108] (1) The present invention aims to solve the problem that the traditional fast charging strategy of aging lithium iron phosphate batteries relies on a fixed current curve, which makes it difficult to balance the charging speed and life and easily causes lithium plating. The present invention proposes a fast charging strategy optimization method based on reinforcement learning. By extracting three-factor features and capacity decay features, combining potential sensitive features to generate gated fusion features, and innovatively designing a potential boundary constraint activation function to calculate the negative electrode potential estimate, the aging battery automatically tightens the estimation range. Furthermore, the present invention extracts the negative electrode potential features through the negative electrode potential estimation sequence, and generates the aging state features in combination with the capacity decay rate. Under the reinforcement learning framework, the aging state features are used as the conditional input of the strategy network, and a comprehensive reward function including a charging efficiency reward function and a lithium plating risk penalty function is designed to achieve multi-objective optimization. The charging current curve finally generated can be dynamically adjusted according to the actual aging state of the battery and the real-time charging parameters, which not only avoids the charging speed limitation of the new battery, but also effectively prevents the risk of lithium plating of the aging battery, significantly improving the charging safety and battery life.
[0109] (2) In order to solve the problems of inaccurate estimation of negative electrode potential in fast charging of aged lithium iron phosphate batteries and sensitivity of aged batteries to charging conditions, the present invention innovatively proposes a gated fusion method of three-factor characteristics and capacity decay characteristics, which effectively associates the dynamic response of the battery with the aging state; by designing potential-sensitive characteristics to intelligently balance the contribution ratio of the three-factor characteristics and capacity decay characteristics, the safety constraints are automatically enhanced under high-risk conditions and the charging efficiency is maintained under safe conditions; the present invention also designs a potential boundary constraint activation function, which dynamically adjusts the negative electrode potential estimation range according to the difference between the capacity decay rate and the capacity decay health state threshold, achieving the technical effect of "the older the battery, the more conservative the negative electrode potential estimation": when the battery is in good health, it is estimated in a conventional way; when the battery aging worsens, the estimation range is automatically tightened.
[0110] (3) Aiming at the technical problem that the static aging state in the charging of aged lithium iron phosphate batteries is difficult to drive dynamic decision-making, the present invention splices the negative electrode potential estimation values into a sequence in chronological order to fully characterize the dynamic evolution law of the charging process; extracts the negative electrode potential features through a one-dimensional convolution layer and a one-dimensional maximum pooling layer, and accurately identifies feature points closely related to the risk of lithium plating, such as the voltage plateau period; at the same time, generates aging modulation features based on the capacity decay rate to reflect the sensitivity of the aged battery to the charging conditions; intelligently fuses the negative electrode potential features with the aging modulation features through feature fusion gate weights to generate risk-sensitive features: giving priority to charging efficiency under safe working conditions, and automatically strengthening safety constraints under high-risk working conditions; finally, generates aging state features that fuse the aging degree and the real-time risk level through the attention mechanism, solving the problem that "static aging state cannot drive real-time decision-making", so that the charging strategy can be dynamically adjusted according to the actual aging state of the battery and the current charging environment, achieving the optimal balance between safety and efficiency.
[0111] (4) In response to the problem that existing reinforcement learning technologies simply treat aging status as a static parameter or an additional dimension of the state space, the present invention innovatively uses aging status characteristics as conditional inputs of the policy network and deeply embeds them into the hidden state generation process; when the aging status characteristics reflect that the battery is in a mild aging state, a high gating weight is maintained, and the policy network outputs a higher charging current to pursue efficiency; when it reflects that the battery has entered a serious aging state, the gating weight is automatically reduced and a negative offset is introduced to adjust the output in a conservative direction, so that the policy function has "adaptive deformation" capabilities, achieving the coordinated optimization of charging speed and battery life, and significantly enhancing the adaptability to the charging scenario of aging lithium iron phosphate batteries.
[0112] (5) In order to solve the problem of difficulty in balancing charging speed and life in fast charging of aged lithium iron phosphate batteries, an innovative comprehensive reward function was designed. Among them, the charging efficiency reward function matches the electrochemical characteristics of the battery through nonlinear design, encourages fast charging in the low SOC area, and realizes natural decay in the high SOC area. The lithium plating risk penalty function is constructed based on the estimated value of the negative electrode potential. When the negative electrode potential approaches the lithium plating critical point of 0.05 volts, the penalty increases exponentially, especially at the lithium plating acceleration critical point of 0.03 volts, which triggers the enhanced penalty mechanism. At the same time, the influence of high SOC area and temperature is considered. The comprehensive reward function achieves a dynamic balance of multiple objectives by subtracting the lithium plating risk penalty from the charging efficiency reward and adding safety constraints of voltage change rate and temperature. The reward function can automatically adjust the direction according to the real-time status, giving priority to speed under safe working conditions and turning to safety under high-risk working conditions. Compared with traditional methods, it is more in line with the physical mechanism of lithium plating and significantly improves the ability to balance charging safety and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0113] Figure 1 Schematic diagram of the algorithm flow for the aging state characteristics of the present invention;
[0114] Figure 2 This is a schematic diagram of the system interface provided by the present invention. DETAILED DESCRIPTION
[0115] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the implementation methods and drawings.
[0116] Example 1
[0117] A method for optimizing a fast-charging strategy for a lithium iron phosphate battery based on reinforcement learning comprises the following steps:
[0118] S1: Use voltage, current, and temperature sensors to collect and preprocess the real-time voltage, current, and temperature data of the aged lithium iron phosphate battery during the charging process to obtain preprocessed voltage, current, and temperature time series data; calculate the capacity attenuation rate; including:
[0119] S11: collecting real-time voltage data, real-time current data, and real-time temperature data of the aged lithium iron phosphate battery during the charging process through a voltage sensor, a current sensor, and a temperature sensor, and performing digital conversion through an analog-to-digital converter to obtain voltage time series data, current time series data, and temperature time series data; the aged lithium iron phosphate battery is a lithium iron phosphate battery with 700 charge and discharge cycles;
[0120] S12: removing high-frequency noise from the voltage time series data, the current time series data, and the temperature time series data by moving average filtering and normalizing them using a minimum-maximum normalization method to obtain preprocessed voltage time series data, preprocessed current time series data, and preprocessed temperature time series data;
[0121] S13: Calculate the capacity attenuation rate of the aged lithium iron phosphate battery using a capacity calibration method based on the historical charge and discharge cycle data stored in the battery management system.
[0122] S2: Based on the pre-processed voltage, current, and temperature time series data and capacity decay rate, the three-factor features and capacity decay features are extracted respectively. Then, the potential-sensitive features and gated fusion features are calculated in sequence. Based on the gated fusion features, a potential boundary-constrained activation function is designed to calculate the estimated negative electrode potential, specifically:
[0123] S21: Based on the preprocessed voltage time series data, the preprocessed current time series data, and the preprocessed temperature time series data, a bidirectional long short-term memory network is used to extract three-factor features. The calculation method is:
[0124] ,
[0125] in, is the three-factor feature at time t, t is the time index, is a bidirectional long short-term memory network, For splicing operation, is the pre-processed voltage time series data at time t, is the pre-processed current time series data at time t, is the pre-processed temperature time series data at time t;
[0126] S22: Based on the capacity decay rate, the capacity decay characteristics are extracted through the embedding layer. The calculation method is:
[0127] ,
[0128] in, is the capacity fading characteristic, is the embedding layer, is the capacity attenuation rate;
[0129] S23: Calculate the potential sensitivity characteristics based on the pre-processed voltage time series data, the pre-processed temperature time series data, and the capacity attenuation rate. The calculation method is:
[0130] ,
[0131] ,
[0132] ,
[0133] in, is the voltage change rate series, is the central difference operator, is the preprocessed voltage time series data, is the temperature change rate series, is the preprocessed temperature time series data, It is a potential sensitive feature. is the Sigmoid function;
[0134] S24: Based on the three-factor characteristics, capacity decay characteristics, and potential-sensitive characteristics, a gated fusion feature is generated through a potential-sensitive gating mechanism. The calculation method is:
[0135] ,
[0136] ,
[0137] in, is the intermediate modulation characteristic at time t, ⊙ is the Hadamard product, is the dynamic characteristic gain coefficient, is the aging modulation amplitude coefficient, is the hyperbolic tangent function, is the fully connected layer, is the gated fusion feature at time t;
[0138] S25: Based on the gated fusion features, a potential boundary constraint activation function is designed to calculate the estimated value of the negative electrode potential. The calculation method is:
[0139] ,
[0140] ,
[0141] in, is the estimated value of the negative electrode potential at time t, is the safe lower limit of negative electrode potential, is the upper limit of the negative electrode potential, is the potential boundary constraint activation function, is the capacity decay health status threshold, is the aging sensitive gain coefficient;
[0142] S3: Based on the estimated value of the negative electrode potential, calculate the negative electrode potential estimation sequence and extract the negative electrode potential characteristics; then combine the capacity decay rate to calculate the aging modulation characteristics, risk sensitivity characteristics and aging state characteristics in sequence. The algorithm flow diagram of the aging state characteristics is as follows: Figure 1 As shown, the specific process is:
[0143] S31: Concatenate the estimated negative electrode potential values at each moment in chronological order to obtain a negative electrode potential estimation sequence ;
[0144] S32: Based on the negative electrode potential estimation sequence, the negative electrode potential features are extracted through a one-dimensional convolution layer and a one-dimensional maximum pooling layer. The calculation method is:
[0145] ,
[0146] in, is the negative electrode potential characteristic, is a one-dimensional convolutional layer, is a one-dimensional maximum pooling layer, is the ReLU function;
[0147] S32: Generate an aging modulation feature based on the capacity decay rate through an aging-sensitive gating mechanism. The calculation method is:
[0148] ,
[0149] in, is the aging modulation characteristic, is a multi-layer perceptron, To obtain the maximum value;
[0150] S33: Generate risk-sensitive features based on the negative electrode potential characteristics and aging modulation characteristics. The calculation method is:
[0151] ,
[0152] ,
[0153] in, is the feature fusion gating weight, is a risk-sensitive characteristic;
[0154] S34: Based on the risk-sensitive features and the negative electrode potential estimation sequence, the aging state features are generated through the attention mechanism. The calculation method is:
[0155] ,
[0156] in, Characteristic of aging state. For the attention mechanism.
[0157] S4: Construct the state space vector and action space of the reinforcement learning model, design the aging policy network and comprehensive reward function, including:
[0158] S41: Construct the state space vector and action space of the reinforcement learning model. The calculation method is:
[0159] ,
[0160] ,
[0161] in, is the state space vector, is the current temperature, is the current charge level, is the current voltage change rate, A is the action space, a is the current charging current action, is the minimum charging current, is the maximum charging current;
[0162] S42: Based on the aging state characteristics, an aging strategy network of the reinforcement learning model is designed to calculate the current charging current. The calculation method is:
[0163] ,
[0164] ,
[0165] ,
[0166] in, is the hidden state of the kth layer of the aging strategy network, k is the layer index, , K is the total number of layers in the aging strategy network, is the optimized weight matrix of the kth layer of the aging strategy network, For aging strategy network The hidden state of the layer, is the optimized bias vector of the kth layer of the aging strategy network, is the hidden state of the final layer of the aging strategy network, For aging strategy network The hidden state of the layer, is the aging gating weight matrix, is the aging gate bias vector, is the bias weight matrix, is the offset bias vector, is the final optimized weight matrix of the Kth layer of the aging strategy network, is the final optimized bias vector of the Kth layer of the aging strategy network;
[0167] S43: Design a comprehensive reward function for the reinforcement learning model based on the state space vector, including:
[0168] S431: Calculate the charging efficiency reward function based on the current charging amount. The calculation method is:
[0169] ,
[0170] in, is the charging efficiency reward function, is the first-order derivative of the current charge, is a natural constant;
[0171] S432: Calculate the lithium plating risk penalty function based on the estimated negative electrode potential, the current charge capacity, and the current temperature. The calculation method is:
[0172] ,
[0173]
[0174] in, is the lithium plating risk penalty function, is the natural logarithm function, is the risk penalty item;
[0175] S433: Calculate the comprehensive reward function based on the charging efficiency reward function, the lithium plating risk penalty function, the current voltage change rate, and the current temperature. The calculation method is:
[0176] ,
[0177] in, is the comprehensive reward function, To take the absolute value.
[0178] S5: Use the aging strategy network and the comprehensive reward function to train the reinforcement learning model until it converges. Based on the converged reinforcement learning model, record the real-time charging current actions in the action space and arrange them in chronological order to obtain the charging current curve. Then, based on the charging current curve, execute the current regulation command through the charging device to complete the charging control of the battery.
[0179] Example 2
[0180] A reinforcement learning-based lithium iron phosphate battery fast charging strategy optimization system, including:
[0181] Data acquisition module: collects the real-time voltage, current, and temperature data of the aging lithium iron phosphate battery during the charging process through voltage, current, and temperature sensors, and pre-processes them to obtain the pre-processed voltage, current, and temperature time series data; calculates the capacity decay rate;
[0182] Negative electrode potential estimation module: Based on the preprocessed voltage, current, and temperature time series data and capacity decay rate, the three-factor features and capacity decay features are extracted respectively. The potential-sensitive features and gated fusion features are then calculated in sequence. Based on the gated fusion features, a potential boundary-constrained activation function is designed to calculate the estimated negative electrode potential.
[0183] For the aged lithium iron phosphate battery of Example 1, if the estimated value of the negative electrode potential is between 0.05V and 0.1V, it is in the critical safety zone, which is acceptable, but requires monitoring. Long-term operation in this range may cause slight lithium deposition; if the estimated value of the negative electrode potential is between 0.02V and 0.05V, it is in the high-risk zone, lithium deposition begins to occur significantly, and long-term operation should be avoided; if the estimated value of the negative electrode potential is <0.02V, it is in the extremely high-risk zone, lithium deposition is intense, lithium dendrites are easily formed, and internal short circuit and thermal runaway may occur;
[0184] Aging state feature extraction module: Based on the estimated negative electrode potential value, the negative electrode potential estimation sequence is calculated to extract the negative electrode potential features. Then, combined with the capacity decay rate, the aging modulation features, risk sensitivity features, and aging state features are calculated in sequence.
[0185] Reinforcement Learning Module: Constructs the state space vector and action space of the reinforcement learning model, designs the aging policy network and comprehensive reward function; then uses the aging policy network and comprehensive reward function to train the reinforcement learning model until convergence;
[0186] Charging control module: Based on the converged reinforcement learning model, the real-time charging current actions in the action space are recorded and arranged in chronological order to obtain the charging current curve. Based on the charging current curve, the current regulation instructions are executed through the charging equipment to complete the charging control of the battery.
[0187] Figure 2 This is a schematic diagram of the system interface provided by the present invention. The human-computer interaction interface of this system only displays the "Data Acquisition Module" and "Charging Control Module" in real time. The "Negative Electrode Potential Estimation Module," "Aging State Feature Extraction Module," and "Reinforcement Learning Module" are intermediate computational steps in deep / reinforcement learning. Their computational processes are not readable to operators and are therefore hidden in the background without any visual display.
[0188] The above description is only a specific embodiment of the present invention. Any feature disclosed in this specification, unless otherwise stated, can be replaced by other equivalent or alternative features with similar purposes; all disclosed features, or all steps in the methods or processes, except for mutually exclusive features and / or steps, can be combined in any way.
Claims
1. A method for optimizing fast charging strategy of lithium iron phosphate batteries based on reinforcement learning, characterized in that: The following steps are involved: S1: Using a voltage sensor, a current sensor, and a temperature sensor, real-time voltage data, real-time current data, and real-time temperature data of the lithium iron phosphate battery during charging are collected, and preprocessed to obtain preprocessed voltage time series data, current time series data, and temperature time series data; At the same time, the capacity attenuation rate is calculated by the capacity calibration method; S2: Extract three-factor features based on the pre-processed voltage time series data, current time series data, and temperature time series data, and extract capacity decay features based on the capacity decay rate; Calculate the potential sensitivity characteristics based on the pre-processed voltage time series data, pre-processed temperature time series data and capacity attenuation rate; Calculate the gated fusion feature based on the three-factor feature, capacity decay feature, and potential sensitivity feature; According to the gated fusion features, a potential boundary constraint activation function is designed to calculate the estimated value of the negative electrode potential; The potential boundary constraint activation function performs a nonlinear transformation on the gated fusion feature based on a sigmoid function and a fully connected layer, and scales the gated fusion feature according to the difference between the capacity decay rate and the capacity decay health state threshold, so that the output range of the negative electrode potential estimate is dynamically tightened; S3: Based on the estimated negative electrode potential, calculate the negative electrode potential estimation sequence and extract the negative electrode potential characteristics. Then, combined with the capacity decay rate, calculate the aging modulation characteristics, risk sensitivity characteristics, and aging state characteristics in sequence. S4: Construct the state space vector and action space of the reinforcement learning model, design the aging policy network and comprehensive reward function; S5: Use the aging strategy network and the comprehensive reward function to train the reinforcement learning model until it converges. Based on the converged reinforcement learning model, record the real-time charging current actions in the action space and arrange them in chronological order to obtain the charging current curve. Then, based on the charging current curve, execute the current regulation command through the charging device to complete the charging control of the battery.
2. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 1, characterized in that: The specific process of step S1 is: S11: Real-time voltage data, real-time current data, and real-time temperature data of the lithium iron phosphate battery during the charging process are collected through a voltage sensor, a current sensor, and a temperature sensor, and digitally converted through an analog-to-digital converter to obtain voltage time series data, current time series data, and temperature time series data; S12: removing high-frequency noise from the voltage time series data, the current time series data, and the temperature time series data by moving average filtering and normalizing them using a minimum-maximum normalization method to obtain preprocessed voltage time series data, preprocessed current time series data, and preprocessed temperature time series data; S13: Calculate the capacity attenuation rate of the aged lithium iron phosphate battery using a capacity calibration method based on the historical charge and discharge cycle data stored in the battery management system.
3. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 2, characterized in that: The lithium iron phosphate battery is a lithium iron phosphate battery whose aging degree reaches a threshold; the threshold is that the number of charge and discharge cycles is greater than 500 times, or its available capacity is less than 80% of the initial nominal capacity, or its DC internal resistance is greater than 120% of the initial internal resistance.
4. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 1, characterized in that: The specific process of step S2 is: S21: Based on the preprocessed voltage time series data, the preprocessed current time series data, and the preprocessed temperature time series data, a bidirectional long short-term memory network is used to extract three-factor features. The calculation method is: , in, is the three-factor feature at time t, t is the time index, is a bidirectional long short-term memory network, For splicing operations, is the pre-processed voltage time series data at time t, is the pre-processed current time series data at time t, is the pre-processed temperature time series data at time t; S22: Based on the capacity decay rate, the capacity decay characteristics are extracted through the embedding layer. The calculation method is: , in, is the capacity fading characteristic, is the embedding layer, is the capacity attenuation rate; S23: Calculate the potential sensitivity characteristics based on the pre-processed voltage time series data, the pre-processed temperature time series data, and the capacity attenuation rate. The calculation method is: , , , in, is the voltage change rate series, is the central difference operator, is the preprocessed voltage time series data, is the temperature change rate series, is the preprocessed temperature time series data, It is a potential sensitive feature. is the Sigmoid function; S24: Based on the three-factor characteristics, capacity decay characteristics, and potential-sensitive characteristics, a gated fusion feature is generated through a potential-sensitive gating mechanism. The calculation method is: , , in, is the intermediate modulation characteristic at time t, ⊙ is the Hadamard product, is the dynamic characteristic gain coefficient, is the aging modulation amplitude coefficient, is the hyperbolic tangent function, is the fully connected layer, is the gated fusion feature at time t; S25: Based on the gated fusion features, a potential boundary constraint activation function is designed to calculate the estimated value of the negative electrode potential. The calculation method is: , , in, is the estimated value of the negative electrode potential at time t, is the safe lower limit of negative electrode potential, is the upper limit of the negative electrode potential, is the potential boundary constraint activation function, is the capacity decay health status threshold, is the aging sensitive gain coefficient.
5. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 1, characterized in that: In the S3 step, the negative electrode potential estimation values at each moment are sequentially spliced in chronological order to obtain a negative electrode potential estimation sequence; the negative electrode potential features are extracted based on the negative electrode potential estimation sequence; the aging modulation features are generated based on the capacity decay rate; the negative electrode potential features and the aging modulation features are fused to generate risk-sensitive features; and the aging state features are generated based on the risk-sensitive features and the negative electrode potential estimation sequence through an attention mechanism.
6. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 4, characterized in that: The specific process of step S3 is: S31: Concatenate the estimated negative electrode potential values at each moment in chronological order to obtain a negative electrode potential estimation sequence ; S32: Based on the negative electrode potential estimation sequence, the negative electrode potential features are extracted through a one-dimensional convolution layer and a one-dimensional maximum pooling layer. The calculation method is: , in, is the negative electrode potential characteristic, is a one-dimensional convolutional layer, is a one-dimensional maximum pooling layer, is the ReLU function; S32: Generate an aging modulation feature based on the capacity decay rate through an aging-sensitive gating mechanism. The calculation method is: , in, is the aging modulation characteristic, is a multi-layer perceptron, To obtain the maximum value; S33: Generate risk-sensitive features based on the negative electrode potential characteristics and aging modulation characteristics. The calculation method is: , , in, is the feature fusion gating weight, is a risk-sensitive characteristic; S34: Based on the risk-sensitive features and the negative electrode potential estimation sequence, the aging state features are generated through the attention mechanism. The calculation method is: , in, Characteristic of aging state. For the attention mechanism.
7. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 1, wherein: In step S4, a state space vector including the current temperature, the current charge amount, the current voltage change rate, and the estimated value of the negative electrode potential, and an action space including a range from the minimum charging current to the maximum charging current are constructed; based on the aging state characteristics, the hidden state of the policy network is modulated through the aging gate weight matrix, the aging gate bias vector, the offset weight matrix, and the offset bias vector, so that the aging state characteristics serve as conditional inputs of the policy network; A comprehensive reward function is designed that includes a charging efficiency reward function and a lithium plating risk penalty function. The charging efficiency reward function is calculated based on the current charge capacity and its first-order derivative, and the lithium plating risk penalty function is calculated based on the estimated negative electrode potential, the current charge capacity, and the current temperature. The comprehensive reward function is further calculated by combining the current voltage change rate and the current temperature.
8. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 6, characterized in that: The specific process of step S4 is: S41: Construct the state space vector and action space of the reinforcement learning model. The calculation method is: , , in, is the state space vector, is the current temperature, is the current charge level, is the current voltage change rate, A is the action space, a is the current charging current action, is the minimum charging current, is the maximum charging current; S42: Based on the aging state characteristics, an aging strategy network of the reinforcement learning model is designed to calculate the current charging current. The calculation method is: , , , in, is the hidden state of the kth layer of the aging strategy network, k is the layer index, , K is the total number of layers in the aging strategy network, is the optimized weight matrix of the kth layer of the aging strategy network, For aging strategy network The hidden state of the layer, is the optimized bias vector of the kth layer of the aging strategy network, is the hidden state of the final layer of the aging strategy network, For aging strategy network The hidden state of the layer, is the aging gating weight matrix, is the aging gate bias vector, is the bias weight matrix, is the offset bias vector, is the final optimized weight matrix of the Kth layer of the aging strategy network, is the final optimized bias vector of the Kth layer of the aging strategy network; S43: Design a comprehensive reward function for the reinforcement learning model based on the state space vector.
9. The method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to claim 8, characterized in that: The specific process of step S43 is: S431: Calculate the charging efficiency reward function based on the current charging amount. The calculation method is: , in, is the charging efficiency reward function, is the first-order derivative of the current charge, is a natural constant; S432: Calculate the lithium plating risk penalty function based on the estimated negative electrode potential, the current charge capacity, and the current temperature. The calculation method is: , , in, is the lithium plating risk penalty function, is the natural logarithm function, is the risk penalty item; S433: Calculate the comprehensive reward function based on the charging efficiency reward function, the lithium plating risk penalty function, the current voltage change rate, and the current temperature. The calculation method is: , in, is the comprehensive reward function, To take the absolute value.
10. The system used in the method for optimizing the fast charging strategy of lithium iron phosphate batteries based on reinforcement learning according to any one of claims 1 to 9, characterized in that: The system comprises: Data acquisition module: collects the real-time voltage, current, and temperature data of the aging lithium iron phosphate battery during the charging process through voltage, current, and temperature sensors, and pre-processes them to obtain the pre-processed voltage, current, and temperature time series data; calculates the capacity decay rate; Negative electrode potential estimation module: Based on the preprocessed voltage, current, and temperature time series data and capacity decay rate, the three-factor features and capacity decay features are extracted respectively. The potential-sensitive features and gated fusion features are then calculated in sequence. Based on the gated fusion features, a potential boundary-constrained activation function is designed to calculate the estimated negative electrode potential. Aging state feature extraction module: Based on the estimated negative electrode potential value, the negative electrode potential estimation sequence is calculated to extract the negative electrode potential features. Then, combined with the capacity decay rate, the aging modulation features, risk sensitivity features, and aging state features are calculated in sequence. Reinforcement Learning Module: Constructs the state space vector and action space of the reinforcement learning model, designs the aging policy network and comprehensive reward function; then uses the aging policy network and comprehensive reward function to train the reinforcement learning model until convergence; Charging control module: Based on the converged reinforcement learning model, the real-time charging current actions in the action space are recorded and arranged in chronological order to obtain the charging current curve. Based on the charging current curve, the current regulation instructions are executed through the charging equipment to complete the charging control of the battery.
Citation Information
Patent Citations
Battery health state comprehensive evaluation method, system and equipment and storage medium
CN119511105A
Safe optimization and rapid charging control method for lithium battery
CN119561182A
Lithium battery performance evaluation system and method based on multi-task deep learning
CN120178047A
Power battery charging and discharging management optimization method and system based on deep learning
CN120257180A
Lithium ion battery pack rapid equalizing charging method considering aging
CN120320449A
Cited By
Sodium ion battery charging optimization method and system
CN121076288A
Electric vehicle battery life prediction method based on deep learning
CN121809623A
Singlechip intelligent power management system based on deep reinforcement learning
CN122239917A