Cylindrical lithium battery charging and discharging strategy optimization method and system based on reinforcement learning
By employing a reinforcement learning-based collaborative optimization mechanism for charging and discharging agents, the performance limitations of existing lithium battery charging and discharging strategies under complex environments and aging conditions are addressed. This enables adaptive charging and discharging strategy optimization, thereby improving battery energy efficiency and lifespan.
Patent Information
- Application Number
- CN202511128369.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-21
AI Technical Summary
Existing lithium battery charge and discharge control strategies are difficult to adapt to complex and changing usage environments and battery aging conditions. They ignore the mutual influence between charging and discharging, resulting in insufficient performance optimization, especially with significant performance degradation during the battery aging stage.
A cooperative optimization mechanism for charging and discharging agents based on reinforcement learning is adopted. By sharing reward functions and state sharing mechanisms, and combining real-time monitoring data and historical data, prediction results are generated to construct an agent cooperative optimization mechanism to optimize the charging and discharging strategy.
The overall optimization of lithium battery charging and discharging strategies has been achieved, adaptively balancing speed and lifespan requirements, improving the adaptability of the control strategy to the battery degradation process, and significantly improving the battery's energy efficiency and lifespan.
Smart Images

Figure CN120996273A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to reinforcement learning technology, and in particular to a cylindrical lithium battery charging and discharging strategy optimization method and system based on reinforcement learning. BACKGROUND
[0002] With the popularity of electric vehicles and portable electronic devices, the performance optimization of lithium battery management system has become a key technical challenge. Current lithium battery charging and discharging control strategies are mainly based on fixed rules or simple models, such as constant current constant voltage charging and threshold triggered discharging control, which are difficult to adapt to complex and variable use environments and battery aging states. In addition, existing control methods usually optimize charging and discharging as independent processes respectively, ignoring the mutual influence between them, resulting in insufficient overall performance optimization. In dynamic working environment, these methods are difficult to balance charging speed, discharging efficiency and battery life at the same time, especially in the battery aging stage, the performance decreases significantly.
[0003] In the prior art, some researches have begun to try to apply machine learning to battery management, but most of the methods only focus on single target optimization, such as maximizing charging speed or prolonging cycle life, lacking multi-objective collaborative optimization mechanism. At the same time, these methods are usually based on fixed models trained offline, which are difficult to cope with the characteristics of dynamic changes of battery parameters over time. In particular, for cylindrical lithium batteries, their unique thermal management and internal polarization characteristics make the charging and discharging process more complex, and factors such as uneven temperature distribution and internal resistance changes need to be considered to comprehensively affect the performance and life of the battery. In addition, existing research lacks a smooth switching mechanism for charging and discharging strategies under different working conditions, which is easy to lead to unstable control and increased battery stress in actual application. SUMMARY
[0004] In view of the deficiencies of the prior art, the present application provides a cylindrical lithium battery charging and discharging strategy optimization method and system based on reinforcement learning, which can solve the problems in the prior art.
[0005] In a first aspect, the application provides a cylindrical lithium battery charging and discharging strategy optimization method based on reinforcement learning, comprising: evaluating the remaining life and performance degradation trend of the battery according to real-time monitoring data and historical monitoring data of the cylindrical lithium battery to generate a prediction result; determining the current charging and discharging state according to the real-time monitoring data; establishing a reinforcement learning framework, wherein the reinforcement learning framework comprises a charging agent and a discharging agent, the state space of the charging agent and the discharging agent contains the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints; constructing an agent collaborative optimization mechanism, wherein the agent collaborative optimization mechanism comprises a shared reward function and a state sharing mechanism, the shared reward function combines the battery life extension indicator and the charging and discharging efficiency indicator as the common optimization goal of the charging agent and the discharging agent, and the state sharing mechanism makes the decision result of the charging agent as the state input of the discharging agent and makes the decision result of the discharging agent as the state input of the charging agent; training the reinforcement learning framework through the agent collaborative optimization mechanism to obtain an optimal charging strategy and an optimal discharging strategy.
[0006] In an optional implementation, the step of evaluating the remaining life and performance degradation trend of the battery according to real-time monitoring data and historical monitoring data of the cylindrical lithium battery to generate a prediction result comprises:
[0007] The real-time monitoring data and the historical monitoring data both comprise voltage data, current data, temperature data and state of charge data of the cylindrical lithium battery; based on the real-time monitoring data and the historical monitoring data, an instantaneous feature vector, a short-term feature vector and a long-term feature vector are respectively constructed, the long-term feature vector is obtained by weighted cumulative calculation of the relative capacity and the relative internal resistance of the lithium battery through a time decay factor; the contribution scores of the instantaneous feature vector, the short-term feature vector and the long-term feature vector to the prediction result are calculated, the corresponding feature weights are determined according to the contribution scores, the feature weights are multiplied by and superimposed with the corresponding feature vectors respectively to obtain a multi-scale feature vector; the multi-scale feature vector is input into a long short-term memory network, the hidden state and the input feature of the long short-term memory network are weighted through a time attention mechanism and a feature attention mechanism to obtain a remaining life prediction value and a performance degradation rate; the performance degradation trend of the lithium battery is predicted according to the performance degradation rate, and the remaining life prediction value and the performance degradation trend are adaptively corrected according to the historical prediction error to obtain the prediction result.
[0008] In an optional implementation, the step of establishing a reinforcement learning framework, wherein the reinforcement learning framework comprises a charging agent and a discharging agent, the state space of the charging agent and the discharging agent contains the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints comprises:
[0009] sliding window on the prediction results and the current state of charge to obtain a time sequence of features, applying a self-attention mechanism to the time sequence of features to calculate the correlation weights between different time points and different state variables, and generating a state representation after dimension reduction by weighted fusion; calculating the internal polarization voltage and temperature distribution of the battery based on an electrochemical dimension reduction model, combining the performance degradation trend to establish an adaptive safety boundary, and setting the charging and discharging power range within the adaptive safety boundary as the action space of the charging agent and the discharging agent, respectively; constructing a hierarchical decision network in the charging agent, the hierarchical decision network including a policy network and an evaluation network, the policy network adopting a double-delay deterministic policy architecture to output a charging power distribution, and the evaluation network evaluating the influence of the charging power on the battery life based on the battery operating characteristics and correcting the policy output; constructing a temperature-sensitive power modulator in the discharging agent, calculating the optimal power trajectory based on the battery equivalent circuit characteristics, and shaping the discharging current by using the pulse width modulation technology to achieve temperature equalization control.
[0010] In an alternative embodiment, the step of constructing a hierarchical decision network in the charging agent, the hierarchical decision network including a policy network and an evaluation network, the policy network adopting a double-delay deterministic policy architecture to output a charging power distribution, and the evaluation network evaluating the influence of the charging power on the battery life based on the battery operating characteristics and correcting the policy output includes:
[0011] The policy network includes an online policy network and a target policy network, the online policy network receiving the state representation, calculating the mean parameter and the standard deviation parameter of the charging power distribution through the hidden layer state generation, and the network parameters of the target policy network being updated by weighting the network parameters of the online policy network through a soft update coefficient; in the evaluation network, calculating a capacity decay rate index, a temperature stress index and a polarization stress index based on the battery operating characteristics, the capacity decay rate index being calculated by the capacity loss under different operating conditions, the temperature stress index being calculated by the deviation of the current temperature from the optimal operating temperature, and the polarization stress index being calculated by the cumulative effect of the polarization voltage; generating a policy correction signal based on the capacity decay rate index, the temperature stress index and the polarization stress index, the policy correction signal being obtained by weighting and combining each index through an adaptive weight coefficient; performing attention processing on the state representation to obtain an attention weight, combining the attention weight with the mean parameter and the standard deviation parameter of the charging power distribution to generate an initial charging power, and correcting the initial charging power according to the policy correction signal to obtain the charging power.
[0012] In an alternative embodiment, the step of constructing an agent cooperative optimization mechanism includes:
[0013] The battery life extension indicator is calculated based on the capacity loss amount, the initial capacity, the working temperature and the optimal temperature, and the charge-discharge efficiency indicator is calculated based on the charging efficiency, the discharging efficiency and the state of charge; the state sharing mechanism includes local state information and shared state information, the state of charge, the charging power, the temperature, the voltage and the internal resistance are input as the shared state information of the discharging agent to the charging agent, and the discharging power, the discharging efficiency and the health state are input as the shared state information of the charging agent to the discharging agent; an adaptive cooperation mechanism is constructed, the adaptive cooperation mechanism constructs a joint value function based on the local state information and the shared state information, the joint value function includes a charging value evaluation, a discharging value evaluation and a shared state value evaluation, and an instant reward is calculated based on the joint value function; the adaptive cooperation mechanism calculates a dynamic weight coefficient through a state evaluation function and a strategy evaluation function, a shared reward function is obtained by weighting and combining the battery life extension indicator and the charge-discharge efficiency indicator through the dynamic weight coefficient, a cooperative learning rate is dynamically adjusted based on a reward change trend and a strategy convergence degree, and the decision-making strategies of the charging agent and the discharging agent are jointly optimized and updated according to the cooperative learning rate.
[0014] In an optional implementation, the step of the adaptive cooperation mechanism calculating a dynamic weight coefficient through a state evaluation function and a strategy evaluation function and dynamically adjusting a cooperative learning rate based on a reward change trend and a strategy convergence degree includes:
[0015] The state evaluation function calculates a state evaluation value based on a state value deviation and a state transition probability, the state value deviation is obtained by weighting and combining the difference between adjacent time state variable evaluation values and state variable importance weights, and the state transition probability is obtained by multiplying the conditional probabilities of state variables at adjacent times; the strategy evaluation function calculates a strategy evaluation value based on a strategy gradient and a strategy entropy, the strategy gradient is calculated by multiplying the logarithm of a parameterized strategy and cumulative rewards, and the strategy entropy is calculated by the information entropy of a strategy distribution; the dynamic weight coefficient is obtained by weighting and combining the state value deviation, the state transition probability, the strategy gradient and the strategy entropy and mapping through a sigmoid function; a reward change trend and a strategy convergence degree are calculated based on the shared reward function, the reward change trend is calculated by the change rate of the shared reward function within a time window, and the strategy convergence degree is calculated by the average value of the difference between adjacent time strategy parameters; the cooperative learning rate is obtained by multiplying an initial learning rate and the product of the absolute value of the reward change trend and the exponential decay function of the strategy convergence degree, according to the dynamic adjustment of the reward change trend and the strategy convergence degree.
[0016] In an optional embodiment, the step of training the reinforcement learning framework through the agent cooperative optimization mechanism to obtain the optimal charging strategy and the optimal discharging strategy comprises:
[0017] An experience replay pool is constructed, the experience replay pool comprising a priority experience replay unit and a shared experience replay unit, the priority experience replay unit performing importance sorting on training samples based on a time difference error, and the shared experience replay unit storing common experience data of the charging agent and the discharging agent; a distributed training framework is constructed based on the experience replay pool, the distributed training framework comprising a parameter server and a plurality of training agents, the parameter server maintaining global policy parameters of the charging agent and the discharging agent, and the training agents performing parallel computation of policy gradients based on local samples; a training evaluation index is constructed, the training evaluation index comprising a policy stability index and a training convergence index, the policy stability index being calculated by a policy output variance, and the training convergence index being calculated by a value function estimation error; when the policy stability index is less than a first preset threshold and the training convergence index is less than a second preset threshold, the optimal charging strategy of the charging agent and the optimal discharging strategy of the discharging agent are obtained.
[0018] In a second aspect, a cylindrical lithium battery charging and discharging strategy optimization system based on reinforcement learning is provided, comprising: a first unit for evaluating battery remaining life and performance degradation trend based on real-time monitoring data and historical monitoring data of the cylindrical lithium battery to generate a prediction result; determining a current charging and discharging state based on the real-time monitoring data; a second unit for establishing a reinforcement learning framework, the reinforcement learning framework comprising a charging agent and a discharging agent, a state space of the charging agent and the discharging agent containing the prediction result and the current charging and discharging state, and an action space being a charging and discharging power range under safety constraints; a third unit for constructing an agent cooperative optimization mechanism, the agent cooperative optimization mechanism comprising a shared reward function and a state sharing mechanism, the shared reward function combining a battery life extension index and a charging and discharging efficiency index into a common optimization goal of the charging agent and the discharging agent, and the state sharing mechanism enabling a decision result of the charging agent to be input as a state of the discharging agent and enabling a decision result of the discharging agent to be input as a state of the charging agent; and a fourth unit for training the reinforcement learning framework through the agent cooperative optimization mechanism to obtain an optimal charging strategy and an optimal discharging strategy.
[0019] In a third aspect, a computer-readable storage medium is provided, having computer program instructions stored thereon, the computer program instructions being executed by a processor to implement the method described above.
[0020] The application realizes the overall optimization of the lithium battery charging and discharging strategy by establishing a reinforcement learning framework for the cooperative optimization of the charging agent and the discharging agent. The shared reward function mechanism takes the battery life and the charging and discharging efficiency as the comprehensive optimization target, so that the charging and discharging strategy can adaptively balance the speed and life demand under different working conditions. The state sharing mechanism enables the two agents to fully consider each other's decision results, effectively coordinates the charging and discharging process, and avoids the suboptimal problem caused by the independent control of charging and discharging in the traditional method. In addition, the dynamic weight coefficient and adaptive learning rate mechanism of the application can adjust the strategy focus in real time according to the battery health state and working environment, significantly improving the adaptability of the control strategy to the battery degradation process. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 The flowchart of the cylindrical lithium battery charging and discharging strategy optimization method based on reinforcement learning of the embodiments of the application is shown in the figure. DETAILED DESCRIPTION
[0022] The technical solutions in the embodiments of the application will be described below with reference to the drawings in the embodiments of the application. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments.
[0023] Figure 1 The flowchart of the cylindrical lithium battery charging and discharging strategy optimization method based on reinforcement learning of the application is shown in the figure. Figure 1 As shown in the figure, the method comprises:
[0024] According to the real-time monitoring data and the historical monitoring data of the cylindrical lithium battery, the remaining life and the performance degradation trend of the battery are evaluated, and a prediction result is generated. According to the real-time monitoring data, the current charging and discharging state is determined, and it is determined whether it is in a single charging state, a single discharging state or a charging and discharging parallel state. A reinforcement learning framework is established, which includes a charging agent and a discharging agent. The state space of the charging agent and the discharging agent includes the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints. An agent cooperative optimization mechanism is constructed, which includes a shared reward function and a state sharing mechanism. The shared reward function combines the battery life extension index and the charging and discharging efficiency index as the common optimization target of the charging agent and the discharging agent. The state sharing mechanism enables the decision result of the charging agent to be used as the state input of the discharging agent, and the decision result of the discharging agent to be used as the state input of the charging agent. The reinforcement learning framework is trained through the agent cooperative optimization mechanism to obtain the optimal charging strategy and the optimal discharging strategy.
[0025] In an alternative embodiment, the step of evaluating the remaining life and performance degradation trend of the cylindrical lithium battery according to the real-time monitoring data and the historical monitoring data of the cylindrical lithium battery, and generating a prediction result comprises:
[0026] The real-time monitoring data and the historical monitoring data both include voltage data, current data, temperature data and state of charge data of the cylindrical lithium battery;
[0027] Based on the real-time monitoring data and the historical monitoring data, an instantaneous feature vector, a short-term feature vector and a long-term feature vector are respectively constructed, and the long-term feature vector is obtained by weighted cumulative calculation of the relative capacity and the relative internal resistance of the lithium battery through a time decay factor;
[0028] The contribution scores of the instantaneous feature vector, the short-term feature vector and the long-term feature vector to the prediction result are calculated, the corresponding feature weights are determined according to the contribution scores, the feature weights are multiplied by and superimposed with the corresponding feature vectors respectively, and a multi-scale feature vector is obtained;
[0029] The multi-scale feature vector is input into a long short-term memory network, the hidden state and the input feature of the long short-term memory network are weighted through a time attention mechanism and a feature attention mechanism, and a remaining life prediction value and a performance degradation rate are obtained; the performance degradation trend of the lithium battery is predicted according to the performance degradation rate, and the remaining life prediction value and the performance degradation trend are adaptively corrected according to the historical prediction error, and a prediction result is obtained.
[0030] For example, the voltage data, current data, temperature data and state of charge data of the cylindrical lithium battery are collected by a multi-channel acquisition module. The voltage data acquisition resolution is 1 mV, and the sampling frequency is 10 Hz; the current data acquisition resolution is 10 mA, and the sampling frequency is 10 Hz; the temperature data is collected by a surface thermistor, with a resolution of 0.1 °C and a sampling frequency of 1 Hz; the state of charge data is obtained by the coulomb counting method combined with open circuit voltage correction, with an accuracy of 1%. The real-time monitoring data is stored in a circular buffer, and the last 72 hours of data are retained; the historical monitoring data is stored in a non-volatile memory, and contains data records of complete charge and discharge cycles of the battery, with a maximum of 1000 cycles.
[0031] Based on the collected real-time monitoring data and historical monitoring data, three time-scale feature vectors are constructed. The instantaneous feature vector contains data features in the last 10 minutes, which are obtained by calculating the statistical features of voltage, current, temperature and state of charge, including mean, standard deviation, maximum, minimum and change rate, forming a 20-dimensional feature vector. The short-term feature vector contains data features in the last 24 hours, which are obtained by segmenting the data, one segment per hour, calculating the statistical features in each segment and extracting the time series pattern, forming a 48-dimensional feature vector. The long-term feature vector is obtained by modeling the long-term change trend of battery capacity and internal resistance. The relative capacity and relative internal resistance values are calculated once a complete charge and discharge cycle is completed. The relative capacity is the ratio of the current releasable capacity to the rated capacity, and the relative internal resistance is the ratio of the current internal resistance to the initial internal resistance. In order to reflect the time cumulative effect, a time decay factor of 0.95 is introduced to weight and accumulate the relative capacity and relative internal resistance. In specific implementation, the relative capacity and relative internal resistance of the last 50 cycles are arranged in chronological order, with the more recent data having higher weight, forming a 100-dimensional feature vector.
[0032] In order to reasonably fuse features of different time scales, the contribution score of each feature vector to the prediction result is calculated. The contribution score calculation uses the feature importance evaluation method based on gradient boosting tree. A preliminary prediction model is trained using historical data, and the prediction performance is evaluated when only using the instantaneous feature vector, the short-term feature vector and the long-term feature vector respectively. In the initial use stage of the battery (cycle number less than 200 times), the contribution scores of the instantaneous feature vector and the short-term feature vector are higher, which are 0.4 and 0.4 respectively, and the contribution score of the long-term feature vector is 0.2; in the middle use stage (cycle number between 200 and 600 times), the contribution scores of the three feature vectors are all 0.33; in the later use stage (cycle number greater than 600 times), the contribution score of the long-term feature vector increases to 0.5, and the contribution scores of the instantaneous feature vector and the short-term feature vector decrease to 0.2 and 0.3 respectively. According to these contribution scores, the corresponding feature weights are determined, the feature weights are multiplied by the corresponding feature vectors respectively and superimposed to obtain the fused multi-scale feature vector, which has a dimension of 168.
[0033] The multi-scale feature vector is input into a long short-term memory network for processing. The network includes two layers of long short-term memory layers, each layer including 128 memory units, the input layer corresponding to a 168-dimensional feature vector, and the output layer having 2 nodes corresponding to the remaining life prediction value and the performance degradation rate. To improve prediction accuracy, a time attention mechanism and a feature attention mechanism are introduced. The time attention mechanism calculates the importance weight of the features at different time points, enabling the network to focus on the information at key time points. The hidden states of the last 30 time points are input into the attention calculation unit to obtain the weight coefficients of each time point, with the value ranging from 0 to 1 and the total weight being 1. In the rapid degradation stage of the battery, higher attention weights are assigned to the most recent time points; in the stable stage of the battery, the attention weights are more evenly distributed. The feature attention mechanism calculates the importance weight of different feature dimensions to enhance the sensitivity of the model to key features. The attention coefficients are calculated for each dimension of the 168-dimensional feature vector, with higher weights being assigned to abnormal features such as voltage jumps and capacity jumps, with typical values ranging from 0.02 to 0.05, and the weights of regular features ranging from 0.002 to 0.008.
[0034] The hidden states and input features of the long short-term memory network are processed by weighting to obtain the remaining life prediction value and the performance degradation rate. The remaining life prediction value is expressed in cycle times, indicating the remaining cycle times required for the battery capacity to drop to 80% of the rated capacity. The performance degradation rate is expressed in percentage of capacity loss per cycle, such as 0.03% / cycle, indicating that the battery capacity is reduced by an average of 0.03% after each complete charge and discharge cycle. According to the performance degradation rate, the performance degradation trend of the lithium battery is predicted, and the degradation rate is divided into low speed (less than 0.02% / cycle), medium speed (0.02% to 0.05% / cycle), and high speed (greater than 0.05% / cycle), and the capacity change curve in the next 50 cycles is predicted.
[0035] The actual error of the last 20 predictions is calculated to obtain the average error and error trend. If it is found that the predicted value is consistently higher or lower than the actual value, the bias term of the prediction model will be automatically adjusted. For example, when the average of the remaining life values predicted for 5 consecutive times is 15% higher than the actual value, the current prediction value will be corrected by 0.85 times; when the error is randomly distributed and the average error is less than 10%, no correction is made. Through this adaptive correction method, the calibrated remaining life prediction value and performance degradation trend are finally output as the prediction result for use by the subsequent charge and discharge strategy optimization module.
[0036] The present application can effectively integrate battery state information at different time scales, adaptively adjust feature weights, dynamically focus on key time points and important features, and significantly improve prediction accuracy and stability. At the same time, through the historical error adaptive correction mechanism, the prediction model can be continuously optimized.
[0037] In an alternative embodiment, a reinforcement learning framework is established, which includes a charging agent and a discharging agent, the state space of the charging agent and the discharging agent contains the prediction results and the current charging and discharging state, and the step of determining the action space as the charging and discharging power range under the safety constraint comprises:
[0038] The prediction results and the current charging and discharging state are time window sliding encoded to obtain a time sequence feature sequence, a self-attention mechanism is applied to the time sequence feature sequence to calculate the correlation weight between different time points and different state variables, and a reduced dimension state representation is generated by weighted fusion as the state space of the charging agent and the discharging agent; the internal polarization voltage and temperature distribution of the battery are calculated based on an electrochemical reduction model, and an adaptive safety boundary is established in combination with the performance degradation trend, the adaptive safety boundary is dynamically adjusted with the performance degradation trend, and the charging and discharging power range within the adaptive safety boundary is set as the action space of the charging agent and the discharging agent, respectively; a hierarchical decision network is constructed in the charging agent, which includes a policy network and an evaluation network, the policy network adopts a double-delay deterministic policy architecture to output a charging power distribution, and the evaluation network evaluates the influence of the charging power on the battery life based on the battery operating characteristics and corrects the policy output; a temperature-sensitive power modulator is constructed in the discharging agent, which calculates the optimal power trajectory based on the battery equivalent circuit characteristics, and shapes the discharging current by using the pulse width modulation technology to realize temperature equalization control.
[0039] For example, the prediction results include the residual cycle life value and the performance degradation rate, and the current charging and discharging state includes the battery voltage, current, temperature distribution and state of charge, the data is time window sliding encoded, a time window of 60 seconds is selected, and the sampling is sliding with a step of 1 second to generate a time sequence feature sequence. The sequence contains the state information of 60 time points, each time point contains 10 state variables, forming a feature matrix of 60x10.
[0040] The self-attention mechanism is applied to the time series of features to calculate the correlation weights between different time points and different state variables. A multi-head self-attention layer is constructed, which contains 8 attention heads, each of which is responsible for capturing different types of correlation patterns. The correlation score is calculated by transforming the query matrix, key matrix, and value matrix, and the correlation score is converted into a weight coefficient through the softmax function, with a value range of 0 to 1 and a total sum of 1. In actual operation, the correlation weight between temperature and current is usually between 0.15 and 0.25, and the correlation weight between voltage and state of charge is between 0.2 and 0.3. Based on the correlation weight, the feature matrix is weighted and fused to generate the reduced state representation. Principal component analysis is used for dimensionality reduction, which compresses the original 600-dimensional features to 64-dimensional features, retaining about 95% of the information. The 64-dimensional state representation is used as the state space input of the charging agent and the discharging agent.
[0041] The internal polarization voltage and temperature distribution of the battery are calculated based on the electrochemical reduced-order model to set a safe action space for the agent. The electrochemical reduced-order model uses an equivalent circuit model structure, which includes an open-circuit voltage source, internal resistance, a first-order RC network, and a thermal resistance and thermal capacity network. The model parameters are obtained through offline calibration, with typical values of internal resistance of 10 to 30 milliohms, polarization resistance of 5 to 15 milliohms, polarization capacitance of 1000 to 3000 farads, thermal resistance of 0.5 to 1.5 degrees Celsius / watt, and thermal capacity of 80 to 120 joules / degree Celsius. The real-time input current value is used to calculate the internal node voltage and temperature values of the battery. The polarization voltage and temperature thresholds are dynamically adjusted based on the performance degradation trend obtained in the previous step. For batteries with low degradation rate (less than 0.02% per cycle), the polarization voltage threshold is set to 200 millivolts, and for batteries with medium (0.02% to 0.05% per cycle) and high (more than 0.05% per cycle) degradation rates, the polarization voltage thresholds are reduced to 180 millivolts and 150 millivolts, respectively. Similarly, the surface temperature threshold and internal temperature threshold are also adjusted with the degradation rate, with values of 45 degrees Celsius and 50 degrees Celsius for low degradation rate, 43 degrees Celsius and 48 degrees Celsius for medium degradation rate, and 40 degrees Celsius and 45 degrees Celsius for high degradation rate.
[0042] An adaptive safety margin is established according to the performance degradation trend, and is dynamically adjusted with the degree of degradation. Not only the current health state is considered, but also the performance degradation rate as a key influencing factor. For batteries with the same health state but different degradation rates, differentiated safety margin strategies are adopted. The battery health state is divided into three levels of high, medium and low, corresponding to the capacity retention rate greater than 90%, 75% to 90% and less than 75%. For the battery with high health state and low degradation rate, the maximum charging power is allowed to be 1.2 times the nominal power, and the maximum discharging power is 1.5 times the nominal power; when the degradation rate is medium, even if the health state is high, the maximum charging power will be reduced to 1.1 times the nominal power; when the degradation rate is high, it will be further reduced to 1.0 times. For the battery with medium health state, when the degradation rate is low, medium and high, the maximum charging power is set to 1.0 times, 0.9 times and 0.8 times the nominal power respectively, and the maximum discharging power is set to 1.2 times, 1.1 times and 1.0 times the nominal power respectively. For the battery with low health state, when the degradation rate is low, medium and high, the maximum charging power is set to 0.8 times, 0.7 times and 0.6 times the nominal power respectively, and the maximum discharging power is set to 1.0 times, 0.9 times and 0.8 times the nominal power respectively. The charging and discharging power ranges in these adaptive safety margins are set as the action space of the charging agent and the discharging agent respectively, and the charging power is discretized into 10 levels and the discharging power is discretized into 15 levels.
[0043] A hierarchical decision-making network is constructed in the charging agent, including a policy network and an evaluation network. The policy network adopts a double-delay deterministic policy architecture, including an online policy network and a target policy network. The online policy network consists of a three-layer fully connected neural network with a structure of 64-128-256-10, an input of 64-dimensional state representation, and an output of probability distribution of 10 charging power levels. The hidden layer uses ReLU activation function, and the output layer uses softmax function to ensure the probability sum to be 1. The structure of the target policy network is the same as that of the online policy network, but the parameter update frequency is lower. A soft update coefficient of 0.01 is adopted, that is, after each training iteration, the target network parameters are only moved to the online network parameters by 1%. This double-delay architecture significantly improves the training stability and reduces the policy oscillation phenomenon.
[0044] The evaluation network evaluates the impact of charging power on battery life based on battery operating characteristics and corrects the strategy output. The evaluation network includes a capacity attenuation evaluation module, a temperature stress evaluation module, and a polarization stress evaluation module. The capacity attenuation evaluation module calculates the capacity loss rate under different charging currents, and the input is the charging current and the battery temperature, and the output is the percentage of capacity loss per cycle. The temperature stress evaluation module calculates the deviation of the current temperature from the optimal operating temperature, which is set to 25 degrees Celsius. The temperature stress index increases by 0.2 for every 5 degrees Celsius of temperature deviation. The polarization stress evaluation module calculates the cumulative effect of the polarization voltage. When the polarization voltage exceeds 150 millivolts, the polarization stress index increases exponentially. The outputs of the three modules are combined to generate a strategy correction signal, with weights of 0.5, 0.3, and 0.2 respectively. The 64-dimensional state representation is processed with attention to obtain attention weights that focus on battery health status and temperature distribution. The initial charging power decision is generated by combining the charging power distribution with the attention weights, and then adjusted according to the strategy correction signal to obtain the final charging power output.
[0045] A temperature-sensitive power modulator is constructed in the discharge agent, which includes a power trajectory generator and a current shaping controller. The power trajectory generator calculates the optimal discharge power curve based on the current state of the battery and the discharge demand. A feedforward neural network structure is used, with the battery state and discharge demand as input, and the power trajectory points for the next 30 seconds as output. The network structure is 64-128-128-30, trained by a regression loss function. The generated power trajectory needs to meet the total energy demand and the maximum power constraint. The typical discharge curve has low power at the beginning, stable power at the optimal efficiency point in the middle, and gradually decreasing power at the end. The current shaping controller uses pulse width modulation technology to shape the discharge current. The basic switching frequency is set to 100Hz, and the duty cycle is dynamically adjusted according to the temperature distribution. For areas with higher temperatures, the duty cycle is reduced to reduce heat generation; for areas with lower temperatures, the duty cycle is increased to improve energy efficiency. The battery surface is divided into 5 temperature monitoring areas, and when the temperature difference between areas exceeds 5 degrees Celsius, differential modulation is started. The duty cycle of the area with the highest temperature can be reduced to 70% of the standard value, and the duty cycle of the area with the lowest temperature can be increased to 120% of the standard value.
[0046] The reinforcement learning framework established by the application effectively extracts the correlation information between the time sequence features through the self-attention mechanism, constructs an accurate safety boundary using the electrochemical reduction model, and realizes the dynamic optimization of the charging and discharging strategy through the mutual cooperation of the hierarchical decision network of the charging agent and the temperature-sensitive power modulator of the discharge agent. This method can adaptively adjust the control strategy according to the battery health status, balance the charging speed and life extension demand, and significantly improve the energy efficiency and service life of the battery.
[0047] In an alternative embodiment, a hierarchical decision network is constructed in the charging agent, which includes a policy network and an evaluation network, the policy network outputs a charging power distribution using a double-delay deterministic policy architecture, and the evaluation network evaluates the impact of the charging power on the battery life based on the battery operating characteristics and corrects the policy output, which includes the steps of:
[0048] The policy network includes an online policy network and a target policy network, the online policy network receives the state representation, calculates the mean parameter and the standard deviation parameter of the charging power distribution through the hidden layer state, and the network parameters of the target policy network are updated by weighting the network parameters of the online policy network through a soft update coefficient; in the evaluation network, the capacity decay rate index, the temperature stress index and the polarization stress index are calculated based on the battery operating characteristics, the capacity decay rate index is calculated by the capacity loss under different operating conditions, the temperature stress index is calculated by the deviation of the current temperature from the optimal operating temperature, and the polarization stress index is calculated by the cumulative effect of the polarization voltage; a policy correction signal is generated based on the capacity decay rate index, the temperature stress index and the polarization stress index, the policy correction signal is obtained by weighting and combining each index through an adaptive weight coefficient; the attention weight is obtained by attention processing of the state representation, the attention weight and the mean parameter and the standard deviation parameter of the charging power distribution are combined to generate an initial charging power, and the initial charging power is corrected according to the policy correction signal to obtain the charging power.
[0049] For example, the hierarchical decision network includes two core components, a policy network and an evaluation network, which work together to achieve intelligent control of charging power. The policy network uses a double-delay deterministic policy architecture, which includes an online policy network and a target policy network. The online policy network directly receives the 64-dimensional state representation generated in the foregoing step, processes it through a four-layer neural network, and the network structure is 64-128-256-128-2, wherein the input layer receives the 64-dimensional state representation, the intermediate hidden layers have 128, 256 and 128 neurons respectively, and the output layer has 2 neurons, corresponding to the mean parameter and the standard deviation parameter of the charging power distribution respectively. The hidden layer uses a LeakyReLU activation function, and the negative half-axis slope of the activation function is set to 0.01 to avoid the problem of gradient disappearance. The output layer does not use an activation function and directly outputs the original value, the effective range of the mean parameter is 0.1 to 1.0, representing the proportion relative to the rated charging power; the effective range of the standard deviation parameter is 0.01 to 0.2, used to control the fluctuation range of the charging power.
[0050] The target policy network structure is completely the same as the online policy network, but the network parameter updating method is different. The target network parameter is updated by a soft updating mechanism to weight the online network parameter, and the soft updating coefficient is set to 0.01. In actual implementation, in each training step, the parameter value of the target network is equal to 0.99 times the parameter value of the target network in the last step plus 0.01 times the parameter value of the current online network. This soft updating mechanism significantly improves the training stability and reduces the policy shock phenomenon. In actual operation, state evaluation and decision calculation are performed every 100 milliseconds, and parameter updating is performed every 10 seconds. The updating frequency of the target network is 10 times lower than that of the online network, and it is updated every 100 seconds. This design further enhances the stability of the decision.
[0051] The evaluation network is responsible for evaluating the impact of charging power on battery life based on battery working characteristics, and generating a policy correction signal. The evaluation network contains three key modules: a capacity attenuation rate evaluation module, a temperature stress evaluation module, and a polarization stress evaluation module. The capacity attenuation rate evaluation module calculates the capacity loss rate under different charging currents, with the input being the charging current and the battery temperature, and the output being the percentage of capacity loss per cycle. Inside the module, a two-layer neural network is used, with a structure of 2-32-64-1, obtained through offline data training. The module takes the current charging current value (expressed in C rate, i.e., the ratio to the rated capacity) and the battery surface temperature as input, and outputs the capacity attenuation rate index. For example, at 25 degrees Celsius, the capacity attenuation rate index corresponding to a 0.5C charging current is 0.018, the capacity attenuation rate index corresponding to a 1.0C charging current is 0.025, and the capacity attenuation rate index corresponding to a 1.5C charging current is 0.042. When the temperature rises to 40 degrees Celsius, the capacity attenuation rate indices under the same current increase to 0.024, 0.035, and 0.065, respectively, reflecting the significant impact of temperature on battery life.
[0052] The temperature stress evaluation module calculates the deviation of the current temperature from the optimal working temperature and converts it into a temperature stress index. The optimal working temperature is set to 25 degrees Celsius, and the temperature stress index increases by 0.05 for every 1 degree Celsius deviation. To reflect the nonlinear impact of temperature deviation on battery life, an exponential conversion function is used for temperature deviation, and the temperature stress index growth rate doubles when the temperature deviation exceeds 10 degrees Celsius. For example, 30 degrees Celsius corresponds to a temperature stress index of 0.25, 35 degrees Celsius corresponds to a temperature stress index of 0.5, 40 degrees Celsius corresponds to a temperature stress index of 1.0, and 45 degrees Celsius corresponds to a temperature stress index of 2.0. The temperature stress evaluation module performs a weighted average of the temperatures at various points on the battery, with a weight of 0.6 for the center region and a weight of 0.4 for the surface region, to comprehensively reflect the temperature state of the battery.
[0053] The polarization stress evaluation module calculates the polarization stress indicator through the cumulative effect of the polarization voltage. Based on the aforementioned electrochemical reduced-order model, the battery polarization voltage is calculated in real time, and the maximum value, average value, and duration of the polarization voltage during the recent charging process are recorded. These three values are compared with the preset threshold to generate the polarization stress indicator. The polarization voltage maximum value threshold is set to 180 millivolts, the average value threshold is set to 120 millivolts, and the duration threshold is set to 300 seconds. When the actual value exceeds the corresponding threshold, the polarization stress indicator increases according to the percentage of excess. For example, the polarization voltage maximum value is 200 millivolts, which is 11.1% higher than the threshold, contributing 0.111 to the polarization stress indicator; the polarization voltage average value is 150 millivolts, which is 25% higher than the threshold, contributing 0.25 to the polarization stress indicator; and the duration is 450 seconds, which is 50% higher than the threshold, contributing 0.5 to the polarization stress indicator. The three contribution values are combined by weighted sum, with weights of 0.3, 0.3, and 0.4, respectively, to obtain the final polarization stress indicator.
[0054] The strategy correction signal is generated based on the capacity decay rate indicator, temperature stress indicator, and polarization stress indicator. The three indicators are combined by weighted sum using adaptive weight coefficients, and the adaptive weights are dynamically adjusted according to the battery health state and working environment. For new batteries (health state greater than 90%), the weights of the capacity decay rate indicator, temperature stress indicator, and polarization stress indicator are 0.2, 0.3, and 0.5, respectively; for medium-term batteries (health state between 75% and 90%), the weights are adjusted to 0.3, 0.4, and 0.3; for aged batteries (health state less than 75%), the weights are adjusted to 0.5, 0.3, and 0.2. This dynamic weight adjustment reflects the change in sensitivity of batteries at different aging stages to various stress factors. The strategy correction signal after weighted combination has a value range of 0 to 1, and the larger the value, the greater the correction degree of the strategy. In actual operation, when the capacity decay rate indicator is 0.02, the temperature stress indicator is 0.3, and the polarization stress indicator is 0.1, the generated strategy correction signal is 0.16 for a battery with a health state of 95%; when the health state decreases to 80%, the generated strategy correction signal increases to 0.22 under the same indicators.
[0055] The 64-dimensional state representation is processed by attention to obtain attention weights. The attention processing uses a single-layer feedforward neural network with a structure of 64-64-64, and the input is the state representation and the output is a 64-dimensional attention weight. The attention weight is normalized by the softmax function to ensure that the sum is 1. In actual operation, the feature dimensions related to the battery health state obtain higher attention weights (usually 0.03 to 0.05), the feature dimensions related to the instantaneous current and voltage obtain medium attention weights (usually 0.01 to 0.03), and the feature dimensions related to environmental factors obtain lower attention weights (usually 0.005 to 0.01). This differentiated attention allocation enables focus on state variables that have a greater impact on decision-making.
[0056] The attention weight is combined with the mean parameter and the standard deviation parameter of the charging power distribution to generate an initial charging power, the state representation is weighted and summed using the attention weight to obtain a core feature value of the state representation, and the mean parameter is fine-tuned according to the core feature value. When the core feature value is greater than a preset threshold 0.6, the mean parameter is increased by 0.05; when the core feature value is less than a preset threshold 0.4, the mean parameter is decreased by 0.05. The standard deviation parameter is dynamically adjusted according to the uncertainty of the state, and when the state is stable, the standard deviation parameter is reduced to 0.05, and when the state fluctuates, the standard deviation parameter is increased to 0.15. The fine-tuned mean parameter and the standard deviation parameter are used to generate a Gaussian distribution, and an initial charging power value is sampled from the distribution.
[0057] The initial charging power is corrected according to the policy correction signal to obtain a final charging power. The correction method is linear reduction, and the corrected charging power is equal to the initial charging power multiplied by (1-policy correction signal). For example, the initial charging power is 0.8 times the rated power, and the policy correction signal is 0.25, then the corrected charging power is 0.8x(1-0.25)=0.6 times the rated power. The corrected charging power is subjected to amplitude limiting processing to ensure that it is within the safety boundary. The final output charging power control signal is transmitted to the power control unit of the battery management to realize accurate adjustment of the charging current.
[0058] The hierarchical decision network in the embodiment realizes stable charging decision through the online policy network and the target policy network, and the evaluation network comprehensively considers the influence of capacity attenuation, temperature stress and polarization stress on the battery life. The attention mechanism enables adaptive attention to key state information, and the policy correction mechanism dynamically balances the charging efficiency and the battery life.
[0059] In an optional implementation, the step of constructing the agent cooperative optimization mechanism comprises:
[0060] The battery life extension index is calculated based on the capacity loss amount, initial capacity, working temperature and optimal temperature, and the charge-discharge efficiency index is calculated based on the charging efficiency, discharging efficiency and state of charge; a state sharing mechanism is constructed, the state sharing mechanism includes local state information and shared state information, the charging agent inputs the state of charge, charging power, temperature, voltage and internal resistance as the shared state information of the discharging agent, and the discharging agent inputs the discharging power, discharging efficiency and health state as the shared state information of the charging agent; an adaptive cooperation mechanism is constructed, the adaptive cooperation mechanism constructs a joint value function based on the local state information and shared state information, the joint value function includes charging value evaluation, discharging value evaluation and shared state value evaluation, the instant reward is calculated based on the joint value function, and the instant reward is used to evaluate the current decision effect of the charging agent and the discharging agent; the adaptive cooperation mechanism calculates a dynamic weight coefficient through a state evaluation function and a strategy evaluation function, combines the battery life extension index and the charge-discharge efficiency index by weighting through the dynamic weight coefficient to obtain a shared reward function, dynamically adjusts the cooperative learning rate based on the reward change trend and the strategy convergence degree, and jointly optimizes and updates the decision strategies of the charging agent and the discharging agent according to the cooperative learning rate.
[0061] For example, the battery life extension index and the charge-discharge efficiency index are defined as evaluation indexes. The battery life extension index is calculated by the capacity loss amount, the initial capacity, the working temperature and the optimal temperature. The capacity loss amount is obtained by combining the coulomb counting method and internal resistance change monitoring, and the available capacity reduction value after each cycle is recorded, with the unit being ampere-hour. The initial capacity is the nominal capacity of the battery when it leaves the factory, which is 3000 milliampere-hours for a typical 18650 type cylindrical lithium battery. The working temperature is obtained by a multi-point temperature sensor, with the accuracy being 0.5 degrees Celsius and the sampling frequency being 1 hertz. The optimal temperature is set to 25 degrees Celsius, at which the battery cycle life and efficiency reach the best balance. The relative capacity loss rate is obtained by dividing the capacity loss amount by the initial capacity, and then the temperature is corrected according to the deviation of the working temperature from the optimal temperature. The temperature correction adopts an exponential decay model, and the life extension index decreases by 5% for every 1 degree Celsius deviation of the temperature from the optimal temperature. In actual calculation, when the working temperature is 30 degrees Celsius, the battery life extension index is 75% of the standard value; when the working temperature is 35 degrees Celsius, the index decreases to 55% of the standard value; and when the working temperature is 40 degrees Celsius, the index further decreases to 40% of the standard value.
[0062] The charge-discharge efficiency indicator is calculated based on the charging efficiency, the discharging efficiency, and the state of charge. The charging efficiency is defined as the ratio of the energy stored by the battery to the input energy, and the discharging efficiency is defined as the ratio of the output energy to the energy released by the battery. The energy flow during the charging and discharging process is measured by a high-precision voltage and current sampling circuit, with a sampling accuracy of 16 bits and a sampling frequency of 1000 Hz. Both the charging efficiency and the discharging efficiency are affected by the current size and the state of charge. For 18650-type lithium batteries, the charging efficiency is typically 92% to 95% at a 0.5C charging current, decreases to 88% to 92% at a 1C charging current, and further decreases to 83% to 87% at a 1.5C charging current. The discharging efficiency is 94% to 97% at a 0.5C discharging current, 90% to 94% at a 1C discharging current, and 85% to 90% at a 2C discharging current. The effect of the state of charge on the efficiency exhibits a U-shaped curve, with the highest efficiency within the range of 20% to 80% state of charge, and a significant decrease in efficiency below 20% or above 80%. Based on the real-time measured charging and discharging efficiency and the current state of charge, the charging-discharging efficiency indicator is calculated by cubic spline interpolation, with a value range of 0 to 1, and a larger value indicating higher efficiency.
[0063] A state sharing mechanism is constructed to enable the charging agent and the discharging agent to access each other's key state information. The state sharing mechanism includes local state information and shared state information. The local state information is the state data that the agent can directly obtain, including the aforementioned 64-dimensional state representation in the state space. The shared state information is the key decision results and state assessments transferred from another agent. The charging agent inputs the state of charge, the charging power, the temperature, the voltage, and the internal resistance as the shared state information of the discharging agent. The state of charge has a precision of 1%, the charging power is expressed as a percentage of the rated power, the temperature includes the surface temperature and the estimated internal temperature, the voltage has a precision of 10 mV, and the internal resistance includes the direct current resistance and the alternating current resistance. The discharging agent inputs the discharging power, the discharging efficiency, and the health state as the shared state information of the charging agent. The discharging power is expressed as a percentage of the rated power, the discharging efficiency includes energy efficiency and temperature efficiency, and the health state is expressed as a percentage of the initial capacity. The shared state information is exchanged in real time through a dedicated communication buffer, with an update frequency of 10 Hz, ensuring that both agents can obtain the latest decision results and state assessments of the other agent in a timely manner.
[0064] An adaptive coordination mechanism is constructed to build a joint value function based on local state information and shared state information. The joint value function includes three components: charging value evaluation, discharging value evaluation, and shared state value evaluation. Charging value evaluation is based on the influence of charging strategy on battery life and efficiency, and is implemented through a three-layer neural network with a structure of 20-40-20-1, input of charging-related state variables, and output of charging value score. Discharging value evaluation is based on the influence of discharging strategy on discharging duration and temperature balance, and is also implemented through a three-layer neural network with a structure of 20-40-20-1, input of discharging-related state variables, and output of discharging value score. Shared state value evaluation is based on the coordination effect of the decisions of the two agents, and evaluates the consistency and complementarity of the decisions of the two agents through processing of shared state information. Shared state value evaluation uses an attention mechanism to weight the shared state information, focusing on battery health state and temperature distribution information, and outputs a shared value score.
[0065] Based on the joint value function, an immediate reward is calculated, which is used to evaluate the current decision effect of the charging agent and the discharging agent. The immediate reward calculation adopts a weighted sum method, and the weights of charging value evaluation, discharging value evaluation, and shared state value evaluation are 0.3, 0.3, and 0.4, respectively. In actual operation, when the charging power and discharging power are set reasonably, the battery temperature is maintained within the optimal range, and the charging and discharging efficiency is high, a high-value immediate reward is generated, with a typical value of 0.8 to 0.95; when the charging power is too high, causing the temperature to rise, or the discharging power fluctuates greatly, causing the efficiency to decrease, a medium-value immediate reward is generated, with a typical value of 0.5 to 0.7; when the charging and discharging strategy causes the battery temperature to be too high or the polarization phenomenon to be serious, a low-value immediate reward is generated, with a typical value of 0.2 to 0.4.
[0066] The adaptive coordination mechanism calculates dynamic weight coefficients through state evaluation functions and strategy evaluation functions. The state evaluation function calculates state evaluation values based on state value deviation and state transition probability. The state value deviation is obtained by the weighted sum of the difference between the evaluation values of adjacent time state variables and the importance weights of state variables. Higher importance weights are given to key state variables such as battery voltage, temperature, and state of charge, with typical values of 0.15 to 0.25; lower importance weights are given to secondary state variables such as ambient temperature and historical usage patterns, with typical values of 0.05 to 0.1. The state transition probability is obtained by the product of the conditional probabilities of state variables at adjacent times, reflecting the stability of state changes. The strategy evaluation function calculates strategy evaluation values based on strategy gradient and strategy entropy. The strategy gradient is calculated by the product of the strategy logarithm and the cumulative reward, reflecting the improvement direction of the current strategy. The strategy entropy is calculated by the information entropy of the strategy distribution, reflecting the exploration degree of the strategy.
[0067] The dynamic weight coefficient is obtained by a weighted combination of state value deviation, state transition probability, policy gradient and policy entropy, and is mapped by a sigmoid function, with a value range of 0 to 1. When the state changes smoothly and the policy gradient direction is clear, the dynamic weight coefficient tends to 1, indicating that more attention is paid to life extension; when the state changes rapidly or the policy uncertainty is high, the dynamic weight coefficient tends to 0.5, indicating a balance between life extension and efficiency improvement; when a rapid response to external demand is needed, the dynamic weight coefficient can be reduced to 0.3, paying more attention to charging and discharging efficiency. The shared reward function is obtained by weighting and combining the battery life extension index and the charging and discharging efficiency index through the dynamic weight coefficient. The calculation method of the shared reward function is: the dynamic weight coefficient multiplied by the battery life extension index plus (1-dynamic weight coefficient) multiplied by the charging and discharging efficiency index.
[0068] The collaborative learning rate is dynamically adjusted based on the reward change trend and the policy convergence degree. The reward change trend is calculated by the change rate of the shared reward function within a time window, and the size of the time window is 100 decision cycles. The policy convergence degree is calculated by the average value of the difference between the policy parameters at adjacent time points, reflecting the stability of policy update. When the reward change trend is positive and the policy convergence degree is high, it indicates that it converges in a positive direction, and the collaborative learning rate remains at a high level, with a typical value of 0.01 to 0.03; when the reward change trend is close to zero and the policy convergence degree is high, it indicates that it has approached the optimal solution, and the collaborative learning rate is reduced to a medium level, with a typical value of 0.005 to 0.01; when the reward change trend is negative or the policy convergence degree is low, it indicates that it may fall into a local optimum or an unstable state, and the collaborative learning rate is further reduced, with a typical value of 0.001 to 0.003. The collaborative learning rate is obtained by multiplying the initial learning rate by the exponential decay function of the absolute value of the reward change trend and the policy convergence degree. The initial learning rate is set to 0.05 and gradually decreases as the training progresses.
[0069] The decision strategies of the charging agent and the discharging agent are jointly optimized and updated according to the collaborative learning rate. The update process uses the batch gradient descent method, and each batch contains 32 state-action-reward samples. The parameter update step is calculated based on the collaborative learning rate, and the network parameters of the agents are updated in the direction of the policy gradient. For the charging agent, the focus is on optimizing the policy network parameters in the hierarchical decision network; for the discharging agent, the focus is on optimizing the parameters of the temperature-sensitive power modulator. The parameter update processes of the two agents are interrelated, and the collaborative optimization is realized through the shared reward function. In actual operation, parameter update is performed every 100 decision cycles, and the update frequency gradually decreases as the training progresses, eventually stabilizing at one update every 500 decision cycles.
[0070] The agent cooperative optimization mechanism in the embodiment realizes the cooperative decision of the charging agent and the discharging agent through the shared reward function and the state sharing mechanism. The adaptive cooperative mechanism can dynamically adjust the optimization target weight and the learning rate according to the battery state. The cooperative optimization method overcomes the limitation of the independent optimization of the traditional charging and discharging control strategy, realizes the global optimal control, and significantly improves the charging and discharging efficiency and the battery life.
[0071] In an optional embodiment, the adaptive cooperative mechanism calculates the dynamic weight coefficient through a state evaluation function and a strategy evaluation function, and the step of dynamically adjusting the cooperative learning rate based on the reward change trend and the strategy convergence degree comprises:
[0072] The state evaluation function calculates the state evaluation value based on the state value deviation and the state transition probability. The state value deviation is obtained by the weighted sum of the difference between the state variable evaluation values of adjacent time points and the state variable importance weight. The state transition probability is obtained by the product of the conditional probability of the state variable at adjacent time points. The strategy evaluation function calculates the strategy evaluation value based on the strategy gradient and the strategy entropy. The strategy gradient is calculated by the product of the logarithm of the parameterized strategy and the cumulative reward. The strategy entropy is calculated by the information entropy of the strategy distribution. The dynamic weight coefficient is obtained by the weighted combination of the state value deviation, the state transition probability, the strategy gradient and the strategy entropy through the sigmoid function mapping. The reward change trend and the strategy convergence degree are calculated based on the shared reward function. The reward change trend is calculated by the change rate of the shared reward function within the time window. The strategy convergence degree is calculated by the average value of the difference between the strategy parameters at adjacent time points.
[0073] The cooperative learning rate is obtained by the product of the initial learning rate and the exponential decay function of the absolute value of the reward change trend and the strategy convergence degree.
[0074] An exemplary state evaluation function is constructed to evaluate the impact of battery state changes on decision making. The state evaluation function calculates a state evaluation value based on a state value deviation and a state transition probability. The state value deviation is obtained by a weighted sum of the difference between the state variable evaluation values at adjacent time instants and the state variable importance weights. The state variables of the battery are classified into three importance levels: high importance variables include battery voltage, current, and temperature, with an importance weight of 0.25; medium importance variables include state of charge, polarization voltage, and internal resistance, with an importance weight of 0.15; and low importance variables include ambient temperature and historical charge-discharge patterns, with an importance weight of 0.05. For a specific case, the battery voltage evaluation values at two adjacent time instants are 3.8 volts and 3.75 volts, with a difference of 0.05 volts, which is multiplied by the weight 0.25 to obtain a weighted difference of 0.0125; the battery temperature evaluation values are 32 degrees Celsius and 34 degrees Celsius, with a difference of 2 degrees Celsius, which is multiplied by the weight 0.25 to obtain a weighted difference of 0.5; and the state of charge evaluation values are 65% and 64%, with a difference of 1%, which is multiplied by the weight 0.15 to obtain a weighted difference of 0.15. The weighted differences of the state variables are summed to obtain the state value deviation, which is 0.6625 in this case.
[0075] The state transition probability is obtained by multiplying the conditional probabilities of the state variables at adjacent time instants. A state transition probability matrix is constructed based on historical data to record the transition frequencies between different states. For the battery voltage, the transition probability from 3.8 volts to 3.75 volts is 0.85; for the battery temperature, the transition probability from 32 degrees Celsius to 34 degrees Celsius is 0.7; and for the state of charge, the transition probability from 65% to 64% is 0.95. These transition probabilities are multiplied to obtain a state transition probability of 0.5653. The state evaluation value is calculated by a weighted combination of the state value deviation and the state transition probability, with weights of 0.6 and 0.4, respectively. In this case, the state evaluation value is 0.6625 x 0.6 + 0.5653 x 0.4 = 0.6236.
[0076] A strategy evaluation function is constructed to evaluate the optimization direction and exploration degree of the current strategy. The strategy evaluation function calculates the strategy evaluation value based on the strategy gradient and the strategy entropy. The strategy gradient is calculated by the product of the logarithm of the parameterized strategy and the cumulative reward. The strategy of the charging agent and the discharging agent is represented as a parameterized probability distribution. For the charging agent, the strategy parameters include the selection probabilities of 10 charging power levels. For the discharging agent, the strategy parameters include the selection probabilities of 15 discharging power levels. The logarithm values of these parameters are calculated and multiplied by the cumulative reward. The cumulative reward is the sum of the rewards from the current time to the future multiple time steps, and the discount factor is set to 0.95, i.e. the future rewards decrease by the power of 0.95. In actual calculation, the probability of the charging agent selecting power level 7 is 0.3, and the logarithm value is -1.2. The cumulative reward of the recent 5 time steps is 4.2. The product of the two is -5.04, which is the strategy gradient component of power level 7. Similarly, the strategy gradient components of all power levels are calculated and summed to obtain the strategy gradient.
[0077] The strategy entropy is calculated by the information entropy of the strategy distribution, which measures the uncertainty and exploration degree of the strategy. For the charging agent, the selection probabilities of the 10 charging power levels are [0.05, 0.1, 0.1, 0.15, 0.1, 0.05, 0.3, 0.05, 0.05, 0.05]. The logarithm values of each probability are calculated and multiplied by the probability itself, and the sum is the strategy entropy. In this example, the strategy entropy is 2.05, indicating a moderate degree of strategy exploration. The strategy evaluation value is calculated by the weighted combination of the normalized value of the strategy gradient and the normalized value of the strategy entropy, with weights of 0.7 and 0.3 respectively. After min-max normalization, the value of the strategy gradient is 0.65, and the value of the strategy entropy is 0.75 after normalization. The combination of the two gives a strategy evaluation value of 0.65x0.7+0.75x0.3=0.68.
[0078] The dynamic weight coefficient is obtained by the weighted combination of the state value deviation, the state transition probability, the strategy gradient and the strategy entropy, and is mapped by the sigmoid function. The weights of the four factors are set to 0.3, 0.2, 0.3 and 0.2 respectively, and the weighted sum is mapped to the range of 0 to 1 by the sigmoid function. In the aforementioned example, the normalized values of the four factors are 0.6625, 0.5653, 0.65 and 0.75 respectively, and the weighted sum is 0.6447, which is mapped to the dynamic weight coefficient 0.656 by the sigmoid function. The weights of the four factors are dynamically adjusted according to the battery state of health and the working environment. When the battery state of health is good (capacity retention rate greater than 90%), the weights of the strategy gradient and the strategy entropy are increased, and more attention is paid to strategy optimization; when the battery state of health decreases (capacity retention rate less than 80%), the weights of the state value deviation and the state transition probability are increased, and more attention is paid to state stability.
[0079] The shared reward function is obtained by dynamically weighting the battery life extension indicator and the charging and discharging efficiency indicator. This shared reward function will be the common optimization goal of the charging agent and the discharging agent, guiding the two agents to make collaborative decisions.
[0080] The reward change trend and the policy convergence degree are calculated based on the shared reward function. The reward change trend is calculated by the change rate of the shared reward function within a time window. The size of the time window is set to 100 decision cycles, and the average change rate of the reward function value within the window is calculated. For example, the reward function value at the start of the window is 0.75, and the reward function value at the end of the window is 0.82, and the time window spans 100 decision cycles, then the reward change trend is (0.82-0.75) / 100=0.0007, indicating that the reward function is slowly increasing. The policy convergence degree is calculated by the average of the difference between the policy parameters at adjacent times. The Euclidean distance of the charging agent and the discharging agent policy parameters is calculated, for example, the average difference between the policy parameters of the charging agent at two times is 0.01, and the average difference between the policy parameters of the discharging agent is 0.008, then the policy convergence degree is 1-(0.01+0.008) / 2=0.991, indicating that the policy has approached the convergence state. The collaborative learning rate is dynamically adjusted according to the reward change trend and the policy convergence degree. The collaborative learning rate is obtained by multiplying the initial learning rate by the absolute value of the reward change trend and the exponential decay function of the policy convergence degree. The initial learning rate is set to 0.05, the reward change trend absolute value influence factor is 10, and the policy convergence degree influence factor is 5. For the reward change trend 0.0007 and the policy convergence degree 0.991, the exponential decay function value is e^(-10×0.0007-5×0.991)=0.0072. The collaborative learning rate is the product of the initial learning rate and the exponential decay function value, i.e. 0.05×0.0072=0.00036. The influence factors are dynamically adjusted according to the battery state and the working environment. When the battery is in the rapid decay stage (capacity decay rate greater than 0.03% / cycle), the reward change trend influence factor is reduced to 5, allowing more aggressive learning; when the environmental temperature fluctuates greatly (temperature change more than 5 degrees Celsius / hour), the policy convergence degree influence factor is reduced to 3, allowing the policy to adapt more flexibly to temperature changes.
[0081] The adaptive collaboration mechanism in this embodiment accurately calculates the dynamic weight coefficient through the state evaluation function and the policy evaluation function, achieving the optimal balance between battery life extension and charging and discharging efficiency. The comprehensive consideration of the reward change trend and the policy convergence degree enables the collaborative learning rate to be adaptively adjusted according to the training state, accelerating policy convergence and avoiding overfitting. This fine collaborative mechanism effectively solves the problem of balancing life and efficiency in traditional methods, providing a reliable decision-making framework for intelligent battery management.
[0082] In an alternative embodiment,
[0083] The step of training the reinforcement learning framework through the agent cooperative optimization mechanism to obtain the optimal charging strategy and the optimal discharging strategy comprises:
[0084] An experience replay pool is constructed, the experience replay pool comprising a priority experience replay unit and a shared experience replay unit, the priority experience replay unit performing importance sorting on training samples based on a time difference error, and the shared experience replay unit storing common experience data of the charging agent and the discharging agent;
[0085] A distributed training framework is constructed based on the experience replay pool, the distributed training framework comprising a parameter server and a plurality of training agents, the parameter server maintaining global policy parameters of the charging agent and the discharging agent, and the training agents calculating policy gradients in parallel based on local samples;
[0086] A training evaluation index is constructed, the training evaluation index comprising a policy stability index and a training convergence index, the policy stability index being calculated by a policy output variance, and the training convergence index being calculated by a value function estimation error;
[0087] The training evaluation index and the dynamic weight coefficient cooperate, the update frequency of the dynamic weight coefficient being adjusted through the training evaluation index, and the sampling strategy of training samples being guided through the dynamic weight coefficient;
[0088] The training evaluation index is fed back to the distributed training framework in real time, the global parameter synchronization frequency of the parameter server being dynamically adjusted based on the policy stability index, and the computing resources and gradient update weights of the training agents being adaptively allocated based on the training convergence index;
[0089] When the policy stability index is less than a first preset threshold and the training convergence index is less than a second preset threshold, the optimal charging strategy of the charging agent and the optimal discharging strategy of the discharging agent are obtained.
[0090] An exemplary reinforcement learning framework is trained by an agent cooperative optimization mechanism to obtain optimal charging strategies and optimal discharging strategies. An experience replay pool is constructed to store state-action-reward-next state data samples during the training process. The experience replay pool includes two core units: a priority experience replay unit and a shared experience replay unit. The priority experience replay unit sorts the training samples based on the temporal difference error. The temporal difference error represents the gap between the actual reward and the expected reward, which is calculated as the immediate reward plus the discounted value estimate of the next state minus the value estimate of the current state. The absolute value of the temporal difference error is calculated for each sample in the experience pool, and the larger the absolute value, the more unexpected and valuable the information contained in the sample. For example, during a charging process, the battery temperature is predicted to rise by 2 degrees Celsius at a certain charging power, but it actually rises by 5 degrees Celsius. The temporal difference error of this sample with large prediction error is 0.15, while the temporal difference error of the sample with accurate prediction is only 0.02. According to the temporal difference error, a sampling probability is assigned to each sample, and the sampling probability of the sample with an error of 0.15 is 0.08, while the sampling probability of the sample with an error of 0.02 is 0.01. The capacity of the priority experience replay unit is set to 10,000 samples, and a ring buffer structure is used. When the buffer is full, the newest sample will overwrite the oldest sample.
[0091] The shared experience replay unit stores common experience data of the charging agent and the discharging agent, including state transitions and reward information generated by the interaction of the two agents. These shared experiences focus on charging and discharging state transitions, such as from charging to discharging, from discharging to charging, and charging and discharging in parallel. The shared experience data structure includes: shared state representation (128-dimensional vector), charging action (10-dimensional vector), discharging action (15-dimensional vector), shared reward value (scalar), next shared state (128-dimensional vector), and state transition flag (scalar). An importance weight is calculated for each shared experience, and the weight is calculated based on state rarity and reward abnormality. The state rarity is calculated by the K-nearest neighbor density estimation method. For states with a frequency less than 0.1% in the experience pool, the rarity is set to 0.9; for states with a frequency between 0.1% and 1%, the rarity is set to 0.6; for common states (with a frequency greater than 1%), the rarity is set to 0.3. The reward abnormality is calculated by the deviation of the reward value from the average reward of the same state. The abnormality of samples with a deviation greater than two standard deviations is set to 0.8, the abnormality of samples with a deviation between one and two standard deviations is set to 0.5, and the abnormality of samples with a deviation less than one standard deviation is set to 0.2. The capacity of the shared experience replay unit is set to 5,000 samples, and the update strategy is to retain the samples with the highest importance weight.
[0092] A distributed training framework is constructed based on the experience replay pool, including a parameter server and multiple training agents. The parameter server maintains the global policy parameters of the charging agent and the discharging agent, including the hierarchical decision network parameters of the charging agent and the temperature-sensitive power modulator parameters of the discharging agent. These parameters are stored in the form of key-value pairs, with the key being the parameter name and the value being the parameter value. The policy network of the charging agent contains about 50,000 parameters, and the evaluation network contains about 30,000 parameters; the power trajectory generator of the discharging agent contains about 40,000 parameters, and the current shaping controller contains about 20,000 parameters. The parameter server is responsible for the storage, update and synchronization of global parameters, and adopts an asynchronous parameter update mechanism. When receiving the gradient update request sent by the training agent, the parameter value is updated according to the gradient information and the global learning rate. The global learning rate is initially set to 0.01 and is reduced by an exponential decay rule every 10,000 training steps, to 0.95 times of the original value.
[0093] The number of training agents is set to 8, each agent independently samples training data from the experience replay pool, performs forward calculation and back propagation, and calculates the policy gradient. The calculation of the policy gradient is based on the partial derivative of the policy objective function with respect to the parameters, which is estimated by sampling method. Each training agent samples 256 data points from the experience replay pool each time, of which 70% comes from the priority experience replay unit and 30% comes from the shared experience replay unit. The training agent performs 5 epochs of training on the sampled data, calculates the parameter gradient, and sends it to the parameter server. The training agents work completely in parallel, and the diversity of the sampled data is ensured by different initialization seeds. To handle the data distribution differences between different agents, an importance weight-based gradient merging mechanism is implemented, with the weight being inversely proportional to the difference between the state distribution observed by the agent and the global state distribution. In actual operation, the gradient weight of the agent that observes more rare states is set to 1.2, and the gradient weight of the agent that observes more common states is set to 0.8.
[0094] The training evaluation indicators include a strategy stability indicator and a training convergence indicator. The strategy stability indicator is calculated by the strategy output variance, reflecting the consistency of the strategy under similar states. The variance of the strategy output in the last 100 decision cycles is calculated to obtain the charging strategy output variance and the discharging strategy output variance. For the charging agent, the average of the standard deviation of the selection probability of the 10 charging power levels in the last 100 decision cycles is taken as the strategy output variance; for the discharging agent, the average of the standard deviation of the selection probability of the 15 discharging power levels is taken as the strategy output variance. In the early stage of training, the strategy output variance is usually between 0.15 and 0.25; as the training progresses, the strategy gradually stabilizes, and the variance decreases to between 0.05 and 0.1; when the training is close to convergence, the variance further decreases to between 0.01 and 0.03. Set the first preset threshold value to 0.03, when the strategy stability indicator is less than this threshold value, it is considered that the strategy has reached a stable state.
[0095] The training convergence indicator is calculated by the value function estimation error, reflecting the accuracy of the value function prediction. The value function estimation error is defined as the squared difference between the predicted state value and the actual cumulative reward. The average of the value prediction error in the last 500 decision cycles is taken as the training convergence indicator. In the early stage of training, the value function estimation error is usually between 0.2 and 0.3; as the training progresses, the estimation error gradually decreases to between 0.1 and 0.15; when the training is close to convergence, the estimation error further decreases to below 0.05. Set the second preset threshold value to 0.05, when the training convergence indicator is less than this threshold value, it is considered that the value function prediction has reached a convergent state.
[0096] The training evaluation indicators and the dynamic weight coefficient work together to adjust the update frequency of the dynamic weight coefficient through the training evaluation indicators, and guide the sampling strategy of the training samples through the dynamic weight coefficient. When the strategy stability indicator is higher than 0.1, it indicates that the strategy is not stable, and the update frequency of the dynamic weight coefficient is set to every 50 decision cycles; when the strategy stability indicator is between 0.03 and 0.1, the update frequency is reduced to every 100 decision cycles; when the strategy stability indicator is less than 0.03, the update frequency is further reduced to every 200 decision cycles, to avoid excessive adjustment. The dynamic weight coefficient in turn affects the sampling strategy of the training samples, when the weight coefficient is greater than 0.7, the samples related to battery life are preferentially sampled from the priority experience replay unit; when the weight coefficient is less than 0.3, the samples related to charging and discharging efficiency are preferentially sampled; when the weight coefficient is between 0.3 and 0.7, a balanced sampling strategy is adopted. This two-way adjustment mechanism ensures that the direction of strategy optimization in the training process is consistent with the current demand.
[0097] The training evaluation index is fed back to the distributed training framework in real time, the global parameter synchronization frequency of the parameter server is dynamically adjusted based on the policy stability index, and the computing resources and gradient update weights of the training agent are adaptively allocated based on the training convergence index. When the policy stability index is higher than 0.1, the global parameter synchronization frequency is set to perform global synchronization once every 10 gradient updates; when the index is between 0.03 and 0.1, the synchronization frequency is reduced to every 20 gradient updates; when the index is lower than 0.03, the synchronization frequency is further reduced to every 50 gradient updates, reducing the communication overhead. For the training convergence index, when the index is higher than 0.15, equal computing resources and gradient update weights are allocated to each training agent; when the index is between 0.05 and 0.15, the resource allocation is adjusted according to the state distribution difference observed by each agent, and the agent that observes more rare states obtains more resources; when the index is lower than 0.05, the weight of the agent that observes the key state transition is further increased to ensure that the key samples are fully learned.
[0098] The policy stability index and the training convergence index are continuously monitored, and when the policy stability index is less than a first preset threshold 0.03 and the training convergence index is less than a second preset threshold 0.05, it is judged that the training has reached a converged state. At this time, the optimal charging strategy of the charging agent and the optimal discharging strategy of the discharging agent are obtained from the parameter server. The optimal charging strategy includes the final parameters of the hierarchical decision network, which is used to guide the charging power control under different battery states and environmental conditions; the optimal discharging strategy includes the final parameters of the temperature-sensitive power modulator, which is used to guide the discharging power control under different demand and temperature conditions. These strategy parameters are encapsulated as a strategy model and deployed to the battery management for real-time control. To ensure the robustness of the model, an additional 1000 decision cycles of verification test are performed after the training is completed, and if both indexes remain below the threshold during the test, the final confirmation is made that the training is successfully completed.
[0099] The distributed training framework and the dual evaluation index mechanism in the embodiment realize efficient and stable reinforcement learning training. The priority sampling of the experience replay pool and the shared experience storage improve the sample utilization efficiency, the distributed architecture accelerates the training process, and the synergistic effect of the evaluation index and the dynamic weight coefficient ensures that the training direction and goal are consistent.
[0100] In this embodiment, to improve training efficiency and policy quality, the charging agent and the discharging agent are trained separately before the collaborative optimization of the agents. During the separate training of the charging agent, a charging-specific reward function is used to evaluate the pros and cons of the charging strategy. The charging-specific reward function consists of three core indicators: charging completion time indicator, charging uniformity indicator, and temperature change indicator caused by charging. The charging completion time indicator represents the time required to charge the battery from the initial state of charge to the target state of charge, compared with the standard charging time. For example, the standard time to charge the battery from 20% state of charge to 80% is 120 minutes, if the actual completion time is 100 minutes, the time indicator is 1.2; if the completion time is 150 minutes, the time indicator is 0.8. The charging uniformity indicator measures the smoothness of the charging current, which is calculated by the standard deviation of the current at 50 consecutive time points. The uniformity indicator is 0.9 when the standard deviation is less than 0.05C, the indicator is 0.7 when the standard deviation is between 0.05C and 0.1C, and the indicator is 0.5 when the standard deviation is greater than 0.1C. The temperature change indicator caused by charging reflects the degree of temperature rise during the charging process, which is represented by the ratio of temperature rise to safety threshold. If the temperature rise during charging is 5 degrees Celsius and the safety threshold is 15 degrees Celsius, the temperature change indicator is 0.67.
[0101] The three indicators are combined using an adaptive weight factor. When the battery health state is good (capacity retention rate is greater than 90%), the charging completion time indicator weight is 0.5, the charging uniformity indicator weight is 0.2, and the temperature change indicator weight is 0.3; in the middle stage of battery use (capacity retention rate is 80% to 90%), the weights of the three indicators are adjusted to 0.4, 0.3 and 0.3 respectively; in the later stage of battery use (capacity retention rate is less than 80%), the weights are further adjusted to 0.3, 0.3 and 0.4, and more attention is paid to temperature control to extend the remaining life.
[0102] In the individual training process of the discharging agent, a discharge-specific reward function is used, including a discharge duration indicator, a discharge temperature balance indicator, and a power output stability indicator. The discharge duration indicator represents the ratio of the time required for the battery to discharge from the initial state of charge to the cutoff state of charge at a specified discharge power to the theoretical time. For example, the theoretical time for discharging from 80% to 20% at a 1C discharge rate is 36 minutes, and if the actual duration is 33 minutes, the duration indicator is 0.92. The discharge temperature balance indicator measures the uniformity of the temperature of each part of the battery during discharging, calculated by measuring the maximum temperature difference between points. The balance indicator is 0.9 when the temperature difference is less than 2 degrees Celsius, 0.7 when the temperature difference is between 2 and 5 degrees Celsius, and 0.5 when the temperature difference is greater than 5 degrees Celsius. The power output stability indicator reflects the fluctuation of the discharge power, calculated by the ratio of the standard deviation of the power at 100 consecutive time points to the average power. The stability indicator is 0.9 when the ratio is less than 3%, 0.7 when the ratio is between 3% and 8%, and 0.5 when the ratio is greater than 8%.
[0103] The discharge indicators are combined using dynamic weight factors. In high-power demand scenarios (discharge power greater than 1.5C), the discharge duration indicator weight is 0.5, the temperature balance indicator weight is 0.3, and the power stability indicator weight is 0.2; in medium-power demand scenarios (discharge power between 0.5C and 1.5C), the weights of the three indicators are 0.4, 0.4, and 0.2, respectively; in low-power long-endurance scenarios (discharge power less than 0.5C), the weights are adjusted to 0.3, 0.4, and 0.3, with more emphasis on temperature balance and power stability.
[0104] The charging agent and the discharging agent are pre-trained using the SAC algorithm based on maximum entropy. The SAC algorithm includes a policy network and a double Q value network structure. The policy network is a three-layer fully connected network with a structure of 64-128-256-output dimension. The output dimension is 10 for the charging agent and 15 for the discharging agent, representing the selection probability of different power levels. The double Q value network contains two networks with the same structure but independent parameters, with a structure of (64+action dimension)-128-256-1. The input is the concatenation of the state vector and the action vector, and the output is the state-action value estimate. The reparameterization trick is used to sample actions, and an entropy regularization term is introduced to encourage policy exploration. The initial value of the entropy regularization coefficient is set to 0.1, and it is dynamically adjusted during training. When the policy entropy is lower than the target entropy 0.5, the regularization coefficient is increased, and when the policy entropy is higher than the target entropy, the regularization coefficient is decreased.
[0105] During the training process, the weight coefficients of each indicator are dynamically adjusted according to the battery operating environment temperature and the current health state. When the environment temperature is lower than 10 degrees Celsius, the weight of the charging temperature change indicator increases by 0.1, and the weight of the discharging temperature balance indicator increases by 0.1; when the environment temperature is higher than 35 degrees Celsius, the weight of the two temperature-related indicators increases by 0.2. When the battery health state is lower than 75%, the weight of the charging uniformity indicator increases by 0.1, and the weight of the power output stability indicator increases by 0.1, to reduce the stress on the aging battery. The batch size for separate training is 256, the learning rate is 0.0003, the Adam optimizer is used, and the training lasts for 50000 environment interaction steps.
[0106] After completing the separate training, the parameters obtained by training are migrated to the collaborative training framework. The policy network parameters of the charging agent are migrated to the online policy network in the hierarchical decision network of the charging agent as the initial parameters; the double Q value network parameters of the charging agent are used to initialize the capacity decay evaluation module, the temperature stress evaluation module and the polarization stress evaluation module in the evaluation network. For the discharging agent, its policy network parameters are used to initialize the power trajectory generator in the temperature-sensitive power modulator, and the value network parameters are used to initialize the current shaping controller. This parameter migration method provides a good initial point for collaborative training and accelerates the convergence process of collaborative optimization.
[0107] In this embodiment, the most suitable application mode is selected in real time according to the current charging and discharging state of the battery. The battery management monitors the charging and discharging state of the battery through the current sensor, and judges the working mode of the battery with 0.05C current as the threshold. When the current value is greater than 0.05C and the current direction is flowing into the battery, it is determined as single charging state; when the current value is greater than 0.05C and the current direction is flowing out of the battery, it is determined as single discharging state; when there are both inflow and outflow currents, or part of the single cells in the battery pack are charging while part of the single cells are discharging, it is determined as charging and discharging in parallel state. In actual application, the electric tool is usually in single charging state when charging during the working gap, and the current value is stable between 0.5C and 2C; the device is in single discharging state when running normally, and the current value is between 0.2C and 3C; the electric vehicle is in charging and discharging in parallel state when energy is recovered during braking and power is supplied, and the charging current is usually between 0.1C and 0.5C, and the discharging current is between 0.3C and 1.5C.
[0108] When the battery is detected to be in a single charging state, the charging agent strategy is mainly applied. The hierarchical decision network is activated to receive the state representation and generate a charging power distribution through the online policy network. After evaluation, the network corrects the output and outputs the final charging power control signal. For example, when the battery temperature is 28 degrees Celsius, the state of charge is 35%, and the state of health is 92%, the charging agent outputs a power of 1.1 times the nominal power; when the temperature rises to 38 degrees Celsius, the power is automatically reduced to 0.8 times the nominal power; when the state of charge reaches 85%, the power is further reduced to 0.4 times the nominal power, realizing multi-stage intelligent charging. When the battery is in a single discharging state, the discharging agent strategy is mainly applied. The temperature-sensitive power modulator generates a power trajectory based on the current discharging demand and the battery state, and realizes precise control through the current shaping controller. For example, the initial discharging power is set to 95% of the demand power, and when the temperature distribution is uniform, it gradually increases to 100%; when local temperature is detected to be too high, the current pulse width of the corresponding area is automatically reduced to achieve temperature balancing control.
[0109] When the battery is detected to be in a single charging state, the charging agent strategy is mainly applied. The hierarchical decision network is activated to receive the state representation and generate a charging power distribution through the online policy network. After evaluation, the network corrects the output and outputs the final charging power control signal. For example, when the battery temperature is 28 degrees Celsius, the state of charge is 35%, and the state of health is 92%, the charging agent outputs a power of 1.1 times the nominal power; when the temperature rises to 38 degrees Celsius, the power is automatically reduced to 0.8 times the nominal power; when the state of charge reaches 85%, the power is further reduced to 0.4 times the nominal power, realizing multi-stage intelligent charging. When the battery is in a single discharging state, the discharging agent strategy is mainly applied. The temperature-sensitive power modulator generates a power trajectory based on the current discharging demand and the battery state, and realizes precise control through the current shaping controller. For example, the initial discharging power is set to 95% of the demand power, and when the temperature distribution is uniform, it gradually increases to 100%; when local temperature is detected to be too high, the current pulse width of the corresponding area is automatically reduced to achieve temperature balancing control.
[0110] In the transition period of strategy switching, the outputs of the strategies before and after switching are interacted through the state sharing mechanism. The transition period is defined as a 5-second time window after state switching. Based on the shared state information, the strategy fusion weight is calculated, and a time decay function is used to smooth the transition. At the initial switching time, the weight of the previous strategy is 0.8 and the weight of the new strategy is 0.2; after 1 second, the weights are 0.6 and 0.4 respectively; after 3 seconds, the weights are 0.3 and 0.7 respectively; after 5 seconds, the new strategy is completely switched. For example, when switching from single discharging to charging and discharging in parallel, the discharging agent strategy weight is 0.8 and the cooperative strategy weight is 0.2 in the first second; the discharging power gradually decreases from 1.2C to 0.3C, and the charging power gradually increases from 0 to 0.3C. The outputs of the individual strategies and the cooperative strategy are smoothly fused through the adaptive weight coefficient, and the calculation method is the sum of the products of the outputs of each strategy and the corresponding weight. When the battery temperature change rate is greater than 1 degree Celsius per second, the transition period is extended to 10 seconds to further smooth the power change curve and prevent temperature fluctuations caused by strategy switching.
[0111] The strategy selection criteria are adjusted online based on actual operation effects. The battery management records the performance of different strategies under various working conditions, and evaluates the effects of different strategy combinations through a joint value function. The joint value function considers four key indicators: battery state of health changes, energy efficiency, temperature uniformity, and power stability. Strategy evaluation is performed every 100 hours of operation, and the average value score of different strategies under each working condition is calculated. For example, if the average value score of the charging agent strategy is 0.85 and the collaborative strategy score is 0.82 under a single charging state, the original strategy selection is maintained. If the collaborative strategy score exceeds the individual strategy under certain conditions (such as environmental temperature above 35 degrees Celsius and battery state of health below 80%), the strategy selection boundary is dynamically adjusted, and the collaborative strategy is preferred even under a single charging state.
[0112] The application scenarios of individual strategies and collaborative strategies are dynamically optimized, and a decision tree structure is constructed to record the optimal strategy selection rules. The decision tree contains four key features: battery state (charging / discharging / parallel), environmental temperature (low / medium / high), state of health (excellent / good / medium / poor), and load characteristics (stable / variable), corresponding to the optimal strategy selection of different feature combinations. The evaluation results are fed back to the state sharing mechanism to update the shared state information. For example, if it is found through evaluation that the collaborative strategy significantly outperforms the individual strategy combination under low-temperature environment (below 10 degrees Celsius) in parallel charging and discharging state, this information is added to the shared state information, increasing the selection probability of the collaborative strategy in similar scenarios, and adjusting the weight of temperature-related parameters in the collaborative strategy to enhance low-temperature adaptability.
[0113] In a second aspect, a cylindrical lithium battery charging and discharging strategy optimization system based on reinforcement learning is provided, comprising:
[0114] A first unit is configured to evaluate the remaining life and performance degradation trend of the battery based on real-time monitoring data and historical monitoring data of the cylindrical lithium battery, and generate a prediction result; and determine the current charging and discharging state based on the real-time monitoring data.
[0115] A second unit is configured to establish a reinforcement learning framework, wherein the reinforcement learning framework includes a charging agent and a discharging agent, and the state space of the charging agent and the discharging agent includes the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints.
[0116] a third unit configured to construct an agent collaborative optimization mechanism, the agent collaborative optimization mechanism comprising a shared reward function and a state sharing mechanism, the shared reward function combining a battery life extension indicator and a charging and discharging efficiency indicator as a common optimization target of the charging agent and the discharging agent, and the state sharing mechanism enabling a decision result of the charging agent as a state input of the discharging agent and enabling a decision result of the discharging agent as a state input of the charging agent;
[0117] a fourth unit configured to train the reinforcement learning framework through the agent collaborative optimization mechanism to obtain an optimal charging strategy and an optimal discharging strategy.
[0118] In a third aspect, a computer readable storage medium is provided, and the computer readable storage medium has stored thereon computer program instructions. The computer program instructions are executed by a processor to implement the method described above.
Claims
1. A method for optimizing the charging and discharging strategy of cylindrical lithium batteries based on reinforcement learning, characterized in that, The application relates to a method for optimizing the charging and discharging strategies of a cylindrical lithium battery. The method comprises the following steps: According to real-time monitoring data and historical monitoring data of the cylindrical lithium battery, the remaining life and performance degradation trend of the battery are evaluated, and a prediction result is generated. The current charging and discharging state is determined according to the real-time monitoring data. An enhanced learning framework is established, which comprises a charging agent and a discharging agent, and the state space of the charging agent and the discharging agent comprises the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints. An agent collaborative optimization mechanism is constructed, which comprises a shared reward function and a state sharing mechanism, the shared reward function combines the battery life extension index and the charging and discharging efficiency index into the common optimization goal of the charging agent and the discharging agent, and the state sharing mechanism makes the decision result of the charging agent as the state input of the discharging agent, and makes the decision result of the discharging agent as the state input of the charging agent.
2. The method of claim 1, wherein, The enhanced learning framework is trained through the agent collaborative optimization mechanism to obtain optimal charging and discharging strategies. The step of evaluating the remaining life and performance degradation trend of the battery according to the real-time monitoring data and the historical monitoring data of the cylindrical lithium battery to generate a prediction result comprises the following steps: The real-time monitoring data and the historical monitoring data both comprise voltage data, current data, temperature data and state of charge data of the cylindrical lithium battery. Based on the real-time monitoring data and the historical monitoring data, an instantaneous feature vector, a short-term feature vector and a long-term feature vector are respectively constructed, and the long-term feature vector is obtained by weighted cumulative calculation of the relative capacity and the relative internal resistance of the lithium battery through a time decay factor. The contribution scores of the instantaneous feature vector, the short-term feature vector and the long-term feature vector to the prediction result are calculated, the corresponding feature weights are determined according to the contribution scores, the feature weights are multiplied by the corresponding feature vectors respectively and are superimposed to obtain a multi-scale feature vector.
3. The method of claim 1, wherein, The multi-scale feature vector is input into a long short-term memory network, the hidden state and the input feature of the long short-term memory network are weighted through a time attention mechanism and a feature attention mechanism to obtain a remaining life prediction value and a performance degradation rate; the performance degradation trend of the lithium battery is predicted according to the performance degradation rate, and the remaining life prediction value and the performance degradation trend are adaptively corrected according to the historical prediction error to obtain the prediction result. The step of establishing an enhanced learning framework, which comprises a charging agent and a discharging agent, and the state space of the charging agent and the discharging agent comprises the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints comprises the following steps: The prediction result and the current charging and discharging state are time window sliding encoded to obtain a time sequence feature sequence, a self-attention mechanism is applied to the time sequence feature sequence to calculate the correlation weight between different time points and different state variables, and a reduced state representation is generated through weighted fusion. The adaptive safety boundary is established by calculating the internal polarization voltage and temperature distribution of the battery based on an electrochemical reduced-order model and combining the performance degradation trend, and a charging power range within the adaptive safety boundary is set as an action space of a charging agent and a discharging agent respectively; A hierarchical decision network is constructed in the charging agent, the hierarchical decision network including a policy network and an evaluation network, the policy network outputting a charging power distribution using a double-delay deterministic policy architecture, and the evaluation network evaluating the influence of the charging power on the battery life based on battery operating characteristics and correcting the policy output; A temperature-sensitive power modulator is constructed in the discharging agent, an optimal power trajectory is calculated based on battery equivalent circuit characteristics, and a pulse width modulation technique is used to shape the discharging current, thereby achieving temperature balancing control.
4. The method of claim 3, wherein, The step of constructing a hierarchical decision network in the charging agent includes: The policy network includes an online policy network and a target policy network, the online policy network receiving the state representation, generating mean parameters and standard deviation parameters of the charging power distribution through hidden layer state calculation, and the network parameters of the target policy network being updated by weighting the network parameters of the online policy network through a soft update coefficient; In the evaluation network, a capacity decay rate index, a temperature stress index, and a polarization stress index are calculated based on battery operating characteristics, the capacity decay rate index being calculated by capacity loss under different operating conditions, the temperature stress index being calculated by the deviation of the current temperature from the optimal operating temperature, and the polarization stress index being calculated by the cumulative effect of the polarization voltage; a strategy correction signal is generated based on the capacity decay rate index, the temperature stress index, and the polarization stress index, the strategy correction signal being obtained by weighting and combining each index through an adaptive weight coefficient; The state representation is processed through attention to obtain attention weights, the attention weights and the mean parameters and standard deviation parameters of the charging power distribution are combined to generate an initial charging power, and the initial charging power is corrected according to the strategy correction signal to obtain a charging power.
5. The method of claim 1, wherein, The step of constructing an agent cooperative optimization mechanism includes: The battery life extension index is calculated based on the capacity loss amount, the initial capacity, the operating temperature, and the optimal temperature, and the charging and discharging efficiency index is calculated based on the charging efficiency, the discharging efficiency, and the state of charge; The state sharing mechanism includes local state information and shared state information, the charging agent inputs the state of charge, the charging power, the temperature, the voltage, and the internal resistance as shared state information of the discharging agent, and the discharging agent inputs the discharging power, the discharging efficiency, and the health state as shared state information of the charging agent; An adaptive coordination mechanism is constructed, which constructs a joint value function based on the local state information and shared state information, the joint value function including a charging value evaluation, a discharging value evaluation and a shared state value evaluation, and calculates an instant reward based on the joint value function; The adaptive coordination mechanism calculates a dynamic weight coefficient through a state evaluation function and a policy evaluation function, combines the battery life extension index and the charging and discharging efficiency index by weighting through the dynamic weight coefficient to obtain a shared reward function, dynamically adjusts a coordination learning rate based on a reward change trend and a policy convergence degree, and jointly optimizes and updates the decision-making strategies of the charging agent and the discharging agent according to the coordination learning rate.
6. The method of claim 5, wherein, The step of calculating a dynamic weight coefficient through a state evaluation function and a policy evaluation function and dynamically adjusting a coordination learning rate based on a reward change trend and a policy convergence degree includes: The state evaluation function calculates a state evaluation value based on a state value deviation and a state transition probability, the state value deviation is obtained by a weighted sum of a difference value of adjacent time state variable evaluation values and a state variable importance weight, and the state transition probability is obtained by a product of conditional probabilities of state variables at adjacent times; the policy evaluation function calculates a policy evaluation value based on a policy gradient and a policy entropy, the policy gradient is calculated by a product of a parameterized policy logarithm and a cumulative reward, and the policy entropy is calculated by an information entropy of a policy distribution; The dynamic weight coefficient is obtained by a weighted combination of the state value deviation, the state transition probability, the policy gradient and the policy entropy through a sigmoid function mapping; A reward change trend and a policy convergence degree are calculated based on the shared reward function, the reward change trend is calculated by a change rate of the shared reward function within a time window, and the policy convergence degree is calculated by an average value of difference values of policy parameters at adjacent times; A coordination learning rate is dynamically adjusted according to the reward change trend and the policy convergence degree, the coordination learning rate is obtained by a product of an initial learning rate and an exponential decay function of absolute values of the reward change trend and the policy convergence degree.
7. The method of claim 5, wherein, The steps of training the reinforcement learning framework through the agent coordination optimization mechanism to obtain optimal charging strategies and optimal discharging strategies include: An experience replay pool is constructed, the experience replay pool including a priority experience replay unit and a shared experience replay unit, the priority experience replay unit performing importance sorting on training samples based on a time difference error, and the shared experience replay unit storing common experience data of the charging agent and the discharging agent; A distributed training framework is constructed based on the experience replay pool, the distributed training framework including a parameter server and multiple training agents, the parameter server maintaining global policy parameters of the charging agent and the discharging agent, and the training agents calculating policy gradients in parallel based on local samples; The training evaluation index includes a strategy stability index and a training convergence index, the strategy stability index is calculated by strategy output variance, and the training convergence index is calculated by value function estimation error; When the strategy stability index is less than a first preset threshold and the training convergence index is less than a second preset threshold, an optimal charging strategy of the charging agent and an optimal discharging strategy of the discharging agent are obtained.
8. A reinforcement learning based cylindrical lithium battery charge and discharge strategy optimization system for implementing the method of any one of the preceding claims 1-7, characterized in that, Comprise: The first unit is used for evaluating the remaining life and performance degradation trend of the battery according to the real-time monitoring data and historical monitoring data of the cylindrical lithium battery, and generating a prediction result; The current charging and discharging state is determined according to the real-time monitoring data; The second unit is used for establishing a reinforcement learning framework, the reinforcement learning framework includes a charging agent and a discharging agent, and the state space of the charging agent and the discharging agent contains the prediction result and the current charging and discharging state, and the action space is the charging and discharging power range under safety constraints; The third unit is used for constructing an agent collaborative optimization mechanism, the agent collaborative optimization mechanism includes a shared reward function and a state sharing mechanism, the shared reward function combines the battery life extension index and the charging and discharging efficiency index into a common optimization target of the charging agent and the discharging agent, and the state sharing mechanism makes the decision result of the charging agent as the state input of the discharging agent, and makes the decision result of the discharging agent as the state input of the charging agent; The fourth unit is used for training the reinforcement learning framework through the agent collaborative optimization mechanism to obtain an optimal charging strategy and an optimal discharging strategy.
9. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions are executed by the processor to realize the method in any one of claims 1 to 7. The computer program instructions are executed by the processor to realize the method in any one of claims 1 to 7.
Citation Information
Cited By
Storage battery charging and discharging parameter intelligent optimization method and system fused with reinforcement learning
CN121417456A
Battery life analysis method and system based on intelligent electric vehicle battery charging pile
CN121805858A
A battery adaptive charging strategy optimization method and system based on reinforcement learning
CN122620724A