Lithium battery charging strategy sample efficiency enhancement method and system based on model short branch deduction
By adopting the optimization framework of reinforcement learning and model short branch deduction in lithium battery charging technology, the contradiction between charging speed improvement and battery safety and durability in the existing technology is solved, and efficient sample utilization and strategy optimization are achieved.
Patent Information
- Application Number
- CN202510207496.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-25
AI Technical Summary
While improving the charging speed, existing lithium battery charging technology is difficult to ensure the safety and durability of the battery. Moreover, the optimization method based on high-precision models has error accumulation problems, which affects the optimization accuracy of the strategy.
The charging strategy optimization framework based on reinforcement learning and model short branch deduction is adopted, and strategy optimization is performed on real batteries through agents to avoid complex modeling processes. The short branch deduction is used to generate additional samples to improve sample efficiency and alleviate error accumulation problems.
The sample efficiency of the lithium battery charging strategy is significantly improved, and solutions with high reward values can be found in fewer training rounds, solving the problem of degradation of convergence accuracy caused by error accumulation in traditional methods.
Smart Images

Figure CN120142941A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of lithium - battery charging technologies, and particularly relates to a method and system for enhancing the sample efficiency of a lithium - battery charging strategy based on model short - branch deduction. Background Art
[0002] With the popularization of electric vehicles, lithium - ion batteries have become the mainstream choice due to their high energy density, long cycle life, and low maintenance cost. However, the disadvantage of electric vehicles in terms of charging speed, especially compared with traditional fuel vehicles, seriously affects the user experience and restricts the further expansion of the market. Although increasing the charging current can speed up the charging process, high current will trigger a series of adverse electrochemical reactions, resulting in battery degradation, reduced charging efficiency, and even potential safety accidents. Therefore, how to improve the charging speed while ensuring the safety and durability of the battery has become an urgent problem to be solved.
[0003] Regarding the optimization of fast - charging strategies, some progress has been made in recent years. Traditional methods are mostly based on preset rules. Although they are simple and easy to implement, they cannot fully consider the dynamic changes inside the battery and are difficult to meet the requirements of efficient charging. Optimization methods based on battery reaction mechanisms, such as equivalent - circuit models and electrochemical models, can more accurately simulate battery behavior, improve charging efficiency, and extend battery life. However, these methods often rely on high - precision models, which limits their reliability and efficiency in practical applications.
[0004] In recent years, reinforcement - learning methods, as an emerging optimization tool, have received extensive attention. Reinforcement learning can improve the performance of charging strategies without relying on battery models through autonomous learning and optimization. However, in practical applications, its sample efficiency is low, resulting in high experimental costs. To address this problem, model - based reinforcement - learning methods improve sample efficiency by introducing an environment model, but the recursive use of the model will lead to the problem of error accumulation, which in turn affects the optimization accuracy of the strategy. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention proposes a method and system for enhancing the sample efficiency of a lithium - battery charging strategy based on model short - branch deduction. It adopts a charging - strategy optimization framework that combines reinforcement learning and model - branch deduction, which can directly optimize the strategy on a real battery, avoiding complex and cumbersome modeling processes. At the same time, the limited utilization of model - deduced data effectively improves the sample efficiency and alleviates the negative impact brought by model - error accumulation.
[0006] On the one hand, an embodiment of the present invention provides a method for enhancing the sample efficiency of a lithium - battery charging strategy based on model short - branch deduction, including the following steps:
[0007] S1. Perform charge and discharge tests on the battery based on the actions taken by the reinforcement learning agent, and sample the charge state transition data;
[0008] The process of sampling the charge state transition data is as follows: Obtain the initial charge state parameters and send the initial charge state parameters to the agent; The agent takes corresponding actions based on the current battery state and simultaneously obtains the corresponding state of the battery at the next moment; The action is to adjust the charging current of the battery.
[0009] S2. Train the battery proxy model using the sampled charge state transition data; Use each state and the next state of the real sampled trajectory data as the input and output of the battery proxy model for model training to minimize the prediction error of the battery proxy model for the charge state transition of the battery.
[0010] S3. On the basis of the real charge state trajectory, perform short-branch deduction using the battery proxy model, that is, starting from a certain charge state in the real sampled trajectory, take an action different from the next action of the real charge state trajectory, and use the prediction of the battery proxy model to obtain the next charge state of the battery, that is, simulate the sampling process of the battery to obtain the branch trajectory data of the deduced charge state transition.
[0011] S4. Combine the real charge state trajectory data with the deduced branch trajectory data to train the agent; And repeat the above steps S1 - S3 until the reinforcement learning training converges to obtain a reinforcement learning agent that can execute the optimal charging strategy.
[0012] Preferably, the agent includes a policy network and a value network, and the value network is used to calculate the Q value; The goal of the policy network is to find a policy that maximizes the weighted sum of the expected cumulative reward and the policy entropy based on the Q value calculated by the value network, the current policy, the charging current, and the charge state of the battery.
[0013] On the other hand, an embodiment of the present invention also provides a lithium battery charging strategy sample efficiency enhancement system based on model short-branch deduction, which is implemented based on the above lithium battery charging strategy sample efficiency enhancement method. The system includes the following modules:
[0014] A sampling module that performs charge and discharge tests on the battery based on the actions taken by the reinforcement learning agent and samples the charge state transition data; The process of sampling the charge state transition data is as follows: Obtain the initial charge state parameters and send the initial charge state parameters to the agent; The agent takes corresponding actions based on the current battery state and simultaneously obtains the corresponding state of the battery at the next moment; The action is to adjust the charging current of the battery.
[0015] A model training module that trains a battery agent model using the sampled charge state transition data; uses each state and the next state of the real sampled trajectory data as the input and output of the battery agent model for model training to minimize the prediction error of the battery agent model for the charge state transition of the battery.
[0016] A deduction module that, based on the real charge state trajectory, uses the battery agent model for short-branch deduction, that is, takes a certain charge state in the real sampled trajectory as the starting point, takes an action different from the next action of the real charge state trajectory, and uses the prediction of the battery agent model to obtain the next charge state of the battery, that is, simulates the sampling process of the battery to obtain the branch trajectory data of the deduced charge state transition.
[0017] An agent training module that combines the real charge state trajectory data with the deduced branch trajectory data to train the agent; and repeats the sampling of the charge state transition data, the training of the battery agent model, and the deduction until the reinforcement learning training converges to obtain a reinforcement learning agent that can execute the optimal charging strategy.
[0018] Compared with the prior art, the beneficial effects achieved by the present invention include:
[0019] The present invention proposes a lithium battery charging strategy optimization framework based on reinforcement learning and model branch deduction. The model learns the charge state transition of the battery through the sampled data of the agent and uses the short-branch deduction method to generate additional samples, thereby accelerating the optimization process of the charging strategy. The present invention is significantly superior to the model-free reinforcement learning method in terms of sample efficiency and can find solutions with high reward values in fewer training rounds; the limited use of model deduction data helps the agent training and effectively solves the problem of the decrease in convergence accuracy caused by error accumulation in traditional model-based methods. Description of the Drawings
[0020] Figure 1 It is a flowchart of the method for enhancing the sample efficiency of the lithium battery charging strategy based on model short-branch deduction in the embodiment of the present invention;
[0021] Figure 2 It is a schematic diagram of the overall framework of the system for enhancing the sample efficiency of the lithium battery charging strategy based on model short-branch deduction in the embodiment of the present invention;
[0022] Figure 3 It is a schematic diagram of the process of using the model for prediction in the embodiment of the present invention;
[0023] Figure 4 It is a schematic diagram of model branch deduction in the embodiment of the present invention;
[0024] Figure 5The curve diagram of the main parameters during the training process of the strategy optimization method adopted in the embodiment of the present invention and the model-free strategy optimization method. Among them, (a) is the curve diagram of the Return value, (b) is the curve diagram of the charging time, (c) is the curve diagram of the negative electrode potential constraint violation, and (d) is the curve diagram of the temperature constraint violation;
[0025] Figure 6 The curve diagram of the best charging strategies found at the initial stage (the first 100 training rounds) and the end of training (1000 training rounds) of the strategy optimization method adopted in the embodiment of the present invention and the model-free strategy optimization method. Among them, (a) is the curve diagram of the charging current, (b) is the curve diagram of the SOC, (c) is the curve diagram of the negative electrode potential, and (d) is the curve diagram of the temperature. Detailed implementation manners
[0026] The following combines the embodiments and the accompanying drawings to describe the present invention in detail. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention. The orientation terms such as left, middle, right, up, and down in the embodiments of the present invention are only relative concepts to each other or are referenced based on the normal use state of the product, and should not be considered restrictive.
[0027] A method for enhancing the sample efficiency of a lithium battery charging strategy based on model short-branch deduction in an embodiment of the present invention, as Figure 1 and Figure 2 shown, includes the following steps:
[0028] Step 1: Perform charge and discharge tests on the battery based on the actions taken by the reinforcement learning agent, and sample the charging state transition data.
[0029] In this embodiment, the battery is discharged until the state of charge SOC is 0, and then the battery is left standing in a room temperature (25 °C) environment until the battery temperature is the same as the ambient temperature, and then subsequent experiments are carried out.
[0030] The sampling process of the charging state transition data is as follows: Obtain the initial charging state parameters and send the initial charging state parameters into the agent. The agent includes a policy network and a value network; the agent takes corresponding actions according to the current battery state (for example, adjust the charging current of the battery, and set the value range of the charging current to 1C to 4C), and at the same time obtain the corresponding state of the battery at the next moment.
[0031] The obtained charging state transition data will be stored in the real trajectory experience pool. The experience pool is constructed in a queue structure to ensure the completion of the update process of the sampled data.
[0032] In the agent, the value network is used to calculate the Q value; the goal of the policy network is to find a policy that maximizes the weighted sum of the expected cumulative reward and the policy entropy based on the Q value calculated by the value network, the optional policies, the charging current, and the charging state of the battery, that is:
[0033]
[0034] In the formula, π * represents the optimal policy, π represents the optional policy, α represents the weight factor, and γ represents the reward discount factor; r(s t , a t ) is the reward at the current time step, which can be abbreviated as r t ; E represents the calculation of the expectation, that is, the Q value estimated by the value network; represents the entropy; a t represents the charging current at the current time step, and its value range in this embodiment is from 1C to 4C; s t is the charging state of the battery at the current time step, which can be specifically expressed as s t =(SoC t , V t , T t , η t ), where SoC t is the state of charge SoC of the battery at the current time step, V t is the battery voltage at the current time step, T t is the battery temperature at the current time step, and η t is the negative electrode potential of the battery at the current time step.
[0035] To reduce the bias of Q value estimation, the value network in this embodiment uses two independent Q networks, namely the first Q network and the second Q network to approximate the state-action value function. At the same time, for training stability, two independent target Q networks are also introduced in the value network, namely the first target Q network and the second target Q network to calculate the target Q value.
[0036] The calculation result of the target Q value is the expected return (i.e., the calculation of the expectation) in the goal of the policy network. The calculation formula of the target Q value is:
[0037]
[0038] In the formula, y t represents the target Q value, represents the calculation of the expectation starting from the next time step, a t+1 represents the charging current at the next time step, st+1 Represents the charging state at the next time step. Represents entropy, which is an index used to measure the randomness of the policy and is defined as:
[0039]
[0040] Represents the calculated expectation under the current policy. The introduction of entropy encourages the policy to maintain greater randomness when choosing actions, thereby enhancing the exploration ability of the policy; the weight factor α is a regularization parameter used to balance the weights between the reward and entropy.
[0041] Step 2: Use the sampled charging state transition data to train the battery agent model; use each state and the next state of the real sampled trajectory data as the input and output of the battery agent model for model training to minimize the prediction error of the battery agent model for the charging state transition of the battery.
[0042] In this embodiment, the battery agent model includes a multi-layer perceptron, with the input being the charging state of the battery and the output being the next charging state of the battery. The training data set of the battery agent model is derived from the charging state transition data obtained in Step 1. In this embodiment, the time step interval for training the model is set to 1000, that is, every 1000 times of sampling of the charging state transition data, the battery agent model is trained once based on all the data in the current real sampled data experience pool (i.e., the real trajectory experience pool). During the entire reinforcement learning cycle, the battery agent model will be continuously trained to maintain adaptability to the transfer of the sampled data distribution. The update of the battery agent model is based on the following formula:
[0043]
[0044] where ω k+1 represents the updated parameters of the battery agent model, ω k represents the current parameters of the battery agent model, represents defining the loss function of the battery agent model as the error between the real next charging state and the predicted next charging state based on the current charging state and action, f ω (s t ,a t ) is the predicted next charging state; represents parameter gradient calculation; represents the calculation expectation on the real trajectory experience pool; D real represents the real trajectory experience pool; η represents the learning rate.
[0045] Step 3: Based on the true charging state trajectory, use the battery surrogate model for short-branch deduction. That is, starting from a certain charging state in the true sampling trajectory, take an action different from the next action of the true charging state trajectory, and use the prediction of the battery surrogate model to obtain the next charging state of the battery, that is, simulate the sampling process of the battery to obtain the branch trajectory data of the deduced charging state transition. Store the deduced branch trajectory data in the deduced branch trajectory data experience pool.
[0046] In this embodiment, a subset of charging states is randomly selected from the true sampling trajectory data experience pool (i.e., the true trajectory experience pool), and the trained state transition model (i.e., the battery surrogate model) is used for short-branch deduction. It is set that the time step interval of the battery surrogate model deduction is also 1000, that is, the battery surrogate model is deduced after each training of the battery surrogate model.
[0047] Specifically, during the deduction process of the battery surrogate model, given the current charging state s t and the action selected by the agent according to the current policy, use the battery surrogate model to generate the predicted state of the next step:
[0048]
[0049] Among them, represents the new action generated by the agent according to the current policy, and the new action is a charging current different from the charging current a t ; represents the next charging state predicted based on the current charging state and the action under the new action; represents the predicted state of the next step generated by the battery surrogate model under the new action.
[0050] Subsequently, based on the deduced predicted state and the next new action generated according to the current policy carry out the next deduction:
[0051]
[0052] Among them, represents the prediction result based on the predicted state of the next step and the new action of the next step, that is, the predicted state of the step after the next step
[0053] Each deduction generates a pair of state and reward:
[0054]
[0055] Among them, represents the predicted state of the k-th step, represents the reward calculated based on the state of the k-th step, represents the reward for the kth deduction.
[0056] This process can be repeated for multiple steps, resulting in short branching trajectories. Finally, the process generates additional simulation results. It constitutes the experience pool of model deduction data and expands the available training data of the intelligent agent. represents the new action taken in the kth step; Represents the predicted state at the k+1th step.
[0057] The schematic diagram of using the above model for prediction is as follows Figure 3 As shown, by utilizing the data in the real trajectory experience pool, this embodiment deduces branch trajectory data. The schematic diagram of branch deduction is shown in Figure 4 shown.
[0058] Step 4: Combine the actual charging state trajectory data with the deduced branch trajectory data to train the SAC (Soft Actor-Critic) agent including the policy network and the value network; and repeat the above steps 1-3 until the reinforcement learning training converges to obtain a reinforcement learning agent that can execute the optimal charging strategy.
[0059] This embodiment uses the real trajectory experience pool D obtained in steps 1-3. real The branch trajectory experience pool D model Train the agent, i.e. the objective function J(π) of the policy network and the objective function J(π) of the value network (i.e. Q network) Q (θ) are optimized based on the data of these two experience pools:
[0060]
[0061] in, represents the computational expectation of the union of the real trajectory experience pool and the deduced branch trajectory experience pool, Q θ (s t ,a t ) represents the calculation result of the value function for the Q value.
[0062] By combining additional generated deduction samples, the agent can explore more state and action combinations based on limited actual interaction data, further explore potential strategy improvement space, achieve better strategy performance, and significantly improve the utilization efficiency of samples.
[0063] This embodiment also provides a lithium battery charging strategy sample efficiency enhancement system based on model short branch deduction. The enhancement system is implemented based on the above lithium battery charging strategy sample efficiency enhancement method, and specifically includes the following modules:
[0064] Sampling module, which conducts charge and discharge tests on the battery based on the actions taken by the reinforcement learning agent, and samples the charge state transition data; the process of sampling the charge state transition data is as follows: obtain the initial charge state parameters and send the initial charge state parameters to the agent; the agent takes corresponding actions according to the current battery state, and at the same time obtains the corresponding state of the battery at the next moment; the action is to adjust the charging current of the battery.
[0065] Model training module, which trains the battery proxy model using the sampled charge state transition data; uses each state and the next state of the real sampled trajectory data as the input and output of the battery proxy model for model training to minimize the prediction error of the battery proxy model for the charge state transition of the battery.
[0066] Deduction module, based on the real charge state trajectory, conducts short-branch deduction using the battery proxy model, that is, takes a different action from the next action of the real sampled trajectory starting from a certain charge state in the real sampled trajectory, and uses the prediction of the battery proxy model to obtain the next charge state of the battery, that is, simulates the sampling process of the battery to obtain the branch trajectory data of the deduced charge state transition.
[0067] Agent training module, combines the real charge state trajectory data and the deduced branch trajectory data to train the agent; and repeats the sampling of the charge state transition data, the training of the battery proxy model and the deduction until the reinforcement learning training converges to obtain a reinforcement learning agent that can execute the optimal charging strategy.
[0068] Taking the model-free reinforcement learning RL method as the baseline method to evaluate the effectiveness of the method in this embodiment. The learning processes of the RL charging method based on model short-branch deduction and the model-free RL charging method are as Figure 5As shown. As the number of training rounds increases, the constraint violation score approaches the boundary (i.e., zero), indicating that the RL method not only learns voltage and temperature constraints but also learns the optimal solution of the charging strategy (because the optimal solution is near the constraint boundary). The method proposed in this embodiment found a charging strategy with a Return value of -6.18 (charging time of 1558 seconds and basically no constraint violation) at the initial stage of training (defined as the first 100 training rounds), while the best charging strategy found by the model-free RL method at this time had a Return value of -9.49 (charging time of 1895 seconds and basically no constraint violation). The charging time of the best strategy found by the method based on the short-branch deduction of the model at the initial stage of training was reduced by 17.8% compared to the model-free method. The charging time of the charging strategy with basically no constraint violation finally found by the method based on the short-branch deduction of the model after 1000 training rounds was 1544 seconds, only a 0.9% reduction compared to the optimal charging strategy found at the initial stage of training. Therefore, at the initial stage of training, the charging strategy found by the method based on the short-branch deduction of the model was basically optimized. The charging time of the charging strategy with basically no constraint violation finally found by the model-free method after 1000 training rounds was 1565 seconds, a 17.4% reduction compared to the best charging strategy found at the initial stage of training, and it can be seen from subfigure (b) of Figure 5 that the optimization of the charging time by the model-free method decreased gradually within 1000 training rounds, that is, the model-free method still requires sufficient training rounds to complete convergence.
[0069] The charging curves of the best charging strategies found by the method based on the short-branch deduction of the model and the model-free method at the initial and late stages of training, including current, SOC, negative electrode potential, and temperature curves are as Figure 6 shown. It can be seen that there is no obvious gap between the final convergence results of the method proposed in this embodiment and the model-free method, which shows that the utilization of the model in this embodiment does not have an obvious impact on the convergence accuracy. At the initial stage of training, the charging strategy found by the method proposed in this embodiment was close to the optimal, while the charging strategy found by the model-free method was not fully optimized, reflecting the superiority of this embodiment in terms of sample efficiency.
[0070] In the process of reinforcement learning of the present invention, a battery proxy model is constructed and the branch deduction of the model is utilized, effectively improving the sample efficiency of charging strategy optimization and effectively alleviating the problem of error accumulation caused by the recursive utilization of the model in traditional model-based optimization methods.
[0071] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of protection of the claims. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A lithium battery charging strategy sample efficiency enhancement method based on model short branch deduction, characterized in that: The following steps are involved: S1. Perform charging and discharging tests on the battery based on the actions taken by the reinforcement learning agent, and sample the charging state transfer data; The process of charging state transfer data sampling is: obtaining the initial charging state parameters and transmitting the initial charging state parameters to the agent; The intelligent agent takes corresponding actions according to the current battery status and obtains the corresponding status of the battery at the next moment; the action is to adjust the charging current of the battery; S2. Train the battery proxy model using the sampled charging state transfer data; use each state and the next state of the real sampling trajectory data as the input and output of the battery proxy model for model training to minimize the prediction error of the battery proxy model for the battery charging state transfer; S3. Based on the real charging state trajectory, a battery proxy model is used to perform short branch deduction, that is, taking a certain charging state in the real sampling trajectory as the starting point, taking an action different from the next action of the real charging state trajectory, and using the prediction of the battery proxy model to obtain the next charging state of the battery, that is, simulating the sampling process of the battery to obtain the branch trajectory data of the deduced charging state transfer; S4. Combine the real charging state trajectory data with the deduced branch trajectory data to train the agent; and repeat the above steps S1-S3 until the reinforcement learning training converges to obtain a reinforcement learning agent that can execute the optimal charging strategy.
2. The lithium battery charging strategy sample efficiency enhancement method according to claim 1, characterized in that: The intelligent agent includes a policy network and a value network. The value network is used to calculate the Q value. The goal of the policy network is to find a strategy that maximizes the weighted sum of the expected cumulative reward and the policy entropy based on the Q value calculated by the value network, the optional strategies, the charging current, and the battery charging status.
3. The lithium battery charging strategy sample efficiency enhancement method according to claim 2, characterized in that: The policy found by the policy network is expressed as: In the formula, π * represents the best strategy, π represents the optional strategy, α represents the weight factor, γ represents the reward discount factor, r(s t ,a t ) is the reward of the current time step; E represents the calculated expectation, that is, the Q value estimated by the value network; represents entropy; a t Indicates the charging current at the current time step; s t is the battery’s state of charge at the current time step, s t =(SoC t ,V t ,T t ,η t ), where SoC t is the battery's state of charge SoC at the current time step, V t is the battery voltage at the current time step, T t is the battery temperature at the current time step, η t is the negative electrode potential of the battery at the current time step.
4. The lithium battery charging strategy sample efficiency enhancement method according to claim 3, characterized in that: The value network of the intelligent agent includes an independent first Q network With the Second Q Network To approximate the state-action value function; the value network also includes two independent first-target Q networks With the second target Q network To calculate the target Q value; the calculation result of the target Q value is the calculation expectation in the strategy network target; The target Q value is calculated as: In the formula, y t represents the target Q value, represents the computational expectation starting from the next time step, a t+1 represents the charging current in the next time step, s t+1 represents the charging state at the next time step; Represents entropy, which is used to measure the randomness of the strategy.
5. The lithium battery charging strategy sample efficiency enhancement method according to claim 2, characterized in that: The battery proxy model includes a multi-layer perceptron, the input of which is the charging state of the battery, and the output of which is the next charging state of the battery; the training data set of the battery proxy model is derived from the charging state transfer data obtained in step S1.
6. The lithium battery charging strategy sample efficiency enhancement method according to claim 5, characterized in that: Step S2 During the training process, the battery proxy model is updated based on: Among them, ω k+1 represents the updated battery agent model parameters, ω k represents the current battery agent model parameters, a t Represents the charging current at the current time step, s t is the battery’s state of charge at the current time step, s t+1 represents the charging state at the next time step; The loss function of the battery proxy model is defined as the error between the actual next charging state and the next charging state predicted based on the current charging state and action, f ω (s t ,a t ) is the predicted next charging state; Indicates parameter gradient calculation represents the computational expectation of the agent on the real trajectory experience pool; D real represents the real trajectory experience pool where the charging state transfer data is stored; η represents the learning rate.
7. The lithium battery charging strategy sample efficiency enhancement method according to claim 5, characterized in that: Step S3 randomly selects a subset of charging states from the real trajectory experience pool and uses the trained battery proxy model to perform short branch deduction; the deduction generates additional simulation results to form a model deduction data experience pool to expand the available training data for the agent.
8. The lithium battery charging strategy sample efficiency enhancement method according to claim 7, characterized in that: Step S3 During the simulation, the battery proxy model takes the current charging state s as a given time step. t And the actions chosen by the agent based on the current strategy, using the battery proxy model to generate the predicted state for the next step: in, represents the new action generated by the agent according to the current strategy, and the new action is related to the charging current a t Different charging currents; Indicates the next charging state predicted based on the current charging state and action under the new action; It represents the next predicted state generated by the battery agent model under the new action; Predicted status based on inference and the next new action generated based on the current strategy Proceed to the next step of deduction: in, Represents the predicted state based on the next step prediction and the prediction result of the next new action, that is, the predicted state of the next step Each game generates a pair of state and reward: in, represents the predicted state at step k, represents the reward calculated based on the state of step k, represents the reward for the kth deduction.
9. The lithium battery charging strategy sample efficiency enhancement method according to claim 2, characterized in that: The real charging state trajectory data is stored in the real trajectory experience pool, and the deduced branch trajectory data is stored in the deduced branch trajectory experience pool; Step S4 During the agent training process, the objective function J(π) of the policy network and the objective function J(π) of the value network Q (θ) are based on the real trajectory experience pool D real The branch trajectory experience pool D model The data is optimized: in, represents the computational expectation of the union of the real trajectory experience pool and the deduced branch trajectory experience pool, Q θ (s t ,a t ) represents the calculation result of the value function for the Q value; y t represents the target Q value, s t is the battery charge state, a t Represents charging current; r t is r(s t ,a t ), which represents the reward of the current time step.
10. A lithium battery charging strategy sample efficiency enhancement system based on model short branch deduction, characterized in that: The method for enhancing the efficiency of a lithium battery charging strategy sample according to any one of claims 1 to 9 is implemented, and the system includes the following modules: The sampling module performs a charge and discharge test on the battery based on the actions taken by the reinforcement learning agent, and samples the charging state transfer data; the charging state transfer data sampling process is: obtaining the initial charging state parameters, and transmitting the initial charging state parameters to the agent; the agent takes corresponding actions according to the current battery state, and obtains the corresponding state of the battery at the next moment; the action is to adjust the charging current of the battery; The model training module uses the sampled charging state transfer data to train the battery proxy model. Each state and the next state of the real sampling trajectory data are used as the input and output of the battery proxy model for model training to minimize the prediction error of the battery proxy model on the battery charging state transfer. The deduction module uses the battery proxy model to perform short branch deduction based on the real charging state trajectory, that is, taking a certain charging state in the real sampling trajectory as the starting point, taking an action different from the next action of the real charging state trajectory, and using the prediction of the battery proxy model to obtain the next charging state of the battery, that is, simulating the sampling process of the battery to obtain the branch trajectory data of the deduced charging state transfer; The agent training module combines the real charging state trajectory data with the deduced branch trajectory data to train the agent; The sampling of charging state transfer data, training and deduction of the battery agent model are repeated until the reinforcement learning training converges, and a reinforcement learning agent that can execute the optimal charging strategy is obtained.
Citation Information
Patent Citations
Intelligent rapid charging method and system for lithium ion battery
CN115632179A
Lithium battery health estimation method and system based on transfer learning and multi-feature fusion
CN119087264A
Lithium battery high-safety rapid charging method based on deep reinforcement learning
CN119171590A
Lithium battery performance prediction method and device based on deep learning, equipment and medium
CN119247142A
System and method for fast charging of lithium-ion batteries
US20210057919A1