System and Method for Optimizing Dynamic Discharge Strategy of Sodium-Ion Batteries Based on Reinforcement Learning

By constructing a battery characteristic model and reinforcement learning algorithm to optimize the battery discharge strategy, combined with local search and real-time monitoring, the safety and reliability problems in the battery charging and discharging strategy are solved, and the battery is efficient, safe and intelligent discharge is achieved.

CN119805252BActive Publication Date: 2025-07-08BEIJING XUNCHAO TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510293890.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-08
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The failure of the prior art to effectively consider changes in the physical and chemical characteristics of the battery has led to hidden dangers in terms of safety and reliability of battery charging and discharging strategies, especially in electric two-wheelers and tricycles, which affect the user experience and scope of use.

Method used

The sodium ion battery dynamic discharge strategy optimization system based on reinforcement learning uses real-time acquisition of battery data, builds a battery characteristic model, optimizes the discharge strategy with reinforcement learning algorithm, and improves the quality of the solution through local searches, and monitors the battery status in real time to ensure safe and stable and efficient discharge.

Benefits of technology

It realizes efficient discharge while ensuring the safety and stability of the battery, avoids safety hazards such as overcharge and overdischarge, and improves the battery's use efficiency and automation level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119805252B_ABST
    Figure CN119805252B_ABST
Patent Text Reader

Abstract

The present invention discloses a system and method for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning, belonging to the technical field of battery discharge. The system includes a data acquisition module, a dynamic adjustment module, and a verification feedback module. The data acquisition module is used to collect relevant data on the discharge of the sodium-ion battery in real time. The dynamic adjustment module is used to construct a reinforcement learning environment and a battery characteristic model, perform reinforcement learning training, and optimize the discharge strategy through local search to improve the quality of the solution. The reinforcement learning training and the local search are continuously executed until the set number of iterations is reached, and the optimized discharge strategy is output to realize the dynamic adjustment and optimization of the discharge strategy. The verification feedback module is used to monitor the battery discharge process in real time and feedback it to the dynamic adjustment module to ensure that the battery works in the best state, improve the use efficiency of the battery, and enhance the automation and intelligence level of battery discharge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of battery discharge, and relates to a system and method for optimizing the dynamic discharge strategy of sodium-ion batteries based on reinforcement learning. Background Art

[0002] Electric two-wheelers and three-wheelers provide a fast and convenient way of traveling with their small and flexible bodies; however, limited by the battery capacity and technical level, their endurance is often limited and they need to be charged frequently, which not only affects the user experience but also limits their scope of use.

[0003] The existing Chinese patent with the authorization announcement number CN118885814B discloses a battery charge and discharge optimization method, system and medium based on deep reinforcement learning, which forms a state set and a charge and discharge action strategy set based on historical data related to charge and discharge, and then trains a target network model; the main network of the target network model includes a prediction actor network and a prediction critic network, and the prediction actor network includes a current policy prediction network and an auxiliary policy prediction network; the current policy prediction network is composed of a diffusion model, which weakens the sensitivity of the model and enhances the policy generation ability and policy exploration ability of the model; based on the comparison learning between the auxiliary policy prediction and the output of the target actor network, it helps to improve the prediction ability of the model and solves the problem of overestimation of action strategies.

[0004] Although the existing technologies can better cope with complex environmental changes, they do not consider the changes in the physical and chemical properties of the battery itself, such as the charge and discharge rate, temperature sensitivity, internal resistance change, self-discharge rate of the battery, as well as the safety and reliability of the discharge strategy in practical applications, and how to avoid safety hazards such as overcharging and over-discharging. Therefore, the present application provides a system and method for optimizing the dynamic discharge strategy of sodium-ion batteries based on reinforcement learning. Summary of the Invention

[0005] Aiming at the deficiencies of the existing technologies, the purpose of the present invention is to provide a system and method for optimizing the dynamic discharge strategy of sodium-ion batteries based on reinforcement learning, which can collect battery data in real time, construct a battery characteristic model, optimize the discharge strategy in combination with the reinforcement learning algorithm, improve the quality of the solution through local search, and monitor the battery state in real time to adjust the discharge strategy in time to ensure efficient discharge of the battery on the premise of safety and stability.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A system for optimizing the dynamic discharge strategy of sodium-ion batteries based on reinforcement learning includes: a data acquisition module, a dynamic adjustment module, and a verification and feedback module;

[0008] The data acquisition module is used to collect relevant data on the discharge of sodium-ion batteries in real time, establish a discharge database, and store historical data and real-time data;

[0009] The dynamic adjustment module is used to construct a reinforcement learning environment and a battery characteristic model. The battery characteristic model is used to estimate the internal state of the battery in real time and serve as the state input of the reinforcement learning environment for reinforcement learning training. Through local search, including SGD and SQP algorithms, the discharge strategy is optimized. The reinforcement learning training and the local search are continuously executed until the set number of iterations is reached, and the optimized discharge strategy is output;

[0010] Among them, the constructed battery characteristic model includes the calculation time step state of charge 、battery temperature and battery internal resistance ;

[0011] The verification and feedback module is used to monitor the battery discharge process in real time. During discharge, it obtains the data of battery parameters in real time, compares them with the preset standard parameter range. When an abnormality is determined, an abnormality alarm is generated, and abnormality adjustment is performed, and the execution result is fed back to the dynamic adjustment module;

[0012] The specific steps of the local search include:

[0013] Obtain the global optimal solution in the reinforcement learning training , and define the loss function as the objective function of the local search , where is the network parameter ;

[0014] In the th iteration of the SGD algorithm, calculate the gradient of the objective function at the point , and calculate according to the gradient;

[0015] Set the gradient iteration threshold to , and continue to use the SGD algorithm for iteration until the gradient value is less than the gradient iteration threshold to obtain the gradient optimization solution ;

[0016] In the th iteration of the SQP algorithm, construct a QP sub-problem according to the point and the constraint conditions;

[0017] Calculate the search direction and , stop the iteration until the maximum number of iterations is reached, and obtain the quadratic optimization solution ;

[0018] Add the quadratic optimization solution to the experience replay buffer and retrain according to the training steps of reinforcement learning;

[0019] Set a loop. Each time the loop is executed, perform one reinforcement learning training and one local search, and perform loop iteration until the set loop iteration upper limit is reached. Stop the loop and output the optimized discharge strategy of the sodium-ion battery.

[0020] Furthermore, constructing the reinforcement learning environment includes defining the state space and the action space , and setting the reward function ;

[0021] Define the state space as , where is the state of charge of the battery, is the battery temperature, is the internal resistance of the battery, is the discharge current;

[0022] Define the set of available discharge current magnitudes as the action space , where is the number of types of available discharge currents, is the action space in the th available discharge current;

[0023] Considering the state of charge, temperature, and discharge efficiency of the sodium-ion battery comprehensively, define the reward function .

[0024] Furthermore, the specific steps of the reinforcement learning training include:

[0025] Initialize the experience replay buffer , network and the target network , , are network parameters;

[0026] For each time step , generate the state vector , and select the action ;

[0027] Execute the action , calculate the next state and the reward ;

[0028] Store the sample into the experience replay buffer ;

[0029] Randomly draw a batch of samples from the experience replay buffer for training, where , , is the number of samples stored in the experience replay buffer ;

[0030] Calculate the target value ;

[0031] Calculate the loss function , and update the parameters of the network by gradient descent, and copy the parameters to the parameters of the target network ;

[0032] Set the iteration threshold to , and judge whether the iteration times of the reinforcement learning reach the upper limit; when , continue the reinforcement learning training; when , perform local search

[0033] Furthermore, the specific steps for monitoring the battery discharge include:

[0034] Transmit the optimized discharge strategy generated by the dynamic adjustment module to the actual sodium-ion battery through the data interface, and control the battery to perform discharge operations according to the optimized discharge strategy;

[0035] Use various sensors equipped in the sodium-ion battery to continuously obtain data of battery parameters;

[0036] Set the standard parameter range of the battery parameters, including the discharge rate range , voltage range , current range , temperature range ; where , are the lower and upper limits of the discharge rate respectively, , are the lower and upper limits of the voltage respectively, , are the lower and upper limits of the current respectively, , are the lower and upper limits of the temperature respectively.

[0037] Furthermore, the specific steps for monitoring the battery discharge further include:

[0038] If , it is determined that the discharge rate is compliant; if or , it is determined that the discharge rate is abnormal; is the discharge rate;

[0039] If , it is determined that the voltage is compliant; if or , it is determined that the voltage is abnormal; is the battery voltage;

[0040] If , it is determined that the current is compliant; if or , it is determined that the current is abnormal;

[0041] If , it is determined that the temperature is compliant; if or , it is determined that the temperature is abnormal;

[0042] Obtain the determination results of the battery parameters. When a parameter is determined to be abnormal, generate an abnormal alarm and perform abnormal adjustment;

[0043] Record the data during the abnormal adjustment process in chronological order to generate an abnormal adjustment data set;

[0044] Use the abnormal adjustment data set as a new training sample and store it in the experience replay buffer and update the

[0045] parameters of the network.

[0046] A method for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning includes:

[0047] Collect relevant data on the discharge of the sodium-ion battery in real time; , and the reward function , and the battery characteristic model includes calculating the state of charge, battery temperature, and battery internal resistance;

[0048] Perform reinforcement learning training and optimize the discharge strategy through local search until the set number of iterations is reached, and output the optimized discharge strategy;

[0049] Monitor the battery discharge process in real time and feedback it to the reinforcement learning training.

[0050] Furthermore, the specific steps of the local search include:

[0051] Obtain the global optimal solution , and define the objective function ;

[0052] In the -th iteration of the SGD algorithm, calculate the objective function at the point to find the gradient , and calculate according to the gradient;

[0053] Set the gradient iteration threshold to , continue to use the SGD algorithm for iteration until the gradient value is less than the gradient iteration threshold to obtain the gradient optimization solution ;

[0054] In the -th iteration of the SQP algorithm, construct a QP sub-problem according to the point and the constraint conditions;

[0055] Calculate the search direction and , until the maximum number of iterations is reached, stop the iteration to obtain the quadratic optimization solution ;

[0056] Add the quadratic optimization solution to the experience replay buffer , and re-train according to the training steps of reinforcement learning;

[0057] Perform cyclic iteration until the set upper limit of cyclic iteration is reached, stop the loop, and output the optimized discharge strategy of the sodium-ion battery.

[0058] Advantages of the present invention:

[0059] Based on the reinforcement learning algorithm and combined with the battery characteristic model, an optimized framework for the discharge strategy with strong adaptability and high learning efficiency is constructed; by introducing a local search strategy, the problem of local optimality that reinforcement learning may fall into is effectively avoided, and the quality of the discharge strategy is further improved; during the iterative process, the discharge strategy is continuously refined and optimized to ensure efficient discharge on the premise of ensuring the safety and stability of the battery; at the same time, the battery discharge process is monitored in real time, and once abnormal parameters are found, an alarm is immediately triggered and adjustments are made to ensure that the battery is always in the best working state, improving the battery's usage efficiency and enhancing the automation and intelligence level of battery discharge. Description of the Drawings

[0060] Figure 1 It is a structural diagram of a sodium-ion battery dynamic discharge strategy optimization system based on reinforcement learning;

[0061] Figure 2 It is a flow chart of reinforcement learning training;

[0062] Figure 3 It is a flow chart of local search;

[0063] Figure 4 It is a flow chart of the method for optimizing the sodium-ion battery dynamic discharge strategy based on reinforcement learning. Detailed Embodiments

[0064] The technical solution of the present invention will be described in detail below through the drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0065] Embodiment 1

[0066] Refer to Figures 1 to 3 As shown, this embodiment introduces a sodium-ion battery dynamic discharge strategy optimization system based on reinforcement learning, including: a data acquisition module, a dynamic adjustment module, a verification feedback module, and a management module;

[0067] The data acquisition module is used to collect relevant data on the discharge of sodium-ion batteries in real time, including but not limited to battery voltage, current, internal resistance, temperature, state of charge, depth of discharge, and establish a discharge database to store historical data and real-time data for subsequent analysis and optimization;

[0068] The dynamic adjustment module optimizes the discharge strategy of sodium-ion batteries based on the reinforcement learning algorithm, constructs a reinforcement learning environment and a battery characteristic model. The battery characteristic model is used to estimate the internal state of the battery in real time and serve as the state input of the reinforcement learning environment, and reinforcement learning training is carried out. In some cases, reinforcement learning may converge to a local optimal solution prematurely, especially when the search space is large and there are multiple local optima. Through local search, the SGD and SQP algorithms are used to further optimize the discharge strategy, thereby improving the quality of the solution. Set the upper limit of the loop iteration, continuously execute reinforcement learning training and local search until the set number of iterations is reached, output the optimized discharge strategy of the sodium-ion battery, and achieve efficient discharge while ensuring the safety and stability of the battery, realizing the dynamic adjustment and optimization of the discharge strategy.

[0069] The verification feedback module is used to monitor and feedback the battery discharge process in real time, which helps to adjust the battery discharge strategy in time, ensure that the battery works in the best state, and realize automatic monitoring and adjustment. During discharge, the data of the battery parameters are obtained in real time and compared with the preset standard parameter range. When an abnormality is determined, an abnormality alarm is generated and abnormal adjustment is carried out, and the execution result is fed back to the dynamic adjustment module.

[0070] The management module is used to integrate and manage resources to realize the coordinated operation of the vehicle and the charging station. Through real-time communication protocols such as MQTT, HTTP / HTTPS, a data interface is established with the charging station, and the detailed information of the charging station such as location, power, and idle status is obtained regularly and stored in the constructed charging database to ensure the real-time and accuracy of the data. The charging reservation function is supported, and a reservation request is sent to the charging station to ensure that the vehicle can charge smoothly when it arrives at the charging station. After the vehicle arrives at the charging station, it is confirmed that the vehicle has arrived and the reserved status of the charging pile is released, allowing the vehicle to start charging.

[0071] Collect information such as the location, power, and idle status of the charging station, intelligently plan the charging route and time according to the vehicle location and battery status, and communicate with the charging station to realize the charging reservation and scheduling functions, reduce the charging waiting time, and improve the charging efficiency.

[0072] Furthermore, constructing the reinforcement learning environment includes defining the state space and the action space , and setting the reward function ;

[0073] According to the physical and chemical characteristics of the sodium-ion battery, the state of the sodium-ion battery is composed of multiple key parameters of the sodium-ion battery, and the state space is defined as , where is the state of charge of the battery, is the battery temperature, is the internal resistance of the battery, is the discharge current;

[0074] Since different discharge currents will have different effects on the discharge process of the battery, according to the actual allowable discharge current range of the sodium-ion battery, the discharge current that can be taken is defined as an action, and at this time, the set of the magnitudes of the discharge currents that can be taken is the action space , where is the number of types of discharge currents that can be taken; is the action space in the th discharge current that can be taken;

[0075] In order to evaluate the pros and cons of an agent taking an action in a certain state, comprehensively considering the state of charge, temperature, and discharge efficiency of the sodium-ion battery, the expression of the reward function is as follows:

[0076] ;

[0077] ;

[0078] ;

[0079] ;

[0080] In the formula, , , are weight coefficients and satisfy , and are used to adjust the importance of different factors; is the reward function of and are respectively the lower and upper safety limits of . When is within the lower and upper safety limits, a positive reward is given, and a negative reward is given when it is too low or too high; is the reward function of temperature, is the safety threshold of temperature, and a negative reward is given when the temperature exceeds ; is the reward function of discharge efficiency, and the higher the discharge efficiency, the higher the reward; , , , are constants.

[0081] Furthermore, the specific steps of the reinforcement learning training include:

[0082] Initialize the experience replay buffer , network and the target Network , 、 are network parameters;

[0083] For each time step , obtain the state information at the current moment from the detection device of the sodium-ion battery to generate a state vector , and use Network and exploration strategy to select an action ;

[0084] Execute the action , calculate the next state according to the battery characteristic model , and calculate the reward according to the reward function ;

[0085] Store the sample in the experience replay buffer , which helps to break the correlation between data and improve the stability of training;

[0086] Randomly sample a batch of samples from the experience replay buffer for training to reduce the data correlation and improve the efficiency and stability of network training, where , , is the number of samples stored in the experience replay buffer ;

[0087] Calculate the target value , and the expression is as follows:

[0088] ;

[0089] In the formula, is the discount factor, which is used to balance short-term and long-term rewards;

[0090] Define the mean square error as the loss function of reinforcement learning , and update the parameters of the network by the gradient descent method , and copy the parameter to the parameter of the target network

[0091] ;

[0092] ;

[0093] ;

[0094] In the formula, is the parameter before network update, is the parameter after network update, is the learning rate of reinforcement learning, is the parameter before target network update, is the parameter after target network update, is the update coefficient;

[0095] Set the iteration threshold to , and judge whether the iteration times of reinforcement learning reach the upper limit, where each iteration includes interactions of multiple time steps; when , continue with reinforcement learning training; when , perform local search.

[0096] Furthermore, the constructed battery characteristic model includes calculating the state of charge , battery temperature and battery internal resistance , and the expressions are as follows:

[0097] ;

[0098] ;

[0099] ;

[0100] ;

[0101] ;

[0102] In the formula, is the state of charge at time step , is the discharge current at time step , is the time step length, is the rated capacity of the sodium-ion battery, is the self-discharge rate, is the battery temperature at time step , is the battery mass, is the battery specific heat capacity, is the internal resistance at time step , is the battery reaction heat generation power, is the heat dissipation coefficient, is the battery heat dissipation area, is the ambient temperature, for The internal resistance of the battery is is the initial internal resistance (internal resistance at reference SOC and reference temperature), and are the internal resistance changes caused by SOC and temperature, , is a constant, and For reference and reference temperature.

[0103] Furthermore, the specific steps of local search include:

[0104] Obtaining the global optimal solution in reinforcement learning training , and define The network’s loss function is the objective function of local search ,in, is the network parameter ;

[0105] The stochastic gradient descent (SGD) algorithm is used to guide the search direction of the global optimal individual position, update the network parameters, avoid the local optimal phenomenon, find the global optimal solution, and improve the quality of the solution; in the SGD algorithm In the iterations, the objective function is calculated At the point Find the gradient , and update the parameters according to the gradient. The expression is as follows:

[0106] ;

[0107] In the formula, For point The updated parameter value, is the learning rate;

[0108] Set the gradient iteration threshold to , continue to use the SGD algorithm to iterate until the gradient value is less than the preset gradient iteration threshold, and the gradient optimization solution is obtained ;

[0109] The sequential quadratic programming (SQP) algorithm is used to further optimize the In the iteration, according to the point The QP subproblem is constructed by using the constraints as shown below:

[0110] ;

[0111] ;

[0112] ;

[0113] In the formula, is the approximate matrix of the Hessian matrix of the Lagrangian function, is the search direction, is the equality constraint function, is the inequality constraint function; is the transpose of the search direction, is the objective function at the point the transpose of the gradient, is the constraint condition satisfied when determining the search direction, is the equality constraint function at the point the transpose of the gradient, is the inequality constraint function at the point the transpose of the gradient;

[0114] Use the QP solver to solve the above QP sub-problem to obtain the search direction , and update the parameters until the maximum number of iterations is reached, then stop the iteration to obtain the quadratic optimization solution , and the expression is as follows:

[0115] ;

[0116] In the formula, is the updated parameter value of the point , is the step size;

[0117] Convert the discharge strategy corresponding to the quadratic optimization solution into a sample that can be used by reinforcement learning, and add it to the experience replay buffer , and re-train according to the training steps of reinforcement learning, including sampling from the experience replay buffer, calculating the target value, and updating the network;

[0118] Set a loop, and perform a reinforcement learning training and a local search each time the loop is executed. Through the loop iteration of reinforcement learning and local search, continuously improve the performance of the discharge strategy and approach the global optimal solution until the set loop iteration upper limit is reached, then stop the loop and output the optimized discharge strategy of the sodium-ion battery.

[0119] Furthermore, the specific steps for monitoring the battery discharge include:

[0120] The optimized discharge strategy generated by the dynamic adjustment module is transmitted to the actual sodium-ion battery through the data interface, and the battery is controlled to perform the discharge operation according to the optimized discharge strategy;

[0121] Using various sensors equipped in the sodium-ion battery, at a certain sampling frequency Collect the data of the sensors to obtain the data of the battery parameters in real time. Among them, the sensors include current sensors, voltage sensors, temperature sensors, and other auxiliary sensors, and the battery parameters include discharge rate 、battery voltage 、current and temperature ;

[0122] Set the standard parameter ranges of the battery parameters, including the discharge rate range 、voltage range 、current range 、temperature range , and compare the battery parameters obtained in real time with the preset standard parameter ranges one by one; among them, 、 are the lower and upper limits of the discharge rate respectively, 、 are the lower and upper limits of the voltage respectively, 、 are the lower and upper limits of the current respectively, 、 are the lower and upper limits of the temperature respectively;

[0123] For the discharge rate , if , it is determined that the discharge rate is compliant; if or , it is determined that the discharge rate is abnormal;

[0124] For the battery voltage , if , it is determined that the voltage is compliant; if or , it is determined that the voltage is abnormal;

[0125] For the current , if , it is determined that the current is compliant; if or , it is determined that the current is abnormal;

[0126] For the temperature , if , it is determined that the temperature is compliant; if or , it is determined that the temperature is abnormal;

[0127] Obtain the determination result of battery parameters. When a parameter determination is abnormal, generate an abnormal alarm and perform abnormal adjustment;

[0128] Obtain the relevant data during the abnormal adjustment process, including the abnormal parameters before adjustment ( , , , ), the adjusted parameters ( , , , ), and the actions taken during the adjustment (such as the proportional coefficient of the adjusted current, the connected load resistance, and the power of the heat dissipation or heating system). Record the data of each adjustment step in chronological order to generate an abnormal adjustment data set; among them, , , , are the discharge rate, voltage, current, and temperature before abnormal adjustment respectively, , , , are the discharge rate, voltage, current, and temperature after abnormal adjustment respectively;

[0129] Use the abnormal adjustment data set as a new training sample and store it in the experience replay buffer to re - perform reinforcement learning training and update the parameters of the network.

[0130] Embodiment 2

[0131] Please refer to Figure 4 , another embodiment provided by the present invention: a method for optimizing the dynamic discharge strategy of sodium - ion batteries based on reinforcement learning, including the following steps:

[0132] S1, Collect the relevant data of the sodium - ion battery discharge in real - time;

[0133] S2, Construct a reinforcement learning environment and a battery characteristic model; among them, the reinforcement learning environment includes a state space , an action space and a reward function . The battery characteristic model is used to estimate the internal state of the battery in real - time and serve as the state input of the reinforcement learning environment, including calculating the state of charge, battery temperature, and battery internal resistance;

[0134] S3, Perform reinforcement learning training, and further optimize the discharge strategy through local search using SGD and SQP algorithms to improve the quality of the solution; set the upper limit of the loop iteration, and continuously execute the reinforcement learning training and local search until the set number of iterations is reached, and output the optimized discharge strategy of the sodium - ion battery;

[0135] S4. Monitor the battery discharge process in real time. During discharge, obtain the data of battery parameters in real time, compare them with the preset standard parameter range. When an abnormality is determined, generate an abnormality alarm and perform abnormality adjustment; and feedback the execution result to the reinforcement learning training.

[0136] Furthermore, the specific steps of local search include:

[0137] Obtain the global optimal solution in the reinforcement learning training , and define the loss function of the network in the reinforcement learning as the objective function of local search , where is the network parameter ;

[0138] Use the Stochastic Gradient Descent (SGD) algorithm to guide the search direction of the global best individual position, update the network parameters, avoid the phenomenon of local optimality, and search for the global optimal solution to improve the quality of the solution; in the th iteration of the SGD algorithm, calculate the objective function at the point to obtain the gradient , and update the parameters according to the gradient. The expression is as follows:

[0139] ;

[0140] In the formula, is the updated parameter value at the point , is the learning rate;

[0141] Set the gradient iteration threshold to , and continue to use the SGD algorithm for iteration until the gradient value is less than the preset gradient iteration threshold to obtain the gradient-optimized solution ;

[0142] Use the Sequential Quadratic Programming (SQP) algorithm for further optimization. In the th iteration, construct the QP sub-problem according to the point and the constraint conditions. The expression is as follows:

[0143] ;

[0144] ;

[0145] ;

[0146] In the formula, is an approximate matrix of the Hessian matrix of the Lagrangian function, is the search direction, is the equality constraint function, is the inequality constraint function;

[0147] Use the QP solver to solve the above QP sub-problem to obtain the search direction , and update the parameters until the maximum number of iterations is reached, then stop the iteration to obtain the quadratic optimization solution , and the expression is as follows:

[0148] ;

[0149] In the formula, is the updated parameter value of point , is the step size;

[0150] Convert the discharge strategy corresponding to the quadratic optimization solution into a sample that can be used by reinforcement learning, add it to the experience replay buffer , and re-train according to the training steps of reinforcement learning, including sampling from the experience replay buffer, calculating the target value, and updating the network;

[0151] Set a loop, and perform a reinforcement learning training and a local search each time the loop is executed. Through the loop iteration of reinforcement learning and local search, continuously improve the performance of the discharge strategy and approach the global optimal solution until the set loop iteration upper limit is reached, then stop the loop and output the optimized discharge strategy of the sodium-ion battery.

[0152] In summary, in the above embodiments, the present invention uses a reinforcement learning algorithm, combines a battery characteristic model, deeply optimizes the discharge strategy, incorporates a local search strategy, and further improves the strategy quality with the help of SGD and SQP algorithms to ensure that the battery maintains its safety and stability while discharging efficiently; at the same time, it monitors the battery discharge state in real time, immediately triggers an alarm and performs adjustments once an abnormality is found, ensures that the battery is always in the best working state, and the execution result is immediately fed back to the dynamic adjustment module to form a closed-loop optimization link.

[0153] The above are only the preferred embodiments of the present invention. The protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.

Claims

1. A sodium-ion battery dynamic discharge strategy optimization system based on reinforcement learning, characterized in that Including: A data acquisition module, a dynamic adjustment module, and a verification feedback module; The data acquisition module is used to collect relevant data on the discharge of the sodium-ion battery in real time, establish a discharge database, and store historical data and real-time data; The dynamic adjustment module is used to construct a reinforcement learning environment and a battery characteristic model. The battery characteristic model is used to estimate the internal state of the battery in real time and serve as the state input of the reinforcement learning environment for reinforcement learning training. Through local search, including SGD and SQP algorithms, the discharge strategy is optimized. The reinforcement learning training and the local search are continuously executed until the set number of iterations is reached, and the optimized discharge strategy is output; Among them, the constructed battery characteristic model includes a calculation time step State of Charge , battery temperature and battery internal resistance ; The verification feedback module is used to monitor the battery discharge process in real time. During discharge, it obtains the data of the battery parameters in real time and compares them with the preset standard parameter range. When an abnormality is determined, an abnormality alarm is generated, and abnormal adjustment is performed, and the execution result is fed back to the dynamic adjustment module; The specific steps of the local search include: Obtain the global optimal solution in the reinforcement learning training , and define a loss function as the objective function for local search , where is the network parameter ; In the -th iteration of the SGD algorithm, calculate the objective function at the point to find the gradient , and calculate according to the gradient; Set the gradient iteration threshold to , and continue to use the SGD algorithm for iteration until the gradient value is less than the gradient iteration threshold to obtain a gradient optimization solution ; In the -th iteration of the SQP algorithm, a QP sub-problem is constructed according to the point and the constraint conditions; Calculate the search direction and , until the maximum number of iterations is reached, stop the iteration, and obtain the quadratic optimization solution ; Add the secondary optimized solution to the experience replay buffer and retrain according to the training steps of reinforcement learning; Set a loop. Each time the loop is executed, a reinforcement learning training and a local search are performed, and loop iteration is carried out until the set loop iteration upper limit is reached, the loop is stopped, and the optimized discharge strategy of the sodium-ion battery is output.

2. The system for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning according to claim 1, wherein: Building a reinforcement learning environment involves defining the state space and the action space , as well as setting the reward function ; Define the state space as , where is the state of charge of the battery, is the battery temperature, is the internal resistance of the battery, is the discharge current; Define the set of magnitudes of the discharge current that can be taken as the action space , where is the number of types of discharge currents that can be taken; is the action space the th discharge current that can be taken; Considering the state of charge, temperature, and discharge efficiency of a sodium-ion battery comprehensively, a reward function is defined .

3. The system for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning according to claim 2, wherein The specific steps of the reinforcement learning training include: Initialize the experience replay buffer and network and target network , and are network parameters; For each time step , generate a state vector , and select an action ; Execute the said action , calculate the next state and the reward ; Store the sample in the experience replay buffer ; Randomly draw a batch of samples from the experience replay buffer for training, where is the number of samples stored in the experience replay buffer , is the experience replay buffer ; Calculation target Value ; Calculate the loss function , and update through gradient descent the parameters of the network , copy the parameters to the target parameters of the network ; Set the iteration threshold to , and judge whether the iteration times of the reinforcement learning reach the upper limit; when , continue the reinforcement learning training; when , perform local search.

4. The system for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning according to claim 3, characterized in that, The specific steps of monitoring the battery discharge include: Transmit the optimized discharge strategy generated by the dynamic adjustment module to the actual sodium-ion battery through a data interface, and control the battery to perform a discharge operation according to the optimized discharge strategy; Use various sensors equipped in the sodium-ion battery to obtain the data of the battery parameters in real time; Set the standard parameter ranges for battery parameters, including the discharge rate range , voltage range , current range , temperature range ; Among them, , are the lower and upper limits of the discharge rate respectively, , are the lower and upper limits of the voltage respectively, , are the lower and upper limits of the current respectively, , are the lower and upper limits of the temperature respectively.

5. The system for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning according to claim 4, wherein The specific steps of monitoring the battery discharge further include: If , it is determined that the discharge rate is compliant; if or , it is determined that the discharge rate is abnormal; is the discharge rate; If , it is determined that the voltage is compliant; if or , it is determined that the voltage is abnormal; is the battery voltage; If , it is determined that the current is compliant; if or , it is determined that the current is abnormal; If , it is determined that the temperature is compliant; if or , it is determined that the temperature is abnormal; Obtain the determination result of the battery parameters. When a parameter is determined to be abnormal, an abnormality alarm is generated, and abnormal adjustment is performed; Record the data during the abnormal adjustment process in chronological order to generate an abnormal adjustment data set; Use the abnormal adjustment data set as a new training sample and store it in the experience replay buffer and update the parameters of the network.

6. The method for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning is implemented based on the system for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning as described in any one of claims 1-5, and is characterized in that Including: Collect relevant data on the discharge of the sodium-ion battery in real time; Build a reinforcement learning environment and a battery characteristic model; among them, the reinforcement learning environment includes a state space , an action space and a reward function , and the battery characteristic model includes calculating the state of charge, battery temperature, and battery internal resistance; Perform reinforcement learning training, and optimize the discharge strategy through local search until the set number of iterations is reached, and output the optimized discharge strategy; Monitor the battery discharge process in real time and feed it back to the reinforcement learning training.

7. The method for optimizing the dynamic discharge strategy of a sodium-ion battery based on reinforcement learning according to claim 6, wherein The specific steps of the local search include: Obtain the global optimal solution , and define the objective function ; In the -th iteration of the SGD algorithm, calculate the objective function at the point to obtain the gradient , and calculate according to the gradient; Set the gradient iteration threshold to , and continue to use the SGD algorithm for iteration until the gradient value is less than the gradient iteration threshold to obtain a gradient optimization solution ; In the -th iteration of the SQP algorithm, a QP sub-problem is constructed according to the point and the constraint conditions; Calculate the search direction and , until the maximum number of iterations is reached, stop the iteration, and obtain the quadratic optimization solution ; Add the secondary optimization solution to the experience replay buffer and retrain according to the training steps of reinforcement learning; Perform loop iteration until the set loop iteration upper limit is reached, stop the loop, and output the optimized discharge strategy of the sodium-ion battery.

Citation Information

Patent Citations

  • Battery charging and discharging optimization method, system and medium based on deep reinforcement learning

    CN118885814B

  • Lithium ion battery pack management strategy optimization method based on deep Q network

    CN118861517A

  • Energy storage EMS system SOC balance control system and method based on artificial intelligence

    CN119298266A