Intelligent energy storage control system and method for unit battery
By building an environment mapping model and a reinforced learning framework, combining battery performance and environmental parameters, dynamically adjusting energy storage strategies, the problems of fluctuations in efficiency and shortening of life of traditional unit batteries in different environments are solved, and more efficient and stable energy storage control is achieved.
Patent Information
- Application Number
- CN202510352647.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The efficiency of traditional battery energy storage control systems in different temperature and humidity environments fluctuates greatly, and lacks intelligence and adaptability, resulting in premature aging of the battery and reduced energy storage efficiency.
The environment mapping model is constructed using the optimized and adjusted deep neural network, combined with the reinforcement learning framework and recurrent neural network, and dynamically adjust the energy storage strategy by collecting battery performance and environmental parameters to optimize the charging and discharging efficiency and life of the battery.
It improves the energy storage efficiency and life of the battery in different environments, achieves more accurate battery state perception and energy storage control, and improves the overall performance of the system.
Smart Images

Figure CN120377413A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of battery data processing, and in particular to an intelligent energy storage control system and method for unit batteries. Background Art
[0002] At present, unit batteries are widely used as energy storage devices in many scenarios, such as distributed energy systems, uninterruptible power supplies (UPS), etc. However, with the growth of energy demand and the increasing requirements for energy storage efficiency and battery life, traditional unit battery energy storage control technologies have gradually exposed many problems.
[0003] Existing energy storage control systems usually directly measure basic parameters such as the charging and discharging efficiency and internal resistance of the battery for energy storage control. However, in different temperature environments, the charging and discharging efficiency of the battery can fluctuate by up to 30%, and the internal resistance can also change by 20%-50%. At the same time, humidity affects the chemical stability of the battery, and in an environment with high humidity, the self-discharge rate of the battery will increase by 15%-25%;
[0004] On the other hand, when traditional control systems adjust the energy storage strategy of the battery, they lack intelligence and adaptability, and cannot dynamically optimize the energy storage scheme according to the real-time state of the battery and environmental changes. In practical applications, the method of using a fixed strategy for energy storage control without adjusting the strategy according to the battery state and environmental changes will cause the battery to age prematurely, the actual service life will be reduced, and the energy storage efficiency will be reduced by 10%-15%. Therefore, an intelligent energy storage control system and method for unit batteries are proposed herein. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention proposes the following technical solutions:
[0006] An intelligent energy storage control system for unit batteries, comprising:
[0007] A data acquisition module: collecting battery performance parameters and environmental parameters of the battery pack;
[0008] Among them, the battery performance parameters at least include charging efficiency, discharging efficiency, and battery internal resistance, and the environmental parameters at least include temperature change parameters and humidity change parameters;
[0009] An environment mapping module: constructing an environment mapping model through an optimized and adjusted deep neural network, using the battery performance parameters and environmental parameters as model input data, and outputting a mapping influence coefficient;
[0010] An intelligent allocation module: using the output of the environment mapping model and the battery performance feedback data as the state and reward of reinforcement learning to construct a reinforcement learning framework, and optimizing the selection strategy of the reinforcement learning framework to obtain an optimal strategy allocation scheme;
[0011] Among them, the optimized and adjusted deep neural network dynamically adjusts the number of hidden layer nodes by obtaining the joint entropy of the model input data; the selection strategy of the optimized reinforcement learning framework is implemented by combining the ∈-greedy strategy with a recurrent neural network.
[0012] The charging efficiency is obtained by measuring the input energy and the effectively stored energy during the battery charging process; the discharging efficiency is obtained by measuring the output energy and the discharging energy within the discharging time through a power sensor; the internal resistance of the battery is obtained by injecting a small AC signal with a known frequency and amplitude into the battery through an internal resistance measurement sensor and measuring the response voltage across the battery.
[0013] The temperature change parameter is obtained by a thermocouple temperature sensor, which is distributed at different positions of the battery pack to convert the temperature change into a voltage signal, and by measuring the voltage value and combining with a calibration curve; the humidity change parameter is obtained by a capacitive humidity sensor, which is placed in the environment where the battery pack is located to measure the change of the capacitance value with the environmental humidity and is obtained through signal processing circuitry.
[0014] The process of constructing the environmental mapping model is as follows:
[0015] The environmental mapping model includes an input layer, an optimized and adjusted hidden layer, and an output layer;
[0016] The number of input layer nodes is determined according to the types of collected environmental parameters and battery performance parameters;
[0017] The number of nodes in the optimized and adjusted hidden layer is set to twice the number of input layer nodes, and a three-layer hidden layer structure is adopted. Each hidden layer uses the ReLU activation function and is adjusted by the joint entropy dynamic adjustment strategy.
[0018] The number of output layer nodes is obtained according to the types of predicted battery mapping data and linear activation is adopted.
[0019] The joint entropy dynamic adjustment strategy includes:
[0020] Dividing the model input data into data groups according to a time window;
[0021] Calculating the joint entropy H(z) of each group of data, where z represents the joint data, z = [η charge , η discharge , R, T, H];
[0022] When the joint entropy H(z) > H upper , increasing the hidden layer by 10% of the current number of nodes;
[0023] When the joint entropy H(z) < H lower , reducing the hidden layer by 5% of the current number of nodes;
[0024] When the joint entropy H lower ≤ H(z) ≤ H upper Keep the number of hidden layer nodes unchanged.
[0025] The battery performance feedback data at least includes the actual charging efficiency, the actual discharging efficiency, and the battery life change rate.
[0026] The reinforcement learning framework takes the output of the environment mapping model and the battery performance feedback data as the state and reward signals of reinforcement learning:
[0027] The reinforcement learning framework includes an agent, an environment V, and a state S;
[0028] Among them, the agent controls the environment mapping model and the weight assignment; the environment V at least includes the charging and discharging processes of the battery at different times, the physical environment it is in, and the environment mapping model itself; the state S is jointly composed of the output of the environment mapping model and the battery performance feedback data.
[0029] Based on the state S, the agent selects an action A using the ∈-greedy strategy combined with a recurrent neural network, comprehensively evaluates the action based on the existing optimal action and a specific probability of accepting the interference action given by the recurrent neural network, and finally determines the action to be executed
[0030] The process of selecting an action using the ∈-greedy strategy combined with a recurrent neural network is as follows:
[0031] Create and initialize a table storing the expected cumulative rewards Q(S, A) for different actions A executed in different states S, and input the current state S through the recurrent neural network t , and output a series of candidate interference actions A adversarial ;
[0032] At time step t, combine the output data of the environment mapping model and the battery performance feedback data into the state S t ;
[0033] Based on the agent querying the Q-table, find the actions corresponding to the specified Q values in the current state S t , and form a preliminary optimal action set A optimal ;
[0034] Randomly select an action from the set of interference actions A provided by the recurrent neural network adversarial as an adversarial action according to a uniform distribution Evaluate the selected adversarial action together with the actions in the preliminary optimal action set A optimal ;
[0035] According to the evaluation results, select the action with the highest score as the action to be finally executed.
[0036] An intelligent energy storage control method for a unit battery, the steps of the method are as follows:
[0037] S1: Collect the battery performance parameters and environmental parameters of the battery pack. The battery performance parameters include charging efficiency, discharging efficiency, and battery internal resistance, and the environmental parameters include temperature change parameters and humidity change parameters.
[0038] S2: Construct an environmental mapping model through an optimized and adjusted deep neural network, use the battery performance parameters and environmental parameters as model inputs, and output a mapping influence coefficient.
[0039] S3: Based on the output of the environmental mapping model and the battery performance feedback data as the state and reward of reinforcement learning, construct a reinforcement learning framework, and optimize the selection strategy of the reinforcement learning framework to obtain an optimal policy allocation scheme.
[0040] The present invention has the following beneficial effects:
[0041] In the present invention, firstly, through the optimized and adjusted deep neural network in the environmental mapping module, the number of hidden layer nodes is dynamically adjusted based on the joint entropy of the input data, which can more accurately capture the complex mapping relationship between environmental parameters (temperature, humidity) and battery performance parameters (charging efficiency, discharging efficiency, internal resistance). The mapping influence coefficient output by the environmental mapping model can clearly reflect that the temperature has a significant impact on the overall performance of the battery, which helps to more accurately predict the change of battery performance and improve the system's perception ability of the battery state.
[0042] Secondly, through the ∈-greedy strategy combined with a recurrent neural network in the intelligent allocation module, making full use of the time series characteristics of the battery's historical charge and discharge data and environmental parameters, the intelligent agent can select the optimal action based on existing experience, and at the same time accept the interference action given by the recurrent neural network with a specific probability to explore new strategies, increasing the diversity of action selection. After multiple iterations of optimization, it can better obtain the optimal policy allocation scheme under different battery states and environments, effectively balancing the charge and discharge efficiency and life of the battery, and improving the overall performance of the energy storage system.
[0043] Finally, the environmental mapping module provides an accurate basis for the energy storage control strategy by capturing the mapping relationship between the environment and battery performance parameters. The intelligent allocation module obtains the optimal strategy allocation scheme through the ∈-greedy strategy combined with the recurrent neural network. The combination of the two can formulate an energy storage control strategy that better suits the real-time state of the battery and environmental conditions based on accurate battery state perception. It can accurately adjust the charging power and discharge strategy according to the significant impact of the environment on battery performance and the characteristics of historical charge and discharge data, enabling the battery to operate efficiently and stably. Brief Description of the Drawings
[0044] Figure 1 It is a system block diagram of an intelligent energy storage control system and method for a unit battery proposed by the present invention.
[0045] Figure 2 It is a method step diagram of an intelligent energy storage control system and method for a unit battery proposed by the present invention. Detailed Embodiments
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Embodiment 1
[0048] As Figure 1 shown, an intelligent energy storage control system for a unit battery proposed by the present invention includes:
[0049] Data acquisition module: Collect battery performance parameters and environmental parameters of the battery pack. The battery performance parameters include charging efficiency, discharging efficiency, and battery internal resistance. The environmental parameters include temperature change parameters and humidity change parameters;
[0050] The process of obtaining battery performance parameters is as follows:
[0051] Charging efficiency sensor: Calculate the charging efficiency η by measuring the input energy and effective stored energy during the battery charging process charge ;
[0052] Discharging efficiency sensor: During the battery discharging process, use a power sensor to measure the output power of the discharging circuit and the discharging power of the battery, and calculate the discharging efficiency η through the output energy and discharging energy during the discharging time discharge ;
[0053] Battery internal resistance measurement sensor: An internal resistance measurement sensor using the principle of AC injection method injects a small AC signal with a known frequency and amplitude into the battery, measures the response voltage across the battery, and calculates the internal resistance R of the battery according to Ohm's law;
[0054] The process of obtaining environmental parameters is as follows:
[0055] Select high-precision thermocouple temperature sensors and install them at different positions of the battery pack to obtain the temperature distribution of the battery pack. The thermocouple, based on the thermoelectric effect, converts the temperature change into a voltage signal. By measuring the voltage value and combining it with the calibration curve, the corresponding temperature value T is obtained;
[0056] Humidity sensor: A capacitive humidity sensor is used and placed in the environment where the battery pack is located to monitor the environmental humidity in real time. The capacitance value of the capacitive humidity sensor changes with the change of environmental humidity. By measuring the capacitance value and converting it through a signal processing circuit into a humidity value H;
[0057] The microcontroller based on the ARM architecture is the core of the data acquisition module, which converts the analog signal output by the sensor into a digital signal and transmits the digital signal to the environment mapping module at the same time.
[0058] Environment mapping module: An environment mapping model is constructed through an optimized and adjusted deep neural network. The battery performance parameters and environmental parameters are used as the model inputs, and the mapping influence coefficient is output;
[0059] The optimized and adjusted deep neural network is based on the hidden layer of the original neural network, and adjusts the hidden layer nodes according to the joint entropy of the battery performance parameters and environmental parameters;
[0060] The process of constructing the environment mapping model is as follows:
[0061] An optimized and adjusted deep neural network with one hidden layer node is used as the basic architecture of the environment mapping model;
[0062] The deep neural network includes an input layer, a hidden layer, and an output layer;
[0063] Then the environment mapping model includes an input layer, an optimized and adjusted hidden layer, and an output layer;
[0064] Constructing the input layer of the environment mapping model: The number of input layer nodes is determined according to the types of collected environmental parameters and battery performance parameters. One type corresponds to one node. Let the input layer node be y;
[0065] Constructing the hidden layer of the environment mapping model: The initial number of hidden layer nodes is set to twice the number of input layer nodes, that is, the initial number of hidden layer nodes is 2y, and a multi-layer hidden layer structure is adopted, with 3 hidden layers set. Each hidden layer uses the ReLU activation function, which is expressed as:
[0066] h j = ReLU(W j x i + b j )
[0067] where j represents the layer index (j = 1, 2, 3), h j is the output of the j-th hidden layer, W j is the weight of the j-th hidden layer, b j is the bias term of the j-th hidden layer, and x i is the data value of the i-th input;
[0068] The process of optimizing and adjusting the hidden layer nodes is as follows:
[0069] Group the input data according to certain features or dimensions;
[0070] For example, for the input data composed of battery performance parameters (charging efficiency, discharging efficiency, battery internal resistance) and environmental parameters (temperature, humidity), divide it into k data windows according to the time series, and each window contains a certain number of sample data;
[0071] For each group, calculate the frequency of occurrence of each data value in it, and use it as the probability estimate of the occurrence of this data value. Suppose there are N samples in a data group, and the i-th data value x i appears n i times, and the probability of x i appearing is
[0072] Calculate the joint entropy of each group of data:
[0073]
[0074] where z represents the joint data, z = [η charge , η discharge , R, T, H], and p z is the probability of the joint data appearing simultaneously;
[0075] By calculating p z (the probability of multiple parameters appearing simultaneously), the joint entropy can reveal the hidden relationship between parameters. For example, under specific humidity and temperature conditions, the distribution law of charging efficiency, helping to discover the influence mechanism of multi-parameter cooperation on battery performance, providing guidance for the deep neural network to capture complex correlation features, and optimizing the network's learning ability for multi-parameter coupling modes;
[0076] According to the calculated joint entropy, formulate an optimization and adjustment strategy for the number of hidden layer nodes;
[0077] Set a threshold: By setting the upper threshold H of the joint entropyupper and the lower threshold H lower , and these two thresholds are used to define different degrees of data complexity;
[0078] For example, when the joint entropy H(z) > H upper , it is considered that the data complexity is high;
[0079] When the joint entropy H(z) < H lower , it is considered that the data complexity is low;
[0080] The upper threshold H lower ≤H(z)≤H upper , it is considered that the data complexity is in a moderate range;
[0081] When the joint entropy H(z) > H upper , that is, when the data complexity is large, increase the number of hidden layer nodes by 10% of the current number of hidden layer nodes;
[0082] When the joint entropy H(z) < H lower , that is, when the data complexity is small, reduce the number of hidden layer nodes. Too many nodes are likely to cause overfitting on simple data. Reducing the number of nodes can make the model more concise and effective, and reduce the number of nodes by 5% of the current number of nodes;
[0083] When H lower ≤H(z)≤H upper , that is, when the data complexity is moderate, keep the number of hidden layer nodes unchanged;
[0084] Construct the output layer of the environment mapping model:
[0085] The number of output layer nodes is obtained according to the types of predicted battery mapping data;
[0086] For example: if predicting 3 battery performance parameters including charging efficiency, discharging efficiency, and battery internal resistance, then the number of output layer nodes is set to 3, because each parameter to be predicted corresponds to a node in the output layer for outputting the predicted value of that parameter;
[0087] Design the output layer structure: The output layer uses linear activation;
[0088] Specifically, the output layer predicts the relationship between the environment and battery performance. Therefore, the number of output layer nodes is set to 2. Let the weight matrix from the third hidden layer to the output layer be W4 and the bias vector be b4. Then the output of the third hidden layer is h3, and the output of the output layer is expressed as Because the battery performance parameters have no restrictions on the value range in practical significance and do not require output within a specific interval, the linear activation of the battery performance parameters can directly output the predicted value, mapping the features extracted by the hidden layer to the predicted values of specific battery mapping data;
[0089] The mean square error is used as the loss function to measure the gap between the output result of the model and the mapping relationship between the actual environment and battery performance. Let the quantization value of the mapping relationship between the actual environment and battery performance be Y, then the loss function
[0090]
[0091] where N is the number of samples in the data grouping;
[0092] Then the output result of the environment mapping model is F1 represents the mapping degree value between the current environment (temperature) and the overall battery performance (charge-discharge efficiency, battery internal resistance) (one-to-many). The larger F1 is, the more significant the influence of the current environmental parameter (temperature) on the overall battery performance. The larger F2 is, the more significant the influence of the current environmental parameter (humidity) on the overall battery performance;
[0093] Example: In a battery environment of a unit:
[0094] Environmental parameters: temperature is 15°C - 30°C, humidity is 30% - 60%;
[0095] Battery performance parameters: charge efficiency is 80% - 90%, discharge efficiency is 75% - 85%, battery internal resistance is 10 mΩ - 20 mΩ;
[0096] Because there are 5 parameters, the number of nodes in the input layer of the environment mapping model is set to 5, and the normalized data is composed into an input vector and input into the model;
[0097] The number of nodes in the initial hidden layer is set to twice the number of nodes in the input layer, that is, 10 nodes, and 3 hidden layers are constructed;
[0098] The input data within one week is divided into windows with 12 samples in each window, and a total of 14 data windows are obtained. Taking one data window as an example, according to the joint entropy formula, the joint entropy H of the joint data within this window is calculated as H = 1.8. Set the upper threshold H upper = 2.0, and the lower threshold H lower = 1.2. Since 1.2 < H = 1.8 < 2.0, it indicates that the complexity of this grouped data is moderate, and the number of nodes in the hidden layer remains unchanged at 10;
[0099] The environment mapping model outputs values through the output layer
[0100] Result Interpretation: The first value, 2.5, is relatively large, indicating that the current environmental parameter (temperature) has a significant impact on the overall performance of the battery. The second value, 1.2, is relatively small, suggesting that in another dimension, the correlation between the environment (humidity) and battery performance (such as humidity's effect on the battery internal resistance) is weak. Based on these results, in the actual intelligent energy storage control system of the unit battery, more attention is paid to the impact of temperature on battery performance to avoid overheating of the battery, which may affect its performance and lifespan.
[0101] Intelligent Allocation Module: Based on the output of the environmental mapping model and the battery performance feedback data as the state and reward for reinforcement learning, construct a reinforcement learning framework, and optimize the selection strategy of the reinforcement learning framework to obtain the optimal strategy allocation plan.
[0102] The battery performance feedback data includes the actual charging efficiency, actual discharging efficiency, and battery life change rate.
[0103] Construct a reinforcement learning framework. Use the output of the model (mapping influence coefficient) and the battery performance feedback data as the state and reward signals of reinforcement learning, and continuously try different output strategies to make the prediction results gradually approach the optimal solution.
[0104] The reinforcement learning framework includes:
[0105] Agent: Control the environmental mapping model and data allocation.
[0106] Environment V: The environment includes the actual operating conditions of the battery, which consists of the charging and discharging processes of the battery at different times, the physical environment it is in (temperature, humidity, etc.), and including the environmental mapping model itself. The environment will change accordingly according to the actions of the agent and provide feedback information.
[0107] State S: The state is jointly composed of the output of the environmental mapping model and the battery performance feedback data. The output of the environmental mapping model is
[0108] The actual battery performance feedback data is Among them, represents the actual charging efficiency, represents the actual discharging efficiency, η * represents the battery life change rate, which is directly obtained through the battery management system (BMS) or relevant monitoring devices.
[0109] Based on the current state S, the agent selects an action A using an improved ε-greedy strategy (combined with a recurrent neural network), the optimal action based on existing experience, and accepts the interfering action given by the recurrent neural network with a specific probability. For example, the agent first determines a preliminary set of optimal actions, then randomly selects an action from the actions provided by the recurrent neural network, and comprehensively evaluates it with the actions in the preliminary set of optimal actions to finally determine the action to be executed.
[0110] The process of the ε-greedy strategy combined with a recurrent neural network is as follows:
[0111] Create a table for storing the expected cumulative reward Q(S, A) for executing different actions A in different states S. In the initial stage, initialize all Q values to a small random value, such as 0.
[0112] Example: Create a Q(S, A) table. Assume that the state S includes two types, S1 and S2, and the action A includes two types, A1 (adjust the charging power) and A2 (adjust the discharging strategy). Initially, Q(S1, A1) = 0, (S1, A2) = 0, Q(S2, A1) = 0, Q(S2, A2) = 0.
[0113] The recurrent neural network inputs the current state S t , where t is the time step, and S t represents the state at time step t, and the output is a series of candidate interfering actions. At the same time, initialize the weights and biases of the recurrent neural network.
[0114] Specifically, the recurrent neural network is good at processing sequence data, capturing long-term dependencies in the time series through its memory unit. For example, analyzing the historical charge and discharge data of the battery, the variation law of environmental parameters over time, etc. When combined, the recurrent neural network can be used to process state inputs (such as time series data of battery performance), extract sequence features, and the ε-greedy strategy decides the exploration or exploitation actions based on the state processed by the recurrent neural network.
[0115] Probability parameter: Set a specific probability of accepting the action of the recurrent neural network and the ε value in the traditional ε-greedy strategy.
[0116] At time step t, the agent collects relevant information of the current environment, including the output of the environmental mapping model and the battery performance feedback data, and integrates them into the state S t , S t includes temperature, humidity, predicted charging efficiency, actual discharging efficiency, and battery internal resistance information.
[0117] Based on the Q-table, determine the preliminary set of optimal actions. The agent queries the Q-table to find the actions in the current state St Actions with a high Q-value below are combined to form a preliminary optimal action set A optimal ;
[0118] For example: If the current state is S t and Q(S t , A1) = 0.2, Q(S t , A2) = 0.2, then the preliminary optimal action set A optimal = A1;
[0119] Specifically, these actions are based on the agent's existing experience and are considered to be able to obtain better expected cumulative rewards in the current state;
[0120] Input the current state S t into the recurrent neural network, and generate a series of interfering actions A adversarial based on the recurrent neural network. The purpose of these actions is to mislead the agent into choosing sub-optimal actions, thereby increasing the diversity and exploratory nature of action selection. For example, the interfering actions output by the recurrent neural network include some parameter adjustments or policy changes that are different from the agent's conventional experience;
[0121] Specifically, after executing action A t , the environment feedback reward r t = 0.3, and it transfers to the new state S t+1 . Use the Q-learning formula to update the Q-value:
[0122] Q(S t , A t ) = Q(S t , A t ) + αr t + γmax A Q(S t+1 , A t ), A - Q(S t , A t )
[0123] Among them, assume that α is the learning rate and the learning rate is 0.2, γ is the discount factor, γ = 0.9, and max t+1 Q(S A , A t+1 ) = 0.4 in the new state S t . Then after the update:
[0124] Q(S t , A1) = 0.2 + 0.2×[0.3 + 0.9×0.4 - 0.2] = 0.292;
[0125] Then from the set of interfering actions A provided by the recurrent neural network adversarialAmong them, a movement is randomly selected as the adversarial movement according to a uniform distribution;
[0126] The selected adversarial movement is evaluated together with the movements in the preliminary optimal movement set A optimal ;
[0127] Specifically, the evaluation method is to query the Q-table again to obtain the Q value of each movement;
[0128] According to the evaluation result, the movement with the highest score is selected as the movement to be finally executed
[0129] For example, if in the comprehensive evaluation, the score of the adversarial movement is higher than all the movements in the preliminary optimal movement set, then otherwise, the movement with the highest score is selected from the preliminary optimal movement set as
[0130] Through multiple iterations (such as 1000 interference movements), the Q-table is continuously updated. When the Q value of the interference movement converges (the change in the Q value is less than 0.01 for 50 consecutive iterations), at this time, all the movement combinations that maximize Q(S, A) in each state in the Q-table are the optimal strategy allocation scheme;
[0131] Finally, A1 is selected in state S1 and A2 is selected in state S2 to form an optimal control strategy for different battery states and environments;
[0132] For example: State S1 (high temperature, high SOC interference movement): Select A1 (reduce the charging power by 50% interference movement) to reduce the loss of the battery due to high temperature and high load;
[0133] State S2 (room temperature, moderate SOC interference movement): Select A2 (maintain the normal charging power) to improve the charging efficiency while ensuring the battery life.
[0134] Embodiment 2
[0135] As Figure 2 shown, an intelligent energy storage control method for a unit battery, and the steps of this method are:
[0136] S1: Collect the battery performance parameters and environmental parameters of the battery pack. The battery performance parameters include the charging efficiency, discharge efficiency, and battery internal resistance, and the environmental parameters include the temperature change parameter and humidity change parameter;
[0137] S2: Construct an environmental mapping model through an optimized and adjusted deep neural network, use the battery performance parameters and environmental parameters as the model input, and output the mapping influence coefficient;
[0138] S3: Based on the output of the environmental mapping model and the battery performance feedback data, construct a reinforcement learning framework as the state and reward of reinforcement learning, and optimize the selection strategy of the reinforcement learning framework to obtain an optimal policy allocation scheme.
[0139] In the application, several formulas involved are calculated by taking their numerical values after dimensionlessization. The establishment of the formulas is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. Some coefficients or weights in the formulas are set by those skilled in the art according to the actual situation, so no more details will be given here.
[0140] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0141] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent energy storage control system for a unit battery, characterized in that, Including: Data acquisition module: Collect battery performance parameters and environmental parameters of the battery pack; Among them, the battery performance parameters at least include charging efficiency, discharging efficiency and battery internal resistance, and the environmental parameters at least include temperature change parameters and humidity change parameters; Environmental mapping module: Construct an environmental mapping model through the optimized deep neural network, use the battery performance parameters and environmental parameters as the model input data, and output the mapping influence coefficient; Intelligent allocation module: Based on the output of the environmental mapping model and the battery performance feedback data as the state and reward of reinforcement learning, construct a reinforcement learning framework, and optimize the selection strategy of the reinforcement learning framework to obtain the optimal strategy allocation scheme; Among them, the optimized deep neural network dynamically adjusts the number of hidden layer nodes by obtaining the joint entropy of the model input data; the selection strategy of the optimized reinforcement learning framework is realized by combining the ∈-greedy strategy with a recurrent neural network.
2. The intelligent energy storage control system of a unit battery according to claim 1, characterized in that, The charging efficiency is obtained by measuring the input energy and the effectively stored energy during the battery charging process; the discharging efficiency is obtained by measuring the output energy and the discharging energy during the discharging time through a power sensor; the battery internal resistance is obtained by an internal resistance measurement sensor injecting a small AC signal with a known frequency and amplitude into the battery and measuring the response voltage at both ends of the battery; The temperature change parameters are obtained by thermocouple temperature sensors distributed at different positions of the battery pack to convert the temperature change into a voltage signal, and are obtained by measuring the voltage value and combining with the calibration curve; the humidity change parameters are obtained by a capacitive humidity sensor placed in the environment where the battery pack is located, measuring the change of the capacitance value with the environmental humidity and converting it through a signal processing circuit.
3. The intelligent energy storage control system for a unit battery according to claim 1, characterized in that The process of constructing the environmental mapping model is as follows: The environmental mapping model includes an input layer, an optimized hidden layer and an output layer; The number of input layer nodes is determined according to the types of collected environmental parameters and battery performance parameters; The number of nodes in the optimized hidden layer is set to twice the number of nodes in the input layer, and a three-layer hidden layer structure is adopted. Each hidden layer uses the ReLU activation function and is adjusted by the joint entropy dynamic adjustment strategy; The number of output layer nodes is obtained according to the types of predicted battery mapping data, and linear activation is adopted.
4. The intelligent energy storage control system for a unit battery according to claim 1, characterized in that The joint entropy dynamic adjustment strategy includes: Dividing the model input data into data groups according to a time window; Calculate the joint entropy H(z) of each group of data, where z represents the joint data, z = [η charge , η discharge , R, T, H]; When the joint entropy H(z) > H upper increase the hidden layer by 10% of the current number of nodes; When the joint entropy H(z) < H lower reduce the hidden layer by 5% of the current number of nodes; When the joint entropy H lower ≤H(z)≤H upper Keep the number of hidden layer nodes unchanged.
5. An intelligent energy storage control system for a unit battery according to claim 1, characterized in that, The battery performance feedback data at least includes the actual charging efficiency, the actual discharging efficiency and the battery life change rate.
6. The intelligent energy storage control system of a unit battery according to claim 1, wherein In the reinforcement learning framework, the output of the environmental mapping model and the battery performance feedback data are used as the state and reward signals of reinforcement learning: The reinforcement learning framework includes an agent, an environment V and a state S; Among them, the agent controls the environmental mapping model and the weight allocation; the environment V at least includes the charging and discharging processes of the battery at different times, the physical environment where it is located, and the environmental mapping model itself; the state S is jointly composed of the output of the environmental mapping model and the battery performance feedback data.
7. The intelligent energy storage control system for a unit battery according to claim 6, characterized in that, The agent selects action A according to state S, using the ∈-greedy strategy in combination with a recurrent neural network. Based on the existing optimal actions and with a specific probability, it accepts the interfering actions given by the recurrent neural network and comprehensively evaluates them with the actions in the preliminary optimal action set to finally determine the action to be executed.
8. The intelligent energy storage control system of a unit battery according to claim 7, characterized in that, The process of using the ∈-greedy strategy combined with a recurrent neural network to select actions is as follows: Create a table of the expected cumulative reward Q(S, A) for performing different actions A in different states S, initialize it, and input the current state S through a recurrent neural network t , and output a series of candidate interference actions A adversarial ; At time step t, combine the output data of the environmental mapping model and the battery performance feedback data into state S t ; Query the Q-table based on the agent to find the current state S t The actions with the specified Q-values are combined to form a preliminary optimal action set A optimal ; From the set of interfering actions A provided by the recurrent neural network adversarial randomly select an action according to a uniform distribution as the adversarial action The selected adversarial action is evaluated together with the actions in the preliminary optimal action set A optimal ; According to the evaluation results, select the action with the highest score as the action to be finally executed 9. An intelligent energy storage control method for a unit battery, using the system according to any one of claims 1 to 8, characterized in that, The steps of this method are: S1: Collect the battery performance parameters and environmental parameters of the battery pack. The battery performance parameters include charging efficiency, discharging efficiency, and battery internal resistance. The environmental parameters include temperature change parameters and humidity change parameters; S2: Construct an environmental mapping model through an optimized deep neural network. Use the battery performance parameters and environmental parameters as model inputs and output mapping influence coefficients; S3: Based on the output of the environmental mapping model and the battery performance feedback data, construct a reinforcement learning framework as the state and reward of reinforcement learning, and optimize the selection strategy of the reinforcement learning framework to obtain an optimal policy allocation scheme.
Citation Information
Patent Citations
Cluster system intelligent operation and maintenance system based on reinforcement learning
CN119090492A
Intelligent optimization control system and method for battery material recovery and preparation process
CN119395996A
Energy storage system power distribution method based on battery health state prediction
CN119651820A
Intelligent fast-charging charger control system based on real-time temperature feedback
CN119651838A
System control method for hybrid electric vehicles based on deep reinforcement learning
JP7624188B1