A method for constructing a control model of an intelligent liquid flow battery, a control method and device

By using a reinforcement learning-based intelligent flow battery control model, the problem of the inability of existing flow battery control methods to accurately control multiple parameters is solved, achieving efficient operation and long-term optimization of the flow battery, and improving battery performance and reliability.

CN119812398BActive Publication Date: 2025-11-28BEIJING HERUI ENERGY STORAGE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411871556.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-11-28
Estimated Expiration
2044-12-18

Smart Images

  • Figure CN119812398B_ABST
    Figure CN119812398B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of liquid flow batteries, and provides a method and device for constructing and controlling an intelligent liquid flow battery control model, which comprises obtaining various parameters of the operating state of a liquid flow battery, and constructing a state space and an action space based on the preprocessed parameters and a preset positive and negative electrode pump frequency; calculating an immediate reward based on a pre-defined reward function; constructing a state transition matrix according to the state space, the action space, the immediate reward and a next state space; wherein the next state space is obtained through the state space and the action space; training a control model based on the state transition matrix and a pre-defined optimization objective function to obtain a constructed control model; and calculating an immediate reward through a pre-defined reward function, so that the control model can accurately evaluate the effect of the current action at each time step and provide immediate feedback.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of flow battery, and particularly relates to an intelligent flow battery control model construction method, a control method and a device. BACKGROUND

[0002] Flow battery is a new type of electrochemical energy storage device, which has the characteristics of high energy conversion efficiency, environmental friendliness, good safety performance and deep discharge. However, the performance of the flow battery is affected by many factors, such as electrolyte temperature, tank and pipeline pressure, stack flow, tank liquid level, stack voltage, stack current, pump frequency feedback, cumulative charge capacity, cumulative discharge capacity, and charge-discharge efficiency. Therefore, how to effectively control these parameters to improve the performance of the flow battery is an important problem faced by the current flow battery technology.

[0003] The existing flow battery control method mainly sets the parameters manually or uses a simple control algorithm, such as PID control, to adjust the running state of the flow battery. However, these methods often cannot achieve accurate control of the flow battery, nor can they achieve coordinated optimization of multiple parameters such as electrolyte temperature, tank and pipeline pressure, stack flow, tank liquid level, stack voltage, stack current, pump frequency feedback, cumulative charge capacity, cumulative discharge capacity, and charge-discharge efficiency, which limits the performance of the battery. SUMMARY

[0004] To solve the problems in the background art, the application provides an intelligent flow battery control model construction method, a control method and a device.

[0005] To achieve the above purpose, the application adopts the following technical solutions:

[0006] In a first aspect, the present disclosure provides an intelligent flow battery control model construction method based on reinforcement learning, comprising:

[0007] Obtaining parameters of the running state of the flow battery, and constructing a state space and an action space based on the pre-processed parameters and a pre-set positive and negative electrode pump frequency;

[0008] Calculating an immediate reward based on a pre-defined reward function; wherein the reward function includes a reward function in each time step; the immediate reward is calculated based on the reward function in each time step;

[0009] Constructing a state transition matrix according to the state space, the action space, the immediate reward and the next state space; wherein the next state space is obtained by the state space and the action space;

[0010] training the control model based on the state transition matrix and a predefined optimization objective function, to obtain a constructed control model.

[0011] Preferably, the method of preprocessing the parameters is normalization.

[0012] Preferably, the method further comprises: taking a restriction measure on the state space.

[0013] Preferably, the reward function further comprises a reward function of each charge-discharge cycle;

[0014] The method further comprises:

[0015] calculating a total reward through the reward function of each charge-discharge cycle;

[0016] evaluating and adjusting the model based on the total reward.

[0017] Preferably, the method further comprises: folding the state space on a time line.

[0018] Preferably, the method further comprises: training the control model through a deep learning network.

[0019] Preferably, the training of the control model based on the state transition matrix and a predefined optimization objective function comprises:

[0020] cutting the state transition matrix into sub-matrices;

[0021] randomly selecting the sub-matrices, training the control model through a backtracking learning method and a predefined optimization objective function.

[0022] In a second aspect, the present disclosure provides an intelligent flow battery construction device based on reinforcement learning, comprising:

[0023] a state space and action space construction module, configured to acquire various parameters of a flow battery operation state, and construct a state space and an action space based on time steps based on the preprocessed parameters and a preset positive and negative electrode pump frequency;

[0024] an immediate reward construction module, configured to calculate an immediate reward based on a predefined reward function;

[0025] a state transition matrix construction module, configured to construct a state transition matrix according to the state space, the action space, the immediate reward, and a next state space; wherein the next state space is obtained through the state space and the action space;

[0026] a control model construction module, configured to train a control model based on the state transition matrix and a predefined optimization objective function, to obtain a constructed control model.

[0027] In a third aspect, the disclosure provides an intelligent flow battery control method based on reinforcement learning, comprising:

[0028] obtaining parameters of the operating state of the flow battery; wherein the parameters include: electrolyte temperature entering the stack, pressure entering the stack, flow entering the stack, tank level, pump frequency feedback, stack voltage and stack current;

[0029] inputting the parameters into the control model constructed in the first aspect to obtain the positive and negative electrode pump frequency given;

[0030] adjusting the pump frequency of the flow battery according to the positive and negative electrode pump frequency given.

[0031] In a fourth aspect, the disclosure provides an intelligent flow battery control device based on reinforcement learning, comprising:

[0032] a parameter acquisition module for obtaining parameters of the operating state of the flow battery; wherein the parameters include: electrolyte temperature entering the stack, pressure entering the stack, flow entering the stack, tank level, pump frequency feedback, stack voltage and stack current;

[0033] a parameter analysis module for inputting the parameters into the control model constructed in the first aspect to obtain the positive and negative electrode pump frequency given;

[0034] an execution module for adjusting the pump frequency of the flow battery according to the positive and negative electrode pump frequency given.

[0035] The beneficial effects of the present application are:

[0036] 1. The method of the present application calculates the immediate reward through the pre-defined reward function, so that the control model can accurately evaluate the effect of the current action at each time step and provide immediate feedback. Not only can it help the control model quickly adjust the strategy, but also it considers multiple performance indicators (such as charging capacity, discharging capacity, fault number, warning number, etc.), avoiding the limitations brought by single indicator optimization, thereby improving the learning efficiency and decision accuracy.

[0037] 2、The method of the present application records the state, action, reward and next time step state of each time step through the construction of the state transition matrix, forming a complete state transition history, not only simplifying the complex problem, making the control model easier to understand and processing system behavior, but also enhancing its dynamic adaptability and generalization ability, enabling the control model to dynamically adapt to changes in the system and timely adjust the strategy to ensure efficient operation under different working conditions. Through the synergistic effect of immediate reward and state transition matrix, the control model not only focuses on short-term effects, but also considers long-term impacts, achieving global optimal solution, thereby significantly improving the performance and reliability of the flow battery energy storage system.

[0038] 3、The method of the present application can realize real-time monitoring of the running state of the flow battery and dynamically adjust the pump frequency by real-time monitoring of the key operating parameters of the flow battery (such as electrolyte temperature, tank and pipeline pressure, stack inlet flow, tank level, stack voltage, stack current, pump frequency feedback, cumulative charge, cumulative discharge and charge-discharge efficiency) and analyzing these operating parameters with the trained control model, which can achieve maximum system-level charge-discharge efficiency, thereby improving the energy utilization efficiency of the flow battery, reducing energy waste, and improving the energy storage capacity of the flow battery, which can achieve maximum balanced charge-discharge efficiency and discharge capacity, thereby improving the service life of the flow battery and reducing maintenance costs.

[0039] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and obtained by the structure indicated in the specification and drawings. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0041] Figure 1 A flow chart of an intelligent flow battery control model construction method based on reinforcement learning of the present application is shown;

[0042] Figure 2 A framework diagram of an intelligent flow battery construction device based on reinforcement learning of the present application is shown;

[0043] Figure 3 A flow chart of an intelligent flow battery control method based on reinforcement learning of the present application is shown;

[0044] Figure 4 A reinforcement learning-based intelligent flow battery control device framework of the application is shown;

[0045] Figure 5 A device structure schematic diagram of the application is shown. DETAILED DESCRIPTION

[0046] To make the objects, technical solutions and advantages of the embodiments of the application clearer, the technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are some but not all of the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the application.

[0047] Referring to Figure 1 A reinforcement learning-based intelligent flow battery control model construction method is shown, and specifically includes the following steps:

[0048] S11, acquiring various parameters of the flow battery operating state, and constructing a state space and an action space based on time steps based on the preprocessed parameters and a preset positive and negative electrode pump frequency given;

[0049] Specifically, the parameters include: stack inlet electrolyte temperature (T), stack inlet pressure (P), stack inlet flow rate (F), tank liquid level (L), pump frequency feedback (QB), stack voltage (V), and stack current (C). It can be understood that these parameters can be historical data extracted from past operation records, or current data collected in real time through a real-time monitoring system.

[0050] Let po represent the positive electrode of the system, ne represent the negative electrode of the system, and t represent the time step (t e N+, the time granularity of the time step depends on the shortest period of data acquisition), then:

[0051] The electrolyte inlet temperature (T) uses two temperature measurement point data of the electrolyte positive and negative main pipes in a set of process systems as state evaluation points, that is:

[0052] The stack inlet pressure (P) uses two pressure measurement point data of the electrolyte positive and negative main pipes in a set of process systems and their difference as state evaluation points, that is: Wherein, represents the pressure difference;

[0053] The stack inlet flow rate (F) uses two flow rate measurement point data of the electrolyte positive and negative main pipes in a set of process systems as state evaluation points, that is:

[0054] The liquid level (L) of the stack storage tank uses the data of the liquid level measuring points of the positive and negative electrolyte storage tanks in a set of process systems as the state evaluation points, namely:

[0055] The pump frequency feedback (QB) uses the data of the liquid level measuring points of the positive and negative electrolyte storage tanks in a set of process systems as the state evaluation points, namely:

[0056] The stack voltage (V) is the real-time voltage of each stack, and the total number of stacks in a set of process systems is k, so the voltage state evaluation point of a set of process systems is:

[0057] The stack current (C) is the real-time current of each stack, and the total number of stacks in a set of process systems is k, so the voltage state evaluation point of a set of process systems is:

[0058] The method for preprocessing the parameters is data normalization. Specifically, to eliminate the influence of the original data magnitude on the algorithm, a normalization algorithm is used for dimensionless, that is, data normalization, and its expression is: In the formula, x represents the real-time value of any state evaluation point, L scal represents the lower limit of the evaluation point range (reference to the instrument range of the collection point), H scal represents the upper limit of the evaluation point range (reference to the instrument range of the collection point), X scal represents the result of x data normalization.

[0059] After preprocessing, the state space based on time steps is:

[0060]

[0061] In the formula, scal represents data normalization processing, and t represents time steps.

[0062] The action space (A) of a set of liquid flow battery energy storage process systems includes: positive and negative pump frequency given (QS). Let po represent the positive electrode of the system, ne represent the negative electrode of the system, and t represent the time step (t∈N+, consistent with the time step of the state space). The action space based on time steps is:

[0063]

[0064] In the formula, represents the lower limit of the positive and negative pump frequency given (reference to the pump frequency control output range), represents the upper limit of the positive and negative pump frequency given (reference to the pump frequency control output range), and the positive and negative pump frequency control output range remains consistent.

[0065] As a preferred embodiment of this disclosure, to prevent excessively large differences between pump frequency control output time steps, or pump frequency settings that are too low leading to inability to charge and discharge, or pump frequency settings that are too high leading to pump damage, it is necessary to impose restrictions on the behavior space, that is, to impose restrictions on the control of the pump frequency for each time step, the expression of which is:

[0066]

[0067] In the formula, a is the given maximum allowable pump frequency; d is the given minimum allowable pump frequency; b is the given maximum inter-step pump frequency difference; and c is the given minimum inter-step pump frequency difference. A t The algorithm calculates the pump frequency value for the current time step t. This represents the actual pump frequency value executed at time step t-1. Calculate the pump frequency value for the current time step t under constraints. This represents the final pump frequency value at the current time step t.

[0068] Specifically, the meaning of the expression for the restrictive measures is as follows:

[0069] If the pump frequency calculation value A at the current time step t t Less than the minimum allowable pump frequency value d, or the current If it is less than the minimum allowable pump frequency value d, then Let's set it to d. This ensures that the pump frequency doesn't get too low, preventing situations where charging and discharging fail.

[0070] If the pump frequency calculation value A at the current time step t t Greater than the maximum allowable pump frequency value 'a', or the current If it is greater than the maximum allowable pump frequency value 'a', then... Let's set it as 'a'. This ensures that the pump frequency won't be too high, preventing damage to the system.

[0071] If the pump frequency calculation value A at the current time step t t Pump frequency execution value compared to the previous time step t-1 If the difference is less than -, then Reducing b can prevent a sudden and excessive drop in pump frequency, which could damage the stability and lifespan of the pump itself, pipelines, and fuel cell stack.

[0072] If the pump frequency calculation value A at the current time step t t Pump frequency execution value compared to the previous time step t-1 If the difference is greater than b, then... Adding b can prevent the pump frequency from rising too suddenly, which could damage the stability and lifespan of the pump itself, pipelines, and fuel cell stack.

[0073] If the pump frequency calculation value A at the current time step t tThe pump frequency setpoint compared to the previous time step t-1 If the difference is between - and c, then keep The difference in pump frequency between steps remains unchanged. If the difference is too small, its impact on the system is negligible, and its effect on algorithm optimization is also negligible. This constraint ensures the stability of the algorithm and reduces unnecessary operational overhead.

[0074] If the pump frequency calculation value A at the current time step t t Pump frequency execution value compared to the previous time step t-1 If the difference is between [-, -c] or [c, b], then execute A. t Pump frequency values ​​within this range do not affect system safety and stability, require no constraint modifications, and can be executed directly. They represent the direct and effective action execution steps of the system and algorithm.

[0075] S12. Calculate the immediate reward based on a predefined reward function;

[0076] Specifically, the reward function (R) of a flow battery energy storage system is defined as follows: under a single charge-discharge cycle, the system's maximum capacity (MaxC), cumulative AC-side charge (CG), cumulative AC-side discharge (DCG), AC-side charge-discharge efficiency (EF), AC-side system efficiency (EP) considering pump energy consumption, and related functions for system alarms (ALM) and faults (ER) are calculated. The AC-side charge (CG) and discharge (DCG) of the energy storage converter are obtained from energy storage converter statistics or meter readings, while pump energy consumption is obtained from the total power supply meter of the pump system. The charge and discharge amounts within a single cycle are normalized.

[0077]

[0078] In the formula, t represents the time step, and CG t -CG t-1 DCG t -DCG t-1 The actual charging and discharging amounts of the system during time steps are normalized to the (0,1) interval based on the ratio of these amounts to the system's maximum capacity MaxC.

[0079] The AC side charge / discharge efficiency (EF) of the energy storage converter is calculated as follows:

[0080]

[0081] Where τ represents the charge / discharge cycle step (τ∈N+, each τ contains multiple time steps t), the charge / discharge amount is calculated based on the cumulative charge / discharge difference at the end of each charge / discharge cycle and at the end of the next cycle, and the charge / discharge efficiency EF under that cycle is also calculated. τ .

[0082] The AC side system efficiency (EP) calculation method considering pump energy consumption is:

[0083]

[0084] wherein, represents the pump energy consumption at the end of the current charge-discharge cycle, represents the pump energy consumption at the end of the previous charge-discharge cycle. The ratio of the charge amount to the discharge amount plus the pump energy consumption is used as the AC side system efficiency EP τ .

[0085] System alarms and faults are set according to the number, with a maximum of M system alarms and K system faults, normalized to the (0, 1) interval, and the calculation method is:

[0086]

[0087] wherein, m t , k t represents the actual number of alarms and faults generated by the system at time step t

[0088] Since efficiency calculation can only be performed in a complete charge-discharge cycle, a charge-discharge cycle contains several time steps, and the system selects an action from the action space to execute at each time step, a relatively immediate reward function is needed to evaluate the pros and cons of the selection, so as to guide the selection and execution of the next action, therefore, the reward function calculation method is as follows:

[0089]

[0090] wherein, R t is the reward function at each time step, R τ is the reward function of each charge-discharge cycle, is a weight coefficient, by setting different weight coefficient values, the key focus indicators of the reward function and the balance relationship of each indicator quantity can be adjusted, and T is the total number of time steps contained in each charge-discharge cycle.

[0091] Specifically, the meaning of the expression of the reward function is: the immediate reward is calculated by R t , and at each step of execution, R t is used as an optimization function, and R t is accumulated in a complete charge-discharge cycle to obtain the total reward, and after the cycle is completed, the total reward is assigned to R t , which is brought into the next charge-discharge cycle for iterative optimization calculation. The total reward is used to evaluate the performance of the entire charge-discharge cycle and serves as a reference to adjust and evaluate the model parameters.

[0092] S13, constructing a state transition matrix according to the state space, the action space, the immediate reward, and the next state space; wherein the next state space is obtained by the state space and the action space;

[0093] Specifically, taking the state space (S) as the input, the action space (A) as the state transition branch point, the next state space (S_), and the reward function immediate calculation (R) as the state transition matrix:

[0094] ((S0,A0,R0,S1),(S1,A1,R1,S2),(S2,A2,R2,S3),…)

[0095] Wherein, the next state space is determined according to the current state S t and the action A t taken.

[0096] As a preferred embodiment of the present disclosure, the state space can be folded on the time line to optimize the problem of insufficient state change caused by slowly changing physical quantities such as pressure and temperature, specifically:

[0097]

[0098] In the formula, y represents the length of the time step folding.

[0099] By combining the states of multiple consecutive time steps into one folded state The change trend of slowly changing physical quantities (such as pressure and temperature) can be more comprehensively captured.

[0100] S14, training a control model based on the state transition matrix and a pre-defined optimization objective function to obtain a constructed control model.

[0101] Specifically, the control model is trained by a deep reinforcement learning algorithm, for example: Q-learning, Deep QNetwork, Double DeepQNetwork, Dueling Deep QNetwork, Deep DeterministicPolicy Gradient, Advantage Actor Critic, Proximal Policy Optimization, Distributed Proximal Policy Optimization, etc.

[0102] The present disclosure takes the minimization of the reward function as the optimization objective of the algorithm, aiming to maximize the charging capacity, discharging capacity, AC side efficiency, pump consumption AC side efficiency, and minimize faults and warnings by optimizing the control strategy, so as to achieve the purpose of optimal control, and taking into account the impact of any behavior on the future, which is expressed as:

[0103]

[0104] In the formula, η represents the discount factor.

[0105] The method of the present disclosure calculates the immediate reward through the pre-defined reward function, so that the control model can accurately evaluate the effect of the current action at each time step and provide immediate feedback, which not only helps the control model to quickly adjust the strategy, but also comprehensively considers multiple performance indicators (such as charging capacity, discharging capacity, number of faults, number of warnings, etc.), avoiding the limitations brought by single indicator optimization, thereby improving the learning efficiency and decision accuracy. And it also constructs a state transition matrix according to the state space, action space, immediate reward and next state space, which records the state, action, reward and next state at each time step, forming a complete state transition history, which not only simplifies complex problems, making it easier for the control model to understand and process system behavior, but also enhances its dynamic adaptability and generalization ability, enabling the control model to dynamically adapt to system changes and adjust strategies in a timely manner to ensure efficient operation under different working conditions. Through the synergistic effect of immediate reward and state transition matrix, the control model not only focuses on short-term effects, but also considers long-term impacts, achieving global optimal solution, thereby significantly improving the performance and reliability of the flow battery energy storage system.

[0106] The present disclosure supports the normalization of other relevant indicators not listed herein to the (0, 1) interval and incorporates them into the reward function through a weight system, thereby achieving the optimization of these indicators. For example, parameters such as positive and negative electrode inlet pressure balance and positive and negative electrode tank internal pressure balance can be optimized. Specifically, after these additional indicators are normalized using pressure average variance, they are used as part of the reward function, which can help the model more comprehensively evaluate and optimize the overall performance of the system.

[0107] In the present disclosure, all time steps t referred to are completely identical, and all charging and discharging cycles τ referred to are completely identical. However, the actual time intervals are not specified to be completely identical, and the present disclosure is applicable to any different actual time intervals of time steps and charging and discharging cycles.

[0108] Referring to Figure 2 Based on the same inventive concept, the present disclosure further proposes an intelligent flow battery construction device based on reinforcement learning, which comprises:

[0109] The state space and action space construction module 110 is configured to acquire parameters of a running state of the flow battery, and construct a time step-based state space and action space based on the preprocessed parameters and a preset positive and negative electrode pump frequency;

[0110] The instant reward construction module 120 is configured to calculate an instant reward based on a predefined reward function;

[0111] The state transition matrix construction module 130 is configured to construct a state transition matrix according to the state space, the action space, the instant reward and a next state space, wherein the next state space is obtained from the state space and the action space;

[0112] The control model construction module 140 is configured to train a control model based on the state transition matrix and a predefined optimization objective function, to obtain a constructed control model.

[0113] Referring to Figure 3 The present disclosure also provides an intelligent flow battery control method based on reinforcement learning, which comprises the following steps:

[0114] S21, acquiring parameters of a running state of the flow battery, wherein the parameters comprise electrolyte temperature entering a stack, pressure entering the stack, flow entering the stack, a tank liquid level, pump frequency feedback, a stack voltage and a stack current;

[0115] S22, inputting the parameters into the constructed control model to obtain a positive and negative electrode pump frequency setpoint;

[0116] S23, adjusting pump frequency of the flow battery according to the positive and negative electrode pump frequency setpoint, to maximize system-level charge and discharge efficiency, maximize discharge capacity, and maximize balanced charge and discharge efficiency and discharge capacity.

[0117] Referring to Figure 4 Based on the same inventive concept, the present disclosure also provides an intelligent flow battery control device based on reinforcement learning, which comprises the following steps:

[0118] The parameter acquisition module 210 is configured to acquire parameters of a running state of the flow battery, wherein the parameters comprise electrolyte temperature entering a stack, pressure entering the stack, flow entering the stack, a tank liquid level, pump frequency feedback, a stack voltage and a stack current;

[0119] The parameter analysis module 220 is configured to input the parameters into the constructed control model to obtain a positive and negative electrode pump frequency setpoint;

[0120] The execution module 230 is configured to adjust pump frequency of the flow battery according to the positive and negative electrode pump frequency setpoint.

[0121] By monitoring the key operating parameters of the flow battery in real time (such as electrolyte temperature, tank and pipeline pressure, stack inlet flow, tank liquid level, stack voltage, stack current, pump frequency feedback, cumulative charging capacity, cumulative discharging capacity, and charging and discharging efficiency), and combining the trained control model to analyze these operating parameters, the real-time monitoring of the flow battery operating state and the dynamic adjustment of the pump frequency can be realized, the system-level charging and discharging efficiency can be maximized, the energy utilization efficiency of the flow battery is improved, the energy waste is reduced, and the energy storage capacity of the flow battery is improved, the balanced charging and discharging efficiency and discharging capacity can be maximized, and the service life of the flow battery is improved, and the maintenance cost is reduced.

[0122] Referring to Figure 5 As shown in the drawings, based on the same inventive concept, the present disclosure also proposes a device comprising a memory and a processor, wherein the memory stores computer instructions capable of running on the processor, and the processor executes the computer instructions to perform the above-mentioned intelligent flow battery control model construction method and / or control method based on reinforcement learning.

[0123] Based on the same inventive concept, the present disclosure also proposes a computer readable storage medium having computer instructions stored thereon, wherein when the computer instructions are run, the above-mentioned intelligent flow battery control model construction method and / or control method based on reinforcement learning can be implemented.

[0124] Wherein, any reference to memory, storage, database or other medium used in each embodiment of the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory.

[0125] It should be noted that in this paper, relationship terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device including the element.

[0126] Although the present application has been described in detail with reference to the foregoing embodiments, it should be understood that modifications can be made to the foregoing embodiments, or additional implementations of the present application can be implemented, without departing from the spirit or scope of the application. Accordingly, the present application is not limited except as by the appended claims.

Claims

1. A method for constructing a control model for an intelligent flow battery based on reinforcement learning, characterized in that, include: S1. Obtain various parameters of the operating state of the flow battery, and construct a state space and behavior space based on the preprocessed parameters and the preset positive and negative electrode pump frequencies. The parameters include: electrolyte temperature at the fuel cell stack, pressure at the fuel cell stack, flow rate at the fuel cell stack, tank level, pump frequency feedback, fuel cell stack voltage, and fuel cell stack current; the behavior space includes positive and negative pump frequency settings. S2. Calculate the instant reward based on a predefined reward function; wherein, the reward function includes a reward function within each time step; the instant reward is calculated based on the reward function within each time step; the predefined reward function is the cumulative charging amount on the AC side of the maximum capacity energy storage converter, the cumulative discharging amount on the AC side of the energy storage converter, the charging and discharging efficiency on the AC side of the energy storage converter, the AC side efficiency taking into account pump energy consumption, and related functions for alarms and faults under a single charge and discharge cycle; S3. Construct a state transition matrix based on the state space, behavior space, immediate reward, and next state space; wherein, the next state space is obtained through the state space and the behavior space; S4. Train the control model based on the state transition matrix and minimize the reward function to obtain the constructed control model.

2. The method for constructing a smart flow battery control model based on reinforcement learning according to claim 1, characterized in that, The method for preprocessing the parameters is normalization.

3. The method for constructing a smart flow battery control model based on reinforcement learning according to claim 1, characterized in that, The step S2 is preceded by: taking restrictive measures on the behavior space.

4. The method for constructing a smart flow battery control model based on reinforcement learning according to claim 1, characterized in that, The reward function also includes a reward function for each charge-discharge cycle; Before step S3, the following also includes: The total reward is calculated using the reward function for each charge-discharge cycle; Based on the aforementioned total reward evaluation and adjustment model.

5. The method for constructing a smart flow battery control model based on reinforcement learning according to claim 1, characterized in that, The step S4 is preceded by: folding the state space on the timeline.

6. The method for constructing a smart flow battery control model based on reinforcement learning according to claim 1, characterized in that, Step S4 is followed by training the control model using a deep learning network.

7. The method for constructing a smart flow battery control model based on reinforcement learning according to claim 1, characterized in that, The training of the control model based on the state transition matrix and minimizing the reward function includes: The state transition matrix is ​​divided into several sub-matrices by segmentation. The submatrix is ​​randomly selected, and the control model is trained using a backtracking learning method and minimizing the reward function.

8. A reinforcement learning-based intelligent flow battery control model construction apparatus for implementing the reinforcement learning-based intelligent flow battery control model construction method according to any one of claims 1-7, characterized in that, include: The state space and behavior space construction module is used to acquire various parameters of the flow battery's operating state, and construct a time-step-based state space and behavior space based on the preprocessed parameters and the preset positive and negative electrode pump frequencies. The parameters include: electrolyte temperature at the fuel cell stack, pressure at the fuel cell stack, flow rate at the fuel cell stack, tank level, pump frequency feedback, fuel cell stack voltage, and fuel cell stack current; the behavior space includes positive and negative pump frequency settings. The instant reward construction module is used to calculate instant rewards based on a predefined reward function. The predefined reward function is the cumulative charging amount on the AC side of the maximum capacity energy storage converter, the cumulative discharging amount on the AC side of the energy storage converter, the charging and discharging efficiency on the AC side of the energy storage converter, the AC side efficiency taking into account pump energy consumption, and related functions for alarms and faults under a single charge and discharge cycle. A state transition matrix construction module is used to construct a state transition matrix based on the state space, behavior space, immediate reward, and next state space; wherein the next state space is obtained through the state space and the behavior space. The control model construction module is used to train the control model based on the state transition matrix and the minimization of the reward function to obtain the constructed control model.

9. A method for controlling an intelligent flow battery based on reinforcement learning, characterized in that, include: SA, acquire various parameters of the operating status of the flow battery; wherein, the parameters include: electrolyte temperature at the battery stack, pressure at the battery stack, flow rate at the battery stack, tank level, pump frequency feedback, battery stack voltage and battery stack current; SB, Input the parameters into the control model constructed according to any one of claims 1-7 to obtain the positive and negative pole pump frequency settings; SC. Adjust the pump frequency of the flow battery according to the given positive and negative electrode pump frequencies.

10. A smart flow battery control device based on reinforcement learning, characterized in that, include: The parameter acquisition module is used to acquire various parameters of the operating status of the flow battery; wherein, the parameters include: electrolyte temperature at the battery stack, pressure at the battery stack, flow rate at the battery stack, tank level, pump frequency feedback, battery stack voltage, and battery stack current; The parameter analysis module is used to input the parameters into the control model constructed according to any one of claims 1-7 to obtain the positive and negative pole pump frequency settings; The execution module is used to adjust the pump frequency of the flow battery according to the given positive and negative electrode pump frequencies.

Citation Information

Patent Citations

  • Hydrogen fuel cell monitoring method and device based on learning model

    CN117276592A

  • Ship multi-energy power system optimization reconstruction method, terminal equipment and medium

    CN118428232A