A method and device for coordinated optimization of operation strategies of a group of energy storage power stations

CN118674109BActive Publication Date: 2026-09-22TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410798959.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-20
Publication Date
2026-09-22
Estimated Expiration
2044-06-20

AI Technical Summary

Technical Problem

[0003]本发明用于解决现有技术中,具有多个被控系统的被控系统群的运行策略优化存在效率低及精度差的问题,未利用不同储能电站决策问题的相似性,未通过迁移学习等方法提升储能电站群的整体学习效率

Benefits of technology

[0013]本发明第四方面提供一种计算机可读存储介质,所述计算机可读存储介质存储有计算机程序,所述计算机程序被计算机设备的处理器执行时实现前述任意一实施例所述方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674109B_ABST
    Figure CN118674109B_ABST
Patent Text Reader

Abstract

The application relates to the field of operation strategy optimization, and provides an operation strategy cooperative optimization method and device for an energy storage power station group, which comprises the following steps: determining Q factor sample variances of each action, performance differences of Q factor fusion values of each optimal action and non-optimal action, and a total consumed sample amount according to simulation results; analyzing the above-mentioned amounts by using a sampling data distribution algorithm to obtain target sample amounts of each action; determining supplementary sampling amounts of each action according to the target sample amounts of each action and the consumed sample amounts of each action; performing supplementary sampling according to the supplementary sampling amounts, and re-determining the Q factor sample variances of each action, the performance differences, and the total consumed sample amount by using supplementary simulation results; adjusting the total consumed sample amount, judging whether the total consumed sample amount is smaller than a preset total sampling amount, recalculating the target sample amount and the following steps if yes, and outputting an optimal action if no. The application cooperatively uses operation data of energy storage equipment with action consistency, and can improve optimization efficiency and performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of operation strategy optimization for energy storage power station clusters in smart grids, and particularly to a method and apparatus for collaborative optimization of operation strategies for energy storage power station clusters. Background Technology

[0002] In existing technologies, there exist controlled system groups with multiple controlled systems, such as energy storage power stations in smart grids. These power stations consist of numerous energy storage devices, and their performance can be improved by rationally scheduling charging and discharging power. With the integration of a large number of new energy power generation systems into the grid, the stable and economical operation of energy storage power stations is becoming increasingly important. The optimization of existing energy storage power station operation strategies involves a large state space (e.g., the power demand of each electrical device, the energy storage level of the energy storage devices, and the power generation capacity of the power generation equipment) and an action space (e.g., the charging or discharging power of each energy storage device), making it difficult to solve accurately using methods such as reinforcement learning. Simplified models are generally used, and the operation strategy is usually based on the data and information of a single energy storage power station for optimization decisions. Since energy storage power stations located in different locations are relatively independent in the strategy optimization process, the learning time for each power station is very long, requiring a large amount of training data. Even general reinforcement learning methods cannot find a high-performance operation strategy after long-term operation, thus requiring manually formulated operation rules to guide the operation of the power station. Summary of the Invention

[0003] This invention addresses the problems of low efficiency and poor accuracy in optimizing the operation strategy of a group of controlled systems with multiple controlled systems in the prior art. It fails to utilize the similarity of decision-making problems of different energy storage power stations and fails to improve the overall learning efficiency of the energy storage power station group through methods such as transfer learning.

[0004] To address the aforementioned technical problems, this invention provides a method for collaborative optimization of the operation strategy of a controlled system group, wherein the controlled system group includes multiple controlled systems, and the method includes: The simulation results of each action of the target controlled system agent are sampled according to the preset sampling amount. Based on the sampling results, the Q factor sample variance of each action, the performance difference of the Q factor fusion value of each optimal action and non-optimal action, and the total amount of samples consumed are determined. The Q factor fusion value of each action is calculated by the Q factor fusion function, which is the Q factor function of the target controlled system agent and its similar controlled system agents under the same state-action. The sampling data allocation algorithm is used to analyze the Q-factor sample variance, performance difference and total amount of consumed samples of each action to obtain the target sample size of each action. The sampling data allocation algorithm includes the determination criteria for the sample size of various actions obtained by the Q-factor fusion function analysis, which is used to progressively maximize the probability of correctly selecting the optimal action. The additional sampling amount for each action is determined based on the target sample size and the consumed sample size for each action. Supplementary sampling is performed based on the supplementary sampling amount for each action, and the Q-factor sample variance, performance difference, and total amount of consumed samples for each action are re-determined using the simulation results of the supplementary sampling. Adjust the total number of consumed samples, and determine whether the total number of consumed samples is less than the preset total sampling amount. If so, recalculate the target sample amount and subsequent steps. If not, output the optimal action.

[0005] In a further embodiment of the present invention, the process of determining the sampling data allocation algorithm includes: Based on the equipment information, system status, control actions, and optimization target information of each controlled system, establish the performance function and action sampling constraints of the control strategy for each controlled system. The performance function of the control strategy of each controlled system is transformed to obtain the Q-factor fusion function of the intelligent agent of each controlled system; Based on the Q-factor fusion function of each controlled system agent, an objective function is established to approximate the optimal action selection probability function. Based on the objective function of the near-optimal action selection probability function and the action sampling amount constraint, the criteria for the sampling data allocation algorithm are determined.

[0006] In a further embodiment of the present invention, the transformation of the performance function of the control strategy of each controlled system to obtain the Q-factor fusion function of each controlled system agent includes: The Q-factor fusion function of the controlled system agent is expressed by the following formula: , ; in, Represents the intelligent agent of the controlled system i With the controlled system intelligent agent j In state s Next action a Q-factor fusion value; s Indicates state; a Indicates an action; This indicates the state and action of the controlled system's intelligent agent i. a The observed mean of the Q factor at that time; This indicates the state and action of the controlled system's intelligent agent j in s. a The observed mean of the Q-factor at time; S represents the state space; Let i represent the action space, where controlled system agent i is similar to controlled system agent j.

[0007] In a further embodiment of the present invention, the objective function for establishing an approximate optimal action selection probability function based on the Q-factor fusion function of each controlled system agent includes: The objective function of the approximate optimal action selection probability function can be expressed by the following formula: ; ; in, APCS This represents the probability function for selecting an approximate optimal action. b Indicates the optimal action; Indicates a non-optimal action; Represents the intelligent agent of the controlled system i With the controlled system intelligent agent j In state s Find the optimal action b Q-factor fusion value; Represents the intelligent agent of the controlled system i With the controlled system intelligent agent j Non-optimal action in state s The Q-factor fusion value indicates that the controlled system agent i is similar to the controlled system agent j.

[0008] In a further embodiment of the present invention, the sampling data allocation algorithm is used to indicate: After the controlled system agent fuses its data with similar controlled system agents, the larger the noise-to-signal ratio of the non-observed optimal action under the same state s, the larger the sampling amount of the non-observed optimal action. In the noise-to-signal ratio, noise refers to the sample standard deviation of the Q factor of the non-observed optimal action, and information refers to the performance difference of the fused Q factor values ​​of the observed optimal action and the non-observed optimal action. The number of samples for the best observed action should be greater than the number of samples for other non-best observed actions under the same condition.

[0009] In a further embodiment of the present invention, the sampling data allocation algorithm includes: Criterion 1: For unobserved optimal actions p and q, , ; in, This represents the sample size of the controlled system agent i when it is not observing the optimal action p; This represents the sample size of agent i in the controlled system when it is not the optimal action q observed. The sample standard deviation of the Q factor represents the Q factor of the controlled system agent i when the action p is not observed optimally. The sample standard deviation of the Q factor represents the Q factor of the controlled system agent i when the action q is not observed optimally. This represents the performance difference between the optimal action b and the non-optimal action p of the controlled system agent i and its similar controlled system agent j. This represents the performance difference between the optimal action b and the non-optimal action q of the controlled system agent i and its similar controlled system agent j. Criterion 2: For observation of optimal action b and non-optimal action b a , ; in, This represents the sampling amount of the controlled system agent i when observing the optimal action b; The sample standard deviation of the Q factor represents the Q factor of the controlled system agent i when observing the optimal action b. This indicates that the controlled system agent i is performing a non-optimal action. a Sampling volume at time; This indicates that the controlled system agent i is performing a non-optimal action. a The sample standard deviation of the Q factor.

[0010] In a further embodiment of the present invention, the target sample size for each action is obtained by analyzing the Q-factor sample variance, performance difference, and total amount of consumed samples for each action using a sampling data allocation algorithm, including: The sampling ratio of each action is obtained by analyzing the Q-factor sample variance and performance differences of each action based on the sampling data allocation algorithm. The number of samples that should be obtained for each action is obtained by multiplying the sampling ratio of each action by the total number of samples consumed; Choose the larger of the number of samples that should be obtained for each action and the number of samples that have actually been used for each action as the target number of samples for each action.

[0011] A second aspect of the present invention provides a device for collaborative optimization of the operation strategy of a controlled system group, the controlled system group including multiple controlled systems, the device comprising: The sampling calculation unit is used to sample the simulation results of each action of the target controlled system agent according to a preset sampling amount, and determine the Q factor sample variance of each action, the performance difference of the Q factor fusion value of each optimal action and non-optimal action, and the total amount of samples consumed based on the sampling results. The Q factor fusion value of each action is calculated by the Q factor fusion function, which is the Q factor function of the target controlled system agent and its similar controlled system agents under the same state-action. The target sampling quantity determination unit is used to analyze the Q-factor sample variance, performance difference and total amount of consumed samples of each action using a sampling data allocation algorithm to obtain the target sample quantity of each action. The sampling data allocation algorithm includes the determination criteria for the sampling quantity of various actions obtained by the analysis of the Q-factor fusion function, which is used to progressively maximize the probability of correctly selecting the optimal action. The supplementary sampling quantity determination unit is used to determine the supplementary sampling quantity for each action based on the target sample quantity for each action and the sample quantity consumed for each action. The supplementary sampling unit is used to perform supplementary sampling based on the supplementary sampling amount of each action, and to redetermine the Q-factor sample variance, performance difference and total amount of samples consumed for each action using the simulation results of the supplementary sampling. The analysis unit is used to adjust the total number of consumed samples and determine whether the total number of consumed samples is less than the preset total sampling amount. If so, the target sampling amount determination unit is used to recalculate the target sample amount. If not, the optimal action is output.

[0012] A third aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in any of the foregoing embodiments.

[0013] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor of a computer device, implements the method described in any of the foregoing embodiments.

[0014] The fifth aspect of the present invention provides a computer program product, which includes a computer program that, when executed by a computer device, implements the method described in any of the foregoing embodiments.

[0015] The present invention provides a method and apparatus for collaborative optimization of the operating strategy of a controlled system group. By pre-designing a Q-factor fusion function for the target controlled system agent, the Q-factor fusion function shares / merges the Q-factor data of the target controlled system agent with that of similar controlled system agents. Similarity between controlled system agents refers to the consistency of their optimal actions in the same state. This can improve the performance of the strategy learned from the shared environment, widen the gap between the Q-factors of the optimal action and other actions, and improve the probability of correctly selecting the optimal action under the same observation uncertainty. By collaboratively using the operating data of controlled systems with consistent actions, the data efficiency of the strategy optimization process is effectively improved.

[0016] A sampling data allocation algorithm is established based on the Q-factor fusion function of the target controlled system agent. Based on the Q-factor fusion data and the sampling data allocation algorithm, the sampling amount for each action is allocated, which can progressively maximize the probability of correctly selecting the optimal action. Simultaneously, this invention can improve the data efficiency of online reinforcement learning during the operation and scheduling of the controlled system by controlling the sampling amount, significantly reducing the time required for online decision-making and decreasing data usage.

[0017] To make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating the criterion determination process in the sampling data allocation algorithm of this embodiment of the invention is shown; Figure 2 A flowchart of the collaborative optimization method for the operation strategy of a controlled system group according to an embodiment of the present invention is shown; Figure 3 A flowchart illustrating the target sample size calculation process according to an embodiment of the present invention is shown; Figure 4 A structural diagram of the collaborative optimization device for the operation strategy of a controlled system group according to an embodiment of the present invention is shown; Figure 5 A structural diagram of a computer device according to an embodiment of the present invention is shown; Figure 6 A flowchart illustrating the program design of the controlled system operation strategy collaborative optimization method according to an embodiment of the present invention is shown.

[0020] Explanation of symbols in the attached drawings: 401. Sampling Calculation Unit; 402. Target sampling quantity determination unit; 403. Supplementary sampling quantity determination unit; 404. Supplementary sampling unit; 405. Analysis Unit; 502. Computer equipment; 504, Processor; 506. Memory; 508. Drive mechanism; 510. Input / output module; 512. Input devices; 514. Output devices; 516. Presentation equipment; 518. Graphical User Interface; 520. Network interface; 522. Communication link; 524. Communication bus. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, apparatus, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0023] This specification provides the operational steps of the methods described in the embodiments or flowcharts, but based on conventional or non-inventive labor, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the methods shown in the embodiments or drawings can be executed sequentially or in parallel.

[0024] It should be noted that the data involved in this invention (including but not limited to data used for analysis, data stored, data displayed, etc.) are all information and data authorized by the user or fully authorized by all parties, and the acquisition, transmission, storage, use and processing of the relevant data comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0025] It should be noted that in the embodiments of the present invention, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of the present invention. However, they do not mean that the applicant has used or necessarily used the solution.

[0026] The controlled system cluster mentioned in this invention refers to a cluster of multiple controlled systems distributed in different regions, and there are certain similarities between the controlled systems, all of which have control requirements, such as an energy storage power station cluster. This invention does not limit the specific type of the controlled system cluster, and for ease of description, the following embodiments will use an energy storage power station cluster as an example.

[0027] In existing technologies, the optimization of operating strategies for controlled system groups with multiple controlled systems suffers from low efficiency and poor accuracy. It fails to utilize the similarity of decision-making problems among different energy storage power stations and fails to improve the overall learning efficiency of energy storage power station groups through methods such as transfer learning.

[0028] To address the aforementioned technical problems, this invention shares the Q-factor data of the target controlled system agent and its similar controlled system agents for the same action under the same state, thus obtaining a Q-factor fusion function for the target controlled system agent (used to fuse the Q-factors of the target controlled system agent and its similar controlled system agents under the same state and action). A sampling data allocation algorithm is established using the Q-factor fusion function, which further widens the gap between the Q-factors of the optimal action and other actions. This facilitates increasing the probability of correctly selecting the optimal action under the same observational uncertainty. By synergistically using the operational data of the controlled system with consistent actions, the data efficiency of the strategy optimization process is effectively improved. Allocating the sampling amount for each action based on the Q-factor fusion data and the sampling data allocation algorithm progressively maximizes the probability of correctly selecting the optimal action. Simultaneously, this invention, through the control of the sampling amount, also improves the data efficiency of online reinforcement learning during the operation and scheduling of the controlled system agent, significantly reducing the time required for online decision-making and decreasing data usage.

[0029] Specifically, before executing the sampling data allocation algorithm, the sampling data allocation algorithm is first established, such as... Figure 1 As shown, the process of establishing the sampling data allocation algorithm includes: Step 101: Based on the equipment information, system status, control actions, and optimization target information of each controlled system, establish the performance function and action sampling constraints of the control strategy for each controlled system.

[0030] Step 102: Transform the performance function of the control strategy of each controlled system to obtain the Q-factor fusion function of each controlled system agent.

[0031] Step 103: Based on the Q-factor fusion function of each controlled system agent, establish an objective function for the approximate optimal action selection probability function.

[0032] Step 104: Determine the criteria for the sampling data allocation algorithm based on the objective function of the near-optimal action selection probability function and the action sampling amount constraint.

[0033] In detail, the equipment information, system status, control actions, and optimization target information in step 101 can be determined according to the specific situation of the controlled system, and this invention does not limit the specific parameters included. A performance function for the control strategy is established based on the equipment information, system status, control actions, and optimization target. Each controlled system is managed by an agent responsible for related strategy optimization, and the performance function of the control strategy for each controlled system can be expressed by the following formula: (1) in, J Represents the performance function. s t express t The system state is constantly monitored. This represents the strategy used by agent i in the controlled system for decision-making. c Represents the target object function. T Indicates the length of the simulation period.

[0034] The goal of the controlled system agent i is to find a strategy d that minimizes the performance function. i , means as follows: (2) The Q-factor (also known as the action value function or action value function) of the controlled system agent i can be expressed by the following formula: (3) in, The Q-factor represents the state s and action a of the controlled system agent i. The strategy is represented by 'c'; the target object function is represented by 'c'. This represents the probability that agent i in the controlled system transitions to state s' from state s and action a; Indicates the state s' and policy of the controlled system agent i. The Q factor at that time. In practical implementation, for the sake of simplicity, it can also be... That is .

[0035] During the learning process, agent i may acquire a policy d. For an action a, agent i will run some simulations to collect data. The collected samples are denoted as... , , This represents the number of samples taken by agent i for action a. In actual implementation, another agent j may have already performed similar work and collected sampling data. , , Let represent the number of samples taken by agent j for action a. Agent j wishes to share this data with agent i. To simplify the discussion, the following assumptions are made: Assuming the same state space We will discuss how to extend the conclusions to situations beyond the above assumptions later.

[0036] Now agent i can generate data based on local samples. , and shared samples , Estimate: (4) in, denoted by s, which represents the Q-factor fusion function of controlled system agent i, i.e., the function obtained by fusing the Q-factors of controlled system agent i with those of its similar controlled system agent j; s represents the state. a represents an action; f represents a fusion function, such as summing the elements contained in f. denoted as the observed mean of the Q factor of the controlled system agent i in state s, action a, and with a sample size of k. This represents the observed mean of the Q-factor of the controlled system agent j in state s and action a when the sample size is k; k represents the sample size. This indicates that the controlled system's intelligent agent i is performing actions. a Time sample size This indicates that the controlled system agent j is performing actions. a Number of samples at time.

[0037] Then we can compare the different actions and select the action that minimizes the estimated Q factor, i.e.: (5) We want to adjust the sampling budget for the entire action space allocation to minimize the Q-factor of the current state, i.e.: (6) Where N is the total number of samples, quantifying the total number of samples that agent i has in this iteration. The problem represented by the previous formula is simplified to P1. When the system reaches the next state, the next iteration begins. Under the convenience assumption of a policy-controlled Markov chain, it can be proven that when there exists infinite sampling ( As the iteration continues, the estimated optimal action in each state converges to the same action as the optimal policy. We want to obtain more details about f, namely how the samples shared by agent j are used by agent i, and how N is allocated to accelerate this learning process.

[0038] For the above formula (4), we consider a specific function. fThis function combines samples from agents i and j to generate... The estimate of that specific function f sample mean and Add them directly, where: (7) (8) Therefore, we can conclude that: (9) We now adopt a Bayesian perspective, assuming that the observed samples are given and fixed. But the true value... It is a random variable that follows a posterior distribution. When the sample size When the value is large, the central limit theorem is proved. and The difference between them follows a Gaussian random variable. Given a Gaussian prior and Gaussian observation noise, the posterior distribution is also Gaussian. Therefore, we can make the following assumption: For all , The prior distribution is as follows .

[0039] The above assumptions basically express our views on... We know nothing except that every value has the same probability. Under the above assumptions, we can give the observations... and , The posterior distribution is ,in: (10) We are interested in maximizing the probability of correctly choosing the optimal action, that is: (11) In the above formula Simplifying to b, we get the following question, abbreviated as P2: (12) That is, after fusing the information from agents i and j, the probability that the optimal action b estimated by agent i in state s is the true optimal action. N a i Represents intelligent agents i Assigned to action a The number of samples. In online operation, the sum of the computational costs allocated to agent i across all actions in the current state should equal the given total computational cost N.

[0040] Problem P2 is relatively easier to solve than the previous problem P1 because it only requires evaluating the probability that the observed best action in the current state is the true optimal action, while the original problem model required evaluating the average performance of a given policy.

[0041] Bonferroni's inequality states that: (13) in, For the event. Next, we can obtain: (14) The right side of the above equation can be called the Approximate Optimal Action Selection Probability Function (APCS). Therefore, the objective function of the approximate optimal action selection probability function can be obtained as follows: (15) (16) The above question can be abbreviated as P3.

[0042] in, APCS denoted as the probability function for selecting the near-optimal action, which is the lower bound of the PCS; b represents the optimal action. Indicates a non-optimal action; Represents the intelligent agent of the controlled system i With the controlled system intelligent agent j In state s Find the optimal action b Q-factor fusion value; Represents the intelligent agent of the controlled system i With the controlled system intelligent agent j Non-optimal action in state s The Q-factor fusion value.

[0043] Problem P3 is easier to solve than problem P2 because the original joint probability is decomposed into a pairwise comparison of the performance of the observed optimal action b in the current state with other actions a. This decomposition provides a basis for parallel computing. However, as will be seen later, this patent presents a faster solution formula.

[0044] In step 104, the criteria for the sampling data allocation algorithm are obtained to provide guidance: After the controlled system and its similar controlled system agents fuse data, the larger the noise-to-signal ratio of the non-observed optimal action under the same state s, the larger the sampling amount of the non-observed optimal action. In the noise-to-signal ratio, noise refers to the sample standard deviation of the Q factor of the non-observed optimal action, and information refers to the performance difference of the fused Q factor values ​​of the observed optimal action and the non-observed optimal action. The number of samples for the best observed action should be greater than the number of samples for other non-best observed actions under the same condition.

[0045] Specifically, the process of determining the criteria for the sampling data allocation algorithm based on the objective function of the near-optimal action selection probability function and the action sampling volume constraints includes: First, we introduce the Lagrange equation: (17) Where λ is a Lagrange multiplier, Then, according to the Karush-Kuhn-Tucker (KKT) conditions (assuming the KKT regularization conditions hold), we have the necessary conditions for a local optimum: (18) (19) (20) (twenty one) For non-observation optimal action We have: (twenty two) For the first term on the right side of formula (22), we can further obtain: (twenty three) Please note ,in: (twenty four) (25) Therefore, we can conclude that: (26) Then we can get: (27) in, (28) Furthermore, we can obtain: (29) Based on similar analysis, for the observed optimal action b, we have: (30) Next, we can analyze The relationship between them.

[0046] First, we analyze two unobserved optimal actions p and q. According to formula (29), we can obtain: (31) (32) Combining formulas (31) and (32), we can obtain: (33) As the sampling size N approaches infinity, it is reasonable to assume that the number of samples observed for the optimal action assignment also increases, i.e. Therefore, we can obtain: (34) Similarly, we can obtain: (35) Next, we can obtain: (36) Furthermore, formula (36) can be adjusted to: (37) Next, taking the logarithm of both sides of formula (37), we can obtain: (38) We can see that as N increases, the logarithmic term grows significantly slower than the other terms; therefore, it can be ignored in the above formula. From this, we can conclude that: (39) Therefore, we can conclude that: (40) Secondly, analysis and The relationship between them can be obtained from formula (29): (41) Substituting formula (41) into formula (30), we get: (42) Then we can get: (43) After readjusting formula (43), we can obtain: (44) In summary, based on the above analysis, the criteria for the sampling data allocation algorithm can be summarized as follows: The sampling data allocation algorithm includes: Criterion 1: For unobserved optimal actions p and q, , ; in, This represents the sample size of the controlled system agent i when it is not observing the optimal action p; This represents the sample size of agent i in the controlled system when it is not the optimal action q observed. The sample standard deviation of the Q factor represents the Q factor of the controlled system agent i when the action p is not observed optimally. The sample standard deviation of the Q factor represents the Q factor of the controlled system agent i when the action q is not observed optimally. This represents the performance difference between the optimal action b and the non-optimal action p of the controlled system agent i and its similar controlled system agent j. This represents the performance difference between the optimal action b and the non-optimal action q of the controlled system agent i and its similar controlled system agent j. Criterion 2: For observing the optimal action b and the non-optimal action a, ; in, This represents the sampling amount of the controlled system agent i when observing the optimal action b; The sample standard deviation of the Q factor represents the Q factor of the controlled system agent i when observing the optimal action b. This indicates that the controlled system agent i is performing a non-optimal action. a Sampling volume at time; This indicates that the controlled system agent i is performing a non-optimal action. a The sample standard deviation of the Q factor.

[0047] ; .

[0048] Q ij This is a specific way of fusing information from agents i and j. Specifically, it involves directly adding the Q-factors of the same action in the same state, estimated by both agents using their respective data. The direct advantage of this approach is that, as long as the optimal actions of the two agents in the same state are identical, this method will further widen the gap between the Q-factors of the optimal action and other actions. This is beneficial for improving the probability of correct selection under the same observational uncertainty. This property can be described as optimal action consistency, i.e., the optimal actions of the two agents in the same state are identical.

[0049] If the state spaces of the subproblems of two agents are not completely identical, data sharing between the two agents can be achieved by normalization or by taking their intersection, selecting a state that satisfies optimal action consistency. In this mode, the agents sharing data may change depending on the state. This aligns well with the characteristics of actual energy storage power station cluster operations, where data is shared only between energy storage power stations that satisfy optimal action consistency under similar states, or by normalizing the data of each energy storage power station according to its load level, energy storage capacity level, and daily power generation level, and then sharing data between energy storage power stations that satisfy optimal action consistency under similar states. Criterion 1 states that after agent i fuses the data of agent j, for two actions that are not observed optimally, the sampling amount of different actions should be proportional to the square of the "noise-to-information ratio" (NNR). Here, noise refers to the sample standard deviation of a single observation of the Q-factor of an action, and information refers to the difference in the sample mean of the Q-factor of the observed optimal action from other actions. Note that fusing the samples of agent j using the aforementioned method further increases this "information" compared to before fusion, which is equivalent to further improving the speed at which the observation order of the superiority or inferiority of two actions converges to the true order.

[0050] Criterion 2 states that for the observed optimal action b, the sample size it should receive should be approximately equal to the sum of the sample sizes of all other actions after normalization according to their respective standard deviations. As a result of this allocation method, with continued sampling, the observed optimal action eventually receives a much larger sample size than other actions. In other words, with the same total sample size, after allocating the sample size using criteria 1 and 2, the observed optimal action converges to the true optimal action the fastest, and fusing data from agent j can further improve this convergence rate.

[0051] According to formula (24), we can obtain: ; The first and second terms on the right side of the above equation are respectively and ,Right now: .

[0052] Then we can get: .

[0053] The second term on the right-hand side of the previous formula quantifies the influence of agent j. Therefore, it can be proven that without the influence of agent j, the above criterion becomes: ; This means that, It is the difference in observational performance between the observation-optimal action b and the non-observation-optimal action a determined by agent j. This is indeed useful if agents i and j share the same optimal action in this state.

[0054] Based on the aforementioned sampling data allocation algorithm, a collaborative optimization method for the operation strategy of the controlled system group can be implemented. Specifically, according to control requirements, the target controlled system is first determined from the controlled system group; then, controlled systems similar to the target controlled system are found from the controlled system group, i.e., controlled systems that satisfy state-action consistency. The process of finding controlled systems similar to the target controlled system can refer to existing technologies and will not be described in detail here; then, as... Figure 2 As shown, perform the following steps: Step 201: Sample the simulation results of each action of the target controlled system agent according to the preset sampling amount, and determine the Q factor sample variance of each action, the performance difference of the Q factor fusion value of each optimal action and non-optimal action, and the total amount of samples consumed based on the sampling results.

[0055] The Q-factor fusion value of each action is calculated by the Q-factor fusion function, which is the Q-factor function of the target controlled system agent and its similar controlled system agents under the same state-action, and is used to realize Q-factor data sharing between the target controlled system and its similar controlled systems.

[0056] The performance difference between the Q-factor fusion values ​​of optimal and non-optimal actions refers to the difference between the Q-factor fusion values ​​of the optimal and non-optimal actions of the target controlled system agent in state s. In some implementations, the performance difference can be expressed by the following formula: ; Where a represents the non-observed optimal action; b represents the observed optimal action; and s represents the current state. The fused mean of the Q-factor observations of controlled system agent i and its similar controlled system agent j in state s and unobserved optimal action a; The fusion mean of the Q-factor observations of controlled system agent i and its similar controlled system agent j in state s and when observing the optimal action b.

[0057] In this step, the action types include optimal actions and non-optimal actions. The sample variance of the Q factor for each type of action is calculated from the sampled Q factor for each type of action. For specific calculation methods, please refer to the existing variance calculation formula.

[0058] The actions in this step can be determined based on the action space. The preset sample size (N0) can be set based on the preset total sample size N, the preset number of loops, the action space, and the newly added sample size Δ. Generally, the preset number of loops is greater than 10, and the preset total sample size N is much larger than the preset sample size N0 multiplied by the number of actions in the action space.

[0059] In this step, the total number of samples consumed is the product of the preset sampling quantity and the number of actions in the action space, plus the newly added sampling quantity. The corresponding calculation formula is: ;in, Indicates the total number of samples. Indicates the preset number of samples. This represents the sampled quantity in the action space of the target controlled system i. This indicates the number of new samples.

[0060] In this step, the actions can be directly applied to the simulation model of the controlled system (e.g., a digital twin model), and the simulation results of each action can be determined through the simulation model. In practice, the execution results of a preset sample size can also be obtained from historical data based on the state and actions, and the obtained execution results can be used as the simulation results.

[0061] Step 202: Analyze the Q-factor sample variance, performance differences, and total amount of consumed samples for each action using the sampling data allocation algorithm to obtain the target sample size for each action.

[0062] The sampling data allocation algorithm includes criteria for determining the sampling amount for various actions. These criteria are derived from the Q-factor fusion function and are used to progressively maximize the probability of correctly selecting the optimal action. The specific content of the criteria is detailed in the aforementioned embodiments and will not be repeated here.

[0063] The specific implementation process of this step can be found in subsequent embodiments, and will not be described in detail here.

[0064] Step 203: Determine the supplementary sampling amount for each action based on the target sample size and the sample size already consumed for each action.

[0065] During this step, the supplementary sampling quantity for each action is obtained by subtracting the sample quantity consumed by each action from the target sample quantity for each action.

[0066] Step 204: Perform supplementary sampling based on the supplementary sampling amount for each action, and use the simulation results of the supplementary sampling to redetermine the Q-factor sample variance, performance difference, and total amount of samples consumed for each action.

[0067] In this step, the calculation process for the Q-factor sample variance and performance difference can be referred to step 201, and will not be repeated here. The total amount of samples consumed in this step is the sum of the actual sampling amounts of the target controlled system for each action.

[0068] Step 205: Adjust the total number of consumed samples. Determine whether the total number of consumed samples is less than the preset total sampling amount. If yes, recalculate the target sample amount and subsequent steps. If no, output the optimal action.

[0069] During this step, the consumed sample quantity is increased by the newly added sample quantity Δ to obtain the adjusted total consumed sample quantity. The action sample quantity constraint is the preset total sample quantity N, which can be set jointly according to the preset initial value of the sample quantity in step 201 and the action space, as long as it can satisfy the preset number of repetitions from step 201 to step 205.

[0070] In this step, if the total number of samples consumed is less than or equal to the action sampling amount constraint, it means that the number of samples consumed after the next iteration has not reached or is exactly equal to the preset total number of samples, and the feasibility of the action can continue to be observed. If the total number of samples consumed is greater than the action sampling amount constraint, it means that the number of samples consumed after the next iteration has exceeded the preset total number of samples, and the feasibility of the action cannot continue to be observed.

[0071] In one specific implementation, the algorithm flow of this embodiment is as follows: Figure 6 As shown. Figure 6 The controlled system agent j is similar to the target controlled system agent i. n 0 For the preset sampling amount, a For action, This is the action space. This represents the performance difference between the Q-factor fusion value of the observed optimal action and the Q-factor fusion value of the unobserved optimal action a for the target controlled system agent i and its similar controlled system agent j in state s. This represents the Q-factor sample variance of action a. Indicates the number of samples consumed, △ represents the number of new samples. N represents the preset total sample size. Indicates action a The required sample size ratio Indicates action a The required sample size Indicates action a Sample size consumed Indicates action a The target sample size Indicates action a Supplement the sampling volume.

[0072] For each state s, the optimal action b(s) is the action in the action space that minimizes the Q-factor fusion value, i.e., satisfies The action.

[0073] This embodiment pre-designs a Q-factor fusion function for the target controlled system agent. This function shares / merges / integrates the Q-factor data of the target controlled system and its similar controlled systems, which can further widen the gap between the Q-factors of the optimal action and other actions. This is beneficial to improving the probability of correctly selecting the optimal action under the same observation uncertainty. By using the operational data of controlled systems with consistent actions, the data efficiency of the strategy optimization process can be effectively improved.

[0074] A sampling data allocation algorithm is established based on the Q-factor fusion function of the target controlled system agent. Based on the Q-factor fusion data and the sampling data allocation algorithm, the sampling amount for each action is allocated, which can progressively maximize the probability of correctly selecting the optimal action. Simultaneously, this invention can improve the data efficiency of online reinforcement learning during the operation and scheduling of the controlled system by controlling the sampling amount, significantly reducing the time required for online decision-making and decreasing data usage.

[0075] In one embodiment of the present invention, as Figure 3 As shown, step 204 above uses a sampling data allocation algorithm to analyze the Q-factor sample variance, performance differences, and total amount of consumed samples for each action to obtain the target sample size for each action, including: Step 301: Analyze the Q-factor sample variance and performance differences of each action according to the sampling data allocation algorithm to obtain the sampling ratio of each action.

[0076] Step 302: Calculate the number of samples required for each action by multiplying the sampling ratio of each action by the total number of samples consumed. The corresponding calculation formula is: .

[0077] Step 303: Select the larger of the number of samples that should be obtained for each action and the number of samples actually used for each action as the target sample number for each action. In specific implementation, the corresponding calculation formula for this step is: ,in, This represents the target number of samples for action a. This indicates the number of samples that action a should obtain. This indicates the actual number of samples used for action a.

[0078] Based on the same inventive concept, this invention also provides a device for collaborative optimization of the operation strategy of a controlled system group, as described in the following embodiments. Since the principle of the device for collaborative optimization of the operation strategy of a controlled system group is similar to that of the method for collaborative optimization of the operation strategy of a controlled system group, the implementation of the device for collaborative optimization of the operation strategy of a controlled system group can refer to the method for collaborative optimization of the operation strategy of a controlled system group, and repeated details will not be elaborated further.

[0079] Specifically, such as Figure 4As shown, the collaborative optimization device for the operation strategy of the controlled system group includes: The sampling calculation unit 401 is used to sample the simulation results of each action of the target controlled system intelligent agent according to the preset sampling amount, and determine the Q factor sample variance of each action, the performance difference of the Q factor fusion value of each optimal action and non-optimal action, and the total amount of samples consumed based on the sampling results.

[0080] The Q-factor fusion value of each action is calculated by the Q-factor fusion function, which is the Q-factor function of the target controlled system agent and its similar controlled system agents under the same state-action.

[0081] The target sampling quantity determination unit 402 is used to analyze the Q-factor sample variance, performance difference and total amount of consumed samples of each action using the sampling data allocation algorithm to obtain the target sample quantity of each action.

[0082] The sampling data allocation algorithm includes a criterion for determining the sampling amount of various actions based on the analysis of the Q-factor fusion function, which is used to progressively maximize the probability of correctly selecting the optimal action.

[0083] The supplementary sampling quantity determination unit 403 is used to determine the supplementary sampling quantity for each action based on the target sample quantity for each action and the sample quantity consumed by each action.

[0084] The supplementary sampling unit 404 is used to perform supplementary sampling based on the supplementary sampling amount of each action, and to redetermine the Q-factor sample variance, performance difference, and total amount of consumed samples of each action using the simulation results of the supplementary sampling.

[0085] Analysis unit 405 is used to adjust the total number of consumed samples and determine whether the total number of consumed samples is less than the preset total sampling amount. If so, the target sampling amount determination unit is used to recalculate the target sample amount. If not, the optimal action is output.

[0086] This embodiment pre-designs a Q-factor fusion function for the target controlled system agent. This function shares / merges / integrates the Q-factor data of the target controlled system and its similar controlled systems. The target controlled system and its similar controlled systems refer to controlled systems in which the controlled system agents have the same optimal action under the same state. This can further widen the gap between the Q-factors of the optimal action and other actions, thereby improving the probability of correctly selecting the optimal action under the same observation uncertainty. By using the operating data of controlled systems with consistent actions, the data efficiency of the strategy optimization process can be effectively improved.

[0087] A sampling data allocation algorithm is established based on the Q-factor fusion function of the target controlled system agent. Based on the Q-factor fusion data and the sampling data allocation algorithm, the sampling amount for each action is allocated, which can progressively maximize the probability of correctly selecting the optimal action. Simultaneously, this invention can improve the data efficiency of online reinforcement learning during the operation and scheduling of the controlled system by controlling the sampling amount, significantly reducing the time required for online decision-making and decreasing data usage.

[0088] When this invention is applied to the collaborative optimization of operation strategies for energy storage power station clusters, it can improve the data efficiency of the energy storage power station cluster operation scheduling strategy optimization process, significantly improve the performance of the obtained strategy, and reduce data usage.

[0089] In one embodiment of the present invention, a computer device is also provided, such as... Figure 5 As shown, computer device 502 may include one or more processors 504, such as one or more central processing units (CPUs), each of which may implement one or more hardware threads. Computer device 502 may also include any memory 506 for storing information of any kind, such as code, settings, data, etc. Non-limitingly, for example, memory 506 may include any type of RAM, any type of ROM, flash memory, hard disk, optical disk, etc. More generally, any memory can use any technology to store information. Furthermore, any memory may provide volatile or non-volatile retention of information. Furthermore, any memory may represent a fixed or removable component of computer device 502. In one case, when processor 504 executes associated instructions stored in any memory or combination of memories, computer device 502 may perform any operation of the associated instructions. Computer device 502 also includes one or more drive mechanisms 508 for interacting with any memory, such as hard disk drive mechanisms, optical disk drive mechanisms, etc.

[0090] Computer device 502 may also include an input / output module 510 (I / O) for receiving various inputs (via input device 512) and providing various outputs (via output device 514). A specific output mechanism may include a presentation device 516 and an associated graphical user interface 518 (GUI). In other embodiments, the input / output module 510 (I / O), input device 512, and output device 514 may be omitted, and the device may function solely as a computer device within a network. Computer device 502 may also include one or more network interfaces 520 for exchanging data with other devices via one or more communication links 522. One or more communication buses 524 couple the components described above together.

[0091] Communication link 522 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 522 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.

[0092] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the above-described method.

[0093] This invention also provides a computer-readable instruction, wherein when a processor executes the instruction, the program therein causes the processor to perform the method as described in any of the foregoing embodiments.

[0094] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0095] It should also be understood that, in the embodiments of the present invention, the term "and / or" is merely a description of the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, in the present invention, the character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0096] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0098] In the embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, or may be electrical, mechanical, or other forms of connection.

[0099] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention, depending on actual needs.

[0100] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A collaborative optimization method for the operation strategy of an energy storage power station cluster, characterized in that, The energy storage power station cluster includes multiple energy storage devices, and the method includes: Based on the equipment information, system status, control actions, and optimization target information of each energy storage device, establish the performance function and action sampling constraints of the control strategy for each energy storage device. The performance function of the control strategy of each energy storage device is transformed to obtain the Q-factor fusion function of each energy storage device agent; Based on the Q-factor fusion function of each energy storage device's intelligent agent, an objective function is established to approximate the optimal action selection probability function. Based on the objective function of the near-optimal action selection probability function and the action sampling amount constraint, the criteria for the sampling data allocation algorithm are determined. Based on control requirements, target energy storage devices and similar energy storage devices are identified from the energy storage power station group. The simulation results of each action of the target energy storage device intelligent agent are sampled according to the preset sampling amount. Based on the sampling results, the Q factor sample variance of each action, the difference between the Q factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed are determined. The Q factor fusion value of each action is calculated by the Q factor fusion function, which is the Q factor function of the target energy storage device intelligent agent and its similar energy storage device intelligent agents under the same state-action. The similarity between energy storage device intelligent agents means that the optimal actions of the energy storage device intelligent agents are consistent under the same state. The sampling data allocation algorithm is used to analyze the Q-factor sample variance of each action, the difference between the Q-factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed to obtain the target sample size of each action. The sampling data allocation algorithm includes the determination criteria for the sample size of various actions obtained from the analysis of the Q-factor fusion function, which is used to progressively maximize the probability of correctly selecting the optimal action. The additional sampling amount for each action is determined based on the target sample size and the consumed sample size for each action. Supplementary sampling is performed based on the supplementary sampling amount for each action. The simulation results of the supplementary sampling are used to redetermine the Q factor sample variance of each action, the difference between the Q factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed. Adjust the total number of consumed samples, and determine whether the total number of consumed samples is less than the preset total sampling amount. If so, recalculate the target sample amount and subsequent steps. If not, output the optimal action. The Q-factor fusion function of each energy storage device's intelligent agent, obtained by transforming the performance function of the control strategy for each energy storage device, includes: The Q-factor fusion function of the energy storage device's intelligent agent is expressed by the following formula: , ; in, Indicates the intelligent entity of energy storage device i With intelligent entities of energy storage devices j In state s Next action a Q-factor fusion value; s Indicates state; a Indicates an action; This indicates the state and actions of the intelligent agent i in the energy storage device. a The observed mean of the Q factor at that time; This indicates the state and actions of the intelligent agent j of the energy storage device. a The observed mean of the Q-factor at time; S represents the state space; The action space is defined, and intelligent agent i of the energy storage device is similar to intelligent agent j of the energy storage device. Among them, the objective function for establishing the approximate optimal action selection probability function based on the Q-factor fusion function of each energy storage device's intelligent agent includes: The objective function of the approximate optimal action selection probability function can be expressed by the following formula: ; ; in, APCS This represents the probability function for selecting an approximate optimal action. b Indicates the optimal action; Indicates a non-optimal action; Indicates the intelligent entity of energy storage device i With intelligent entities of energy storage devices j In state s Find the optimal action b Q-factor fusion value; Indicates the intelligent entity of energy storage device i With intelligent entities of energy storage devices j Non-optimal action in state s The Q-factor fusion value of energy storage device intelligent agent i is similar to that of energy storage device intelligent agent j. N a i Represents intelligent agents i Assigned to action a The number of samples.

2. The method as described in claim 1, characterized in that, The sampling data allocation algorithm is used to indicate: After data fusion between an energy storage device intelligent agent and its similar energy storage device intelligent agents, the larger the noise-to-signal ratio of the non-observed optimal action under the same state s, the larger the sampling amount of the non-observed optimal action. In the noise-to-signal ratio, noise refers to the sample standard deviation of the Q factor of the non-observed optimal action, and information refers to the difference between the fused Q factor values ​​of the observed optimal action and the non-observed optimal action. The number of samples for the best observed action should be greater than the number of samples for other non-best observed actions under the same condition.

3. The method as described in claim 2, characterized in that, The sampling data allocation algorithm includes: Criterion 1: For unobserved optimal actions p and q , , ; in, This represents the sampling amount of the intelligent agent i of the energy storage device when it is not the optimal action p observed; This represents the sampling amount of the intelligent agent i of the energy storage device when it is not the optimal action q observed; The sample standard deviation of the Q factor represents the energy storage device intelligent agent i when performing a non-observed optimal action p; The sample standard deviation of the Q factor represents the energy storage device intelligent agent i when the action q is not observed optimally. This represents the difference between the Q-factor fusion values ​​of the optimal action b and the non-optimal action p of energy storage device intelligent agent i and its similar energy storage device intelligent agent j; This represents the difference between the Q-factor fusion values ​​of the optimal action b and the non-optimal action q of energy storage device intelligent agent i and its similar energy storage device intelligent agent j; Criterion 2: For observation of optimal action b and non-optimal action b a , ; in, This represents the sampling amount of the intelligent agent i of the energy storage device when observing the optimal action b; The sample standard deviation of the Q factor represents the energy storage device intelligent agent i when observing the optimal action b; This indicates that the intelligent agent i of the energy storage device is performing a non-optimal action. a Sampling volume at time; This indicates that the intelligent agent i of the energy storage device is performing a non-optimal action. a The sample standard deviation of the Q factor.

4. The method as described in claim 1, characterized in that, The target sample size for each action is obtained by analyzing the Q-factor sample variance of each action, the difference between the Q-factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed using a sampling data allocation algorithm. This includes: The sampling ratio of each action is obtained by analyzing the sample variance of the Q factor of each action and the difference between the fusion values ​​of the Q factors of each optimal action and non-optimal action based on the sampling data allocation algorithm. The number of samples that should be obtained for each action is obtained by multiplying the sampling ratio of each action by the total number of samples consumed; Choose the larger of the number of samples that should be obtained for each action and the number of samples that have actually been used for each action as the target number of samples for each action.

5. A collaborative optimization device for the operation strategy of an energy storage power station cluster, characterized in that, The energy storage power station cluster includes multiple energy storage devices, and the devices include: The unit that performs the following operations: Based on the equipment information, system status, control actions, and optimization target information of each energy storage device, establish the performance function and action sampling constraints of the control strategy for each energy storage device; The performance function of the control strategy of each energy storage device is transformed to obtain the Q-factor fusion function of each energy storage device agent; Based on the Q-factor fusion function of each energy storage device's intelligent agent, an objective function is established to approximate the optimal action selection probability function. Based on the objective function of the near-optimal action selection probability function and the action sampling amount constraint, the criteria for the sampling data allocation algorithm are determined. The unit performs the following operation: determining the target energy storage device and similar energy storage devices from the energy storage power station group according to control requirements; The sampling and calculation unit is used to sample the simulation results of each action of the target energy storage device intelligent agent according to a preset sampling amount. Based on the sampling results, it determines the Q-factor sample variance of each action, the difference between the Q-factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed. The Q-factor fusion value of each action is calculated by the Q-factor fusion function, which is the Q-factor function of the target energy storage device intelligent agent and its similar energy storage device intelligent agents under the same state-action. The similarity between energy storage device intelligent agents refers to the consistency of the optimal action of the energy storage device intelligent agents under the same state. The target sampling quantity determination unit is used to analyze the Q-factor sample variance of each action, the difference between the Q-factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed to obtain the target sample quantity of each action using a sampling data allocation algorithm. The sampling data allocation algorithm includes determination criteria for various action sampling quantities obtained based on the Q-factor fusion function analysis, which is used to progressively maximize the probability of correctly selecting the optimal action. The supplementary sampling quantity determination unit is used to determine the supplementary sampling quantity for each action based on the target sample quantity for each action and the sample quantity consumed for each action. The supplementary sampling unit is used to perform supplementary sampling based on the supplementary sampling amount of each action, and to redetermine the Q factor sample variance of each action, the difference between the Q factor fusion values ​​of each optimal action and non-optimal action, and the total amount of samples consumed using the simulation results of the supplementary sampling. The analysis unit is used to adjust the total number of consumed samples and determine whether the total number of consumed samples is less than the preset total sampling amount. If so, the target sampling amount determination unit is used to recalculate the target sample amount. If not, the optimal action is output. The Q-factor fusion function of each energy storage device's intelligent agent, obtained by transforming the performance function of the control strategy for each energy storage device, includes: The Q-factor fusion function of the energy storage device's intelligent agent is expressed by the following formula: , ; in, Indicates the intelligent entity of energy storage device i With intelligent entities of energy storage devices j In state s Next action a Q-factor fusion value; s Indicates state; a Indicates an action; This indicates the state and actions of the intelligent agent i in the energy storage device. a The observed mean of the Q factor at that time; This indicates the state and actions of the intelligent agent j of the energy storage device. a The observed mean of the Q-factor at time; S represents the state space; The action space is defined, and intelligent agent i of the energy storage device is similar to intelligent agent j of the energy storage device. Among them, the objective function for establishing the approximate optimal action selection probability function based on the Q-factor fusion function of each energy storage device's intelligent agent includes: The objective function of the approximate optimal action selection probability function can be expressed by the following formula: ; ; in, APCS This represents the probability function for selecting an approximate optimal action. b Indicates the optimal action; Indicates a non-optimal action; Indicates the intelligent entity of energy storage device i With intelligent entities of energy storage devices j In state s Find the optimal action b Q-factor fusion value; Indicates the intelligent entity of energy storage device i With intelligent entities of energy storage devices j Non-optimal action in state s The Q-factor fusion value of energy storage device intelligent agent i is similar to that of energy storage device intelligent agent j. N a i Represents intelligent agents i Assigned to action a The number of samples.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor of the computer device, it implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-energy system collaborative scheduling method and device based on electric car access

    CN107769237A

  • Multi-time-scale multi-agent reinforcement learning method and device

    CN112163690A