Multi-power station game method and device based on random generalized Nash equilibrium search
By adopting a method based on random generalized Nash equilibrium search in multi-power station game scenarios, and using technologies such as variance reduction operator splitting algorithm and multiplication sub-graphs, the problem of the inability to achieve distributed optimization decisions in a random environment in the existing technology is solved, and more efficient and reliable energy allocation and system stability are achieved.
Patent Information
- Application Number
- CN202410570586.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-05-09
AI Technical Summary
The prior art cannot realize distributed power station optimization decisions and rapid convergence under the conditions of random environmental uncertainty in the multi-power station game scenario.
A multi-power station game method based on random generalized Nash equilibrium search is adopted to construct a distributed random generalized Nash equilibrium problem through the variance reduction operator splitting algorithm, and a multiplication graph and weighted adjacency matrix are used to update the strategy to achieve the goal strategy solution of the target power station.
It improves the decision-making accuracy and reliability of the power station in an uncertain environment, optimizes energy allocation, enhances the response speed and flexibility of the power system, and ensures the stability and overall performance of the system.
Smart Images

Figure CN118473019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power station market competition, and specifically to a multi-power station game method and device based on stochastic generalized Nash equilibrium search. Background Art
[0002] In the current power station market competition environment, power stations face multiple challenges such as large fluctuations in market prices and uncertainties in the supply chain. The factors of fluctuations and uncertainties seriously affect the operation decisions and market competitiveness of power stations. When there are market price fluctuations and supply chain disruptions, the competition problem among power stations can be approximately transformed into a generalized Nash equilibrium problem (GNEP). Although the existing forward-backward (FB) algorithm can theoretically handle the generalized Nash equilibrium problem (GNEP), the distributed computing ability of the algorithm is insufficient and cannot meet the requirements of real-time and accuracy.
[0003] In addition, traditional algorithms also have deficiencies in terms of information locality and privacy protection, which limit the timeliness and efficiency of power stations when making decisions.
[0004] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a multi-power station game method and device based on stochastic generalized Nash equilibrium search, which solves the problem that in the multi-power station game scenario of the existing technology, the existing algorithms cannot achieve distributed power station optimization decisions and fast convergence under the uncertainty of the stochastic environment.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a multi-power station game method based on stochastic generalized Nash equilibrium search, wherein the method includes:
[0008] Obtain the constraint conditions and power station costs of a number of power stations, and construct the energy allocation problem of the number of power stations based on the constraint conditions and power station costs of the number of power stations. Among them, the energy allocation problem of the number of power stations is a distributed stochastic generalized Nash equilibrium problem, and the state variables of the stochastic generalized Nash equilibrium problem include decision variables, dual variables, and auxiliary variables;
[0009] Solve the energy allocation problem of the number of power stations through the variance reduction operator splitting algorithm to obtain the target strategy of the target power station. Among them, the process of solving the energy allocation problem of the number of power stations through the variance reduction operator splitting algorithm specifically includes:
[0010] Construct a multiplier graph of a number of power stations and determine the weighted adjacency matrix;
[0011] For power station i, based on the multiplier graph, obtain the first set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the first set of auxiliary variables and the first set of dual variables of the first multiplier neighbor power stations corresponding to power station i;
[0012] Based on the first set of policy information, determine the random estimate of the first decision variable of power station i;
[0013] Based on the random estimate of the first decision variable, determine the first updated policy of power station i through the projected stochastic gradient descent method;
[0014] Based on the first set of dual variables and the weighted adjacency matrix, update the auxiliary variables of power station i to obtain the first updated auxiliary variables of power station i;
[0015] Based on the first updated policy, the first set of auxiliary variables, and the first set of dual variables, update the dual variables of power station i to obtain the first updated dual variables of power station i;
[0016] For power station i, based on the multiplier graph, obtain the second set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the second set of auxiliary variables and the second set of dual variables of the second multiplier neighbor power stations corresponding to power station i;
[0017] Based on the second set of policy information, determine the random estimate of the second decision variable of power station i;
[0018] Based on the random estimate of the first decision variable and the random estimate of the second decision variable, determine the second updated policy of power station i;
[0019] Based on the first updated auxiliary variables, the second set of dual variables, and the weighted adjacency matrix, update the auxiliary variables of power station i to obtain the second updated auxiliary variables of power station i;
[0020] Based on the first updated policy, the second set of auxiliary variables, and the second set of dual variables, update the dual variables of power station i to obtain the second updated dual variables of power station i;
[0021] Re-execute the step of, for power station i, based on the multiplier graph, obtaining the first set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the first set of auxiliary variables and the first set of dual variables of the first multiplier neighbor power stations corresponding to power station i, until the number of re-executions reaches the preset number of times to obtain the target policy of the target power station.
[0022] In one implementation, the process of constructing the energy allocation problem of a number of power stations based on the constraint conditions and power generation costs of the number of power stations includes:
[0023] Define a variational equilibrium problem based on the constraint conditions of the number of power stations and the power station costs;
[0024] Use a distributed environment to construct the variational equilibrium problem to obtain a distributed variational equilibrium problem;
[0025] Add a random factor to the distributed variational equilibrium problem to obtain a stochastic generalized Nash equilibrium problem. In one implementation, the process of determining the random estimate of the first decision variable of power station i based on the first policy information set can be expressed as:
[0026]
[0027] where u k represents the first policy information set, represents the random event t, S k represents the sampling rate at iteration step k, represents the power station cost of power station i, represents the cumulative gradient of the objective function of power station i with respect to its policy under the random event t.
[0028] In one implementation, the process of determining the first updated policy of power station i based on the random estimate of the first decision variable by the projected stochastic gradient descent method can be expressed as:
[0029]
[0030] where, represents the proximal operator, γ i represents the step size, g i represents a non-smooth function, represents the transpose of the coupling constraint coefficient matrix related to power station i, u i,k represents the policy variable of power station i before the k-th iteration, λ i,k represents the dual variable of the power station at the k-th iteration.
[0031] In one implementation, the process of determining the second updated policy of power station i based on the random estimate of the first decision variable and the random estimate of the second decision variable can be expressed as:
[0032]
[0033] where, represents the random estimate of the first decision variable, represents the random estimate of the second decision variable, Denote the first updated dual variable of power station i.
[0034] In one implementation, the process of updating the auxiliary variable of power station i based on the first updated auxiliary variable, the second set of dual variables, and the weighted adjacency matrix to obtain the second updated auxiliary variable of power station i can be expressed as:
[0035]
[0036] where μ i,k+1 Denote the second updated auxiliary variable of power station i, Denote the first updated auxiliary variable of power station i, σ i Denote the step size parameter of power station i.
[0037] In one implementation, the process of updating the dual variable of power station i based on the first update strategy, the second set of auxiliary variables, and the second set of dual variables to obtain the second updated dual variable of power station i can be expressed as:
[0038]
[0039] where λ i,k+1 Denote the second updated dual variable of power station i, Denote the value of the first updated dual variable of power station i, τ i Denote the step size parameter of power station i for dual variable update, D i Denote the strategy constraint matrix of power station i, w i,j Denote the weight of the adjacency matrix between power station i and its neighboring power station j.
[0040] Second, the embodiment of the present invention also provides a construction device for a multi-power station game based on random generalized Nash equilibrium search. The device includes the following components:
[0041] Problem construction module, which obtains the constraint conditions and objective functions of several power stations, and constructs the energy allocation problem of several power stations based on the constraint conditions and objective functions of the several power stations. The energy allocation problem of the several power stations is a distributed random generalized Nash equilibrium problem, and the state variables of the random generalized Nash equilibrium problem include decision variables, dual variables, and auxiliary variables;
[0042] Problem solving module, which solves the energy allocation problem of the several power stations through the variance reduction operator splitting algorithm to obtain the target strategy of the target power station. The process of solving the energy allocation problem of the several power stations through the variance reduction operator splitting algorithm specifically includes:
[0043] Construct the multiplier graph of several power stations and determine the weighted adjacency matrix;
[0044] For power station i, obtain the first set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the first set of auxiliary variables and the first set of dual variables of the first multiplier neighbor power stations corresponding to power station i based on the multiplier graph;
[0045] Determine the first random estimate of the decision variable of power station i based on the first set of policy information;
[0046] Based on the first random estimate of the decision variable, determine the first updated policy of power station i by the projected stochastic gradient descent method;
[0047] Update the auxiliary variables of power station i based on the first set of dual variables and the weighted adjacency matrix to obtain the first updated auxiliary variables of power station i;
[0048] Update the dual variables of power station i based on the first updated policy, the first set of auxiliary variables, and the first set of dual variables to obtain the first updated dual variables of power station i;
[0049] For power station i, obtain the second set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the second set of auxiliary variables and the second set of dual variables of the second multiplier neighbor power stations corresponding to power station i based on the multiplier graph;
[0050] Determine the second random estimate of the decision variable of power station i based on the second set of policy information;
[0051] Determine the second updated policy of power station i based on the first random estimate of the decision variable and the second random estimate of the decision variable;
[0052] Update the auxiliary variables of power station i based on the first updated auxiliary variables, the second set of dual variables, and the weighted adjacency matrix to obtain the second updated auxiliary variables of power station i;
[0053] Update the dual variables of power station i based on the first updated policy, the second set of auxiliary variables, and the second set of dual variables to obtain the second updated dual variables of power station i;
[0054] Re-execute the step of, for power station i, obtaining the first set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the first set of auxiliary variables and the first set of dual variables of the first multiplier neighbor power stations corresponding to power station i based on the multiplier graph until the number of re-executions reaches the preset number of times to obtain the target policy of the target power station.
[0055] In a third aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a memory, a processor, and a multi-power station game program stored in the memory and executable on the processor. When the processor executes the multi-power station game program, the steps of the multi-power station game method based on random generalized Nash equilibrium search as described above are implemented.
[0056] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where a multi-power station game program is stored on the computer-readable storage medium. When the multi-power station game program is executed by a processor, the steps of the multi-power station game method based on random generalized Nash equilibrium search as described above are implemented.
[0057] Beneficial effects:
[0058] The present invention provides a multi-power station game method and device based on random generalized Nash equilibrium search. The method constructs a distributed game framework, and each power station only needs to process information directly related to it, thereby reducing the need for centralized data processing, enhancing information security and privacy protection, optimizing the use of computing resources, and improving the operating efficiency of the entire system. The variance reduction operator splitting algorithm is used to handle random factors, reducing the error and instability in the calculation of decision variables under the influence of random events, thereby improving the accuracy and reliability of the strategy and ensuring that the optimization of energy distribution is more in line with actual needs. Secondly, by using iterative steps to gradually update strategies, auxiliary variables, and dual variables, each power station is allowed to refine its strategy in each round of calculation, making energy distribution more accurate, and allowing the system to adapt to market and environmental changes, improving the response speed and flexibility of the entire power system. Thirdly, the weighted adjacency matrix is used to consider and optimize the interaction between neighboring power stations, ensuring the coordination and consistency of the entire power network, not only optimizing local decisions but also enhancing the overall performance of the entire network, and ensuring the stability of the system in the face of large-scale adjustments. Finally, by introducing random factors to simulate uncertainties in the real world, such as price fluctuations or supply problems, the algorithm can effectively operate under incomplete information, ultimately improving the adaptability and robustness of decisions. Description of the Drawings
[0059] Figure 1 It is a flowchart of a multi-power station game method based on random generalized Nash equilibrium search provided by the present application.
[0060] Figure 2 It is a schematic diagram of a market example of a multi-power station game method based on random generalized Nash equilibrium search provided by the present application.
[0061] Figure 3A decision diagram for a multi-power station game method based on random generalized Nash equilibrium search provided by this application.
[0062] Figure 4 A graph showing the algorithm convergence results of a multi-power station game method based on random generalized Nash equilibrium search provided by this application at different sampling rates.
[0063] Figure 5 The first comparison graph of the convergence results of a multi-power station game method based on random generalized Nash equilibrium search provided by this application with the SFB algorithm.
[0064] Figure 6 The second comparison graph of the convergence results of a multi-power station game method based on random generalized Nash equilibrium search provided by this application with the SFB algorithm.
[0065] Figure 7 A structural schematic diagram of a construction device for a multi-power station game based on random generalized Nash equilibrium search provided by this application.
[0066] Figure 8 A structural schematic diagram of the terminal device provided by this application. Detailed implementation manners
[0067] This application provides a multi-power station game method and device based on random generalized Nash equilibrium search. To make the purpose, technical solutions and effects of this application clearer and more definite, the following further elaborates this application with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain this application and are not used to limit this application.
[0068] Those skilled in the art of this technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of this application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0069] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which this application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as here.
[0070] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.
[0071] In the existing power station market competition, power stations need to make optimal decisions in a dynamically changing market. The decisions of power stations must take into account various factors such as fuel costs, demand forecasts, environmental regulations, and market price fluctuations. In such an environment, the main problems faced by power stations include:
[0072] The price of the electricity market is affected by various factors, including but not limited to weather conditions, policy changes, and uncertainties in market demand;
[0073] Any interruption in the supply chain, from fuel supply to the acquisition of maintenance parts, can affect power generation efficiency and costs;
[0074] When making decisions, power stations usually only have access to local information, such as their own operating costs and output capabilities, while global information about the market is often unavailable;
[0075] Under rapidly changing market conditions, traditional decision-making algorithms are difficult to provide real-time optimal responses, which limits the competitiveness of power stations in the market.
[0076] The competition problem among the power stations can generally be reduced to a generalized Nash equilibrium problem and solved by the Forward-Backward (FB) algorithm. However, in the face of random factors such as market price fluctuations and supply chain disruptions, the distributed computing ability of the Forward-Backward algorithm is insufficient, making it difficult to ensure real-time performance and accuracy. In addition, under relatively weak assumptions, the convergence speed of the Forward-Backward algorithm is slow, and the strategy generation takes a long time, being unable to cope in the face of the need for emergency strategy adjustment.
[0077] To solve the above problems, in the embodiments of the present application, the constraint conditions and power station costs of several power stations are obtained, and an energy allocation problem of several power stations is constructed based on the constraint conditions and power station costs of the several power stations. Among them, the energy allocation problem of the several power stations is a distributed stochastic generalized Nash equilibrium problem, and the state variables of the stochastic generalized Nash equilibrium problem include decision variables, dual variables, and auxiliary variables; the energy allocation problem of the several power stations is solved by a variance reduction operator splitting algorithm to obtain the target strategy of the target power station. The method of the embodiment improves the accuracy and robustness of gradient estimation in the decision-making process through stochastic generalized Nash equilibrium (SGNEP) search. When facing the fluctuations of chips in the game situation among power stations, the present invention optimizes the gradient estimation by using variance reduction technology to ensure that power stations can make quick responses based on more accurate information. In addition, the operator splitting and projection technology can effectively handle the non-smooth elements in the problem, such as the irregularity of power station costs caused by unstable power generation demands, and realize distributed computing, so as to achieve the Nash equilibrium of the game process without centralized information sharing. In addition, through the improved forward-backward iteration steps, the convergence rate of the variance reduction operator splitting algorithm is further accelerated, and the applicable conditions of the variance reduction operator splitting algorithm are broadened, providing high flexibility and adaptability for each power station to find equilibrium in different environments and more extensive game scenarios.
[0078] The following further describes the application content by describing the embodiments in conjunction with the accompanying drawings.
[0079] This embodiment provides a multi-power station game method based on stochastic generalized Nash equilibrium search.
[0080] First of all, the power station generally refers to an individual acting in a strategic environment where power stations compete with each other and has the ability to choose different actions. In the field of distributed games of power stations, each power station has its own purpose (the purpose can be expressed as an objective function) and constraint conditions, and makes independent power generation strategies.
[0081] In the energy allocation problem of several power stations, the energy allocation problem of the several power stations essentially belongs to a non-cooperative game problem. In the embodiments of the present application, the constraint conditions of the power station usually refer to the rules or mathematical expressions that limit the power generation behavior of each power station, and the constraint conditions can be resource limitations, behavior rules, etc. The objective function of the power station is a mathematical function that characterizes the objective of the power station.
[0082] Such as Figure 1 shown, the multi-power station game method based on stochastic generalized Nash equilibrium search specifically includes:
[0083] S10. Obtain the constraint conditions and objective functions of several power stations, and construct the energy allocation problem of several power stations based on the constraint conditions and objective functions of the several power stations. Among them, the energy allocation problem of the several power stations is a distributed stochastic generalized Nash equilibrium problem, and the state variables of the stochastic generalized Nash equilibrium problem include decision variables, dual variables, and auxiliary variables.
[0084] In non - cooperative games, the objective function of each power station describes the benefits or costs of its ideal strategy. The strategic goal of a power station is to maximize or minimize the value of its objective function. The energy allocation problem of the several power stations is a type of problem in game theory, where the involved power stations make strategic actions independently without cooperation with other power stations. In the energy allocation problem of several power stations, each power station tries to optimize its own objective function, sometimes even at the expense of the objectives of other power stations.
[0085] The stochastic generalized Nash equilibrium problem is a game problem in the case where power stations make decentralized strategies. Each power station may be affected by random variables, and the goal is to find an equilibrium state where each power station cannot further optimize its objective function by unilaterally changing its strategy.
[0086] The state variables of the stochastic generalized Nash equilibrium problem are variables used by algorithms to find optimal strategies or equilibrium points, such as decision variables, dual variables, and auxiliary variables, etc. The decision variable is a term in optimization and game theory used to refer to the gradient of the objective function with respect to the strategy or strategy variable, usually used to refer to relatively rough estimates or gradient calculations that do not include all influencing factors. The dual variable is usually associated with the constraint conditions of the original problem in optimization problems and can be used in optimization algorithms to incorporate the constraint conditions into the objective function. The auxiliary variable is an additional variable introduced in the algorithm, whose purpose is to simplify the problem or decompose the problem into more easily handled parts. In game theory problems and optimization algorithms, auxiliary variables can be used to temporarily store values in calculations or promote algorithm convergence.
[0087] Specifically, the solution process of the stochastic generalized Nash equilibrium problem can be expressed as: each participating power station minimizes the objective function \(J(u^i, u^{-i})\) under the condition of satisfying the constraint \(u^i\in C(u^i)\), where the objective function and constraints of each participating power station depend on other power stations, and the objective function contains random factors. i ∈C i (u -i ) i (u i ,u -i )
[0088] In one implementation, the process of defining a multi-power station game problem based on the constraint conditions and objective functions of the several power stations includes:
[0089] S11. Define a variational equilibrium problem based on the constraint conditions and objective functions of the several power stations;
[0090] S12. Use a distributed environment to construct the variational equilibrium problem to obtain a distributed variational equilibrium problem;
[0091] S13. Add a random factor to the distributed variational equilibrium problem to obtain a stochastic generalized Nash equilibrium problem.
[0092] Specifically, the process of constructing the energy allocation problem of several power stations based on the constraint conditions and objective functions of the several power stations includes three stages.
[0093] The first stage (step S11): Definition of the problem description:
[0094] Consider a set of power stations \(i = \{1,\ldots,N\}\), and each power station selects a strategy in a non-cooperative game For each power station \(i\), its local objective function is It can be expressed as:
[0095] J i (u i ,u -i ) = f i (u i ,u -i ) + g i (u i )...(1)
[0096] where \(u -i =\text{col}(u_1,\ldots,u i-1 ,u i+1 ,\ldots,u N )\) is the vector of strategy combinations of other power stations except power station \(i\), is the non-smooth part in the objective function, such as the indicator function and the penalty function, etc., and satisfies:
[0097] (1) The function \(g i \) in the formula is a lower semi-continuous convex function, and for each \(i\) in \(g i \), its domain is a non-empty convex compact set.
[0098] And the smooth part \(f i (u i ,u -i )\) depends on the strategy \(u i \) of the power station itself and the strategies of some opponent power stations where is the interactive neighbor power station of i.
[0099] And it satisfies that the function f in formula (1) i (·, u -i ) is a convex function and is continuously differentiable for all i ∈ I and u -i .
[0100] The collective feasible region is defined as:
[0101] C = {u ∈ U | Du - b ≤ 0 m}...(2)
[0102] where U = U1 × … × U N , D = [D|…|D N ∈ R m×d D = [D1|…|D N ∈ R m×d , and each block reflects how power station i participates in the coupling constraints.
[0103] In addition, the non - empty feasible region C in formula (2) satisfies the Slater condition.
[0104] Then the game problem can be defined as:
[0105]
[0106] And denote the decision variable as:
[0107]
[0108] And it satisfies that the mapping F in formula (4) is monotonic and Lipchitz continuous.
[0109] Among them, the Karush - Kuhn - Tucker condition is a necessary and sufficient condition for u* to be a generalized Nash equilibrium. Among all generalized Nash equilibria, the goal is to achieve a variational equilibrium (VE), that is, a generalized Nash equilibrium is achieved when the dual variables of all power stations are consistent.
[0110] The target variational equilibrium satisfies the following rules:
[0111]
[0112] The second stage (step S12): Based on the variational equilibrium problem, construct a distributed structure:
[0113] Since the strategies of power stations usually rely on other power stations, it is necessary to construct a distributed environment that allows each power station to use only its local information for strategies, thereby solving the variational equilibrium problem. In other words, through a distributed network structure, multiple power stations can interact and communicate within the network and process the local information and resources of each power station. The construction process of the distributed structure includes:
[0114] Each power station i maintains its own local information U i 、D i and b i , and controls the strategy u i 、the local estimate of the dual variable and the auxiliary variable μ introduced to achieve consensus on the dual variable i ∈R m . The power stations communicate through an undirected connected graph G, and the weighted adjacency matrix of the undirected connected graph is W = [w i,j ∈R N×N , where w i,j > 0 if and only if (i, j) is an edge of the undirected connected graph G. The multiplier neighbor of power station i is denoted as Define the Laplacian matrix of the undirected connected graph as
[0115] Let u = col(u1,...,u N ), μ = col(μ1,...,μ N ), λ = col(λ1,...,λ N ), and take x = (u, μ, λ) ∈ X := R n ×R mN ×R mN as the state variable, and define the maximal monotone operator
[0116]
[0117] where D = diag{D1,...,D N}, Transform the variational equilibrium problem into the problem of finding x* ∈ zer(V + T).
[0118] The third stage (step S13): Introduce random factors into the variational equilibrium problem.
[0119] Since the uncertainties in the real world are inevitable, the above process of solving the variational equilibrium problem based on the distributed structure must also be able to handle random factors. The above variational problem is transformed into a stochastic generalized Nash equilibrium problem. To find the variational equilibrium by estimating the gradient for the stochastic characteristics of the objective function, the specific process can be expressed as follows:
[0120] Stochastic variable The objective function under the influence of Under the influence of the stochastic variable, F(u) and V(x) are respectively denoted as and Denote ξ k =col(ξ i,k ) i∈I .
[0121] Wherein, is an unbiased estimate of V(x), and the variance is bounded,
[0122] Since the distribution of the stochastic variable is unknown, the estimated value is needed to approximate the true gradient F(u). For each iteration step k, in equation (6), the estimated value is used to replace the true value to obtain the operator, which is expressed as:
[0123]
[0124] Denote the stochastic error
[0125] It should be noted that in the process of constructing the energy distribution problem of the several power stations, relevant information of the several power stations and the game scenario will be obtained. For example, for a market with a non-fixed size, there are a non-fixed number of power stations competing and cooperating.
[0126] The game problem can be applied to the real scenario. For example, the power stations can be several power stations participating in market competition. The real scenario can specifically be that N participating power stations C1,…,C N , and m markets M1,…,M m , and the participating power station i participates in the competition of n i markets by putting power products into the market. The market participation of all participating power stations is shown in the appendix Figure 2 . Each power station M j has a capacity limit r j , and the market price P j (u,ξ)=q j +p j (ξ)[S j (u)] σ is the inverse demand function, Sj (u) = ∑ i∈I [u i j is the total amount of power generation products for market M j . The production cost of power station i is There is an upper bound constraint on the power generation output, satisfying 0 < x i < Θ i .
[0127] In a specific real - world scenario, the number of participating power stations N = 20, the number of markets m = 7, and the random variable p j (ξ) follows a Gaussian distribution, with the mean randomly determined within [5, 7] and the variance being 1; q j is randomly determined within [2, 4]; each component of Θ i is randomly determined from [1, 1.5]; r j is randomly determined from [0.5, 1]; s i is randomly determined from [1, 8]; t i is randomly determined within [0.1, 0.6].
[0128] S20. Solve the energy allocation problem of the several power stations through the variance reduction operator splitting algorithm to obtain the target strategy of the target power station.
[0129] Specifically, the variance reduction operator splitting algorithm is an algorithm for the stochastic generalized Nash equilibrium problem, aiming to find an equilibrium strategy in the power station organization based on a preset number of iterative operations when the objective functions of each power station are affected by randomness. The algorithm includes a variance reduction strategy and an operator splitting strategy.
[0130] The variance reduction strategy is a statistical technical strategy that can reduce the gradient estimation error caused by random perturbations. In the process of searching for the Nash equilibrium, due to the uncertainty of the environment, it is necessary to accurately estimate the gradient of the objective function. The algorithm combines multiple estimated values through the variance reduction strategy, reduces the influence of randomness, improves the accuracy of gradient estimation, and thus more reliably guides the algorithm to find the equilibrium.
[0131] Specifically, in each iteration step k, power station i extracts a set of random samples where each element ξ i are independently and identically distributed, and the sample size S k > 1. Perform a mean estimation operation on the sampling, and the operation can be expressed as:
[0132]
[0133] To balance computational efficiency and accuracy, the variance reduction strategy adaptively adjusts the sampling rate at each iteration step to converge the gradient estimation, thereby replacing complex probability distributions and gradient calculations to achieve the convergence of the algorithm and improve the efficiency of multi-power plant game analysis.
[0134] By selecting a sampling rate sequence that satisfies and for any \(i\in I\) and \(u\in\mathbb{R}\) d , Equation (9) is an unbiased estimator of the local gradient. Substituting it into Equation (8) gives a stochastic estimate of the operator \(V\), and we have For
[0135] The operator splitting strategy allows the algorithm to separately handle the smooth and non-smooth parts of the objective function. By first dealing with the simpler smooth part and then considering the more complex non-smooth part, it simplifies the solution of the problem. The operator splitting strategy can be implemented using different mathematical operations. In the specific implementation of this embodiment, gradient descent and proximal mapping are used for implementation.
[0136] Specifically, the variance reduction operator splitting algorithm based on the operator splitting strategy is expressed using operator theory as follows:
[0137]
[0138] where \(X\) k \(=(u\) k ,\(\mu\) k ,\(\lambda\) k ), The matrix
[0139] \(\Psi=\text{diag}(\gamma\) -1 ,\(\sigma\) -1 ,\(\tau\) -1 )...(11)
[0140] contains step size information for three parts. Among them, \(\sigma,\tau\) are the same.
[0141] The stochastic generalized Nash equilibrium algorithm converges to the Nash equilibrium through a fully distributed manner using neighbor communication and two-step update iterations. First, in the first step of Equation (10), the proposed variance reduction stochastic approximation strategy is used to calculate which ensures the convergence of the gradient estimation with a smaller sample size without sacrificing computational efficiency. And the resolvent of the non-smooth part \(T\) is processed using projection thus solving the non-smoothness of the objective function. Based on the first-step FB iteration, a difference forward step is introduced, that is, the second step in Equation (10), and a stochastic estimation of the gradient is performed again using the result of the first step And through algebraic operations, it ensures that no excessive computational burden is introduced, achieving the effect of accelerating convergence, enabling the algorithm to reach a convergence rate of O(1 / k) under the monotonic assumption and a convergence rate of O(q k ) under the strong monotonic assumption, and relaxing the assumption condition from a strong monotonic operator to a monotonic operator, realizing adaptability under weak assumption conditions.
[0142] In addition, the results of the variance reduction algorithm splitting algorithm compared with other operator splitting algorithms are as follows:
[0143] The convergence rate of the stochastic forward-backward algorithm in the stochastic scenario under the strong monotonic assumption is O(1 / k);
[0144] The convergence rate of the stochastic forward-backward algorithm in the stochastic scenario under the monotonic assumption is
[0145] The convergence rate of the variance reduction operator splitting algorithm in the stochastic scenario under the strong monotonic assumption is O(q k );
[0146] The convergence rate of the variance reduction algorithm splitting algorithm in the stochastic scenario under the monotonic assumption is O(1 / k).
[0147] From the above comparison results, it can be known that the variance reduction operator splitting algorithm described in this application can not only handle randomness problems, but also be applicable to weaker assumption conditions and achieve a faster convergence rate.
[0148] In addition, the convergence analysis of the variance reduction operator splitting algorithm is as follows:
[0149] First, the operator V: X → X is maximally monotone and satisfies l V =(l + 2κ + |D|)-Lipschitz continuous; and the operator X is maximally monotone.
[0150] Among them, the operator V can be divided into V = V1(x) + V2(x), Both are maximally monotone operators. At the same time, V1 is l1=(l + κ)-Lipschitz continuous, and V2 is l2=(|D| + κ)-Lipschitz continuous. Therefore, V is l1 + l2 = l V -Lipschitz continuous.
[0151] zer(V + T) is the variational equilibrium of the game problem that satisfies the KKT condition (5).
[0152] Denote the stochastic process Define the σ-algebra F k : = σ(x0, ξ0, …, ξ k-1 , η0, …, ηk-1 ),G k := σ(F k ∪ σ(ξ k ))。The central error process U k := A k - E[A k |F k , W k := B k - E[B k ∣G k . Therefore, it can be obtained that:
[0153] E[U k ∣F k = E[W k ∣F k = 0,
[0154] Secondly, the residual function For each Ψ > 0,
[0155] Since r Ψ (x) = 0 if and only if which is equivalent to:
[0156] x - Ψ -1 V(x) ∈ (I + Ψ -1 T)(x)
[0157] That is, it is equivalent to:
[0158] 0 ∈ V(x) + T(x)
[0159] Thirdly, v k , u k , δ k , ψ k are random variables defined on the σ - algebra F k , and the following formula expressions are obtained:
[0160]
[0161] where, and v ≥ 0 is a random variable.
[0162] In addition, considering a sequence {X k} generated by equation (10), there exists X* ∈ zer(V + T), then for all k ≥ 0, there is
[0163]
[0164] From Y k and Xk+1 By the definition of
[0165]
[0166] where \(u\) k \(= V(X\) k ). Since \(0\in V(X^*) + T(X^*)\), we have \(u^*+v^* = 0\), where \(u^* = V(X^*)\) and \(v^*\in T(X^*)\). Then we obtain the equation
[0167]
[0168] Substituting into the definition of the residual function we get
[0169]
[0170] where the last inequality holds because is a non-expansive operator. Thus, we have
[0171]
[0172] According to equation (12), we have
[0173]
[0174] The above equation is further relaxed to:
[0175]
[0176] Substituting into equation (14), we get
[0177]
[0178] Taking the conditional expectation under the condition of \(F\) k and according to the tower property, we obtain:
[0179]
[0180] Based on this, for any \(X^*\in\mathcal{X}^*\), \(\{\|X\) k - X^*\| \} converges almost surely, and is almost surely bounded.
[0181] Therefore \(\{X\) k \} is almost surely bounded and has a convergent subsequence. Considering any convergent subsequence of \(\{X\) k \}, with the index set \(K\) and the limit According to the continuity of \(r\) Ψ (\(\cdot\)), we have Then is a solution of 0 ∈ T(X), and thus satisfies
[0182] Since {||X k - X*||} converges almost surely for any X* ∈ X*, thus converges almost surely, with the unique limit being 0. Therefore, the entire sequence {X k} converges almost surely to
[0183] In summary, in the variance reduction operator splitting algorithm described above, for a sequence {X k} generated by Equation (10), take a non-decreasing sequence of sampling rates {S k} that satisfies Take the step-size matrix Ψ to satisfy Then {X k} converges almost surely to the variational equilibrium.
[0184] In one implementation, the solution of the energy allocation problem of the several power stations by the variance reduction operator splitting algorithm specifically includes:
[0185] S2001. Construct the multiplier graph of the several power stations and determine the weighted adjacency matrix.
[0186] Specifically, the multiplier graph represents the relationship graph formed between power stations due to shared constraints. In the multiplier graph, nodes represent power stations, and edges are used to represent the relationship of mutual influence between power stations due to shared constraints. The multiplier graph is used to coordinate the dual variables (also called multipliers) in distributed optimization, and the dual variables are used to represent shared resources or other forms of coupling constraints in the generalized Nash equilibrium problem.
[0187] The weighted adjacency matrix can be expressed as W = [w i,j ∈ R N×N , where the weights of the matrix are evenly distributed.
[0188] S2002. For power station i, based on the multiplier graph, obtain the first strategy information set of the interactive neighbor power stations corresponding to power station i, as well as the first auxiliary variable set, the first dual variable set of the first multiplier neighbor power stations corresponding to power station i.
[0189] Among them, the first auxiliary variable set contains several first auxiliary variables, the first dual variable set contains several dual variables, and the first strategy information set contains several first strategy information, where several dual variables are Lagrange multiplier values.
[0190] Specifically, the i-th power station obtains the first strategy information u from the interactive neighbor power stations j,k and obtains it from the multiplier neighbor obtain the multiplier information λ j,k and μ j,k , where λ j,k is the first dual variable of power station j, which is a neighbor of power station i, in the current iteration round k, and μ j,k is the corresponding first auxiliary variable of neighbor j of power station i.
[0191] S2003. Determine the random estimate of the first decision variable of power station i based on the first policy information set.
[0192] Specifically, the step S2003 can be expressed as:
[0193]
[0194] where represents the decision variable estimated by power station i according to the first policy information set u of all power stations and the current random event ξ k in the iteration step k, S i,k represents the sampling rate of iteration step k, that is, the number of random samples used in this step, k represents the normalization factor, which is used to ensure that the decision variable estimate is the average value, represents the gradient of the policy of power station i, represents the objective function of power station i, u represents the first policy information set of power station i, k and represents a column of randomly drawn variables.
[0195] In other words, step S2003 means that in iteration step k, in order to estimate the decision variable of power station i, calculate S k gradient samples of its policy, each sample corresponding to a set of random events take the average value of the sample gradients as the estimate of the decision variable. Calculate the decision variable in an environment with random fluctuations through the above estimation method, and use it to iteratively update the policy set of the power station.
[0196] S2004. Determine the first updated policy of power station i through the projected stochastic gradient descent method based on the random estimate of the first decision variable.
[0197] Specifically, the step S2004 can be expressed as:
[0198]
[0199] where represents the operation of the proximal operator, which is used to handle functions g containing non-smooth (such as regularization terms) iMethod for variables in the function of, u i,k Denotes the policy variable before the k-th iteration, γ i Denotes the step size Denotes the transpose of the coupling constraint coefficient matrix related to power station i, λ i,k Denotes the dual variable related to the coupling constraint, which is used to incorporate the influence of the constraint condition during the optimization process
[0200] In other words, the said step means that in each iteration, power station i first estimates according to the current policy vector u k and the first decision variable to calculate a gradient descent step of its policy. Then, considering the influence of the coupling constraint reflecting the influence of the policies of other power stations on the policy of power station i. After that, the proximal operator is used to adjust the result to ensure that the updated policy meets the requirements of the non-smooth term g i Finally, the obtained is the updated intermediate policy value
[0201] S2005. Update the auxiliary variable of power station i based on the said first set of dual variables and the weighted adjacency matrix to obtain the first updated auxiliary variable of power station i
[0202] Specifically, the said step S2005 can be expressed as
[0203]
[0204] where Denotes the updated value of the auxiliary variable of power station i after the k-th iteration (the first updated auxiliary variable), μ i,k Denotes the auxiliary variable of power station i before the k-th iteration, σ i Denotes a positive proportionality coefficient or step size parameter used by power station i, w i,j Denotes the weight of the weighted adjacency matrix between power stations i and j, expressing the strength of the relationship between power stations, λ j,k and λ i,j Denotes the dual variables corresponding to power stations j and i at the k-th iteration (the first set of dual variables)
[0205] In other words, the said step S2005 means that in the k-th iteration, the first updated auxiliary variable of power station i is determined according to the difference between the dual variables of all neighbors j and the multiplier of power station i The difference of each dual variable (λ j,k -λ i,j ) is multiplied by the weight w between neighbors i,j , then all the weighted differences are added up, and finally multiplied by the step size σ iAnd add it to the original auxiliary variable μ i,k to obtain the first updated auxiliary variable The above steps ensure that when the entire multi-power station system executes the algorithm to seek the optimal solution, the behaviors of each power station are coordinated under the global constraints.
[0206] S2006. Update the dual variable of power station i based on the first update strategy, the first set of auxiliary variables, and the first set of dual variables to obtain the first updated dual variable of power station i.
[0207] Specifically, step S2006 can be expressed as:
[0208]
[0209] where represents the value of the first updated dual variable of power station i after the k-th iteration, λ i,k represents the dual variable of power station i at the k-th iteration (the first set of dual variables), τ i represents the positive step size parameter used by power station i, D i represents the coupling constraint coefficient matrix related to power station i, u i,k represents the first update strategy of power station i at the k-th iteration, b i represents the constraint value vector related to power station i, τ represents the global step size parameter, u i,k and u j,k represent the auxiliary variables of power stations i and j at the k-th iteration, {} represents the projection operation on the result within the brackets to make it in the non-negative m-dimensional real space The projection operation ensures that the dual variable is under the non-negativity constraint (usually, in an optimization problem, the dual variable (Lagrange multiplier) usually represents the price of a certain resource or commodity and must be non-negative).
[0210] In other words, the above steps mainly include calculating three parts:
[0211] 1. τ i (D i u i,k -b i ): Correct the error between the first update strategy and the inherent constraints of power station i;
[0212] 2. τ∑ j w i,j [(μ i,k -μ j,k )-(λ i,k -λ j,k )]: Correct the differences between the auxiliary variables and the dual variables of power station i and its neighbors;
[0213] 3. For λ i,k Sum the results of Step 1 and Step 2: Add the current dual variable of power station i to the correction terms of Step 1 and Step 2. Finally, project the result of the above sum onto the non - negative m - dimensional real - number space to ensure that the value of the dual variable does not become negative, thus satisfying the constraints in the dual space of the optimization problem and further assisting the primal problem in finding its optimal solution.
[0214] S2007. For power station i, obtain the second policy information set of the interactive neighbor power stations corresponding to power station i, as well as the second auxiliary variable set and the second dual variable set of the second multiplier neighbor power stations corresponding to power station i based on the multiplier graph.
[0215] Specifically, in Step S2007, the i - th power station obtains the second policy information set again from the interactive neighbor power stations and obtains the second auxiliary variable set from the multiplier neighbors and the second dual variable set where the second policy information set contains several second policy information, the second auxiliary variable set contains several second auxiliary variables, and the second dual variable set contains several second dual variables.
[0216] S2008. Determine the second decision variable random estimate of power station i based on the second policy information set.
[0217] Specifically, Step S2008 can be expressed as:
[0218]
[0219] where represents the decision variable estimate of the objective function of power station i (the second decision variable random estimate) in the current iteration k, represents the second policy information set of power station i in the current iteration step, η i,k represents the random variable or random perturbation introduced for power station i in the current iteration step, and the random variable or random perturbation affects the objective function of power station i, S k represents the number of random samples used for power station i in the current iteration step k, represents the t - th random sample of power station i, represents the gradient of the objective function of power station i under the t - th random sample.
[0220] In other words, in step S2008, the gradient of each random sample of power station i is first calculated, and the gradients of all random samples are summed up and then divided by the total number of samples to obtain the average estimated value of the gradient. The method of estimating the gradient by sampling in this step is often used to solve optimization problems when the objective function contains randomness, such as noisy data or changing market conditions. Through this step, more stable gradient information can be obtained, so as to better find the optimal solution or reach an equilibrium state.
[0221] S2009. Determine the second update strategy of power station i based on the first decision variable random estimate and the second decision variable random estimate.
[0222] Specifically, step S2009 can be expressed as:
[0223]
[0224] where u i,k+1 represents the strategy of power station i at iteration step k + 1 (the second update strategy), represents the second strategy information at iteration step k, γ i represents the step size coefficient of power station i, which is used to control the size of the update step, represents the current estimated value of the gradient of the objective function of power station i under the first strategy information and random variables obtained through decision variable estimation, represents the estimated value of the gradient of the objective function of power station i under the second strategy information set and random variables, represents the transpose of the coupling constraint coefficient matrix associated with power station i, λ i,k and respectively represent the dual variables of power station i in the first and second.
[0225] In other words, in this step, the strategy variable of power station i is updated by using the decision variable estimation of the current step and the intermediate step as well as the change of the dual variable. The update step includes adding a correction amount obtained by comparing the current and intermediate decision variable estimations and the correction amount brought by the constraint conditions to the intermediate strategy variable value. In this way, power station i adjusts its strategy according to the objective function, constraint conditions and random influences in each iteration, in order to find the optimal or balanced solution in the whole system.
[0226] S2010. Update the auxiliary variable of power station i based on the first update auxiliary variable, the second dual variable set and the weighted adjacency matrix to obtain the second update auxiliary variable of power station i.
[0227] Specifically, step S2010 can be expressed as:
[0228]
[0229] Among them, μ i,k+1 represents the updated value of the auxiliary variable (the second updated auxiliary variable) of power station i at the (k + 1)-th iteration step, and represents the value of the second auxiliary variable of power station i at the k-th iteration step.
[0230] In other words, the said step includes the following process:
[0231] 1. Between each pair of power stations i and j, calculate the change amount of the dual variable, that is, λ i,k minus the difference between their respective intermediate values j,k from the difference between λ
[0232] 2. Multiply each pair of change amounts calculated in the first step by the corresponding weight w i,j , and then accumulate the results of all power station neighbors;
[0233] 3. Multiply the sum of the above accumulations by the step size σ i ;
[0234] 4. Finally, add this product to the intermediate value of the current auxiliary variable to obtain the value of the auxiliary variable μ i,k+1 for the next iteration.
[0235] The purpose of updating the auxiliary variable is to consider the constraint conditions during the optimization process, make the strategies of power stations more coordinated, and thus help find the optimal or equilibrium strategy that satisfies the constraints.
[0236] S2011. Update the dual variable of power station i based on the first update strategy, the second auxiliary variable set, and the second dual variable set to obtain the second updated dual variable of power station i.
[0237] Specifically, the step S2011 can be expressed as:
[0238]
[0239] Among them, λ i,k+1 represents the second updated dual variable of power station i at step (k + 1), represents the value of the second dual variable of power station i at the k-th iteration step, τ i represents the step size parameter for the update of the dual variable of power station i, D i represents the constraint matrix related to the strategy variable of power station i, u i,k represents the value of the strategy variable of power station i at the k-th iteration step, represents the value of the second strategy variable of power station i at the k-th iteration step, Denotes the sum of all multiplier neighbors j of power station i.
[0240] In other words, the said steps include the following process:
[0241] 1. Starting from the second dual variable, first multiply the difference between the intermediate and current policy variables by the constraint matrix and the step size, and subtract the product from the intermediate dual variable;
[0242] 2. Then consider the update of the auxiliary variable of each neighbor j related to the coupling constraint. Subtract the product of the difference between the current and intermediate auxiliary variables multiplied by the connection weight from the result of the previous step;
[0243] 3. Finally, multiply the result of step 2 by the influence and weight related to the dual variable.
[0244] Through the above steps, the value of the second updated dual variable of power station i at the iteration step k + 1 is obtained. The said steps combine the direct influence brought by the constraint matrix and the indirect influence between neighbors introduced through the auxiliary variable and the dual variable to ensure that the algorithm follows the constraints and gradually reaches the optimal or balanced state.
[0245] S2012. Re - execute the step of obtaining the first policy information set of the interactive neighbor power stations corresponding to power station i, the first auxiliary variable set and the first dual variable set of the first multiplier neighbor power stations corresponding to power station i based on the multiplier graph for power station i until the number of re - executions reaches the preset number of times to obtain the target policy of the target power station.
[0246] Specifically, the preset number of times is the preset number of iterations. When the total number of iterations reaches the preset number of iterations, the state variable values of all power stations are obtained, and the target variational equilibrium is obtained. Among them, the policy values of some power stations are as shown in the appendix Figure 3 as follows.
[0247] In addition, the proof process of the convergence effect of the variance reduction operator splitting algorithm is as follows:
[0248] The relationship between the convergence effect of the variance reduction operator splitting algorithm and the sampling sequence is as shown in the appendix Figure 4 where the vertical coordinate is the residual of the state variable of the power station, and it can be observed from the appendix Figure 4 that the convergence effect of the variance reduction operator splitting algorithm is not sensitive to the selection of the sampling sequence. Therefore, good convergence effects can be obtained with a relatively small sample size.
[0249] At the same time, the comparison results of the variance reduction operator splitting algorithm with the stochastic forward - backward algorithm under two preset situations of strongly monotone operator and monotone operator are as shown in the appendix Figure 5 and the appendixFigure 6 As shown in the appended Figures 5 - 6 It can be observed that the variance reduction operator splitting algorithm not only adapts to weaker assumption conditions, but also achieves a faster convergence rate under both the strongly monotone operator and monotone operator preset cases.
[0250] In summary, the present application discloses a multi-power station game method based on stochastic generalized Nash equilibrium search. First, the variance reduction stochastic approximation method introduced by the method improves the accuracy of gradient estimation in an uncertain environment and enhances the robustness in practical scenarios with large fluctuations of random variables. Second, the method effectively handles the non-smooth objective function problem that often appears in game problems, decouples and processes complex non-smooth parts through operator splitting and projection techniques; in addition, the method is no longer limited to strongly monotone assumption conditions, and its adaptability and flexibility are significantly enhanced; the improved forward-backward iteration step and the introduced difference forward step enable the method to work effectively under more general conditions. Finally, the method not only maintains a fast convergence rate under monotone assumption conditions without increasing additional computational burden, but also achieves a higher convergence efficiency under strongly monotone assumption conditions. Overall, the method provides a fast and reliable solution method for game problems in multi-power station systems, and greatly improves the ability to seek coordinated action strategies in dynamic and complex environments.
[0251] Based on the above multi-power station game method based on stochastic generalized Nash equilibrium search, this embodiment provides a construction device for multi-power station game based on stochastic generalized Nash equilibrium search, as Figure 7 shown, the device includes:
[0252] A problem construction module 100, which obtains the constraint conditions and objective functions of several power stations, and constructs an energy allocation problem of several power stations based on the constraint conditions and objective functions of the several power stations, wherein the energy allocation problem of the several power stations is a distributed stochastic generalized Nash equilibrium problem, and the state variables of the stochastic generalized Nash equilibrium problem include decision variables, dual variables and auxiliary variables;
[0253] A problem solving module 200, which solves the energy allocation problem of the several power stations through a variance reduction operator splitting algorithm to obtain the target strategy of the target power station, wherein solving the energy allocation problem of the several power stations through the variance reduction operator splitting algorithm specifically includes:
[0254] Construct a multiplier graph of several power stations and determine a weighted adjacency matrix;
[0255] For power station i, based on the multiplier graph, obtain the first set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the first set of auxiliary variables and the first set of dual variables of the first multiplier neighbor power stations corresponding to power station i;
[0256] Based on the first set of policy information, determine a random estimate of the first decision variable of power station i;
[0257] Based on the random estimate of the first decision variable, determine the first updated policy of power station i by the projected stochastic gradient descent method;
[0258] Based on the first set of dual variables and the weighted adjacency matrix, update the auxiliary variables of power station i to obtain the first updated auxiliary variables of power station i;
[0259] Based on the first updated policy, the first set of auxiliary variables, and the first set of dual variables, update the dual variables of power station i to obtain the first updated dual variables of power station i;
[0260] For power station i, based on the multiplier graph, obtain the second set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the second set of auxiliary variables and the second set of dual variables of the second multiplier neighbor power stations corresponding to power station i;
[0261] Based on the second set of policy information, determine a random estimate of the second decision variable of power station i;
[0262] Based on the random estimate of the first decision variable and the random estimate of the second decision variable, determine the second updated policy of power station i;
[0263] Based on the first updated auxiliary variables, the second set of dual variables, and the weighted adjacency matrix, update the auxiliary variables of power station i to obtain the second updated auxiliary variables of power station i;
[0264] Based on the first updated policy, the second set of auxiliary variables, and the second set of dual variables, update the dual variables of power station i to obtain the second updated dual variables of power station i;
[0265] Re-execute the step of, for power station i, based on the multiplier graph, obtaining the first set of policy information of the interactive neighbor power stations corresponding to power station i, as well as the first set of auxiliary variables and the first set of dual variables of the first multiplier neighbor power stations corresponding to power station i, until the number of re-executions reaches a preset number to obtain the target policy of the target power station.
[0266] Based on the above multi-power-plant game method based on random generalized Nash equilibrium search, this embodiment provides a computer-readable storage medium. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the multi-power-plant game method based on random generalized Nash equilibrium search as described in the above embodiment.
[0267] Based on the above multi-power-plant game method based on random generalized Nash equilibrium search, the present application also provides a terminal device, as Figure 8 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a preset user guidance interface in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.
[0268] In addition, when the logical instructions in the above-mentioned memory 22 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0269] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the method in the embodiment of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the method in the above embodiment.
[0270] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes can also be transient storage media.
[0271] In addition, the specific processes of loading and executing multiple instructions by the above-mentioned storage medium and the terminal device have been described in detail in the above method, and will not be repeated here one by one.
[0272] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A multi-power station game method based on random generalized Nash equilibrium search, characterized in that: The method comprises: Obtaining constraints and costs of several power stations, and constructing energy allocation problems of several power stations based on the constraints and costs of the several power stations, wherein the energy allocation problem of the several power stations is a distributed stochastic generalized Nash equilibrium problem, and the state variables of the stochastic generalized Nash equilibrium problem include decision variables, dual variables, and auxiliary variables; Solving the energy allocation problem of the plurality of power plants by using a variance reduction operator splitting algorithm to obtain a target strategy of a target power plant, wherein the process of solving the energy allocation problem of the plurality of power plants by using a variance reduction operator splitting algorithm specifically includes: Construct a multiplier graph of several power stations and determine the weighted adjacency matrix; For power station i, based on the multiplier graph, a first strategy information set of an interactive neighbor power station corresponding to the power station i, and a first auxiliary variable set and a first dual variable set of a first multiplier neighbor power station corresponding to the power station i are obtained; Determining a first decision variable stochastic estimate for the power plant i based on the first strategy information set; Determining a first update strategy for the power station i by projected stochastic gradient descent method based on the first decision variable stochastic estimation; updating the auxiliary variables of the power plant i based on the first dual variable set and the weighted adjacency matrix to obtain first updated auxiliary variables of the power plant i; updating the dual variables of the power plant i based on the first update strategy, the first auxiliary variable set, and the first dual variable set to obtain first updated dual variables of the power plant i; For power station i, based on the multiplier graph, obtaining a second strategy information set of an interactive neighbor power station corresponding to the power station i, and a second auxiliary variable set and a second dual variable set of a second multiplier neighbor power station corresponding to the power station i; Determining a second decision variable stochastic estimate for the power plant i based on the second strategy information set; Determining a second update strategy for the power plant i based on the first decision variable stochastic estimate and the second decision variable stochastic estimate; updating the auxiliary variables of the power plant i based on the first updated auxiliary variables, the second dual variable set and the weighted adjacency matrix to obtain second updated auxiliary variables of the power plant i; updating the dual variables of the power plant i based on the first updating strategy, the second auxiliary variable set and the second dual variable set to obtain second updated dual variables of the power plant i; Re-execute the step of obtaining, for power station i, a first strategy information set of an interactive neighbor power station corresponding to the power station i, and a first auxiliary variable set and a first dual variable set of a first multiplier neighbor power station corresponding to the power station i based on the multiplier graph, until the number of re-executions reaches a preset number, so as to obtain a target strategy for the target power station; The process of determining the random estimation of the first decision variable of the power plant i based on the first strategy information set can be expressed as: Among them, u k represents the first policy information set, represents a random event t, S k represents the sampling rate of iteration step k, represents the power station cost of power station i, represents the cumulative gradient of the power station cost of power station i with respect to its strategy under random event t; The process of determining the first update strategy of the power plant i by projected stochastic gradient descent method based on the random estimation of the first decision variable can be expressed as: in, represents the proximal operator, γ i represents the step length, g i represents a non-smooth function, represents the transpose of the coupling constraint coefficient matrix associated with power station i, u i,k represents the strategy variable of power station i before the kth iteration, λ i,k represents the dual variable of power station i at the kth iteration; The process of determining the second update strategy of the power plant i based on the first decision variable stochastic estimation and the second decision variable stochastic estimation can be expressed as: in, represents the random estimate of the first decision variable, represents the random estimate of the second decision variable, represents the first updated dual variable of power station i; The process of updating the auxiliary variables of the power plant i based on the first updated auxiliary variables, the second dual variable set and the weighted adjacency matrix to obtain the second updated auxiliary variables of the power plant i can be expressed as: Among them, μ i,k+1 represents the second updated auxiliary variable of power station i, represents the first updated auxiliary variable of power station i, σ i represents the step size parameter of power station i; The process of updating the dual variables of the power plant i based on the first update strategy, the second auxiliary variable set and the second dual variable set to obtain the second updated dual variables of the power plant i can be expressed as: Among them, λ i,k+1 represents the second updated dual variable of power station i, represents the first update of the dual variable value of power station i, τ i represents the step size parameter of power station i regarding the dual variable update, D i represents the strategic constraint matrix of power station i, w i,j Represents the adjacency matrix weight between power station i and its neighbor power station j.
2. A multi-power station game method based on random generalized Nash equilibrium search according to claim 1, characterized in that: The process of constructing the energy allocation problem of the plurality of power plants based on the constraints and power generation costs of the plurality of power plants comprises: defining a variational equilibrium problem based on constraints of the plurality of power plants and power plant costs; Using a distributed environment to construct the variational equilibrium problem to obtain a distributed variational equilibrium problem; A random factor is added to the distributed variational equilibrium problem to obtain a random generalized Nash equilibrium problem.
3. A device for constructing a multi-power station game based on random generalized Nash equilibrium search, characterized in that: The device comprises the following components: A problem construction module, which obtains constraints and objective functions of several power stations, and constructs energy allocation problems of several power stations based on the constraints and objective functions of the several power stations, wherein the energy allocation problem of the several power stations is a distributed random generalized Nash equilibrium problem, and the state variables of the random generalized Nash equilibrium problem include decision variables, dual variables and auxiliary variables; The problem solving module solves the energy allocation problem by using a variance reduction operator splitting algorithm to obtain a target strategy for a target power plant, wherein the process of solving the energy allocation problem of the plurality of power plants by using a variance reduction operator splitting algorithm specifically includes: Construct a multiplier graph of several power stations and determine the weighted adjacency matrix; For power station i, based on the multiplier graph, a first strategy information set of an interactive neighbor power station corresponding to the power station i, and a first auxiliary variable set and a first dual variable set of a first multiplier neighbor power station corresponding to the power station i are obtained; Determining a first decision variable stochastic estimate for the power plant i based on the first strategy information set; Determining a first update strategy for the power station i by projected stochastic gradient descent method based on the first decision variable stochastic estimation; updating the auxiliary variables of the power plant i based on the first dual variable set and the weighted adjacency matrix to obtain first updated auxiliary variables of the power plant i; updating the dual variables of the power plant i based on the first update strategy, the first auxiliary variable set, and the first dual variable set to obtain first updated dual variables of the power plant i; For power station i, based on the multiplier graph, obtaining a second strategy information set of an interactive neighbor power station corresponding to the power station i, and a second auxiliary variable set and a second dual variable set of a second multiplier neighbor power station corresponding to the power station i; Determining a second decision variable stochastic estimate for the power plant i based on the second strategy information set; Determining a second update strategy for the power plant i based on the first decision variable stochastic estimate and the second decision variable stochastic estimate; updating the auxiliary variables of the power plant i based on the first updated auxiliary variables, the second dual variable set and the weighted adjacency matrix to obtain second updated auxiliary variables of the power plant i; updating the dual variables of the power plant i based on the first updating strategy, the second auxiliary variable set and the second dual variable set to obtain second updated dual variables of the power plant i; Re-execute the step of obtaining, for power station i, a first strategy information set of an interactive neighbor power station corresponding to the power station i, and a first auxiliary variable set and a first dual variable set of a first multiplier neighbor power station corresponding to the power station i based on the multiplier graph, until the number of re-executions reaches a preset number, so as to obtain a target strategy for the target power station; The process of determining the random estimation of the first decision variable of the power plant i based on the first strategy information set can be expressed as: Among them, u k represents the first policy information set, represents a random event t, S k represents the sampling rate of iteration step k, represents the power station cost of power station i, represents the cumulative gradient of the power station cost of power station i with respect to its strategy under random event t; The process of determining the first update strategy of the power plant i by projected stochastic gradient descent method based on the random estimation of the first decision variable can be expressed as: in, represents the proximal operator, γ i represents the step length, g i represents a non-smooth function, represents the transpose of the coupling constraint coefficient matrix associated with power station i, u i,k represents the strategy variable of power station i before the kth iteration, λ i,k represents the dual variable of power station i at the kth iteration; The process of determining the second update strategy of the power plant i based on the first decision variable stochastic estimation and the second decision variable stochastic estimation can be expressed as: in, represents the random estimate of the first decision variable, represents the random estimate of the second decision variable, represents the first updated dual variable of power station i; The process of updating the auxiliary variables of the power plant i based on the first updated auxiliary variables, the second dual variable set and the weighted adjacency matrix to obtain the second updated auxiliary variables of the power plant i can be expressed as: Among them, μ i,k+1 represents the second updated auxiliary variable of power station i, represents the first updated auxiliary variable of power station i, σ i represents the step size parameter of power station i; The process of updating the dual variables of the power plant i based on the first update strategy, the second auxiliary variable set and the second dual variable set to obtain the second updated dual variables of the power plant i can be expressed as: Among them, λ i,k+1 represents the second updated dual variable of power station i, represents the first update of the dual variable value of power station i, τ i represents the step size parameter of power station i regarding the dual variable update, D i represents the strategic constraint matrix of power station i, w i,j Represents the adjacency matrix weight between power station i and its neighbor power station j.
4. A terminal device, characterized in that: The terminal device includes a memory, a processor, and a multi-power station game program stored in the memory and executable on the processor. When the processor executes the multi-power station game program, the steps of the multi-power station game method based on random generalized Nash equilibrium search as described in any one of claims 1-2 are implemented.
5. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multi-power station game program, and when the multi-power station game program is executed by the processor, the steps of the multi-power station game method based on random generalized Nash equilibrium search as described in any one of claims 1-2 are implemented.
Citation Information
Patent Citations
Distributed market commodity supply scheduling method based on non-cooperative game
CN116050732A
Supply market production regulation and control method based on asynchronous distributed Nash equilibrium algorithm
CN116797251A