Method for solving non-zero and random game equilibrium of smart power grid

By constructing a multi-stage dynamic grid game model between power companies and users, iteratively update strategies and record experience distribution, the problem of traditional algorithms not convergence is solved, and the interaction between power companies and user strategies in smart grids is realized, accurately solving and balanced, and promoting efficient energy utilization and sustainable economic and social development.

CN120414486APending Publication Date: 2025-08-01ACAD OF MATHEMATICS & SYSTEMS SCIENCE - CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510484546.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

It is difficult for the existing technology to effectively solve non-zero and random game equilibrium in a smart grid environment. Traditional algorithms rely on difficult-to-verify functional structural properties, resulting in non-convergence problems.

Method used

Build a multi-stage dynamic grid game model between power companies and ordinary users, update the strategy and experience distribution through iterative algorithms, record the empirical distribution during strategy conversion, and directly output approximate Nash equilibrium under specific conditions, or use a linear model to estimate the linear equation to calculate intersection points.

Benefits of technology

Accurately solve the approximate Nash equilibrium of the power grid stochastic game, reasonably plan power production and consumption, promote efficient energy utilization, and promote sustainable economic and social development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120414486A_ABST
    Figure CN120414486A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a method for solving non-zero and random game equilibrium of a smart power grid, and the scheme can comprise the steps: constructing a multi-stage dynamic power grid game model between an electric power company and a common user, and defining game participants, states, actions, strategies and long-term income definitions. Through the steps of setting an initial value, iteratively updating a participant strategy and empirical distribution, recording empirical distribution during strategy conversion and the like, after iteration is finished, according to an empirical distribution set condition, or directly outputting approximate Nash equilibrium, or utilizing a linear model to estimate a linear equation and calculating an intersection point, the approximate Nash equilibrium is obtained. The method can effectively solve the random game equilibrium of the smart power grid, and promotes efficient utilization and sustainable development of energy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of optimal power system dispatching, and particularly to a method for solving the equilibrium of non-zero-sum stochastic games in smart grids. Background Art

[0003] The environment of the smart grid is constantly changing dynamically. Stochastic games in game theory have become an important theoretical framework for depicting the interactions of participants in a dynamic environment and are widely used in the research fields of economics, engineering, and artificial intelligence multi-agent reinforcement learning. In a stochastic game, players independently and simultaneously select actions at each time step and obtain stage rewards. The state transitions according to the current state and the probabilities of the selected actions, which highly matches the dynamic characteristics of the power grid.

[0004] From the perspective of the composition of the smart grid, it is necessary to consider both the user side and the supply side, involving heterogeneous participants such as energy service providers and users, and different participants have different impacts on the dynamic environment. Regarding the energy supply as the system state, a power grid stochastic game in which the energy provider determines the state change belongs to a single-controller stochastic game.

[0005] In a stochastic game model, the Nash equilibrium is a reasonable concept for predicting the behavior of participants, but its solution is quite challenging. Currently, the solution of the stochastic game equilibrium mainly relies on learning algorithms, such as the gradient descent / ascent OGDA, the extra-gradient method EG, the regularization method FTRL, the fictitious play FP, etc. However, the existing learning algorithms have limitations. Their feasibility highly depends on the structure of the long-term convergence function of the game, such as zero-sum, potential, convexity and concavity, etc., and these properties are difficult to verify in practice. Even for a simple 2×2×2 single-controller stochastic game, the evolution trajectory obtained by the fictitious play algorithm may not converge and may fall into a cycle, making it difficult to clarify the evolution law of its dynamic system. In view of this, there is an urgent need to propose a more efficient algorithm to solve the equilibrium. Summary of the Invention

[0006] Embodiments of this specification provide a method for solving the equilibrium of non-zero-sum stochastic games in smart grids to solve at least one of the above-mentioned technical problems.

[0007] To solve the above technical problems, the embodiments of this specification are implemented as follows:

[0008] Embodiments of the present invention provide a method for solving the equilibrium of non-zero-sum stochastic games in smart grids, including:

[0009] The method is implemented in a multi-stage dynamic power grid game model between a power company and ordinary users, and the method includes the following steps:

[0010] S1. Construct a power grid game model, specifically including:

[0011] S11. Define the game players. Player A is the power company, responsible for the production and supply of electricity. Its action selection determines the power supply status of the power grid at the next moment. Player B is an ordinary user, consuming electricity. Its action selection and that of Player A jointly determine the benefits of both parties.

[0012] S12. Define the power grid state sets s1 and s2, where the symbol s1 represents sufficient power supply, and the symbol s2 represents insufficient power supply.

[0013] S13. Define the action set of Player A as high-power power generation and low-power power generation, and the action set of Player B as buying more electricity and buying less electricity.

[0014] S14. Define the state transition probability function p(s'|s,a), which represents the probability that the power grid state transfers to s' after the power company selects action a in the power grid state s.

[0015] S15. Define the benefit functions and respectively represent the real-time benefits of both parties when the power company selects action a and the user selects action b in the power grid state s.

[0016] S2. Define the strategies of the players, specifically including:

[0017] S21. The stationary strategy of the power company is where the symbol represents the probability of selecting high-power power generation in the sufficient power supply state, represents the probability of selecting high-power power generation in the insufficient power supply state;

[0018] S22. The stationary strategy of the user is where represents the probability of selecting to buy more electricity in the sufficient power supply state, represents the probability of selecting to buy more electricity in the insufficient power supply state; The pure stationary strategies of Players A and B belong to the set H = {h1=(0,0), h2=(0,1), h3=(1,1), h4=(1,0)}, and these pure stationary strategies respectively correspond to the deterministic action combinations of the power company and the ordinary user in two states. Under the stationary strategy, the long-term benefits of each of Players A and B are defined as That is, the average expected benefit of the player within an infinite time span;

[0019] S3. Solve the approximate Nash equilibrium through an iterative algorithm, specifically including:

[0020] S31. Set the initial value The number of iterations T; where, denotes the initial empirical distribution of the power company, and the symbol denotes the initial empirical distribution of ordinary users;

[0021] S32. For t = 1, …, T, the player A selects the pure stationary strategy that optimizes its own payoff from the set H of pure stationary strategies according to the formula . If there are multiple pure stationary strategies that optimize the payoff simultaneously, the strategy with the smaller serial number is selected. Subsequently, update the empirical distribution according to . Among them, the symbol H = {h1 = (0, 0), h2 = (0, 1), h3 = (1, 1), h4 = (1, 0)} is the set of pure strategies;

[0022] The player B selects a pure strategy through . If there are multiple optimal pure stationary strategies, the one with the smaller serial number is selected, and then update the empirical distribution according to . Among them, are the empirical distributions of players A and B at time t respectively, are the pure strategies of players A and B at time t, and γ i (·) is the payoff function of player i, i = A, B;

[0023] S33. During the process of repeated games, if in a certain round t', player A or player B observes that the opponent switches from the i-th pure stationary strategy h i to the j-th pure stationary strategy h j , i ≠ j, then it records its own empirical distribution or at this time. The set composed of all such empirical distributions is denoted as

[0024] S34. If after the iteration ends, the number of types of the data set of a certain player is less than or equal to 1, directly output the empirical distributions of the two players in the last step, and the combination of the empirical distributions is the approximate Nash equilibrium;

[0025] S35. Otherwise, according to the data in , use the linear model to estimate the straight-line equation According to the multiple straight-line equations of each person calculate the intersection points and <q respectively, then output the result as the approximate Nash equilibrium of the power grid stochastic game.

[0026] One embodiment of the specification can at least achieve the following beneficial effects:

[0027] The technical solution of this application constructs a multi-stage dynamic power grid game model between power companies and ordinary users, and clarifies the elements of the game. During the solution process, by setting initial values, iteratively updating the strategies and empirical distributions of both parties, and recording the empirical distributions when strategy conversions occur. When the iteration ends, according to the situation of the empirical distribution set, if specific conditions are met, the approximate Nash equilibrium is directly output; otherwise, a linear model is used to estimate the straight-line equation and calculate the intersection point to obtain the approximate Nash equilibrium. Considering from the perspective of model construction, the technical solution of this application accurately depicts the strategic interaction and revenue relationship between power companies and users in the smart grid, making the game analysis fit the actual scenario. From the perspective of the solution process, by recording the empirical distribution and subsequent processing, it does not rely on the properties of function structures that are difficult to verify, effectively solves the problem of non-convergence of traditional algorithms, and can accurately solve the equilibrium. And accurate equilibrium results contribute to the reasonable planning of power production and consumption, promote the efficient utilization of energy, and drive the sustainable development of the economic society. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Figure 1 It is a schematic diagram of the strategy evolution trajectory of the traditional fictitious play algorithm in the stochastic game of the smart grid;

[0030] Figure 2 It is a schematic diagram of finding the intersection point by estimating the straight-line equation through collecting the empirical distribution in a method for solving the equilibrium of a non-zero-sum stochastic game in the smart grid provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the following will clearly and completely describe the technical solutions of one or more embodiments of this specification in conjunction with the specific embodiments and corresponding accompanying drawings of this specification. Obviously, the described embodiments are only some of the embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by one or more embodiments of this specification.

[0032] It should be understood that although terms such as first, second, and third may be used in this application document to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other.

[0033] Those of ordinary skill in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily essential for implementing the present invention.

[0034] Consider the following power grid game model, which is a multi-stage dynamic game between power companies and ordinary users. The game model is set as follows:

[0035] (1) Game players: Denote them as A and B. Among them, player A represents the power company, which can choose to generate electricity with different powers. Its different decisions will completely determine the power supply at the next moment. Therefore, it is in a controlling position in the game. For example, the power company can decide to generate electricity with high power or low power according to its own plan and grid demand, thus affecting the power supply status of the entire power grid. Player B represents ordinary users such as residents and factories, and its actions are the electricity purchase volume and electricity consumption volume at the current moment. It should be noted that the decision of player B will not affect the game state (i.e., the power supply status of the power grid), but will jointly determine its own income with the decision of player A. For example, a factory decides to purchase more or less electricity according to its production demand. This decision does not change the state of the power grid being well-supplied or lacking power, but will affect the factory's own electricity consumption cost and income together with the power generation decision of the power company.

[0036] (2) Game state: Divide the power supply situation of the power grid into two different states s1 and s2. Among them, state s1 represents the state of sufficient power supply, at which time the power grid power supply can meet the user demand, and state s2 represents the state of lacking power supply, that is, the power grid power supply is insufficient and cannot fully meet the user demand.

[0037] (3) Actions of game players: In each state, both player A and player B can choose two actions. The two actions of player A respectively represent that the power company decides to generate electricity with high power or low power. Among them, generating electricity with high power means increasing power output, which may cause the power grid to change from the state of lacking power supply to the state of sufficient power supply, or further improve the power supply when the power supply is sufficient. Generating electricity with low power means reducing power output, which may cause the power grid to change from the state of sufficient power supply to the state of lacking power supply, or maintain a lower power supply level when the power supply is lacking.

[0038] The two actions of player B respectively represent that the user buys more electricity or less electricity. Buying more electricity: In the state of sufficient power supply or lacking power supply, the user chooses to purchase more electricity to meet the high electricity demand for its own production or life. Buying less electricity means that the user chooses to purchase less electricity, corresponding to its own low electricity demand.

[0039] After player A chooses action a and player B chooses action b in the current state s, the state will transfer from s to state s' with a probability of p(s'|s,a), and the two players respectively obtain the current game income, which is and That is, Participant A obtains (reflecting the power company's revenue under this state-action combination, such as the comprehensive revenue of power generation cost and power sales revenue), and the participant obtains (reflecting the user's revenue under this state-action combination, such as the comprehensive revenue of electricity consumption cost and production benefit).

[0040] (4) Participant's strategy: Since the system state will randomly transition, the participant's strategy is complete, that is, the action choice of the participant must be determined in each state, usually characterized by a stationary strategy. The stationary strategy of Participant A is where represents the probability that the power company selects the first action, i.e., high-power power generation, in the first state, i.e., sufficient power supply, represents the probability that it selects the first action in the second state, i.e., power shortage. Similarly, the stationary strategy of Participant B, i.e., the user, is where represents the probability that the user selects the first action, i.e., buys more electricity, in the first state, i.e., sufficient power supply, represents the probability that the user selects the first action, i.e., buys more electricity, in the second state, i.e., power shortage. The pure stationary strategies of Participants A and B refer to: in each state, the participant will deterministically take a certain action. Therefore, each person's pure stationary strategy belongs to the set H = {h1 = (0,0), h2 = (0,1), h3 = (1,1), h4 = (1,0)}, where h1 = (0,0) means that the power company does not select the first action in both states (i.e., does not generate electricity at high power when the power supply is sufficient and does not generate electricity at high power when the power supply is short), and the user does not select the first action in both states (i.e., does not buy more electricity when the power supply is sufficient and does not buy more electricity when the power supply is short); and so on for other elements, and each element corresponds to the deterministic action selection combination of the participant in both states.

[0041] (5) Participant's long-term revenue: Under the stationary strategy, the long-term revenue of each person A and B is defined as the average value of their game revenues over an infinite time:

[0042]

[0043] This formula represents that for participant i (i = A is the power company, i = B is an ordinary user) under strategies σ A and σ B , as the number of game rounds T approaches infinity, the lower limit of the limit of its average expected revenue, which is used to measure the comprehensive revenue level of the participant in the long-term game process.

[0044] Suppose that in the t-th round of the game, the pure stationary strategy selected by Participant A is What Participant B chooses is After the t-round game ends, denote the empirical distribution of Participant A as The empirical distribution of Participant B is These empirical distributions will be used for strategy adjustment and analysis in the subsequent game process to gradually seek the equilibrium state of the game.

[0045] Through the above settings, this power grid game model comprehensively depicts the strategy choices, interaction relationships, and revenue situations of power companies and ordinary users in a multi-stage dynamic environment.

[0046] The algorithm proceeds in the following steps:

[0047] Step 1: Set the initial values The number of iterations T;

[0048] In this step and respectively represent the initial empirical distributions or initial strategy values of Participant A (such as a power company) and Participant B (such as an ordinary user) in the game. They are the starting points for subsequent iterative updates, symbolizing the initial definitions of the strategies of both parties at the beginning of the game. Subsequently, they will be gradually adjusted based on the revenue feedback in the game to approach the optimal strategy. T is used to limit the maximum number of rounds for the repeated execution of the game process and is a control parameter for the algorithm operation. By setting T, it is ensured that the game completes the strategy iteration within a finite number of steps, avoiding unlimited calculations, balancing the computational cost and result accuracy, and finally outputting an approximate equilibrium result when reaching T.

[0049] Step 2: For t = 1, …, T, Participant A selects his own action according to the following formula, that is, selects the pure strategy that can obtain the optimal revenue when facing If there are two pure stationary strategies that reach the optimum simultaneously, select the one with the smaller serial number:

[0050] After that, Participant A updates his empirical distribution according to the following formula

[0051] Similarly, the action selection and empirical distribution update formulas for Participant B are

[0052]

[0053] For the iteration rounds t = 1, …, T, Participants A and B optimize their strategies based on the current strategy states (empirical distributions) of each other in each round, and gradually approach the game equilibrium through repeated iterations.

[0054] First, the strategy selection and update of Participant A will be elaborated below. The participant uses the formula Select an action from the set of pure stationary strategies H. The logic is as follows: when facing the current empirical distribution of player B select the pure stationary strategy that maximizes its own long-term payoff γ i If multiple pure strategies achieve the optimal payoff simultaneously, determine the unique selection according to the rule of "prioritize the smaller serial number" to ensure the certainty of strategy selection. Update the empirical distribution using the weighted average method: where the newly selected strategy has a weight of and the historical experience has a weight of This mechanism reflects the "integration of new and old strategies" and gradually optimizes the strategy distribution.

[0055] Now, elaborate on the symmetric operation of player B: Player B selects a strategy through The logic is the same as that of player A, that is, based on the current empirical distribution of player A to maximize its own long-term payoff. At the same time, player B also updates through the weighted average: Realize the iterative optimization of player B's strategy experience, reflecting the dynamic interaction and strategic coordination adjustment process between the two parties in the game.

[0056] In summary, in step two, through the cycle of "strategy selection - experience update", players A and B optimize their own behaviors according to the strategies of the other party in each round of the game, reflecting the strategic interaction and optimization logic of the participants in the dynamic game.

[0057] Step three: In the process of repeated games, if in a certain round t', player A or player B observes that the opponent switches from the i-th pure stationary strategy h i to the j-th pure stationary strategy h j , i≠j, then he will record his own empirical distribution or at this time. The set composed of all such empirical distributions is denoted as Then, for different evolutionary processes, some will never appear; it is also possible that all will never appear.

[0058] In the process of repeated games, if in a certain round t′, player A or B observes that the opponent makes a strategy switch - that is, the opponent switches from the i-th pure stationary strategy h i to the j-th pure stationary strategy h j (i≠j), this is the trigger condition for the recording behavior. When the trigger condition is met, the player (A or B) who observes the strategy switch will record his own current empirical distribution (i≠j), this is the trigger condition for the recording behavior. When the trigger condition is met, the player (A or B) who observes the strategy switch will record his own current empirical distribution or All the empirical distributions recorded due to the strategy conversion of the opponent h i →h j respectively form sets and These sets retain the strategy states related to the opponent's strategy conversion and provide data support for analyzing the evolution of the game.

[0059] During different game evolution processes, the occurrences of the set are different: some may never appear due to the characteristics of the game (the opponent does not have the corresponding strategy conversion). In extreme cases, all may never appear, that is, the opponent's strategy is always stable and no conversion of i≠j occurs.

[0060] Step 4: If, after the iteration ends, the number of types of the data set of a certain participant is less than or equal to 1, directly output the empirical distributions of the two participants in the last step, and the combination of the empirical distributions is the approximate Nash equilibrium.

[0061] Taking the end of the iteration as the time point, check the number of types of the data set of a certain participant (such as Participant A). Here, the number of types refers to the number of different empirical distribution sets formed due to the opponent's strategy conversion (h i →h j , i≠j). If this value is less than or equal to 1, it means that the opponent's strategy conversion is small (or even no effective conversion occurs). After meeting the above conditions, directly output the empirical distributions of both sides in the last step( and ), and their combination is recognized as the "approximate Nash equilibrium", that is, if the number of types is extremely small, it indicates that the complexity of the strategy adjustment of both sides in the game is low, and the final empirical distribution tends to be stable. At this time, the strategy combination of both sides is close to the optimal response to each other, meeting the "mutual optimality of strategies" feature of the Nash equilibrium, so it is output as an approximate solution.

[0062] Step 5: Otherwise, according to the data in , use the linear model to estimate the straight-line equation According to the multiple straight-line equations of each person calculate the intersection points and respectively, and then output the result as the approximate Nash equilibrium of the power grid stochastic game.

[0063] This step is the processing logic when the determination condition in Step 4 is not met( the number of types is greater than 1), and it is based on the Data (including the empirical distributions of Participants A and B when the opponent's strategy changes), and a linear model is used to perform fitting analysis on the data. By mining the patterns in the data, the straight-line equations are estimated. These equations characterize the changing trends or equilibrium boundaries of the participants' own empirical distributions in scenarios where the opponent's strategy changes (h i →h j ). For each participant (A, B), their multiple straight-line equations constitute a system of equations. By solving the intersection points and of the system of equations, the equilibrium solutions of the strategies of both sides are obtained. The physical meaning of this intersection point is that in the scenario of the power grid stochastic game, and represent the strategy combinations of the power company and ordinary users respectively, enabling both sides to achieve the optimal revenue response under the opponent's strategy and satisfying the core feature of the Nash equilibrium that "the optimal pure strategy revenues are equal". Finally, is output as the approximate Nash equilibrium of the power grid stochastic game. This result is a mathematical abstraction of the complex game process. Through linear model fitting and solving the system of equations, the historical data of the dynamic strategy adjustments of both sides of the game is transformed into quantifiable equilibrium strategies, providing a theoretical solution for the strategy optimization of both the supply and demand sides of the power grid.

[0064] The technical solution of this application constructs a multi-stage dynamic power grid game model between the power company and ordinary users, clarifying the elements of the game. In the solution process, by setting initial values, the strategies and empirical distributions of both sides are iteratively updated, and the empirical distributions at the time of strategy conversion are recorded. When the iteration ends, according to the situation of the set of empirical distributions, if specific conditions are met, the approximate Nash equilibrium is directly output; otherwise, the straight-line equations are estimated using a linear model and the intersection points are calculated to obtain the approximate Nash equilibrium. From the perspective of model construction, the technical solution of this application accurately depicts the strategic interaction and revenue relationship between the power company and users in the smart grid, making the game analysis fit the actual scenario. From the solution process, by recording the empirical distributions and subsequent processing, it does not rely on the properties of function structures that are difficult to verify, effectively solving the problem of non-convergence of traditional algorithms and being able to accurately solve the equilibrium. And the accurate equilibrium results help to reasonably plan power production and consumption, promote the efficient utilization of energy, and drive the sustainable development of the economic society.

[0065] The above technical solution is illustrated below with a specific case. Consider a specific game, whose payoff matrix is presented in the following table form. Each cell in this table has two numbers, representing the payoffs of Participant 1 (such as the power company) and Participant 2 (such as ordinary users) respectively. The values in the parentheses below (such as (0.9; 0.1)) represent a certain probability distribution.

[0066]

[0067] There exists a unique Nash equilibrium in this game, namely (the strategy of player 1), (the strategy of player 2).

[0068] Figure 1 is a schematic diagram of the strategy evolution trajectory of the traditional fictitious play algorithm in the stochastic game of the smart grid; Figure 2 is a schematic diagram of finding the intersection point by collecting the empirical distribution to estimate the straight-line equation in a method for solving the equilibrium of a non-zero-sum stochastic game of the smart grid provided by the present invention. When the traditional fictitious play algorithm is adopted, the game evolution trajectories obtained are shown in the two lower left and right figures in the figure. The left figure is "frequency of player1" (the frequency of player 1), and the right figure is "frequency of player2" (the frequency of player 2). It can be seen that the trajectories of the traditional method do not converge to the Nash equilibrium of the game but keep circling, indicating that the traditional fictitious play algorithm cannot effectively solve the equilibrium in this game

[0069] According to step three, by marking the empirical distribution of oneself (points of different colors in the figure) when the opponent selects different actions, the empirical distribution when changing actions is obtained, and then the corresponding straight-line equation (the gray straight line in the figure) is restored. For example:

[0070]

[0071] By calculating the intersection points of these straight-line equations, the output is Similarly, can be output. They constitute the Nash equilibrium of the game, reflecting the improvement of the defects of the traditional algorithm by this method and the process of effectively solving the equilibrium.

[0072] The above advantages are generated because after the traditional fictitious play (steps one to two) process, a more detailed analysis of the game process is added. In a 2×2×2 single-controller stochastic game, the Nash equilibrium equivalence property is that the payoffs obtained from some actions selected by the opponent are the same. By recording which mixed strategies the player changes strategies when facing, the equation of the straight line where the points making the payoffs of the opponent's two actions equal can be restored. Then, by finding the intersection points of these straight lines, the cyclic state faced by the fictitious play when it does not converge can be jumped out, and thus the Nash equilibrium of the game can be obtained.

[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for solving the equilibrium of a non-zero-sum stochastic game in a smart grid, characterized in that, The method is implemented in a multi-stage dynamic power grid game model between power companies and ordinary users, and the method includes the following steps: S1. Construct a power grid game model, specifically including: S11. Define the game players. Among them, player A is the power company, responsible for the production and supply of electricity, and its action selection determines the power supply state of the power grid at the next moment; player B is an ordinary user, consuming electricity, and its action selection and player A jointly determine the benefits of both parties; S12. Define the power grid state sets s1 and s2, where the symbol s1 represents sufficient power supply, and the symbol s2 represents power shortage; S13. Define the action set of player A as high-power power generation and low-power power generation, and the action set of player B as more power purchase and less power purchase; S14. Define the state transition probability function p(s'|s,a), which represents the probability that the power grid state transfers to s' after the power company selects action a in the power grid state s; S15. Define the revenue function and respectively represent the real-time revenues of both the power company and the user when the power company selects action a and the user selects action b under the grid state s; S2. Define the strategies of the players, specifically including: S21. The stable strategy of the power company is where the symbol represents the probability of selecting high-power power generation in the state of sufficient power supply, represents the probability of selecting high-power power generation in the state of power shortage; S22. The user's stationary strategy is where represents the probability of choosing to buy more electricity under the condition of sufficient power supply, represents the probability of choosing to buy more electricity under the condition of power shortage; the pure stationary strategies of players A and B belong to the set H = {h1=(0,0), h2=(0,1), h3=(1,1), h4=(1,0)}, and the pure stationary strategies respectively correspond to the deterministic action combinations of the power company and ordinary users in two states; under the stationary strategy, the long-term benefits of each of players A and B are defined as That is, the average expected payoff of the participant over an infinite time horizon; S3. Solve the approximate Nash equilibrium through an iterative algorithm, specifically including: S31. Set the initial value The number of iterations T; where represents the initial empirical distribution of the power company, and the symbol represents the initial empirical distribution of ordinary users; S32. For \(t = 1,\ldots,T\), Participant A selects the pure stationary strategy that optimizes its own payoff according to the formula from the set \(H\) of pure stationary strategies. If there are multiple pure stationary strategies that optimize the payoff simultaneously, the strategy with the smaller index is selected. Subsequently, update the empirical distribution according to where the symbol \(H=\{h_1=(0,0), h_2=(0,1), h_3=(1,1), h_4=(1,0)\}\) is the set of pure strategies. Participant B passes through selecting a pure strategy. If there are multiple pure stationary strategy optima, the one with the smaller serial number is selected, and then according to update the empirical distribution, where are the empirical distributions of Participants A and B at time t respectively, are the pure strategies of Participants A and B at time t, and γ i (·) is the payoff function of Participant i, i = A, B; S33. During the process of repeated games, if in a certain round t', player A or player B observes that the opponent switches from the i-th pure stationary strategy h i to the j-th pure stationary strategy h j , where i ≠ j, then they record their own empirical distribution at this time or . The set composed of all such empirical distributions is denoted as S34. If, after the iteration ends, the number of types of the data set of a certain participant is less than or equal to 1, directly output the empirical distributions of the two participants in the last step, and the combination of the empirical distributions is the approximate Nash equilibrium. S35. Otherwise, according to the data in, use the linear model to estimate the straight line equation According to the multiple straight line equations of each person calculate the intersection points respectively and Then output the result which is the approximate Nash equilibrium of the power grid stochastic game.