Method for determining power market behavior strategy of multi-element user under new power system

By establishing a two-level decision optimization model for multiple users and a deep reinforcement learning algorithm, the problem of insufficient consideration of user collaborative behavior in the electricity market was solved, and the optimal allocation of power resources and reduction of carbon emissions were achieved in the new power system.

CN118944048BActive Publication Date: 2025-12-26NORTH CHINA ELECTRIC POWER UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410910678.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2025-12-26
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

Existing technologies have failed to fully consider the collaborative market behavior of diverse users in the electricity market, especially in new power systems with a high proportion of renewable energy integration, resulting in insufficient market operation mechanisms and difficulty in effectively absorbing renewable energy and reducing carbon emissions.

Method used

A two-level decision optimization model for multiple users is established. A clearing model for the wholesale and retail electricity market is constructed using a deep reinforcement learning algorithm. By combining the Markov decision process, the market behavior strategies of each entity are determined, including the reporting strategies of power generators and aggregators.

Benefits of technology

By making precise market-driven decisions, we can optimize the allocation of power resources, promote the stable operation and sustainable development of the new power system, reduce carbon emissions, give full play to the role of price signals, and guide users to increase production and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118944048B_ABST
    Figure CN118944048B_ABST
Patent Text Reader

Abstract

The application relates to the field of power systems, and particularly discloses a method for determining the power market behavior strategy of multiple users in a new power system, wherein the power market comprises a power wholesale market and a power retail market, the main bodies of the retail market include retail users on the demand side and aggregated users on the power supply side, the main bodies of the wholesale market include wholesale users on the demand side and power suppliers on the power supply side, and the wholesale users include the aggregated users on the power supply side of the retail market. The method comprises the following steps: constructing a double-layer decision optimization model of the power suppliers and the wholesale users in the power wholesale market and a double-layer decision optimization model of the aggregated users in the power retail market, with the maximum social welfare as a lower target and the maximum income as an upper target; then, a corresponding Markov decision process is established, and each model is solved based on the DQN algorithm to obtain the behavior decisions of each main body. The application fully considers the multiple participating main bodies and multiple market levels of the power market, and can promote the construction of the new power system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of power systems, and particularly relates to a method for determining power market behavior strategies of multi-element users under a new power system. BACKGROUND

[0002] With the proposal of the carbon peak and carbon neutralization targets, the construction of a new power system mainly based on new energy has become the development form of future energy and power. The high proportion of new energy access to the power system poses new challenges, and the market needs to play a more important role in resource allocation to promote the safe and stable operation and sustainable development of the system. Among them, under the background of the deepening of power market reform, the coordinated market behavior of users increasingly deeply affects the consumption of new energy and the efficiency and stability of the new power system, and the user side has gradually become the focus of market construction.

[0003] However, the construction of a unified power market in China is still in its infancy, and there are still many deficiencies in the market operation mechanism, especially on the user side. SUMMARY

[0004] Therefore, the present application provides a method for determining power market behavior strategies of multi-element users under a new power system, in an attempt to solve or at least alleviate the above problems.

[0005] According to one aspect of the present application, there is provided a method for determining power market behavior strategies of multiple users in a new power system, the power market including a power wholesale market and a power retail market, subjects of the power retail market including retail users on the demand side and aggregation users on the supply side, subjects of the power wholesale market including wholesale users on the demand side and power generators on the supply side, and the wholesale users including the aggregation users on the supply side of the power retail market, the method comprising: establishing a power wholesale market clearing model and a power retail market clearing model with the maximum social welfare as the target, and for each subject of the power wholesale market, constructing a power wholesale market transaction strategy model thereof with the maximum revenue as the target, and for each aggregation user, constructing a power retail market transaction strategy model thereof with the maximum revenue as the target; for each subject of the power wholesale market, constructing a two-level decision optimization model thereof in the power wholesale market by taking the power wholesale market transaction strategy model thereof as an upper model and the power wholesale market clearing model as a lower model, and for each aggregation user, constructing a two-level decision optimization model thereof in the power retail market by taking the power retail market transaction strategy model thereof as an upper model and the power retail market clearing model as a lower model; for any subject, establishing a Markov decision process of each wholesale user and power generator in the power wholesale market and a Markov decision process of each aggregation user in the power retail market by taking a bidding coefficient of each transaction period thereof as an action of each transaction period thereof, taking a revenue of a unit declared power of each transaction period thereof as a reward of each transaction period thereof, and determining a state of each transaction period thereof based on a declared price, a declared power, a winning power, and a clearing price of the power wholesale market of each transaction period thereof; and solving each two-level decision optimization model by using a deep reinforcement learning algorithm based on the established Markov decision processes to obtain a declaration strategy of each wholesale user and power generator in the power wholesale market and a declaration strategy of each aggregation user in the power retail market, the declaration strategy including a declared price and a declared power of each transaction period.

[0006] According to another aspect of the present application, there is provided a computing device comprising at least one processor and a memory storing program instructions configured to be executed by the at least one processor, the program instructions comprising instructions for implementing the method for determining power market behavior strategies of multiple users in a new power system according to the present application.

[0007] According to another aspect of the present application, there is provided a readable storage medium storing program instructions, which, when read and executed by a computing device, cause the computing device to implement the method for determining power market behavior strategies of multiple users in a new power system according to the present application.

[0008] In summary, the application fully considers the multi-element subject and multi-level market of the electricity market, takes the user side as the research object, focuses on the collaborative market behavior of the user, establishes a double-decision optimization model of multi-element users, and solves it by using a deep reinforcement learning algorithm to accurately and efficiently obtain the market behavior decision of each subject. In this way, carbon emissions can be effectively reduced and new energy can be consumed through the electricity market. And for the construction of the electricity market, the role of the price signal can be more fully played to guide users to actively increase production and efficiency, green power consumption and optimize power consumption methods, thereby promoting the optimal allocation of power resources and providing guidance for the construction of new power systems and electricity market mechanisms. In addition, the double-decision optimization model of the application considers factors such as carbon prices and demand response, so that the market behavior decisions of each subject obtained are more accurate and more representative. BRIEF DESCRIPTION OF DRAWINGS

[0009] To the accomplishment of the foregoing and related ends, certain illustrative aspects are described herein in connection with the following description and the annexed drawings. These aspects are indicative of various ways in which the principles disclosed herein can be practiced and all aspects and equivalents thereof are intended to be within the scope of the claimed subject matter. The above- and other objectives, features and advantages of the disclosure will be apparent from the following detailed description in conjunction with the accompanying drawings. Throughout the disclosure, like reference numerals are generally utilized to refer to like elements or features.

[0010] Figure 1 A structural block diagram of a computing device 100 according to an embodiment of the application is shown;

[0011] Figure 2 A flowchart of a method 200 for determining the electricity market behavior strategy of multi-element users under a new power system according to an embodiment of the application is shown;

[0012] Figure 3 A schematic diagram of the structure of the electricity market according to an embodiment of the application is shown;

[0013] Figure 4 A schematic diagram of the clearing of the centralized bidding market according to an embodiment of the application is shown;

[0014] Figure 5 A schematic diagram of the decision-making behavior flow of the subject under the DQN algorithm according to an embodiment of the application is shown;

[0015] Figure 6 A schematic diagram of the market operation flow of the new power system according to an embodiment of the application is shown;

[0016] Figure 7 A schematic diagram of the change in user neural network training error according to an embodiment of the application is shown;

[0017] Figure 8 Fig. 4 shows a schematic diagram of the iteration of the user's offer according to an embodiment of the present application;

[0018] Figure 9 Fig. 5 shows a schematic diagram of the average offer of the supply and demand sides and the average clearing price of the market according to an embodiment of the present application;

[0019] Figure 10 Fig. 6 shows a schematic diagram of the monthly market electricity supply and demand ratio and the wholesale market transaction electricity price according to an embodiment of the present application;

[0020] Figure 11 Fig. 7 shows a schematic diagram of the monthly retail market pricing according to an embodiment of the present application;

[0021] Figure 12 Fig. 8 shows a schematic diagram of the output of various types of units according to an embodiment of the present application;

[0022] Figure 13 Fig. 9 shows a schematic diagram of the impact of new energy generation on the market clearing price according to an embodiment of the present application;

[0023] Figure 14 Fig. 10 shows a schematic diagram of the proportion of new energy generation in the total generation and actual electricity purchase, as well as the total new energy generation and total electricity purchase according to an embodiment of the present application. DETAILED DESCRIPTION

[0024] Exemplary embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings; however, they are not limited to the embodiments set forth herein but can be implemented in various forms. The present embodiments are provided so that this disclosure will be thorough, and will fully convey the scope of the disclosure to those skilled in the art.

[0025] The new power system uses new energy such as wind power and photovoltaic power as the main power source, gradually replacing thermal power generation, and uses thermal power with a changed function, energy storage, and flexible load as flexible adjustment resources to ensure safe and stable operation of the system. Among them, the user side is a potential driver of carbon emissions in the new power system. When the electricity user participates in the transaction, the environmental impact of the demand side resources of the electricity market gradually becomes prominent. If there is no active participation of electricity users, it is difficult to effectively reduce carbon emissions and consume new energy through the electricity market. User participation in the electricity market can reduce their own electricity costs, and can more fully play the role of price signals in the construction of the electricity market, guide users to actively increase production and efficiency, green electricity, and optimize electricity usage, in order to promote the optimal allocation of power resources.

[0026] However, the existing research on market bidding mostly focuses on the power generation side, and mostly aims at a single market subject, without reflecting the user side bidding strategy. Moreover, the existing power user market strategy research mostly only considers a single market process, and few researches involve the whole process of power market transaction and the connection of markets at all levels. Therefore, the application provides a method for determining the power market behavior strategy of multiple users under a new power system, which fully considers the multiple participating subjects, multiple markets and transaction modes of the power market, and focuses on the strategic bidding behavior of users participating in the market.

[0027] The method for determining the power market behavior strategy of multiple users under a new power system can be executed in a computing device. Figure 1 A block diagram of the physical components (i.e., hardware) of the computing device 100 is shown. In a basic configuration, the computing device 100 includes at least one processing unit 102 and a system memory 104. According to an aspect, depending on the configuration and type of the computing device, the processing unit 102 can be implemented as a processor. The system memory 104 includes, but is not limited to, volatile (e.g., random access memory (RAM)), non-volatile (e.g., read-only memory (ROM)), flash memory, or any combination thereof. According to an aspect, the system memory 104 includes an operating system 105 and one or more program modules 106, which include a behavior strategy determination module 120 configured to execute the method 200 for determining the power market behavior strategy of multiple users under a new power system.

[0028] According to an aspect, the operating system 105 is suitable for controlling the operation of the computing device 100, for example. Furthermore, the example is practiced in conjunction with a graphics library, other operating systems, or any other application programs, and is not limited to any particular application or system. In Figure 1 The basic configuration is illustrated in FIG. 1 by those components within the dashed line 108. According to an aspect, the computing device 100 has additional features or functionality. For example, according to an aspect, the computing device 100 includes additional data storage devices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or tape. Such additional storage is illustrated in FIG. 1 by the removable storage 109 and the non-removable storage 110. Figure 1 According to an aspect, the removable storage 109 and the non-removable storage 110 are examples of the storage memory 104.

[0029] As stated above, according to an aspect, program modules are stored in the system memory 104. According to an aspect, the program modules can include one or more applications, and the application is not limited to the type of application, for example, the application can include: an email and contact application, a word processing application, a spreadsheet application, a database application, a slide show application, a drawing or computer-aided application, a web browser application, etc.

[0030] According to an aspect, examples can be practiced with electronic circuitry integrated on a single integrated circuit chip, with separate electronic components or in a computer system having a separate display and input device(s) 112, such as a computer or the like. According to an aspect, examples can be practiced with other computer system configurations, including hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. According to an aspect, examples can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices. Figure 1 According to an aspect, examples can be practiced with electronic circuitry integrated on a single integrated circuit chip, with separate electronic components or in a computer system having a separate display and input device(s) 112, such as a computer or the like. According to an aspect, examples can be practiced with other computer system configurations, including hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. According to an aspect, examples can also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.

[0031] According to an aspect, the computing device 100 can also have one or more input device(s) 112, such as a keyboard, a mouse, a pen, a voice input device, a touch input device, etc. Output device(s) 114, such as a display, speakers, a printer, etc. can also be included. The aforementioned devices are examples and others can also be used. The computing device 100 can include one or more communication connections 116 allowing communications with other computing devices 118. Examples of suitable communication connections 116 include, but are not limited to: RF transmitter, receiver, and / or transceiver circuitry; universal serial bus (USB), parallel, and / or serial ports.

[0032] The term computer readable media as used herein includes computer storage media. Computer storage media can include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, or program modules. The system memory 104, the removable storage 109, and the non-removable storage 110 are all computer storage media examples (i.e., memory storage.) Computer storage media can include Random Access Memory (RAM), Read Only Memory (ROM), Electronically Erasable Programmable Read Only Memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article of manufacture which can be used to store information and which can be accessed by the computing device 100. According to an aspect, any such computer storage media can be part of the computing device 100. Computer storage media does not include a carrier wave or other propagated data signal.

[0033] According to an aspect, communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. According to an aspect, the term "modulated data signal" describes a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0034] Figure 2 A flowchart of a method 200 for determining power market behavior strategies of multiple users under a new power system according to an embodiment of the present application is shown, which is suitable for being executed in a computing device (e.g. Figure 1 The computing device 100 shown).

[0035] It is to be noted that the power market of the present application includes a power wholesale market and a power retail market, and the subjects of the power market include subjects of the power wholesale market and subjects of the power retail market. The subjects of the power retail market include retail users on the demand side (i.e., users purchasing power from the power retail market) and aggregated users on the power supply side, and the subjects of the power wholesale market include wholesale users on the demand side (i.e., users purchasing power from the power wholesale market) and power suppliers on the power supply side, and the wholesale users on the demand side of the power wholesale market include aggregated users selling power to retail users in the power retail market. That is, the present embodiment is composed of two-level markets of the power wholesale market and the power retail market, and the two-level markets are connected through aggregated users.

[0036] Further, the aggregated user can include a power selling company, a micro-grid, a distributed park, and the retail user can include a small and medium-sized power user, and the wholesale user can also include a large power user. Specifically, the large power user can be a power user with a power supply voltage greater than or equal to a first predetermined value, and the small and medium-sized power user can be a user with a power supply voltage less than the first predetermined value. Wherein, the first predetermined value can be 10kV, of course, which is only an example. In addition, in some embodiments, the large power user and the small and medium-sized power user can also be divided by the annual electricity consumption, such as the user with the annual electricity consumption greater than or equal to a second predetermined value as a large power user, and the user with the annual electricity consumption less than the second predetermined value as a small and medium-sized power user, which is not limited by the present application. In addition, the aggregated user is described herein, which is different from the ordinary power user. The subject such as the power selling company and the grid-connected micro-grid usually contains multiple users, and accordingly can contain multiple equipment resources. The embodiment unifies the modeling of the subject as an aggregated user, and further unifies the market participation mode of the power grid enterprise such as the micro-grid and the power selling company, and models the subject as an aggregated user.

[0037] It can be seen that the power user of the present application includes large, medium, small ordinary power users, and aggregated users such as power selling companies and micro-grids. It should be noted that the user can only participate in one of the power wholesale transaction or retail transaction in a single transaction period, and therefore the operation structure of the market in the whole year is designed as follows: the large power user only participates in the power wholesale market transaction, and freely declares the price and quantity; the small and medium-sized power user only participates in the power retail market transaction, and only reports the quantity; the aggregated user participates in the power wholesale market, and its declaration mode is consistent with that of the large power user, and in the retail market, the aggregated user sells electricity to the retail user as a power selling party with unilateral pricing, and the distributed power supply is sold by the aggregated user. In addition, the power wholesale transaction in the power wholesale market includes medium and long-term transaction (monthly market) and spot transaction (day-ahead market and real-time market). The connection mode of the medium and long-term transaction and the spot transaction can adopt a decentralized market mode, mainly based on physical execution of medium and long-term contracts, with part of the electricity quantity (deviation electricity quantity) participating in the spot market, and the power generation and consumption parties themselves determining the power generation and consumption curve in the day-ahead stage, and adjusting the deviation electricity quantity through day-ahead and real-time balance transactions. For the power market structure of the present application, please refer to Figure 3 .

[0038] After clearly understanding the power market structure of the present application, next, the method 200 for determining the power market behavior strategy of the multi-element user under the new type of power system of the present application is described. As shown in Figure 2 , the method 200 starts from 210.

[0039] In 210, the power wholesale market clearing model and the power retail market clearing model are established with the maximum social welfare as the target, and for each subject of the power wholesale market, a power wholesale market transaction strategy model is constructed with the maximum income as the target, and for each aggregated user, a power retail market transaction strategy model is also constructed with the maximum income as the target.

[0040] First of all, it is explained here. The application adopts centralized bidding mode in power wholesale transaction (monthly medium and long-term transaction and spot transaction), and the clearing mode adopts bid by both supply and demand, and the queuing method is used for unified clearing. The main steps include: the power suppliers and users report the power and price within a specified time, the management subject accumulates the reported power to form the supply and demand step curve in the order of low to high selling price and high to low buying price, the large price difference is preferentially traded, a transaction pair is formed, the negative price difference cannot be traded, the arithmetic average of the reported price difference of the last traded buying side and selling side in all traded price difference pairs is the unified clearing price, and finally the unified settlement is made according to the price. If the reported power of the supply and demand sides is the same, a transaction pair is directly formed; if the reported power is not completely matched, the side with larger reported power continues to find the next transaction object to form a new transaction pair until all reported buying power or selling power is zero, or the user's bid is lower than the power supplier. Regarding the power suppliers, it is explained here that the power suppliers are mainly composed of thermal power generating units and new energy power generating units, and the new energy power generating units mainly include photovoltaic power generating units and wind power generating units, which are not limited by the application.

[0041] Under the previous mode of reporting power without price by the user side, the user enters the market as a price accepter, compared with which, under the power reporting and pricing mode of the embodiment, the autonomy of the user is greater, so that the role of the market can be fully played. In addition, in the case that the power suppliers and users both report price and power, the market management subject can maximize the social welfare as the target for clearing, and calculate the unified clearing price of the market, and the schematic diagram of the step curve formed by the bids of the power suppliers and sellers can be seen from Figure 4 .

[0042] Based on this, according to an embodiment of the application, a power wholesale market clearing model is constructed to represent the clearing process of centralized bidding of bid by both supply and demand in a single centralized bidding market transaction. Specifically, the maximum social welfare in the power wholesale market can be the maximum of the reported price difference of the wholesale users and power suppliers (i.e. the maximum of the reported price difference of the supply and demand sides), so that the power wholesale market clearing model can be specifically as follows.

[0043]

[0044] In the formula, y represents the reported price difference of the wholesale users and power suppliers, which can also be represented as y (f i ,d je k ), B Uj , B Ok , B Gi respectively represent the bidding price of large power users Uj, aggregated users Ok and power generators Gi in any trading period of the power wholesale market, N U , N O , N G respectively represent the number of large power users, aggregated users and power generators in the power wholesale market.

[0045] It has been pointed out above that the power wholesale transaction in the power wholesale market includes medium and long-term transaction (monthly market) and spot transaction (day-ahead market, real-time market), so in the power wholesale market, the present embodiment contains three time scales: monthly (month), day-ahead (day) and real-time (hour), so the present embodiment contains power wholesale market clearing models of three time scales. It should be noted that the power wholesale market clearing model of any time scale can be represented by the above formula, only the meaning of the trading period is different. For example, for the monthly, the trading period refers to the month (any trading period refers to any month), for the day-ahead, the trading period refers to the day (any trading period refers to any day), and for the real-time, the trading period refers to the hour (any trading period refers to any hour). Among them, the various models in the power wholesale market involved in the following are the same, and the subsequent will not be repeated. In addition, for the power retail market, the present embodiment only contains day-ahead time scale, so the trading period in the various models of the power retail market involved in the following refers to the day (any trading period refers to any day).

[0046] Next, the bidding price of power generators, large power users and aggregated users in any trading period of the power wholesale market will be specifically explained (the present embodiment adopts single segment pricing mode).

[0047] Power generators: In the present embodiment, capacity retention is not considered, and power generators declare power according to actual output, that is, when power generators have multiple units, they quote prices based on the weighted average marginal cost. In addition, in the actual power generation process, the marginal cost of renewable energy power generation is much smaller than that of thermal power, so the benchmark price in the present embodiment does not consider renewable energy. Based on this, the bidding price of power generators in any trading period of the power wholesale market can be specifically represented by the following formula.

[0048] B Gi (f i ) = f i C Gi

[0049]

[0050] In the formula, B Gi (f irepresents the declared price of the power supplier Gi in any trading period of the power wholesale market (i.e. the actual power bidding price), C Gi represents the power bidding reference price of the power supplier Gi, f i represents the power bidding coefficient (i.e. the bidding coefficient) of the power supplier Gi in any trading period of the power wholesale market, f min ≤f i ≤f max , when f i =1, the power supplier is based on the marginal cost bidding, f max , f min respectively represent the upper and lower limits of the power supplier bidding coefficient, C gx (p gx ) represents the marginal cost of the thermal power unit gx, p gx represents the output upper limit of the thermal power unit gx, m gx , n gx , l gx are respectively the quadratic coefficient, the linear coefficient and the constant term reflecting the characteristics of the thermal power unit gx. Among them, the value range of x is 1 to X, and X is the total number of thermal power units in the power supplier Gi. In addition, B Gi (f i ) is the same as the above B Gi .

[0051] Large power users: Large power users mainly purchase power from the power wholesale market, first conduct monthly transactions, and then declare based on the deviation between the monthly contract decomposed power and the actual demand power in the day-ahead market and the real-time market to realize real-time power balance. When the user participates in the power market transaction, the demand power and the willingness price are declared. Since the transaction power of the user is largely affected by the bidding and the actual load demand, only the declared price of the large power user is modeled. Assuming that the declared power is based on the actual demand power and the willingness price is a fixed value, the bidding of the large power user can be expressed as follows:

[0052] B Uj (d j ) = d j C Uj

[0053] In the formula, B Uj (d j ) represents the declared price of the large power user Uj in any trading period of the power wholesale market, C Uj represents the reference price of the large power user Uj in the power wholesale market, d j represents the bidding coefficient of the large power user in any trading period of the power wholesale market, wherein d min ≤d j ≤dmax When d j = 1, the user quotes according to the benchmark price, d max , d min respectively represent the upper and lower limits of the large power user quotation coefficient, wherein, B Uj (d j ) is the same as the above B Uj .

[0054] Aggregated user: The aggregated user first participates in the power wholesale market transaction, in the order of monthly transaction, day-ahead transaction, real-time transaction, and then sells the purchased power to small and medium-sized power users in the spot retail market. The actual demand power and demand response power of the aggregated user can be obtained by aggregating the power of the agent subject, and can be regarded as an ordinary power user when participating in the wholesale market transaction. The bidding model can be seen in the following formula.

[0055] B Ok (e k ) = e k C Ok

[0056] In the formula, B Ok (e k ) represents the declared price of the aggregated user Ok in any trading period of the power wholesale market, C Ok represents the benchmark price of the aggregated user Ok in the power wholesale market, e k represents the quotation coefficient of the aggregated user Ok in any trading period of the power wholesale market, wherein, e min ≤ e k ≤ e max When e k = 1, the aggregated user quotes according to the benchmark price, e max , e min respectively represent the upper and lower limits of the quotation coefficient of the aggregated user Ok in the power wholesale market. Among them, B Ok (e k ) is the same as the above B Ok .

[0057] In addition, the bidding price of the aggregation user in any trading period of the electricity retail market (i.e., the retail pricing) is described again. The bidding price of the aggregation user in the wholesale market affects the clearing price of the wholesale market, and the electricity price of the wholesale contract needs to be referred to when the retail electricity price is determined, so as to transmit the market signal of the wholesale market to the retail market. Therefore, the retail pricing strategy of the aggregation user is affected by the bidding strategy of the aggregation user in the wholesale market, and since the procurement electricity quantity of the aggregation user is mainly in the form of medium and long-term trading in a month, in some embodiments, the retail pricing of the aggregation user can be represented based on the monthly market clearing price of the aggregation user, and the bidding price of the aggregation user in any trading period of the electricity retail market can be specifically represented as the following formula.

[0058] b Ok (h k )=h k λ'

[0059] In the formula, b Ok (h k ) represents the bidding price of the aggregation user Ok in any trading period of the electricity retail market, λ' represents the electricity wholesale market clearing price of the trading period associated with the trading period in the electricity retail market, and specifically, it has been pointed out above that the time scale of the electricity retail market is day-ahead, and therefore λ' represents the electricity wholesale market clearing price of the last month of the month to which the trading period (i.e., any trading period mentioned above) belongs. h k represents the pricing coefficient (i.e., the bidding coefficient) of the aggregation user Ok in any trading period of the electricity retail market, wherein h min ≤h k ≤h ma x, h ma x, h min respectively represent the upper and lower limits of the pricing coefficient of the aggregation user Ok (i.e., the upper and lower limits of the bidding coefficient of the aggregation user Ok in the retail market). In addition, it has been pointed out above that the aggregation user purchases electricity in the electricity wholesale market and sells electricity in the electricity retail market, and therefore the number of aggregation users in the electricity retail market is the same as the number of aggregation users in the electricity wholesale market.

[0060] The above is the objective function of the electricity wholesale market clearing model. The electricity wholesale market clearing model also includes constraints, which are as follows.

[0061]

[0062] In the formula, X Gi , X Uj , X Ok respectively represent the winning electricity quantity of the power generator Gi, the large power user Uj, and the aggregation user Ok in any trading period of the electricity wholesale market, H Gi , H Uj , HOk Q represents the declared electricity volume of power generator Gi, large power user Uj, and aggregator Ok at any trading time in the wholesale electricity market. Uj Q Ok p represents the actual electricity demand of large electricity user Uj and aggregate user Ok at any trading time in the wholesale electricity market. min p max q represents the lower and upper limits of the bid prices set by the electricity wholesale market for participating entities, respectively. min q max These represent the lower and upper limits on the declared electricity volume stipulated by the electricity wholesale market for participating entities.

[0063] Thus, the electricity wholesale market clearing model, consisting of the objective function and constraints, is obtained. The electricity retail market clearing model is explained below. The electricity retail market also adopts a competitive market model, where multiple aggregated users participate in bidding. A market pricing mechanism is used, with aggregated users determining retail prices based on wholesale market prices, and retail users choosing retailers based on their needs and preferences. To reflect the pricing mechanism of the electricity retail trading market, a listing trading model is adopted. That is, aggregated users (on the sales side) submit unilateral bids and quantities, which are then listed on the electricity retail platform. Retail users submit quantity bids but do not quote prices. Retail contracts are formed according to the retail prices from low to high, and the contract price is the retail price of the aggregated users. This reflects the market behavior of retail users prioritizing the lowest-priced electricity retailers to achieve the lowest electricity cost. This process can be described as the retail market management entity clearing the market with the goal of minimizing electricity purchase costs. Based on this, according to an embodiment of the present invention, an electricity wholesale market clearing model is constructed to characterize a single clearing process, and the maximum social welfare in the electricity retail market is achieved by minimizing the electricity costs of retail users. The specific electricity retail market clearing model can be expressed as follows.

[0064]

[0065] In the formula, y' represents the electricity cost for retail users, which can also be expressed as y'(h k ), b Ok This represents the electricity price declared by aggregate user Ok during any trading period in the electricity retail market (as described in b above). Ok (h k (same), N O This indicates the number of aggregated users in the electricity retail market.

[0066] Furthermore, the electricity retail market clearing model also includes constraints, as detailed below.

[0067]

[0068] wherein X Ok represents the winning electricity quantity of the aggregated users Ok in the power retail market in any trading period, i.e. any trading period, N u represents the number of retail users in the power retail market, further, i.e. the number of small and medium-sized power users in the power retail market, Q uj' , Q u ' j' respectively represent the actual demand electricity quantity and the demand response electricity quantity of the retail user uj' in any trading period of the power retail market, p min , p max respectively represent the lower limit and the upper limit of the bid price of the power retail market prescribed by the power retail market participant.

[0069] Next, the market transaction strategy model will be described one by one.

[0070] The power wholesale market transaction strategy model of the power generator: in the present embodiment, the model considers the carbon cost of the power generator, and is specifically as follows.

[0071]

[0072] wherein M Gi represents the revenue of the power generator Gi in any trading period of the power wholesale market, λ represents the clearing price of any trading period of the power wholesale market, C Gi represents the benchmark price of the power generator Gi, X Gi represents the winning electricity quantity of the power generator Gi in any trading period of the power wholesale market, p CO2 represents the carbon price, B CO2 represents the carbon quota coefficient per unit of electricity generation, represents the weighted average carbon emission coefficient of all units of the power generator Gi.

[0073] Further, the power wholesale market transaction strategy model of the power generator also includes the constraint conditions corresponding to the above objective function, and is specifically as follows.

[0074] f min C Gi ≤B Gi ≤f max C Gi

[0075] The power wholesale market transaction strategy model of the large power user: since the users participating in the power wholesale market are mostly large industrial users, the electricity utilization maximization target is quantified as the output maximization target, and the subtraction of the cost is the revenue maximization target, and the upper limit of the electricity quantity allowed to be demanded by the user in response is set as a fixed proportion of the maximum electricity load. Based on this, the power wholesale market transaction strategy model of the large power user can be expressed as the following formula.

[0076] max M Uj = (I Uj - λ) X Uj + δ Q' Uj

[0077] wherein, M Uj represents the revenue of large power user Uj in any trading period of power wholesale market, I Uj represents the benefit per unit of electricity (electricity output) of large power user Uj, i.e. the benefit that large power user Uj can obtain by purchasing unit electricity, λ represents the clearing price in any trading period of power wholesale market, X Uj represents the winning electricity of large power user Uj in any trading period of power wholesale market, δ represents the compensation coefficient when participating in demand response to obtain compensation, and Q' Uj represents the demand response electricity of large power user Uj in any trading period of power wholesale market.

[0078] Further, the power wholesale market transaction strategy model of large power user also includes constraint conditions corresponding to the above objective function, which are as follows.

[0079]

[0080] wherein, η represents the upper limit of the proportion of the declared demand response electricity to the maximum electricity load.

[0081] The power wholesale market transaction strategy model of aggregated user: the bidding strategy of aggregated user in power wholesale market is the same as that of large power user, which is as follows.

[0082] max M Ok = (I Ok - λ) X Ok + δ Q' Ok

[0083]

[0084] wherein, M Ok represents the revenue of aggregated user Ok in any trading period of power wholesale market, I Ok represents the benefit per unit of electricity of aggregated user Ok, λ represents the clearing price in any trading period of power wholesale market, X Ok represents the winning electricity of aggregated user Ok in any trading period of power wholesale market, δ represents the compensation coefficient when participating in demand response to obtain compensation, and Q' Ok represents the demand response electricity of aggregated user Ok in any trading period of power wholesale market, and η represents the upper limit of the proportion of the declared demand response electricity to the maximum electricity load.

[0085] The aggregated user's electricity retail market transaction strategy model: the aggregated user with higher bidding in the electricity wholesale market has more stable electricity transaction volume, but may increase the clearing price of the wholesale market, resulting in the decrease of the difference between the wholesale market and the retail market, thereby affecting the income, while the electricity purchase volume may decrease when the wholesale market bidding is lower, thereby affecting the income, but the difference is larger, so that the aggregated user has more bidding space in the retail market. Therefore, the aggregated user needs to balance the electricity purchase price, electricity sale price, electricity purchase volume and electricity transaction volume to determine the strategy bidding coefficient. In addition, after the monthly and day-ahead electricity wholesale market transaction is completed, the aggregated user participates in the retail transaction. Specifically, the aggregated user's electricity retail market transaction strategy model is as follows.

[0086] max M Ok '=(b Ok -λ')X' Ok

[0087] In the formula, M Ok ' represents the income of the aggregated user Ok in any transaction period of the electricity retail market, X' Ok represents the winning electricity volume of the aggregated user Ok in any transaction period of the electricity retail market, b Ok represents the bidding price of the aggregated user Ok in any transaction period of the electricity retail market, and λ' represents the clearing price of the electricity wholesale market associated with the transaction period of the electricity retail market, further, λ' represents the clearing price of the electricity wholesale market of the last month of the month to which the transaction period of the electricity retail market belongs.

[0088] Further, the aggregated user's electricity retail market transaction strategy model also includes the constraint condition corresponding to the above objective function, which can be seen as follows.

[0089] h min λ'≤b Ok ≤h max λ'

[0090] At this point, the power wholesale market clearing model, the power retail market clearing model, the power retail market transaction strategy model of each aggregated user, and the power wholesale market transaction strategy model of each power supplier, large power user, and aggregated user are obtained. Subsequently, at 220, for each subject in the power wholesale market, the power wholesale market transaction strategy model of the subject is taken as an upper model, and the power wholesale market clearing model is taken as a lower model to construct a two-level decision optimization model of the subject in the power wholesale market. For each aggregated user, the power retail market transaction strategy model of the aggregated user is taken as an upper model, and the power retail market clearing model is taken as a lower model to construct a two-level decision optimization model of the aggregated user in the power retail market. The two-level decision optimization model of each power supplier in the power wholesale market, the two-level decision optimization model of each large power user in the power wholesale market, the two-level decision optimization model of each aggregated user in the power wholesale market, and the two-level decision optimization model of each aggregated user in the power retail market are obtained.

[0091] Subsequently, at 230, for any subject, the bidding coefficient of each transaction period of the subject is taken as the action of each transaction period of the subject, the revenue of each unit of declared power of each transaction period of the subject is taken as the reward of each transaction period of the subject, and the state of each transaction period of the subject is determined based on the declared power price, declared power, winning power, and clearing price of the power wholesale market. The Markov decision process of each wholesale user and power supplier in the power wholesale market, and the Markov decision process of each aggregated user in the power retail market are established.

[0092] Here, the Markov decision process is first described. It is a modeling of behavior process on discrete time series. The model mainly includes five elements: agent, environment, state, action, and reward. The agent is the subject, the environment is the system in which the agent is located (in the present application, it is a new power system including subject objects and environment objects), the state generally refers to the state of the entire system, the action is the action taken by the subject, and the reward is the feedback given by the environment to the agent.

[0093] Specifically, at each discrete time step t (in the present application, it is monthly, day-ahead, and real-time), the subject obtains the environment observation value (in the present application, it is market information and equipment information), determines the state s t ∈ S (state set), performs the action a t ∈ A (action set) according to the behavior strategy, and then moves to a new state s ( ∈ S with a state transition probability P = P ) s’|s,a t+1 ∈ S, obtains the immediate reward r t= R(s, a) and update the policy parameters and environment information, thereby forming a Markov chain such as s0, a0, r0, s1, a1, r1, s2, a2, r2,... This Markov decision process can be defined as (S, A, P, R, γ), where γ is a discount factor.

[0094] According to the above analysis, the behavior process of the agent can be divided into three stages of state observation, policy selection and state update. Meanwhile, the behavior of the agent in the new power system has obvious hierarchy and diversity. There are three types of actions: actions made immediately after obtaining the observation value of the environment, actions made independently after policy selection, and actions made according to policy selection after multi-agent collaboration and game. Therefore, the behavior hierarchy of the agent can be divided into the stress layer, the autonomous layer and the collaboration layer.

[0095] Next, the Markov decision process established in this embodiment will be described.

[0096] The Markov decision process of the power generator in the power wholesale market includes:

[0097]

[0098] a it = [f it ]

[0099]

[0100] In the formula, s it represents the state of the power generator Gi at the tth trading period in the power wholesale market, represents the declared price of the power generator Gi at the tth trading period in the power wholesale market, represents the maximum value in the declared prices of all power generators at the tth trading period in the power wholesale market, λ t represents the clearing price of the tth trading period in the power wholesale market, respectively represent the winning electric quantity and the declared electric quantity of the power generator Gi at the tth trading period in the power wholesale market, a it represents the action of the power generator Gi at the tth trading period in the power wholesale market, f jt represents the bidding coefficient of the power generator Gi at the tth trading period in the power wholesale market, r it represents the reward of the power generator Gi at the tth trading period in the power wholesale market, represents the revenue of the power generator Gi at the tth trading period in the power wholesale market, P i (s i(t+1) |s it ,a it) represents the state transition probability of power generator Gi in the tth trading period of the electricity wholesale market, further, that is, the power generator Gi performs action a in the tth trading period of the electricity wholesale market it The probability of transitioning from state s it to state s i(t+1) , in addition, the following is described about the transaction pair, which refers to whether the power generator Gi and the large power user form a transaction pair, whether the power generator Gi and the aggregation user form a transaction pair.

[0101] The Markov decision process of the large power user in the electricity wholesale market includes:

[0102]

[0103] a jt =[d jt ]

[0104]

[0105]

[0106] Wherein, s jt represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, represents the bidding price of the large power user Uj in the tth trading period of the electricity wholesale market, represents the minimum value of the bidding price of all large power users in the tth trading period of the electricity wholesale market, λ t represents the clearing price of the tth trading period of the electricity wholesale market, respectively represent the winning electric quantity, the actual demand electric quantity, the demand response electric quantity of the large power user Uj in the tth trading period of the electricity wholesale market, a jt represents the action of the large power user Uj in the tth trading period of the electricity wholesale market, d jt represents the bidding coefficient of the large power user Uj in the tth trading period of the electricity wholesale market, for the convenience of research, the bidding coefficient can be uniformly taken in its value set, and the bidding and bidding coefficient is represented as a discrete value, r jt represents the reward of the large power user Uj in the tth trading period of the electricity wholesale market, wherein the monthly transaction transaction electric quantity is much larger than the day-ahead and real-time spot market transaction electric quantity, so as to train the agent in the whole process of monthly transaction, day-ahead transaction and real-time transaction, the reward can be represented as the reward of unit reportable electric quantity, specifically, the obtained income is divided by the actual demand electric quantity of the user in the tth trading period (or the tth transaction), I Uj represents the degree of electric benefit of the large power user Uj, and δ represents the compensation coefficient when participating in demand response, P j(s j(t+1) |s jt ,a jt ) represents the state transition probability of large power user Uj in the tth trading period of the power wholesale market, and similarly, whether a transaction pair is formed here refers to whether a transaction pair is formed between large power user Uj and power generator.

[0107] The Markov decision process of the aggregated user in the power wholesale market includes:

[0108]

[0109] a kt =[e kt ]

[0110]

[0111] In the formula, s kt represents the state of the aggregated user Ok in the tth trading period of the power wholesale market, represents the declared price of the aggregated user Ok in the tth trading period of the power wholesale market, represents the minimum value of the declared price of all aggregated users in the tth trading period of the power wholesale market, λ t represents the clearing price of the tth trading period of the power wholesale market, respectively represent the winning electric quantity, actual demand electric quantity and demand response electric quantity of the aggregated user Ok in the tth trading period of the power wholesale market, a kt represents the action of the aggregated user Ok in the tth trading period of the power wholesale market, e kt represents the bidding coefficient of the aggregated user Ok in the tth trading period of the power wholesale market, r kt represents the reward of the aggregated user Ok in the tth trading period of the power wholesale market, I Ok represents the degree of electric benefit of the aggregated user Ok, δ represents the compensation coefficient when participating in demand response, P k (s k(t+1) |s kt ,a kt ) represents the state transition probability of the aggregated user Ok in the tth trading period of the power wholesale market, and similarly, whether a transaction pair is formed here refers to whether a transaction pair is formed between the aggregated user Ok and the power generator.

[0112] The Markov decision process of the aggregated user in the power retail market includes:

[0113]

[0114] a kt '=[h kt ]

[0115]

[0116] wherein s kt denotes the state of the aggregated user Ok in the tth trading period of the electricity retail market, denotes the bidding price of the aggregated user Ok in the tth trading period of the electricity retail market, denotes the maximum value among the bidding prices of all the aggregated users in the tth trading period of the electricity retail market, λ t denotes the clearing price of the transaction period associated with the tth trading period of the electricity retail market in the electricity wholesale market (further, the clearing price of the electricity wholesale market in the last month of the month to which the tth trading period of the electricity retail market belongs), denotes the winning electricity quantity of the aggregated user Ok in the tth trading period of the electricity retail market, k denotes the winning electricity quantity of the aggregated user Ok in the tth trading period of the electricity retail market, denotes the winning electricity quantity of the aggregated user Ok in the tth trading period of the electricity retail market, k denotes the winning electricity quantity of the aggregated user Ok in the tth trading period of the electricity retail market, k denotes the winning electricity quantity of the aggregated user Ok in the tth trading period of the electricity retail market, the target transaction period being the last month of the month to which the tth trading period of the electricity retail market belongs, or the last month of the month to which the tth trading day of the electricity retail market belongs), kt denotes the action of the aggregated user Ok in the tth trading period of the electricity retail market, h kt denotes the bidding coefficient of the aggregated user Ok in the tth trading period of the electricity retail market, r kt denotes the reward of the aggregated user Ok in the tth trading period of the electricity retail market, P k (s k(t+1) |s kt ,a kt denotes the state transition probability of the aggregated user Ok in the tth trading period of the electricity retail market, and it is understood that whether a transaction is formed depends on whether the aggregated user Ok and the retail user form a transaction pair.

[0117] At this point, the establishment of the Markov decision process is completed. Subsequently, at 240, based on the established Markov decision process, the deep reinforcement learning algorithm is used to solve each bi-level decision optimization model, and the bidding strategies of each wholesale user and power supplier in the electricity wholesale market and the bidding strategies of each aggregated user in the electricity retail market are obtained.

[0118] In order to solve the double-layer decision optimization model of each power supplier in the power wholesale market, the double-layer decision optimization model of each large power user in the power wholesale market, the double-layer decision optimization model of each aggregation user in the power wholesale market, and the double-layer decision optimization model of each aggregation user in the power retail market using the deep reinforcement learning algorithm (Deep Q-Network, DQN), the embodiment designs a decision-making behavior process of the subject under the deep reinforcement learning algorithm, which is specifically as follows.

[0119] In the first step, the network parameters of the evaluation network, the exploration factor, the target network update interval, and the current state of the subject are initialized. Further, in some embodiments, the update step of the evaluation network, the discount factor, the memory capacity, and the experience replay batch size also need to be initialized.

[0120] In the second step, the greedy strategy (ε-greedy strategy) is used to determine the declared price of the subject in the current trading period, and the exploration factor is updated. Specifically, the ε-greedy strategy for determining the declared price of the subject in the current trading period includes: generating a random number β∈[0,1], if the random number β is less than the exploration factor, a random bid is selected in the bid interval, otherwise the optimal bid is selected according to the evaluation network. In addition, the exploration factor can be updated by the following formula.

[0121]

[0122] wherein εt represents the exploration factor of the tth step (i.e., the exploration factor of the tth trading period); ε0 represents the initial exploration factor; εT represents the exploration factor at the end of the simulation, which can be a preset value; and T is the total number of iterations (steps), i.e., the total number of trading periods, which can be 8759. t start end

[0123] In the third step, the action of the declared price determined in the second step is executed to obtain the state of the next trading period of the subject and the reward, and the current state, action, reward, and state of the next trading period are saved as a sample to the memory bank. When saving the sample, it is first determined whether the current number of samples in the memory bank has reached the memory bank capacity. If not, the storage is directly performed. If it has reached, the sample stored in the memory bank the earliest is deleted, and then the storage is performed.

[0124] ​​​In the fourth step, the network parameters of the evaluation network are updated based on the first prediction value and the second prediction value. The first prediction value is the prediction value of the evaluation network for the current state and the current action, and the second prediction value is the maximum value of the prediction values of the target network for the subject state and all actions in the action set in the selected sample from the memory library. A batch of experience replay samples can be selected from the memory library. Further, the network parameters of the evaluation network can be updated by the following formula.

[0125]

[0126] L(w) = E s,a,r,s' [(r + γmax A Q(s', a'; w') - Q(s, a; w)) 2 ]

[0127] where w' and w respectively represent the updated network parameters and the current network parameters, a represents the update step, L(w) represents the loss function, E s,a,r,s' [] represents the expected value for a batch of experience replay samples (i.e., the expected value of a batch of experience replay samples), r represents the reward of the action of executing the determined bidding price, γ represents the discount factor, Q(s, a; w) represents the prediction value of the evaluation network when the input is the current state s and the current action a (i.e., the first prediction value), max A Q(s', a'; w') represents the maximum value of all prediction values obtained by inputting the state s' (which is the state in the selected sample from the memory library) and each action in the action set A in the target network (i.e., the second prediction value).

[0128] In the fifth step, it is detected whether the current step number reaches an integer multiple of the target network update interval. If not, no processing is performed. If yes, the network parameters of the evaluation network are copied to the target network, and it is judged whether the current step number reaches a preset value (whether the current step number reaches an integer multiple of the target network update interval or not, the judgment of whether the preset value is reached is performed).

[0129] In the sixth step, if the preset value is not reached, the current state is updated to the state of the next trading period, the current step number is increased by one, and the steps of bidding price determination, exploration factor updating, action execution, sample saving, network parameter updating, network parameter copying, and step number reaching preset value judgment (i.e., the above-mentioned second to fifth steps) are continued until the current step number reaches the preset value, and the process is ended.

[0130] For the above-mentioned double-layer decision optimization model, based on the established Markov decision process, through the above-mentioned process, the monthly, day-ahead, and hourly bidding price and bidding capacity of each large power user, aggregated user, and power generator in the power wholesale market, and the day-ahead bidding price and bidding capacity of each aggregated user in the power retail market can be obtained. Specifically, in some embodiments, the transaction period corresponding to each step can be set to one hour, and then for monthly transactions, the last hour of the last day of each month is carried out (the monthly double-layer decision optimization model is run), and for daily transactions, the last hour of each day is carried out (the daily double-layer decision optimization model is run).

[0131] In order to better understand the decision-making behavior process of the designed agent under the DQN algorithm, the following will be specifically described in combination with Figure 5 the established Markov decision process. For the model constructed by the present application, the agent state s, action a and reward r in the process can be seen from the above-mentioned Markov decision process, and the specific steps of the algorithm process are as follows:

[0132] (1) Generate evaluation network parameters w, if the current step step = 0, copy the evaluation network parameters to the target network, generate the exploration factor ε, update the step size α, the discount factor γ, the experience replay batch size k, the memory bank capacity Z and the network update interval N, if the current step step = 0, initialize the current state s of the agent.

[0133] (2) Action selection is performed using the ε-greedy strategy, that is, the agent selects the bid. The ε-greedy strategy is applied in the reinforcement learning algorithm, which is used to determine the probability of two action modes of random exploration of the environment and decision-making using the strategy set. The specific process is as follows: first, define an exploration rate ε ∈ [0, 1], and each time in the bid selection link, generate a random number β ∈ [0, 1], if the random number is less than ε, select a random bid in the bid range; otherwise, select the optimal bid according to the neural network. The purpose of using the ε-greedy strategy is to enable the agent to fully explore the environment at the beginning of iteration, so as to accumulate experience in the memory bank as much as possible. With the increase of experience, the exploration rate is gradually reduced, and the possibility of using the learned knowledge is increased. In this way, the needs of exploration and utilization can be balanced, and the learning and convergence process of the strategy is continuously optimized. Therefore, ε is reduced in a linear manner in each training round, and in the initial stage, the agent will tend to explore, and then with the learning, the proportion of using known information is gradually increased, that is:

[0134]

[0135] wherein, εt represents the exploration factor of the tth step, ε0 represents the initial exploration factor, ε represents the final exploration factor, and α represents the learning rate. t start end ​​T is the total number of iterations, and T is the exploration factor at the end of the simulation.

[0136] (3) Perform action a, i.e., report the electricity price. Participating in market bidding belongs to the collaborative layer behavior, because the result of performing action a needs to be embodied in the market clearing result obtained through multi-agent collaboration. According to the market clearing model, the clearing result is obtained, and then the next state s' and the reward r are obtained. Save the information [s, a, r, s'] to the memory bank, and if the memory bank data reaches the capacity Z, remove the earliest data put into the memory bank. It should be noted that all the embodiments belong to the collaborative layer behavior.

[0137] (4) Gradient descent process. The core of the DQN algorithm process is to train the neural network parameters w to approximate the value function. In this paper, gradient descent method is used to minimize the loss function, i.e., the deviation between the target value and the evaluation network output value, to continuously train and adjust the network parameters. For state s, the predicted value of the evaluation network output is Q(s, a; w), the reward of performing action a is r, the state transition is s', the target network output is Q(s', a; w'), k samples are selected from the memory bank, and if there are less than k samples, all information in the memory bank is taken out, then the loss function can be expressed as:

[0138] L(w)=E s,a,r,s' [(r+γmax a Q(s',a;w')-Q(s,a;w)) 2 ]

[0139] Next, the gradient of w is calculated, and finally w is updated:

[0140]

[0141] (5) Update the environment information, step+1, when the step reaches the network update interval N, copy the evaluation network parameters to the target network. In some embodiments, updating the environment information can specifically update the electricity of the agent, and further, the electricity in the contract signed by the agent. In the operation process of the new power system, the power supply and demand parties record the interaction results in the electricity contract as the medium for the implementation of the behavior between each other. At the same time, the interaction process between the agent and the environment is also recorded and executed in the form of generating and updating the electricity contract. In addition, in some embodiments, updating the environment information also includes updating the capacity of the unit and the load.

[0142] (6) Repeat steps (2)-(5) above until step=8759, the simulation period of one year ends, and the training ends.

[0143] Further, the simulation process of the market operation of the new power system is also designed, such as Figure 6The cycle of a single simulation is one year, the simulation step is refined to the hour level, the simulation process is pushed forward according to the time scale, and the whole process of market participants participating in the market transaction of the new power system is simulated, covering medium and long-term transactions, spot transactions, retail / distributed transactions, and auxiliary services and demand response. According to the time scale, medium and long-term transactions (monthly transactions), spot transactions, retail / distributed transactions, and finally auxiliary services and demand response are carried out. In each level of the transaction market, the transaction process can be refined into market participants' information declaration, market clearing, contract execution and settlement, real-time power determination and other behaviors. Among them, information declaration is the autonomous layer behavior of the subject, market clearing is the collaborative layer behavior, which is carried out according to the above established subject bidding model and market clearing model, contract execution and settlement, real-time power determination is the stress layer behavior of the subject, in each transaction link, multiple subjects identify external information, call behavior function to participate in the simulation process, and its behavior process is as shown in the above Figure 5 Since the simulation cycle is one year, annual transactions are not considered, and the predicted monthly output, daily output and real-time output curves of units and loads are obtained from empirical data. The specific overall development process of the power market is as follows:

[0144] (1) First, input the initial data, carry out monthly centralized bidding transactions, and the declared power of power suppliers and users is declared according to 80% of the monthly predicted power generation and consumption, and the price is determined according to the above-mentioned model and the deep reinforcement learning algorithm. The information is summarized and cleared by the management subject, the transaction results are recorded in the contract object, and the transaction is completed. After the transaction is completed, the environment information and the neural network parameters of the subject are updated, and the whole year has 12 monthly transactions, which are carried out when the step is the last hour of the last day of each month.

[0145] (2) The dispersed market mode is adopted to link medium and long-term transactions and spot transactions, and the physical execution of medium and long-term contracts is mainly used, and the deviation power of spot market is used. That is, in the day-ahead stage, the monthly contract power, the predicted power generation and the predicted demand power are decomposed into days, the contract power is decomposed according to the proportion of demand power of each day in the month, and the day-ahead centralized bidding transaction is carried out after the determination of power generation and demand power. The declared power is the day deviation power that is not satisfied after the decomposition of the contract, the price is obtained through the deep reinforcement learning algorithm, and the contract information and the neural network parameters are updated after the transaction is completed. The whole year has 365 day-ahead transactions, which are carried out when the step is the last hour of each day.

[0146] (3) After the day-ahead wholesale market transaction, the day-ahead retail market transaction is carried out, the wholesale power is retailed to the retail user by the aggregation user, if there is a distributed power in the aggregation user, the power thereof is retailed by the aggregation user agent, the price signal of the wholesale market is transmitted to the retail market in a market pricing manner, the retail user is a single listing pricing on the selling side, and the retail contract is formed in the order from low to high according to the pricing of the retail user. The retail contract information and the neural network parameters of the retail user are updated after the transaction.

[0147] (4) Further refine the transaction scale and carry out real-time spot transaction. First, the predicted power generation and the predicted demand power are still decomposed to hours, then the day-ahead contract and the monthly contract are decomposed to the power of the day, the real-time predicted power generation and the real-time predicted power consumption of the power generation and consumption parties are determined, the power is declared, the bid is also obtained through the deep reinforcement learning algorithm, the contract information and the neural network parameters are updated after the transaction is completed, and a total of 8760 real-time transactions are carried out throughout the year, and each step is carried out once.

[0148] (5) Considering the real-time output fluctuation of the unit and the load, the actual output and the actual load at the time point are determined, and the smaller value of the two is taken as the actual executable power of the contract, that is, the actual power purchase of the user. By comparing the contract power and the actual executable power, the contract power is adjusted, the executable power of all contracts at the time point is compared with the demand power at the time point, if the supply and demand are unbalanced, the thermal power generator and the user carry out auxiliary service and demand response behavior, the real-time balancing behavior is preferentially carried out by the user, the real-time supply and demand power balance is maintained, the maximum power of the auxiliary service and the demand response is obtained from the climbing rate of the unit and the load, and meanwhile the output of the equipment needs to be ensured within the allowed output upper and lower limits.

[0149] In addition, the model and the algorithm are applied to the following examples of a certain regional power system.

[0150] 1. Parameter and scheme setting

[0151] The example test is based on NVIDIA RTX 3060Intel(R)Core(TM)i7-12650h CPU. The deep reinforcement learning DQN algorithm, the market bidding model, the subject strategy model and the code compilation of the Markov decision process are mainly completed in an object-oriented programming manner using Python, wherein the neural network of the DQN is generated using pytorch. In addition, a database background of the simulation system is built based on an SQLServer database, and a visual interface of the case is built by connecting the Anylogic software through a pypeline plug-in. In order to prove the superiority and effectiveness of the DQN algorithm, the same example is solved by using a Q-learning algorithm, and finally the training effects are compared.

[0152] (1) Simulation experiment setting: The purposes include two aspects, firstly, to conduct algorithm comparison to prove the effectiveness and superiority of DQN deep reinforcement learning algorithm for user bidding strategy selection, and secondly, to conduct macro-micro combined analysis, to analyze the changes of user's decision-making process under DQN algorithm training and the influence on market operation results. Among them, in order to conduct algorithm comparison, experiment one uses DQN deep reinforcement learning algorithm for simulation, as the test group, and experiment two uses Q-learning algorithm for simulation, as the control group. The main parameter settings of the model algorithm are shown in Table 1.

[0153] Table 1

[0154]

[0155] Among them, the action set, i.e. the bidding coefficient set of power suppliers and users, is defined as follows: the bidding lower limit of ordinary users is set to 430 yuan / MW·h, and the minimum bidding coefficient is 0.86, the bidding upper limit is 627.2 yuan / MW·h, and the maximum bidding coefficient is 1.2544, 28 points are evenly taken in the upper and lower limit interval of the bidding coefficient, forming a discrete action interval with 30 bidding coefficients, corresponding to 30 bidding values. The bidding lower limit of the aggregated user in the power wholesale market is set to 410 yuan / MW·h, and the upper limit is 607.2 yuan / MW·h, the bidding coefficient upper and lower limit is [0.854, 1.265], and the pricing coefficient upper and lower limit of the power retail market is [1.01, 1.3]; the bidding lower limit of the power supplier is set to (C Gi -60) yuan / MW·h, and the upper limit is (C Gi +144) yuan / MW·h, and its action space is also defined as a set containing 30 discrete bidding coefficients. For the Q-learning algorithm, because it cannot analyze high-dimensional states, the state set is simplified to the bidding of the subject in the last training step.

[0156] (2) Experimental parameter setting: The installed capacity ratio of renewable energy in the power system is 65%, the renewable energy generation capacity ratio is 40%, the total power generation capacity is 58700000 MWh, and the total supply and demand balance in the system. The simulation experiment simulates the operation of the system within one year, and the single-step operation time is 1 hour, and the total time is 8760 hours. The experimental parameter inputs of the two groups of experiments are the same, and the main experimental data are shown in Table 2.

[0157] Table 2

[0158]

[0159] The power generators in the experiment include 5 thermal power generation subjects, 2 wind power generation subjects, 2 photovoltaic power generation subjects and 3 distributed photovoltaic power generation subjects; the demand subjects include 15 small and medium-sized power consumption subjects and 7 large power consumption subjects; the aggregated user subjects include 2 power selling companies and one distributed park. Among them, the thermal power generating units include 330MW, 600MW, 660MW and 1000MW coal-fired units, the related marginal cost technical parameters are obtained by fitting the actual data, and the marginal cost of the unit is finally calculated. The annual utilization hours, the reportable electricity quantity and the annual electricity quantity decomposition curve of each unit are obtained according to the empirical data, and the specific information of the power generators and users is shown in Tables 3 and 4, respectively.

[0160] Table 3

[0161]

[0162] Table 4

[0163]

[0164]

[0165] 2. Analysis of user market collaborative bidding strategy

[0166] (1) Algorithm effectiveness analysis: The market collaborative behavior of users is a series of behaviors that users actively collaborate with other subjects to participate in the power market under market signals, including self-reporting of electricity quantity and price, participating in wholesale and retail transactions, demand response behavior, etc. In the model constructed in this embodiment, the adaptive bidding behavior of users and power generators in the market bidding process, as well as the deep reinforcement learning process under environmental changes and market information feedback, are mainly studied. The transaction of one year is simulated, with an hour as the simulation step. After the transaction of each hour is completed, training iteration is performed once, a total of 8760 iterations. For each agent, it may not participate in all time-scale market transactions at the same time, so the number of iterations and training it receives is about 8000 times.

[0167] ① Convergence analysis: For the two types of agents, users and power generators, their double Q networks and memory banks are set, as shown in Figure 7 The training error value of the user DQN neural network under 8760 steps is shown in FIG. 6. It can be seen that, with the iteration, the training error of the double-layer Q neural network gradually decreases, and finally tends to be stable, and the network parameters are basically stable with the iteration.

[0168] Through the bidding of users, the convergence of the algorithm can be more intuitively analyzed. In order to prove the superiority of the DQN algorithm, the bidding convergence of users under the Q-learning algorithm and the DQN algorithm is compared. In the power wholesale market transaction, the bidding convergence of some users is shown inFigure 8 The results are shown (Q-learning on the left and DQN on the right).

[0169] From the user's bid distribution, it can be seen that the Q-learning algorithm is unstable in convergence and is prone to local optimal solution. With the increase of state quantity, the algorithm cannot fully traverse the state and action, so it is difficult to achieve good results. Under the DQN algorithm, the user also experiences a random exploration process, but because of the form of neural network, it is not necessary to traverse all the state and action space, and finally the experience of the memory bank can be used to train the bid strategy to convergence. From the overall trend of user bidding, it can be seen that with the iteration, the user's bid gradually converges under the training of the DQN neural network: in the case of insufficient market transaction times, the user's bid selection is relatively random, and the image shows that it is relatively scattered. With the progress of the transaction, the user neural network is basically trained, and the user's bid gradually converges to the strategy that can obtain the maximum benefit. Overall, the convergence speed and effect of the DQN deep reinforcement learning algorithm are obviously better than those of the Q-learning algorithm, and the bid at the final convergence is less than that of the latter, which is conducive to the user to obtain lower wholesale market electricity purchase price to obtain more benefits.

[0170] 2. Algorithm effectiveness comparison: In order to exclude the influence of randomness on the experimental results, 50 experiments were conducted for each of experiment one and experiment two, and single-factor variance analysis was conducted on the relevant power market operation indexes obtained in each experiment to verify the significant influence of algorithm application on the experimental results. The settings are shown in Table 5.

[0171] Table 5

[0172]

[0173] The sample data obtained by repeating the simulation of the two experiments (each for 50 times) was subjected to single-factor variance analysis, and the hypothesis was as follows: H0: μ1= μ2 (algorithm selection has no significant effect on the average clearing price of the market); H1: μ1≠ μ2 (algorithm selection has a significant effect on the average clearing price of the market). The SPSS software was used to perform single-factor variance analysis on the algorithm selection factor, and the market average clearing price was used as the comparison index. The results of normality test and variance homogeneity test are shown in Tables 6 and 7, respectively.

[0174] Table 6

[0175]

[0176] Table 6 shows the results of descriptive statistics and normality test on the market average clearing price, including median, mean, etc. for testing the normality of the data. Due to the small sample size, the results of Shapiro-Wilk test show that the significance P value is 0.294, which is not significant at the level, and the null hypothesis (the data is normally distributed) cannot be rejected, so the data meets the normal distribution.

[0177] Table 7

[0178]

[0179] The results of the homogeneity of variance test of Table 7 show that for the clearing price, the significance P value is 0.535, which is not significant at the level, and the null hypothesis (the null hypothesis: the homogeneity of variance is met) cannot be rejected, so the data meets the homogeneity of variance. The final ANOVA results are shown in Table 8 below.

[0180] Table 8

[0181]

[0182] The mean values of scheme one and scheme two on the clearing price are 429.269 and 437.392, respectively. Since the homogeneity of variance is met, a single sample variance test is used (significance level 5%), and the ANOVA results P≤0.05, so the statistical results are significant, and the null hypothesis H0: μ1= μ2 (the algorithm selection has no significant impact on the market average clearing price) is rejected, indicating that different schemes result in significant differences in the market average clearing price, i.e. different user market behavior decision-making methods have a significant impact on the market clearing result. Among them, μ1= 429.26, which is significantly lower than μ2= 437.39, indicating that under the DQN algorithm, the training effect of the user market bidding strategy is better, the market clearing efficiency is high, the average price is low, and it is beneficial to improve the user's revenue. At the same time, the market average clearing price standard deviation of scheme one is lower than that of scheme two, indicating that the algorithm has lower randomness and the experimental results are more stable.

[0183] (2) User bidding strategy and market clearing price analysis: In the case of double-sided bidding, users participating in the electricity market need to balance the transaction volume and clearing price in response to price signals. When bidding is high, it can stabilize the required power, but it may increase the market clearing price and thus the electricity cost. When bidding is low, it may reduce the market clearing price, but it cannot fully meet the demand for power for production and life. For power suppliers, user participation in bidding will inevitably affect their overall bidding strategy, and thus affect the market clearing result under the joint action of supply and demand. The above has analyzed the annual market transaction bidding of some users and power suppliers. Next, the overall trend and market clearing price under double-sided bidding will be analyzed, and finally the revenue results of the main bidding strategy will be analyzed.

[0184] ①Power wholesale market bidding strategy analysis: In the power wholesale market, the average value of each level of market bidding by each power producer and user and the average value of each level of market clearing price are shown in the simulation experiment results as shown in Table 1. Figure 9 First, the user side bidding is analyzed. It can be seen that although the initial benchmark value of the bidding is relatively close, the bidding behavior of the users has differentiated as the training progresses. The average bidding of some users has decreased and stabilized, while the average bidding of some users has gradually increased. This is mainly due to the different degree of benefit of each user. U4, U6 and U7 are mainly due to their high degree of benefit, so they have more bidding space and more room for concessions. The total benefit is mainly related to the winning power, and in order to maximize the benefit, the user will tend to give up part of the benefit to the power side and bid a higher price to ensure power purchase to reduce demand response power. O2, U1 and U3 are mainly due to their low degree of benefit, and they will have to choose between winning power and power cost, gradually reducing the bidding to approach the market clearing price to reduce power cost. Similarly, the bidding of power producers is mainly affected by marginal cost. The higher the marginal cost, the more likely it is to bid a higher price, and the more likely it is to become a marginal unit. The units with cost advantage choose a lower bid price to stabilize the transaction volume. From the bidding of G1 and G4, which have higher prices, it can be seen that their bidding gradually decreases and stabilizes, which is mainly due to the continuous reduction of the minimum bid price of the user side. The power producers reduce their bidding to ensure winning volume. Wind power sellers G6 and G7 and photovoltaic power sellers G8 and G9 all bid at a lower price, not only because their benchmark bid price is low, but also because of the volatility and instability of their output. In order to get as much stable power as possible, renewable energy power producers have adopted lower bid prices.

[0185] From the overall trend of the bidding curve and the clearing price curve, it can be observed that in the process of participating in the market, the power producers and users in the market train and iterate their bidding strategies according to market information and transaction results. Under the influence of bidding on both supply and demand sides, the market clearing price gradually stabilizes and has a slight downward trend, and the up and down fluctuations are mainly related to the monthly power supply and demand ratio. Finally, the strategic bidding behavior of each user and power producer for maximizing their own benefit gradually converges, and both parties have found the bidding strategy with the highest expected return. The final convergence of all users' bidding is shown in Table 9.

[0186] Table 9

[0187]

[0188] It can be seen that the bid price is closely related to the degree of electricity benefit, which is consistent with the analysis result of the average bid price. The user with greater degree of electricity benefit has greater bidding space, and the final converged bid price is higher. Therefore, in general, the user with high degree of electricity benefit tends to take high bid price, which has less impact on the unified clearing price of the market. The user with low degree of electricity benefit is likely to form the last pair of transactions, and the same is true for the generation side. The unit with greater marginal cost is more likely to become the marginal unit, which affects the market clearing price.

[0189] ②Analysis of power wholesale market clearing price: the market transaction price directly affects the power purchase cost of the subject, as shown in Figure 10 The market power supply and demand ratio and the wholesale market transaction price of each month are shown in the figure. As can be seen from the figure, under the influence of the bidding behavior of both supply and demand sides, the market price does not continuously increase or decrease. The overall trend is basically determined by the supply and demand ratio. In months with high supply and demand ratio, such as November and December, the market supply is sufficient, and the overall price is low. In months with low supply and demand ratio, such as August and September, the overall market price is high, reflecting the general commodity characteristics of electricity price affected by supply and demand relationship in the free bidding market of both supply and demand sides. In addition, from the training effect of the model algorithm, with the increase of the number of iterations, the market price gradually tends to be stable. In months with similar supply and demand ratio, such as April and June, March and July, and May and November, it can be seen that in the case of similar market supply and demand ratio, the later the month, the smaller the market price fluctuation, and the lower the average market clearing price. This reflects that with the training, the market clearing price has a downward trend and gradually tends to be stable, and the power cost of the user is reduced. In the experiment, the total average market clearing price λ C = 419.17 yuan / MW·h, and the detailed market average clearing price is shown in the following table 10.

[0190] Table 10

[0191]

[0192] ③Analysis of power retail market bidding strategy and market price: in the retail market, the posted transaction mode is adopted, and the power supply side of the aggregation user determines the price on one side, and the retail user independently selects the order to purchase. The market clearing price is shown in Figure 11 . It can be seen that since the market pricing is determined by the power supply side on one side, with the convergence of the bidding strategy of the aggregation user, the retail pricing of the market also converges obviously. In the model, the monthly market clearing price is used as the bidding reference value of the retail market in the month. It can be seen that the trend of the final market clearing result is consistent with the overall trend of the wholesale market, which reflects the transmission of the power wholesale to retail market price. The market pricing of each month and the bidding coefficient of the aggregation user are shown in the following table 11.

[0193] Table 11

[0194]

[0195]

[0196] It can be seen that in the unilateral pricing transaction, the retail user basically does not implement market behavior, and the pricing coefficient of the aggregation user as the power supply side is continuously improved in order to pursue maximum benefit, and the finally converged pricing coefficient is 1.4, reaching the upper limit, which reflects that the pricing strategy is optimized under the reinforcement learning algorithm. Compared with the experimental results of the bilateral centralized bidding of the wholesale market, it can be seen that the demand side market behavior can well limit the supply side bidding, stabilize the market clearing price, and play an important role in reflecting the supply and demand relationship of electric energy.

[0197] 3. Analysis of green synergy of user market behavior

[0198] The user green behavior is the purchase of new energy power generation by the user, and the user cooperative behavior is a series of behaviors of the user in the market signal, including self-reporting of power and price, participating in power wholesale and retail transactions, demand response behavior, and the like. Before analyzing the power purchase amount of the user, the total power generation amount of the system is analyzed first, and the predicted power generation amount curves of various resources are fitted from experience data, and the final actual output result is obtained from the market transaction, as shown in Figure 12 It can be seen that the output fluctuation of the thermal power unit is the smallest, and the operation is stable and reliable, which is the most stable power supply unit in the new power system. The output fluctuation of photovoltaic power generation is affected by the day and night sunlight conditions, and the output curve of wind power basically has no regularity, and both of them can only be generated under certain conditions, and the output is 0 under no wind or no light conditions. Therefore, with the large-scale integration of renewable energy into the power grid, the safe and stable operation of the new power system is greatly challenged, and the user needs to actively respond to the price signal to purchase new energy power generation, and actively participate in auxiliary services and demand response.

[0199] On this basis, the new energy power generation consumption amount of the user in the energy wholesale market is analyzed. In this embodiment, the minimum marginal cost is set for the renewable energy according to the actual situation, and the carbon price is considered in the calculation of the power generation business income, so that the new energy power generation business has a lower bidding price in the market transaction process throughout the year under the training of the reinforcement learning algorithm, and the finally converged average bidding price is 283.9 yuan / MW·h, which is obviously lower than 343.9 yuan / MW·h of the thermal power generation business. Under the condition that the demand amount of the user is fixed, the market clearing price is generally low in the transaction period with high new energy power generation amount, and the user adjusts the bidding strategy under the guidance of the price signal to improve the overall power purchase amount. The process in which the new energy power generation amount affects the market clearing price is shown in Figure 13 .

[0200] Since the demand of the user fluctuates every month, the proportion of the new energy generation capacity purchased by the user in the total purchased capacity and the proportion of the new energy expected generation capacity in the total expected generation capacity are analyzed, such as Figure 14 It can be seen that the new energy generation capacity purchased by the user gradually approaches the actual generation capacity trend of the renewable energy with the increase of the training times, and the deviation gradually decreases in the last stage of training. Under the training of the DQN algorithm, the new energy generation consumption capacity of the user basically coincides with the new energy output trend, indicating that under the training of the deep reinforcement learning algorithm, the instability of the cleared capacity caused by the randomness of the supply and demand bid can be avoided as much as possible, and the consumption effect of the new energy generation is improved. However, under the Q-learning algorithm, the non-convergence of the supply and demand bid strategy leads to a large fluctuation of the transaction capacity.

[0201] In addition, the total new energy generation purchase amount of each month is also shown, such as Figure 14 It can be seen that under the DQN algorithm, the new energy generation purchase amount of the user not only coincides with the expected output of the unit in the trend, but also has a higher total new energy generation purchase amount and a higher ratio of the total amount to the actual generation capacity of the new energy. The utilization rate θ of the new energy is calculated R , wherein is the actual on-grid capacity of the new energy generator Gri' after the energy transaction and auxiliary service and demand response process in the transaction period t; is the real-time maximum generation capacity of the new energy generator Gri' in the transaction period t, N Gr is the number of new energy generators. Under the DQN condition, θ R = 100%, while under the Q-learning condition, this value is reduced to 98.79%, indicating that the deviation capacity of the new energy exceeds the limit of the auxiliary service and demand response provided by the thermal power and the user side, and the wind and light curtailment phenomenon occurs. The experiment shows that the adaptive learning behavior of the user has a profound influence on the electricity market and the new type of power system, embodies the green collaborative market behavior of the user actively responding to the price signal to purchase new energy generation, and fully proves the important role of the user side in participating in the electricity market.

[0202] In addition, the embodiments of the present application also include: A8, the method of any one of A3-A7, wherein the electricity wholesale market transaction strategy model of the user is aggregated, including:

[0203] max M Ok = (I Ok - λ)X Ok + δQ' Ok

[0204] wherein, MOk represents the revenue of the aggregated user Ok in any trading period of the electricity wholesale market, I Ok represents the revenue of the aggregated user Ok in any trading period of the electricity wholesale market, I Ok represents the revenue of the aggregated user Ok in any trading period of the electricity wholesale market, I Ok represents the revenue of the aggregated user Ok in any trading period of the electricity wholesale market, I

[0205] A9. The method of any one of A3-A8, wherein the electricity retail market trading strategy model of the aggregated user comprises:

[0206] max M Ok ' = (b Ok - λ ') X ' Ok

[0207] wherein M Ok ' represents the revenue of the aggregated user Ok in any trading period of the electricity retail market, X ' Ok represents the revenue of the aggregated user Ok in any trading period of the electricity retail market, X ' Ok represents the revenue of the aggregated user Ok in any trading period of the electricity retail market, X ' jt represents the revenue of the aggregated user Ok in any trading period of the electricity retail market, X '

[0208] A11. The method of any one of A3-A10, wherein the Markov decision process of the large power user in the electricity wholesale market comprises:

[0209]

[0210]

[0211] wherein s jt represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, t represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, jt represents the state of the large power user Uj in the tth trading period of the electricity wholesale market, jtdenotes the bidding coefficient of large power consumer Uj in the tth trading period of the electricity wholesale market, r jt denotes the reward of large power consumer Uj in the tth trading period of the electricity wholesale market, I Uj denotes the degree of electricity benefit of large power consumer Uj, and δ denotes the compensation coefficient when participating in demand response to obtain compensation.

[0212] A12. The method of any one of A3-A11, wherein the Markov decision process of the aggregated consumer in the electricity wholesale market comprises:

[0213]

[0214]

[0215]

[0216] wherein s kt denotes the state of aggregated consumer Ok in the tth trading period of the electricity wholesale market, denotes the declared price of aggregated consumer Ok in the tth trading period of the electricity wholesale market, denotes the minimum value of the declared price of all aggregated consumers in the tth trading period of the electricity wholesale market, λ t denotes the clearing price of the tth trading period of the electricity wholesale market, denotes the winning electricity quantity, the actual demand electricity quantity, and the demand response electricity quantity of aggregated consumer Ok in the tth trading period of the electricity wholesale market, respectively, a kt denotes the action of aggregated consumer Ok in the tth trading period of the electricity wholesale market, e kt denotes the bidding coefficient of aggregated consumer Ok in the tth trading period of the electricity wholesale market, r kt denotes the reward of aggregated consumer Ok in the tth trading period of the electricity wholesale market, I Ok denotes the degree of electricity benefit of aggregated consumer Ok, and δ denotes the compensation coefficient when participating in demand response to obtain compensation.

[0217] A13. The method of any one of A3-A12, wherein the Markov decision process of the aggregated consumer in the electricity retail market comprises:

[0218]

[0219] a kt ' = [h kt ]

[0220]

[0221] wherein s kt' denotes the state of the aggregated user Ok at the tth trading period of the electricity retail market, ' denotes the bidding price of the aggregated user Ok at the tth trading period of the electricity retail market, ' denotes the maximum value among the bidding prices of all aggregated users at the tth trading period of the electricity retail market, λ t ' denotes the clearing price of the trading period associated with the tth trading period of the electricity retail market in the electricity wholesale market, ' denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, k ' denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, ' denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, k ' denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, kt ' denotes the action of the aggregated user Ok at the tth trading period of the electricity retail market, h kt ' denotes the bidding coefficient of the aggregated user Ok at the tth trading period of the electricity retail market, r kt ' denotes the reward of the aggregated user Ok at the tth trading period of the electricity retail market.

[0222] A16. A readable storage medium storing program instructions, which, when read and executed by a computing device, cause the computing device to perform the method of any one of A1-A14.

[0223] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order to not obscure the understanding of this description. It will be appreciated, that the scope of the

[0224] While the application has been described in accordance with the various embodiments shown and described, it is to be understood that the application is not limited to those precise embodiments, and that various modifications and changes can be made by those skilled in the art without departing from the scope of the present application. Furthermore, the language used in this specification has been principally selected for readability and instructional purposes and can not have been selected to delineate or circumscribe the patent rights to which it refers. Accordingly, the present application is intended to be illustrative, but not limiting, of the scope of the application, which is set forth with particularity in the claims that follow.

Claims

1. A method for determining the electricity market behavior strategies of multiple users in a novel power system, wherein the electricity market includes a wholesale electricity market and a retail electricity market, the main participants in the retail electricity market include retail users on the demand side and aggregated users on the supply side, the main participants in the wholesale electricity market include wholesale users on the demand side and power generators on the supply side, and the wholesale users include aggregated users on the supply side of the retail electricity market, the method comprising: A power wholesale market clearing model and a power retail market clearing model are established with the goal of maximizing social welfare. For each entity in the power wholesale market, a power wholesale market trading strategy model is constructed with the goal of maximizing its own profits. For each aggregated user, a power retail market trading strategy model is also constructed with the goal of maximizing its own profits. For each entity in the wholesale electricity market, a two-layer decision optimization model is constructed, using its own wholesale electricity market trading strategy model as the upper-layer model and the aforementioned wholesale electricity market clearing model as the lower-layer model. Similarly, for each aggregated user, a two-layer decision optimization model is constructed, using its own retail electricity market trading strategy model as the upper-layer model and the aforementioned retail electricity market clearing model as the lower-layer model. For any entity, the Markov decision-making process of each wholesale user and power generator in the electricity wholesale market and the Markov decision-making process of each aggregate user in the electricity retail market are established by taking the bidding coefficient of each trading period as its action in each trading period, taking the revenue of the unit declared electricity volume in each trading period as its reward in each trading period, and determining its status in each trading period based on the declared electricity price, declared electricity volume, winning bid electricity volume, and clearing price of the electricity wholesale market in each trading period. Based on the established Markov decision process, a deep reinforcement learning algorithm is used to solve each two-level decision optimization model to obtain the bidding strategies of each wholesale user and power generator in the electricity wholesale market and the bidding strategies of each aggregate user in the electricity retail market. The bidding strategies include the bid price and bid volume for each trading period.

2. The method of claim 1, wherein, The trading period includes monthly, day-ahead, and hourly periods in the wholesale electricity market and day-ahead periods in the retail electricity market. Accordingly, the reporting strategy includes the reported electricity price and reported electricity volume of each wholesale user and power generator in the wholesale electricity market for each month, day-ahead, and hour, as well as the reported electricity price and reported electricity volume of each aggregate user in the retail electricity market for each day.

3. The method of claim 1 or 2, wherein, The wholesale users also include large power users, the retail users include small and medium-sized power users, the aggregated users include power sales companies, microgrids, and distributed parks, the large power users are power users whose supply voltage is greater than or equal to a first predetermined value, and the small and medium-sized power users are the remaining power users other than the aggregated users and large power users.

4. The method of claim 3, wherein, The greatest social welfare in the wholesale electricity market is achieved when the price difference between the declared electricity prices of wholesale users and power generators is maximized, and the aforementioned wholesale electricity market clearing model includes: where y denotes the difference between the declared price of the wholesale user and the power producer, B Uj , B Ok , B Gi denote the declared price of the large power user Uj, the aggregated user Ok, and the power producer Gi at any trading period of the power wholesale market, N U , N O , N G denote the number of large power users, aggregated users, and power producers in the power wholesale market, respectively.

5. The method of claim 3, wherein, The social welfare in the electricity retail market is maximized when the electricity cost for retail users is minimized, and the electricity retail market clearing model includes: where y' represents the electricity cost of the retail user, b Ok represents the declared price of the aggregated user Ok at any trading period in the electricity retail market, N O represents the number of aggregated users in the electricity wholesale market.

6. The method of claim 3, wherein, The power wholesale market transaction strategy model of the power generator, comprising: where M Gi represents the profit of the power producer Gi in any trading period of the power wholesale market, λ represents the clearing price of any trading period of the power wholesale market, C Gi represents the benchmark price of the power producer Gi, X Gi represents the winning power quantity of the power producer Gi in any trading period of the power wholesale market, represents the carbon price, represents the carbon quota coefficient per unit of power generation, represents the weighted average carbon emission coefficient of all units of the power producer Gi.

7. The method of claim 3, wherein, The power wholesale market transaction strategy model of the large power consumer, comprising: max M Uj = (I Uj - λ) X Uj + δQ' Uj wherein M Uj represents the income of large power consumer Uj in any trading period of the power wholesale market, I Uj represents the income of large power consumer Uj in any trading period of the power wholesale market, I Uj represents the income of large power consumer Uj in any trading period of the power wholesale market, I Uj represents the income of large power consumer Uj in any trading period of the power wholesale market, I 8. The method of claim 3, wherein, The power wholesale market transaction strategy model of the aggregation user, comprising: maxM Ok = (I Ok - λ) X Ok + δQ' Ok where M Ok represents the revenue of the aggregated user Ok in any trading period of the electricity wholesale market, I Ok represents the benefit per unit of electricity of the aggregated user Ok, λ represents the clearing price in any trading period of the electricity wholesale market, X Ok represents the winning electricity quantity of the aggregated user Ok in any trading period of the electricity wholesale market, δ represents the compensation coefficient when participating in demand response to obtain compensation, Q' Ok represents the demand response electricity quantity of the aggregated user Ok in any trading period of the electricity wholesale market.

9. The method of claim 3, wherein, The power retail market transaction strategy model of the aggregation user, comprising: maxM Ok '=(b Ok -λ')X' Ok where M Ok represents the revenue of the aggregated user Ok in any trading period of the electricity retail market, X Ok represents the winning electricity quantity of the aggregated user Ok in any trading period of the electricity retail market, b Ok represents the declared price of the aggregated user Ok in any trading period of the electricity retail market, and λ' represents the clearing price of the trading period associated with the trading period in the electricity retail market in the electricity wholesale market.

10. The method of claim 3, wherein, The Markov decision process of the power generator in the power wholesale market, comprising: a it = [f it ] where s it denotes the state of the power generator Gi at the tth trading period of the power wholesale market, denotes the declared price of the power generator Gi at the tth trading period of the power wholesale market, denotes the maximum value among the declared prices of all power generators at the tth trading period of the power wholesale market, λ t denotes the clearing price of the power wholesale market at the tth trading period, denote the winning electric quantity and the declared electric quantity of the power generator Gi at the tth trading period of the power wholesale market, a it denotes the action of the power generator Gi at the tth trading period of the power wholesale market, f jt denotes the bidding coefficient of the power generator Gi at the tth trading period of the power wholesale market, r it denotes the reward of the power generator Gi at the tth trading period of the power wholesale market, denotes the revenue of the power generator Gi at the tth trading period of the power wholesale market.

11. The method of claim 3, wherein, The Markov decision process of the large power consumer in the power wholesale market, comprising: a jt = [d jt ] wherein s jt denotes the state of large power consumer Uj at the tth trading period of the electricity wholesale market, denotes the bidding price of large power consumer Uj at the tth trading period of the electricity wholesale market, denotes the minimum value among the bidding prices of all large power consumers at the tth trading period of the electricity wholesale market, λ t denotes the clearing price of the tth trading period of the electricity wholesale market, denotes the winning electricity quantity, actual demand electricity quantity, demand response electricity quantity of large power consumer Uj at the tth trading period of the electricity wholesale market, respectively, a jt denotes the action of large power consumer Uj at the tth trading period of the electricity wholesale market, d jt denotes the bidding coefficient of large power consumer Uj at the tth trading period of the electricity wholesale market, r jt denotes the reward of large power consumer Uj at the tth trading period of the electricity wholesale market, I Uj denotes the degree of electricity benefit of large power consumer Uj, and δ denotes the compensation coefficient when compensation is obtained by participating in demand response.

12. The method of claim 3, wherein, The Markov decision process of the aggregation user in the power wholesale market, comprising: a kt = [e kt ] where s kt denotes the state of the aggregated user Ok at the tth trading period of the electricity wholesale market, denotes the bidding price of the aggregated user Ok at the tth trading period of the electricity wholesale market, denotes the minimum value of the bidding price of all aggregated users at the tth trading period of the electricity wholesale market, λ t denotes the clearing price of the tth trading period of the electricity wholesale market, denotes the winning electricity quantity, the actual demand electricity quantity, and the demand response electricity quantity of the aggregated user Ok at the tth trading period of the electricity wholesale market, respectively, a kt denotes the action of the aggregated user Ok at the tth trading period of the electricity wholesale market, e kt denotes the bidding coefficient of the aggregated user Ok at the tth trading period of the electricity wholesale market, r kt denotes the reward of the aggregated user Ok at the tth trading period of the electricity wholesale market, I Ok denotes the degree-of-electricity benefit of the aggregated user Ok, and δ denotes the compensation coefficient when participating in demand response to obtain compensation.

13. The method of claim 3, wherein, The Markov decision process of the aggregation user in the power retail market, comprising: a kt '=[h kt ] where s kt denotes the state of the aggregated user Ok at the tth trading period of the electricity retail market, denotes the bidding price of the aggregated user Ok at the tth trading period of the electricity retail market, denotes the maximum value among the bidding prices of all aggregated users at the tth trading period of the electricity retail market, λ t denotes the clearing price of the trading period associated with the tth trading period of the electricity retail market in the electricity wholesale market, denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, k denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, k denotes the winning electricity quantity of the aggregated user Ok at the tth trading period of the electricity retail market, kt denotes the action of the aggregated user Ok at the tth trading period of the electricity retail market, h kt denotes the bidding coefficient of the aggregated user Ok at the tth trading period of the electricity retail market, r kt denotes the reward of the aggregated user Ok at the tth trading period of the electricity retail market.

14. The method of claim 1 or 2, wherein, The declaration strategy of each subject is obtained by the following decision behavior mode constructed under a deep reinforcement learning algorithm: The network parameters of the evaluation network, the exploration factor, the target network update interval, and the current state of the subject are initialized; The greedy strategy is used to determine the declared price of the subject in the current transaction period, and the exploration factor is updated; The action of declaring the determined declared price is performed, the state and the reward of the next transaction period of the subject are obtained, and the current state, the action, the reward, and the state of the next transaction period are saved as a sample to the memory bank; Based on the first prediction value and the second prediction value, the network parameters of the evaluation network are updated, the first prediction value is the prediction value of the evaluation network for the current state and the current action, and the second prediction value is the maximum value in the prediction values of the target network for all actions in the action set of the subject state in the selected sample from the memory bank; If the current step number reaches an integer multiple of the target network update interval, the network parameters of the evaluation network are copied to the target network, and it is judged whether the current step number reaches a preset value; If the preset value is not reached, the current state is updated to the state of the next transaction period, the current step number is incremented by one, and the steps of declared price determination, exploration factor updating, action execution, sample saving, network parameter updating, network parameter copying, and step number reaching the preset value are continued until the current step number reaches the preset value, and the process is ended.

15. A computing device, comprising: at least one processor; and a memory storing program instructions configured to be executed by the at least one processor, the program instructions comprising instructions for performing the method of any one of claims 1-14.

16. A readable storage medium storing program instructions, when the program instructions are read and executed by a computing device, causing the computing device to perform the method of any one of claims 1-14.

Citation Information

Patent Citations

  • Power purchasing and selling joint strategy optimization method based on multi-task deep reinforcement learning

    CN116029415A

  • Renewable energy consumption method based on game model

    CN117726418A