A Agent-Based Electricity Spot Market Strategy Processing Method and Computer Equipment
By employing an agent-based approach to electricity spot market strategy processing, and utilizing DDPG and DQN algorithms to generate and select strategies, the problems of subjectivity and response lag in existing electricity spot market strategies are solved. This approach achieves scientific rigor, accuracy, and efficiency in the strategies, ensuring maximum returns in the electricity spot market.
Patent Information
- Application Number
- CN202510118132.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-24
AI Technical Summary
The formulation of electricity spot market strategies in existing technologies is highly subjective, resulting in poor accuracy of day-ahead strategies and difficulty in responding to dynamic changes in the electricity spot market and personalized needs of users in a timely manner, thus limiting the efficiency and effectiveness of electricity spot trading.
The method of processing electricity spot market strategies based on intelligent agents generates multiple candidate strategies using a first target intelligent agent and determines the acceptance probability of specified strategy factors through a second target intelligent agent. The target strategy that best fits the current state of the electricity spot market is then selected. The neural network of the intelligent agent is constructed by combining DDPG and DQN algorithms to generate and select strategies.
It improves the scientific rigor, accuracy, and efficiency of strategy generation and selection, ensuring the flexibility and adaptability of the strategies, and enabling them to maximize returns in the electricity spot market.
Smart Images

Figure CN120069597B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electricity market technology, and in particular to a method and computer equipment for processing electricity spot market strategies based on intelligent agents. Background Technology
[0002] The electricity spot market typically reflects dynamic changes in market supply and demand through real-time price fluctuations in the day-ahead and intraday markets, thereby incentivizing market participants to adjust their production and consumption behaviors. The day-ahead market, as a crucial component of the spot market, is primarily responsible for optimizing and arranging electricity trading plans one day in advance to ensure the safe, economical, and efficient operation of the power system.
[0003] In the day-ahead market, power plants typically formulate day-ahead strategies and participate in trading based on their own power generation capacity forecasts, market price fluctuations, and load demand. In most related technologies, day-ahead strategies are determined by establishing rules based on experience, and then the planned power generation or consumption is determined based on these strategies. However, this approach is highly subjective, leading to poor accuracy of day-ahead strategies. Furthermore, adjustments can usually only be made through backtesting after the strategy is formulated, making it difficult for day-ahead strategies to respond promptly to dynamic changes in the electricity spot market and personalized user needs, thus limiting the efficiency and effectiveness of electricity spot trading. Summary of the Invention
[0004] This application provides a power spot market strategy processing method and computer device based on intelligent agents. It solves the technical problem that current strategy formulation methods are highly subjective, resulting in poor accuracy of day-ahead strategies and difficulty in timely responding to dynamic changes in the power spot market and personalized needs of users. The embodiments of this application generate multiple candidate strategies corresponding to different specified strategy factors through a first target intelligent agent, and generate the acceptance probability of each specified strategy factor in the power spot market under the current environment through a second target intelligent agent. Then, the target strategy that best fits the current power spot market state is selected from the candidate strategies, which greatly improves the scientificity, accuracy and efficiency of strategy generation and selection, and helps the target individual to maximize its benefits in the power spot market.
[0005] To achieve the above objectives, the main technical solutions adopted in this application include:
[0006] In a first aspect, embodiments of this application provide a method for processing electricity spot market strategies for intelligent agents, the method comprising:
[0007] The system acquires the current individual status data of the target individual and the current environmental status data of the electricity spot market. The current individual status data is used to characterize the power generation forecast of the target individual on the current operating day, and the current environmental status data is used to characterize the total load and total unit output forecast in the electricity spot market on the current operating day.
[0008] The current individual state data and the current environment state data are input into the first target agent to generate multiple candidate strategies; wherein, each candidate strategy corresponds to a specified strategy factor.
[0009] The current environmental state data is input into the second target agent to determine the acceptance probability corresponding to the specified strategy factor;
[0010] The multiple candidate strategies are screened based on the acceptance probability to obtain the target strategy for the target individual to conduct electricity spot trading on the current operating day.
[0011] The electricity spot market strategy processing method proposed in this application involves a first target agent perceiving and making decisions based on current individual state data and current environmental state data, thereby generating candidate strategies corresponding to different specified strategy factors. Simultaneously, a second target agent determines the acceptance probability of each specified strategy factor based on environmental state data. Finally, based on the acceptance probability, the most suitable candidate strategy corresponding to the strategy factor is selected to obtain the target strategy. Therefore, this application embodiment uses a reinforcement learning-based agent to perceive the real-time state data of the target individual and the electricity spot market, providing comprehensive and real-time input information for subsequent strategy generation and selection. Furthermore, the first agent can generate candidate strategies corresponding to different specified strategy factors, allowing for different focuses among multiple candidate strategies, improving the flexibility of strategy generation. Combined with the acceptance probability output by the second agent, it responds to changes in the electricity spot market, thereby achieving adaptive selection of candidate strategies. Compared with traditional strategy selection methods based on fixed rules or experience, this method significantly improves the scientific rigor, accuracy, and efficiency of strategy generation and selection, ensuring that the target strategy best suited to the current electricity spot market state is selected, which is beneficial for the target individual to achieve expected returns in the electricity spot market.
[0012] Optionally, in some embodiments of this application, the designated strategy factor is at least one of the following: load factor of the electricity spot market, wind power output factor, photovoltaic power output factor, external power transmission factor, and power generation factor of the target individual.
[0013] Optionally, in some embodiments of this application, the first target agent includes multiple simple policy generation layers and a baseline policy generation layer, wherein the simple policy generation layers correspond one-to-one with the specified policy factors.
[0014] This application embodiment sets up multiple simple policy generation layers in the first intelligent agent to independently generate corresponding candidate policies for multiple specified factors, so that the policy generation stage has multiple different focuses, providing a wider range of options for subsequent policy screening and improving the flexibility and comprehensiveness of candidate policies.
[0015] Optionally, in some embodiments of this application, the plurality of simple strategy generation layers include a first simple strategy generation layer, a second simple strategy generation layer, a third simple strategy generation layer, and a fourth simple strategy generation layer;
[0016] The step involves inputting the current individual state data and the current environment state data into the first target agent to generate multiple candidate strategies, including:
[0017] The predicted total power generation of all units in the current individual state data and the current predicted total load in the current environmental state data are input into the first simple strategy generation layer to generate a first candidate strategy corresponding to the load factor.
[0018] The predicted wind power generation of the wind turbine in the current individual state data and the predicted total wind power output of the wind turbine in the current environmental state data are input into the second simple strategy generation layer to generate a second candidate strategy corresponding to the wind power output factor.
[0019] The predicted photovoltaic power generation of the photovoltaic unit in the current individual state data and the predicted total photovoltaic output of the photovoltaic unit in the current environmental state data are input into the third simple strategy generation layer to generate a third candidate strategy corresponding to the photovoltaic output factor.
[0020] The predicted individual power transmission volume of all the units in the current individual state data and the predicted total power transmission volume of all the units in the current environmental state data are input into the fourth simple strategy generation layer to generate a fourth candidate strategy corresponding to the power transmission volume factor.
[0021] The predicted total power generation of all units in the current individual status data and the installed capacity of the target individual are input into the benchmark strategy generation layer, and the ratio between the predicted total power generation and the installed capacity is determined, so as to generate a benchmark candidate strategy corresponding to the power generation factor based on the ratio.
[0022] The plurality of candidate strategies include the first candidate strategy, the second candidate strategy, the third candidate strategy, the fourth candidate strategy, and the benchmark candidate strategy.
[0023] This application embodiment generates candidate strategies for load conditions, wind power processing conditions, photovoltaic processing conditions, and external power transmission conditions by setting a first simple strategy generation layer, a second simple strategy generation layer, a third simple strategy generation layer, and a fourth simple strategy generation layer, respectively. At the same time, a benchmark strategy generation layer is set to generate corresponding benchmark candidate strategies. Thus, the benchmark candidate strategies reflect the passive trading behavior of the target individual that only considers its own power generation. Therefore, this application embodiment can generate strategies separately for different strategy factors through different strategy generation layers, thereby improving the flexibility and comprehensiveness of strategy generation.
[0024] Optionally, in some embodiments of this application, the simple policy generation layer is a neural network constructed based on the DDPG algorithm, and each simple policy generation layer includes an actor neural network using a sigmoid activation function and a critic neural network using a tanh activation function. The actor neural network is used to generate multiple simple policies corresponding to the specified policy factor, and the critic neural network is used to determine the candidate policy corresponding to the specified policy factor based on the returns of the multiple simple policies.
[0025] This application embodiment constructs a simple policy generation layer in the first target agent based on the DDPG (Deep Deterministic Policy Gradient) algorithm, thereby generating corresponding simple policies through an actor neural network, and using the returns of the simple policies as the evaluation criteria through a critic neural network, thereby generating candidate policies with the best returns under specified policy factors.
[0026] Optionally, in some embodiments of this application, the second target agent includes an actor neural network, which employs the DQN algorithm and the sigmoid activation function; the step of inputting the current environmental state data into the second target agent to determine the acceptance probability corresponding to the specified strategy factor includes:
[0027] The current environmental state data is input into the actor neural network of the second target agent to score the applicability of each specified strategy factor under the current environmental state data based on the DQN algorithm.
[0028] The applicability score is normalized using the sigmoid activation function to obtain the acceptance probability for each specified strategy factor.
[0029] This application embodiment constructs a second target agent based on the DQN (Deep Q-Network) algorithm, uses the DQN algorithm to score the applicability of each specified policy factor, and then uses the sigmoid activation function to normalize the score, thereby determining the acceptance probability of each policy factor. This achieves accurate calculation of the acceptance probability, which is beneficial for the subsequent precise screening of candidate policies.
[0030] Optionally, in some embodiments of this application, the step of filtering the plurality of candidate strategies based on the acceptance probability to obtain the target strategy for the target individual to conduct electricity spot trading on the current operating day includes:
[0031] The target strategy factor corresponding to the maximum value of the acceptance probability is determined, and the candidate strategy corresponding to the target strategy factor is selected from the plurality of candidate strategies as the target strategy.
[0032] This application embodiment uses the acceptance probability to reflect the feasibility and expected return of each candidate strategy under the current environmental state data. The target strategy factor is determined based on the maximum value of the acceptance probability, avoiding the limitations of manual intervention or fixed rules. This ensures that the target strategy corresponding to the target strategy factor is the most likely strategy to succeed under the current electricity spot market environment, which is conducive to maximizing the returns of the target individual in electricity spot trading.
[0033] Optionally, in some embodiments of this application, the training process of the first target agent includes:
[0034] Obtain the historical strategy return and historical benchmark return for historical operating days, wherein the historical strategy return is the actual return of the target individual on the historical operating day, and the historical benchmark return is the return obtained by the target individual from electricity spot trading based on the historical unit's predicted power generation on the historical operating day;
[0035] The initial reward is determined based on the difference between the historical strategy reward and the historical benchmark reward, and the first initial agent is trained based on the initial reward to obtain the first target agent.
[0036] This application embodiment determines the original reward by the difference between the historical strategy return and the historical benchmark return, thereby eliminating the data noise caused by the historical benchmark return corresponding to passive trading behavior without taking any strategy, reducing the possibility of overfitting in the first initial agent during training, and thus effectively improving the training effect of the first target agent.
[0037] Optionally, in some embodiments of this application, the training process of the second target agent includes:
[0038] The continuous loss indicator for the target individual's earnings is determined based on the original reward.
[0039] When the continuous loss indicator reaches a preset limit, the original reward is proportionally adjusted according to the target individual's preference adjustment hyperparameter, and the adjusted reward value is obtained.
[0040] The second initial agent is trained according to the adjusted reward value to obtain the second target agent;
[0041] The continuous loss indicator includes at least one of the following: number of consecutive loss days, maximum historical number of consecutive loss days, amount of consecutive loss, and maximum historical amount of consecutive loss.
[0042] In this embodiment, the original reward is proportionally adjusted by adjusting the hyperparameters based on the target individual's preferences in order to train the second initial agent. Therefore, the user's preference for strategy selection is introduced during the training process of the second initial agent, so that the trained second target agent can respond to the user's preference and output a more accurate acceptance probability. This integrates the user's strategy preference into the strategy generation process, thereby ensuring that the selected target strategy maximizes the benefits of electricity spot trading while satisfying the user's preference.
[0043] Secondly, embodiments of this application provide a computer device, including:
[0044] The system includes a memory and a processor, which are interconnected. The memory stores computer instructions, and the processor executes these computer instructions to perform the electricity spot market strategy processing method described in the above embodiments.
[0045] The computer device proposed in this application can perceive the real-time state data of the target individual and the electricity spot market through the reinforcement learning-based intelligent agent of the above-mentioned electricity spot market strategy processing method, providing comprehensive and real-time input information for subsequent strategy generation and screening. Simultaneously, the first intelligent agent can generate candidate strategies corresponding to different specified strategy factors, allowing multiple candidate strategies to have different focuses, improving the flexibility of strategy generation. Furthermore, by combining the acceptance probability output by the second intelligent agent, it responds to changes in the electricity spot market, thereby achieving adaptive screening of candidate strategies. Compared with traditional strategy selection methods based on fixed rules or experience, this significantly improves the scientific rigor, accuracy, and efficiency of strategy generation and screening, ensuring that the target strategy best suited to the current electricity spot market state is selected, which is beneficial for the target individual to maximize its returns in the electricity spot market. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0047] Figure 1 This is a flowchart illustrating a method for processing electricity spot market strategies by an intelligent agent, as proposed in an embodiment of this application.
[0048] Figure 2 This is a schematic diagram of the neural network architecture of the first target intelligent agent proposed in the embodiments of this application;
[0049] Figure 3 This is a schematic diagram of the neural network architecture of the second target intelligent agent proposed in the embodiments of this application;
[0050] Figure 4 This is a schematic diagram of the power spot market strategy processing device proposed in the embodiments of this application;
[0051] Figure 5 This is a schematic diagram of the structure of the computer device proposed in the embodiments of this application. Detailed Implementation
[0052] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] In the electricity spot market, the day-ahead market, as a crucial component, is primarily responsible for optimizing and arranging electricity trading plans one day in advance to ensure the safe, economical, and efficient operation of the power system. Therefore, power plants need to formulate reasonable day-ahead strategies to determine their day-ahead electricity demand to maximize their profits. In related technologies, most rely on expert experience to assess profit opportunities in the electricity spot market and determine corresponding day-ahead strategies, subsequently deciding on planned day-ahead power generation or consumption based on these strategies.
[0054] In some application scenarios, with the rapid popularization of new energy power generation, wind power, photovoltaic power generation and other forms of power generation fluctuate greatly due to external conditions such as weather and terrain. The access of these intermittent energy sources often leads to a large difference between actual power generation and predicted values. This volatility significantly increases the complexity and uncertainty of application strategies.
[0055] In other application scenarios, power plants typically aim to maximize profits. Under this goal, different power plants have different preferences regarding profitability. For example, some power plants tend to have lower total profits but still need to ensure profitability every month, while others tend to minimize losses and have less concern about total profits.
[0056] However, relying on expert experience to determine day-ahead strategies is highly subjective and can usually only be adjusted through backtesting after the day-ahead strategy has been formulated. This results in poor accuracy of day-ahead strategies and makes it difficult to respond in a timely manner to dynamic changes in the electricity spot market and personalized needs of users, thereby limiting the efficiency and effectiveness of electricity spot trading.
[0057] The agent-based power spot market strategy processing method provided in this specification can be applied to agent architectures based on reinforcement learning (RL). Specifically, it can be set up as an application in a computer device, which may include industrial control equipment, servers, or distributed computing terminals. These computer devices generate and select strategies by running the agent application.
[0058] An intelligent agent is an agent that can perceive its environment and take actions to achieve specific goals. Reinforcement learning provides the learning foundation and algorithmic support for the training phase of an intelligent agent, enabling it to perceive changes in the environment (such as through sensors or data input), make judgments and decisions based on its learned knowledge and algorithms, and then execute actions to achieve predetermined goals.
[0059] According to an embodiment of this application, an embodiment of a method for processing electricity spot market strategies by an intelligent agent is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0060] This embodiment provides a method for processing electricity spot market strategies for intelligent agents, which can be used in the aforementioned computer equipment, such as industrial control equipment, servers, or distributed computing terminals. Figure 1 This is a flowchart of a power spot market strategy processing method according to an embodiment of this application, such as... Figure 1 As shown, the process includes the following steps:
[0061] Step S1: Obtain the current individual status data of the target individual and the current environmental status data of the electricity spot market. The current individual status data is used to characterize the power generation forecast of the target individual on the current operating day, and the current environmental status data is used to characterize the total load and total unit output forecast in the electricity spot market on the current operating day.
[0062] In this embodiment, the target individual can be a power plant participating in electricity spot trading or an electricity-consuming enterprise with its own generating units. This embodiment utilizes current individual status data to obtain the target individual's operational capabilities and production status from an individual perspective, and utilizes current environmental status data to obtain the trading environment status of the entire electricity spot market from an external perspective, thereby providing a data foundation from both individual and external dimensions for subsequently generating multiple candidate strategies.
[0063] Step S3: Input the current individual state data and the current environment state data into the first target agent to generate multiple candidate policies; wherein, each candidate policy corresponds to a specified policy factor.
[0064] In this embodiment, a first target agent performs computational processing on two types of data: individual and external dimensions. This allows the agent to generate multiple candidate strategies corresponding to different specified strategy factors by utilizing its ability to perceive the state. These specified strategy factors represent different focuses of the strategy, and the candidate strategies are generated based on the current individual state and market environment, thus enabling a more comprehensive generation of candidate strategies.
[0065] Step S5: Input the current environmental state data into the second target agent to determine the acceptance probability corresponding to the specified strategy factor.
[0066] This application embodiment uses a second target intelligent agent to calculate and process data from external dimensions, thereby utilizing the agent's ability to perceive the state to generate acceptance probabilities corresponding to specified strategy factors. The acceptance probabilities can reflect the applicability of each specified strategy factor under the current environmental state data, that is, the probability of maximizing the profit by selecting the candidate strategy corresponding to the specified strategy factor to participate in electricity spot trading, thus providing a scientific and quantitative screening basis for subsequent candidate strategies.
[0067] Step S7: Screen multiple candidate strategies based on the acceptance probability to obtain the target strategy for the target individual to conduct electricity spot trading on the current operating day.
[0068] The embodiments of this application select a target strategy from candidate strategies based on the acceptance probability, and then execute relevant power dispatch instructions based on the target strategy to ensure that the target individual achieves the expected returns in the electricity spot market.
[0069] The electricity spot market strategy processing method provided in this embodiment uses a first target agent to perceive and make decisions based on the current individual state data and the current environmental state data, thereby generating candidate strategies corresponding to different specified strategy factors. Simultaneously, a second target agent determines the acceptance probability of each specified strategy factor based on the environmental state data. Finally, based on the acceptance probability, the most suitable candidate strategy corresponding to the strategy factor is selected to obtain the target strategy. Therefore, this embodiment uses a reinforcement learning-based agent to perceive the real-time state data of the target individual and the electricity spot market, providing comprehensive and real-time input information for subsequent strategy generation and selection. Furthermore, the first agent can generate candidate strategies corresponding to different specified strategy factors, allowing for different focuses among multiple candidate strategies, improving the flexibility of strategy generation. Combined with the acceptance probability output by the second agent, it responds to changes in the electricity spot market, thereby achieving adaptive selection of candidate strategies. Compared with traditional strategy selection methods based on fixed rules or experience, this method significantly improves the scientific rigor, accuracy, and efficiency of strategy generation and selection, ensuring that the target strategy most suitable for the current electricity spot market state is selected, which is beneficial for the target individual to achieve its expected returns in the electricity spot market.
[0070] In some embodiments of this application, the specified strategy factors are at least one of the following: load factors of the electricity spot market, wind power output factors, photovoltaic power output factors, external power transmission factors, and power generation factors of the target individual.
[0071] Specifically, this application embodiment predetermines multiple strategy factors, including load factors in the electricity spot market, wind power output factors, photovoltaic power output factors, external transmission power factors, and the power generation factors of the target individual. Then, based on the current trading environment of the electricity spot market, some or all of these strategy factors are selected. For example, if neither the electricity-consuming enterprise nor the power-generating enterprise participates in electricity spot trading with external transmission power, then only the load factor, wind power output factor, photovoltaic power output factor, and power generation factor can be selected as the aforementioned designated strategy factors.
[0072] Furthermore, in some embodiments of this application, the power generation forecast data of the target individual on the current operating day is obtained according to the power generation factor of the target individual, and the hyperparameter rolling_n is set to represent the number of days for rolling standardization processing. The average value of the power generation forecast data is calculated using the hyperparameter rolling_n, and then the percentage of the average value of the power generation forecast data relative to the installed capacity is calculated to achieve standardized dimensional processing.
[0073] Similarly, for the load factors, wind power output factors, photovoltaic power output factors, and external power transmission factors corresponding to the electricity spot market, relevant data published by the trading center are obtained, including the regional load forecast data, regional wind power output forecast data, regional photovoltaic power output forecast data, and regional external power transmission forecast data for the current operating day. The average value of these data is calculated using the hyperparameter rolling_n to achieve standardization.
[0074] In some embodiments of this application, the first target agent includes multiple simple policy generation layers and a baseline policy generation layer, wherein the simple policy generation layers correspond one-to-one with the specified policy factors.
[0075] This application embodiment sets up multiple simple policy generation layers in the first intelligent agent to independently generate corresponding candidate policies for multiple specified factors, so that the policy generation stage has multiple different focuses, providing a wider range of options for subsequent policy screening and improving the flexibility and comprehensiveness of candidate policies.
[0076] In some embodiments of this application, the multiple simple policy generation layers include a first simple policy generation layer, a second simple policy generation layer, a third simple policy generation layer, and a fourth simple policy generation layer.
[0077] Furthermore, step S3 above includes:
[0078] The predicted total power generation of all units in the current individual state data and the current predicted total load in the current environmental state data are input into the first simple strategy generation layer to generate the first candidate strategy corresponding to the load factor. Therefore, the first candidate strategy is generated only based on the above predicted total power generation and the above predicted total load, so that the focus of the first candidate strategy is on the load situation of participating in the electricity spot market, and the total load is used as the basis for determining the day-ahead declared electricity volume.
[0079] The predicted wind power generation from the current individual state data and the predicted total wind power output from the current environmental state data are input into the second simple strategy generation layer to generate a second candidate strategy corresponding to the wind power output factor. Therefore, the second candidate strategy is generated only based on the above predicted wind power generation and the above predicted total wind power output, making the focus of the second candidate strategy on the wind power output of wind turbines participating in the electricity spot market, and using the wind turbine output as the basis for determining the day-ahead declared electricity volume.
[0080] The predicted photovoltaic (PV) power generation from the current individual state data and the predicted total PV output from the current environmental state data are input into the third simple strategy generation layer to generate a third candidate strategy corresponding to the PV output factor. Therefore, the third candidate strategy is generated solely based on the aforementioned predicted PV power generation and total PV output, making the focus of the third candidate strategy on the PV unit output situation participating in the electricity spot market, and using the PV unit output situation as the basis for determining the day-ahead declared electricity volume.
[0081] The predicted individual power transmission volume of all units in the current individual state data and the predicted total power transmission volume of all units in the current environmental state data are input into the fourth simple strategy generation layer to generate a fourth candidate strategy corresponding to the power transmission volume factor. Therefore, the fourth candidate strategy is generated only based on the above predicted individual power transmission volume and the above predicted total power transmission volume, so that the focus of the fourth candidate strategy is on the power transmission volume participating in the electricity spot market, and the power transmission volume is used as the basis for determining the day-ahead declared power volume.
[0082] The predicted total power generation of all units in the current individual status data and the installed capacity of the target individual are input into the benchmark strategy generation layer. The ratio between the predicted total power generation and the installed capacity is determined, and a benchmark candidate strategy corresponding to the power generation factor is generated based on this ratio. Therefore, the benchmark candidate strategy is generated only based on the above-mentioned predicted total power generation and installed capacity. This makes the benchmark strategy generation layer focus on the power generation situation of the target individual itself, and uses the power generation situation at the individual level as the basis for determining the day-ahead declared electricity volume. The benchmark candidate strategy describes the passive trading behavior of the target individual, which does not consider external market factors and only considers its own power generation situation.
[0083] The plurality of candidate strategies include the first candidate strategy, the second candidate strategy, the third candidate strategy, the fourth candidate strategy, and the benchmark candidate strategy.
[0084] It should be noted that in this embodiment, the above-mentioned candidate strategy is a mapping S from the state data set to the day-ahead reported electricity. At the same time, since the installed capacity of different power plants is different, their power generation will inevitably have scale differences. Therefore, the candidate strategy is normalized, so the candidate strategy can be expressed as q_percent = q_ahead / capacity = S(d), where q_percent represents the proportion of the day-ahead reported electricity of the target individual to the installed capacity, q_ahead represents the day-ahead reported electricity of the target individual, capacity represents the installed capacity of the target individual, which is usually a constant for the target individual, and d represents the data set of the current individual state data and the current environmental state data.
[0085] Therefore, this embodiment of the application generates candidate strategies for load conditions, wind power processing conditions, photovoltaic processing conditions, and external power transmission conditions by setting a first simple strategy generation layer, a second simple strategy generation layer, a third simple strategy generation layer, and a fourth simple strategy generation layer, respectively. At the same time, a benchmark strategy generation layer is set to generate corresponding benchmark candidate strategies. Thus, the benchmark candidate strategies reflect the passive trading behavior of the target individual that only considers its own power generation. Therefore, this embodiment of the application can generate strategies separately for different strategy factors through different strategy generation layers, thereby improving the flexibility and comprehensiveness of strategy generation.
[0086] Furthermore, the simple policy generation layer is a neural network constructed based on the DDPG algorithm, and each simple policy generation layer includes an actor neural network using the sigmoid activation function and a critic neural network using the tanh activation function. The actor neural network is used to generate multiple simple policies corresponding to the specified policy factor, and the critic neural network is used to determine the candidate policy corresponding to the specified policy factor based on the payoff of the multiple simple policies.
[0087] The DDPG algorithm is a deep reinforcement learning-based algorithm, particularly suitable for solving problems in continuous action spaces. In this embodiment, a simple policy generation layer is constructed in the first target agent based on the DDPG algorithm. This layer generates corresponding simple policies through an actor neural network, and a critic neural network evaluates the returns of these simple policies to generate candidate policies with the best returns under specified policy factors.
[0088] Specifically, in some embodiments of this application, such as Figure 2 As shown, the first, second, third, and fourth simple policy generation layers all employ a continuous action space (DDPG) agent. Each DDPG agent includes an actor neural network and a critic neural network. The actor neural network is responsible for selecting and executing actions based on the current state. It acts as the agent's "decision-maker," typically using a policy function to determine the action to take in each state; this action generates a simple policy. The critic neural network evaluates the value of the actions selected by the actor neural network. It acts as the agent's "evaluator," providing a value judgment based on the payoff, informing the actor neural network whether the selected action is good or bad. Therefore, the higher the payoff of the simple policy, the better the evaluation result of the critic neural network, thus guiding the actor neural network to output the candidate policy with the highest payoff.
[0089] Furthermore, the first, second, third, and fourth simple policy generation layers all employ a two-layer Multi-Layer Perceptron (MLP) network, with a neural network using sigmoid as the activation function as the actor neural network, and a two-layer MLP network with a neural network using tanh as the activation function as the critic neural network. The actor neural network generates simple policies based on the input data, and the critic neural network evaluates the value of the simple policies generated by the actor neural network, thereby guiding the actor neural network to output corresponding candidate policies. The benchmark policy generation layer directly uses the aforementioned ratio of predicted total power generation to installed capacity as the output to represent negative trading behavior.
[0090] In some embodiments of this application, such as Figure 3 As shown, the second target agent is a DQN agent employing a discrete action space, and the second target agent includes an actor neural network, which uses the DQN algorithm and the sigmoid activation function. Specifically, a 3-layer MLP network using the sigmoid activation function serves as the actor neural network.
[0091] The DQN algorithm is an algorithm that combines deep learning and reinforcement learning. It uses deep neural networks to approximate the Q-function and is suitable for solving problems in discrete action spaces.
[0092] Furthermore, step S5 above includes:
[0093] The current environmental state data is input into the actor neural network of the second target agent to score the applicability of each specified policy factor under the current environmental state data based on the DQN algorithm.
[0094] The applicability score is normalized using the sigmoid activation function to obtain the acceptance probability for each specified policy factor.
[0095] In this embodiment of the application, the actor neural network of the second target intelligent agent is implemented using the DQN (Deep Q-Network) algorithm. Its goal is to score the applicability of different specified strategy factors in a specific environment through a deep learning network. The specific environment is the trading environment of the electricity spot market corresponding to the current environmental state data.
[0096] First, the current environmental state data is input into the actor neural network of the second target agent. The second target agent uses this data, combined with the trained network weights, to calculate the applicability score of each specified strategy factor under the current environmental state. Specifically, the DQN algorithm extracts features from the current environmental state data through the actor neural network, mapping the complex input state into an applicability score in a high-dimensional strategy space. Each applicability score corresponds to the potential feasibility and benefit of a specified strategy factor. For example, the applicability score can characterize value indicators such as return stability and loss risk tolerance. Each applicability score reflects the priority or suitability of these specified strategy factors in the current electricity spot market trading environment. The higher the applicability score, the higher the probability that the corresponding specified strategy factor will be selected as a subsequent factor.
[0097] The suitability score is then normalized using the sigmoid activation function, converting the score into an acceptance probability within the range of [0,1]. The purpose of normalization is to standardize the suitability score, allowing for comparison of acceptance probabilities across different policy factors. The non-linear nature of the sigmoid function ensures that policy factors with higher scores correspond to larger acceptance probabilities, while those with lower scores correspond to smaller acceptance probabilities, thus providing a clear basis for subsequent policy selection.
[0098] Therefore, this application embodiment constructs a second target agent based on the DQN (Deep Q-Network) algorithm, uses the DQN algorithm to score the applicability of each specified policy factor, and then uses the sigmoid activation function to normalize the score, thereby determining the acceptance probability of each policy factor. This achieves accurate calculation of the acceptance probability, which is beneficial for the subsequent accurate screening of candidate policies.
[0099] Furthermore, in some embodiments of this application, step S7 includes: determining the target strategy factor corresponding to the maximum value of the acceptance probability, and selecting the candidate strategy corresponding to the target strategy factor from multiple candidate strategies as the target strategy.
[0100] Therefore, this application embodiment uses the acceptance probability to reflect the feasibility and expected returns of each candidate strategy under the current environmental state data, and determines the target strategy factor based on the maximum value of the acceptance probability. This avoids the limitations of manual intervention or fixed rules, thereby ensuring that the target strategy corresponding to the target strategy factor is the most likely strategy to succeed under the current electricity spot market environment, which is conducive to maximizing the returns of the target individual in electricity spot trading.
[0101] Therefore, in some embodiments of this application, the synergistic effect of the DDPG Agent and the DQN Agent is utilized to output the target policy. The DDPG Agent generates diverse candidate policies, ensuring the comprehensiveness and adaptability of the policy space; simultaneously, the DQN Agent evaluates the specified policy factors of these candidate policies and calculates the acceptance probability, thereby selecting the optimal target policy. The combination of DDPG and DQN achieves a division of labor between policy generation and evaluation, enabling the system to generate rich policy alternatives in complex environments and accurately identify and select the optimal policy in dynamic markets. This synergy significantly improves decision-making efficiency and accuracy, while taking into account both policy diversity and the scientific nature of the optimization process, ensuring that the target policy maximizes the benefits of the current environment and meets personalized needs.
[0102] In some embodiments of this application, the training process of the first target agent includes:
[0103] Obtain the historical strategy return and historical benchmark return for historical operating days. The historical strategy return is the actual return of the target individual on the historical operating day, and the historical benchmark return is the return obtained by the target individual from spot electricity trading based on the historical unit's predicted power generation on the historical operating day.
[0104] The initial reward is determined based on the difference between the historical strategy reward and the historical benchmark reward, and the first initial agent is trained based on the initial reward to obtain the first target agent.
[0105] Specifically, the historical strategy payoff is determined according to the following formula (1):
[0106] profit_1=(q_ahead-q_real)*(p_ahead-p_real)
[0107] limit = abs(0.4 * q_real * (p_ahead - p_real)) Formula (1)
[0108] profit = min(profit_1, limit)
[0109] In the formula, profit_1 represents the actual profit of the output historical target strategy, q_ahead represents the day-ahead declared electricity volume corresponding to the historical strategy, q_real represents the actual power generation of the target individual on the historical operating day, p_ahead represents the day-ahead electricity price on the historical operating day, and the day-ahead electricity price is determined by the declaration behavior of market participants. The trading center's program calculates the electricity price based on the declaration behavior of market participants, p_real represents the real-time electricity price on the historical operating day, and the real-time electricity price is determined by the declaration behavior of market participants and the real-time supply and demand situation, and profit represents the final historical strategy profit.
[0110] Similarly, the historical benchmark return is determined according to the following formula (2):
[0111] profit_2=(q_ahead_b-q_real)*(p_ahead-p_real)
[0112] limit = abs(0.4 * q_real * (p_ahead - p_real)) Formula (2)
[0113] baseline_profit=min(profit_2,limit)
[0114] In the formula, profit_2 represents the actual return of the historical benchmark strategy, q_ahead_b represents the day-ahead declared electricity volume corresponding to the historical benchmark strategy, q_real represents the actual power generation of the target individual on the historical operating day, p_ahead represents the day-ahead electricity price on the historical operating day, and the day-ahead electricity price is determined by the declaration behavior of market participants. The trading center's program calculates the electricity price based on the declaration behavior of market participants, p_real represents the real-time electricity price on the historical operating day, and the real-time electricity price is determined by the declaration behavior of market participants and the real-time supply and demand situation, and baseline_profit represents the final historical benchmark return.
[0115] The original reward is then calculated according to the following formula (3):
[0116] reward = profit - baseline_profit formula (3)
[0117] In this embodiment, the first initial agent is trained using the original reward to obtain the first target agent.
[0118] It's important to note that power generation forecasting at power plants is often objectively challenging in the electricity spot market. For example, in some cases, inaccurate weather forecasts can lead to significant discrepancies between predicted and actual power generation. This discrepancy not only impacts power generation plans but also directly affects market trading profits or losses. Without a specific trading strategy, these profits or losses are caused by the forecasting error itself, rather than by the optimization process of the trading strategy. Therefore, these profits or losses cannot be effectively learned from the environmental state and may interfere with the training of the agent.
[0119] To eliminate data noise caused by power prediction errors and prevent overfitting of the model during training due to these uncontrollable factors, this application introduces baseline_profit. By subtracting baseline_profit from the training data, the gains or losses caused by power prediction errors are removed from the environmental feedback, thereby ensuring that the model training focuses more on the optimization effect of the trading strategy itself.
[0120] Therefore, this embodiment determines the original reward by the difference between the historical strategy return and the historical benchmark return, thereby eliminating the data noise caused by the historical benchmark return corresponding to passive trading behavior without taking any strategy, reducing the possibility of overfitting in the first initial agent during training, and thus effectively improving the training effect of the first target agent.
[0121] Furthermore, in some embodiments of this application, the training process of the second target agent includes:
[0122] The continuous loss index of the target individual is determined based on the original reward; the continuous loss index includes at least one of the following: continuous loss days factor_1, historical maximum continuous loss days factor_2, continuous loss amount factor_3, and historical maximum continuous loss amount factor_4.
[0123] Specifically, factor_1 represents the number of consecutive days with a reward < 0 ending with the current step, factor_2 represents the maximum number of consecutive loss days for each running day before the current step, factor_3 represents the sum of rewards for consecutive days with a reward < 0 ending with the current step, and factor_4 represents the minimum amount of consecutive losses for each running day before the current step. It should be noted that because the reward is negative in the case of consecutive losses, factor_4 takes the minimum value.
[0124] When the revenue consecutive loss indicator factor_i reaches the preset limit threshold_i, adjust the hyperparameters according to the preferences of the target individual to proportionally adjust the original reward, and obtain the adjusted reward value.
[0125] Specifically, i = {1, 2, 3, 4}. When factor_i < threshold_i, adjust the subsequent k trading days according to the following formula (4):
[0126] reward’ = reward - reward * adjust_percent Formula (4)
[0127] Train the second initial agent according to the adjusted reward value reward’ to obtain the second target agent.
[0128] Therefore, in the embodiment of the present application, the hyperparameters are adjusted according to the preferences of the target individual to proportionally adjust the original reward, so as to train the second initial agent. Therefore, the preference of the user for strategy selection is introduced in the training process of the second initial agent, so that the trained second target agent can output a more accurate acceptance probability in response to the user's preference, so that the user's strategy preference is integrated into the strategy generation process, and further ensure that the selected target strategy maximizes the revenue of the electricity spot trading on the premise of meeting the user's preference.
[0129] In addition, in the embodiment of the present application, a two-agent structure is adopted, and the first initial agent is trained with the original reward reward before adjustment, so that the goal of each candidate strategy generated by the trained first target agent is to pursue maximum revenue. Since maximizing revenue is a relatively general strategy goal, the candidate strategies can be通用 on different target individuals, reducing the difficulty and workload of training. Then, train the second initial agent with the adjusted reward value reward’, and use the adjusted reward method to quantitatively respond to the specific strategy preferences of the target individual, thereby further improving the accuracy and applicability of the target strategy. If the first initial agent is directly trained with the adjusted reward value reward’, the risk of overfitting the candidate strategy to some special artificial adjustments will increase.
[0130] In addition, it should be noted that in the embodiment of the present application, in the application process of the first target agent and the second target agent, the original reward and the adjusted reward value in the above training process can be referred to, and the first target agent and the second target agent can be optimized in real time respectively. In this process, the original reward reward is the difference between the actual revenue of the previous transaction and the revenue corresponding to the benchmark strategy generated last time.
[0131] Correspondingly, please refer to Figure 4This application provides an intelligent agent's electricity spot market strategy processing device, the device comprising:
[0132] The acquisition module 100 is used to acquire the current individual status data of the target individual and the current environmental status data of the electricity spot market. The current individual status data is used to characterize the power generation forecast of the target individual on the current operating day, and the current environmental status data is used to characterize the total load and total unit output forecast in the electricity spot market on the current operating day.
[0133] The strategy generation module 200 is used to input the current individual state data and the current environment state data into the first target agent to generate multiple candidate strategies; wherein, the candidate strategies correspond to specified strategy factors;
[0134] The strategy screening module 300 is used to input the current environmental state data into the second target agent to determine the acceptance probability corresponding to the specified strategy factor, and to screen the multiple candidate strategies according to the acceptance probability to obtain the target strategy for the target agent to conduct electricity spot trading on the current operating day.
[0135] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0136] In this embodiment, the power spot market strategy processing device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0137] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application, such as... Figure 5 As shown, the computer device includes:
[0138] One or more processors 10, a memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The memory 20 and the processor 10 are communicatively connected to each other. The memory 10 stores computer instructions, and the processor 10 executes the aforementioned computer instructions to perform the electricity spot market strategy processing method described in the above embodiments.
[0139] It should be noted that the various components described above communicate and connect with each other using different buses, and can be installed on a common motherboard or otherwise as needed. The processor can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to an interface). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 10 as an example.
[0140] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.
[0141] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.
[0142] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0143] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0144] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.
[0145] This application also provides a computer-readable storage medium. The methods described in this application can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code downloaded over a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code. When the software or computer code is accessed and executed by the computer, processor, or hardware, the methods shown in the above embodiments are implemented.
[0146] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the method of any embodiment of this application.
[0147] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
[0148] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0149] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0150] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0151] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0152] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0153] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0154] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0155] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0156] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
[0157] Although embodiments of this application have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of this application, and all such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A method for processing electricity spot market strategies for intelligent agents, characterized in that, The method includes: The system acquires the current individual status data of the target individual and the current environmental status data of the electricity spot market. The current individual status data is used to characterize the power generation forecast of the target individual on the current operating day, and the current environmental status data is used to characterize the total load and total unit output forecast in the electricity spot market on the current operating day. The current individual state data and the current environmental state data are input into the first target intelligent agent to generate multiple candidate strategies; wherein, each candidate strategy corresponds to a specified strategy factor, and the candidate strategy is used to determine the daily reported electricity volume; The current environmental state data is input into the second target agent to determine the acceptance probability corresponding to the specified strategy factor; the acceptance probability is obtained by scoring the applicability of each specified strategy factor under the current environmental state data and performing normalization processing. The multiple candidate strategies are screened based on the acceptance probability to obtain the target strategy for the target individual to conduct electricity spot trading on the current operating day; The specified strategy factors are at least one of the following: load factors of the electricity spot market, wind power output factors, photovoltaic power output factors, external power transmission factors, and power generation factors of the target individual. The first target intelligent agent includes multiple simple strategy generation layers and a benchmark strategy generation layer. The simple strategy generation layers correspond one-to-one with the specified strategy factors. The plurality of simple strategy generation layers include a first simple strategy generation layer, a second simple strategy generation layer, a third simple strategy generation layer, and a fourth simple strategy generation layer; The step involves inputting the current individual state data and the current environment state data into the first target agent to generate multiple candidate strategies, including: The predicted total power generation of all units in the current individual state data and the current predicted total load in the current environmental state data are input into the first simple strategy generation layer to generate a first candidate strategy corresponding to the load factor. The predicted wind power generation of the wind turbine in the current individual state data and the predicted total wind power output of the wind turbine in the current environmental state data are input into the second simple strategy generation layer to generate a second candidate strategy corresponding to the wind power output factor. The predicted photovoltaic power generation of the photovoltaic unit in the current individual state data and the predicted total photovoltaic output of the photovoltaic unit in the current environmental state data are input into the third simple strategy generation layer to generate a third candidate strategy corresponding to the photovoltaic output factor. The predicted individual power transmission volume of all the units in the current individual state data and the predicted total power transmission volume of all the units in the current environmental state data are input into the fourth simple strategy generation layer to generate a fourth candidate strategy corresponding to the power transmission volume factor. The predicted total power generation of all units in the current individual status data and the installed capacity of the target individual are input into the benchmark strategy generation layer, and the ratio between the predicted total power generation and the installed capacity is determined, so as to generate a benchmark candidate strategy corresponding to the power generation factor based on the ratio. The plurality of candidate strategies include the first candidate strategy, the second candidate strategy, the third candidate strategy, the fourth candidate strategy, and the benchmark candidate strategy.
2. The electricity spot market strategy processing method according to claim 1, characterized in that, The simple policy generation layer is a neural network built based on the DDPG algorithm, and each simple policy generation layer includes an actor neural network with a sigmoid activation function and a critic neural network with a tanh activation function. The actor neural network is used to generate multiple simple policies corresponding to the specified policy factor, and the critic neural network is used to determine the candidate policy corresponding to the specified policy factor based on the payoff of the multiple simple policies.
3. The electricity spot market strategy processing method according to claim 1, characterized in that, The second target agent includes an actor neural network, which employs the DQN algorithm and a sigmoid activation function; the step of inputting the current environmental state data into the second target agent to determine the acceptance probability corresponding to the specified strategy factor includes: The current environmental state data is input into the actor neural network of the second target agent to score the applicability of each specified strategy factor under the current environmental state data based on the DQN algorithm. The applicability score is normalized using the sigmoid activation function to obtain the acceptance probability for each specified strategy factor.
4. The electricity spot market strategy processing method according to claim 1, characterized in that, The step of filtering the multiple candidate strategies based on the acceptance probability to obtain the target strategy for the target individual to conduct electricity spot trading on the current operating day includes: The target strategy factor corresponding to the maximum value of the acceptance probability is determined, and the candidate strategy corresponding to the target strategy factor is selected from the plurality of candidate strategies as the target strategy.
5. The method for handling electricity spot market strategies according to claim 1, characterized in that, The training process of the first target agent includes: Obtain the historical strategy return and historical benchmark return for historical operating days, wherein the historical strategy return is the actual return of the target individual on the historical operating day, and the historical benchmark return is the return obtained by the target individual from electricity spot trading based on the historical unit's predicted power generation on the historical operating day; The initial reward is determined based on the difference between the historical strategy reward and the historical benchmark reward, and the first initial agent is trained based on the initial reward to obtain the first target agent.
6. The electricity spot market strategy processing method according to claim 5, characterized in that, The training process for the second target agent includes: The continuous loss indicator for the target individual's earnings is determined based on the original reward. When the continuous loss indicator reaches a preset limit, the original reward is proportionally adjusted according to the target individual's preference adjustment hyperparameter, and the adjusted reward value is obtained. The second initial agent is trained according to the adjusted reward value to obtain the second target agent; The continuous loss indicator includes at least one of the following: number of consecutive loss days, maximum historical number of consecutive loss days, amount of consecutive loss, and maximum historical amount of consecutive loss.
7. A computer device, characterized in that, include: A memory and a processor are interconnected, the memory stores computer instructions, and the processor executes the power spot market strategy processing method according to any one of claims 1 to 6 by executing the computer instructions.
Citation Information
Patent Citations
Electric power transaction auxiliary decision-making method and device for spot market
CN115293802A
Method and device for determining spot electricity market transaction strategy of energy storage power station
CN116307100A