Demand side response electricity market transaction decision-making method

By combining chaos theory, particle swarm optimization algorithm and deep reinforcement learning algorithm, a demand-side response power market transaction decision-making method is formed, solving the problems of the existing technology in the face of complexity and uncertainty of the power market, and achieving more efficient and flexible trading decisions.

CN120218984APending Publication Date: 2025-06-27HONGHE POWER SUPPLY BUREAU OF YUNNAN POWER GRID
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510197904.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When the existing trading decision-making methods of the power market are faced with the complexity and uncertainty of the power market, it is difficult to effectively formulate trading strategies, and the calculation complexity is high and the adaptability is poor.

Method used

The integrated method is adopted, including chaos theory for price prediction, particle swarm optimization algorithm optimization trading strategies, and deep reinforcement learning algorithm learning and adjusting trading strategies, forming a demand-side response method for the power market trading decision-making.

Benefits of technology

By accurately predicting the electricity market prices, optimizing trading strategies, reducing trading risks, improving decision-making efficiency, improving transaction profitability, and suitable for complex and dynamic electricity market environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218984A_ABST
    Figure CN120218984A_ABST
Patent Text Reader

Abstract

The invention relates to a demand side response electricity market transaction decision-making method, and belongs to the technical field of electricity market prediction. The method comprises five steps of collecting and processing data, predicting price by using a chaos theory, designing a power transaction strategy, optimizing the transaction strategy by using a particle swarm optimization algorithm, and learning and adjusting the transaction strategy by using a deep reinforcement learning algorithm. According to the method, power market data can be collected and processed in real time, prediction and decision making can be rapidly made, and the method is particularly important for the power market with rapid demand change. The method is not only suitable for the electricity market, but also can be applied to other similar demand side response markets, such as fuel gas, heating power and the like, and has very high reusability and wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power market forecasting, and particularly relates to a power market trading decision-making method for demand-side response. Background Art

[0002] Power market trading is an important link in modern power systems, and it plays a key role in aspects such as power supply-demand balance, power system planning, and power price decision-making. However, due to the high complexity and uncertainty of the power market, such as the randomness of power demand, the volatility of power prices, and the uncertainty of renewable energy, it poses great challenges to formulating effective trading strategies in the power market.

[0003] In traditional power trading decision-making methods, methods such as economic decentralization, power system operation optimization, and risk management are often used to solve this problem. However, these methods are often based on certain model assumptions, require a large amount of calculations, and have poor adaptability to market dynamic changes. For example, the economic decentralization method needs to model the entire power market, with high computational complexity and high requirements for the prediction accuracy of key parameters such as prices; the power system operation optimization method needs to model the operation status and equipment parameters of the power system in detail, with high computational complexity and poor adaptability; the risk management method often needs to accurately estimate market risks, which is often difficult to achieve in practical applications.

[0004] In recent years, with the development of artificial intelligence and data analysis technologies, more and more advanced methods have been proposed and applied to power market trading decision-making. These methods usually have stronger adaptability and flexibility, and can better adapt to the dynamic changes and uncertainties of the power market. For example, methods such as chaos theory, optimization algorithms, and deep reinforcement learning have been widely applied to the prediction of power market prices and the optimization of trading strategies. These methods can usually achieve higher returns and risk control. However, their implementation processes are often more complex, require a large amount of computing resources, and have high requirements for data quality and consistency.

[0005] Therefore, how to overcome the deficiencies of the existing technology is an urgent problem to be solved in the current technical field of power market forecasting. Summary of the Invention

[0006] The purpose of the present invention is to solve the deficiencies of the existing technology and provide a power market trading decision-making method for demand-side response.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0008] A power market trading decision-making method for demand-side response includes the following steps:

[0009] Step (1), data collection and processing: Collect historical and real-time data of the electricity market, and then perform preprocessing; the data includes price, demand and supply volume, and weather data;

[0010] Step (2), predicting price using chaos theory: Adopt the time series analysis method of chaos theory to analyze the historical data of electricity market prices, predict their future trends, and obtain price prediction results;

[0011] Step (3), designing an electricity trading strategy: According to the price prediction results, design a preliminary electricity trading strategy to allocate the electricity volume for purchase and sale;

[0012] Step (4), optimizing the trading strategy using the particle swarm optimization algorithm: Regard the trading strategy designed in step (3) as the solution to the optimization problem, and use the particle swarm optimization algorithm to find the optimal solution in the trading strategy space to obtain the optimal electricity trading strategy;

[0013] Step (5), learning and adjusting the trading strategy using the deep reinforcement learning algorithm: Use the optimal electricity trading strategy obtained by the particle swarm optimization algorithm as the initial solution to train the deep reinforcement learning model; the deep reinforcement learning model continuously tries new electricity trading strategies, adjusts the network weights according to the feedback, and finally obtains a more optimized electricity trading strategy.

[0014] Furthermore, preferably, in step (1), the preprocessing includes data cleaning, missing value processing, and normalization;

[0015] Among them, when cleaning the data, remove the outliers in the data;

[0016] When processing missing values, use the method of linear interpolation to fill in the missing values;

[0017] The normalization adopts the min-max normalization method.

[0018] Furthermore, preferably, the specific method of step (2) is:

[0019] (2.1) Reconstruct the phase space: Utilize the Takens theorem to embed the one-dimensional time series data into a high-dimensional phase space through the delay coordinate method; the specific steps are as follows:

[0020] Select appropriate time delay τ and embedding dimension m;

[0021] Construct vectors in the phase space:

[0022] X(t) = [x(t), x(t + τ), x(t + 2τ),..., x(t + (m - 1)τ)]

[0023] Among them, \(x(t)\) is the value of the time series at time \(t\), which is the electricity price corresponding to time \(t\); \(\tau\) is the time delay, and \(m\) is the embedding dimension, both of which are determined by the method of phase space reconstruction;

[0024] (2.2) Calculate the maximum Lyapunov exponent:

[0025] Select two adjacent points \(X_1\) and \(X_2\) in the initial phase space, and calculate their initial distance \(d(0)\);

[0026] As time evolves, calculate their distance \(d(t)\) at time \(t\);

[0027] The calculation formula for the maximum Lyapunov exponent \(\lambda\) is:

[0028]

[0029] Among them, \(N\) is the number of time steps;

[0030] (2.3) Predict future prices: Use the local nearest neighbor prediction method to predict future prices.

[0031] Furthermore, preferably, the specific method of step (3) is:

[0032] (3.1) Determine the trading objective: The trading objective includes profit maximization and risk control; profit maximization is achieved by maximizing the net income, and risk control is achieved by setting the risk tolerance to control risks; the risk tolerance is the maximum loss limit or volatility limit;

[0033] (3.2) Analyze the prediction results: Analyze the price prediction results obtained in step (2) to identify the peak and trough periods of prices; the peak period is suitable for selling electricity, while the trough period is suitable for buying electricity;

[0034] (3.3) Establish a trading model: Based on the predicted price data, establish an electricity trading model;

[0035] The electricity trading model aims to maximize the net income, and its objective function is:

[0036]

[0037] \(P\) t sell is the electricity price for selling electricity at time \(t\);

[0038] is the amount of electricity planned to be sold at time \(t\);

[0039] \(P\) t buy is the electricity price for buying electricity at time \(t\);

[0040] Electric power quantity planned to be purchased at time t;

[0041] The constraint conditions of the described power trading model include:

[0042] Power balance constraint:

[0043]

[0044] Among them, is the available electric power quantity at time t;

[0045] Risk control constraint:

[0046]

[0047] Step (3.4), strategy optimization: Solve the above model using linear programming to obtain a preliminary power trading strategy.

[0048] Furthermore, preferably, the specific method of step (4) is:

[0049] (4.1) Initialization: Randomly initialize the particle swarm. Each particle represents a possible trading strategy, that is, a combination of the power purchase quantity and the power sale quantity; Initialize the velocity and position of each particle; Set the number of particles as N, and randomly allocate the purchase quantity and sale quantity of each particle within a predetermined range; Initialize the velocity of each particle to zero;

[0050] (4.2) Evaluate fitness: Calculate the fitness value of each particle, that is, calculate the total revenue of the current trading strategy according to the objective function as the fitness value;

[0051] (4.3) Update velocity and position:

[0052] Update the particle velocity:

[0053]

[0054] Among them, v i (t) is the velocity of particle i at time t, ω is the inertia weight, c1 and c2 are acceleration constants, r1 and r2 are random numbers, is the historical best position of particle i, g best is the global best position;

[0055] Update the particle position

[0056] x i (t + 1) = x i (t) + v i (t + 1)

[0057] Among them, x i(t) is the position of particle i at time t;

[0058] (4.4) Update the best position: If the fitness of the current particle is better than its historical best position, update its historical best position; if the fitness of the current particle is better than the global best position, update the global best position;

[0059] (4.5) Iteration: Repeat steps (4.2) to (4.4) until a predetermined number of iterations is reached or the fitness gain is lower than a set threshold;

[0060] (4.6) Output the optimal strategy: After multiple iterations, output the trading strategy corresponding to the global best position, which is the optimal power trading strategy.

[0061] Furthermore, preferably, the specific method of step (5) is:

[0062] (5.1) Initial solution setting: Use the optimal power trading strategy obtained by the particle swarm optimization algorithm as the initial solution;

[0063] (5.2) Deep reinforcement learning training:

[0064] (5.2.1) Environment modeling: Define the simulation environment of the power market, including the state space, action space, and reward function; the state space includes the current electricity price, historical price trend, and available electricity; the action space includes the electricity purchase volume and sales volume in each time period;

[0065] (5.2.2) Select the model: The deep reinforcement learning model selects the deep Q network;

[0066] (5.2.3) Training process:

[0067] Initialization: Initialize the weights of the deep Q network and use the optimal power trading strategy obtained by the particle swarm optimization algorithm as the initial solution;

[0068] Interaction: In the simulation environment, the model selects an action according to the current strategy, applies the action to the environment, and the environment feedbacks the next state and immediate reward;

[0069] Experience replay: Store the results of each interaction in the experience replay buffer;

[0070] Update the strategy: Use the gradient descent method to update the network weights; for the deep Q network, the goal is to minimize the Bellman error, and the objective function is:

[0071]

[0072] where r is the immediate reward; γ is the discount factor; θ and θ′ are the parameters of the current and target networks respectively; s is the state; a is the action;

[0073] Policy update: Update the policy by optimizing the objective function to ensure that the difference between the new and old policies does not exceed a threshold to stabilize training:

[0074] (5.2.4) Policy optimization: During training, the model continuously adjusts the policy to maximize the cumulative reward;

[0075]

[0076] where, G t is the cumulative reward at time t, r t+k is the immediate reward at time t + k, and γ is the discount factor;

[0077] (5.2.5) Evaluation and adjustment: Regularly evaluate the performance of the training policy in the simulation environment and adjust the model parameters and training policy according to the evaluation results.

[0078] Furthermore, preferably, in step (4.3), ω is an inertial weight with a value range between [0.4, 0.9].

[0079] Furthermore, preferably, in step (4.3), both c1 and c2 are set to 2.0; both r1 and r2 are random numbers between 0 and 1.

[0080] Furthermore, preferably, in step (4.5), the number of iterations is 1000 times.

[0081] In step (1) of the present invention, price data is used to predict future electricity price trends; demand and supply data are used to evaluate market supply and demand balance; weather data is used to consider renewable energy power generation and its impact on the market.

[0082] In the electricity market, how to effectively formulate trading strategies to achieve maximum profit and minimum risk is an important but complex issue. Especially in the demand-side response (DSR) environment, due to the volatile characteristics of electricity demand and price, as well as the influence of various uncertain factors, formulating effective trading strategies becomes even more complex. Therefore, a trading decision-making method that can handle this complexity and adapt to the dynamic changes of the electricity market is needed.

[0083] The present invention proposes a trading decision-making method for the demand-side response electricity market. By integrating chaos theory, particle swarm optimization algorithm, and deep reinforcement learning algorithm, it makes full use of the historical and real-time data of the electricity market to predict prices and formulate and adjust trading strategies according to the prediction results, thereby realizing the intelligence and optimization of electricity market trading

[0084] Compared with the prior art, the beneficial effects of the present invention are:

[0085] This method aims to maximize the profit by optimizing the electricity trading strategy, specifically, how to allocate the amount of electricity purchased and sold.

[0086] By accurately predicting the electricity market price and formulating an electricity trading strategy based on it, the trading risk can be reduced. Meanwhile, the model also takes into account risk control factors when designing the trading strategy, further reducing the risk brought by market fluctuations.

[0087] Through the automated collection and processing of electricity market data and the optimization of the deep learning model, accurate trading decisions can be provided in a short time, greatly improving the efficiency of decision-making.

[0088] This method uses the deep reinforcement learning algorithm to self-adjust and learn the trading strategy, enabling the model to gradually optimize the strategy after multiple transactions and improve the profitability of trading.

[0089] This method is not only applicable to the electricity market but also can be applied to other similar demand response markets, such as gas and heat, etc., with high reusability and broad application prospects.

[0090] This method can collect and process electricity market data in real time and make predictions and decisions quickly, which is particularly important for the electricity market with rapid demand changes and can help power companies gain an invincible position in the ever-changing market. Description of the Drawings

[0091] Figure 1 It is a flowchart of the electricity market trading decision-making method for demand response of the present invention. Detailed Embodiments

[0092] The present invention will be further described in detail below in conjunction with the embodiments.

[0093] Those skilled in the art will understand that the following embodiments are only used to illustrate the present invention and should not be construed as limiting the scope of the present invention. For those not specified in the embodiments regarding specific technologies or conditions, they shall be carried out according to the technologies or conditions described in the literature in this field or according to the product specifications. For those materials or equipment not specified regarding the manufacturer, they are all conventional products that can be obtained through purchase.

[0094] Embodiment 1

[0095] A method for making electricity market trading decisions for demand response includes the following steps:

[0096] Step (1), data collection and processing: Collect historical and real-time data of the electricity market and then perform preprocessing; the data includes price, demand and supply volume, and weather data.

[0097] Step (2), predicting price using chaos theory: Using the time series analysis method of chaos theory, analyze the historical data of the electricity market price to predict its future trend and obtain the price prediction result;

[0098] Step (3), designing an electricity trading strategy: According to the price prediction result, design a preliminary electricity trading strategy to allocate the amounts of electricity purchased and sold;

[0099] Step (4), optimizing the trading strategy using the particle swarm optimization algorithm: Consider the trading strategy designed in step (3) as the solution to the optimization problem, and use the particle swarm optimization algorithm to find the optimal solution in the trading strategy space to obtain the optimal electricity trading strategy;

[0100] Step (5), learning and adjusting the trading strategy using the deep reinforcement learning algorithm: Use the optimal electricity trading strategy obtained by the particle swarm optimization algorithm as the initial solution to train the deep reinforcement learning model; The deep reinforcement learning model continuously tries new electricity trading strategies and adjusts the network weights according to the feedback to finally obtain a more optimized electricity trading strategy.

[0101] Embodiment 2

[0102] An electricity market trading decision-making method for demand response includes the following steps:

[0103] Step (1), data collection and processing: Collect the historical and real-time data of the electricity market and then perform preprocessing; The data includes price, demand and supply amount, and weather data;

[0104] Step (2), predicting price using chaos theory: Using the time series analysis method of chaos theory, analyze the historical data of the electricity market price to predict its future trend and obtain the price prediction result;

[0105] Step (3), designing an electricity trading strategy: According to the price prediction result, design a preliminary electricity trading strategy to allocate the amounts of electricity purchased and sold;

[0106] Step (4), optimizing the trading strategy using the particle swarm optimization algorithm: Consider the trading strategy designed in step (3) as the solution to the optimization problem, and use the particle swarm optimization algorithm to find the optimal solution in the trading strategy space to obtain the optimal electricity trading strategy;

[0107] Step (5), learning and adjusting the trading strategy using the deep reinforcement learning algorithm: Use the optimal electricity trading strategy obtained by the particle swarm optimization algorithm as the initial solution to train the deep reinforcement learning model; The deep reinforcement learning model continuously tries new electricity trading strategies and adjusts the network weights according to the feedback to finally obtain a more optimized electricity trading strategy.

[0108] In step (1), the preprocessing includes data cleaning, missing value handling, and normalization;

[0109] Among them, when cleaning the data, the outliers in the data are removed;

[0110] When handling missing values, the method of linear interpolation is used to fill in the missing values;

[0111] The normalization adopts the min-max normalization method.

[0112] The specific method of step (2) is as follows:

[0113] (2.1) Reconstruct the phase space: Using the Takens theorem, the one-dimensional time series data is embedded into the high-dimensional phase space through the delay coordinate method; the specific steps are as follows:

[0114] Select appropriate time delay τ and embedding dimension m;

[0115] Construct the vector in the phase space:

[0116] X(t) = [x(t), x(t + τ), x(t + 2τ),..., x(t + (m - 1)τ)]

[0117] Among them, x(t) is the value of the time series at time t, which is the electricity price corresponding to time t; τ is the time delay, and m is the embedding dimension, both of which are determined by the method of phase space reconstruction;

[0118] (2.2) Calculate the maximum Lyapunov exponent:

[0119] Select two adjacent points X1 and X2 in the initial phase space and calculate their initial distance d(0);

[0120] As time evolves, calculate their distance d(t) at time t;

[0121] The calculation formula for the maximum Lyapunov exponent λ is:

[0122]

[0123] Among them, N is the number of time steps;

[0124] (2.3) Predict future prices: The local nearest neighbor prediction method is used to predict future prices.

[0125] The specific method of step (3) is as follows:

[0126] (3.1) Determine the trading objectives: The trading objectives include profit maximization and risk control; profit maximization is achieved by maximizing the net income, and risk control is achieved by setting the risk tolerance to control risks; the risk tolerance mentioned is the maximum loss limit or volatility limit;

[0127] (3.2) Prediction result analysis: Analyze the price prediction results obtained in step (2) to identify the peak and trough periods of prices; the peak period is suitable for selling electricity, while the trough period is suitable for purchasing electricity;

[0128] (3.3) Establish a trading model: Based on the predicted price data, establish an electricity trading model;

[0129] The described electricity trading model aims to maximize the net profit, and its objective function is:

[0130]

[0131] P t sell is the electricity price for selling electricity at time t;

[0132] is the amount of electricity planned to be sold at time t;

[0133] P t buy is the electricity price for purchasing electricity at time t;

[0134] is the amount of electricity planned to be purchased at time t;

[0135] The constraints of the described electricity trading model include:

[0136] Power balance constraint:

[0137]

[0138] Among them, is the available power at time t;

[0139] Risk control constraint:

[0140]

[0141] Step (3.4), Strategy optimization: Use linear programming to solve the above model to obtain a preliminary electricity trading strategy.

[0142] The specific method of step (4) is:

[0143] (4.1) Initialization: Randomly initialize the particle swarm. Each particle represents a possible trading strategy, that is, a combination of the electricity purchase amount and the electricity sale amount; initialize the velocity and position of each particle; set the number of particles to N, and randomly allocate the purchase amount and sale amount of each particle within a predetermined range; initialize the velocity of each particle to zero;

[0144] (4.2) Evaluate fitness: Calculate the fitness value of each particle, that is, calculate the total revenue of the current trading strategy according to the objective function as the fitness value;

[0145] (4.3) Update velocity and position:

[0146] Update the particle velocity:

[0147]

[0148] where, v i (t) is the velocity of particle i at time t, ω is the inertia weight, c1 and c2 are acceleration constants, r1 and r2 are random numbers, is the historical best position of particle i, and g best is the global best position;

[0149] Update the particle position

[0150] x i (t + 1) = x i (t) + v i (t + 1)

[0151] where, x i (t) is the position of particle i at time t;

[0152] (4.4) Update the best position: If the fitness of the current particle is better than its historical best position, update its historical best position; if the fitness of the current particle is better than the global best position, update the global best position;

[0153] (4.5) Iteration: Repeat steps (4.2) to (4.4) until a predetermined number of iterations is reached or the fitness gain is lower than the set threshold;

[0154] (4.6) Output the optimal strategy: After multiple iterations, output the trading strategy corresponding to the global best position, which is the optimal power trading strategy.

[0155] The specific method of step (5) is:

[0156] (5.1) Initial solution setting: Use the optimal power trading strategy obtained by the particle swarm optimization algorithm as the initial solution;

[0157] (5.2) Deep reinforcement learning training:

[0158] (5.2.1) Environment modeling: Define the simulation environment of the power market, including the state space, action space, and reward function; the state space includes the current electricity price, historical price trend, and available power; the action space includes the electricity purchase volume and sales volume in each time period;

[0159] (5.2.2) Selection of Model: The deep reinforcement learning model selects the Deep Q-Network;

[0160] (5.2.3) Training Process:

[0161] Initialization: Initialize the weights of the Deep Q-Network and use the optimal electricity trading strategy obtained by the particle swarm optimization algorithm as the initial solution;

[0162] Interaction: In the simulation environment, the model selects an action according to the current strategy, applies the action to the environment, and the environment feedbacks the next state and the immediate reward;

[0163] Experience Replay: Store the results of each interaction in the experience replay buffer;

[0164] Update Strategy: Update the network weights using the gradient descent method; for the Deep Q-Network, the goal is to minimize the Bellman error, and the objective function is:

[0165]

[0166] where r is the immediate reward; γ is the discount factor; θ and θ′ are the parameters of the current and target networks respectively; s is the state; a is the action;

[0167] Policy Update: Update the policy by optimizing the objective function to ensure that the difference between the old and new policies does not exceed a threshold to stabilize the training:

[0168] (5.2.4) Policy Optimization: During the training process, the model continuously adjusts the policy to maximize the cumulative reward;

[0169]

[0170] where G t is the cumulative reward at time t, r t+k is the immediate reward at time t + k, and γ is the discount factor;

[0171] (5.2.5) Evaluation and Adjustment: Regularly evaluate the performance of the training policy in the simulation environment and adjust the model parameters and training policy according to the evaluation results.

[0172] Example 3

[0173] A decision-making method for electricity market transactions in demand response includes the following steps:

[0174] Step (1), Data Collection and Processing: Collect historical and real-time data of the electricity market, including price, demand and supply volume, and weather data. Then perform preprocessing, including data cleaning, missing value processing, and normalization.

[0175] In power market trading decisions, data collection and processing are crucial first steps.

[0176] First, various data sources in the power market need to be collected:

[0177] Price data: Obtain historical and real-time power market price data. This data can be obtained from power exchanges or market operators.

[0178] Demand and supply data: Collect historical and real-time power demand and supply data. This data can be obtained from power companies or relevant market reports.

[0179] Weather data: Include temperature, humidity, wind speed, and precipitation, as weather changes have a significant impact on power demand and supply. This data can be obtained from meteorological service providers.

[0180] Data preprocessing is a key step to ensure data quality and consistency, including the following aspects:

[0181] Data cleaning: Remove noise and outliers from the data.

[0182] Missing value handling: For missing values in the data, use the method of linear interpolation to supplement the missing values;

[0183] Data normalization: Scale the data to a specific range, such as [0,1], to eliminate the influence between different dimensions. The specific method is to use the min-max normalization method.

[0184] Through the above steps, ensure the integrity, consistency, and applicability of the data, laying a solid foundation for subsequent analysis and prediction. These preprocessing steps can significantly improve the performance of the model and the accuracy of prediction.

[0185] Step (2), use chaos theory to predict prices: Adopt the time series analysis method of chaos theory to analyze the historical data of power market prices and predict their future trends.

[0186] In the power market, price fluctuations usually have complex dynamic characteristics. Chaos theory provides an effective method to analyze and predict this complex time series behavior. Chaos theory is a theory used to study irregular and unpredictable behaviors in deterministic nonlinear dynamic systems. Although chaotic systems exhibit randomness in the short term, they have a certain degree of determinacy and predictability in the long term. Power market prices, as time series data, often exhibit chaotic characteristics.

[0187] The specific implementation process is as follows:

[0188] (2.1) Phase space reconstruction: Using Takens' theorem, one-dimensional time series data is embedded into a high-dimensional phase space through the method of delayed coordinates. This process can reveal the internal structure of the time series. The specific steps are as follows:

[0189] Select appropriate time delay τ and embedding dimension m;

[0190] Construct vectors in the phase space:

[0191] X(t) = [x(t), x(t + τ), x(t + 2τ),..., x(t + (m - 1)τ)]

[0192] where x(t) is the value of the time series at time t, which is the electricity price corresponding to time t; τ is the time delay, and m is the embedding dimension, both determined by the method of phase space reconstruction;

[0193] (2.2) Calculate the maximum Lyapunov exponent: The maximum Lyapunov exponent is used to measure the degree of chaos of the system. If the maximum Lyapunov exponent is positive, it indicates that the system has chaotic characteristics.

[0194] Here, the system refers to the dynamic change system of the electricity market price. By reconstructing the phase space and calculating the maximum Lyapunov exponent of the historical price data, its chaotic characteristics are analyzed; those skilled in the art should know that this system refers to the system in mathematics, not the system used for control and management in reality.

[0195] The calculation steps include:

[0196] Select two adjacent points X1 and X2 in the initial phase space and calculate their initial distance d(0);

[0197] As time evolves, calculate their distance d(t) at time t; The calculation formula for the maximum Lyapunov exponent λ is:

[0198]

[0199] where N is the number of time steps;

[0200] (2.3) Predict future prices: Use the evolution law of points in the phase space for short-term prediction. The commonly used method is the local nearest neighbor prediction method, that is, find the historical state closest to the current state in the phase space and predict the future value through its evolution trend.

[0201] First, find the historical state closest to the current state in the phase space, that is, the local nearest neighbor. Then, based on the evolution trend of this nearest neighbor point, predict the electricity price value at the next time step. The specific steps are as follows:

[0202] Calculate the position of the current state in the phase space.

[0203] Search for the N nearest neighboring points in the historical phase space to the current state.

[0204] Based on the future evolution trends of these neighboring points, use weighted average or other methods to predict the future price value. For example, use the future price average of these neighboring points as the future price prediction value of the current state.

[0205] Through the above steps, using chaos theory for time series analysis can effectively capture the dynamic characteristics of electricity market prices and provide a reliable prediction basis for formulating trading strategies.

[0206] Step (3), design an electricity trading strategy: According to the price prediction results, design a preliminary electricity trading strategy on how to allocate the electricity quantity for purchase and sale.

[0207] In the electricity market, designing an effective trading strategy is the key to maximizing profits and minimizing risks. Based on the price prediction results of chaos theory in step (2), the goal of step (3) is to formulate a reasonable electricity trading strategy to optimize the allocation of electricity purchase and sale. The following is the detailed implementation process and the involved mathematical formulas.

[0208] Trading Strategy Design

[0209] (3.1) Determine the trading objectives: The trading objectives include profit maximization and risk control; profit maximization is achieved by maximizing the net income, and risk control is achieved by setting a risk tolerance (such as a maximum loss limit or volatility limit) to control risks and balance the benefits and risks, that is, risk control means that the loss does not exceed the maximum loss limit, or the volatility does not exceed the volatility limit;

[0210] Specifically, the objective function for profit maximization is expressed as:

[0211]

[0212] P t sell : That is, the electricity selling price at time t, usually expressed in yuan per megawatt-hour (yuan / MWh).

[0213] The electricity quantity sold at time t. That is, the electricity quantity planned to be sold at time t, usually expressed in megawatt-hour (MWh).

[0214] P t buy The electricity purchase price at time t. That is, the electricity purchase price at time t, usually expressed in yuan per megawatt-hour (yuan / MWh).

[0215] The electricity purchase volume at time t. That is, the amount of electricity planned to be purchased at time t, usually expressed in megawatt-hours (MWh).

[0216] T is the total time;

[0217] (3.2) Prediction result analysis: Analyze the price prediction results obtained in step (2) to identify the peak and trough periods of prices; it is suitable to sell electricity during the peak period and purchase electricity during the trough period; for other periods, according to the overall market trend and supply and demand situation, a default strategy can be set, such as maintaining the original electricity volume or adjusting according to the risk tolerance;

[0218] (3.3) Establish a trading model: Based on the predicted price data, establish an electricity trading model. Common methods include dynamic programming and optimization models. The following is a simple example of an optimization model:

[0219] Objective function: Maximize the net profit

[0220]

[0221] P t sell is the electricity price for selling electricity at time t;

[0222] is the amount of electricity planned to be sold at time t;

[0223] P t buy is the electricity price for purchasing electricity at time t;

[0224] is the amount of electricity planned to be purchased at time t;

[0225] Constraints:

[0226] Power balance constraint: Ensure that within each time period, the electricity sold and the electricity purchased satisfy the power balance.

[0227]

[0228] where, is the available power at time t.

[0229] Risk control constraint: Limit the maximum trading volume within a certain time period to reduce the risk brought by market fluctuations.

[0230]

[0231] Step (3.4), Strategy Optimization: Solve the above model using linear programming to obtain a preliminary electricity trading strategy. Step (4) then optimizes the preliminary electricity trading strategy through the particle swarm optimization algorithm to further improve the effectiveness of the strategy.

[0232] Through the above steps, a set of electricity trading strategies based on price prediction can be designed to reasonably allocate the electricity purchase and sales volumes, thereby maximizing the revenue and minimizing the risks in an uncertain market environment.

[0233] Step (4), Optimize the trading strategy using the particle swarm optimization algorithm: Consider the trading strategy as the solution to the optimization problem and use the particle swarm optimization algorithm to search for the optimal solution in the strategy space, that is, to maximize the total revenue of electricity trading.

[0234] In the electricity market, optimizing the trading strategy to maximize the revenue is a complex task. The particle swarm optimization (PSO) algorithm is a powerful global optimization tool that can effectively search for the optimal solution in the strategy space. The following is the detailed implementation process of considering the electricity trading strategy as an optimization problem and using the PSO algorithm for optimization.

[0235] Based on the model established in step (3), model the electricity trading strategy as an optimization problem. The goal is to maximize the total revenue, defined as the revenue from selling electricity minus the cost of purchasing electricity. In this model, the strategy parameters include the electricity purchase volume and sales volume for each time period.

[0236] PSO is an optimization algorithm based on swarm intelligence that simulates the foraging behavior of bird flocks and optimizes the problem through information sharing among individuals. The following is the application process of the PSO algorithm in optimizing the electricity trading strategy:

[0237] (4.1) Initialization: Randomly initialize the particle swarm. Each particle represents a possible trading strategy, that is, a combination of the electricity purchase volume and sales volume; initialize the velocity and position of each particle; set the number of particles as N, and randomly allocate the purchase volume and sales volume of each particle within a predetermined range; initialize the velocity of each particle to zero;

[0238] (4.2) Evaluate the fitness: Calculate the fitness value of each particle, that is, calculate the total revenue of the current trading strategy according to the objective function as the fitness value;

[0239] The fitness function is the objective function, specifically:

[0240] (4.3) Update the velocity and position:

[0241] Update the particle velocity:

[0242]

[0243] where v i (t) is the velocity of particle i at time t; ω is the inertia weight (ranging from 0.4 to 0.9); c1 and c2 are acceleration constants (both with a value of 2.0); r1 and r2 are random numbers (random numbers between 0 and 1); is the historical best position of particle i, which is updated when its fitness is better than the previously recorded best fitness; g best is the global best position, which is the position of the particle with the best fitness among all current particles;

[0244] Update the particle position

[0245] x i (t + 1) = x i (t) + v i (t + 1)

[0246] where x i (t) is the position of particle i at time t.

[0247] (4.4) Update the best position: If the fitness of the current particle is better than its historical best position, then update its historical best position If the fitness of the current particle is better than the global best position, then update the global best position g best ;

[0248] (4.5) Iteration: Repeat steps (4.2) to (4.4) until a predetermined number of iterations (e.g., 1000 times) is reached or the fitness gain is lower than a set threshold (e.g., 0.001).

[0249] (4.6) Output the optimal strategy: After multiple iterations, output the trading strategy corresponding to the global best position, which is the optimal power purchase and sale allocation plan.

[0250] Through the PSO algorithm, it is possible to effectively search for the optimal solution in a complex strategy space, optimize the power trading strategy to maximize the profit, while satisfying the constraints of power balance and risk control.

[0251] Step (5), use the deep reinforcement learning algorithm to learn and adjust the trading strategy: Use the optimal strategy obtained by the particle swarm optimization algorithm as the initial solution to train the deep reinforcement learning model. The model continuously tries new trading strategies, adjusts the network weights according to the feedback, and finally obtains a more optimized strategy.

[0252] In the optimization of power trading strategies, Deep Reinforcement Learning (DRL) provides a powerful approach to learn the optimal strategy through interaction with the environment. Combining the optimal strategy obtained by the Particle Swarm Optimization (PSO) algorithm in step (4), the goal of step (5) is to further learn and adjust the trading strategy using the deep reinforcement learning algorithm to achieve higher returns. Specifically as follows:

[0253] (5.1) Initial solution setting: Before the start of deep reinforcement learning training, the optimal strategy obtained by the PSO algorithm is used as the initial solution. This strategy provides a good starting point for the DRL model, helping the model converge to an effective strategy faster. This combination can accelerate the learning process and improve the quality of the initial strategy.

[0254] (5.2) Deep reinforcement learning training:

[0255] (5.2.1) Environment modeling: Define the simulation environment of the power market, including the state space, action space, and reward function; the state space includes the current electricity price, historical price trend, and available power; the action space includes the electricity purchase volume and sales volume in each time period;

[0256] (5.2.2) Select model: Select a suitable deep reinforcement learning algorithm, such as Deep Q-Network (DQN). DQN is suitable for discrete action spaces.

[0257] (5.2.3) Training process:

[0258] Initialization: Initialize the weights of the deep Q-network, and use the optimal power trading strategy obtained by the particle swarm optimization algorithm as the initial solution;

[0259] Interaction: In the simulation environment, the model selects actions (electricity purchase volume and sales volume) according to the current strategy, applies the actions to the environment, and the environment feedbacks the next state (new electricity price, supply and demand situation) and immediate reward (net income or risk index);

[0260] Experience replay (for DQN): Store the results of each interaction in the experience replay buffer for subsequent random sampling for training.

[0261] Update strategy: Update the network weights using the gradient descent method; for the deep Q-network, the goal is to minimize the Bellman error:

[0262]

[0263] Among them, r is the immediate reward; γ is the discount factor; θ and θ′ are the parameters of the current and target networks respectively; s is the state. Specifically, s and s’ are the states of the current and target networks respectively; a is the action. Specifically, a and a’ are the actions of the current and target networks respectively;

[0264] The state s should comprehensively reflect the dynamic changes in the power trading environment; the action a should cover all possible trading decision options. When designing the reward function for the reward r, the benefits and risks should be considered comprehensively;

[0265] Policy update: Update the policy by optimizing the objective function (minimizing the Bellman error) to ensure that the difference between the new and old policies does not exceed a threshold for stable training:

[0266]

[0267] Among them, L PPO (θ) is the loss function of PPO; is the calculation of the expected value, obtained by sampling; r t (θ) is the probability ratio, defined as the probability ratio of the new policy to the old policy for the action a; is the advantage estimate, measuring how good the action a taken at time step t is compared to the average level; clip(r t (θ), 1 - ∈, 1 + ∈) is the clipping function, used to limit the amplitude of policy update, and ∈ is the clipping range parameter;

[0268] (5.2.4) Policy optimization: During the training process, the model continuously adjusts the policy to maximize the cumulative reward;

[0269]

[0270] Among them, G t is the cumulative reward at time t, r t+k is the immediate reward at time t + k, and γ is the discount factor;

[0271] (5.5) Evaluation and adjustment: Regularly evaluate the performance of the training policy in the simulation environment, and adjust the model parameters and training policy according to the evaluation results to ensure the robustness and stability of the model.

[0272] For example, regularly (every 1000 steps) evaluate the performance of the training policy in the simulation environment, and evaluate it through metrics such as cumulative reward and maximum drawdown. According to the evaluation results, adjust the model parameters (learning rate, discount factor) to improve the model performance.

[0273] The symbol a (action) represents the more optimized power trading strategy described in step (5). Through the deep Q-network, the agent can select the optimal action a in different states s, thereby realizing the optimization of the overall power trading strategy.

[0274] Through deep reinforcement learning, the model can autonomously learn and adjust trading strategies in the dynamic environment of the power market. Combining the initial solution provided by PSO, the DRL model can converge to a high-yield strategy faster, realizing the intelligence and optimization of power trading.

[0275] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A demand-side responsive power market transaction decision-making method, characterized in that: The steps include: Step (1), data collection and processing: collecting historical and real-time data of the electricity market and then preprocessing it; the data includes price, demand and supply, and weather data; Step (2), using chaos theory to predict prices: using the time series analysis method of chaos theory to analyze the historical data of electricity market prices, predict its future trend, and obtain price prediction results; Step (3), designing a power trading strategy: based on the price forecast results, designing a preliminary power trading strategy to allocate the amount of power purchased and sold; Step (4), using particle swarm optimization algorithm to optimize the trading strategy: the trading strategy designed in step (3) is regarded as the solution of the optimization problem, and the particle swarm optimization algorithm is used to find the optimal solution in the trading strategy space to obtain the optimal power trading strategy; Step (5), using deep reinforcement learning algorithm to learn and adjust trading strategies: the optimal electricity trading strategy obtained by the particle swarm optimization algorithm is used as the initial solution to train the deep reinforcement learning model; the deep reinforcement learning model continuously tries new electricity trading strategies, adjusts the network weights according to feedback, and finally obtains a more optimized electricity trading strategy.

2. The power market transaction decision-making method based on demand-side response according to claim 1, characterized in that: In step (1), preprocessing includes data cleaning, missing value processing, and normalization; Among them, when cleaning the data, outliers in the data are removed; When dealing with missing values, linear interpolation is used to fill in the missing values; Normalization uses the minimum-maximum normalization method.

3. The power market transaction decision-making method based on demand-side response according to claim 1 is characterized in that: The specific method of step (2) is: (2.1) Reconstructing the phase space: Using Takens’ theorem, the one-dimensional time series data is embedded into the high-dimensional phase space through the delayed coordinate method. The specific steps are as follows: Choose appropriate time delay τ and embedding dimension m; Construct a vector in phase space: X(t)=[x(t),x(t+τ),x(t+2τ),...,x(t+(m-1)τ)] Where x(t) is the value of the time series at time t, that is, the electricity price corresponding to time t; τ is the time delay, and m is the embedding dimension, both of which are determined by the phase space reconstruction method; (2.2) Calculate the maximum Lyapunov exponent: Select two adjacent points X1 and X2 in the initial phase space and calculate their initial distance d(0); Evolving over time, calculate their distance d(t) at time t; The maximum Lyapunov exponent λ is calculated as: Where N is the number of time steps; (2.3) Predicting future prices: Use the local nearest neighbor prediction method to predict future prices.

4. The power market transaction decision-making method of demand-side response according to claim 1 is characterized in that: The specific method of step (3) is: (3.1) Determine trading objectives: Trading objectives include profit maximization and risk control; profit maximization is achieved by maximizing net returns, and risk control is achieved by setting risk tolerance; the risk tolerance is the maximum loss limit or volatility limit; (3.2) Forecast result analysis: Analyze the price forecast results obtained in step (2) to identify the peak and trough periods of prices; the peak period is suitable for selling electricity, while the trough period is suitable for purchasing electricity; (3.3) Establishing a trading model: Establishing an electricity trading model based on the predicted price data; The electricity trading model aims to maximize net profit, and its objective function is: P t sell is the electricity price for electricity sales transactions at time t; is the planned amount of electricity sold at time t; P t buy is the electricity price for electricity purchase transaction at time t; is the planned purchase amount of electricity at time t; The constraints of the power trading model include: Power balance constraints: in, is the amount of electricity available at time t; Risk control constraints: Step (3.4), strategy optimization: Use linear programming to solve the above model and obtain a preliminary electricity trading strategy.

5. The power market transaction decision-making method of demand-side response according to claim 4 is characterized in that: The specific method of step (4) is: (4.1) Initialization: Randomly initialize the particle swarm, each particle represents a possible trading strategy, that is, a combination of electricity purchase and sales; initialize the speed and position of each particle; set the number of particles to N, and randomly distribute the purchase and sales of each particle within a predetermined range; initialize the speed of each particle to zero; (4.2) Evaluate fitness: Calculate the fitness value of each particle, that is, calculate the total profit of the current trading strategy as the fitness value according to the objective function; (4.3) Update speed and position: Update particle velocity: Among them, v i (t) is the velocity of particle i at time t, ω is the inertia weight, c1 and c2 are acceleration constants, r1 and r2 are random numbers, is the best historical position of particle i, g best is the global optimal position; Update particle position x i (t+1)=x i (t)+v i (t+1) Among them, x i (t) is the position of particle i at time t; (4.4) Update the best position: If the fitness of the current particle is better than its historical best position, update its historical best position; if the fitness of the current particle is better than the global best position, update the global best position; (4.5) Iteration: Repeat steps (4.2) to (4.4) until the predetermined number of iterations is reached or the fitness gain is lower than the set threshold; (4.6) Output the optimal strategy: After multiple iterations, the trading strategy corresponding to the global optimal position is output, which is the optimal electricity trading strategy.

6. The power market transaction decision-making method of demand-side response according to claim 1 is characterized in that: The specific method of step (5) is: (5.1) Initial solution setting: The optimal power trading strategy obtained by the particle swarm optimization algorithm is used as the initial solution; (5.2) Deep reinforcement learning training: (5.2.1) Environment modeling: Define the simulation environment of the electricity market, including state space, action space, and reward function; the state space includes current electricity prices, historical price trends, and available electricity; The action space includes the amount of electricity purchased and sold in each time period; (5.2.2) Model selection: Deep reinforcement learning model selection deep Q network; (5.2.3) Training process: Initialization: Initialize the weights of the deep Q network and use the optimal power trading strategy obtained by the particle swarm optimization algorithm as the initial solution; Interaction: In a simulated environment, the model selects actions based on the current strategy, applies the actions to the environment, and the environment feeds back the next state and immediate rewards; Experience replay: Store the results of each interaction in the experience replay buffer; Update strategy: Use gradient descent method to update network weights; For the deep Q network, the goal is to minimize the Bellman error, and the objective function is: Where r is the immediate reward; γ is the discount factor; θ and θ′ are the parameters of the current and target networks, respectively; s is the state; a is the action; Policy update: Update the policy by optimizing the objective function to ensure that the difference between the new and old policies does not exceed a threshold to stabilize training: (5.2.4) Strategy optimization: During the training process, the model continuously adjusts its strategy to maximize the cumulative reward; Among them, G t is the cumulative reward at time t, r t+k is the immediate reward at time t+k, γ is the discount factor; (5.2.5) Evaluation and Adjustment: Regularly evaluate the performance of the training strategy in the simulation environment and adjust the model parameters and training strategy based on the evaluation results.

7. The power market transaction decision-making method of demand-side response according to claim 5 is characterized in that: In step (4.3), ω is the inertia weight and its value range is [0.4, 0.9].

8. The power market transaction decision-making method of demand-side response according to claim 5 is characterized in that: In step (4.3), c1 and c2 are both 2.0; r1 and r2 are both random numbers between 0 and 1.

9. The power market transaction decision-making method of demand-side response according to claim 5 is characterized in that: In step (4.5), the number of iterations is 1000.

Citation Information

Cited By

  • Electric power transaction strategy generation method and system based on Monte Carlo simulation and deep learning

    CN120707198A

  • Electric power market transaction optimization system and method based on hybrid prediction

    CN120894076A