Intelligent agent quotation method, device and equipment in out-of-area power purchase auxiliary decision-making and medium

By applying the actor critic algorithm based on recurrent neural network improvement and extreme gradient enhancement model in the out-of-regional power purchase decisions, the problem of lack of targeted and adaptable power purchase strategies in the existing technology is solved, and more accurate transaction prediction and strategy optimization are achieved.

CN120013589AInactive Publication Date: 2025-05-16STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510183217.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is difficult to comprehensively and in-depth analysis of influencing factors in the decision-making of out-of-regional power purchases, resulting in a lack of targeted and adaptable power purchase strategies and failure to effectively simulate the game behavior and competitive relationship between the supply and demand parties of the transaction.

Method used

Agent modeling is performed using the actor critic algorithm (RNN-Actor-Critic) based on recurrent neural network improvement, simulates trading scenarios, predicts transaction prices through extreme gradient boosting (Xgboost) model, and optimizes the declaration price strategy.

Benefits of technology

By capturing the long-term dependence in time series data and optimizing quotation strategies, the targetedness and adaptability of power purchase decisions are improved, and the simulation and prediction capabilities of trading scenarios are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013589A_ABST
    Figure CN120013589A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power system power purchase decision making, in particular to an agent quotation method and device in an out-of-area power purchase auxiliary decision making and a medium, and the method comprises the steps: setting a transaction scene and a declaration strategy, and determining a sending end and receiving end agent modeling range and an output data item; acquiring historical data according to the modeling range and the output data item, and preprocessing the historical data; performing agent modeling by using an actor commentator algorithm RNN-Actor-Critic improved on the basis of a recurrent neural network, and simulating a transaction by using an agent on the basis of a transaction scene; according to the change trend of agent boundary data, prediction boundary data of a transaction target time period is calculated, then an extreme gradient is utilized to promote an Xgboost model, and the predicted boundary data is input to calculate an expected transaction price; and calculating the declaration price based on the expected transaction price, and finally outputting the predicted transaction price, the predicted boundary data, the prediction model and the predicted declaration price.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power system power purchase decision-making, and in particular to an intelligent body quotation method, device equipment and medium in out-of-region power purchase auxiliary decision-making. Background Art

[0002] As an important way to ensure the supply of regional and provincial power grids, out-of-region power purchase has always occupied an important position in the power market in recent years. However, the current formulation of out-of-region power purchase strategies still faces many challenges. First, there are many provinces and trading entities at the sending and receiving ends, and their supply and demand situation and power supply characteristics are complex and changeable. The traditional manual experience method is difficult to comprehensively and deeply analyze these influencing factors, resulting in the lack of pertinence and adaptability in the formulation of power purchase strategies. Secondly, with the gradual opening of the power market and the gradual formation of a unified national power market system, out-of-region power purchases are transitioning from traditional bilateral negotiations to market-oriented methods, which requires the formulation of power purchase strategies to be more flexible and efficient to meet the needs of market-oriented transactions.

[0003] The existing implementation schemes include systems that improve trading decision support capabilities by integrating multiple data sources and applying basic forecasting models to forecast load changes, transaction prices, and inter-provincial market supply and demand trends. These systems may provide basic data management functions and can perform simple intelligent forecasts of supply and demand from the perspective of a single trading entity, but they do not simulate the game behavior and competitive relationships in actual trading scenarios, and do not consider the different influencing factors of both sides of the transaction supply and demand, as well as the impact of the existence of trading competitors on the transaction.

[0004] The information disclosed in this background technology section is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as acknowledging or suggesting in any form that the information constitutes the prior art already known to those skilled in the art. Summary of the invention

[0005] The present invention provides a method, device and medium for intelligent body quotation in auxiliary decision-making for out-of-region electricity purchase, thereby effectively solving the problems in the background technology.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is: a method for intelligent agent quotation in auxiliary decision-making for out-of-region electricity purchase, comprising the following steps:

[0007] Set transaction scenarios and declaration strategies, determine the modeling scope and output data items of the sending and receiving agents;

[0008] Acquire historical data according to the modeling scope and output data items, and preprocess the historical data;

[0009] Using the improved actor-critic algorithm RNN-Actor-Critic based on recurrent neural network to model the intelligent agent, and based on the transaction scenario, using the intelligent agent to simulate the transaction;

[0010] The predicted boundary data of the target trading period is deduced through the change trend of the boundary data of the intelligent agent, and then the extreme gradient boosting Xgboost model is used to input the predicted boundary data to calculate the expected transaction price;

[0011] The current declared price is calculated based on the expected transaction price, and the predicted transaction price, predicted boundary data, prediction model and predicted declared price are finally output.

[0012] Furthermore, the transaction scenario includes: transaction type, transaction method, forecast time range, risk preference coefficient, randomness coefficient, and strategy goal;

[0013] The declaration strategy includes: modeling scope of sending and receiving intelligent agents, and output data items.

[0014] Furthermore, the preprocessing of the historical data comprises the following steps:

[0015] Identify and process errors, missing values, and outliers in the data; use the linear interpolation method to fill in missing values, that is, use the data outside the two ends of the missing period to perform linear regression modeling, predict missing values, and fill them in;

[0016] Data normalization and standardization, scaling the data to a set range;

[0017] Perform correlation analysis on the input data, select important features or create new features until the data meets the requirements for model training.

[0018] Furthermore, the predicting missing values ​​and filling them in includes:

[0019] For entities that lack historical data, the average historical data of other entities in the same province will be used as training samples for the entities involved in this transaction;

[0020] If the historical data of other subjects in the province to which the subject belongs is missing, the mean historical data of subjects in provinces across the country is used as the training sample.

[0021] Furthermore, the agent modeling using the improved actor-critic algorithm RNN-Actor-Critic based on recurrent neural network includes the following steps:

[0022] When the agent receives the current state s t After that, the actor network outputs a probability distribution or action strategy based on the state, from which an action a is selected.t implement;

[0023] The environment is based on the action a performed t Returns an instant reward r t and the next state s t+1 At the same time, the recurrent neural network RNN ​​in the critic network and the feedforward neural network learn simultaneously, taking the current state s t and action a t As input, estimate the value function under the current strategy, and add the value functions obtained by the two neural networks to obtain the final true value function;

[0024] The accuracy of the value function estimate is measured by calculating the TD error;

[0025] During the update process, high-value samples are extracted from the experience replay pool and provided to the Actor network and the Critic network for training; this alternating update process is repeated continuously, so that the Actor gradually learns the optimized action selection strategy, and the Critic network gradually approaches the true value function estimate, achieving strategy optimization and modeling of environmental dynamics.

[0026] Furthermore, the using the intelligent agent to simulate transactions includes:

[0027] During the simulated transaction process of the actor network, the generated predicted transaction price, predicted boundary data, predicted model and predicted declared price are added as samples to the experience replay pool of the improved algorithm. When the experience replay pool reaches the storage limit, the dual critic Critic network starts working.

[0028] In each transaction, the priority replay cache mechanism in the improved algorithm is used to extract high-value sample data from the experience replay pool for the training of the actor recurrent neural network.

[0029] When the sample data in the experience replay pool reaches the upper limit, the dual critic Critic network inside the agent starts learning, and the Actor network starts receiving the strategy value from the dual critic network;

[0030] In each iteration, the dual critic network also uses the priority replay cache mechanism to extract high-value sample data from the experience replay pool for learning;

[0031] During the prediction process, the dual critic network obtains the real boundary data, transaction results, and output of the actor network in the corresponding simulated transaction scenario. The predicted transaction price is obtained by using the prediction output of the actor network when the real boundary data is obtained. The two networks calculate the comprehensive deviation rate of the transaction price predicted by the actor network and the critic network respectively, and calculate the comprehensive deviation rate between the transaction price under the real boundary and the transaction price predicted by the critic network.

[0032] Based on the comprehensive deviation rate, the Critic network calculates the evaluation value Reward for the Actor network's predicted performance;

[0033] For the buyer, if the declared price is higher than the actual transaction price, the strategy coefficient will be lowered; after updating the strategy, the agent will end this round of simulation trading and output the latest strategy coefficient K t轮更新策略 , enter the next simulation transaction.

[0034] Furthermore, the calculation of the comprehensive deviation rate between the transaction price under the actual boundary and the transaction price predicted by the Critic network includes:

[0035] Predicted transaction price under the real boundary The predicted transaction price under the predicted boundary Deviation MAPE f The calculation method is as follows:

[0036]

[0037] Predicted transaction price under the real boundary and the actual transaction price P 实际成交 The deviation rate MAPE p The calculation method is as follows:

[0038]

[0039] Furthermore, based on the comprehensive deviation rate, the Critic network calculates a reward for the predicted performance of the Actor network, including:

[0040] Reward t =α*MAPE f +β*MAPE p ;

[0041] In the formula, α and β are the update step sizes. The larger the data prediction deviation, the larger the strategy coefficient needs to be modified.

[0042] The agent uses the change in the reward of the past three months to calculate the reward of round t tAfter the correction is completed, the calculation method is as follows:

[0043] Reward t =δΔ t-1 +δ 2 Δ t-2 +δ 3 Δ t-3 ;

[0044] In the formula, Δ i-1 is the change in the Reward value from the i-1th month to the tth month, that is, Reward i -Reward i-1 ; δ t-i is the weight coefficient. The weight decays exponentially with the lag time, that is, the older the change, the lower the contribution. Assuming that round t is the first three rounds of simulation, then Δ t-i All are supplemented with 0 values.

[0045] Furthermore, the priority playback cache mechanism includes:

[0046] The experience tuple (s) generated when the agent interacts with the environment t ,a t ,r t ,s t+1 ) is stored in a fixed-size buffer and a mini-batch of data is randomly sampled from it for learning during each training;

[0047] During training, samples are extracted from the buffer, the absolute value of the TD error is used as the priority of the sample, and the priority p of the experience tuple is calculated. t =|δ t |+∈, where ∈ is a small constant to prevent the priority from being zero when the TD error is zero;

[0048] Normalize the priorities of all samples so that the priority of each sample conforms to the probability distribution:

[0049]

[0050] Where α controls the degree of influence of priority on sampling.

[0051] Furthermore, the predicted boundary data of the target transaction period is calculated through the change trend of the boundary data of the intelligent agent, and then the extreme gradient boosting Xgboost model is used to input the predicted boundary data to calculate the expected transaction price, including the following steps:

[0052] Use the regression model or weighted moving average model to deduce the predicted boundary data for the target trading period, and then use the Xgboost model to input the predicted boundary data to calculate the expected transaction price;

[0053] By superimposing the declaration strategy, the declaration price is calculated based on the expected transaction price, and the predicted transaction price, predicted boundary data, prediction model and predicted declaration price are finally output;

[0054] The calculation method of the declaration strategy is as follows:

[0055] P t轮申报 =P t轮Actor预期成交 *Risk*K t轮 ;

[0056] Risk represents risk preference, which is three different preference types of aggressive, conservative and conventional that are manually input, and the coefficients are randomly selected in different intervals; K is the strategy coefficient itself.

[0057] Furthermore, it also includes: when the transaction method is centralized bidding, the influencing factors of the historical inter-provincial spot prices of the sending end and other receiving provinces are added to consider, and finally the predicted declared price of the sending end after the strategy coefficient is updated, and the predicted declared price of the province and other receiving entities is output.

[0058] Furthermore, the predicted bid price of the sending end after the strategy coefficient is finally output, and the predicted bid price of the province and other receiving entities is predicted based on the inter-provincial spot transaction volume of each sending entity, the proportion of power source types and installed capacity of each sending province, and the factors of changes in the supply and demand situation of each sending province to obtain the centralized bidding bid volume of each sending transaction entity;

[0059] The receiving-end centralized bidding power declaration is obtained by combining the supply and demand patterns, economic indexes, meteorological conditions, changing trends in installed power generation capacity and the province's forecast of centralized bidding power declaration.

[0060] Furthermore, the centralized bidding power reported by each sending-end trading entity is obtained by combining the factors of supply and demand situation changes in each sending-end province, including:

[0061] The sending agent quotes:

[0062] Step 1: Combine the power type of the sending end and the marginal cost of this type of power to give the sending end quotation P1;

[0063] Step 2: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the sending end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price forecast model, the inter-provincial spot prices and bilateral negotiated prices within the trading window are forecasted to obtain the corresponding day-ahead spot clearing price P. d and bilateral negotiation forecast price P l ;

[0064] Step 3: Combine P d , P l, according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean;

[0065] Step 4: Assign different weights n1 and m2 to the quotations P1 and P2 generated in Step 1 and Step 3 respectively, and generate the final quotation P3;

[0066] Step 5: Output P1, P2 or P3 according to user selection;

[0067] The sending end intelligent agent reports:

[0068] The sending-end intelligent agent's reported quantity data is fitted and allocated through the multivariate linear regression method combined with the artificial rule method to the intelligent agent's historical inter-provincial spot transaction quantity. Combined with the decision maker's risk preference setting, the corresponding intelligent agent's reported quantity strategy is output.

[0069] Furthermore, the combination of supply and demand forms, economic indexes, meteorological conditions, the changing trend of power installed capacity and the forecast of the centralized bidding declared power in the province to obtain the receiving end centralized bidding declared power includes:

[0070] Receiving agent quote:

[0071] Step 1: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the receiving end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price model, the inter-provincial spot prices and bilateral negotiated prices within the transaction window are predicted to obtain the corresponding predicted spot price P d and bilaterally negotiated price P l ;

[0072] Step 2: Combine P d , P l , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean.

[0073] Receiver intelligent body report quantity:

[0074] The sending-end intelligent agent's reported quantity data is fitted and allocated through the multivariate linear regression method combined with the artificial rule method to the intelligent agent's historical inter-provincial spot transaction quantity. Combined with the decision maker's risk preference setting, the corresponding intelligent agent's reported quantity strategy is output.

[0075] The present invention also includes an intelligent quotation device for auxiliary decision-making of out-of-region electricity purchase, using the above method, the device includes:

[0076] The setting unit is used to set the transaction scenario and declaration strategy, determine the modeling scope and output data items of the sending and receiving agents;

[0077] A collection unit, used for acquiring historical data according to the modeling scope and output data items, and preprocessing the historical data;

[0078] An agent modeling unit, used for performing agent modeling using an improved actor-critic algorithm RNN-Actor-Critic based on a recurrent neural network, and simulating transactions using the agent based on the transaction scenario;

[0079] A prediction unit, used to calculate the predicted boundary data of the target transaction period according to the change trend of the boundary data of the intelligent agent, and then use the extreme gradient boosting Xgboost model to input the predicted boundary data to calculate the expected transaction price;

[0080] The output unit is used to calculate the current declared price based on the expected transaction price, and finally output the predicted transaction price, predicted boundary data, prediction model and predicted declared price.

[0081] The present invention also includes a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described above is implemented.

[0082] The present invention also includes a storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented.

[0083] The beneficial effects of the present invention are as follows: by introducing an improved Actor-Critic algorithm based on RNN, it is used for learning and predicting bilateral negotiation quotation agents and centralized bidding quotation agents. The algorithm captures long-term dependencies in time series data through a recurrent neural network RNN, simulates game behaviors and competitive relationships in actual trading scenarios, and improves the stability and accuracy of value function estimation through a Critic network, thereby optimizing the agent's quotation strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0085] Figure 1 is a flow chart of the method in Example 1;

[0086] Figure 2 It is a structural schematic diagram of the device in Example 1;

[0087] Figure 3 The intelligent agent architecture in Example 2;

[0088] Figure 4 It is a framework diagram of the RNN-Actor-Critic algorithm in Example 2;

[0089] Figure 5 The centralized bidding agent structure in Example 2;

[0090] Figure 6 It is a schematic diagram of the structure of the computer device of the present invention. DETAILED DESCRIPTION

[0091] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0092] Embodiment 1:

[0093] like Figure 1 As shown: A method for intelligent agent quotation in auxiliary decision-making for out-of-region electricity purchase, comprising the following steps:

[0094] Set transaction scenarios and declaration strategies, determine the modeling scope and output data items of the sending and receiving agents;

[0095] Obtain historical data according to the modeling scope and output data items, and pre-process the historical data;

[0096] Use the improved actor-critic algorithm RNN-Actor-Critic based on recurrent neural network for agent modeling, and use agents to simulate transactions based on trading scenarios;

[0097] The predicted boundary data of the target trading period is deduced through the changing trend of the agent boundary data, and then the extreme gradient boosting Xgboost model is used to input the predicted boundary data to calculate the expected transaction price;

[0098] The declared price is calculated based on the expected transaction price, and the predicted transaction price, predicted boundary data, prediction model and predicted declared price are finally output.

[0099] By introducing an improved Actor-Critic algorithm based on RNN, it is used for the learning and prediction of bilateral negotiation quotation agents and centralized bidding quotation agents. The algorithm captures the long-term dependencies in time series data through the recurrent neural network RNN, simulates the game behavior and competitive relationships in actual trading scenarios, and improves the stability and accuracy of value function estimation through the Critic network, thereby optimizing the agent's quotation strategy.

[0100] In this embodiment, the transaction scenario includes: transaction type, transaction method, forecast time range, risk preference coefficient, randomness coefficient, and strategy target;

[0101] The declaration strategy includes: modeling scope of sending and receiving intelligent agents, and output data items.

[0102] The preprocessing of historical data includes the following steps:

[0103] Identify and process errors, missing values, and outliers in the data; use the linear interpolation method to fill in missing values, that is, use the data outside the two ends of the missing period to perform linear regression modeling, predict missing values, and fill them in;

[0104] Data normalization and standardization, scaling the data to a set range;

[0105] Perform correlation analysis on the input data, select important features or create new features until the data meets the requirements for model training.

[0106] As a preferred embodiment of the above, predicting missing values ​​and filling in missing values ​​includes:

[0107] For entities that lack historical data, the average historical data of other entities in the same province will be used as training samples for the entities involved in this transaction;

[0108] If the historical data of other subjects in the province to which the subject belongs is missing, the mean historical data of subjects in provinces across the country is used as the training sample.

[0109] In this embodiment, the agent modeling is performed using the RNN-Actor-Critic algorithm based on the improved recurrent neural network, which includes the following steps:

[0110] When the agent receives the current state s t After that, the actor network outputs a probability distribution or action strategy based on the state, from which an action a is selected. t implement;

[0111] The environment is based on the action a performed t Returns an instant reward r t and the next state st+1 At the same time, the recurrent neural network RNN ​​in the critic network and the feedforward neural network learn simultaneously, taking the current state s t and action a t As input, estimate the value function under the current strategy, and add the value functions obtained by the two neural networks to obtain the final true value function;

[0112] The accuracy of the value function estimate is measured by calculating the TD error;

[0113] During the update process, high-value samples are extracted from the experience replay pool and provided to the Actor network and the Critic network for training; this alternating update process is repeated continuously, so that the Actor gradually learns the optimized action selection strategy, and the Critic network gradually approaches the true value function estimate, achieving strategy optimization and modeling of environmental dynamics.

[0114] Use agents to simulate trading, including:

[0115] During the simulated transaction process of the actor network, the generated predicted transaction price, predicted boundary data, predicted model and predicted declared price are added as samples to the experience replay pool of the improved algorithm. When the experience replay pool reaches the storage limit, the dual critic Critic network starts working.

[0116] In each transaction, the priority replay cache mechanism in the improved algorithm is used to extract high-value sample data from the experience replay pool for the training of the actor recurrent neural network.

[0117] When the sample data in the experience replay pool reaches the upper limit, the dual critic Critic network inside the agent starts learning, and the Actor network starts receiving the strategy value from the dual critic network;

[0118] In each iteration, the dual critic network also uses the priority replay cache mechanism to extract high-value sample data from the experience replay pool for learning;

[0119] During the prediction process, the dual critic network obtains the real boundary data, transaction results, and output of the actor network in the corresponding simulated transaction scenario. The predicted transaction price is obtained by using the prediction output of the actor network when the real boundary data is obtained. The two networks calculate the comprehensive deviation rate of the transaction price predicted by the actor network and the critic network respectively, and calculate the comprehensive deviation rate between the transaction price under the real boundary and the transaction price predicted by the critic network.

[0120] Based on the comprehensive deviation rate, the Critic network calculates the reward for the Actor network's predicted performance;

[0121] For the buyer, if the declared price is higher than the actual transaction price, the strategy coefficient will be lowered; after updating the strategy, the agent will end this round of simulation trading and output the latest strategy coefficient K t轮更新策略 , enter the next simulation transaction.

[0122] Among them, the comprehensive deviation rate between the transaction price under the actual boundary and the transaction price predicted by the Critic network is calculated, including:

[0123] Predicted transaction price under the real boundary The predicted transaction price under the predicted boundary Deviation MAPE f The calculation method is as follows:

[0124]

[0125] Predicted transaction price under the real boundary and the actual transaction price P 实际成交 The deviation rate MAPE p The calculation method is as follows:

[0126]

[0127] In this embodiment, based on the comprehensive deviation rate, the Critic network calculates the evaluation value Reward for the Actor network's predicted performance, including:

[0128] Reward t =α*MAPE f +β*MAPE p ;

[0129] In the formula, α and β are the update step sizes. The larger the data prediction deviation, the larger the strategy coefficient needs to be modified.

[0130] The agent uses the change in the reward of the past three months to calculate the reward of round t t After the correction is completed, the calculation method is as follows:

[0131] Reward t =δΔ t-1 +δ 2 Δ t-2 +δ 3 Δ t-3 ;

[0132] In the formula, Δ i-1is the change in the Reward value from the i-1th month to the tth month, that is, Reward i -Reward i-1 ; δ t-i is the weight coefficient. The weight decays exponentially with the lag time, that is, the older the change, the lower the contribution. Assuming that round t is the first three rounds of simulation, then Δ t-Bi All are supplemented with 0 values.

[0133] The priority playback cache mechanism includes:

[0134] The experience tuple (s) generated when the agent interacts with the environment t ,a t ,r t ,s t+1 ) is stored in a fixed-size buffer and a mini-batch of data is randomly sampled from it for learning during each training;

[0135] During training, samples are extracted from the buffer, the absolute value of the TD error is used as the priority of the sample, and the priority p of the experience tuple is calculated. t =|δ t |+∈, where ∈ is a small constant to prevent the priority from being zero when the TD error is zero;

[0136] Normalize the priorities of all samples so that the priority of each sample conforms to the probability distribution:

[0137]

[0138] Where α controls the degree of influence of priority on sampling.

[0139] As a preferred embodiment of the above, the predicted boundary data of the target transaction period is calculated through the change trend of the agent boundary data, and then the extreme gradient boosting Xgboost model is used to input the predicted boundary data to calculate the expected transaction price, including the following steps:

[0140] Use the regression model or weighted moving average model to deduce the predicted boundary data for the target trading period, and then use the Xgboost model to input the predicted boundary data to calculate the expected transaction price;

[0141] By superimposing the declaration strategy, the declaration price is calculated based on the expected transaction price, and the predicted transaction price, predicted boundary data, prediction model and predicted declaration price are finally output;

[0142] The calculation method of the declaration strategy is as follows:

[0143] P t轮申报 =P t轮Actor预期成交 *Risk*K t轮 ;

[0144] Risk represents risk preference, which is three different preference types of aggressive, conservative and conventional that are manually input, and the coefficients are randomly selected in different intervals; K is the strategy coefficient itself.

[0145] In this embodiment, it also includes: when the transaction method is centralized bidding, the influencing factors of the historical inter-provincial spot prices of the sending end and other receiving provinces are added to consider, and finally the predicted declared price of the sending end after the strategy coefficient is updated, and the predicted declared price of the province and other receiving entities are output.

[0146] Among them, the final output is the predicted bid price of the sending end after the strategy coefficient is updated, and the predicted bid price of the province and other receiving entities is based on the inter-provincial spot transaction volume of each sending entity, the proportion of power source types and installed capacity of each sending province, combined with the factors of changes in the supply and demand situation of each sending province, to predict the centralized bidding bid volume of each sending transaction entity;

[0147] The receiving-end centralized bidding power declaration is obtained by combining the supply and demand patterns, economic indexes, meteorological conditions, changing trends in installed power generation capacity and the province's forecast of centralized bidding power declaration.

[0148] Combined with the factors of supply and demand changes in each sending province, the centralized bidding power reported by each sending transaction entity is predicted, including:

[0149] The sending agent quotes:

[0150] Step 1: Combine the power type of the sending end and the marginal cost of this type of power to give the sending end quotation P1;

[0151] Step 2: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the sending end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price forecast model, the inter-provincial spot prices and bilateral negotiated prices within the trading window are forecasted to obtain the corresponding day-ahead spot clearing price P. d and bilateral negotiation forecast price P l ;

[0152] Step 3: Combine P d , P l , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean;

[0153] Step 4: Assign different weights m1 and m2 to the quotations P1 and P2 generated in Step 1 and Step 3 respectively, and generate the final quotation P3;

[0154] Step 5: Output P1, P2 or P3 according to user selection;

[0155] The sending end intelligent agent reports:

[0156] The sending-end intelligent agent's reported quantity data is fitted and allocated through the multivariate linear regression method combined with the artificial rule method to the intelligent agent's historical inter-provincial spot transaction quantity. Combined with the decision maker's risk preference setting, the corresponding intelligent agent's reported quantity strategy is output.

[0157] As a preferred embodiment of the above, the receiving-end centralized bidding declared electricity quantity is obtained by combining the supply and demand form, economic index, meteorological conditions, the change trend of power installed capacity and the forecast of the centralized bidding declared electricity quantity of the province, including:

[0158] Receiving agent quote:

[0159] Step 1: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the receiving end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price model, the inter-provincial spot prices and bilateral negotiated prices within the transaction window are predicted to obtain the corresponding predicted spot price P d and bilaterally negotiated price P l ;

[0160] Step 2: Combine P d , P l , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean.

[0161] Receiver intelligent body report quantity:

[0162] The sending-end intelligent agent's reported quantity data is fitted and allocated through the multivariate linear regression method combined with the artificial rule method to the intelligent agent's historical inter-provincial spot transaction quantity. Combined with the decision maker's risk preference setting, the corresponding intelligent agent's reported quantity strategy is output.

[0163] like Figure 2 As shown, this embodiment also includes an intelligent quotation device for auxiliary decision-making in out-of-region electricity purchase, using the above method, the device includes:

[0164] The setting unit is used to set the transaction scenario and declaration strategy, determine the modeling scope and output data items of the sending and receiving agents;

[0165] A collection unit, used for acquiring historical data according to the modeling scope and output data items, and preprocessing the historical data;

[0166] The agent modeling unit is used to perform agent modeling using the RNN-Actor-Critic algorithm based on the improved recurrent neural network, and use agents to simulate transactions based on trading scenarios;

[0167] The prediction unit is used to calculate the predicted boundary data of the target trading period through the change trend of the agent boundary data, and then use the extreme gradient boosting Xgboost model to input the predicted boundary data to calculate the expected transaction price;

[0168] The output unit is used to calculate the current declared price based on the expected transaction price, and finally output the predicted transaction price, predicted boundary data, prediction model and predicted declared price.

[0169] Embodiment 2:

[0170] By analyzing historical data and combining the business scenarios of purchased electricity in Jiangsu Province, we abstracted the economic characteristics and behavioral characteristics of the main business users, combined with the historical inter-provincial medium- and long-term transaction prices of the province, the historical inter-provincial medium- and long-term transaction prices of the sending end entities, the historical inter-provincial spot prices of the sending end entities, meteorological data, energy price data and other related factors, and built an artificial trading agent. Through the improved Actor-Critic algorithm, we predicted the transaction volume and price of the sending end and the declared price of the receiving end. The specific process is as follows:

[0171] (1) Model configuration

[0172] According to the simulation plan, set the transaction type, transaction method, forecast time range, risk preference coefficient, randomness coefficient, strategy target and other parameters for this modeling, and then determine the modeling scope and output data items of the sending and receiving agents according to the simulation plan.

[0173] (2) Data acquisition and processing

[0174] Read relevant archival information from the database, including the sending and receiving transaction unit ID and name, power generation type, region, province, rated capacity, power type, etc., as well as inter-provincial power grid topology information such as the interconnection line ID and name, line starting province, line terminal province, regional power grid, voltage level, line loss rate, transmission and distribution price and other data, obtain historical transaction data, historical and predicted meteorological data, power supply and demand ratio, power installed capacity, power marginal cost, power generation and consumption, energy price index, economic index, etc.

[0175] The following processing is performed on the input data: (1) Data cleaning: Identify and process errors, missing values, outliers, etc. in the data; for missing values, the algorithm uses the linear interpolation method to fill in the missing values, that is, use the data outside the two ends of the missing period to perform linear regression modeling, predict the missing values ​​and fill in the missing values; for the lack of historical data, the entity participating in this transaction will refer to the mean of the historical data of other entities in the same province as the training sample; if there is no historical data of other entities in the province to which the entity belongs, then the algorithm will refer to the mean of the historical data of entities in all provinces across the country as the training sample. (2) Data standardization: Including data normalization and standardization, scaling the data to a specific range (usually 0 to 1 or with unit variance and zero mean). (3) Feature engineering: Perform correlation analysis on the input data, select important features or create new features until the data meets the requirements of model training.

[0176] (3) Agent learning and prediction

[0177] The agent is designed around the improved Actor-Critic algorithm framework and combined with human business experience. Based on the transaction scenario of this simulation solution, the Actor network inside the agent will participate in a series of continuous simulated transactions. First, through the trend of boundary data, the predicted boundary data of the transaction target period is inferred using a regression model or a weighted moving average model. Then, using the Xgboost model, the predicted boundary data is input to calculate the expected transaction price. By superimposing the declaration strategy, the current declaration price is calculated based on the expected transaction price, and finally the predicted transaction price, predicted boundary data, prediction model and predicted declaration price are output.

[0178] The calculation method of the declaration strategy is as follows:

[0179] P t轮申报 =P t轮Actor预期成交 *Risk*K t轮

[0180] Risk represents risk preference, which will randomly select coefficients within different intervals based on the three different preference types of aggressive, conservative and conventional input by humans; K is the strategy coefficient itself.

[0181] In the process of simulated trading by the Actor network, the generated predicted transaction price, predicted boundary data, prediction model and predicted declared price are added as samples to the experience replay pool of the improved algorithm. When the experience replay pool reaches the storage limit, the dual critic network starts working. In the process of each transaction, the priority replay cache mechanism in the improved algorithm is used to extract high-value sample data from the experience replay pool for training the Actor recurrent neural network. Due to the introduction of the recurrent neural network, the Actor network can record simulated transaction data for a longer period of time.

[0182] When the sample data in the experience replay pool reaches the upper limit, the dual critic network inside the agent starts learning, and the actor network starts receiving the strategy value from the dual critic network. In each iteration, the dual critic network also uses the priority replay cache mechanism to extract high-value sample data from the experience replay pool for learning. In the prediction process, the dual critic network can know the real boundary data, transaction results, and output of the actor network in the corresponding simulated transaction scenario, and use the Xgboost model output by the actor network to predict the predicted transaction price obtained when the real boundary data is obtained. The two networks calculate the comprehensive deviation rate of the transaction price predicted by the actor network and the critic network respectively, and calculate the comprehensive deviation rate between the transaction price under the real boundary and the transaction price predicted by the critic network.

[0183] Predicted transaction price under the real boundary The predicted transaction price under the predicted boundary Deviation MAPE f The calculation method is as follows:

[0184]

[0185] Predicted transaction price under the real boundary and the actual transaction price P 实际成交 The deviation rate MAPE p The calculation method is as follows:

[0186]

[0187] In the early stages of agent learning and prediction, the deviation rate calculation of the dual critic network gives the feedforward neural network a higher weight, focusing on feedback on the recent learning status of the actor network. As the training time increases, the weight of the recurrent neural network increases, allowing the agent to obtain longer training information. Based on these two deviation rates, the critic network calculates the reward for the actor network's prediction performance:

[0188] Reward t =α*MAPE f +β*MAPE p

[0189] Among them, α and β are update steps, which are set to 0.01 by default. The larger the data prediction deviation, the larger the change in the strategy coefficient. Finally, the agent will use the change in the reward of the past three months to adjust the reward of round t. t After the correction is completed, the calculation method is as follows:

[0190] Reward t =δΔ t-1 +δ 2 Δ t-2 +δ 3 Δ t-3

[0191] Δ i-1 is the change in the Reward value from the i-1th month to the tth month, that is, Reward i -Reward i-1 δ t-i is the weight coefficient, the default value is 0.8, and the weight decays exponentially with the lag time, that is, the older the change, the lower the contribution. Assuming that round t is the first three rounds of simulation, then Δ t-i For the buyer, if the declared price is higher than the actual transaction price, it means that there is still room for the bid to be compressed, so the next bid can be relatively lower, that is, the strategy coefficient is adjusted lower. After updating the strategy, the agent will end this round of simulation trading and output the latest strategy coefficient K t轮更新策略 , enter the next simulation transaction.

[0192] Bilateral Negotiation Agent Architecture:

[0193] like Figure 3 As shown, in the data input stage, historical transaction prices and related influencing factors are obtained. The required data are mainly divided into macro data, market data and meteorological data. Macro data include the provincial consumer price index, provincial gross domestic product, etc.; market data include transaction data such as bilateral negotiated historical transaction prices, marginal cost of power supply at the sending and receiving ends; meteorological data include data affecting the prediction such as the temperature at the sending and receiving ends and the humidity at the sending and receiving ends. In the data processing stage, the data is cleaned, standardized, and features are extracted. After the data is processed, the intelligent agent learning stage is entered, and the improved Actor-Critic algorithm proposed in this embodiment is used to predict the transaction price, calculate the deviation from the predicted transaction price and the actual transaction price deviation, and comprehensively update the single-month and multi-month strategy update coefficients, and iterate and update month by month. Finally, the prediction stage is entered, and the transaction price model output in the learning stage is used to output the related factor data to predict the expected transaction price in the next transaction window; then the strategy and status output in the learning stage are used to make corrections to the expected transaction price, and the final declared price is output.

[0194] Improved dual evaluation network Actor-Critic algorithm based on RNN;

[0195] Actor-Critic algorithm and its flaws:

[0196] The Actor-Critic algorithm is a reinforcement learning method that combines policy optimization and value function estimation, in which the Actor is responsible for directly outputting the policy and generating the probability distribution of the action based on the state, while the Critic evaluates the pros and cons of the current policy by estimating the value function (such as the state value function or the action value function) and generates a feedback signal to guide the Actor to update the policy parameters. The algorithm optimizes the Actor and the Critic alternately through the temporal difference (TD) error generated by the Critic, which not only retains the efficiency of the policy gradient method, but also utilizes the stability of the value function method, thereby ensuring better performance in continuous and high-dimensional action spaces.

[0197] However, the Actor-Critic algorithm has defects in the actual prediction and learning process. First, the accuracy of the Critic network's estimation of the value function greatly affects the optimization effect of the Actor strategy. A single Critic network is prone to estimation bias and high variance, which in turn affects the stability of the strategy update. The traditional Actor-Critic structure usually uses Feedforward Neural Networks (FNNs) as a function approximator, but FNNs have limited modeling capabilities for sequence data and cannot fully capture the temporal dependencies and long-term information that may exist in the environment, especially in environments with partial observability or complex dynamic characteristics. This limitation causes the Actor-Critic algorithm to perform poorly when dealing with long-term decision-making or memory-related tasks. To address these problems, this paper proposes an improved method that combines recurrent neural networks and dual evaluation networks to enhance the model's ability to capture temporal dependencies, and improves the stability and accuracy of value function estimation by introducing dual Critic networks.

[0198] This embodiment introduces the following key modules to the original Actor-Critic algorithm: 1. Long-term learning of recurrent neural network: The neural network used for action selection and value evaluation in the original algorithm is replaced by a recurrent neural network, thereby improving the performance of the algorithm in processing time series data and tasks with long-term dependencies. 2. Dual Critic network strategy value evaluation: On the basis of improving the Critic network to a recurrent neural network, an additional feedforward neural network is introduced to perform the same learning, and a weighted strategy value evaluation is given by combining the learning results of the two neural networks. 3. Priority playback cache mechanism: Introduce a playback buffer to extract previously trained data for secondary training when the algorithm is trained. Sample data that contributes greatly to model training is given a higher probability of being extracted, thereby ensuring the training effect of the neural network.

[0199] Improved dual-network RNN-Actor-Critic algorithm;

[0200] (1) Recurrent Neural Network:

[0201] A recurrent neural network is a neural network architecture specifically designed to process sequence data. Its innovation lies in the ability to use the output of the previous time step as the input of the current time step through a cyclically connected hidden layer, thereby establishing a memory mechanism within the network. This design enables RNN to capture the temporal dependencies and contextual information in sequence data. Compared with traditional feedforward neural networks, RNN is more suitable for processing time series, natural language processing, and other tasks involving sequential characteristics. Unlike FNN, which can only process independent samples, RNN can implicitly consider the dynamic characteristics of the input data in the model, thereby showing stronger capabilities in scenarios where long-term and short-term dependencies need to be modeled.

[0202] (2) Improved algorithm flow:

[0203] 1) Dual Critic network strategy value evaluation and recurrent neural network long-term learning;

[0204] The operation process of the original Actor-Critic algorithm: Starting from the interaction with the environment, the agent receives the current state s t After that, the Actor network outputs a probability distribution or action strategy based on the state, from which an action a is selected. t Execute. Then, the environment is based on the action a executed t Returns an instant reward r t and the next state s t+1 At the same time, the Critic network is in the current state s t and action a t As input, estimate the value function under the current policy and calculate the temporal difference (TD) error δ t =r t +γV(s t+1 )-V(s t ) to measure the accuracy of the value function estimate. During the update process, the Actor network uses the TD error generated by the Critic network as the weight of the policy gradient to optimize its strategy to maximize the cumulative reward; the Critic network updates its value function estimate by minimizing the mean square error based on the TD error. This alternating update process is repeated continuously, allowing the Actor to gradually learn the optimized action selection strategy, while the Critic network gradually approaches the true value function estimate, thereby achieving strategy optimization and accurate modeling of environmental dynamics. The core of the entire process lies in the collaboration between the Actor network and the Critic network. The Critic network provides feedback to the Actor network, and the Actor network relies on the Critic network's estimate to adjust its strategy, ultimately achieving efficient learning of the environment.

[0205] The improved algorithm proposed in this embodiment replaces the learning network of Actor and Critic with a recurrent neural network, transmits time series information through hidden states, and introduces a memory mechanism. For the Actor network, RNN inputs the current state s at each time step t. t and the hidden state of the previous time step Output the policy distribution of the current action:

[0206]

[0207] in is the hidden state of the Actor network at time step t, which can capture long-term dependencies in sequence data.

[0208] For the Critic network, RNN takes the current state s t 、Action a t and the hidden state at the previous time step As input, the output value function estimates:

[0209]

[0210] in By preserving historical information through the recurrent unit, the Critic network can estimate the value function more accurately, especially in non-Markov decision processes or partially observable environments.

[0211] Based on the introduction of RNN, this embodiment introduces a common feedforward neural network into the Critic network to learn recent data. The state value functions obtained by the two neural networks are weighted and added in a certain proportion to obtain the final output value function. This method ensures that the Critic network can comprehensively learn long-term and recent sample data. Suppose the state value function obtained by the feedforward neural network is V′(s t θ Critic ), then the true value function V finally output by the Critic network is final (s t θ critic )=V(s t θ Critic )+V′(s t θ Critic ).

[0212] 2) Priority playback cache mechanism:

[0213] Experience Replay Buffer is a commonly used technology in deep reinforcement learning, especially in deep Q network (DQN). It uses the experience tuple (st ,a t ,r t ,s t+1 ) is stored in a fixed-size buffer and a small batch of data is randomly sampled from it for learning each time training. This can break the temporal correlation between samples, improve the stability of the training process, and significantly improve sample utilization by repeatedly using past experience samples. In addition, the experience replay area can also smooth the changes in target values ​​and reduce the drastic fluctuations in target value updates, thereby accelerating the convergence of the model.

[0214] This embodiment introduces the experience replay area into the Actor-Critic algorithm and extracts samples from the experience replay area during training. In order to ensure that the experience replay pool can select samples with higher value, thereby accelerating model training, this paper performs weighted sampling on the samples. First, the absolute value of the TD error is used as the priority of the sample, and the priority p of the experience tuple is calculated. t =|δ t |+∈, where ∈ is a small constant that prevents the priority from being zero when the TD error is zero.

[0215] Next, the priorities of all samples are normalized to facilitate the calculation of sampling probability so that the priority of each sample conforms to the probability distribution:

[0216]

[0217] Among them, α controls the degree of influence of priority on sampling. Through this mechanism, the algorithm can give priority to samples with higher learning value during training, thereby accelerating convergence and improving model performance.

[0218] 3) Improve the algorithm process:

[0219] The improved algorithm process is as follows: the agent receives the current state s t After that, the Actor network outputs a probability distribution or action strategy based on the state, from which an action a is selected. t Execute. Then, the environment is based on the action a executed t Returns an instant reward r t and the next state s t+1 At the same time, the RNN neural network in the Critic network and the feedforward neural network learn simultaneously, taking the current state s t and action a tAs input, estimate the value function under the current strategy, and add the value functions obtained by the two neural networks to obtain the final true value function. After that, the accuracy of the value function estimation is measured by calculating the TD error. During the update process, high-value samples are extracted from the experience replay pool and provided to the Actor network and the Critic network for training. This alternating update process is repeated continuously, so that the Actor gradually learns the optimized action selection strategy, and the Critic network gradually approaches the true value function estimate, thereby achieving strategy optimization and accurate modeling of environmental dynamics. The framework of the algorithm is as follows Figure 4 shown.

[0220] Centralized bidding agent architecture:

[0221] When the transaction method is centralized bidding, the forecasting model will add influencing factors such as the historical inter-provincial spot prices of the sending end and other receiving provinces, and finally output the predicted bid price of the sending end after the strategy coefficient is updated, and the predicted bid price of the province and other receiving entities. At the same time, based on the inter-provincial spot transaction volume of each sending entity, the proportion of power source types in each sending province, installed capacity, etc., combined with factors such as changes in the supply and demand situation in each sending province, the centralized bidding bid power of each sending transaction entity is predicted; combined with the supply and demand form, economic index, meteorological conditions, power installed capacity and other trends and the province's predicted centralized bidding bid power (externally acquired data), the centralized bidding bid power of the receiving end (in competition with the province) is predicted. The structure of the centralized bidding intelligent agent is as follows: Figure 5 shown.

[0222] Centralized bidding agent quotation prediction process;

[0223] (1) Quotation of the sending agent:

[0224] The factors considered in the centralized bidding price quotation of the sending-end intelligent entity are: 1. The marginal cost of different power types; 2. The inter-provincial spot price level of the province where it is located; 3. The monthly bilateral negotiated price, etc. A certain strategy declaration is made comprehensively, and the specific process is as follows:

[0225] Step 1: First, based on the type of power supply at the sending end and the marginal cost of this type of power supply, the sending end quotation P1 is given.

[0226] Step 2: Secondly, based on the historical inter-provincial spot prices and bilateral negotiated prices of the province where the sending end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price forecast model, the inter-provincial spot prices and bilateral negotiated prices within the trading window are forecasted to obtain the corresponding day-ahead spot clearing price P. d (96 points) and bilateral negotiation forecast price P l .

[0227] Step 3: Combine P d , Pl , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Average (no historical declaration data, no reference).

[0228] Step 4: Assign different weights m1 and m2 (both are 0.5 by default) to the quotations (P1 and P2) generated in Step 1 and Step 3 respectively, and generate the final quotation P3.

[0229] Step 5: Output P1, P2 or P3 according to user selection.

[0230] (2) Sending-end intelligent agent reporting quantity:

[0231] Under the centralized bidding transaction mode, the sending-end intelligent body's reported quantity data is fitted and allocated to the intelligent body's historical inter-provincial spot transaction quantity through the multivariate linear regression method combined with the artificial rule method, and the corresponding intelligent body's declared quantity strategy is output in combination with the decision maker's risk preference setting (converted into the indicator coefficient set in advance). At the same time, the algorithm supports the use of the declared quantity strategy of each intelligent body input by manual declaration.

[0232] (3) Quote from the receiving agent:

[0233] The quotation generation of other receiving agents mainly considers the inter-provincial spot transaction price and bilateral negotiated price of the receiving agent, and generates the quotation strategy based on the above two factors. The specific process is as follows:

[0234] Step 1: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the receiving end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price model, the inter-provincial spot prices and bilateral negotiated prices within the transaction window are predicted to obtain the corresponding predicted spot price P d and bilaterally negotiated price P l ;

[0235] Step 2: Combine P d , P l , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean.

[0236] (4) Quantity reported by the receiving intelligent agent:

[0237] Under the centralized bidding transaction mode, the sending-end intelligent body's reported quantity data is fitted and allocated to the intelligent body's historical inter-provincial spot transaction quantity through the multivariate linear regression method combined with the artificial rule method, and the corresponding intelligent body's declared quantity strategy is output in combination with the decision maker's risk preference setting (converted into the indicator coefficient set in advance). At the same time, the algorithm supports the use of the declared quantity strategy of each intelligent body input by manual declaration.

[0238] In this embodiment, an improved dual evaluation network Actor-Critic algorithm based on RNN is introduced for learning and prediction of bilateral negotiation quotation agents and centralized bidding quotation agents. The algorithm captures long-term dependencies in time series data through a recurrent neural network (RNN), and improves the stability and accuracy of value function estimation through a dual critic network, thereby optimizing the agent's quotation strategy. The protection points of this technology include the specific implementation steps of the improved dual evaluation network Actor-Critic algorithm based on RNN, including the combined use of recurrent neural networks (RNN) and feedforward neural networks (FNN), the introduction of dual critic networks and the application of a priority playback cache mechanism, and the application of this algorithm in agent quotation, including the design and implementation of bilateral negotiation quotation agents and centralized bidding quotation agents.

[0239] See also Figure 6 A computer device 400 provided in an embodiment of the present application includes: a processor 410 and a memory 420, wherein the memory 420 stores a computer program executable by the processor 410, and when the computer program is executed by the processor 410, the above method is executed.

[0240] The embodiment of the present application further provides a storage medium 430 on which a computer program is stored. When the computer program is run by the processor 410, the above method is executed.

[0241] Among them, the storage medium 430 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable red-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, disk or optical disk.

[0242] In the description of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. "Multiple" means two or more, unless otherwise clearly and specifically defined.

[0243] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0244] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are contradictory.

[0245] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.

[0246] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.

[0247] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0248] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0249] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present invention. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for intelligent agent quotation in auxiliary decision-making for out-of-region electricity purchase, characterized in that: The steps include: Set transaction scenarios and declaration strategies, determine the modeling scope and output data items of the sending and receiving agents; Acquire historical data according to the modeling scope and output data items, and preprocess the historical data; Using the improved actor-critic algorithm RNN-Actor-Critic based on recurrent neural network to model the intelligent agent, and based on the transaction scenario, using the intelligent agent to simulate the transaction; The predicted boundary data of the target trading period is deduced through the change trend of the boundary data of the intelligent agent, and then the extreme gradient boosting Xgboost model is used to input the predicted boundary data to calculate the expected transaction price; The current declared price is calculated based on the expected transaction price, and the predicted transaction price, predicted boundary data, prediction model and predicted declared price are finally output.

2. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 1 is characterized in that: The transaction scenario includes: transaction type, transaction method, forecast time range, risk preference coefficient, randomness coefficient, and strategy target; The declaration strategy includes: modeling scope of sending and receiving intelligent agents, and output data items.

3. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 1 is characterized in that: The preprocessing of the historical data comprises the following steps: Identify and process errors, missing values, and outliers in the data; use the linear interpolation method to fill in missing values, that is, use the data outside the two ends of the missing period to perform linear regression modeling, predict missing values, and fill them in; Data normalization and standardization, scaling the data to a set range; Perform correlation analysis on the input data, select important features or create new features until the data meets the requirements for model training.

4. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 3 is characterized in that: The predicted missing values ​​are filled in, including: For entities that lack historical data, the average historical data of other entities in the same province will be used as training samples for the entities involved in this transaction; If the historical data of other subjects in the province to which the subject belongs is missing, the mean historical data of subjects in all provinces across the country is used as the training sample.

5. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 1 is characterized in that: The agent modeling using the improved actor-critic algorithm RNN-Actor-Critic based on the recurrent neural network includes the following steps: When the agent receives the current state s t After that, the actor network outputs a probability distribution or action strategy based on the state, from which an action a is selected. t implement; The environment is based on the action a performed t Returns an instant reward r t and the next state s t+1 At the same time, the recurrent neural network RNN ​​in the critic network and the feedforward neural network learn simultaneously, taking the current state s t and action a t As input, estimate the value function under the current strategy, and add the value functions obtained by the two neural networks to obtain the final true value function; The accuracy of the value function estimate is measured by calculating the TD error; During the update process, high-value samples are extracted from the experience replay pool and provided to the Actor network and the Critic network for training; this alternating update process is repeated continuously, so that the Actor gradually learns the optimized action selection strategy, and the Critic network gradually approaches the true value function estimate, achieving strategy optimization and modeling of environmental dynamics.

6. The intelligent agent quotation method in the auxiliary decision-making of out-of-region electricity purchase according to claim 5 is characterized in that: The method of using the intelligent agent to simulate transactions includes: During the simulated transaction process of the actor network, the generated predicted transaction price, predicted boundary data, predicted model and predicted declared price are added as samples to the experience replay pool of the improved algorithm. When the experience replay pool reaches the storage limit, the dual critic Critic network starts working. In each transaction, the priority replay cache mechanism in the improved algorithm is used to extract high-value sample data from the experience replay pool for the training of the actor recurrent neural network. When the sample data in the experience replay pool reaches the upper limit, the dual critic Critic network inside the agent starts learning, and the Actor network starts receiving the strategy value from the dual critic network; In each iteration, the dual critic network also uses the priority replay cache mechanism to extract high-value sample data from the experience replay pool for learning; During the prediction process, the dual critic network obtains the real boundary data, transaction results, and output of the actor network in the corresponding simulated transaction scenario. The predicted transaction price is obtained by using the prediction output of the actor network when the real boundary data is obtained. The two networks calculate the comprehensive deviation rate of the transaction price predicted by the actor network and the critic network respectively, and calculate the comprehensive deviation rate between the transaction price under the real boundary and the transaction price predicted by the critic network. Based on the comprehensive deviation rate, the Critic network calculates the evaluation value Reward for the Actor network's predicted performance; For the buyer, if the declared price is higher than the actual transaction price, the strategy coefficient will be lowered; after updating the strategy, the agent will end this round of simulation trading and output the latest strategy coefficient K t轮更新策略 , enter the next simulation transaction.

7. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 6 is characterized in that: The calculation of the comprehensive deviation rate between the transaction price under the actual boundary and the transaction price predicted by the Critic network includes: Predicted transaction price under the real boundary The predicted transaction price under the predicted boundary Deviation MAPE f The calculation method is as follows: Predicted transaction price under the real boundary and the actual transaction price P 实际成交 The deviation rate MAPE p The calculation method is as follows:

8. The intelligent agent quotation method in the auxiliary decision-making of out-of-region electricity purchase according to claim 6 is characterized in that: Based on the comprehensive deviation rate, the Critic network calculates the evaluation value Reward for the Actor network's predicted performance, including: Reward t =α*MAPE f +β*MAPE p ; In the formula, α and β are the update step sizes. The larger the data prediction deviation, the larger the strategy coefficient needs to be modified. The agent uses the change in the reward of the past three months to calculate the reward of round t t After the correction is completed, the calculation method is as follows: Reward t =sD t-1 +d 2 D t-2 +d 3 D t-3 ; In the formula, Δ i-1 is the change in the Reward value from the i-1th month to the tth month, that is, Reward i -Reward i-1 ; δ t-i is the weight coefficient. The weight decays exponentially with the lag time, that is, the older the change, the lower the contribution. Assuming that round t is the first three rounds of simulation, then Δ t-i All are supplemented with 0 values.

9. The intelligent agent quotation method in the auxiliary decision-making of out-of-region electricity purchase according to claim 6 is characterized in that: The priority playback cache mechanism includes: The experience tuple (s) generated when the agent interacts with the environment t ,a t ,r t ,s t+1 ) is stored in a fixed-size buffer and a mini-batch of data is randomly sampled from it for learning during each training; During training, samples are extracted from the buffer, the absolute value of the TD error is used as the priority of the sample, and the priority p of the experience tuple is calculated. t =|δ t |+∈, where ∈ is a small constant to prevent the priority from being zero when the TD error is zero; Normalize the priorities of all samples so that the priority of each sample conforms to the probability distribution: Where α controls the degree of influence of priority on sampling.

10. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 1 is characterized in that: The method calculates the predicted boundary data of the target transaction period by the change trend of the boundary data of the intelligent agent, and then uses the extreme gradient boosting Xgboost model to input the predicted boundary data to calculate the expected transaction price, including the following steps: Use the regression model or weighted moving average model to deduce the predicted boundary data for the target trading period, and then use the Xgboost model to input the predicted boundary data to calculate the expected transaction price; By superimposing the declaration strategy, the declaration price is calculated based on the expected transaction price, and the predicted transaction price, predicted boundary data, prediction model and predicted declaration price are finally output; The calculation method of the declaration strategy is as follows: P t轮申报 =P t轮Actor预期成交 *Risk*K t轮 ; Risk represents risk preference, which is three different preference types of aggressive, conservative and conventional that are manually input, and the coefficients are randomly selected in different intervals; K is the strategy coefficient itself.

11. The intelligent agent quotation method in the auxiliary decision-making of out-of-region power purchase according to claim 1 is characterized in that: Also includes: When the transaction method is centralized bidding, the influencing factors of the historical inter-provincial spot prices of the sending end and other receiving provinces are added to consider, and the final output is the predicted declared price of the sending end after the strategy coefficient is updated, and the predicted declared price of the province and other receiving entities.

12. The intelligent agent quotation method in the auxiliary decision-making of out-of-region electricity purchase according to claim 11 is characterized in that: The final output includes the predicted bid price of the sending end after the strategy coefficient is updated, the predicted bid price of the province and other receiving entities, and the centralized bidding bid power of each sending end trading entity is predicted based on the inter-provincial spot transaction power of each sending end entity, the proportion of power source types and installed capacity of each sending end province, and the factors of changes in the supply and demand situation of each sending end province; The receiving-end centralized bidding power declaration is obtained by combining the supply and demand patterns, economic indexes, meteorological conditions, changing trends in installed power generation capacity and the province's forecast of centralized bidding power declaration.

13. The intelligent agent quotation method in the auxiliary decision-making of out-of-region electricity purchase according to claim 12 is characterized in that: The centralized bidding power reported by each sending-end trading entity is obtained by combining the factors of supply and demand situation changes in each sending-end province, including: The sending agent quotes: Step 1: Combine the power type of the sending end and the marginal cost of this type of power to give the sending end quotation P1; Step 2: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the sending end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price forecast model, the inter-provincial spot prices and bilateral negotiated prices within the trading window are forecasted to obtain the corresponding day-ahead spot clearing price P. d and bilateral negotiation forecast price P l ; Step 3: Combine P d , P l , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean; Step 4: Assign different weights m1 and m2 to the quotations P1 and P2 generated in Step 1 and Step 3 respectively, and generate the final quotation P3; Step 5: Output P1, P2 or P3 according to user selection; The sending end intelligent agent reports: The sending-end intelligent agent's reported quantity data is fitted and allocated through the multivariate linear regression method combined with the artificial rule method to the intelligent agent's historical inter-provincial spot transaction quantity. Combined with the decision maker's risk preference setting, the corresponding intelligent agent's reported quantity strategy is output.

14. The intelligent agent quotation method in the auxiliary decision-making of out-of-region electricity purchase according to claim 12 is characterized in that: The centralized bidding power reported by the receiving end is obtained by combining the supply and demand forms, economic indexes, meteorological conditions, the changing trend of power installed capacity and the forecast of the centralized bidding power reported by the province, including: Receiving agent quote: Step 1: Based on the historical inter-provincial spot prices and bilateral negotiated prices in the province where the receiving end is located, and on the premise of establishing the inter-provincial spot forecast price and bilateral negotiated price model, the inter-provincial spot prices and bilateral negotiated prices within the transaction window are predicted to obtain the corresponding predicted spot price P d and bilaterally negotiated price P l ; Step 2: Combine P d , P l , according to the user's risk preference, the sender quote P2 is converted according to the preference range, and P2 should satisfy: P l ≤P2≤P d , the default is P d , P l Mean; Receiver intelligent body report quantity: The sending-end intelligent agent's reported quantity data is fitted and allocated through the multivariate linear regression method combined with the artificial rule method to the intelligent agent's historical inter-provincial spot transaction quantity. Combined with the decision maker's risk preference setting, the corresponding intelligent agent's reported quantity strategy is output.

15. An intelligent quotation device for auxiliary decision-making of out-of-region electricity purchase, characterized in that: Using the method according to any one of claims 1 to 14, the device comprises: The setting unit is used to set the transaction scenario and declaration strategy, determine the modeling scope and output data items of the sending and receiving agents; A collection unit, used for acquiring historical data according to the modeling scope and output data items, and preprocessing the historical data; An agent modeling unit, used for performing agent modeling using an improved actor-critic algorithm RNN-Actor-Critic based on a recurrent neural network, and simulating transactions using the agent based on the transaction scenario; A prediction unit, used to calculate the predicted boundary data of the target transaction period according to the change trend of the boundary data of the intelligent agent, and then use the extreme gradient boosting Xgboost model to input the predicted boundary data to calculate the expected transaction price; The output unit is used to calculate the current declared price based on the expected transaction price, and finally output the predicted transaction price, predicted boundary data, prediction model and predicted declared price.

16. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 14 is implemented.

17. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 14 is implemented.

Citation Information

Cited By

  • Medium and short term electric power spot price prediction method based on multi-agent simulation

    CN120707186A

  • Logistics quotation management method and system

    CN120725567A