Aggregator optimal bidding method based on deep reinforcement learning and hybrid intelligent architecture
By employing an aggregator-based optimal bidding method based on deep reinforcement learning and a hybrid intelligent architecture, the problems of low accuracy and poor robustness of bidding strategies in quasi-linear demand response are addressed. This method achieves efficient decision-making and high returns in abnormal scenarios, thereby improving the accuracy and robustness of bidding strategies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI UNIVERSITY OF ELECTRIC POWER
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-10
Smart Images

Figure CN122367594A_ABST
Abstract
Description
Technical Field
[0001] This invention discloses an optimal bidding method for aggregators based on deep reinforcement learning and a hybrid intelligent architecture, which involves demand response market bidding optimization and demand response resource scheduling technology, and belongs to the technical field of demand response scheduling and intelligent decision control. Background Technology
[0002] With the advancement of new power system construction, demand response, as a flexible adjustment resource, plays an increasingly crucial role in balancing power supply and demand and improving grid efficiency. Aggregators, as the link between user-side resources and the demand response market, directly impact the effectiveness of demand response and their own revenue through the rationality of their bidding strategies. Among these, quasi-linear demand response, due to the linear correlation characteristics of load regulation and the existence of boundary constraints, places higher demands on aggregators' bidding accuracy and dynamic adaptation capabilities.
[0003] Currently, existing aggregator bidding strategies largely rely on single-decision-making or pure machine algorithms. While single-decision-making can incorporate market experience, it is inefficient in processing massive amounts of user load data and struggles to accurately quantify linear constraint boundaries. Pure machine algorithms, while capable of rapid data processing, lack the flexibility to adapt to sudden changes in market rules and uncertainties in user responses. Furthermore, existing methods are inadequate in addressing real-time electricity price fluctuations and user response deviations in the demand response market. This can easily lead to a disconnect between bidding proposals and actual demand, resulting in issues such as over- or under-bidding capacity, lower-than-expected returns, or low response execution rates. A bidding strategy that balances efficiency, accuracy, and flexibility has yet to be developed. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of the aforementioned background technologies by providing an optimal bidding method for load aggregators based on deep reinforcement learning and a hybrid intelligent architecture. By analyzing the response mechanism of load aggregators and user uncertainties, a hybrid intelligent architecture is constructed, and a multi-agent deep reinforcement learning algorithm (MATD3) improved by VMD-FFCM-LSTM-iTransformer is used for training and decision-making. This hybrid intelligent architecture overcomes the limitations of pure machine decision-making in abnormal scenarios, solving the technical problems of low accuracy and poor robustness of bidding strategies for aggregators in quasi-linear demand response.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0006] The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture includes the following steps:
[0007] S1. Obtain the load baseline, demand response price, and user response data released by the power grid demand response center, and perform preprocessing. Standardize the load curve and use Euclidean distance to measure the similarity between the load curve and the load baseline, which serves as the basis for calculating subsidy revenue.
[0008] S2. Based on the principles of consumer psychology and uncertainty models, characterize user response behavior, including the minimum perceptible threshold, saturation response price, and the normal distribution of response quantity deviation.
[0009] S3. Construct the optimal bidding model for aggregators, including the objective function and constraints. The objective function focuses on maximizing the aggregator's profits, comprehensively considering subsidy revenue, user incentive costs, intervention costs, and penalty costs for failing to meet targets. The constraints cover linear response, bid price, ranking price, subsidy adjustment, and limits on the extent of manual intervention.
[0010] S4. A multi-agent deep reinforcement learning algorithm is adopted, combined with the improved multi-agent dual-delay deep deterministic policy gradient algorithm of VMD-FFCM-LSTM-iTransformer, to train the optimal bidding model of the aggregator.
[0011] S5. Design a hybrid intelligent architecture, including a human-computer interaction layer, a decision fusion layer, and an execution feedback layer, to realize a collaborative mechanism of "machines leading routine decisions and humans intervening in abnormal scenarios". The human-computer interaction layer provides real-time data monitoring and manual adjustment interfaces. The decision fusion layer adopts a dynamic weight allocation mechanism to integrate machine and human decisions. The execution feedback layer provides feedback updates.
[0012] S6. At the decision fusion layer, the fusion weight of machine and human decision-making is dynamically adjusted according to the level of abnormal alarms. In normal scenarios, the machine takes the lead, while in abnormal scenarios, human intervention is enhanced. The fusion results are constrained and verified to ensure the feasibility of the bidding strategy.
[0013] S7. By executing the feedback layer, the data is fed back to the multi-agent deep reinforcement learning model to update the Actor-Critic network parameters and optimize future human intervention strategies.
[0014] S8. Train and test the model using historical transaction data and simulation environment. Evaluate the algorithm performance through reward function and profit index to verify the effectiveness and superiority of the method in normal and abnormal scenarios.
[0015] As a preferred technical solution of the present invention: In step S1, the load curve is normalized as follows:
[0016] set up For the set of intervals divided into linear response periods, T = {t1,…,t7}. After the power grid demand response center releases the load curve, it first needs to standardize the load curve within period T. The purpose is to simplify calculations and provide a standardized load baseline target for aggregators. The formula is as follows:
[0017] (1);
[0018] In the formula, For the load curve The per-unit value.
[0019] As a preferred technical solution of the present invention: In step S2, when modeling user response behavior based on consumer psychology theory, two key parameters, the minimum perceptible threshold and the saturation response price, are introduced to establish a piecewise linear incentive-response model to characterize the change of user response quantity with incentive price, as follows:
[0020] When the incentive price is below the minimum perceptible threshold, users are not sensitive to the incentive, and no response behavior has been triggered. When the incentive price exceeds this threshold, the user response volume increases approximately linearly with the increase of the incentive price. As the incentive price increases further, limited by the user's own adjustment ability, the response volume gradually approaches saturation and remains stable. Its change no longer changes significantly with the increase of the incentive price. Its mathematical expression is as follows:
[0021] (2);
[0022] In the formula, This represents the response volume of user n in time period t, calculated using a consumer psychology model. This represents the maximum response capability of user n in time period t; , Let these represent the initial response price and the saturation response price for user n in time period t, respectively.
[0023] As a preferred technical solution of the present invention: In step S2, based on the idea of uncertainty modeling, the randomness in user demand response behavior is characterized. It is assumed that the deviation of user response quantity follows a normal distribution, and its standard deviation is negatively correlated with the incentive price, as follows:
[0024] As the incentive price increases, the uncertainty of user response behavior gradually decreases, and the standard deviation of the response bias decreases accordingly. When the incentive price is below a preset threshold, the volatility of user response increases, and the level of uncertainty rises, as shown in the following formula:
[0025] (3);
[0026] In the formula, , Let represent the expected value and standard deviation of the deviation of user n response volume in time period t, respectively.
[0027] As a preferred technical solution of the present invention: In step S3, when constructing the optimal bidding model for aggregators, the objective function is to maximize total profit, the subsidy revenue is calculated based on the demand response price and curve similarity, the user cost is determined by the incentive price and the response volume, the intervention cost is only generated during manual adjustments, and the penalty cost is set for the non-compliance response volume, as shown in the following formula:
[0028] (4);
[0029] In the formula, Q is the aggregator's total profit, and W... a (t) represents the subsidy revenue that the aggregator receives from the power grid during the response period T; C U (t) is the incentive cost that the aggregator has to pay to the user; C P (t) represents the cost of penalties incurred for failing to meet the standards; The intervention cost of human-machine hybrid decision-making is incurred only when decisions are adjusted manually.
[0030] As a preferred technical solution of the present invention: In step S3, the constraints ensure that the total electricity consumption of users remains unchanged, the bidding price is within the limit, the ranking price meets the response performance index, and the extent of manual intervention is limited by an upper limit, as shown in the following formula:
[0031] (5);
[0032] (6);
[0033] (7);
[0034] In the formula, To establish a unified market clearing price; P d This represents the number of bid responses from aggregators. and Let be the per-unit load values for the aggregator and the power grid company, respectively, in time period t. The Euclidean distance is used to measure the similarity between the per-unit load curve of the aggregator and the load baseline published by the power grid company, where d is the Euclidean distance between the market participant and the load baseline. This serves as a similarity indicator between demand response market participants and load lines. The similarity coefficient. This is the subsidy adjustment coefficient;
[0035] Human intervention is based on machine decision-making, as shown in the following formula:
[0036] (8);
[0037] In the formula The manual adjustment range is limited to 0.2 by default, meaning the adjustment range cannot exceed 20% of the machine's initial value.
[0038] As a preferred technical solution of the present invention: In step S4, when using the MADRL algorithm, the bidding strategy problem is modeled as a Markov game, defining a state space and an action space. The state space includes the load curve, clearing price, and user bias, while the action space includes the bid quantity, bid price, and a reward function based on profit and completion rate, as detailed below:
[0039] The state space representation of the i-th aggregator at time step j is:
[0040] (9);
[0041] In the formula, It is the shape of the load curve published by the power grid at time step j, t j For the current time, This is the market clearing price at this stage. It is a vector representing the uncertainty deviation between the response of the i-th aggregator and the response of all users aggregated by itself at time step j;
[0042] The action space of the i-th aggregator at time step j is the bidding strategy under the current circumstances, as shown in the following formula:
[0043] (10);
[0044] In the formula, It is the bid volume of the i-th aggregator at time step j. It is its bid price;
[0045] (11);
[0046] It is the revenue of the i-th aggregator. It is the i-th yes The subsidy price for time period t has a base value of 3 yuan for peak shaving and 1.2 yuan for valley filling; K t The completion adjustment factor for time period t is linked to the response completion rate; V t The actual response capacity for time period t is expressed in kilowatts; the regulation cost is the compensation paid by the aggregator to the user and the cost of equipment wear and tear, which is positively correlated with the load regulation amount, where U is the set of users. user The unit adjustment cost is expressed in yuan / kilowatt and reflects equipment wear or comfort compensation. User The load adjustment amount during time period t, expressed in kilowatts. The target response capacity for time period t is determined by the grid curve command; the penalty factor for insufficient completion is applied only when... Activated at that time.
[0047] As a preferred technical solution of the present invention: in step S5...
[0048] When designing a hybrid intelligent architecture, the human-computer interaction layer displays key indicators through a real-time dashboard, and the anomaly alarm module triggers manual intervention.
[0049] The decision fusion layer sets fusion weights based on alarm levels;
[0050] The execution feedback layer updates the machine model parameters and generates an intervention effect report, forming a closed-loop optimization.
[0051] As a preferred technical solution of the present invention: In step S6, the decision fusion layer fuses the initial machine decision and the manually adjusted decision to output the final bidding strategy. The fusion algorithm adopts a dynamic weight allocation mechanism, as shown in the following formula:
[0052] (12);
[0053] In the formula For the final bid amount, For weight fusion.
[0054] As a preferred technical solution of the present invention: In step S6, the abnormal scenarios include mild anomalies and severe anomalies, and the fusion weight settings are as follows:
[0055] Normal scenario: The fusion weight is 0.8, which is machine-dominated. Mild anomaly: The fusion weight is 0.5, which is human-machine balanced. Severe anomaly: The fusion weight is 0.3, which is human-dominated.
[0056] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0057] 1. Address the issue of decision robustness and improve adaptability in abnormal scenarios. Through a hybrid intelligent architecture, a collaborative mechanism is achieved between machine-led routine decision-making and human intervention in abnormal scenarios. This compensates for the biases of pure machine decision-making in situations such as sudden changes in electricity prices or abnormal data distribution deviations in user responses, avoiding profit losses and substandard responses due to decision biases. In abnormal scenarios, human-machine hybrid decision-making will increase average profits and improve response completion compared to pure machine decision-making.
[0058] 2. Improve the accuracy and profitability of bidding strategies. The improved MADRL algorithm (VMD-FFCM-LSTM-iTransformer-MATD3) using VMD-FFCM-LSTM-iTransformer effectively captures the long-term time-series dependencies of load curves, enhancing the accuracy and market competitiveness of bidding strategies. In typical scenarios, compared to pure machine decision-making, the human-machine hybrid approach increases average profit and improves response completion; further optimization of revenue and reduction of penalties for non-compliance are achieved through manual fine-tuning of bid prices and quantities.
[0059] 3. Provides more intelligent decision support while optimizing decision-making efficiency and response speed. Through a layered hybrid intelligent architecture consisting of a human-computer interaction layer, a decision fusion layer, and an execution feedback layer, the decision-making task is decomposed into machine initial decision-making and manual adjustment, reducing the subjectivity and computational complexity of purely manual decision-making. Although human intervention is introduced, through dynamic weight allocation and constraint verification, the decision-making time is controlled within 2-3 minutes, meeting the power grid response time requirements. Compared to purely manual decision-making, efficiency is significantly improved, while ensuring the real-time nature and feasibility of the decision. Attached Figure Description
[0060] Figure 1 A flowchart of the aggregator's optimal bidding method based on deep reinforcement learning and hybrid intelligent architecture; Figure 2 This is a training result diagram of the intelligent optimization algorithm. Detailed Implementation
[0061] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention.
[0062] like Figure 1 As shown, the optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture proposed in this invention includes the following steps:
[0063] S1. Obtain the load baseline, demand response price, and user response data released by the power grid demand response center, and perform preprocessing. Standardize the load curve and use Euclidean distance to measure the similarity between the load curve and the load baseline, which serves as the basis for calculating subsidy revenue.
[0064] in:
[0065] The per-unit normalization process for the load curve is as follows:
[0066] set up For the set of intervals divided into linear response periods, T = {t1,…,t7}. After the power grid demand response center releases the load curve, it first needs to standardize the load curve within period T. The purpose is to simplify calculations and provide a standardized load baseline target for aggregators. The formula is as follows:
[0067] (1);
[0068] In the formula, For the load curve The per-unit value.
[0069] S2. Based on the principles of consumer psychology and uncertainty models, characterize user response behavior, including the minimum perceptible threshold, saturation response price, and the normal distribution of response quantity deviation.
[0070] Specifically:
[0071] When modeling user response behavior based on consumer psychology theory, two key parameters, the minimum perceptible threshold and the saturation response price, are introduced to establish a piecewise linear incentive-response model to characterize the change in user response quantity with the incentive price, as follows:
[0072] When the incentive price is below the minimum perceptible threshold, users are not sensitive to the incentive, and no response behavior has been triggered. When the incentive price exceeds this threshold, the user response volume increases approximately linearly with the increase of the incentive price. As the incentive price increases further, limited by the user's own adjustment ability, the response volume gradually approaches saturation and remains stable. Its change no longer changes significantly with the increase of the incentive price. Its mathematical expression is as follows:
[0073] (2);
[0074] In the formula, This represents the response volume of user n in time period t, calculated using a consumer psychology model. This represents the maximum response capability of user n in time period t; , Let $\mathbf$ represent the initial response price and the saturation response price for user $n$ in time period $t$.
[0075] Based on the concept of uncertainty modeling, this paper characterizes the randomness in user demand response behavior. It assumes that the deviation of user response volume follows a normal distribution, and that its standard deviation is negatively correlated with the incentive price, as detailed below:
[0076] As the incentive price increases, the uncertainty of user response behavior gradually decreases, and the standard deviation of the response bias decreases accordingly. When the incentive price is below a preset threshold, the volatility of user response increases, and the level of uncertainty rises, as shown in the following formula:
[0077] (3);
[0078] In the formula, , Let represent the expected value and standard deviation of the deviation of user n response volume in time period t, respectively.
[0079] S3. Construct the optimal bidding model for aggregators, including the objective function and constraints. The objective function focuses on maximizing the aggregator's profits, comprehensively considering subsidy revenue, user incentive costs, intervention costs, and penalty costs for failing to meet targets. The constraints cover linear response, bid price, ranking price, subsidy adjustment, and limits on the extent of manual intervention.
[0080] Specifically:
[0081] When constructing the optimal bidding model for aggregators, the objective function is to maximize total profit. Subsidy revenue is calculated based on demand response price and curve similarity. User costs are determined by incentive price and response volume. Intervention costs are only incurred during manual adjustments. Penalty costs are set for non-compliance response volume, as shown in the following formula:
[0082] (4);
[0083] In the formula, Q is the aggregator's total profit, and W... a (t) represents the subsidy revenue that the aggregator receives from the power grid during the response period T; C U (t) is the incentive cost that the aggregator has to pay to the user; C P (t) represents the cost of penalties incurred for failing to meet the standards; The intervention cost of human-machine hybrid decision-making is incurred only when decisions are adjusted manually.
[0084] The constraints ensure that the total electricity consumption of users remains unchanged, the bidding price is within the limit, the ranked price meets the response performance indicators, and the degree of manual intervention is limited to the upper limit. The formula is as follows:
[0085] (5);
[0086] (6);
[0087] (7);
[0088] In the formula, To establish a unified market clearing price; P d This represents the number of bid responses from aggregators. and Let be the per-unit load values for the aggregator and the power grid company, respectively, in time period t. The Euclidean distance is used to measure the similarity between the per-unit load curve of the aggregator and the load baseline published by the power grid company, where d is the Euclidean distance between the market participant and the load baseline. This serves as a similarity indicator between demand response market participants and load lines. The similarity coefficient. This is the subsidy adjustment coefficient;
[0089] Human intervention is based on machine decision-making, as shown in the following formula:
[0090] (8);
[0091] In the formula The manual adjustment range is limited to 0.2 by default, meaning the adjustment range cannot exceed 20% of the machine's initial value.
[0092] S4. The optimal bidding model of the aggregator is trained by using a multi-agent dual-delay deep deterministic policy gradient algorithm improved by deep learning algorithm.
[0093] Specifically:
[0094] When using the MADRL algorithm, the bidding strategy problem is modeled as a Markov game, defining a state space and an action space. The state space includes the load curve, clearing price, and user bias, while the action space includes the bid quantity, bid price, and a reward function based on profit and completion rate, as detailed below:
[0095] The state space representation of the i-th aggregator at time step j is:
[0096] (9);
[0097] In the formula, It is the shape of the load curve published by the power grid at time step j, t j For the current time, This is the market clearing price at this stage. It is a vector representing the uncertainty deviation between the response of the i-th aggregator and the response of all users aggregated by itself at time step j;
[0098] The action space of the i-th aggregator at time step j is the bidding strategy under the current circumstances, as shown in the following formula:
[0099] (10);
[0100] In the formula, It is the bid volume of the i-th aggregator at time step j. It is its bid price;
[0101] (11);
[0102] It is the revenue of the i-th aggregator. It is the i-th yes The subsidy price for time period t has a base value of 3 yuan for peak shaving and 1.2 yuan for valley filling; K t The completion adjustment factor for time period t is linked to the response completion rate; V t The actual response capacity for time period t is expressed in kilowatts; the regulation cost is the compensation paid by the aggregator to the user and the cost of equipment wear and tear, which is positively correlated with the load regulation amount, where U is the set of users. user The unit adjustment cost is expressed in yuan / kilowatt and reflects equipment wear or comfort compensation. User The load adjustment amount during time period t, expressed in kilowatts. The target response capacity for time period t is determined by the grid curve command; the penalty factor for insufficient completion is applied only when... Activated at that time.
[0103] S5. Design a hybrid intelligent architecture, including a human-computer interaction layer, a decision fusion layer, and an execution feedback layer, to realize a collaborative mechanism of "machines leading routine decisions and humans intervening in abnormal scenarios". The human-computer interaction layer is responsible for detecting abnormal data, the decision fusion layer adopts a dynamic weight allocation mechanism to integrate machine and human decisions, and the execution feedback layer provides feedback updates.
[0104] in:
[0105] When designing a hybrid intelligent architecture, the human-computer interaction layer displays key indicators through a real-time dashboard, and the anomaly alarm module triggers manual intervention.
[0106] The decision fusion layer sets fusion weights based on alarm levels;
[0107] The execution feedback layer updates the machine model parameters and generates an intervention effect report, forming a closed-loop optimization.
[0108] S6. At the decision fusion layer, the fusion weight of machine and human decision-making is adjusted according to the degree of abnormal alarms. In normal scenarios, the machine takes the lead, while in abnormal scenarios, human intervention is enhanced. The fusion results are constrained and verified to ensure the feasibility of the bidding strategy.
[0109] in:
[0110] The decision fusion layer integrates the initial machine decision and the manually adjusted decision to output the final bidding strategy. The fusion algorithm uses a dynamic weight allocation mechanism, as shown in the following formula:
[0111] (12);
[0112] In the formula For the final bid amount, For weight fusion.
[0113] Abnormal scenarios include minor anomalies and severe anomalies, and the specific weighting settings for fusion are as follows:
[0114] Normal scenario: The fusion weight is 0.8, which is machine-dominated. Mild anomaly: The fusion weight is 0.5, which is human-machine balanced. Severe anomaly: The fusion weight is 0.3, which is human-dominated.
[0115] S7. By executing the feedback layer, the data is fed back to the multi-agent deep reinforcement learning model to update the Actor-Critic network parameters and optimize future human intervention strategies.
[0116] S8. Train and test the model using historical transaction data and simulation environment. Evaluate the algorithm performance through reward function and profit index to verify the effectiveness and superiority of the method in normal and abnormal scenarios.
[0117] The invention will be illustrated below with three examples.
[0118] Example 1: The VMD-FFCM-LSTM-iTransformer was trained by combining it with the MATD3 and MADDPG algorithms respectively. The training results are as follows: Figure 2 As shown.
[0119] As shown in the figure, the reward value gradually increases with the progression of training rounds. The VMD-FFCM-LSTM-iTransformer-MATD3 combination converges first around 500 rounds, and its reward value stabilizes in the 93.1-93.5 range after 900-1000 rounds. Compared to the 91.8-93.3 range of VMD-FFCM-LSTM-iTransformer-MADDPG, this combination is not only more stable and has smaller fluctuations, but also has a higher cumulative reward value. This fully demonstrates the superior stability and good convergence of the VMD-FFCM-LSTM-iTransformer-MATD3 algorithm in the later stages. Conversely, the VMD-FFCM-LSTM-iTransformer-MADDPG algorithm exhibits more significant fluctuations. This is likely due to MADDPG's less precise and nuanced control over parameter updates during training, or its tendency to get trapped in local optima when handling interactions between multiple agents, thus negatively impacting overall stability and the final reward value. Therefore, the good performance of the VMD-FFCM-LSTM-iTransformer-MATD3 algorithm is due to the interplay of several factors: VMD first decomposes the original data into multiple stationary modes, eliminating noise and non-stationarity, and providing high-quality input for subsequent processing; FFCM extracts and filters features from each decomposed mode, enhancing data representation capabilities and reducing dimensionality; LSTM captures the temporal dependencies of each mode, learns local change patterns, and achieves multi-scale feature fusion; iTransformer models the relationships between variables from a global perspective, captures long-term dependencies, and provides more comprehensive prediction basis; MATD3 has significant advantages in multi-agent collaboration and training stability; and these factors, along with their good adaptability to specific tasks, result in a complex interplay of these factors.
[0120] Example 2: Set up three comparison scenarios to verify the superiority of the hybrid intelligent architecture:
[0121] Option 1: Pure MADRL decision-making, without VMD-FFCM-LSTM-iTransformer improvements, and without human intervention;
[0122] Option 2: VMD-FFCM-LSTM-iTransformer-MADRL decision-making, with improvements over VMD-FFCM-LSTM-iTransformer, requiring no human intervention;
[0123] Solution 3: VMD-FFCM-LSTM-iTransformer-MADRL with human-machine hybrid decision-making, which is the solution proposed in this paper;
[0124] Option 4: Purely manual decision-making, based on the experience of dispatchers, without machine support.
[0125] Under typical scenarios such as stable electricity prices and a user response rate of ≥90%, the average profit and response completion rate of the aggregators for the four schemes are shown in Table 1.
[0126]
[0127] Table 1
[0128] The results show that, compared with Scheme 1, Scheme 2 improved profits by 15.9% with the VMD-FFCM-LSTM-iTransformer, indicating that long-range dependency capture is effective; compared with Scheme 2, the human-machine hybrid scheme 3 further improved profits by 2.1% because the bidding price was optimized by human fine-tuning; Scheme 4, which was purely manual decision-making, had the lowest profits and the longest time, proving the necessity of machine-led decision-making.
[0129] Example 3: Set up an abnormal scenario where the clearing price suddenly increases to 12 yuan / kW in the fourth period, exceeding the upper limit by 20%, and the user response rate drops to 60%. The results of the four schemes are shown in Table 2.
[0130]
[0131] Table 2
[0132] Example 3 demonstrates that the VMD-FFCM-LSTM-iTransformer-MADRL hybrid human-machine decision-making solution possesses strong scenario adaptability and risk resistance. Its average profit of 82,000 yuan and response completion rate of 88.7% far exceed those of pure MADRL decision-making, solutions with only algorithm improvements and no human intervention, and pure human decision-making solutions. This not only reflects that the improved VMD-FFCM-LSTM-iTransformer algorithm can provide basic technical support for decision-making in abnormal scenarios and make up for the shortcomings of pure machine decision-making under data distribution deviations, but also proves that the mechanism of "human intervention in abnormal scenarios" in the hybrid intelligent architecture can effectively optimize bidding strategies. At the same time, the decision-making time of this solution in 2 minutes is much lower than that of pure human decision-making in 3 minutes, achieving a balance between decision-making effect and decision-making efficiency in abnormal scenarios. Pure machine decision-making is unable to flexibly adapt to sudden market and user changes, and pure human decision-making is subject to strong subjectivity and low efficiency, both of which are difficult to achieve the comprehensive benefits and response standards of human-machine hybrid decision-making.
[0133] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.
Claims
1. An optimal bidding method for aggregators based on deep reinforcement learning and a hybrid intelligent architecture, characterized in that, Includes the following steps: S1. Obtain the load baseline, demand response price, and user response data released by the power grid demand response center, and perform preprocessing. Standardize the load curve and use Euclidean distance to measure the similarity between the load curve and the load baseline, which serves as the basis for calculating subsidy revenue. S2. Based on the principles of consumer psychology and uncertainty models, characterize user response behavior, including the minimum perceptible threshold, saturation response price, and the normal distribution of response quantity deviation. S3. Construct the optimal bidding model for aggregators, including the objective function and constraints. The objective function focuses on maximizing the aggregator's profits, comprehensively considering subsidy revenue, user incentive costs, intervention costs, and penalty costs for failing to meet targets. The constraints cover linear response, bid price, ranking price, subsidy adjustment, and limits on the extent of manual intervention. S4. The optimal bidding model of the aggregator is trained by using a multi-agent dual-delay deep deterministic policy gradient algorithm improved by deep learning algorithm. S5. Design a hybrid intelligent architecture, including a human-computer interaction layer, a decision fusion layer, and an execution feedback layer, to realize a collaborative mechanism of "machines leading routine decisions and humans intervening in abnormal scenarios". The human-computer interaction layer is responsible for detecting abnormal data, the decision fusion layer adopts a dynamic weight allocation mechanism to integrate machine and human decisions, and the execution feedback layer provides feedback updates. S6. At the decision fusion layer, the fusion weight of machine and human decision-making is adjusted according to the degree of abnormal alarms. In normal scenarios, the machine takes the lead, while in abnormal scenarios, human intervention is enhanced. The fusion results are constrained and verified to ensure the feasibility of the bidding strategy. S7. By executing the feedback layer, the data is fed back to the multi-agent deep reinforcement learning model to update the Actor-Critic network parameters and optimize future human intervention strategies. S8. Train and test the model using historical transaction data and simulation environment. Evaluate the algorithm performance through reward function and profit index to verify the effectiveness and superiority of the method in normal and abnormal scenarios.
2. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S1, The per-unit normalization process for the load curve is as follows: set up For the set of intervals divided into linear response periods, T = {t1,…,t7}. After the power grid demand response center releases the load curve, it first needs to standardize the load curve within period T. The purpose is to simplify calculations and provide a standardized load baseline target for aggregators. The formula is as follows: (1); In the formula, For the load curve The per-unit value.
3. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S2, When modeling user response behavior based on consumer psychology theory, two key parameters, the minimum perceptible threshold and the saturation response price, are introduced to establish a piecewise linear incentive-response model to characterize the change in user response quantity with the incentive price, as follows: When the incentive price is below the minimum perceptible threshold, users are not sensitive to the incentive and the response behavior has not yet been triggered; when the incentive price exceeds the threshold, the number of user responses increases approximately linearly with the increase of the incentive price. As the incentive price increases further, the response volume gradually reaches saturation and stabilizes due to limitations in users' own adjustment capabilities. Its variation no longer changes significantly with further increases in the incentive price, as shown in the following formula: (2); In the formula, This represents the response volume of user n in time period t, calculated using a consumer psychology model. This represents the maximum response capability of user n in time period t; , Let represent the initial response price and the saturation response price for user n in time period t, respectively.
4. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S2, Based on the concept of uncertainty modeling, this paper characterizes the randomness in user demand response behavior. It assumes that the deviation of user response volume follows a normal distribution, and that its standard deviation is negatively correlated with the incentive price, as detailed below: As the incentive price increases, the uncertainty of user response behavior gradually decreases, and the standard deviation of the response bias decreases accordingly. When the incentive price is below a preset threshold, the volatility of user response increases, and the level of uncertainty rises, as shown in the following formula: (3); In the formula, , Let represent the expected value and standard deviation of the deviation of user n response volume in time period t, respectively.
5. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S3, When constructing the optimal bidding model for aggregators, the objective function is to maximize total profit. Subsidy revenue is calculated based on demand response price and curve similarity. User costs are determined by incentive price and response volume. Intervention costs are only incurred during manual adjustments. Penalty costs are set for non-compliance response volume, as shown in the following formula: (4); In the formula, Q is the total profit of the aggregator, and W... a (t) represents the subsidy revenue that the aggregator receives from the power grid during the response period T; C U (t) is the incentive cost that the aggregator has to pay to the user; C P (t) represents the cost of penalties incurred for failing to meet the standards; The intervention cost of human-machine hybrid decision-making is incurred only when decisions are adjusted manually.
6. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S3, The constraints ensure that the total electricity consumption of users remains unchanged, the bidding price is within the limit, the ranked price meets the response performance indicators, and the degree of manual intervention is limited to the upper limit. The formula is as follows: (5); (6); (7); In the formula, To establish a unified market clearing price; P d This represents the number of bid responses from aggregators. and Let be the per-unit load values for the aggregator and the power grid company, respectively, in time period t. The Euclidean distance is used to measure the similarity between the per-unit load curve of the aggregator and the load baseline published by the power grid company, where d is the Euclidean distance between the market participant and the load baseline. This serves as a similarity indicator between demand response market participants and load lines. The similarity coefficient. This is a subsidy adjustment factor; Human intervention is based on machine decision-making, as shown in the following formula: (8); In the formula The manual adjustment range is limited to 0.2 by default, meaning the adjustment range cannot exceed 20% of the machine's initial value.
7. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S4, The bidding strategy problem is modeled as a Markov game, defining a state space and an action space. The state space includes the load curve, clearing price, and user bias, while the action space includes the bid quantity, bid price, and a reward function based on profit and completion rate, as detailed below: The state space representation of the i-th aggregator at time step j is: (9); In the formula, It is the shape of the load curve published by the power grid at time step j, t j For the current time, This is the market clearing price at this point in time. It is a vector representing the uncertainty deviation between the response of the i-th aggregator and the response of all users aggregated by itself at time step j; The action space of the i-th aggregator at time step j is the bidding strategy under the current circumstances, as shown in the following formula: (10); In the formula, It is the bid volume of the i-th aggregator at time step j. It is its bid price; (11); It is the revenue of the i-th aggregator. It is the i-th yes The subsidy price for time period t has a base value of 3 yuan for peak shaving and 1.2 yuan for valley filling; K t The completion adjustment factor for time period t is linked to the response completion rate; V t This is the actual response capacity for time period t, in kilowatts; Adjustment costs are the compensation paid by the aggregator to users and the cost of equipment depreciation, and are positively correlated with the load adjustment amount; where U is the set of users. user The unit adjustment cost, expressed in yuan / kilowatt, reflects equipment wear and tear or comfort compensation. User The load adjustment amount during time period t, expressed in kilowatts. The target response capacity for time period t is determined by the grid curve command; the penalty factor for insufficient completion is applied only when... Activated at that time.
8. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S5, When designing a hybrid intelligent architecture The human-computer interaction layer triggers manual intervention through alarms for abnormal key indicators; The decision fusion layer sets fusion weights based on alarm levels; The feedback layer updates the machine model parameters, forming a closed-loop optimization.
9. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S6, The decision fusion layer integrates the initial machine decision and the manually adjusted decision to output the final bidding strategy. The fusion algorithm uses a dynamic weight allocation mechanism, as shown in the following formula: (12); In the formula For the final bid amount, For weight fusion.
10. The optimal bidding method for aggregators based on deep reinforcement learning and hybrid intelligent architecture according to claim 1, characterized in that, In step S6, Abnormal scenarios include minor anomalies and severe anomalies, and the specific weighting settings for fusion are as follows: In a typical scenario, the fusion weight is 0.8, in which case the machine is dominant. Mild anomaly: The fusion weight is 0.5, which is a human-machine balance. Severe anomaly: The fusion weight is 0.3, indicating human dominance.