A virtual power plant distributed load response prediction method based on deep federated learning
By optimizing incentive allocation through deep federated learning and reinforcement learning, the problem of poor load response caused by the complexity of enterprise production processes in virtual power plants is solved, achieving a balance between energy-saving benefits and production quality, and improving enterprise participation and scheduling efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU ZHONGKE ZHIXUN TECH CO LTD
- Filing Date
- 2025-06-11
- Publication Date
- 2026-04-24
AI Technical Summary
Existing methods fail to effectively adapt to the complexity of enterprise production processes when coordinating the participation of industrial enterprises in virtual power plants, resulting in poor load response, difficulty in balancing equipment start-up and shutdown losses with energy-saving benefits, and increased complexity of enterprise decision-making and insufficient willingness to participate.
By employing a deep federated learning approach, enterprise data is aggregated with privacy protection to construct a trade-off function between energy-saving benefits and equipment start-up and shutdown losses. Reinforcement learning is used to optimize incentive allocation, and supply and demand balance is monitored in real time to dynamically adjust scheduling instructions. Incentive schemes are further optimized in conjunction with production quality constraints.
It has enabled an effective balance between energy-saving benefits and production quality in virtual power plants, improved dispatch efficiency and enterprise participation, and ensured supply and demand balance and production stability.
Smart Images

Figure CN120875307B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a method for predicting distributed load response in a virtual power plant based on deep federated learning. Background Technology
[0002] Virtual power plants, as a key technology driving energy transition and power system flexibility, achieve power supply and demand balance by aggregating distributed resources, and have significant application value in the industrial sector. Industrial loads, due to their high energy consumption and high controllability, are a crucial component of virtual power plants. However, existing methods for coordinating industrial enterprises' participation in virtual power plants often overlook the complexity of their production processes, leading to poor load response. In particular, current solutions are mostly based on uniform load dispatch models, which are difficult to adapt to the production characteristics of different industrial enterprises, resulting in low response efficiency or insufficient enterprise participation. In practical applications, the constraints of industrial enterprises' production processes on load elasticity become a core challenge. Adjustments to production plans are limited by the uninterrupted nature of key processes, requiring load regulation to be completed within a strict time window; otherwise, it may lead to production interruptions or quality degradation. Due to the rigid constraints of key processes, the balance point between equipment start-up and shutdown losses and energy-saving benefits is difficult to determine precisely, leading enterprises to face a dilemma between economic efficiency and production stability when responding to virtual power plant dispatch. This dynamic change in the balance point further exacerbates the boundary limitations of load regulation depth, requiring enterprises to weigh production quality against response benefits when conducting deep regulation, increasing the complexity of decision-making for participating in virtual power plants. Therefore, designing a dynamically optimized incentive mechanism within the federated learning framework of a virtual power plant to adapt to the uninterrupted constraints of industrial enterprise production processes, balance equipment start-up and shutdown losses with energy-saving benefits, and overcome the limitations of production quality on load regulation depth has become a key issue in improving the participation and response reliability of industrial enterprises. Summary of the Invention
[0003] This invention provides a distributed load response prediction method based on deep federated learning virtual power plants, mainly comprising:
[0004] The system acquires production plan data from industrial enterprises, uses a pre-defined federated learning framework to aggregate the distributed data of the virtual power plant with privacy protection, combines the uninterruptible time window with equipment start-up and shutdown loss parameters, and calculates the load adjustment boundary of each process in the time dimension through a dynamic programming algorithm to obtain the dispatchable load range of the virtual power plant.
[0005] Based on the dispatchable load range of the virtual power plant, a trade-off function between power saving revenue and equipment start-up and shutdown losses is constructed. The loss cost is defined as the sum of equipment restart energy consumption and maintenance costs, and the revenue is the sum of electricity price difference and virtual power plant subsidies. A reinforcement learning algorithm is used to optimize the trade-off strategy and obtain a preliminary incentive allocation scheme.
[0006] If the load adjustment depth of any enterprise in the initial incentive allocation scheme exceeds the production quality limit, the incentive weight of the virtual power plant is adjusted based on the quality constraint threshold. The reinforcement learning model is updated iteratively, and the incentive allocation is recalculated to obtain an optimized incentive scheme that satisfies the quality constraint. If the load adjustment depth of any enterprise in the initial incentive allocation scheme does not exceed the production quality limit, the initial incentive allocation scheme is adopted as the optimized incentive scheme.
[0007] The optimized incentive scheme parameters are distributed to each industrial enterprise node participating in the virtual power plant. The feedback on the enterprises' willingness to respond is collected and judged to see if it meets the preset participation rate threshold, and the response willingness assessment results are obtained.
[0008] If the response willingness assessment result is lower than the preset threshold, the preference features of the enterprise's willingness to participate in the virtual power plant are extracted through the federated learning framework, the incentive allocation is re-optimized, and an incentive mechanism to enhance participation willingness is generated. If the response willingness assessment result is higher than the preset threshold, the optimized incentive scheme is adopted as the incentive mechanism.
[0009] The system monitors the supply and demand balance of the virtual power plant in real time, obtains the grid load demand curve and the actual response data of the enterprise, predicts the short-term supply and demand gap, and determines whether the incentive mechanism needs to be adjusted. If adjustment is required, the system returns to re-optimize the incentive allocation; if no adjustment is required, the system generates a supply and demand balance dispatch instruction.
[0010] In response to supply and demand balance scheduling instructions, the global model in the federated learning framework is updated, the latest response data and predicted gaps of each enterprise are integrated, the load adjustment depth of enterprises is adjusted, a load scheduling plan is formed, and it is issued to each industrial enterprise node. After execution, load adjustment data and production quality data are collected to obtain the final scheduling execution result.
[0011] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0012] This invention discloses a distributed load response prediction method for virtual power plants based on deep federated learning. The method first aggregates distributed industrial enterprise data with privacy protection using a federated learning framework to calculate the dispatchable load range of the virtual power plant. Then, it constructs a trade-off function between energy-saving benefits and equipment start-up and shutdown losses, and optimizes the incentive allocation scheme using reinforcement learning. Next, it uses a deep reinforcement learning model to generate schemes that include equipment status and revenue / cost information, and adjusts them according to production quality constraints. Finally, it distributes incentive schemes through the federated learning framework, collects enterprise response intentions, and optimizes them using game theory algorithms when necessary. This invention also monitors supply and demand balance in real time, predicts short-term gaps, and dynamically adjusts dispatch instructions. This method effectively balances energy-saving benefits and production quality, improving the dispatch efficiency of the virtual power plant and enterprise participation. Attached Figure Description
[0013] Figure 1This is a flowchart of a distributed load response prediction method for a virtual power plant based on deep federated learning, according to the present invention. Detailed Implementation
[0014] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] like Figure 1 This embodiment of a distributed load response prediction method based on deep federated learning virtual power plants may specifically include:
[0016] Step S101: Obtain industrial enterprise production plan data, use a preset federated learning framework to aggregate the distributed data of the virtual power plant with privacy protection, combine the uninterruptible time window and equipment start-up and shutdown loss parameters, and use a dynamic programming algorithm to calculate the load adjustment boundary of each process in the time dimension to obtain the dispatchable load range of the virtual power plant.
[0017] Production plan data from industrial enterprises is acquired, and a production process time sequence diagram is constructed based on equipment type. A federated learning framework is used to perform local gradient calculations on the load data of each enterprise node. While keeping the original data within its domain, only the gradient parameters are transmitted to the central aggregator. The aggregator assigns corresponding weight coefficients based on the data scale of each node and performs a weighted average calculation to generate a virtual power plant aggregated load model. Based on the aggregated load model, the baseline power curves of each process are extracted. A state transition cost matrix is established by combining equipment start-up and shutdown loss parameters. The start-up cost includes preheating energy consumption and mechanical wear conversion values, while the shutdown cost covers cooling process energy consumption and restart preparation time conversion values. By setting an uninterrupted time window constraint for each process, the minimum load value that the equipment must maintain during that time period is determined. A dynamic programming algorithm combined with the state transition cost matrix is used to optimize the load adjustment of each process in the time dimension. The state variable is defined as the operating power Pi(t) of process i at time t, where t represents the time point and i represents the process number. The decision variable is the power adjustment amount ΔPi(t). The state transition equation is:
[0018] Pi(t+1) = Pi(t) + ΔPi(t). Under the conditions of uninterrupted time window constraints and load lower limit, the optimal load adjustment path for each time period is calculated to obtain the upper and lower boundary curves of the adjustable load for each process. The upper and lower boundary curves of the adjustable load for each process are superimposed. Based on the production process time sequence diagram, the upstream and downstream relationships between processes are identified. When the load adjustment of a certain process affects the material supply of downstream processes, the adjustment amount transmission coefficient is calculated according to the material balance constraint, and the adjustable range of downstream processes is adjusted accordingly. The adjusted adjustable ranges of each process are summarized to form the overall dispatchable load range of the virtual power plant in each time period.
[0019] In one possible implementation, acquiring production planning data for industrial enterprises involves real-time monitoring and historical data integration across multiple manufacturing stages. In steel enterprises, key equipment such as blast furnaces, converters, and rolling mills form continuous production processes, each with its specific power load characteristics. The production process timeline diagram reflects the flow sequence and time intervals of materials between different processes. Molten iron produced from the blast furnace must be fed into the converter within a specific temperature range; this timeline constraint directly affects the flexibility of load adjustments for each process. The application of a federated learning framework resolves the contradiction between data privacy protection and collaborative optimization across multiple enterprises. Each enterprise node trains a local model based on its local historical load data. The calculated gradient parameters only contain information about the model update direction and magnitude, without involving the specific values of the original production data. After receiving the gradients from each node, the central aggregator assigns weights based on factors such as enterprise size and data quality. Large enterprises receive higher weights due to their abundant and stable data. This weighted averaging mechanism ensures the accuracy and representativeness of the aggregated load model.
[0020] Specifically, after the aggregated load model is generated, the baseline power curves for each process under different production conditions can be extracted. The power requirements of the rolling mill vary significantly when rolling different specifications of steel; the power requirement for thin plates is approximately 60% of that for thick plates. These differentiated power characteristics form the basis for load regulation. Quantifying the equipment start-up and shutdown loss parameters is crucial. Starting and preheating a large electric arc furnace consumes approximately 2000 kWh of electricity, and frequent start-ups and shutdowns accelerate the thermal fatigue damage of the furnace lining refractory material, shortening the equipment's service life.
[0021] In one embodiment, the setting of the uninterruptible time window is based on the physical constraints of the process flow. Once the continuous casting process begins, it must run continuously for 4-6 hours until the entire furnace of molten steel is poured; interruption will cause the molten steel to solidify and block the equipment. This process characteristic dictates that the lower limit of the load during this period must be maintained above the minimum operating power of the equipment, providing a hard constraint for the subsequent dynamic programming algorithm. The dynamic programming algorithm achieves global optimum through inverse solution in the time dimension. Starting from the end point of the production plan, it progressively calculates the optimal load configuration for each time point. The state transition cost matrix plays a crucial role in this process, quantifying the cost of switching between different load levels, including costs from multiple dimensions such as increased energy consumption, equipment wear and tear, and product quality fluctuations. Under the premise of satisfying the uninterruptible constraint, the algorithm finds the load adjustment path with the minimum total cost.
[0022] It should be noted that the coupling relationship between processes is reflected through material balance and energy conservation. A reduction in upstream blast furnace production directly affects the raw material supply to downstream converters, forcing the converters to correspondingly reduce their production load. This cascading effect is quantitatively described through material flow relationships. When the blast furnace load is reduced by 20%, considering the buffer time for molten iron, the converter needs to adjust its load accordingly after 2 hours. The correlation coefficient reflects the strength and time-delay characteristics of this transmission relationship.
[0023] Preferably, the superposition of the adjustable load boundary curves of each process is not a simple numerical addition. Due to shared resource constraints between processes, such as power line capacity and cooling water system capacity, the actual adjustable range is smaller than the sum of the independent adjustable ranges of each process. By identifying these constraint bottlenecks and reasonably correcting the superposition result, the final virtual power plant dispatchable load range is closer to the actual operating capacity, providing a reliable decision-making basis for power grid dispatch.
[0024] Step S102: Based on the dispatchable load range of the virtual power plant, construct a trade-off function between power saving revenue and equipment start-up and shutdown losses. Define the loss cost as the sum of equipment restart energy consumption and maintenance costs, and the revenue as the sum of electricity price difference and virtual power plant subsidies. Use reinforcement learning algorithm to optimize the trade-off strategy and obtain a preliminary incentive allocation scheme.
[0025] Based on the dispatchable load range of the virtual power plant, the upper and lower limits of load regulation for each time period are extracted. Real-time electricity price data and peak-valley price differences are obtained. Historical operating data is used to statistically analyze the number of equipment start-ups and shutdowns and the corresponding restart energy consumption. Combined with equipment maintenance records, the incremental maintenance costs generated by a single start-up or shutdown are calculated, constructing a trade-off function that includes energy-saving benefits and start-up / shutdown losses. Based on this trade-off function, the loss cost is defined as the sum of equipment restart energy consumption and maintenance costs, and the benefit is defined as the sum of the electricity price difference and virtual power plant subsidies. Backtesting on historical dispatch data verifies the rationality of the coefficients for the energy-saving benefits and start-up / shutdown losses in the trade-off function, determines the value range of each coefficient, and generates a trade-off function for decision evaluation. A reinforcement learning algorithm is used to optimize the decision-making process of the trade-off function. Using the current load level and electricity price information as input, the load regulation amount as output, and the net benefit value as the optimization objective, iterative learning obtains the optimal dispatch decision under different scenarios. The net benefit value generated by each participating enterprise during the execution of the dispatch decision is recorded, obtaining the benefit contribution data of each enterprise. The incentive allocation ratio is calculated based on the revenue contribution data of each enterprise. The response speed index, adjustment accuracy index and historical performance rate data of each enterprise are obtained. The three indicators are weighted according to the preset weight coefficient to obtain the comprehensive score of each enterprise. Based on the comprehensive score and revenue contribution, the incentive amount allocation coefficient of each enterprise is determined, and a preliminary incentive allocation plan including the incentive amount and execution period of each enterprise is output.
[0026] In one possible implementation, the extractable load range of the virtual power plant is based on the aggregated results of the load boundary curves of each process obtained in the early stages. A virtual power plant in a chemical industrial park includes multiple enterprises. During the morning peak hours of 8:00-10:00, the overall adjustable load has an upper limit of 50 MW and a lower limit of 20 MW. This range directly determines the boundary conditions for subsequent economic evaluation. Real-time electricity price data is updated every 15 minutes through the power trading platform interface. Peak-hour electricity prices can reach 1.2 yuan / kWh, while off-peak prices are only 0.3 yuan / kWh. This significant price difference creates economic space for load adjustment. Quantifying equipment start-up and shutdown losses involves cost accounting across multiple dimensions. Large compressor units require 30 minutes of preheating time each time they start, consuming approximately 500 kWh of electricity during this period. Simultaneously, the starting inrush current accelerates the aging of the motor winding insulation. Maintenance records show that the annual maintenance cost of equipment with frequent start-ups and shutdowns is about 40% higher than that of continuously operating equipment. Statistical analysis of this historical data provides a reliable parameter basis for constructing the trade-off function.
[0027] Specifically, the construction process of the trade-off function reflects the dynamic balance between benefits and costs. The energy-saving benefit term includes not only the direct economic benefits from peak-valley electricity price differences but also government subsidies for virtual power plants' capacity and dispatch. For example, a region provides an additional subsidy of 0.1 yuan / kWh to industrial users participating in demand response, a mechanism that significantly enhances enterprises' enthusiasm for participation. The loss cost term comprehensively considers equipment physical losses and opportunity costs. When a production line restarts after a shutdown, in addition to additional energy consumption, fluctuations in product quality may also lead to an increase in the defect rate.
[0028] In one embodiment, backtesting using historical scheduling data revealed the impact of different coefficient values on the decision-making results. Analysis of scheduling records from the past year showed that the simulated decision best matched the actual optimal decision when the coefficient for energy-saving benefits was set to 1.2 and the coefficient for start-stop losses was set to 0.8. This parameter calibration method based on actual operating data ensures that the trade-off function accurately reflects the real economic relationship. The reinforcement learning algorithm demonstrated adaptive learning capabilities during the optimization process. The algorithm uses the load level and current electricity price for each scheduling period as the environmental state, and the selectable load adjustment amounts constitute the action set. Through continuous interaction with the simulated environment, the algorithm gradually learns the optimal adjustment strategy under different electricity price levels.
[0029] For example, when it is predicted that electricity prices will continue to rise in the next two hours, the algorithm tends to reduce the load in advance to avoid equipment damage caused by large adjustments during periods of high electricity prices.
[0030] It should be noted that the calculation of revenue contribution fully considers the differentiated characteristics of each enterprise. One enterprise may have a large adjustment capacity but a slow response time, requiring one hour to complete load adjustment; while another data center has a fast response time, completing adjustment within five minutes, but with limited adjustment capacity. This difference is quantitatively represented through a multi-dimensional evaluation system that includes response speed indicators and adjustment accuracy indicators.
[0031] Preferably, the weighting coefficients of the comprehensive score reflect the power grid's emphasis on different capability dimensions. The response speed weight is set at 0.4, reflecting the power grid's high demand for rapid response capabilities; the regulation accuracy weight is 0.3, ensuring the accurate execution of dispatch instructions; and the historical performance rate weight is 0.3, encouraging enterprises to maintain stable and reliable response performance. The incentive allocation scheme calculated based on this multi-dimensional evaluation system not only embodies the principle of fairness but also fully mobilizes the participation enthusiasm of different types of enterprises, forming a virtuous cycle of incentives.
[0032] Based on load fluctuation data within the dispatchable load range of the virtual power plant, a set of equipment start-up and shutdown thresholds is generated. This set of thresholds, along with the electricity price difference parameter, is input into a pre-trained deep reinforcement learning model. The model outputs a weighted allocation matrix of loss costs and energy-saving benefits, generating equipment status markers for each time period and corresponding revenue-cost differences. The model then determines whether the revenue-cost difference meets the preset net revenue growth condition. If it does, a preliminary incentive allocation scheme is generated, including a mapping table of equipment identifiers and incentive amounts. If it does not meet the condition, the model returns to adjust the trade-off function between loss costs and energy-saving benefits, re-optimizing the trade-off strategy until the net revenue growth condition is met.
[0033] Load fluctuation data for each time period is extracted from the dispatchable load range of the virtual power plant. Statistical analysis is used to determine the time points when the load change rate exceeds a preset threshold. Based on the equipment start-up and shutdown characteristic curves, the corresponding power thresholds are calculated, generating a set of equipment start-up and shutdown thresholds containing time series and power values. Simultaneously, the peak-valley electricity price difference for each time period is obtained to form an electricity price difference parameter sequence. A pre-trained deep reinforcement learning model is used to process the set of equipment start-up and shutdown thresholds and the electricity price difference parameter sequence. Through feature extraction and decision networks within the model, the weight values of loss cost items and energy-saving benefit items under different scenarios are calculated, outputting a weight allocation matrix for each time period. Based on the weight allocation matrix and actual load data and electricity price information, a calculation is performed to obtain equipment status labels and corresponding revenue-cost differences for each time period. The revenue-cost difference is judged time-by-time. If the revenue-cost difference for a certain time period exceeds a preset net revenue growth threshold, the equipment status and corresponding difference for that time period are marked. The equipment identifiers and cumulative revenue values for all time periods that meet the conditions are summarized. The incentive amount for each equipment is determined based on the proportion of cumulative revenue value, generating a preliminary incentive allocation scheme containing a mapping table of equipment identifiers and incentive amounts. If the revenue-cost difference does not reach the preset net revenue growth threshold, the parameters of the tradeoff function are adjusted based on the current weight allocation matrix. The calculation rules of the tradeoff function are changed by modifying the loss cost coefficient and the power saving revenue coefficient. The adjusted tradeoff function is used to regenerate the set of equipment start-up and shutdown thresholds. The new weight allocation matrix and revenue-cost difference are then input into the deep reinforcement learning model again. The process continues until the revenue-cost difference meets the net revenue growth threshold.
[0034] In one possible implementation, load fluctuation data is extracted based on power time-series data collected by a virtual power plant real-time monitoring system. Load data from a virtual power plant in an industrial park over a 24-hour period shows that the load rapidly increased from 30 MW to 45 MW between 6:00 AM and 8:00 AM, a rate of change of 7.5 MW per hour. When the load change rate exceeds a preset threshold of 5 MW / hour, the system identifies a significant load adjustment demand during that period. Equipment start-up and shutdown characteristic curves reflect the starting power requirements of different types of equipment. Large motors can generate 5-7 times their rated power at startup; this characteristic dictates that the setting of start-up and shutdown thresholds must consider instantaneous power surges.
[0035] Specifically, the formation process of the electricity price difference parameter sequence reflects the time-varying characteristics of the electricity market. Weekday electricity prices exhibit a clear peak-valley distribution, with peak prices reaching 1.5 yuan / kWh from 10:00 AM to 12:00 PM, while off-peak prices are only 0.25 yuan / kWh from 2:00 AM to 5:00 AM. This six-fold price difference creates significant economic incentives for load regulation. The equipment start-up and shutdown threshold set includes power thresholds for each time period. For example, a lower start-up threshold of 30 MW is set during periods of rising electricity prices, while a higher shutdown threshold of 40 MW is set during periods of falling electricity prices, forming a hysteresis control characteristic to avoid frequent start-ups and shutdowns. The pre-training process of the deep reinforcement learning model integrates a large amount of historical operating data and market transaction information. The model's feature extraction network identifies multi-dimensional features such as electricity price change patterns, load fluctuation patterns, and equipment operating constraints, while the decision network calculates the optimal weight allocation based on these features. The weight of the loss cost item reflects the importance of equipment protection. When the equipment is aging, the model tends to assign a higher weight to the loss cost to reduce the frequency of start-ups and shutdowns. The weight of the power saving benefit item is positively correlated with the magnitude of the electricity price difference. The larger the price difference, the higher the benefit weight.
[0036] In one embodiment, the weighting matrix exhibits time-varying characteristics. During the early morning hours, the weight for loss costs is 0.7, and the weight for energy-saving revenue is 0.3, because electricity prices are low and equipment is operating at low load. During the midday peak hours, the weighting is adjusted to 0.3 for loss costs and 0.7 for energy-saving revenue, fully utilizing the economic benefits of high electricity prices. Based on this dynamic weighting configuration, the system calculates the difference between revenue and cost for each time period. For example, a chemical company's revenue-cost difference reached 8,000 yuan per hour at 11:00 AM, far exceeding the preset net revenue growth threshold of 5,000 yuan.
[0037] It should be noted that the initial incentive allocation plan follows the principle of contribution. The calculation of cumulative revenue value takes into account the actual response of enterprises at different times; enterprises that respond quickly and execute scheduling instructions accurately receive higher cumulative revenue values. For example, although one enterprise has a large adjustment capacity, its response delay reaches 45 minutes due to process limitations, resulting in a lower cumulative revenue value than a data center that only requires a 10-minute response time. The incentive amount is determined using a cumulative revenue value percentage method, with the total incentive fund pool allocated according to the proportion of each enterprise's cumulative revenue value.
[0038] Preferably, the iterative adjustment mechanism of the trade-off function parameters ensures the system's adaptive optimization capability. When the initially calculated difference between revenue and cost fails to meet the target, the system analyzes the reasons for the insufficient difference. If it is due to excessive loss costs, the loss cost coefficient is reduced from 1.2 to 0.8; if it is due to insufficient calculation of energy-saving benefits, the energy-saving benefit coefficient is increased from 0.9 to 1.3. The adjusted trade-off function re-evaluates the start-up and shutdown decisions for each time period, generating a new set of equipment start-up and shutdown thresholds. This iterative optimization process typically converges within 3-5 rounds, obtaining a balanced solution that protects the equipment while maximizing economic benefits, achieving a dual improvement in the operating efficiency and economy of the virtual power plant.
[0039] Step S103: If the load adjustment depth of any enterprise in the preliminary incentive allocation scheme exceeds the production quality limit, the incentive weight of the virtual power plant is adjusted based on the quality constraint threshold. The reinforcement learning model is iteratively updated, and the incentive allocation is recalculated to obtain an optimized incentive scheme that satisfies the quality constraint. If the load adjustment depth of any enterprise in the preliminary incentive allocation scheme does not exceed the production quality limit, the preliminary incentive allocation scheme is adopted as the optimized incentive scheme.
[0040] The process involves acquiring load adjustment depth data for each enterprise in the preliminary incentive allocation scheme, extracting production quality constraint parameters for each process from enterprise production monitoring records (including allowable fluctuation ranges for key process parameters and product quality qualification standards), and comparing the relationship between load adjustment depth and production quality constraint parameters to determine whether each enterprise's adjustment depth exceeds the allowable range. A quality inspection result table containing enterprise identifiers and over-limit status markers is then generated. For enterprises marked as exceeding limits in the quality inspection result table, the difference between their corresponding quality constraint threshold and actual adjustment depth is extracted. Based on the magnitude of the difference, the reduction ratio of the enterprise's incentive weight is calculated. By reducing the incentive weight of over-limit enterprises and proportionally increasing the incentive weight of non-over-limit enterprises, an incentive weight adjustment matrix that satisfies the quality constraint conditions is formed. The parameter configuration of the reinforcement learning model is updated based on the incentive weight adjustment matrix. By modifying the weight parameters related to incentive calculation in the model, the model automatically considers quality constraint factors when outputting incentive allocation results. After updating the model parameters, an optimized reinforcement learning model incorporating quality constraints is obtained. An optimized reinforcement learning model is used to recalculate the incentive allocation for each enterprise. By inputting the current load data, electricity price information, and incentive weight adjustment matrix, the recalculated incentive allocation result that meets the production quality requirements is obtained. If the load adjustment depth of all enterprises in the result is within the quality limit, the result is output as the optimized incentive scheme. If there are still enterprises in the result that exceed the quality limit, the incentive weight adjustment matrix is returned for further adjustment until all enterprises meet the quality constraints.
[0041] In one possible implementation, the acquisition of load adjustment depth data reflects the actual production flexibility of an enterprise. A certain enterprise's blast furnace is initially required to reduce its load from 100% to 60% in the incentive allocation scheme, implying a 40% adjustment depth. However, the production characteristics of the blast furnace dictate that its temperature must be maintained above 1500 degrees Celsius; excessive load reduction will lead to a drop in furnace temperature, affecting the quality of molten iron. Production monitoring records show that when the load is below 70%, the sulfur content in the molten iron exceeds the standard, failing to meet the raw material requirements of downstream processes. The extraction of production quality constraint parameters involves key indicators from multiple process stages. Chemical enterprises' reactors require specific temperature and pressure conditions; temperature fluctuations exceeding ±5 degrees Celsius or pressure fluctuations exceeding ±0.2 MPa will affect the conversion rate of chemical reactions. Food processing enterprises have even stricter cold chain systems; cold storage temperatures must be maintained between -18 degrees Celsius and -22 degrees Celsius, and any fluctuations outside this range may lead to product spoilage. These quality parameters constitute the hard constraints of load adjustment.
[0042] Specifically, the generation process of the quality inspection result table embodies a multi-dimensional evaluation mechanism. The system compares each enterprise's actual adjustment depth with its quality limit parameters item by item. For example, a textile enterprise's weaving workshop has a load adjustment depth of 35%, while its tension control system's maximum allowable fluctuation is only 20%, so the system marks this enterprise as exceeding the limit. Meanwhile, a cement enterprise's ball mill has a load adjustment depth of 25%, within its allowable 30% range, and is marked as normal. The calculation of the incentive weight adjustment matrix is based on a quantitative analysis of the degree of exceeding the limit. Enterprises exceeding the limit by 15% have their incentive weight reduced to 0.7 times the original value; enterprises exceeding the limit by 30% have their weight reduced to 0.4 times. This gradient adjustment mechanism reflects both the emphasis on production quality and the preservation of enterprises' enthusiasm for participating in scheduling. The weight of enterprises not exceeding the limit is correspondingly increased, ensuring that resources are tilted towards enterprises with good quality control while maintaining the total incentive amount unchanged.
[0043] It's important to note that the parameter update process of the reinforcement learning model involves a redesign of the reward mechanism. The original model primarily considered maximizing economic returns, while the updated model adds a quality penalty term to the reward function. When a firm's load adjustments cause the quality parameter to approach the limit boundary, the model will award a lower reward value even if the economic returns are high. This mechanism guides the model to automatically balance economic benefits and production quality during decision-making.
[0044] Preferably, the application of the optimized reinforcement learning model achieves dynamic equilibrium. After recalculation, a paper manufacturing company reduced its load adjustment depth from the initial 45% to 28%. Although this reduced some economic benefits, it ensured that the tensile strength and whiteness of the paper met national standards. Through multiple rounds of iterative learning, the model mastered the quality sensitivity characteristics of different companies, providing technical support for refined scheduling. The convergence of the iterative optimization process reflects the stability of the algorithm. Typically, after 2-3 rounds of adjustment, the load adjustment depth of all companies converges to within the allowable quality range. In one actual operation, the initial plan had 5 companies exceeding the limit; after the first round of adjustment, this number decreased to 2; and after the second round of adjustment, all companies met the requirements.
[0045] Step S104: The optimized incentive scheme parameters are sent to each industrial enterprise node participating in the virtual power plant. The feedback on the enterprises' willingness to respond is collected and judged to see if it meets the preset participation rate threshold, and the response willingness evaluation result is obtained.
[0046] The incentive amount and adjustment requirement parameters of each enterprise in the optimized incentive scheme are obtained. The scheme data is formatted and encapsulated using the distributed communication mechanism of the federated learning framework. Sensitive incentive amount information is protected for privacy using an encrypted transmission protocol, generating an incentive parameter data packet that can be securely transmitted between enterprise nodes. Based on the incentive parameter data packet, the incentive scheme information is pushed to each industrial enterprise node through the federated learning framework. Each node receives the data packet, parses it locally, and extracts the corresponding incentive amount and load adjustment requirements. Enterprises assess the feasibility of implementation based on their own production plans and equipment status, generating response willingness feedback information containing participation intentions. The response willingness feedback information returned by each enterprise node is collected and aggregated. The proportion of enterprises willing to participate in scheduling to the total number of enterprises is calculated as the enterprise participation rate. Simultaneously, the proportion of the total load adjustment capacity of willing participating enterprises to the total adjustment capacity of the virtual power plant is calculated as the capacity participation rate, resulting in two participation rate indicators. Determine whether the enterprise participation rate and capacity participation rate have simultaneously reached the preset participation rate threshold. If both participation rate indicators meet the threshold requirements, mark the current enterprise response status as meeting the participation conditions, and output the response willingness assessment result containing the list of participating enterprises, the corresponding adjustment capacity, and the participation rate value. If either participation rate indicator fails to reach the threshold, record the type of participation rate that failed to meet the threshold and the specific value in the response willingness assessment result.
[0047] In one possible implementation, the acquisition of the optimized incentive scheme marks a crucial transition from scheme formulation to actual implementation. This scheme includes specific incentive amounts for each enterprise. One enterprise, due to its 50 MW regulation capacity and excellent response record, is allocated an incentive standard of 200 yuan per MW; while another data center, although with a regulation capacity of only 10 MW, receives a differentiated incentive of 300 yuan per MW due to its rapid response capability within 5 minutes. This capability-based incentive design fully reflects the balance between fairness and incentive. The distributed communication mechanism of the federated learning framework plays a dual role in protecting privacy and ensuring efficient transmission at this stage. Traditional centralized schemes may expose the incentive information of all enterprises, leading to the leakage of commercially sensitive information. However, through the federated framework, each enterprise node only receives its own incentive parameters, while information from other enterprises remains encrypted during transmission. The incentive parameter data packets are end-to-end encrypted, ensuring that even if intercepted during network transmission, the specific incentive amount and regulation requirements cannot be deciphered.
[0048] Specifically, the formatted encapsulation process transforms complex incentive schemes into standardized data structures. Each data packet contains key fields such as enterprise identification code, incentive amount, adjustment period, and response time limit. A data packet received by a chemical company shows that from 14:00 to 17:00 the following day, the load needs to be reduced from 80 MW to 60 MW, for which an incentive of 12,000 yuan can be obtained; response confirmation must be completed before 18:00 on the same day. This clear information structure facilitates rapid understanding and decision-making by the enterprise. The local evaluation process at the enterprise node involves collaborative analysis across multiple departments. The production department assesses the impact of the adjustment requirement on output; a cement company calculates that reducing the load by 20 MW will reduce cement production by 100 tons. The finance department compares the incentive gains with the production losses; the company's incentive gain is 4,000 yuan, while the profit loss from 100 tons of cement is approximately 3,000 yuan, resulting in a positive net gain. The equipment department confirms that the relevant equipment can be safely adjusted. Based on these analyses, the company forms a response decision of whether to participate or not.
[0049] It's important to note that the collection of response willingness feedback information reflects decentralization and timeliness. Each enterprise transmits its decision-making results through a secure channel, and the system has a 2-hour response time limit to ensure the timeliness of the scheduling plan. In one actual operation, 16 out of 20 participating enterprises responded within 1 hour, and the remaining 4 also responded within the time limit. This efficient information collection mechanism laid the foundation for subsequent participation rate calculations. The dual-dimensional design of the participation rate indicator ensures the reliability of virtual power plant scheduling. The enterprise participation rate reflects the acceptance of the scheme; when 16 out of 20 enterprises are willing to participate, the participation rate reaches 80%. The capacity participation rate measures the actual dispatchable resources; the total regulating capacity of these 16 enterprises is 180 MW, accounting for 90% of the total virtual power plant capacity of 200 MW. The two dimensions complement each other, avoiding the one-sidedness of focusing solely on quantity while ignoring capacity or focusing solely on capacity while ignoring coverage.
[0050] Preferably, the preset participation rate thresholds take into account both the minimum requirements of grid dispatch and economic balance. The enterprise participation rate threshold is typically set at 70% to ensure that most enterprises accept the scheme; the capacity participation rate threshold is set at 85% to guarantee sufficient regulation resources. When an evaluation result shows that the enterprise participation rate is only 65%, even if the capacity participation rate reaches 88%, the system still marks it as unsatisfactory, indicating that the incentive scheme needs to be optimized to improve enterprise acceptance. This strict dual standard ensures the stability and sustainability of the virtual power plant operation.
[0051] Step S105: If the response willingness assessment result is lower than the preset threshold, the preference features of the enterprise's willingness feedback virtual power plant are extracted through the federated learning framework, the incentive allocation is re-optimized, and an incentive mechanism to enhance participation willingness is generated. If the response willingness assessment result is higher than the preset threshold, the optimized incentive scheme is adopted as the incentive mechanism.
[0052] Based on the comparison between the participation rate values in the response willingness assessment results and preset thresholds, if either the enterprise quantity participation rate or capacity participation rate is lower than the corresponding threshold, the incentive expectation value and acceptable adjustment period data of the enterprises that refuse to participate are extracted from the enterprise feedback information. A federated learning framework is used to perform feature analysis on the feedback data of each enterprise, resulting in an enterprise preference feature vector containing incentive amount sensitivity, response period preference, and load adjustment constraints. A multi-party participation payoff matrix is constructed based on the enterprise preference feature vector, where rows represent different enterprises, columns represent different incentive amount levels, and matrix element values are the net payoff values of each enterprise under the corresponding incentive amount. By combining incentive amount sensitivity with enterprise load adjustment costs, the values of each element are calculated to generate a payoff matrix reflecting the enterprise incentive response relationship. A game theory algorithm is used to solve the payoff matrix for equilibrium, iteratively calculating to find an incentive allocation state that all enterprises are willing to accept. In this state, unilateral changes in decision by any enterprise will not result in higher payoffs, while simultaneously satisfying the total incentive limit of the virtual power plant, thus obtaining a new incentive allocation scheme that balances the interests of all parties. Based on the new incentive allocation scheme, an incentive mechanism is generated that includes the incentive amount and execution period for each enterprise. If the participation rate indicators in the response willingness assessment results are all higher than the preset threshold, the current optimized incentive scheme is directly used as the final incentive mechanism, and a complete incentive mechanism file is output to guide the actual operation of the virtual power plant.
[0053] In one possible implementation, the assessment of response willingness reflects the management philosophy of precise policy implementation. When the participation rate of enterprises reaches only 65%, below the preset threshold of 70%, it indicates that the existing incentive plan has failed to fully mobilize the enthusiasm of enterprises. A case study from a chemical industrial park shows that 7 out of 20 enterprises refused to participate, mainly due to the low incentive amount and conflicting adjustment periods. In this case, simply increasing the overall incentive level is not the optimal solution; a deeper analysis of the specific demands of each enterprise is needed. The extraction process of enterprise preference feature vectors reveals the essence of differentiated needs. A steel company reported that its expected incentive amount was 250 yuan per megawatt-hour, while the existing plan only provided 180 yuan; a pharmaceutical company stated that 2-5 pm is a critical production period, making load adjustment impossible. The federated learning framework identifies three key preference dimensions through feature analysis of these feedback data: incentive sensitivity reflects the degree to which enterprises value economic compensation, and enterprises with high sensitivity are usually in industries with a high proportion of energy costs; response time preference reflects the time constraints of production processes, and enterprises with continuous production tend to adjust during shift changes; and load adjustment constraints are determined by equipment characteristics and product quality requirements.
[0054] Specifically, the construction of the payoff matrix transforms abstract preferences into calculable numerical relationships. Each element of the matrix represents the net payoff of a specific firm at a specific incentive level, which equals the incentive revenue minus the production adjustment cost. For example, a textile firm with an incentive of 200 yuan / MWh can earn 2000 yuan by adjusting 10 MW, but incurs 1500 yuan in production adjustment costs, resulting in a net payoff of 500 yuan. With an incentive of 250 yuan / MWh, the net payoff increases to 1000 yuan. This quantification method allows for the comparison and balancing of the interests of different firms within the same framework. The application of game theory algorithms solves the problem of coordinating the interests of multiple parties. In the virtual power plant scenario, firms are both partners and competitors for resources. The algorithm iteratively calculates and finds an equilibrium state where each firm receives a relatively satisfactory incentive allocation. In one calculation, the initial scheme gave large firms a uniform incentive standard, leading to low participation from small, fast-response firms. After three rounds of iterative adjustments, a differentiated incentive pattern was formed: large firms received stable but moderate incentives, while small firms received higher unit incentives due to their flexibility.
[0055] It should be noted that the total incentive limit ensured the economic feasibility of the plan. The virtual power plant operator set a daily incentive budget cap of 1 million yuan, and the game theory algorithm optimized the allocation under this constraint. Through clever incentive design, a higher participation rate was achieved with the same budget: transferring some incentives from low-sensitivity enterprises to high-sensitivity enterprises resulted in little change in the former's willingness to participate, while the latter's willingness to participate increased significantly.
[0056] Preferably, the final incentive mechanism reflects the results of dynamic optimization. The new scheme abandons a one-size-fits-all incentive standard, instead forming a tiered and categorized incentive system. Rapid response enterprises receive an incentive of 280-320 yuan per megawatt-hour, conventional response enterprises 200-240 yuan, and basic response enterprises 150-180 yuan. Simultaneously, floating coefficients are set for different time periods, with incentives increasing by 20% during periods of urgent grid demand. This refined incentive mechanism increases the participation rate from 65% to 85% and the capacity participation rate from 82% to 92%, effectively matching enterprise intentions with grid demand and laying a solid foundation for the stable operation of virtual power plants.
[0057] Step S106: Monitor the supply and demand balance of the virtual power plant in real time, obtain the grid load demand curve and the actual response data of the enterprise, predict the short-term supply and demand gap, and determine whether the incentive mechanism needs to be adjusted. If adjustment is required, return to re-optimize the incentive allocation; if no adjustment is required, generate a supply and demand balance dispatch instruction.
[0058] Based on the incentive standards and execution periods determined in the incentive mechanism for each enterprise, the current load demand value of the power grid and the predicted load curve for future periods are obtained through the data acquisition module. Simultaneously, actual load regulation is collected from each enterprise node, and the real-time deviation between the total adjustable capacity of the virtual power plant and the power grid demand is calculated. A time-series dataset is constructed based on the real-time deviation value and historical response data. An autoregressive moving average model is used to predict the supply-demand gap within a preset future period. By analyzing the trend of load demand changes and the fluctuation pattern of enterprise response capabilities, short-term supply-demand forecast results, including the supply-demand gap value and its duration, are obtained. Threshold judgments are applied to the supply-demand gap value and duration in the short-term supply-demand forecast results. If the gap value exceeds a preset proportion of the total adjustable capacity of the virtual power plant or the duration exceeds a preset threshold, the current state is marked as requiring adjustment, and the incentive mechanism is recalculated for optimization. If both the gap value and duration are within the preset thresholds, a scheduling scheme is generated based on the current supply-demand state. Based on the current supply-demand deviation and the adjustable capacity of each enterprise, a list of enterprises participating in the scheduling is determined in descending order of incentive amount. A specific load increase or decrease is assigned to each enterprise, and a supply-demand balance scheduling instruction containing enterprise identifier, load adjustment amount and execution time is generated and sent to each enterprise node for execution through a distributed communication mechanism.
[0059] In one possible implementation, the execution monitoring of the incentive mechanism embodies the core concept of dynamic management. The data acquisition module updates the grid load demand data every 5 minutes. At 10:00 AM on a certain weekday, the system detected a grid load demand of 850 MW, while the virtual power plant's current adjustable capacity was only 800 MW, creating a supply-demand gap of 50 MW. The calculation of this real-time deviation value not only considers the absolute value but also the load change trend. If the load is rising rapidly, even if the current gap is small, advance response preparation is necessary. The collection of actual load adjustment by enterprises reflects the deviation between the execution effect and the plan. One enterprise promised to adjust 30 MW, but due to production line switchover delays, only 25 MW was actually adjusted. This 5 MW deviation directly affects the overall supply-demand balance. By comparing the promised and actual values of each enterprise, the adjustable capacity data is dynamically updated, providing an accurate basis for subsequent forecasting.
[0060] Specifically, the construction of the time series dataset integrates multi-dimensional information. The autoregressive moving average model utilizes the autocorrelation and moving average characteristics of historical data for prediction. Model inputs include load demand data from the past 24 hours, enterprise response records, and weather factors. On a hot summer day, based on data from similar historical dates, the model predicted that load would continue to climb over the next two hours, potentially reaching a peak of 920 MW, while adjustable capacity might drop to 780 MW due to production restrictions at some enterprises, resulting in a predicted shortfall of 140 MW. Analysis of the duration of the supply-demand gap reveals the urgency of the problem. Temporary supply-demand imbalances can be resolved through rapid enterprise response, but persistent gaps require systemic adjustments. One forecast indicated a two-hour shortfall of over 100 MW between 2 PM and 4 PM; in such cases, existing incentive mechanisms alone are insufficient, necessitating an incentive optimization process.
[0061] It's important to note that the dual criteria for threshold judgment ensure the rationality of the decision. The gap value threshold is typically set at 15% of the total regulation capacity; for a virtual power plant with a total capacity of 200 MW, the threshold would be 30 MW. The duration threshold is set at 1 hour to avoid frequent adjustments due to short-term fluctuations. When the predicted gap of 140 MW exceeds the 30 MW threshold and lasts for 2 hours, exceeding the 1-hour threshold, the system determines that the incentive mechanism needs to be re-optimized, potentially requiring an increase in incentive standards for certain key enterprises to enhance their willingness to participate. The incentive priority ranking mechanism reflects the principle of economic efficiency. When resources are limited, resources with lower unit costs are prioritized. While a data center may have a high incentive of 300 yuan per MWh, its 10 MW rapid response capability is irreplaceable in emergencies; whereas an enterprise with an incentive of only 150 yuan / MWh can provide 40 MW of regulation capacity, and should be prioritized in regular scheduling.
[0062] Preferably, the generation of supply and demand balance dispatch instructions enables precise regulation. Based on a real-time deviation of 50 MW, three enterprises are selected to participate in the regulation: the cement enterprise reduces its load by 20 MW, the chemical enterprise by 20 MW, and the steel enterprise by 10 MW. Each instruction contains specific execution parameters, such as "Enterprise A, starting at 10:15, will reduce its load from 80 MW to 60 MW within 30 minutes." This specific instruction ensures operability and, through a distributed communication mechanism, is transmitted in real time, allowing each enterprise to respond immediately, achieving dynamic supply and demand balance and ensuring the stable operation of the power grid.
[0063] Step S107: For the supply and demand balance scheduling instruction, update the global model in the federated learning framework, integrate the latest response data and predicted gaps of each enterprise, adjust the load adjustment depth of enterprises, form a load scheduling plan, issue it to each industrial enterprise node, collect the load adjustment data and production quality data after execution, and obtain the final scheduling execution result.
[0064] To address the load adjustment requirements of each enterprise in the supply-demand balance scheduling command, the latest response data of the enterprises, including actual adjustment amount and execution deviation rate, is extracted. Combined with the current predicted gap value, the weight parameters in the global model are updated using the gradient aggregation method of the federated learning framework. The execution characteristics of each enterprise are integrated into the model, resulting in an updated global model reflecting the latest response capabilities. Based on the output parameters of the updated global model, the load adjustment depth of each enterprise is recalculated. The adjustment amount is corrected according to the enterprise's historical execution deviation rate. Enterprises with an execution deviation rate below a preset threshold maintain their original adjustment depth, while enterprises with an execution deviation rate above the preset threshold increase their adjustment amount proportionally to the deviation. This forms a load scheduling scheme that includes the corrected load target value and execution time requirements for each enterprise. The load scheduling scheme is distributed to each industrial enterprise node through the distributed communication channel of the federated learning framework. Enterprise nodes execute load adjustments according to the scheme and upload execution data in real time. The collected data includes actual load change curves and production quality parameters, including product qualification rate and process stability indicators. The deviation rate between the actual response and the target adjustment is calculated based on the collected actual load change curves. At the same time, the production quality parameters are checked to see if they are within the preset range. If the deviation rate is lower than the response efficiency threshold and the production quality parameters are qualified, it is determined that the virtual power plant response efficiency requirements are met. The dispatch execution results containing the compliance status of each enterprise are output. If the requirements are not met, the non-compliant enterprises and their deviation data are marked in the dispatch execution results.
[0065] In one possible implementation, the execution feedback mechanism of supply and demand balancing scheduling instructions embodies the concept of closed-loop control. When a company receives a scheduling instruction to reduce load by 30 MW, due to fluctuations in blast furnace operating conditions, only 27 MW of load reduction is actually implemented, resulting in an execution deviation of 3 MW. This deviation data is collected through a real-time monitoring system and becomes a key input for optimizing the global model. The calculation of the execution deviation rate considers not only the absolute value but also the response time. If the company completes 90% of the adjustment within the specified 15 minutes, its response quality can still be considered good. The gradient aggregation method of the federated learning framework plays a core role in model updates. Each enterprise node calculates gradient information based on local execution data; these gradients reflect the difference between model predictions and actual execution. After collecting the gradients from all nodes, the central server aggregates them using a weighted average method, with weights determined based on the enterprise's data quality and historical reliability. A certain chemical enterprise receives a weight coefficient of 0.15 due to its stable execution record, while newly added enterprises have a weight of only 0.05. This differentiated processing ensures the robustness of model updates.
[0066] Specifically, the parameter adjustments after the global model update directly affect load allocation decisions. The model includes the response speed coefficient, execution accuracy coefficient, and equipment constraint parameters for each enterprise. For example, the execution accuracy coefficient of a textile enterprise was updated from 0.85 to 0.82, reflecting a recent increase in its execution deviation. Based on this change, the system will increase redundancy when calculating the enterprise's regulation depth; the original allocation of a 20 MW regulation task will be adjusted to 22 MW to compensate for possible execution deviations. The process of correcting the load regulation depth reflects the concept of refined management. For high-performing enterprises with a consistently low execution deviation rate (below 5%), the system maintains their original regulation depth and may even appropriately increase scheduling tasks; while for enterprises with a deviation rate exceeding 10%, the regulation amount needs to be increased proportionally. For instance, a cement enterprise with a historical average deviation rate of 12% automatically increased its 50 MW regulation task to 56 MW to ensure that the expected target is achieved after actual execution.
[0067] It is important to note that monitoring production quality parameters constitutes a crucial constraint on scheduling execution. Product qualification rate is the most direct quality indicator. For example, after implementing load adjustment, the temperature stability index of a key process in a pharmaceutical company increased from ±0.5 degrees Celsius to ±0.8 degrees Celsius. While still within the allowable range, this was close to the upper limit of quality control. Process stability indicators are measured by the standard deviation of process parameters, reflecting the degree of fluctuation in the production process. Analysis of actual load change curves reveals the company's true response characteristics. Ideally, the load should transition from the initial value to the target value along a linear path, but actual curves often exhibit step-like or oscillating characteristics. One company's load adjustment curve showed a 5-minute delay at the beginning of the adjustment, followed by a rapid decrease, and a slight overshoot before stabilizing near the target value. This response characteristic was recorded and used to optimize subsequent scheduling strategies.
[0068] Preferably, the comprehensive evaluation mechanism for scheduling execution results ensures continuous system improvement. The response efficiency threshold is typically set at a deviation rate of 8%, which guarantees the basic achievement of scheduling objectives while providing enterprises with reasonable execution flexibility. In one scheduling operation, 13 out of 15 participating enterprises met the standards, while 2 exceeded the deviation limit due to equipment failure. The system not only recorded the specific deviation data of the non-compliant enterprises but also analyzed the causes of the deviations, providing a basis for the next round of scheduling optimization. This data-driven continuous optimization mechanism enables the overall response efficiency of the virtual power plant to steadily improve, effectively supporting the flexible scheduling needs of the power grid.
[0069] Based on the real-time load fluctuation characteristics and predicted gaps in the enterprise response data, adjustment gradient parameters are generated. The dynamic aggregation window is adjusted according to the adjustment gradient parameters. The differences in load adjustment depth among multiple enterprises are integrated to generate power allocation instructions for each node, which are then sent to the edge computing units of each industrial enterprise node.
[0070] Based on the real-time load fluctuation characteristics in enterprise response data, the maximum deviation value and change cycle of load changes are extracted. Combined with the predicted supply-demand gap value, the gap compensation coefficient is obtained by calculating the ratio of the actual total load adjustment to the gap amount. This coefficient reflects the degree to which the current load adjustment covers the gap, generating gap compensation coefficient data for subsequent adjustments. Adjustment gradient parameters are calculated based on the gap compensation coefficient data. When the compensation coefficient is lower than the preset compensation target value, the gradient parameter value is increased to accelerate the adjustment speed; when the compensation coefficient is close to the compensation target value, the gradient parameter value is decreased to avoid overshoot. The duration of the dynamic aggregation window is determined by multiplying the gradient parameter by the federated learning update cycle, obtaining dynamic aggregation window parameters adapted to the current adjustment needs. The dynamic aggregation window parameters are used to control the model synchronization frequency of the federated learning framework. When the window duration is short, the synchronization frequency is increased to respond quickly to changes; when the window duration is long, the synchronization frequency is decreased to reduce communication overhead. Based on the calculation results of the synchronized federated learning global model, the differences in load adjustment depth among enterprises are integrated. Specific power adjustment amounts are allocated to each node according to the maximum adjustment capacity and historical response time of each node, generating power allocation instructions containing target power values and execution times. Power allocation commands are sent to the edge computing units of each industrial enterprise node through a distributed communication network. The edge computing units control local devices to complete load adjustment and collect actual adjustment data. Based on the deviation between the actual adjustment amount and the target value reported by each node, the enterprise response capability coefficient and adjustment accuracy coefficient in the global model are updated to form optimized global model parameters for use in the next round of federated learning training.
[0071] In one possible implementation, the extraction of real-time load fluctuation characteristics reveals the dynamic patterns of enterprise electricity consumption behavior. A certain enterprise's electric arc furnace exhibits periodic load fluctuations during the smelting process, with each smelting cycle lasting approximately 45 minutes, and the load varying between 80 and 120 MW, with a maximum deviation of 40 MW. This fluctuation characteristic interacts with the power grid's supply-demand gap. When the predicted gap is 50 MW, the enterprise's natural fluctuations can cover 80% of the gap, resulting in a calculated gap compensation coefficient of 0.8. This coefficient directly reflects the degree to which the enterprise's current regulation capacity meets system demand. The application of the gap compensation coefficient embodies the concept of adaptive control. When the coefficient is below 0.6, it indicates that the existing regulation is insufficient, and the system needs to accelerate its response; when the coefficient is close to 1.0, it indicates that the regulation is close to the target, and the adjustment should be slowed down to avoid over-response. The regulation gradient parameter changes dynamically according to this principle; when the compensation coefficient is 0.4, the gradient parameter is set to 2.5, while when the compensation coefficient is 0.9, the gradient parameter drops to 0.5. This nonlinear mapping relationship ensures a reasonable response of the system under different states.
[0072] Specifically, the calculation of the dynamic aggregation window incorporates considerations from multiple time dimensions. The federated learning update cycle is typically 5 minutes. When the gradient parameter is 2.0, the dynamic aggregation window is calculated every 10 minutes; when the gradient parameter drops to 0.5, the window extends to 20 minutes. The rationale for this design is that frequent synchronization is needed for rapid adjustments in emergency situations, while the synchronization frequency can be reduced during stable operation to save communication resources. In a chemical industrial park, during periods of rapid load increases, the aggregation window was shortened to 8 minutes to ensure timely response; during periods of stable load at night, the window was extended to 25 minutes, reducing unnecessary communication overhead. The control mechanism of model synchronization frequency directly affects the system's response performance and communication efficiency. High-frequency synchronization means that each enterprise node must upload its local model parameters every few minutes, and the central server quickly aggregates and distributes updates, which is particularly important when the load changes drastically. In one instance, a power grid failure caused a sharp increase in load; the system increased the synchronization frequency from every 20 minutes to every 5 minutes, enabling scheduling decisions to quickly adapt to the new situation.
[0073] It should be noted that the calculation results of the federated learning global model provide a scientific basis for power allocation. The model output includes each enterprise's predicted response capability, current adjustable capacity, and execution reliability score. A textile enterprise, although having a maximum adjustment capacity of only 30 MW, has a historical average response time of only 8 minutes and a reliability score of 0.92, thus receiving a higher allocation priority in emergency scheduling. A cement enterprise, with an adjustment capacity of 60 MW, has a response time of 25 minutes, making it more suitable for planned scheduling. The application of edge computing units enables rapid localized execution of scheduling commands. These computing devices deployed on-site are directly connected to the production control system, immediately converting power allocation commands into specific equipment control signals upon receipt. For example, upon receiving a 20 MW load reduction command, one enterprise's edge computing unit automatically calculates the adjustment ratio for each production line, prioritizing the adjustment of auxiliary system loads to ensure the stable operation of the main production line.
[0074] Preferably, the global model parameter update mechanism ensures continuous system optimization. The enterprise responsiveness coefficient is dynamically adjusted based on actual execution performance. For example, an enterprise with an average deviation rate of only 3% across five consecutive scheduling operations saw its responsiveness coefficient increase from 0.85 to 0.95; while another enterprise, due to aging equipment leading to increased execution deviation, had its coefficient adjusted downwards. The adjustment precision coefficient reflects the enterprise's control fineness and is calculated by comparing the standard deviation between the target value and the actual value. This parameter update mechanism based on actual data enables the federated learning model to accurately characterize the true capabilities of each enterprise, laying a solid foundation for the next round of scheduling optimization.
[0075] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A distributed load response prediction method based on deep federated learning virtual power plant, characterized in that, The method includes: The system acquires production plan data from industrial enterprises, aggregates distributed data with privacy protection through a federated learning framework, combines uninterrupted time windows and equipment start-up and shutdown loss parameters, and calculates the load adjustment boundary of each process in the time dimension through a dynamic programming algorithm to determine the dispatchable load range of the virtual power plant. A preliminary incentive allocation scheme is generated based on the dispatchable load range. This scheme is generated in one of two ways: Method 1 involves extracting the upper and lower limits of load regulation for each time period, obtaining real-time electricity price data and the peak-valley price difference, statistically analyzing the number of equipment start-ups and shutdowns and the energy consumption during restart, calculating the incremental maintenance cost per start-up and shutdown in conjunction with maintenance records, constructing a trade-off function that includes energy-saving benefits and start-up / shutdown losses, defining the loss cost as the sum of equipment restart energy consumption and maintenance costs, and the benefit as the sum of the electricity price difference and virtual power plant subsidies, determining the range of coefficient values, and using a reinforcement learning algorithm to optimize the trade-off function, with the current load level and electricity price information as input and the load regulation amount as output. The first method involves extracting load fluctuation data to generate a set of equipment start-up and shutdown thresholds. This is then combined with electricity price difference parameters and input into a preset deep reinforcement learning model. The model outputs a weighted distribution matrix of loss cost and energy-saving revenue. The model calculates the equipment status markers and revenue-cost differences for each time period and determines whether the revenue-cost differences meet the net revenue growth conditions. If they do, a preliminary incentive allocation scheme containing a mapping table of equipment identifiers and incentive amounts is generated. If not, the set of equipment start-up and shutdown thresholds is regenerated, input into the deep reinforcement learning model, and iteratively calculated until the net revenue growth conditions are met. Determine whether the load adjustment depth of each enterprise in the preliminary incentive allocation scheme exceeds the production quality limit. If it does, adjust the incentive weight based on the quality constraint threshold to generate an optimized incentive scheme that meets the quality constraint. If it does not exceed the limit, use the preliminary incentive allocation scheme as the optimized incentive scheme. The optimized incentive scheme is obtained, and the incentive scheme parameters are distributed to each industrial enterprise node through the federated learning framework. Feedback on the enterprises' willingness to respond is collected, it is determined whether the preset participation rate threshold is met, and a response willingness evaluation result is generated. If the response willingness assessment result does not meet the preset participation rate threshold, the incentive allocation will be re-optimized and an incentive mechanism will be generated. If it meets the threshold, the optimized incentive scheme will be used as the incentive mechanism. Based on the incentive mechanism, real-time data on grid load demand and actual enterprise response are obtained, and short-term supply and demand gaps are predicted using time series analysis. Supply and demand balance scheduling instructions are then generated based on the predicted short-term supply and demand gaps. Based on the supply and demand balance scheduling instructions, the global model of the federated learning framework is updated to generate a load scheduling scheme.
2. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, The process of acquiring industrial enterprise production plan data, aggregating distributed data with privacy protection using a federated learning framework, combining uninterrupted time windows and equipment start-up and shutdown loss parameters, and calculating the load adjustment boundaries of each process in the time dimension using a dynamic programming algorithm to determine the dispatchable load range of the virtual power plant includes: Acquire industrial enterprise production plan data, construct production process time sequence diagram, perform local gradient calculation on the load data of each enterprise node through federated learning framework, transmit only gradient parameters to the preset central aggregator, the aggregator allocates weight coefficients according to the data scale, performs weighted average calculation, and generates aggregated load model. Based on the aggregated load model, the baseline power curves of each process are extracted, and a state transition cost matrix is established in combination with the equipment start-up and shutdown loss parameters. Uninterrupted time window constraints are set, and the lower limit of the load for each process is determined. Dynamic programming algorithm is used to optimize the load regulation of each process. The state variable is defined as the process operating power, and the decision variable is the power adjustment amount. The optimal load adjustment path for each time period is calculated to form the overall dispatchable load range of the virtual power plant.
3. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, The step of determining whether the load adjustment depth of each enterprise in the preliminary incentive allocation scheme exceeds the production quality limit, and if so, adjusting the incentive weights based on the quality constraint threshold to generate an optimized incentive scheme that satisfies the quality constraint, includes: For enterprises that exceed production quality limits, calculate the difference between the quality constraint threshold and the adjustment depth, determine the reduction ratio of incentive weights, and adjust the incentive weights to form an incentive weight adjustment matrix; The reinforcement learning model parameters are updated based on the incentive weight adjustment matrix, the incentive allocation is recalculated, and an optimized incentive scheme that meets the quality constraints is generated.
4. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, The process of obtaining the optimized incentive scheme involves distributing the incentive scheme parameters to each industrial enterprise node through the distributed gradient update mechanism of the federated learning framework, collecting feedback on enterprise response willingness, determining whether a preset participation rate threshold is met, and generating a response willingness evaluation result, including: Obtain the incentive amount and adjustment requirement parameters in the optimized incentive scheme, and generate an incentive parameter data packet using an encrypted transmission protocol through a federated learning framework; The incentive parameter data package is pushed to each enterprise node through the federated learning framework, and the response willingness feedback information is generated by parsing. Collect the feedback information on response willingness, calculate the participation rate of enterprise number and the participation rate of capacity, determine whether the preset participation rate threshold is met, and generate response willingness assessment results.
5. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, If the response willingness assessment result does not meet the preset participation rate threshold, the incentive allocation will be re-optimized to generate an incentive mechanism, including: Compare the response willingness assessment results with the preset participation rate threshold. If the threshold is not met, extract the incentive expectation value and acceptable adjustment period data of the enterprises that refuse to participate, and generate an enterprise preference feature vector. A payoff matrix is constructed based on the aforementioned preference feature vectors, and a game theory algorithm is used to solve for the equilibrium state to generate a new incentive allocation scheme. An incentive mechanism containing incentive amounts and execution periods is generated based on the new incentive allocation scheme.
6. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, The process involves acquiring real-time grid load demand and actual enterprise response data based on the incentive mechanism, predicting short-term supply-demand gaps using time series analysis, and generating supply-demand balance dispatch instructions, including: Based on the incentive mechanism, obtain grid load demand and actual enterprise response data, and calculate the real-time deviation value; Based on the aforementioned deviation value and historical response data, time series analysis is used to predict the short-term supply and demand gap. Determine whether the supply-demand gap exceeds a preset threshold. If it does, re-optimize the incentive allocation. If it does not, generate a supply-demand balance scheduling instruction.
7. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, The step of updating the global model of the federated learning framework and generating a load scheduling scheme according to the supply and demand balance scheduling instruction includes: Extract the enterprise load adjustment requirements from the supply and demand balance scheduling instructions, integrate the latest response data and the predicted gap, and update the global model weight parameters; Based on the global model, the load adjustment depth is recalculated to generate a load scheduling scheme.
8. The distributed load response prediction method based on deep federated learning virtual power plant according to claim 1, characterized in that, The update of the global model of the federated learning framework includes: Based on enterprise response data, load fluctuation characteristics and predicted gaps are extracted, gap compensation coefficients are calculated, and adjustment gradient parameters are generated. The dynamic aggregation window is adjusted based on the aforementioned adjustment gradient parameters to obtain dynamic aggregation window parameters that adapt to the current adjustment requirements. The model synchronization frequency of the federated learning framework is controlled by dynamic aggregation window parameters; Based on the calculation results of the synchronized federated learning global model, the load regulation depth differences of each enterprise are integrated, and a specific power regulation amount is allocated to each enterprise according to its maximum regulation capacity and historical response time, generating a power allocation instruction. Power allocation commands are sent to the edge computing units of various industrial enterprises through a distributed communication network.
Citation Information
Patent Citations
Virtual power plant industrial load optimization scheduling method considering load characteristics
CN117578490A
Virtual power plant resource scheduling method based on federated learning
CN118677031A