A park virtual power plant resource coordination scheduling and demand response optimization method and system

By employing a multi-agent collaborative scheduling and a three-layer fusion optimization architecture, the problems of weak adaptability of the global optimization model and insufficient professionalism of intelligent decision-making in the virtual power plant of the park are solved, thus achieving efficient and stable collaborative scheduling of resources and optimization of demand response in the virtual power plant of the park.

CN122371209APending Publication Date: 2026-07-10SHANDONG ZHENGCHEN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG ZHENGCHEN TECH CO LTD
Filing Date
2026-03-23
Publication Date
2026-07-10

Smart Images

  • Figure CN122371209A_ABST
    Figure CN122371209A_ABST
Patent Text Reader

Abstract

This invention discloses a resource collaborative scheduling and demand response optimization system and method for virtual power plants in industrial parks. It achieves global collaborative scheduling through a two-layer Actor-Critic network architecture, and uses a fusion of particle swarm optimization and fuzzy control to achieve grid-connected stability control. A three-layer fusion optimization architecture is constructed to complete the continuous amplitude calculation of demand response and the generation of equipment-level instructions. A unified algorithm support is provided by an intelligent optimization module. This invention effectively solves the problems of weak adaptability of existing global optimization models for virtual power plants in industrial parks, low level of intelligent decision-making in specialized sub-scenarios, and difficulty in coordinating multi-timescale scheduling and refined execution. It can achieve robust optimization and refined control in uncertain scenarios, balancing system economy, stability, and operational reliability. It is suitable for virtual power plants in industrial parks to efficiently participate in the power market and resource collaborative scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of smart grids and power systems, specifically relating to a method and system for resource collaborative scheduling and demand response optimization of virtual power plants in industrial parks. Background Technology

[0002] In existing technologies, intelligent control of virtual power plants typically relies on the aggregation and coordinated regulation of various heterogeneous resources such as distributed power sources, energy storage, and controllable loads. This is achieved by constructing optimization decision models and combining them with predictive information to perform scheduling calculations, thereby realizing multi-objective optimizations such as economy, stability, and reliability. However, existing coordinated control methods still have significant limitations in addressing high-dimensional nonlinear coupling constraints, multi-timescale coordinated decision-making, and dynamic responses to electricity market incentive signals.

[0003] Traditional global optimization methods often employ linear programming models, relying on deterministic forecast data and linearized constraints. They fail to adequately characterize the random fluctuations in renewable energy output, market electricity prices, and the nonlinear characteristics of equipment operation, resulting in weak robustness of dispatch schemes under uncertain scenarios and insufficient global optimization adaptability. While multi-agent-based control methods can achieve distributed decision-making, their action space and reward function designs are relatively general and do not fully integrate the strong constraints of grid-connected stability control and demand response execution. This leads to problems such as low policy learning efficiency and insufficient control precision, making it difficult to quickly generate equipment-level fine-grained instructions that meet actual operational needs.

[0004] It is evident that existing collaborative scheduling technologies for virtual power plants in industrial parks generally suffer from problems such as weak adaptability of the global optimization model, insufficient specialization of intelligent decision-making for specific control sub-problems, and difficulty in coordinating multi-timescale scheduling and refined execution. These issues fail to fully meet the actual needs of efficient, stable, and economical operation of virtual power plants in industrial parks.

[0005] In view of this, it is very necessary to provide a method and system for resource collaborative scheduling and demand response optimization of virtual power plants in industrial parks to solve the above-mentioned defects in the prior art. Summary of the Invention

[0006] The purpose of this invention is to address the shortcomings of existing technologies, such as weak adaptability of global optimization models, insufficient specialization of intelligent decision-making for specific control sub-problems, and difficulty in coordinating multi-timescale scheduling and refined execution. This invention provides a method and system for collaborative scheduling and demand response optimization of virtual power plant resources in a design park, thereby solving the aforementioned technical problems.

[0007] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a resource collaborative scheduling and demand response optimization system for virtual power plants in industrial parks, comprising: The data acquisition and forecasting module is used to collect real-time status information of energy resources within the park and information on the external environment and electricity market, and generate corresponding forecasts. The resource digital twin and constraint modeling module constructs digital twin models and full-dimensional operation constraint models for various energy resources within the park based on real-time status information of energy resources within the park and information on the external environment and electricity market. The multi-agent collaborative scheduling module maps various energy resources and market interaction links within the park into multiple agents. Based on a two-layer Actor-Critic network architecture and combined with a global collaboration mechanism, it outputs the optimal joint action that satisfies the full-dimensional operational constraint model. The grid connection stability and energy management control module is used to execute a fuzzy control strategy with parameters tuned by the particle swarm optimization algorithm based on the control error, output the energy storage power compensation amount, and generate the final charge and discharge power command after merging with the energy storage plan charge and discharge power. The power market and demand response optimization module is used to build a three-layer integrated optimization architecture. It obtains the optimal load regulation amplitude and the optimal discrete execution set based on particle swarm optimization algorithm and ant colony algorithm, and generates equipment-level demand response instructions through a weighted allocation strategy. The intelligent optimization control module provides algorithm training, parameter tuning, and collaborative solution support for the multi-agent collaborative scheduling module, the grid connection stability and energy management control module, and the power market and demand response optimization module. The execution and verification module is used to execute the final charge and discharge power commands and device-level demand response commands, and to monitor and provide feedback on the execution status.

[0008] Furthermore, the data acquisition and prediction module collects real-time status information of energy resources within the park and information on the external environment and electricity market, and generates various prediction quantities, specifically including: The data acquisition and prediction module collects real-time status information of energy resources within the park and information on the external environment and electricity market. The collected data are real-time measured values, and the acquisition frequency matches the system scheduling time step to ensure the real-time performance and effectiveness of the data, meeting the needs of dynamic scheduling of the virtual power plant in the park. The real-time status information of energy resources within the park includes operating status data of renewable energy equipment, operating status data of energy storage systems, operating status data of distributed power sources, operating status data of controllable loads, and operating status data of grid connection points; the external environment and electricity market information includes external meteorological environment data and real-time electricity market signal data; The renewable energy equipment operation status data includes the actual output power of photovoltaic renewable energy power generation equipment, the actual output power of wind power renewable energy power generation equipment, equipment operating conditions, and fault status. The energy storage system operating status data includes the energy storage system's state of charge (SOC), current charging and discharging power, available power boundary, remaining energy storage capacity, and equipment operating temperature; The distributed power source operating status data includes the operating parameters of the controllable distributed power source, including current output, available capacity, start / stop status, and ramp rate. The controllable load operation status data includes the real-time power consumption, adjustable upper limit, equipment operation constraints, and comfort / process constraint thresholds of various controllable loads within the park. The grid connection point operation status data includes the real-time voltage V(t), frequency f(t), and grid-connected switching power of the virtual power plant in the park and the public power grid connection point. ; The external meteorological environment data includes real-time monitoring data of key meteorological factors that affect the output of photovoltaic and wind power, such as light intensity, wind speed, temperature, and humidity. The real-time electricity market signal data includes the real-time electricity price at time t. The time-of-use pricing standard and the demand response directive D(t) issued by the electricity market include the power to be responded to and the power to be reduced.

[0009] Based on big data analysis, time series forecasting, and machine learning forecasting algorithms, the collected basic data is preprocessed and predictive analyzed to generate various forecast quantities adapted to the scheduling cycle of the virtual power plant in the park. The forecast cycle of the forecast quantities is set to match the discretization of the system scheduling cycle, covering the scheduling time range of t=1,2,…,T. The specific forecast quantities generated include renewable energy output forecast, park load forecast, demand response command forecast, and electricity price trend forecast. The renewable energy output forecast includes the photovoltaic power forecast at generation time t. Wind power forecast This provides data support for calculating the equivalent grid-connected power of renewable energy and predicting power output fluctuations. The park load forecast includes the total predicted load power of the park at time t. It accurately predicts the load change trend in the park, providing a basis for power balance constraint calculation and load adjustment decision-making; The predicted demand response command includes combining historical power market commands and real-time market signals to predict the trend of demand response commands or capacity demand changes at each moment within the dispatch cycle, assisting the power market and demand response optimization module in formulating response strategies in advance. The electricity price trend prediction includes the prediction of electricity price change trends at each moment within the dispatch cycle based on real-time electricity price π(t) and time-of-use pricing rules, providing data reference for the park's virtual power plant's electricity purchase / grid connection decisions and economic optimization.

[0010] Furthermore, the resource digital twin and constraint modeling module includes: For all energy resources involved in the scheduling within the park, corresponding digital twin models are constructed. The parameters of the digital twin models are linked in real time with the data acquisition and prediction modules and are dynamically updated according to the actual operating status of the equipment, realizing the synchronous evolution of physical entities and digital twin models. The constructed digital twin models include physical sub-models of distributed power sources, renewable energy power generation equipment, energy storage systems, controllable loads, and grid connection point interaction. The physical sub-model of the distributed power source is designed for controllable distributed power sources. It constructs a physical model that includes output characteristics, start-stop characteristics, ramp rate, energy consumption characteristics, and operating efficiency, and accurately simulates the actual output capacity and operating status of the distributed power source under different operating conditions. The physical sub-model of renewable energy power generation equipment, for photovoltaic and wind power equipment, constructs a power output physical model based on meteorological parameters and equipment parameters, simulates the influence of environmental factors such as sunlight and wind speed on the actual power output of photovoltaic renewable energy power generation equipment and wind power renewable energy power generation equipment, and provides model support for the calculation of equivalent grid-connected power of renewable energy. The physical sub-model of the energy storage system is constructed, including the energy storage capacity. State of charge (SOC) and charge / discharge efficiency A physical model of energy storage with charge and discharge power boundaries accurately simulates the energy conversion law and state change process of energy storage system under different operating conditions such as charging, discharging and standby, matching the calculation requirements of SOC state equation; The controllable load physical sub-model is constructed for various controllable loads within the park, and includes power consumption characteristics, adjustable range, adjustment response speed, and process / comfort constraints according to load type, clarifying the adjustment potential and adjustment limits of different loads; the load types include industrial process loads, commercial lighting loads, and HVAC loads; The aforementioned grid connection point interaction physical sub-model constructs an electrical characteristic model of the virtual power plant in the park and the public power grid connection point, simulating the grid connection point voltage V(t), frequency f(t), and grid connection exchange power. The interaction patterns are matched to the model requirements for grid-connected stability control; Based on the constructed digital twin model, a full-dimensional operational constraint model is established, which specifically includes a power boundary constraint sub-model, a ramp constraint sub-model, a SOC constraint sub-model, a grid-connected power constraint sub-model, a renewable reduction constraint sub-model, a controllable load adjustment range constraint sub-model, and a time-series constraint sub-model. The power boundary constraint sub-model defines upper and lower limits for the power output / consumption / charging / discharging of various energy resources, including the maximum charging power of energy storage. Maximum discharge power Distributed power generation available output boundary, controllable load adjustable upper limit in real time. Define the power regulation boundaries of individual devices; The ramp constraint sub-model defines upper and lower limits for the rate of change of power for distributed power sources and grid-connected renewable energy sources, limiting the sudden increase in equipment output and avoiding equipment loss and system impact caused by rapid power fluctuations. The SOC constraint sub-model, for energy storage systems, defines the boundary range of the State of Charge (SOC). To avoid equipment lifespan loss and safety risks caused by overcharging and over-discharging of energy storage, this model directly addresses the calculation constraints of the energy storage SOC state equation. The grid-connected power constraint sub-model defines the boundary of grid-connected exchange power based on the grid-connected capacity of the park and the power grid and contractual constraints, clarifies the power range of the park's power purchase / grid connection, and ensures that grid-connected behavior complies with power grid requirements; The renewable energy reduction constraint sub-model defines the boundary of renewable energy power reduction, limits the proportion of renewable energy curtailment, and takes into account both system stability and renewable energy consumption needs. The controllable load adjustment range constraint sub-model defines the boundary of the load adjustment amount, constrains the reduction / peak shift of the controllable load, and avoids excessive adjustment from affecting the normal production and life in the park; The time-constraint sub-model defines time-dimensional operational constraints for energy storage SOC, equipment start-up and shutdown times, and load adjustment duration, such as constraints on the rate of change of energy storage charging and discharging power and constraints on the continuous operating time of equipment, to match the time-series scheduling requirements of system rolling optimization.

[0011] Furthermore, the multi-agent cooperative scheduling module includes: Based on the energy resource types and dispatch decision-making needs of the virtual power plant in the park, various energy resources and market interaction links involved in dispatch within the park are abstracted into intelligent agents with local observation, autonomous decision-making, and collaborative interaction capabilities, thus constructing a complete multi-agent ensemble. This invention achieves a one-to-one mapping between physical resources and digital intelligent agents. Each intelligent agent independently undertakes the scheduling and decision-making tasks of its corresponding resources / links. At the same time, information sharing and collaborative decision-making among intelligent agents are realized through system interaction channels. The intelligent agents constructed by this invention include renewable energy intelligent agents, energy storage intelligent agents, controllable load intelligent agents, and market / aggregation intelligent agents. Based on real-time energy resource status information within the park, external environment and electricity market information, as well as a full-dimensional operational constraint model, the local observation status and autonomous decision-making actions of each intelligent agent are standardized and integrated to form a system joint action, providing a standardized input and output carrier for subsequent constraint coordination mechanisms and reinforcement learning training; The local observation state refers to the local observation state of each agent i at time t. All settings are based on the operational characteristics of their mapped resources and key global influencing factors; The autonomous decision-making action refers to the decision-making action output by each intelligent agent based on its own local observation state, which conforms to the operational constraints of its own mapped resources. The autonomous decision-making actions of the intelligent agents include energy storage intelligent agent actions, controllable load intelligent agent actions, and renewable energy intelligent agent actions; the autonomous decision-making actions of all intelligent agents in the multi-agent set are integrated to form a system joint action; To avoid global constraint conflicts caused by distributed autonomous decision-making among agents and to ensure that the joint actions of the system satisfy the full-dimensional operational constraint model, a global coordination mechanism is set up, including a centralized constraint projector mechanism and a Lagrange penalty coordination mechanism. These two mechanisms operate on the decision output stage and training stage of reinforcement learning, respectively, to collaboratively ensure action compliance. The global feasible region... It is composed of power balance constraints, equipment boundary constraints, SOC boundary constraints, and grid-connected power boundary constraints defined by the resource digital twin and constraint modeling module. The centralized constraint projector mechanism employs a feasible region projection operator in the sense of Euclidean distance. The candidate action vectors output by each agent Centralized projection processing will exceed the global feasible region. Candidate actions are mapped to the feasible region to obtain the optimal joint action; The Lagrange penalty coordination mechanism incorporates a constraint violation penalty term into the reward function of each agent, transforming the global constraint satisfaction problem into a reinforcement learning reward optimization problem. Through the reinforcement learning reward mechanism, it guides each agent to autonomously avoid constraint violations, ensuring constraint compliance of decision actions from the source, and constructing a reward function containing a penalty term.

[0012] in, To constrain the penalty coefficient, To constrain the amount of violation; c(t) is the stage cost. Let k be the k-th inequality constraint, and ; Agents that violate constraints will receive negative rewards, and through training, the agents will gradually learn decision-making strategies that conform to the constraints.

[0013] Define the stage cost c(t) at time t, which includes electricity cost, power curtailment loss, load regulation impact, system fluctuation loss, and energy storage loss; The basic reward function takes the negative of the stage cost, guiding the agent to learn in the direction of minimizing the stage cost. The expression is: The reward function with a penalty term, combined with the Lagrange penalty coordination mechanism, incorporates a constraint violation penalty term, and its expression is: ; To balance the overall optimization goals of the current moment and future time periods, a discount reward is defined. By weighting the rewards for each future period according to a discount factor, a balance is achieved between short-term decision-making and long-term optimization. The mathematical expression is as follows:

[0014] in, It is a discount factor, and 0 < <1; T is the total scheduling period; A two-layer Actor-Critic network architecture with centralized evaluation and distributed execution is constructed. The evaluation network (Critic) adopts a centralized design, which integrates global observation and joint actions to evaluate value. The policy network (Actor) adopts a distributed design, with each agent i making autonomous decisions. The two work together to complete reinforcement learning training, and combine the Lagrange penalty coordination mechanism to achieve the goals of constraint compliance and global optimization. With the goal of maximizing the Q-value of the network output, the network parameters are updated using a deterministic gradient approach. The updated formula is:

[0015] in, For the objective function Regarding policy network parameters The gradient; For policy function Regarding its own parameters The gradient; The Q-value function with respect to the current action of agent i The gradient; Through iterative training, combined with the constraint guidance of the Lagrange penalty coordination mechanism, each agent gradually learns a decision-making strategy that satisfies global constraints and is globally optimal. During the online operation phase, each agent, based on its own local observation state... The policy network completed through training Output candidate decision actions and integrate them to form candidate joint actions. Constraints are corrected through a centralized constraint projector mechanism and mapped to the global feasible region. Within this framework, the optimal joint action of the system that ultimately satisfies all global constraints is obtained. The output is sent to the grid connection stability and energy management control module and the power market and demand response optimization module, serving as the basis for subsequent control and optimization dispatch instructions; Furthermore, the grid connection stability and energy management control module includes: The grid connection stability and energy management control module receives real-time operation data, scheduling layer decision data, and algorithm and model support data; The real-time operating data includes the real-time voltage V(t) and frequency f(t) at the grid connection point. Actual output of renewable energy , Real-time load power of the park Real-time SOCs(t) and charge / discharge power boundaries of the energy storage system; The scheduling layer decision data includes the energy storage plan's charging and discharging power. Reference value for equivalent grid-connected power of renewable energy, and planned value for power exchange between the park and the power grid. ; The algorithm and model support data include the optimal fuzzy control parameters after particle swarm optimization tuning. and full-dimensional operational constraint model; The standardized definition of control error serves as both a core input for fuzzy control and a quantitative indicator for constructing the fitness function in particle swarm optimization. Control error includes voltage error and power error, adapting to different grid-connected control requirements. The grid connection stability and energy management control module can flexibly select input combinations according to the priority of grid connection control in the park: voltage error and voltage error change rate can be selected as the core input of fuzzy control, or it can be replaced with power error and power error change rate, or a combination of multiple error inputs can be used to adapt to the grid connection control needs of different parks. The fuzzy control strategy shown specifically includes: The input variables are fuzzified; the set of input linguistic variables is defined as {NB, NS, ZO, PS, PB}, representing negative large, negative small, zero, positive small, and positive large, respectively; a corresponding membership function is defined for each linguistic variable. This process converts continuous, precise error quantities into fuzzy linguistic values, thus fuzzifying the input quantities; the input quantities include voltage errors. and voltage error change rate Combination or power error and power error change rate The combination; Construct a fuzzy inference rule base based on grid-connected operation experience and control requirements, with rules uniformly expressed as follows: :like for and for ,but for ,in , For a fuzzy subset of the input linguistic variables, Output energy storage power compensation amount The fuzzy subset of the rule base covers all typical scenarios of voltage / power fluctuations at the grid connection point, ensuring the comprehensiveness of the reasoning; Based on the constructed fuzzy inference rule base, the Mamdani inference method is used to infer the fuzzified input, and the activation weight of each rule is calculated. The activation weight calculation formula is as follows:

[0016] The weighted average method is used to defuzzify the fuzzy inference results, transforming the fuzzy linguistic values ​​into continuous and precise energy storage power compensation quantities. The calculation formula is as follows:

[0017] Where M is the total number of fuzzy rules. This represents the output value corresponding to the m-th rule; The energy storage power compensation amount output by fuzzy control is compared with the planned charging and discharging power of the energy storage system. By combining the results, the final charge and discharge power command of the energy storage system is obtained, as shown in the formula:

[0018] in, The planned power output by the scheduling layer. This is the final charge / discharge power command; The final charging and discharging power command is sent to the execution and verification module, which directly drives the energy storage device to operate; Furthermore, the particle swarm optimization algorithm is used to tune the key parameters of the fuzzy control, specifically including: All key parameters of fuzzy control are packaged into a position vector x for particle swarm optimization, with the core including the input-output scaling factor. , , ; With the goals of minimizing voltage and power fluctuations and optimizing energy consumption control, a comprehensive quantitative fitness function F(x) is constructed to balance grid-connected stability control performance with energy storage control costs. The standard particle swarm optimization (PSO) iterative update formula is used to perform a global optimum search on the parameter vector x. The iterative formula is as follows:

[0019]

[0020] in, Let be the velocity vector of the i-th particle in the k-th iteration. For position vectors, This is the optimal position in the particle's history. This is the best historical position for the group. For inertial weights, , As a learning factor, , 2 is a uniformly random number in the interval [0,1]. The optimal parameters for fuzzy control are obtained through iterative search. It is directly applied to online fuzzy controllers; at the same time, based on the real-time changes in the operating status of the grid connection point, small dynamic fine-tuning of parameters is carried out through particle swarm optimization to ensure that the fuzzy control always adapts to the changing operating conditions of the park's grid connection operation.

[0021] Furthermore, the electricity market and demand response optimization module includes: A three-layer fusion optimization architecture is constructed, including a continuous decision-making layer, a discrete decision-making layer, and a collaborative allocation layer; The continuous decision-making layer employs a particle swarm optimization algorithm to solve for the optimal load adjustment amplitude at each moment within the scheduling cycle in the time dimension, addressing the global economic optimization problem of continuous quantities. The discrete decision-making layer employs an ant colony algorithm to select the controllable load execution set participating in demand response at each moment in the device dimension, addressing the practical executability problem of discrete resources. The collaborative allocation layer adopts a weighted allocation strategy to scientifically allocate the continuous optimal amplitude obtained from particle swarm optimization within the discrete execution set selected by the ant colony algorithm, forming device-level executable instructions and achieving dual optimization of economy and executability. For the continuous decision variable of demand response adjustment magnitude, a particle swarm optimization algorithm is used to find the global optimum. With the objective of maximizing the net demand response benefit of the industrial park, the optimal load adjustment magnitude at each time point is determined within constraints. Specifically, this includes: The demand response magnitude at each moment within the scheduling period T is constructed into a continuous decision vector; By comprehensively calculating the total revenue and various costs of demand response, a net revenue calculation system is formed, providing a quantitative basis for constructing the fitness function of particle swarm optimization; The goal of maximizing net profit is transformed into minimizing it. The fitness function of particle swarm optimization is defined. The smaller the fitness function value, the better the economy of the demand response solution. Using the continuous decision vector y as the position vector of the particle swarm, a global optimum search is performed using the standard particle swarm iteration update formula, which is:

[0022]

[0023] in, Let be the velocity vector of the i-th particle in the k-th iteration. For position vectors, This is the optimal position in the particle's history. This represents the group's historically optimal position. The position vector obtained in each iteration is truncated to ensure that all elements satisfy 0 ≤ ≤ The final output is the optimal load adjustment amplitude within the scheduling cycle. ; For the discrete decision problem of device / loop level execution selection, an ant colony algorithm is used to select the optimal set of controllable load executions at each time point, balancing executability and optimal unit adjustment cost, including: All controllable loads within the park are constructed as a discrete set: Each element corresponds to a controllable load loop / device; the decision objective of the ant colony algorithm at time t is defined as selecting a subset from set L. As the execution device that participates in demand response at that moment; Define the basic parameters of the ant colony algorithm to provide a decision-making basis for the selection of discrete sets, specifically including: Define pheromone concentration , represents the pheromone concentration at time t that selects load j to participate in the demand response. The higher the concentration, the greater the probability of being selected by the ants. Define heuristic factors This reflects the degree of matching between the benefits and costs of selecting load j, and is defined as the ratio of unit benefit to unit impact cost, as shown in the formula:

[0024] in, Let j be the specific impact cost coefficient for load j. To prevent small positive numbers from being divided by zero; Pheromones volatile coefficient The pheromone constant Q and the heuristic factor weights α / β are all parameters calibrated offline by the intelligent optimization control module and dynamically fine-tuned according to real-time operating conditions. The probability that the k-th ant will choose load j to participate in the demand response at time t is determined by both the pheromone concentration and the heuristic factor. The formula for calculating the ant selection probability is:

[0025] Where α is the pheromone weight and β is the heuristic factor weight, which respectively determine the degree of influence of pheromone and heuristic factor on the selection probability.

[0026] After each iteration, the pheromone is globally updated based on the cost of the scheme corresponding to the execution subset selected by each ant, including pheromone evaporation and pheromone incremental replenishment. Through multiple iterations, the pheromone converges to the optimal execution subset, ultimately outputting the optimal discrete execution set of the demand response at time t. This ensures that the equipment set meets the adjustment capacity while minimizing the unit adjustment cost. The optimal load adjustment amplitude was obtained through particle swarm optimization. Discrete optimal execution set of ant colony algorithm Then, a weighted allocation strategy is used to determine the optimal load adjustment at each time point. In executing collection The equipment within the facility is scientifically allocated to generate equipment-level demand response instructions. The specific allocation rules are as follows: The allocation weight is based on the proportion of each equipment's adjustable upper limit or the reciprocal of its unit impact cost. Priority is given to allocating more adjustment to equipment with a large adjustable upper limit and a low unit impact cost, taking into account both adjustment potential and the impact on park operation. Let time t be the execution set There are m devices, and the adjustable upper limit of device j is . The weights are assigned as follows Then the adjustment amount of device j is:

[0027] The allocated adjustment amount for each device is then verified a second time to ensure that 0 ≤ ≤ If the constraint is exceeded, a truncation correction is performed, and the remaining adjustment is redistributed to other devices to ensure that the total adjustment is consistent with the constraint. Consistent.

[0028] Secondly, this invention provides a method for resource collaborative scheduling and demand response optimization of virtual power plants in a park. It adopts a rolling time-domain online operation mode, and by setting a rolling window length H, it achieves dynamic online optimization of resource collaborative scheduling and demand response of virtual power plants in the park, balancing the global optimality of the scheduling strategy with the real-time adaptability of execution. Specifically, it includes the following steps: Step S1: At the current running time t, collect real-time status information of energy resources in the park and real-time measurement values ​​of external environment and electricity market information, and generate predicted quantities within the rolling window [t, t+H]. Step S2: Based on the real-time measurement and prediction values ​​from Step S1, feasible joint scheduling actions are obtained through multi-agent collaborative scheduling decisions. The feasible joint scheduling actions include planned energy storage charging and discharging power and initial load adjustment decisions; Step S3: Based on the energy storage plan charging and discharging power output in step S2, and combined with the fuzzy control strategy with parameters tuned by the particle swarm optimization algorithm, calculate the energy storage power compensation amount at the current time t, and fuse them to generate the final charging and discharging power command. Step S4: Based on the initial load adjustment decision output in step S2, within the rolling window [t, t+H], the optimal load adjustment amplitude is solved by the particle swarm optimization algorithm, and the optimal discrete execution set is selected by the ant colony algorithm. The equipment-level load adjustment command at the current time t is generated by the weighted allocation strategy. Step S5: Send the final charging and discharging power command generated in step S3 and the equipment-level load adjustment command generated in step S4 to the field equipment and execute only the command at the current time t, while monitoring and recording the command execution status and deviation. Step S6: Use the execution result and deviation information from step S5 as feedback, update the system status, advance the running time to t+1, and return to step S1.

[0029] The beneficial effects of this invention are as follows: By adopting a cooperative scheduling framework based on multi-agent reinforcement learning, it overcomes the problem of weak model adaptability in existing technologies, enabling the scheduling scheme to exhibit stronger stability and economy when facing prediction errors and random disturbances. Addressing the deficiency of general reinforcement learning frameworks in terms of insufficient accuracy for specialized control problems, this invention employs a fuzzy control strategy with parameters tuned by particle swarm optimization, constructing a three-layer fusion architecture to solve for the optimal load adjustment amplitude, select the optimal set of execution devices, and generate device-level instructions, thereby enhancing the decision-making accuracy and efficiency for specific sub-problems. Furthermore, by adopting a hierarchical optimization and cooperative allocation strategy, it achieves efficient coordination between scheduling layer planning and control layer compensation, and between global optimization and device-level execution. This solves the problems of weak adaptability of global optimization models, insufficient specialization of intelligent decision-making for specific control sub-problems, and difficulty in coordinating multi-timescale scheduling and refined execution.

[0030] Furthermore, the design principle of this invention is reliable, the structure is simple, and it has a very wide range of application prospects.

[0031] Therefore, it is evident that the present invention has substantial features and progress compared with the prior art, and the beneficial effects of its implementation are also obvious. Attached Figure Description

[0032] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0033] Figure 1 This invention provides a principle block diagram of a virtual power plant resource collaborative scheduling and demand response optimization system for industrial parks.

[0034] Figure 2The present invention provides a flowchart of a method for resource collaborative scheduling and demand response optimization of a virtual power plant in a park.

[0035] The module comprises: 1-Data Acquisition and Prediction Module, 2-Resource Digital Twin and Constraint Modeling Module, 3-Multi-Agent Cooperative Scheduling Module, 4-Intelligent Optimization Control Module, 5-Grid Connection Stability and Energy Management Control Module, 6-Electricity Market and Demand Response Optimization Module, and 7-Execution and Verification Module. Detailed Implementation

[0036] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The following embodiments are explanations of the present invention, but the present invention is not limited to the following implementation methods.

[0037] Example 1: Firstly, this invention provides a resource collaborative scheduling and demand response optimization system for virtual power plants in industrial parks, such as... Figure 1 As shown, it includes: Data acquisition and forecasting module 1 is used to collect real-time status information of energy resources in the park and information on the external environment and electricity market, and generate corresponding forecasts. Resource Digital Twin and Constraint Modeling Module 2: Based on real-time status information of energy resources within the park and information on the external environment and electricity market, construct digital twin models and full-dimensional operational constraint models for various energy resources within the park; The multi-agent collaborative scheduling module 3 maps various energy resources and market interaction links in the park into multiple agents. Based on a two-layer Actor-Critic network architecture and combined with a global collaborative mechanism, it outputs the optimal joint action that satisfies the full-dimensional operation constraint model. The intelligent optimization control module 4 is used to provide algorithm training, parameter tuning and collaborative solution support for the multi-agent collaborative scheduling module 3, the grid connection stability and energy management control module 5 and the power market and demand response optimization module 6. The grid connection stability and energy management control module 5 is used to execute a fuzzy control strategy with parameters tuned by the particle swarm optimization algorithm based on the control error, output the energy storage power compensation amount, and generate the final charge and discharge power command after merging with the energy storage plan charge and discharge power. The power market and demand response optimization module 6 is used to build a three-layer integrated optimization architecture. It obtains the optimal load regulation amplitude and the optimal discrete execution set based on particle swarm optimization algorithm and ant colony algorithm, and generates equipment-level demand response instructions through a weighted allocation strategy. The execution and verification module 7 is used to execute the final charge and discharge power command and the device-level demand response command, and to monitor and provide feedback on the execution status.

[0038] This invention follows the basic physical laws governing the operation of a virtual power plant within the park, maintaining power balance within the park at all times:

[0039] in, This represents the exchange power between the park and the public power grid. A positive value indicates that electricity is purchased from the grid, while a negative value indicates that electricity is fed back to the grid. To provide power output for controllable distributed power sources within the park; The equivalent grid-connected power of photovoltaic renewable energy and wind power renewable energy; The charging and discharging power of the energy storage system is defined by... >0 indicates discharge. <0 indicates charging; This represents the total load power of the park. This represents the load adjustment amount in demand response; a positive value indicates load reduction.

[0040] Equivalent grid-connected power of renewable energy The actual output of photovoltaic and wind power is determined by both the power reduction commanded by the system, and their relationship is as follows:

[0041] in, and These represent the actual power output of photovoltaic and wind power at time t, respectively. The renewable energy power reduction is given for system optimization decisions; Furthermore, the data acquisition and prediction module 1 specifically includes: The data acquisition and prediction module 1 collects real-time status information of energy resources within the park and information on the external environment and electricity market; the collected data are real-time measured values, and the acquisition frequency matches the system scheduling time step. The real-time status information of energy resources within the park includes operating status data of renewable energy equipment, operating status data of energy storage systems, operating status data of distributed power sources, operating status data of controllable loads, and operating status data of grid connection points; the external environment and electricity market information includes external meteorological environment data and real-time electricity market signal data; The renewable energy equipment operation status data includes the actual output power, equipment operating conditions, and fault status of photovoltaic renewable energy power generation equipment and wind power renewable energy power generation equipment. The energy storage system operating status data includes the energy storage system's state of charge (SOC), current charging and discharging power, available power boundary, remaining energy storage capacity, and equipment operating temperature; The distributed power source operating status data includes the current output, available capacity, start / stop status, and ramp rate of the controllable distributed power source. The controllable load operation status data includes the real-time power consumption, adjustable upper limit, equipment operation constraints, and comfort / process constraint thresholds of various controllable loads within the park. The grid connection point operation status data includes the real-time voltage V(t), frequency f(t), and grid-connected switching power of the virtual power plant in the park and the public power grid connection point. ; The external meteorological environment data includes real-time monitoring data of key meteorological factors that affect the output of photovoltaic and wind power, such as light intensity, wind speed, temperature, and humidity. The real-time electricity market signal data includes the real-time electricity price at time t. The time-of-use pricing standard and the demand response directive D(t) issued by the electricity market include the power to be responded to and the power to be reduced.

[0042] Based on big data analysis, time series forecasting, and machine learning forecasting algorithms, the collected basic data is preprocessed and predictive analyzed to generate various forecast quantities adapted to the scheduling cycle of the virtual power plant in the park. The forecast cycle of the forecast quantities is set to match the discretization of the system scheduling cycle, covering the scheduling time range of t=1,2,…,T. The specific forecast quantities generated include renewable energy output forecast, park load forecast, demand response command forecast, and electricity price trend forecast. The renewable energy output forecast includes the photovoltaic power forecast at generation time t. Wind power forecast This provides data support for calculating the equivalent grid-connected power of renewable energy and predicting power output fluctuations. The park load forecast includes the total predicted load power of the park at time t. It accurately predicts the load change trend in the park, providing a basis for power balance constraint calculation and load adjustment decision-making; The predicted demand response command includes combining historical power market commands and real-time market signals to predict the trend of demand response commands or capacity demand changes at each moment within the dispatch cycle, assisting the power market and demand response optimization module 6 in formulating response strategies in advance. The electricity price trend prediction includes the prediction of electricity price change trends at each moment within the dispatch cycle based on real-time electricity price π(t) and time-of-use pricing rules, providing data reference for the park's virtual power plant's electricity purchase / grid connection decisions and economic optimization.

[0043] After standardizing the collected real-time measured values ​​and generated forecasts, the data is output to other modules in a standardized data format according to the input requirements of each subsequent module of the system. The output data includes renewable energy related data, load related data, electricity market and demand response related data, equipment status data, grid connection point operation data and environmental auxiliary data. The renewable energy-related data includes photovoltaic power forecasts. Wind power forecast Real-time values ​​of actual renewable energy output and data on the fluctuation trend of renewable energy output; The load-related data includes the predicted power of the park's load. Real-time power of controllable load, adjustable upper limit of controllable load ; The electricity market and demand response related data include the electricity price π(t) at time t, the demand response order or capacity demand D(t), the electricity price trend forecast, and the demand response order forecast. The device status data includes energy storage. Available power boundary of energy storage, available capacity of distributed power sources, and controllable load operation constraint parameters; The grid connection point operation data includes grid connection point voltage V(t), frequency f(t), and grid-connected switching power. ; The environmental auxiliary data, including sunlight and wind speed, provides auxiliary basis for correcting renewable energy output forecasts and dynamically adjusting dispatch strategies.

[0044] Furthermore, the resource digital twin and constraint modeling module 2 includes: For all energy resources involved in the scheduling within the park, a corresponding digital twin model is constructed, and the model parameters are dynamically updated according to the actual operating status of the equipment. The constructed digital twin model includes a physical sub-model of distributed power sources, a physical sub-model of renewable energy power generation equipment, a physical sub-model of energy storage systems, a physical sub-model of controllable loads, and a physical sub-model of grid connection point interaction. The physical sub-model of the distributed power source is designed for controllable distributed power sources. It constructs a physical model that includes output characteristics, start-stop characteristics, ramp rate, energy consumption characteristics, and operating efficiency to simulate the actual output capacity and operating status of the distributed power source under different operating conditions. The physical sub-model of renewable energy power generation equipment, for photovoltaic and wind power equipment, constructs a power output physical model based on meteorological parameters and equipment parameters, simulates the influence of environmental factors such as sunlight and wind speed on the actual power output of photovoltaic renewable energy power generation equipment and wind power renewable energy power generation equipment, and provides model support for the calculation of equivalent grid-connected power of renewable energy. The physical sub-model of the energy storage system is constructed, including the energy storage capacity. State of charge (SOC) and charge / discharge efficiency A physical model of energy storage with charge and discharge power boundaries accurately simulates the energy conversion law and state change process of energy storage system under different operating conditions such as charging, discharging and standby, matching the calculation requirements of SOC state equation; The SOC state equation is used to describe the update law of the state of charge of the energy storage system with time step, and the specific mathematical expression is as follows:

[0045] in, and These represent the energy storage state of charge at the current time and the next time, respectively; Δt is the scheduling time step. This refers to the rated capacity of the energy storage. For energy storage charging and discharging power, the following is agreed upon >0 indicates discharge. <0 indicates charging; and These are charging efficiency and discharging efficiency, respectively. The controllable load physical sub-model is constructed for various controllable loads within the park, and includes power consumption characteristics, adjustable range, adjustment response speed, and process / comfort constraints according to load type, clarifying the adjustment potential and adjustment limits of different loads; the load types include industrial process loads, commercial lighting loads, and HVAC loads; The aforementioned grid connection point interaction physical sub-model constructs an electrical characteristic model of the virtual power plant in the park and the public power grid connection point, simulating the grid connection point voltage V(t), frequency f(t), and grid connection exchange power. The interaction patterns are matched to the model requirements for grid-connected stability control; Based on the constructed digital twin model, a full-dimensional operational constraint model is established, including power boundary constraint sub-model, ramp constraint sub-model, SOC constraint sub-model, grid-connected power constraint sub-model, renewable load reduction constraint sub-model, controllable load adjustment range constraint sub-model, and time-series constraint sub-model. The power boundary constraint sub-model defines upper and lower limits for the power output / consumption / charging / discharging of various energy resources, including the maximum charging power of energy storage. Maximum discharge power Distributed power generation available output boundary, controllable load adjustable upper limit in real time. Define the power regulation boundaries of individual devices; The ramp constraint sub-model defines upper and lower limits for the rate of change of power for distributed power sources and grid-connected renewable energy sources, limiting the sudden increase in equipment output and avoiding equipment loss and system impact caused by rapid power fluctuations. The SOC constraint sub-model, for energy storage systems, defines the boundary range of the State of Charge (SOC). To avoid equipment lifespan loss and safety risks caused by overcharging and over-discharging of energy storage, this model directly addresses the calculation constraints of the energy storage SOC state equation. The grid-connected power constraint sub-model defines the boundary of grid-connected exchange power based on the grid-connected capacity and contractual constraints of the park and the power grid, clarifies the power range for the park to purchase / connect electricity, and ensures that grid-connected behavior complies with power grid requirements; the specific expression is:

[0046] in, For grid-connected power exchange, To achieve the minimum grid-connected switching power, Maximum grid-connected switching power; The renewable energy reduction constraint sub-model defines the boundary of renewable energy power reduction, limits the proportion of renewable energy curtailment, and balances system stability with renewable energy consumption needs; the specific expression is:

[0047] in, , These are the actual power output values ​​for photovoltaic and wind power, respectively. To reduce power; The controllable load adjustment range constraint sub-model defines the boundary of the load adjustment amount, constrains the reduction / peak shifting amplitude of the controllable load, and avoids excessive adjustment from affecting the normal production and life in the park; the specific expression is:

[0048] The time-constraint sub-model defines time-dimensional operational constraints for energy storage SOC, equipment start-up and shutdown times, and load adjustment duration, such as constraints on the rate of change of energy storage charging and discharging power and constraints on the continuous operating time of equipment, to match the time-series scheduling requirements of system rolling optimization.

[0049] Furthermore, the multi-agent cooperative scheduling module 3 specifically includes: Based on the energy resource types and dispatch decision-making needs of the virtual power plant in the park, various energy resources and market interaction links involved in dispatch within the park are abstracted into intelligent agents to construct a complete multi-agent ensemble. This system achieves a one-to-one mapping between physical resources and digital intelligent agents. Each intelligent agent independently undertakes the scheduling and decision-making tasks for its corresponding resources / links. At the same time, it realizes information sharing and collaborative decision-making among intelligent agents through system interaction channels. The constructed intelligent agents include renewable energy intelligent agents, energy storage intelligent agents, controllable load intelligent agents, and market / aggregation intelligent agents. The renewable energy intelligent agent maps to the photovoltaic and wind power generation equipment within the park and is responsible for reducing renewable energy power consumption. The decision-making process and the setting of the equivalent grid-connected power of renewable energy should take into account both the renewable energy consumption target and the system grid connection stability requirements. The energy storage intelligent agent is mapped to the park's energy storage system and is responsible for the energy storage charging and discharging power. The decision-making process dynamically adjusts the charging and discharging state of energy storage based on the system power balance requirements and the fluctuation of renewable energy output, thereby realizing the spatiotemporal transfer of energy and buffering of system fluctuations. The controllable load agent is mapped to various controllable load clusters within the park and is responsible for load regulation. Based on the power market demand response instructions and the park load forecast trend, the decision-making process formulates load adjustment strategies within the load adjustable constraint range. The market / aggregation agent interacts with the power market of the virtual power plant in the park and maps the global resource aggregation process. It is responsible for the fusion processing of electricity price signals, demand response instructions and power grid transaction boundaries. At the same time, it undertakes the distribution of global constraint information, the aggregation of decision information from various agents, and the coordination of various agents to achieve global goals.

[0050] Based on real-time energy resource status information within the park, external environment and electricity market information, as well as a full-dimensional operational constraint model, the local observation status and autonomous decision-making actions of each intelligent agent are standardized and integrated to form a system joint action, providing a standardized input and output carrier for subsequent constraint coordination mechanisms and reinforcement learning training; The local observation state refers to the local observation state of each agent i at time t. All settings are based on the operational characteristics of their mapped resources and key global influencing factors, and the specific expressions are as follows:

[0051] in, Forecasted output of renewable energy = + ; For energy storage SOC, For real-time electricity prices, , These are the grid connection point voltage and frequency, respectively. The autonomous decision-making action refers to the decision-making action output by each intelligent agent based on its own local observation state, which conforms to the operational constraints of its own mapped resources. The autonomous decision-making actions of the intelligent agents include energy storage intelligent agent actions, controllable load intelligent agent actions, and renewable energy intelligent agent actions. The actions of the energy storage intelligent agent are defined as follows:

[0052] in, For energy storage charging and discharging power; The controllable load agent action is defined as:

[0053] in, This refers to the load adjustment amount; The actions of the renewable energy intelligent agent are defined as follows:

[0054] in, Reduce power consumption for renewable energy sources; Integrating the autonomous decision-making actions of all agents in a multi-agent ensemble to form a joint system action, the specific expression is as follows:

[0055] To ensure that the system's joint actions meet the full-dimensional operational constraints model, a global coordination mechanism is set up, including a centralized constraint projector mechanism and a Lagrange penalty coordination mechanism; this mechanism operates in the decision output and training phases of reinforcement learning, collaboratively ensuring action compliance; among which, the global feasible region It consists of a full-dimensional operational constraint model defined by the resource digital twin and constraint modeling module 2; The centralized constraint projector mechanism employs a feasible region projection operator in the sense of Euclidean distance. The candidate action vectors output by each agent Centralized projection processing will exceed the global feasible region. The candidate actions are mapped to the feasible region to obtain the optimal joint action, which can be expressed mathematically as:

[0056] The Lagrange penalty coordination mechanism incorporates a constraint violation penalty term into the reward function of each agent, transforming the global constraint satisfaction problem into a reinforcement learning reward optimization problem. Through the reinforcement learning reward mechanism, it guides each agent to autonomously avoid constraint violations, ensuring constraint compliance of decision actions from the source, and constructing a reward function containing a penalty term.

[0057] in, To constrain the penalty coefficient, To constrain the amount of violation; c(t) is the stage cost. Let k be the k-th inequality constraint, and ; Agents that violate constraints will receive negative rewards, and through training, the agents will gradually learn decision-making strategies that conform to the constraints.

[0058] Based on the global optimization objective, the stage cost c(t) at time t is defined, including electricity cost, power curtailment loss, load regulation impact, system fluctuation loss, and energy storage loss. The specific expression is as follows:

[0059] in, For electricity price; For electricity purchase / grid connection capacity; Represents electricity costs, if A negative value can represent revenue from internet access; This serves as a penalty coefficient for curtailment, promoting the consumption of renewable energy. The demand response comfort / process impact coefficient is used to suppress the impact of excessive load reduction. This refers to the voltage deviation at the grid connection point. = - ;in The desired voltage can be either the rated voltage or the target voltage for dispatching. This refers to the frequency deviation at the grid connection point. = - ,in The rated frequency; , This is a stability penalty coefficient; This is the penalty coefficient.

[0060] The basic reward function takes the negative of the stage cost, guiding the agent to learn in the direction of minimizing the stage cost. The expression is: The reward function with a penalty term, combined with the Lagrange penalty coordination mechanism, incorporates a constraint violation penalty term, and its expression is: ; To balance the overall optimization goals of the current moment and future time periods, a discount reward is defined. By weighting the rewards for each future period according to a discount factor, a balance is achieved between short-term decision-making and long-term optimization. The mathematical expression is as follows:

[0061] in, It is a discount factor, and 0 < <1; T is the total scheduling period; A two-layer Actor-Critic network architecture is constructed. The evaluation network (Critic) adopts a centralized design, which integrates global observation and joint actions to evaluate value. The policy network (Actor) adopts a distributed design, with each agent i making autonomous decisions. The two work together to complete reinforcement learning training, and combine the Lagrange penalty coordination mechanism to achieve the goals of constraint compliance and global optimization. The evaluation network is defined as follows:

[0062] in, For global observation / mosaic observation, For system-wide coordinated actions; Update network parameters by minimizing the mean squared error loss. The loss function is:

[0063] The target value is:

[0064] in, For the target network parameters, The combined action for the policy output in the next moment; Each agent i corresponds to an independent policy network:

[0065] With the goal of maximizing the Q-value of the network output, the network parameters are updated using a deterministic gradient approach. The updated formula is:

[0066] in, For the objective function Regarding policy network parameters The gradient; For policy function Regarding its own parameters The gradient; The Q-value function with respect to the current action of agent i The gradient; Through iterative training using the aforementioned two-layer Actor-Critic network architecture of centralized evaluation and distributed execution, combined with the constraint guidance of the Lagrange penalty coordination mechanism, each agent gradually learns a decision-making strategy that satisfies global constraints and is globally optimal. During the online operation phase, each agent, based on its own local observation state... The policy network completed through training Output candidate decision actions and integrate them to form candidate joint actions. Constraints are corrected through a centralized constraint projector mechanism and mapped to the global feasible region. Within this framework, the optimal joint action of the system that ultimately satisfies all global constraints is obtained. The output is sent to the grid connection stability and energy management control module 5 and the power market and demand response optimization module 6, serving as the basis for subsequent control and optimization dispatch instructions; Furthermore, the grid connection stability and energy management control module 5 specifically includes: The grid connection stability and energy management control module 5 receives real-time operation data, scheduling layer decision data, and algorithm and model support data; The real-time operating data includes the real-time voltage V(t) and frequency f(t) at the grid connection point. Actual output of renewable energy , Real-time load power of the park Real-time SOCs(t) and charge / discharge power boundaries of the energy storage system; The scheduling layer decision data refers to the scheduling instructions calculated and output by the multi-agent collaborative scheduling module 3, which serve as plans or reference values, including the planned charging and discharging power of energy storage. Reference value for equivalent grid-connected power of renewable energy, and planned value for power exchange between the park and the power grid. ; The algorithm and model support data include the optimal fuzzy control parameters after particle swarm optimization tuning. and full-dimensional operational constraint model; The standardized definition of control error serves as both a core input for fuzzy control and a quantitative indicator for constructing the fitness function in particle swarm optimization. Control error includes voltage error and power error, adapting to different grid-connected control requirements. The voltage error is defined as:

[0067] in, The reference value for the grid connection point voltage can be the grid rated voltage or the scheduling target voltage set by the multi-agent collaborative scheduling module 3; It reflects the degree of deviation between the actual voltage at the grid connection point and the target voltage; The power error is defined as:

[0068] in, For the desired grid-connected switching power or power smoothing target value, It reflects the degree of deviation between the actual grid-connected power and the planned power at the grid connection point; The grid connection stability and energy management control module 5 can flexibly select input combinations according to the priority of grid connection control in the park: voltage error and voltage error change rate can be selected as the core input of fuzzy control, or it can be replaced with power error and power error change rate, or a combination of multiple error inputs can be used to adapt to the grid connection control requirements of different parks. As the core of real-time response to grid connection fluctuations, the fuzzy control strategy adopts the classic execution process of fuzzification-fuzzy inference-defuzzification, and outputs the energy storage power compensation amount Δ. The power output of the energy storage plan from the scheduling layer is corrected to achieve rapid suppression of fluctuations, including: The input variables are fuzzified; the set of input linguistic variables is defined as {NB, NS, ZO, PS, PB}, representing negative large, negative small, zero, positive small, and positive large, respectively; a corresponding membership function is defined for each linguistic variable. This process converts continuous, precise error quantities into fuzzy linguistic values, thus fuzzifying the input quantities; the input quantities include voltage errors. and voltage error change rate Combination or power error and power error change rate The combination; Construct a fuzzy inference rule base based on grid-connected operation experience and control requirements, with rules uniformly expressed as follows: :like for and for ,but for ,in , For a fuzzy subset of the input linguistic variables, Output energy storage power compensation amount The fuzzy subset of the rule base covers all typical scenarios of voltage / power fluctuations at the grid connection point, ensuring the comprehensiveness of the reasoning; Based on the constructed fuzzy inference rule base, the Mamdani inference method is used to infer the fuzzified input, and the activation weight of each rule is calculated. The activation weight calculation formula is as follows:

[0069] The weighted average method is used to defuzzify the fuzzy inference results, transforming the fuzzy linguistic values ​​into continuous and precise energy storage power compensation quantities. The calculation formula is as follows:

[0070] Where M is the total number of fuzzy rules. This represents the output value corresponding to the m-th rule; The energy storage power compensation amount output by fuzzy control is compared with the planned charging and discharging power of the energy storage system. By combining the results, the final charge and discharge power command of the energy storage system is obtained, as shown in the formula:

[0071] in, The planned power output by the scheduling layer. This is the final charge / discharge power command; The final charging and discharging power command is sent to the execution and verification module 7, which directly drives the energy storage device to operate; Furthermore, to achieve the globally optimal configuration of fuzzy control parameters and eliminate the subjectivity and limitations of manual experience-based tuning, a particle swarm optimization algorithm is used to tune the key parameters of the fuzzy control, including offline global search and online dynamic fine-tuning of the parameters, including: All key parameters of fuzzy control are packaged into a position vector x for particle swarm optimization, with the core including the input-output scaling factor. , , And the center and width of the membership functions of each language variable, to achieve unified parameter optimization; the specific expression is:

[0072] With the goals of minimizing voltage and power fluctuations and optimizing control energy consumption, a comprehensive quantitative fitness function F(x) is constructed, which balances grid-connected stability control performance with energy storage control costs. The formula is as follows:

[0073] in, , , These are weighting coefficients, corresponding to the weights of voltage fluctuation, power fluctuation, and energy storage compensation power, respectively, and can be dynamically adjusted according to the control priority of the park. The standard particle swarm optimization (PSO) iterative update formula is used to perform a global optimum search on the parameter vector x. The iterative formula is as follows:

[0074]

[0075] in, Let be the velocity vector of the i-th particle in the k-th iteration. For position vectors, This is the optimal position in the particle's history. This is the best historical position for the group. For inertial weights, , As a learning factor, , 2 is a uniformly random number in the interval [0,1]. The optimal parameters for fuzzy control are obtained through iterative search. It is directly applied to online fuzzy controllers; at the same time, based on the real-time changes in the operating status of the grid connection point, the parameters are dynamically fine-tuned through particle swarm optimization to ensure that the fuzzy control always adapts to the changes in the operating conditions of the park's grid connection. The grid connection stability and energy management control module 5 integrates energy management strategies while achieving grid connection stability control. It combines grid connection control with renewable energy consumption and energy storage optimization to achieve synergistic optimization of control and management. The core strategies include renewable energy consumption priority strategy, energy storage energy state constraint strategy, and grid connection power smoothing and load balancing synergistic strategy. The aforementioned renewable energy consumption priority strategy prioritizes absorbing excess power through energy storage systems when renewable energy output surges, thereby reducing renewable energy reduction. Only when energy storage reaches the charging power boundary or the SOC limit should the output of renewable energy be moderately reduced to ensure that the local consumption of renewable energy is maximized. The energy storage state constraint strategy generates energy storage power compensation. At the same time, strictly adhere to the energy storage SOC boundary defined by the resource digital twin and constraint modeling module. and charge / discharge power boundary If the compensation power exceeds the constraint, it will be truncated and corrected to avoid overcharging, over-discharging or over-power operation of the energy storage. The proposed grid-connected power smoothing and load balancing coordination strategy, combined with real-time load changes in the park, suppresses grid-connected power fluctuations while balancing power supply and demand within the park through energy storage power compensation. This reduces power exchange fluctuations between the park and the public power grid, lowers grid regulation pressure, and enhances the park's energy system's self-balancing capability.

[0076] Furthermore, the electricity market and demand response optimization module 6 includes: The Electricity Market and Demand Response Optimization Module 6 receives electricity market signal data, park load and equipment data, upstream dispatch and algorithm data, and constraint model data; all data are updated synchronously in real time, providing comprehensive and accurate basic support for demand response optimization decisions; The electricity market signal data includes the real-time electricity price π(t) at time t from the data acquisition and prediction module 1, the time-of-use electricity price standard, the demand response subsidy unit price ρ(t), the demand response instruction D(t) issued by the power grid, and the trend prediction of electricity price and subsidy. The park load and equipment data include real-time park load data from data acquisition and prediction module 1. Controllable load with adjustable upper limit The operating status of each controllable load loop / equipment, as well as the controllable load adjustment range constraints and process / comfort constraint parameters defined by the resource digital twin and constraint modeling module; The upstream scheduling and algorithm data includes the initial load adjustment decision Δ The offline training parameters and online solution support for particle swarm optimization and ant colony algorithm from intelligent optimization control module 4; The constraint model data includes controllable load adjustment range constraints. Energy conservation constraints for peak-shifting demand response, and penalty rules for failure to complete response instructions.

[0077] The three-layer fusion optimization architecture consists of a continuous decision-making layer, a discrete decision-making layer, and a collaborative allocation layer. The continuous decision-making layer employs a particle swarm optimization algorithm to solve for the optimal load adjustment amplitude at each moment within the scheduling cycle in the time dimension, addressing the global economic optimization problem of continuous quantities. The discrete decision-making layer employs an ant colony algorithm to select the controllable load execution set participating in demand response at each moment in the device dimension, addressing the practical executability problem of discrete resources. The collaborative allocation layer adopts a weighted allocation strategy to scientifically allocate the continuous optimal amplitude obtained from particle swarm optimization within the discrete execution set selected by the ant colony algorithm, forming device-level executable instructions and achieving dual optimization of economy and executability. For the continuous decision variable of demand response adjustment magnitude, a particle swarm optimization algorithm is used to find the global optimum. With the objective of maximizing the net demand response benefit of the industrial park, the optimal load adjustment magnitude at each time point is determined within constraints, including: The demand response magnitude at each time point within the scheduling period T is constructed into a continuous decision vector:

[0078] Each element in the vector The load adjustment at time t, satisfying the basic constraint 0 ≤ Δ ≤ If it is a peak-shifting demand response, meaning the reduction is compensated for at other times, then an additional energy conservation constraint is added:

[0079] in, This is the adjustment amount for the peak-shifting equivalent power. By comprehensively calculating the total revenue and various costs of demand response, a net revenue calculation system is formed, providing a quantitative basis for constructing the fitness function of particle swarm optimization: Total revenue It consists of demand response subsidy revenue and electricity cost savings revenue, as shown in the formula:

[0080] in, To subsidize the unit price, Electricity cost savings resulting from load reduction; Considering the impact of load regulation on the park's production processes and living comfort, the load impact cost is defined. The formula is:

[0081] in, Comfort / Process Influence Coefficient; Penalties are set for failure to fulfill grid demand response instructions, and the cost of default penalties is defined. The formula is:

[0082] in, This is the penalty coefficient for breach of contract. The power quantity of the incomplete response; The net income expression is defined as follows:

[0083] The goal of maximizing net profit is transformed into minimizing it. A fitness function is defined for particle swarm optimization. A smaller fitness function value indicates better economic efficiency of the demand response solution. The specific expression for the fitness function is:

[0084] Using the continuous decision vector y as the position vector of the particle swarm, a global optimum search is performed using the standard particle swarm iteration update formula, which is:

[0085]

[0086] in, Let be the velocity vector of the i-th particle in the k-th iteration. For position vectors, This is the optimal position in the particle's history. This represents the group's historically optimal position. The position vector obtained in each iteration is truncated to ensure that all elements satisfy 0 ≤ ≤ The final output is the optimal load adjustment amplitude within the scheduling cycle. ; For the discrete decision problem of device / loop level execution selection, an ant colony algorithm is used to select the optimal set of controllable load executions at each time point, balancing executability and optimal unit adjustment cost. Specifically, this includes: All controllable loads within the park are constructed as a discrete set: Each element corresponds to a controllable load loop / device; the decision objective of the ant colony algorithm at time t is defined as selecting a subset from set L. As the execution device that participates in demand response at that moment; Define the basic parameters of the ant colony algorithm to provide a decision-making basis for the selection of discrete sets, specifically including: Define pheromone concentration , represents the pheromone concentration at time t that selects load j to participate in the demand response. The higher the concentration, the greater the probability of being selected by the ants. Define heuristic factors This reflects the degree of matching between the benefits and costs of selecting load j, and is defined as the ratio of unit benefit to unit impact cost, as shown in the formula:

[0087] in, Let j be the specific impact cost coefficient for load j. To prevent small positive numbers from being divided by zero; Pheromones volatile coefficient The pheromone constant Q and the heuristic factor weights α / β are all parameters calibrated offline by the intelligent optimization control module 4 and dynamically fine-tuned according to real-time operating conditions. The probability that the k-th ant will choose load j to participate in the demand response at time t is determined by both the pheromone concentration and the heuristic factor. The formula for calculating the ant selection probability is:

[0088] Where α is the pheromone weight and β is the heuristic factor weight, which respectively determine the degree of influence of pheromone and heuristic factor on the selection probability.

[0089] After each iteration, the pheromone concentration is globally updated based on the cost of the chosen subset of actions for each ant, including pheromone evaporation and incremental replenishment. The formula for updating the pheromone concentration is as follows:

[0090] in, This is the pheromone evaporation coefficient, used to reduce the pheromone concentration of ineffective selections and prevent the algorithm from getting trapped in local optima; The pheromone increment of the k-th ant on load j is defined as:

[0091] Where Q is the pheromone constant. The cost of the solution corresponding to the selected subset of execution for the k-th ant; the smaller the solution cost, the greater the pheromone increment. Through multiple iterations, the pheromone converges to the optimal execution subset, ultimately outputting the optimal discrete execution set of the demand response at time t. This ensures that the equipment set meets the adjustment capacity while minimizing the unit adjustment cost. The optimal load adjustment amplitude was obtained through particle swarm optimization. Discrete optimal execution set of ant colony algorithm Then, a weighted allocation strategy is used to determine the optimal load adjustment at each time point. In executing collection The equipment within the facility is scientifically allocated to generate equipment-level demand response instructions. The specific allocation rules are as follows: The allocation weight is based on the proportion of each equipment's adjustable upper limit or the reciprocal of its unit impact cost. Priority is given to allocating more adjustment to equipment with a large adjustable upper limit and a low unit impact cost, taking into account both adjustment potential and the impact on park operation. Let time t be the execution set There are m devices, and the adjustable upper limit of device j is . The weights are assigned as follows Then the adjustment amount of device j is:

[0092] The allocated adjustment amount for each device is then verified a second time to ensure that 0 ≤ ≤ If the constraint is exceeded, a truncation correction is performed, and the remaining adjustment is redistributed to other devices to ensure that the total adjustment is consistent with the constraint. Consistent.

[0093] Furthermore, the intelligent optimization control module 4 integrates deep reinforcement learning, particle swarm optimization, fuzzy control, and ant colony algorithm to achieve unified algorithm training, collaborative solution, and dynamic parameter tuning. It outputs the optimal algorithm parameters, scheduling strategies, and control schemes adapted to the park's operating conditions for each business module, and serves as a supporting carrier for realizing intelligent optimization of the entire process of the park's virtual power plant. The input data for the intelligent optimization control module 4 includes basic data for algorithm training, system constraint model data, algorithm requirement data for each business module, and real-time operating condition feedback data. Among them, the basic data for algorithm training includes historical operating data of park energy resources, historical data of the electricity market, and historical monitoring data of meteorological environment; the system constraint model data includes power boundary, energy storage SOC constraint, grid-connected power constraint, and controllable load regulation constraint; the algorithm requirement data for each business module includes reinforcement learning training, parameter tuning, and joint solution requirements; and the real-time operating condition feedback data includes data on algorithm execution effect and instruction execution deviation.

[0094] The intelligent optimization control module 4 encapsulates each algorithm into a deep reinforcement learning unit, a particle swarm optimization unit, a fuzzy control unit, and an ant colony algorithm unit. Each algorithm unit has independent offline training and online solving functions, and achieves collaborative calling and joint solving through standardized interfaces to adapt to the differentiated needs of different business modules.

[0095] The deep reinforcement learning unit adopts a centralized evaluation and distributed execution architecture, providing policy training and inference support for the multi-agent collaborative scheduling module 3; The particle swarm optimization unit is used to achieve fuzzy control parameter tuning and continuous amplitude optimization of demand response. The fuzzy control unit is used for real-time compensation control of grid connection fluctuations; The ant colony algorithm unit is used for demand response device-level discrete execution set optimization.

[0096] The intelligent optimization control module 4 adopts an operation mechanism that prioritizes offline training and supplements it with online fine-tuning. It completes basic algorithm training and parameter optimization based on historical data, forming an algorithm model library and a parameter scenario library. After the system goes online, it dynamically fine-tunes the algorithm parameters based on real-time operating conditions and execution feedback data to ensure algorithm adaptability and solution accuracy.

[0097] Simultaneously, a multi-algorithm collaborative solution strategy is constructed, including the collaboration between particle swarm optimization and fuzzy control, the collaboration between particle swarm optimization and ant colony algorithm, and the collaboration between deep reinforcement learning and other algorithms, so as to realize parameter interoperability and result verification among various algorithms and improve the overall optimization efficiency and control accuracy.

[0098] The intelligent optimization control module 4 outputs training models, network parameters, optimal control parameters, and optimization results to the multi-agent collaborative scheduling module 3, the grid connection stability and energy management control module 5, and the power market and demand response optimization module 6, respectively. All outputs use standardized interfaces and are seamlessly connected with each business module.

[0099] Furthermore, the execution and verification module 7 receives control commands output from the grid connection stability and energy management control module 5 and the power market and demand response optimization module 6, connects to various control devices and monitoring terminals in the park, realizes standardized issuance of control commands, real-time monitoring of equipment execution status, accurate detection and correction of execution deviations, and feeds back the execution data of the whole process to the upstream modules to form a closed-loop control. The execution and verification module 7 receives upper-level control commands, basic parameters of field equipment, system constraint boundary data, and real-time monitoring data. The upper-level control commands include final charge / discharge power commands for energy storage, power reduction commands for renewable energy, and equipment-level load adjustment commands. The basic parameters of field equipment include the model, interface protocol, and adjustment accuracy of various energy devices and controllers. The system constraint boundary data comes from the resource digital twin and constraint modeling module 2, serving as rigid boundaries for command verification and deviation correction. The real-time monitoring data comes from field sensors and monitoring terminals, providing a basis for deviation detection. In response to the differentiated characteristics of field device interfaces, the execution and verification module 7 constructs an execution-driven process that standardizes instruction parsing, adapts and converts multiple protocols, and precisely targets and distributes instructions. This process transforms upper-level digital instructions into control signals that can be recognized by the field controller, ensuring the accuracy, real-time performance, and compatibility of instruction distribution. At the same time, the entire instruction distribution process is recorded for traceability, and synchronous triggering control is implemented for the collaborative execution of instructions by multiple devices to avoid operational risks caused by timing differences.

[0100] The execution and verification module 7 constructs a comprehensive execution status monitoring system, which collects the device command execution status and safe operation status in real time, records the entire process data of command execution, forms a traceable execution archive, and displays it through a visual interface, supporting manual intervention and emergency response by operation and maintenance personnel; if the device malfunctions or exceeds the safety threshold, it immediately triggers an early warning and suspends the execution of relevant commands.

[0101] After deviation correction, the correction effect is monitored in real time until the deviation is eliminated; at the same time, the deviation detection and correction-related data are fed back to the upstream modules to provide a basis for prediction model updates, scheduling strategy optimization and algorithm parameter tuning, thereby promoting the overall optimization capability of the system.

[0102] Example 2: This invention provides a method for resource collaborative scheduling and demand response optimization of virtual power plants in a park. It employs a rolling time-domain online operation mode, and by setting a rolling window length H, it achieves dynamic online optimization of resource collaborative scheduling and demand response for virtual power plants in the park, balancing the global optimality of the scheduling strategy with the real-time adaptability of execution; such as Figure 2 As shown, the specific steps include: Step S1: At the current running time t, the data acquisition and prediction module 1 completes real-time data acquisition and updates the predicted values ​​within the rolling window, providing the latest data support for subsequent optimization decisions. Collect real-time measurements of the park's energy system, including grid connection point voltage / frequency V(t) and f(t), energy storage SOCs(t), and actual renewable energy output. Real-time load of the park Grid-connected switching power ; Generate predicted values ​​within a rolling window [t, t+H], update the prediction model based on the latest measured data, and output the predicted values ​​of renewable energy output within the window. Forecast values ​​of park load power The real-time / time-of-use electricity price forecast π(t:t+H) for the electricity market, and the demand response command D(t:t+H) issued by the power grid; After unifying the format of the measured values ​​and the predicted values, they are synchronously output to the multi-agent collaborative scheduling module 3, the grid connection stability and energy management control module 5, and the electricity market and demand response optimization module 6. Step S2: The multi-agent cooperative scheduling module 3, based on the measured / predicted data from step S1, completes global cooperative decision-making within a rolling window [t, t+H], and only outputs feasible joint actions at the current time t: Each agent is based on its local observation state Output the candidate action sequence within the scrolling window. (t:t+H), extract candidate actions at the current time. ; Using a centralized constraint projector Or a Lagrange penalty coordination mechanism, Mapping to the feasible domain Ω defined in the resource digital twin and constraint modeling module 2, we obtain the feasible joint action a(t) at the current moment that satisfies the full-dimensional operational constraint model; Decompose a(t) into the energy storage plan's charging and discharging power. Initial decision-making for load regulation Renewable energy power reduction The outputs are respectively sent to the grid connection stability and energy management control module 5 and the electricity market and demand response optimization module 6; Step S3: Based on the planned energy storage power from step S2, and combined with the fuzzy controller tuned by particle swarm optimization (PSO), the grid connection stability and energy management control module 5 completes the current-moment grid connection fluctuation compensation and final energy storage command generation. Calculate the voltage error at the grid connection point Power error Based on the optimal fuzzy control parameters after offline PSO tuning and online fine-tuning, fuzzy inference is performed on the error to output the energy storage power compensation amount at the current moment. ; Integrating planned and compensated power, the system generates a final charge / discharge power command for energy storage, verifies it, and ensures that the command meets the energy storage SOC and charge / discharge power boundary constraints. , Output to Execution and Verification Module 7.

[0103] Step S4: Based on the initial load regulation decision in step S2, the electricity market and demand response optimization module 6 completes the continuous amplitude and discrete resource optimization of the demand response at the current moment within a rolling window. With the objective of maximizing net demand response revenue, the PSO algorithm is used to solve for the optimal continuous amplitude sequence within a rolling window. (t:t+H), extract the optimal amplitude at the current moment. With the goal of minimizing unit adjustment cost, the ant colony optimization (ACO) algorithm is used to select the optimal execution subset from the controllable load set at the current moment. ;Will exist Internal weighted allocation generates equipment-level load adjustment commands. Verify and ensure that load regulation constraints are met; Will ,Right now Output to execution and verification module 7; Step S5: Execution and verification module 7 only executes the control instructions at the current time t, and completes the execution status monitoring and result recording. It does not execute the instructions in the time period from t+1 to t+H within the scrolling window. Will , , The command is converted into a control signal that can be recognized by the field equipment and sent to the field equipment; the field equipment strictly executes the command at time t and suspends the execution of commands in subsequent time periods within the window; the actual execution parameters of the equipment are collected in real time, and the command execution deviation, equipment operating status, and grid-connected electrical quantity changes are recorded; Step S6: Feed back the execution results and deviation information of step S5 to each module, update the system status variables, and repeat steps S1 to S5 at the next time step t+1 to achieve full-process online closed-loop rolling optimization. The execution deviation, actual equipment operating constraints, and grid-connected stable state are synchronously fed back to the data acquisition and prediction module 1, the multi-agent collaborative scheduling module 3, and the intelligent optimization control module 4. Each module updates its own state variables based on the feedback data. When the system clock advances to t+1, steps S1 to S5 are repeated. Based on the updated measured / predicted data and state variables, the optimization strategy within the rolling window [t+1, t+1+H] is recalculated. Only the instruction at time t+1 is executed. This process is repeated to achieve continuous online rolling optimization.

[0104] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The methods disclosed in the embodiments are described simply because they correspond to the systems disclosed in the embodiments; relevant details can be found in the method section.

[0105] In the embodiments provided by this invention, it should be understood that the disclosed systems, methods, and approaches can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0106] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit.

[0107] Similarly, in the various embodiments of the present invention, each processing unit can be integrated into a functional module, or each processing unit can exist physically, or two or more processing units can be integrated into a functional module.

[0108] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0109] The above-disclosed embodiments are merely preferred embodiments of the present invention, but the present invention is not limited thereto. Any non-creative variations that can be conceived by those skilled in the art, as well as any improvements and modifications made without departing from the principles of the present invention, should fall within the protection scope of the present invention.

Claims

1. A resource collaborative scheduling and demand response optimization system for virtual power plants in a park, characterized in that, include: The data acquisition and forecasting module is used to collect real-time status information of energy resources within the park and information on the external environment and electricity market, and to generate forecasts. The resource digital twin and constraint modeling module constructs digital twin models and full-dimensional operation constraint models for various energy resources in the park based on real-time status information of energy resources within the park and information on the external environment and electricity market. The multi-agent collaborative scheduling module maps various energy resources and market interaction links within the park to multiple agents, constructs a two-layer Actor-Critic network architecture, and, combined with a global collaborative mechanism, outputs the optimal joint action that satisfies the full-dimensional operational constraint model. The optimal joint action includes the energy storage plan's charging and discharging power. The grid connection stability and energy management control module defines the control error amount, executes the fuzzy control strategy with parameters tuned by the particle swarm optimization algorithm, outputs the energy storage power compensation amount, and generates the final charge and discharge power command after merging with the energy storage plan charge and discharge power. The power market and demand response optimization module is used to build a three-layer integrated optimization architecture. It obtains the optimal load regulation amplitude and the optimal discrete execution set based on particle swarm optimization algorithm and ant colony algorithm, and generates equipment-level demand response instructions through a weighted allocation strategy. The intelligent optimization control module provides algorithm training, parameter tuning, and collaborative solution support for the multi-agent collaborative scheduling module, the grid connection stability and energy management control module, and the power market and demand response optimization module. The execution and verification module is used to execute the final charge and discharge power commands and device-level demand response commands, and to monitor and provide feedback on the execution status.

2. The system according to claim 1, characterized in that, In the data acquisition and prediction module, the real-time status information of energy resources in the park includes the operating status data of renewable energy equipment, the operating status data of energy storage system, the operating status data of distributed power source, the operating status data of controllable load, and the operating status data of grid connection point; The external environment and electricity market information includes external meteorological environment data and real-time electricity market signal data; The forecasts include renewable energy output forecasts, park load forecasts, demand response command forecasts, and electricity price trend forecasts. The renewable energy output forecast includes photovoltaic power forecast and wind power forecast; the park load forecast includes the total park load forecast.

3. The system according to claim 2, characterized in that, In the resource digital twin and constraint modeling module, the full-dimensional operational constraint model includes a power boundary constraint sub-model, a ramp constraint sub-model, a SOC constraint sub-model, a grid-connected power constraint sub-model, a renewable reduction constraint sub-model, a controllable load adjustment range constraint sub-model, and a time-series constraint sub-model. The power boundary constraint sub-model defines the upper and lower limits of energy resource power, including the maximum charging power and maximum discharging power of energy storage, the available output boundary of distributed power sources, and the real-time adjustable upper limit of controllable loads. The SOC constraint sub-model defines the boundary range of the State of Charge (SOC); the grid-connected power constraint sub-model defines the boundary of grid-connected exchange power; the controllable load regulation range constraint sub-model defines the boundary of load regulation amount; and the renewable energy reduction constraint sub-model defines the boundary of renewable energy power reduction.

4. The system according to claim 3, characterized in that, In the multi-agent cooperative scheduling module, the global cooperative mechanism includes a centralized constraint projector mechanism and a Lagrange penalty coordination mechanism; The centralized constraint projector mechanism employs a feasible region projection operator in the sense of Euclidean distance. The candidate action vectors output by each agent Centralized projection processing will exceed the global feasible region. The candidate actions are mapped to the feasible region to obtain the optimal joint action, which is expressed mathematically as follows: The Lagrange penalty coordination mechanism incorporates a constraint violation penalty term into the agent's reward function. Through a reinforcement learning reward mechanism, it guides each agent to autonomously avoid constraint violations, constructing a reward function that includes the penalty term. in, To constrain the penalty coefficient, To constrain the amount of violation; c(t) is the stage cost. Let be the k-th inequality constraint, and .

5. The system according to claim 4, characterized in that, The execution process of the fuzzy control strategy, after parameter tuning by the particle swarm optimization algorithm, includes: The input data is fuzzed. Construct a fuzzy reasoning rule base; Based on the fuzzy inference rule base, the Mamdani inference method is used to infer the fuzzy input and calculate the activation weight of each rule. The weighted average method is used to defuzzify the fuzzy inference results. The calculation formula is as follows: Where M is the total number of fuzzy rules. This represents the output value corresponding to the m-th rule. This is the energy storage power compensation amount. To activate the weights.

6. The system according to claim 5, characterized in that, The fuzzy control strategy employs a particle swarm optimization algorithm to tune key parameters, specifically including: All key parameters of fuzzy control are packaged into a position vector x for particle swarm optimization, including input and output scaling factors. , , With the goals of minimizing voltage and power fluctuations and optimizing energy consumption, a comprehensive quantitative fitness function is constructed. The standard particle swarm optimization formula is used to perform a global optimum search on the parameter vector x. The specific expression is as follows: in, Let be the velocity vector of the i-th particle in the k-th iteration. For position vectors, This is the optimal position in the particle's history. This is the best historical position for the group. For inertial weights, , As a learning factor, , 2 is a uniformly random number in the interval [0,1]. The optimal parameters for fuzzy control are obtained through iterative search. .

7. The system according to claim 6, characterized in that, The three-layer fusion optimization architecture described in the electricity market and demand response optimization module includes a continuous decision-making layer, a discrete decision-making layer, and a collaborative allocation layer; The continuous decision-making layer uses a particle swarm optimization algorithm to solve for the optimal load adjustment amplitude at each time point within the scheduling cycle. The discrete decision-making layer uses the ant colony algorithm to select the optimal discrete execution set that participates in the demand response at each time. The collaborative allocation layer adopts a weighted allocation strategy, which scientifically allocates the optimal load adjustment amplitude obtained by particle swarm optimization within the optimal discrete execution set selected by ant colony algorithm to form equipment-level demand response instructions.

8. The system according to claim 7, characterized in that, The continuous decision-making layer employs a particle swarm optimization algorithm to solve for the optimal load adjustment amplitude at each time point within the scheduling period, specifically including: The demand response magnitude at each moment within the scheduling period T is constructed into a continuous decision vector; A fitness function is constructed with the goal of maximizing net revenue, where net revenue is the difference between total revenue and various costs. The total revenue includes demand response subsidy revenue and electricity cost savings. The various costs include load regulation impact costs and default penalty costs. The continuous decision vector y is used as the position vector for particle swarm optimization, and the global optimal search is performed using the standard particle swarm iterative update formula. The position vector obtained in each iteration is truncated at the boundary, and the resulting position vector satisfies 0 ≤ ≤ The final output is the optimal load adjustment amplitude within the scheduling cycle. .

9. The system according to claim 7, characterized in that, The discrete decision-making layer employs an ant colony algorithm to select the optimal discrete execution set for participating in demand response at each time point, specifically including: All controllable loads within the park are constructed as a discrete set: Each element corresponds to a controllable load circuit or device. Defined at time t, the decision objective of the ant colony algorithm is to select a subset from set L. As the execution device that participates in demand response at that moment; Define the basic parameters of the ant colony algorithm, including pheromone concentration and heuristic factor; calculate the probability of ants selecting load based on pheromone concentration and heuristic factor; After each iteration, the pheromone concentration is globally updated based on the fitness value of the selected execution subset of each ant, including pheromone evaporation and pheromone incremental replenishment. Through multiple iterations, the pheromone concentration converges to the optimal execution subset, ultimately outputting the optimal discrete execution set of demand response at time t. .

10. A method for resource collaborative scheduling and demand response optimization of a virtual power plant in a park, characterized in that, Includes the following steps: Step S1: At the current running time t, collect real-time status information of energy resources in the park and real-time measurement values ​​of external environment and electricity market information, and generate predicted quantities within the rolling window [t, t+H]. Step S2: Based on the real-time measurement and prediction values ​​from Step S1, feasible joint scheduling actions are obtained through multi-agent collaborative scheduling decisions. The feasible joint scheduling actions include planned energy storage charging and discharging power and initial load adjustment decisions; Step S3: Based on the energy storage plan charging and discharging power output in step S2, and combined with the fuzzy control strategy with parameters tuned by the particle swarm optimization algorithm, calculate the energy storage power compensation amount at the current time t, and fuse them to generate the final charging and discharging power command. Step S4: Based on the initial load adjustment decision output in step S2, within the rolling window [t, t+H], the optimal load adjustment amplitude is solved by the particle swarm optimization algorithm, and the optimal discrete execution set is selected by the ant colony algorithm. The equipment-level load adjustment command at the current time t is generated by the weighted allocation strategy. Step S5: Send the final charging and discharging power command generated in step S3 and the equipment-level load adjustment command generated in step S4 to the field equipment and execute only the command at the current time t, while monitoring and recording the command execution status and deviation. Step S6: Use the execution result and deviation information from step S5 as feedback, update the system status, advance the running time to t+1, and return to step S1.