Virtual power plant regulation method, device and equipment

CN122823583APending Publication Date: 2026-09-25ZHUHAI SIMUWAN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610919963.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

虚拟电厂运行数据具有强时序性、强波动性特征,现有方案仅依托实时数据制定瞬时调度决策,无法深度挖掘历史数据的时序关联特征,难以精准预判动态运行态势,工况感知精度不足,无法适配新能源出力、负荷动态波动的复杂场景

Benefits of technology

[0017]本发明提供的技术方案中,在融合长短期记忆网络与时序数据处理技术的基础上,根据进化柔性动作与评价网络开展智能决策,一方面通过对多源运行数据预处理、历史时序特征提取以及实时量测数据融合,综合考量时序运行规律、实时工况与设备运行边界约束,充分挖掘运行信息,保障状态感知的全面性与精准度;另一方面借助进化柔性动作与评价网络输出调度指令并配套校验修正环节,既发挥智能算法在复杂场景下的动态调度优势,又能确保调度指令合规可执行,提升了虚拟电厂多源可控资源调度的智能化水平、响应效率与运行可靠性,适配虚拟电厂复杂多变的运行调控需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122823583A_ABST
    Figure CN122823583A_ABST
Patent Text Reader

Abstract

The present application relates to the field of power system intelligent regulation, and discloses a virtual power plant regulation method, device and equipment, which is used for optimizing intelligent flexible regulation of a virtual power plant, comprising: constructing an evolutionary flexible action and evaluation network, and defining state space dimensions and action space of the evolutionary flexible action and evaluation network; obtaining multi-source operation data in the virtual power plant, and preprocessing to obtain a time series input matrix; inputting historical time series data in the time series input matrix into a pre-trained long short-term memory network to obtain a feature vector of an operation state within a preset time length; performing scale normalization processing and splicing on device constraint parameters, the feature vector and real-time measurement data in the time series input matrix to form a state vector conforming to the state space dimensions, and inputting the state vector into the evolutionary flexible action and evaluation network to obtain a scheduling instruction; and checking and correcting the scheduling instruction to generate an executable scheduling instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control of power systems, and in particular to a virtual power plant control method, apparatus and equipment. Background Technology

[0002] Against the backdrop of the construction of new power systems and the advancement of power market reforms, distributed resources such as wind power, photovoltaics, energy storage, and flexible loads are being connected to the grid on a large scale, and the power system is gradually transforming into a coordinated and flexible control model of "source, grid, load, and storage". Virtual power plants can aggregate various decentralized and controllable resources, coordinate and optimize the dispatch of distributed energy, effectively improve the renewable energy absorption rate and grid operation flexibility, and are a core technology supporting intelligent control and market-oriented operation of the power system. Their application scope continues to expand.

[0003] Currently, virtual power plant (VPS) control relies heavily on traditional optimization algorithms and basic machine learning models, resulting in rigid control methods and significant technical shortcomings. VPS operational data exhibits strong temporal and volatile characteristics. Existing solutions rely solely on real-time data to make instantaneous dispatch decisions, failing to deeply mine the temporal correlations of historical data, making it difficult to accurately predict dynamic operating conditions, and resulting in insufficient precision in condition perception. This makes them unsuitable for complex scenarios involving dynamic fluctuations in renewable energy output and load. Furthermore, traditional models have rigid state and action space definitions, weak flexibility, and often generate dispatch instructions with insufficient rationality and poor implementation. Moreover, existing technologies lack a comprehensive instruction verification and correction mechanism, failing to optimize instructions based on actual equipment and grid operating constraints. This easily leads to mismatches between dispatch instructions and on-site resource conditions, significantly reducing the efficiency of multi-source resource collaborative dispatch, and potentially causing grid fluctuations and control failures, thus hindering the stability of VPS operation.

[0004] In summary, existing control technologies suffer from problems such as weak time-series feature mining capabilities, low sensing accuracy, poor decision-making flexibility, and insufficient command executability, making it difficult to meet the high-precision, high-reliability, and flexible real-time control requirements of virtual power plants in scenarios with a high proportion of renewable energy access.

[0005] Therefore, existing technologies still need improvement and development. Summary of the Invention

[0006] This invention provides a virtual power plant control method, apparatus, and equipment for optimizing intelligent and flexible control of virtual power plants.

[0007] The first aspect of this invention provides a virtual power plant control method, which involves constructing an evolutionary flexible action and evaluation network and defining the state space dimension and action space of the network; acquiring multi-source operating data within the virtual power plant and preprocessing it to obtain a time-series input matrix; inputting historical time-series data from the time-series input matrix into a pre-trained long short-term memory network to obtain feature vectors of the operating state within a preset time period; acquiring equipment constraint parameters; performing scale normalization and concatenation on the equipment constraint parameters, the feature vectors, and real-time measurement data from the time-series input matrix to construct a state vector conforming to the state space dimension; and inputting the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions; and verifying and correcting the scheduling instructions to generate executable scheduling instructions.

[0008] Preferably, the construction of the evolutionary flexible action and evaluation network, and the definition of the state space dimension and action space of the evolutionary flexible action and evaluation network, includes: constructing the network structure of the evolutionary flexible action and evaluation network, the network structure including a policy network, a first value network and a second value network, the policy network being used to output actions based on the input state vector, the first value network and the second value network being used to evaluate the expected reward of the action respectively, and the parameter update of the policy network being guided by the smaller of the expected reward output by the first value network and the second value network; setting the state space dimension according to a predetermined value; and defining the action space according to the range of values ​​and discrete or continuous attributes of the controllable resource's adjustment power, response time, energy storage charging and discharging power or control level.

[0009] Preferably, the construction of the evolutionary flexible action and evaluation network, and the definition of the state space dimension and action space of the evolutionary flexible action and evaluation network, further includes: initializing the parameters of the policy network, the parameters of the first value network, and the parameters of the second value network, and generating initial individuals of the evolutionary population, each individual corresponding to a set of initial parameters or hyperparameters of the policy network; after initialization, the evolutionary iteration process begins, in each round of evolutionary iteration, the policy network corresponding to each individual in the population interacts with the environment according to the input state vector, and obtains cumulative rewards according to a preset reward function, using the cumulative rewards as the individual fitness, and obtaining the fitness evaluation results of each individual in the population; based on the fitness evaluation results, selection, crossover, and mutation operations are performed to generate the next generation of individuals, and an elite retention strategy is adopted to retain the individuals with the highest fitness to the next generation; the above evolutionary iteration process is repeated until a preset convergence condition is reached, and the parameters of the optimal individual are used as the parameters of the policy network.

[0010] Preferably, the step of acquiring multi-source operation data within the virtual power plant and preprocessing it to obtain a time-series input matrix includes: acquiring historical time-series data and real-time measurement data. The historical time-series data includes photovoltaic output, wind power output, industrial interruptible load power, grid frequency, node voltage, active power, reactive power, energy storage state of charge, time-of-use electricity price, meteorological data, historical dispatch instructions, equipment operating status, and execution feedback data. The real-time measurement data includes real-time power, real-time frequency, real-time voltage, and real-time energy storage state of charge. The historical time-series data and real-time measurement data are then processed by timestamp unification and sampling frequency reconstruction to obtain aligned data. Outliers in the aligned data are detected and removed using a physical threshold method to obtain cleaned data. The cleaned data is arranged in chronological order to form a time-series input matrix, where each row of the time-series input matrix corresponds to a sampling time, and each column corresponds to a data feature.

[0011] Preferably, the step of inputting the historical time-series data in the time-series input matrix into a pre-trained long short-term memory network to obtain the feature vector of the running state within a preset time period includes: extracting columns belonging to the historical time-series data from the time-series input matrix to form a historical time-series submatrix; and inputting the historical time-series submatrix as an input sequence into the pre-trained long short-term memory network to obtain the feature vector of the running state within a preset time period.

[0012] Preferably, the step of obtaining device constraint parameters, performing scale normalization processing and concatenation of the device constraint parameters, the feature vector, and the real-time measurement data in the time-series input matrix to construct a state vector conforming to the state space dimension, and inputting the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions, includes: extracting real-time measurement data at the current moment from the time-series input matrix to form a real-time measurement vector; obtaining device constraint parameters and organizing the device constraint parameters into a constraint parameter vector, wherein the device constraint parameters include the operation constraint parameters of various resources in the controllable resources; performing scale normalization processing on the constraint parameter vector, the feature vector, and the real-time measurement vector to obtain three normalized vectors; concatenating the three normalized vectors end to end to form a one-dimensional state vector; verifying whether the dimension of the one-dimensional state vector is equal to the defined state space dimension, and if not, performing zero-padding processing; and inputting the zero-padding one-dimensional state vector as the current state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions.

[0013] Preferably, the step of inputting the zero-padding one-dimensional state vector as the current state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions includes: inputting the zero-padding one-dimensional state vector as the current state vector into the policy network of the evolutionary flexible action and evaluation network to obtain the probability distribution of actions; sampling the original action values ​​from the probability distribution of actions; scaling or mapping the original action values ​​according to the range of the action space to obtain scheduling actions, wherein the scheduling actions include industrial interruptible load adjustment power, response time, energy storage charging and discharging power or control level; and generating scheduling instructions based on the scheduling actions.

[0014] Preferably, the step of verifying and correcting the scheduling instruction to generate an executable scheduling instruction includes: verifying the scheduling instruction against industrial interruptible load constraints, including verifying whether the interruptible power is within a preset range, whether the duration of a single interruption does not exceed the upper limit, whether the number of daily interruptions does not exceed the upper limit, and whether the current time is a critical process period that cannot be interrupted; verifying the scheduling instruction against power grid safety constraints, including verifying whether the frequency deviation is within the allowable range, whether the node voltage exceeds the limit, whether the line power flow exceeds the limit, and whether the system power is balanced; if the scheduling instruction violates any constraint, the parameter value of the violated constraint in the scheduling instruction is adjusted to the nearest feasible boundary value of the violated constraint to obtain a verification instruction; and using a moving average filter to smooth and correct the verification instruction, and using the smoothed and corrected verification instruction as an executable scheduling instruction.

[0015] A second aspect of the present invention provides a virtual power plant control device, comprising: a construction module for constructing an evolutionary flexible action and evaluation network, and defining the state space dimension and action space of the evolutionary flexible action and evaluation network; a preprocessing module for acquiring multi-source operating data within the virtual power plant and performing preprocessing to obtain a time-series input matrix; a prediction module for inputting historical time-series data from the time-series input matrix into a pre-trained long short-term memory network to obtain a feature vector of the operating state within a preset time period; a scheduling module for acquiring equipment constraint parameters, performing scale normalization processing and concatenation of the equipment constraint parameters, the feature vector, and real-time measurement data from the time-series input matrix to construct a state vector conforming to the state space dimension, and inputting the state vector into the evolutionary flexible action and evaluation network to obtain a scheduling instruction; and a correction module for verifying and correcting the scheduling instruction to generate an executable scheduling instruction.

[0016] A third aspect of the present invention provides a virtual power plant control device, comprising: a memory and at least one processor, wherein the memory stores computer-readable instructions, and the memory and the at least one processor are interconnected via a line; the at least one processor invokes the computer-readable instructions in the memory to cause the virtual power plant control device to perform the various steps of the virtual power plant control method described above.

[0017] The technical solution provided by this invention, based on the integration of long short-term memory networks and time-series data processing technology, conducts intelligent decision-making according to evolutionary flexible action and evaluation networks. On the one hand, by preprocessing multi-source operating data, extracting historical time-series features, and fusing real-time measurement data, it comprehensively considers the time-series operating rules, real-time operating conditions, and equipment operating boundary constraints, fully mining operating information and ensuring the comprehensiveness and accuracy of state perception. On the other hand, by using the evolutionary flexible action and evaluation network to output scheduling instructions and matching them with verification and correction links, it not only leverages the dynamic scheduling advantages of intelligent algorithms in complex scenarios, but also ensures that scheduling instructions are compliant and executable, improving the intelligence level, response efficiency, and operational reliability of multi-source controllable resource scheduling in virtual power plants, and adapting to the complex and ever-changing operation and control needs of virtual power plants. Attached Figure Description

[0018] Figure 1 A flowchart of the virtual power plant control method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of the virtual power plant control device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the virtual power plant control equipment provided in an embodiment of the present invention. Detailed Implementation

[0019] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” or “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0020] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 A virtual power plant control method in this embodiment of the invention includes: S101. Construct an evolutionary flexible action and evaluation network, and define the state space dimension and action space of the evolutionary flexible action and evaluation network; S102. Obtain multi-source operation data within the virtual power plant and perform preprocessing to obtain the time-series input matrix; S103. Input the historical time series data in the time series input matrix into the pre-trained long short-term memory network to obtain the feature vector of the running state within the preset time period; S104. Obtain device constraint parameters, perform scale normalization processing and splicing on the device constraint parameters, the feature vector and the real-time measurement data in the time series input matrix to construct a state vector that conforms to the state space dimension, and input the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions; S105. Verify and correct the scheduling instruction to generate an executable scheduling instruction.

[0021] It is understood that the executing entity of this invention can be a virtual power plant control device, a terminal, or a server; no specific limitation is made here. This embodiment of the invention will be described using a server as an example.

[0022] This embodiment provides a virtual power plant control method. Based on the integration of long short-term memory networks and time-series data processing technology, it conducts intelligent decision-making based on an evolutionary flexible action and evaluation network. On the one hand, by preprocessing multi-source operating data, extracting historical time-series features, and fusing real-time measurement data, it comprehensively considers the time-series operating rules, real-time operating conditions, and equipment operating boundary constraints to fully mine operating information and ensure the comprehensiveness and accuracy of state perception. On the other hand, by using the evolutionary flexible action and evaluation network to output scheduling instructions and provide corresponding verification and correction links, it not only leverages the dynamic scheduling advantages of intelligent algorithms in complex scenarios but also ensures that scheduling instructions are compliant and executable. This improves the intelligence level, response efficiency, and operational reliability of multi-source controllable resource scheduling in virtual power plants, adapting to the complex and ever-changing operation and control needs of virtual power plants.

[0023] In this embodiment, step S101 involves constructing an evolutionary flexible action and evaluation network and defining its state space dimension and action space. This includes: constructing the network structure of the evolutionary flexible action and evaluation network, which includes a policy network, a first value network, and a second value network. The policy network outputs actions based on the input state vector, and the first and second value networks evaluate the expected returns of the actions. The parameter updates of the policy network are guided by the smaller of the expected returns output by the first and second value networks. The state space dimension is set according to a predetermined value. The action space is defined based on the range of values ​​and discrete or continuous attributes of the controllable resource's adjustment power, response time, energy storage charging and discharging power, or control level. The parameters of the policy network, the first value network, and the second value network are initialized, and initial individuals of the evolutionary population are generated, with each individual corresponding to a set of initial parameters or hyperparameters of the policy network.

[0024] In this embodiment, the Evolutionary Flexible Action and Evaluation Network (SAC) is a deep reinforcement learning network architecture that integrates evolutionary policies and flexible action evaluation algorithms. The construction of the network structure first requires determining that the deep reinforcement learning framework used is Flexible Action and Evaluation (SAC). This framework encourages the maximization of policy entropy while maximizing cumulative reward, thereby improving exploration capabilities. Specifically, a policy network (Actor) and a first value network and a second value network (Critic1 and Critic2) are established. The number of nodes in the input layer of the policy network is determined by the state space dimension, and the number of nodes in the output layer is determined by the action space dimension. The intermediate hidden layers are typically set to 2 to 3 fully connected layers, each containing 256 or 512 neurons, with ReLU or SELU activation functions. The input of the value network is a concatenation of the state vector and the action vector, and the output is a scalar value representing the expected reward for performing the action in the current state. The two value networks have identical structures but their parameters are initialized independently. During training, the smaller value of the two outputs is taken as the target Q-value to mitigate overestimation bias. After the network structure is constructed, the weight dimension, activation function type, and optimizer configuration (such as the Adam learning rate) of each layer need to be recorded. This structured definition ensures that the network has sufficient expressive power and training stability.

[0025] In this embodiment, the setting of the state space dimension directly affects the scale of the network input and the computational complexity of subsequent training. This embodiment employs a deterministic pre-defined value setting method. First, the dimension of the feature vector output by the Long Short-Term Memory (LSTM) network is determined. This dimension is equal to the number of neurons in the LSTM hidden layer, typically set to 64, 128, or 256, with the specific value chosen based on the time window length and feature richness of historical data. Then, the number of features in the real-time measurement data is counted, such as real-time power (1-dimensional), real-time frequency (1-dimensional), real-time voltage (3-dimensional), real-time energy storage SOC (1-dimensional), and other possible rapidly changing quantities (such as temperature and wind speed), assuming a total of M dimensions. Next, the dimension of the equipment constraint parameter vector is counted, denoted as C, which includes the interruptible power range of the industrial interruptible load, the upper limit of a single interruption duration, the upper limit of the number of daily interruptions, the uninterruptible periods for critical processes, the upper and lower limits of energy storage SOC, and the charging and discharging power limits, etc. The state space dimension is the sum of the LSTM feature vector dimension, the real-time measurement dimension, and the constraint parameter dimension. If M changes at certain times in actual engineering (e.g., sensor malfunction leading to partial data loss), the dimension is kept constant during the stitching stage by padding with zeros or interpolation. This pre-defined value setting method avoids the inefficiency of determining the state space dimension through repeated trials in traditional methods, and ensures that the complementarity between the depth features extracted by LSTM and the original measurement information is not destroyed by the dimensionality reduction operation. At the code implementation level, this dimension value is passed in as a constant when building the network and hard-coded in the input layer of the policy network.

[0026] In this embodiment, the action space is defined based on the adjustable power, response time, energy storage charging / discharging power, or the range of controllable levels of controllable resources. Controllable resources include at least one of industrial interruptible loads, energy storage, charging piles, air conditioning loads, and distributed power sources. The definition of the action space needs to accurately map the adjustment capabilities of each controllable resource in the virtual power plant. First, for industrial interruptible loads (ILES), their power adjustment range needs to be clearly defined. For example, the interruptible power range of a certain air compressor is [0, 200] kW, with an adjustment step size of 1 kW (continuous action) or only 5 levels (discrete action). The response time refers to the time from the issuance of the interruption command to the actual reduction in power, generally set to the second or minute level, as a dimension of the action. Second, the range of charging / discharging power of the energy storage system is determined by the capacity of the energy storage converter, for example, [-100 kW, 100 kW] (negative values ​​indicate charging, and positive values ​​indicate discharging). For control levels, if there are multi-level control devices, they can be defined as a discrete action space, such as {0, 1, 2} representing no control, level one control, and level two control, respectively. When defining this space, it's necessary to clarify whether each dimension is a continuous or discrete variable. For continuous actions, a Beta or Gaussian distribution is typically used for sampling, constrained to [-1, 1] using a tanh function, and then linearly transformed to the target interval. For discrete actions, a Softmax layer is used to output the probabilities of each class. Ultimately, all action dimensions combine to form a multi-dimensional action space, with the total dimension being the number of control variables for all independently controllable resources. This step also requires recording the physical unit, value range boundary values, and default values ​​for each dimension for subsequent scaling mapping and constraint verification.

[0027] In this embodiment, after defining the network structure and space, parameter initialization is performed. The weights of the policy network, the first value network, and the second value network are initialized using Xavier or He initialization to ensure the stability of the variance of activation values ​​during forward propagation; the bias term is usually initialized to 0. To support evolutionary optimization, a population needs to be generated, and the population size can be set according to computing resources (e.g., 20 individuals). Each individual includes a set of initial parameters (i.e., weights and biases) for the policy network. These parameters can be completely identical and then have a small amount of noise added, or they can be independently initialized using different random seeds. In addition, individuals can also include hyperparameter mutations, such as a learning rate randomly sampled in the range of [1e-5, 1e-3], and an initial standard deviation of exploration noise randomly sampled in the range of [0.1, 0.5]. During the training phase of the evolutionary flexible action and evaluation network, a preset reward function is defined to guide the optimization direction of the policy network. The reward function includes at least one of the following optimization objectives: renewable energy absorption rate, regulation cost, response deviation, grid safety violation penalty, and equipment constraint violation penalty. The renewable energy absorption rate is used to evaluate the degree of utilization of renewable energy generation by the virtual power plant. The renewable energy absorption rate is calculated as the proportion of actually absorbed renewable energy power generation to the available renewable energy power generation. The higher the absorption rate, the greater the reward. When the energy storage system in the virtual power plant is charging or industrial interruptible loads are increased to absorb surplus wind and solar power, the absorption rate increases, and the reward increases positively. This goal aims to incentivize virtual power plants to prioritize the use of clean energy such as wind and solar power, thereby reducing wind and solar curtailment rates.

[0028] Regulation costs are used to assess the economic costs required to execute dispatch instructions. The calculation of regulation costs includes: compensation costs for production losses caused by power outages to industrial interruptible loads, lifetime depreciation costs due to energy storage system charge-discharge cycles, and incremental fuel or maintenance costs related to distributed power output regulation. Lower regulation costs result in greater rewards, guiding the network to choose the most economically optimal regulation scheme while meeting regulation requirements.

[0029] Response bias is used to evaluate the degree of deviation between the actual execution effect and the scheduling command. Response bias is calculated as the weighted sum of the absolute values ​​of the differences between the actual power change of each controllable resource and the target value of the command; the smaller the bias, the greater the reward. This objective constrains the feasibility of scheduling commands, preventing the network from outputting command values ​​that exceed the actual response capabilities of the devices, while simultaneously encouraging the network to learn the accurate response characteristics of each resource during training.

[0030] Grid safety violation penalties are used to impose negative rewards for potential grid safety risks arising after the execution of dispatch instructions. Grid safety violations include: frequency deviation at the grid connection point exceeding the allowable range (e.g., 49.5Hz–50.5Hz), voltage amplitude exceeding limits at critical nodes (e.g., below 0.93 pu or above 1.07 pu), and feeder power flow exceeding thermal stability limits. When any of the above violations is predicted or measured, a negative reward value proportional to the violation amount is applied; the more severe the violation, the greater the penalty, to guide the network to prioritize ensuring grid operation safety.

[0031] Equipment constraint violation penalties are used to impose negative rewards for behaviors that violate the operational constraints of industrial interruptible loads or energy storage systems themselves. Equipment constraints include: the interruptible power range of industrial interruptible loads, the upper limit of a single interruption duration, the upper limit of the number of daily interruptions, the uninterruptible periods for critical processes, and the upper and lower limits of the state of charge and the charging and discharging power limits of energy storage systems. When a dispatch instruction violates any of the above constraints, a negative reward related to the degree of violation is imposed to guide the network to fully consider the physical limitations of the equipment and production process requirements when making decisions.

[0032] In this embodiment, after construction, the evolutionary flexible action and evaluation network needs to be trained. This involves constructing the evolutionary flexible action and evaluation network and defining its state space dimension and action space. The process then includes: initializing the parameters of the policy network, the first value network, and the second value network; generating initial individuals for the evolutionary population, with each individual corresponding to a set of initial parameters or hyperparameters for the policy network; after initialization, an evolutionary iteration process begins. In each iteration, the policy network corresponding to each individual in the population interacts with the environment based on the input state vector and obtains a cumulative reward according to a preset reward function. This cumulative reward is used as the individual's fitness, resulting in a fitness evaluation result for each individual in the population. Based on the fitness evaluation result, selection, crossover, and mutation operations are performed to generate the next generation of individuals. An elite retention strategy is used to retain the individuals with the highest fitness in the next generation. This evolutionary iteration process is repeated until a preset convergence condition is met. The parameters of the optimal individual are then used as the parameters of the policy network, completing the evolutionary optimization process. This initialization and evolution strategy enables the evolutionary flexible action and evaluation network to find the optimal configuration under various initial conditions, significantly improving the robustness and convergence of the algorithm.

[0033] In this embodiment, a dual-value network structure is adopted, taking the smaller value of the two outputs as the target Q value. This can alleviate the problem of single-network overestimation of action value and improve the accuracy of value assessment. At the same time, the state space dimension is set based on the sum of the LSTM feature vector dimension, real-time measurement dimension, and constraint parameter dimension, ensuring that the state representation does not lose key information or introduce redundancy. In addition, by defining the action space in close conjunction with actual controllable resources, it supports the hybrid expression of continuous and discrete control. Furthermore, the training process of the evolutionary flexible action and evaluation network not only includes conventional parameter initialization but also introduces an evolutionary population generation mechanism. Each individual corresponds to a set of initial parameters or hyperparameters of the policy network, which allows subsequent population-level optimization of the network through evolutionary strategies. This effectively avoids traditional reinforcement learning from getting stuck in local optima and significantly improves the robustness and convergence of the algorithm.

[0034] In this embodiment, step S102 involves collecting multi-source operation data within the virtual power plant and preprocessing it to obtain a time-series input matrix. This includes: acquiring historical time-series data and real-time measurement data. The historical time-series data includes photovoltaic output, wind power output, industrial interruptible load power, grid frequency, node voltage, active power, reactive power, energy storage state of charge, time-of-use electricity price, meteorological data, historical dispatch instructions, equipment operating status, and execution feedback data. The real-time measurement data includes real-time power, real-time frequency, real-time voltage, and real-time energy storage state of charge. The historical time-series data and real-time measurement data are then processed with timestamp unification and sampling frequency reconstruction to obtain aligned data. Anomalies in the aligned data are detected and removed using a physical threshold method to obtain cleaned data. The cleaned data is then arranged in chronological order to form a time-series input matrix. Each row of the time-series input matrix corresponds to a sampling time, and each column corresponds to a data feature.

[0035] In this embodiment, the historical time-series data encompasses the entirety of the virtual power plant's operation, including renewable energy output (power curves for photovoltaic and wind power), real-time power of industrial interruptible loads (reflecting production status), grid-side measurements (frequency, node voltage, active / reactive power), energy storage system status (state of charge, SOC), market information (time-of-use pricing), meteorological data (irradiance, wind speed), control history (historical dispatch commands), and execution feedback (actual equipment response). This data is typically stored in a historical database or SCADA system at second- or minute-level intervals.

[0036] Real-time measurement data focuses on rapidly changing quantities in the current instant, including real-time power (total active and total reactive power), real-time frequency, real-time three-phase voltage, and real-time energy storage SOC. This data typically comes from a PMU or RTU, with sampling frequencies down to the millisecond level, used for rapid response to sudden disturbances. In implementation, each data stream needs a unified naming identifier and unit, and data is retrieved from different subsystems via industrial protocols such as OPCUA, Modbus, or IEC104. To reduce network load, an incremental retrieval method can be used, acquiring only data newly generated since the last acquisition and storing the timestamp as the primary key.

[0037] In this embodiment, due to differences in sampling frequencies among different data sources (e.g., meteorological data every 15 minutes, electrical measurements multiple times per second), time alignment must be performed before the data is fed into the model. First, a unified baseline sampling period is determined, such as 1 second or 5 seconds, depending on the timescale of the virtual power plant's control. Then, for all data streams, linear interpolation is used to interpolate asynchronous data points onto the baseline time grid. For example, the irradiance value for each second between two adjacent points in meteorological data can be linearly interpolated. For switching quantities or discrete states (such as equipment start / stop flags), forward padding (keeping the previous valid value) is used until a new state arrives. Timestamp unification also needs to consider time zone issues and daylight saving time switching, uniformly converting to UTC time. In the code implementation, the `resample` and `interpolate` methods of the Pandas library can be used for this. After resampling, all data streams have the same number of rows on the time axis, with each row corresponding to the same time point, and columns corresponding to different features.

[0038] In this embodiment, industrial operation data frequently contains outliers caused by sensor drift, communication packet loss, and equipment failure. These outliers, if left untreated, can severely interfere with the training of LSTM and reinforcement learning. The physical threshold method is an anomaly detection technique based on prior knowledge. Specifically, reasonable minimum and maximum values ​​are defined for each physical quantity. For example, photovoltaic output should be between [0, installed capacity], with negative values ​​or values ​​exceeding 20% ​​of capacity considered anomalies; the grid frequency should be between 49.5Hz and 50.5Hz under normal operating conditions, with values ​​outside this range potentially indicating measurement errors; node voltage is typically between 0.9 and 1.1 times the rated value; and the energy storage SOC must be between 0% and 100%. For detected outliers, one of the following strategies is used: if the outlier is isolated and the preceding and following values ​​are normal, linear interpolation can be used for replacement; if multiple outliers occur consecutively, the data segment is marked as missing and filled with forward or backward padding; for critical measurements (such as frequency), a predictive model (such as Kalman filtering) can be used to estimate the current expected value. The advantage of the physical threshold method is that it is simple and efficient, and the threshold has a clear physical meaning, which makes it easy for engineers to debug.

[0039] In this embodiment, after outlier removal and data cleaning, all valid data are stored according to a unified time base, with each sampling time corresponding to a set of feature vectors. The core task of this step is to organize these discrete, time-indexed data records into a standardized two-dimensional matrix (i.e., the time-series input matrix), which serves as the standard input format for the subsequent LSTM network.

[0040] In practical implementation, the row and column structure of the matrix is ​​first determined. Rows correspond to the time axis, starting from the earliest sampling time and arranged in ascending order until the latest time. The sampling time is determined by the resampling frequency determined in the preprocessing stage. Columns correspond to feature dimensions, including all historical time-series features and real-time measurement features, such as photovoltaic output, wind power output, industrial load power, grid frequency, node voltage, energy storage SOC, time-of-use pricing, meteorological data, historical dispatch instructions, equipment status, execution feedback, and real-time power, frequency, and voltage. Each feature occupies one column, and a fixed column order is maintained between features; for example, historical time-series feature columns are arranged first, followed by real-time measurement feature columns, so that subsequent modules can extract them as needed.

[0041] During data filling, each sampling time is traversed, and the feature values ​​corresponding to that time are written into the corresponding row and column cells of the matrix. If a feature is missing at a certain time due to sensor failure or communication packet loss, a predefined default value (such as 0 or the value of the previous time) is used to fill it, ensuring the integrity of the matrix and avoiding sparse rows.

[0042] The resulting time-series input matrix is ​​a two-dimensional array of shape T×F, where T is the total number of sampling points and F is the total number of features. Each row of this time-series input matrix represents a snapshot of the system at a given time, and each column represents the trajectory of a physical quantity over time. This time-series input matrix can be directly used as the sequence input for subsequent LSTM operations. The LSTM will process the data row by row, extracting the temporal dependencies. In practical engineering, this time-series input matrix is ​​typically stored as a NumPy array or a PyTorch tensor.

[0043] In this embodiment, the problem of data alignment due to inconsistent sampling rates of different sensors is solved by unifying timestamps and reconstructing sampling frequencies, enabling cross-frequency data to be used collaboratively on the same time base. Moreover, the physical threshold method is used to detect and remove outliers, which is more engineering interpretable than pure statistical methods and can effectively identify dirty data caused by sensor drift and communication packet loss.

[0044] In this embodiment, step S103 involves inputting historical time-series data from the time-series input matrix into a pre-trained long short-term memory network to obtain a feature vector of the running state within a preset time period. This includes: extracting columns belonging to historical time-series data from the time-series input matrix to form a historical time-series submatrix; and inputting the historical time-series submatrix as an input sequence into the pre-trained long short-term memory network to obtain a feature vector of the running state within a preset time period.

[0045] In this embodiment, the time-series input matrix includes both historical time-series data and real-time measurement data as features. However, LSTM only needs historical time-series data to extract dynamic features. Therefore, column extraction is required based on feature names or a preset column index range. Specifically, a feature list is maintained during the data preprocessing stage, where historical time-series feature fields are labeled "history" and real-time measurement fields are labeled "real_time". In the code, data corresponding to these columns is extracted from the time-series input matrix by filtering by column name or by integer index slicing. Note that the time dimension of the time-series input matrix is ​​T (window length), and the feature dimension is F. total After extraction, a new matrix is ​​obtained, namely the historical time series submatrix, with dimensions T×F. history , where F history This refers to the number of historical time-series features. This step involves no numerical calculations; it's simply a view partitioning of the data. In practical engineering, to improve efficiency, historical time-series features and real-time measurement features can be pre-stored in adjacent column blocks. This way, extraction only requires a simple slicing operation, avoiding dynamic column lookups. Furthermore, for online real-time inference, only the latest window of data is needed each time, so slicing can be done directly from the output of the sliding window.

[0046] In this embodiment, the historical time-series submatrix is ​​input into a pre-trained LSTM network. The LSTM network has been trained offline using a large amount of historical data, and its parameters are frozen and do not participate in online gradient updates. During computation, a row of feature vectors is input into the LSTM unit row by row (i.e., at each time step). The internal computation of the LSTM includes a forget gate that determines which old information to discard, an input gate that determines which new information to store, and an output gate that determines the output of the current hidden state. After T time steps, the last hidden state h_T of the LSTM has incorporated the sequence information within the entire window and is typically used as the output feature vector. Alternatively, another strategy can be adopted, such as calculating the mean or maximum pooling of the hidden states across all time steps to obtain a global feature vector. The choice of method depends on the specific task; if the scheduling decision relies more on the final state, the last hidden state is used; if comprehensive information from the entire window is required, pooling is used. Regardless of the method, the dimension of the output feature vector is equal to the number of LSTM hidden units, for example, 128 dimensions. This feature vector is then passed to the subsequent reinforcement learning module. Since LSTMs are pre-trained and have fixed parameters, forward computation is very fast, meeting millisecond-level online decision-making requirements. During implementation, the context manager of deep learning frameworks (such as PyTorch or TensorFlow) can be used to disable gradient computation, further reducing memory overhead.

[0047] In this embodiment, an offline pre-trained LSTM is used to specifically learn the nonlinear dynamic relationship between wind and solar power output, industrial load, and grid status. Parameters are frozen during online scheduling, ensuring the stability and computational efficiency of feature extraction. Compared with directly inputting the original time series data into reinforcement learning, this scheme significantly reduces the dimensionality of the state space, for example, compressing it from 1500 dimensions to 128 dimensions, which greatly alleviates the sample efficiency problem in high-dimensional state spaces.

[0048] In this embodiment, step S104 involves obtaining device constraint parameters, performing scale normalization processing and concatenating the device constraint parameters, the feature vector, and the real-time measurement data in the time-series input matrix to construct a state vector conforming to the state space dimension, and inputting the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions. This includes: extracting real-time measurement data from the time-series input matrix to form a real-time measurement vector; obtaining device constraint parameters and organizing them into a constraint parameter vector, where the device constraint parameters include the operational constraint parameters of various resources in the controllable resources; performing scale normalization processing on the constraint parameter vector, the feature vector, and the real-time measurement vector to obtain three normalized vectors; concatenating the three normalized vectors end-to-end to form a one-dimensional state vector; verifying whether the dimension of the one-dimensional state vector is equal to the defined state space dimension; if not, padding with zeros; and inputting the zero-padding one-dimensional state vector as the current state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions.

[0049] In this embodiment, each time step in the time-series input matrix corresponds to a set of feature values. The real-time measurement data at the current moment is the most critical information describing the current operating condition of the system. Therefore, it is necessary to extract the feature columns belonging to the real-time measurement type from the last row of the time-series input matrix. A column index list real_time_indices can be predefined, and then the values ​​of these columns can be extracted from the last row of the matrix (index -1) to form a one-dimensional array, i.e., the real-time measurement vector, with a length of M. Each element in this vector has a clear physical meaning, such as [real-time power, real-time frequency, real-time voltage_phase A, real-time voltage_phase B, real-time voltage_phase C, real-time energy storage SOC]. Since the sampling frequency of real-time measurement data is usually higher than the resampling frequency of historical data, the existence of the current moment is guaranteed during sliding window slicing. If a real-time measurement is not received at a certain moment due to a communication failure, it is necessary to fill it with the valid value of the previous moment or replace it with the predicted value. The length of the output vector of this step is denoted as M.

[0050] In this embodiment, for industrial interruptible loads, the constraint parameters include the interruptible power range, the upper limit of a single interruption duration, the upper limit of the number of daily interruptions, and the non-interruptible period for critical processes; for energy storage systems, the constraint parameters include upper and lower limits of state of charge and charging / discharging power limits; for charging piles, air conditioning loads, or distributed power sources, the constraint parameters include at least one of their respective power adjustment range, response time, and operating status limits. These parameters are organized into a one-dimensional constraint parameter vector with a length denoted as C. Switching constraints (such as whether it is in a critical period) are represented by 0 / 1.

[0051] Then, the three vectors are normalized. The feature vector (length L) output by the LSTM is batch normalized or layer normalized to make its mean 0 and variance 1. The real-time measurement vector is standardized using the mean and standard deviation of historical statistics using Z-score. The constraint parameter vector, since its numerical range is known (e.g., the power range boundary is 0~200kW), is directly linearly mapped to the [0,1] interval. In this embodiment, the total state space dimension is set to D=L+M+C, and no dimensionality adjustment is performed, only normalization is applied.

[0052] In this embodiment, after normalization, the vectors are concatenated end-to-end in the order from LSTM feature vectors to real-time measurement vectors and then to constraint parameter vectors, forming a longer one-dimensional vector. The concatenation order is adjusted according to the actual situation. The final state vector has a dimension of L+M+C (assuming no dimensionality adjustment is performed). This state vector integrates historical dynamic trends (refined by LSTM), the current instantaneous operating condition (real-time measurement), and the hard constraint boundaries of the device and system (constraint parameters). These three elements are organized in parallel without explicit feature crossing. In the subsequent policy network, the first fully connected layer will automatically capture the nonlinear interaction between these three elements through learned weights.

[0053] In this embodiment, during actual operation, sensor malfunctions, abnormal data acquisition, or configuration file changes may cause the dimension of the concatenated state vector to differ from the dimension D of the state space defined when constructing the network. Without verification, this dimension mismatch will directly lead to a dimension error in the neural network's forward computation. Therefore, this step adds a defensive verification mechanism.

[0054] The actual length of the state vector, len(state), is compared with D. If len(state) < D, several zeros are appended to the end of the vector to make the length D; if len(state) > D, the excess part is truncated (usually from the tail, because features at the end of real-time measurements may have lower priority). While this zero-padding or truncation operation introduces some information loss, it ensures the robustness of the system and prevents the entire scheduling process from crashing due to a small amount of data anomalies. In engineering practice, logs of dimensionality mismatches are recorded, and alarms are triggered so that operations personnel can investigate data source problems. Alternatively, a certain number of redundant dimensions can be reserved when building the network (e.g., D is 10% larger than actually expected). Missing features are padded with zeros, and extra features are absorbed through a learnable mapping layer.

[0055] In this embodiment, the zero-padding one-dimensional state vector is input as the current state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions. This includes: inputting the zero-padding one-dimensional state vector as the current state vector into the policy network of the evolutionary flexible action and evaluation network to obtain the probability distribution of actions; sampling the original action values ​​from the probability distribution of actions; scaling or mapping the original action values ​​according to the range of the action space to obtain scheduling actions, wherein the scheduling actions include industrial interruptible load adjustment power, response time, energy storage charging and discharging power, or control level; and generating scheduling instructions based on the scheduling actions.

[0056] In this embodiment, the policy network is the Actor part responsible for decision-making in the evolutionary SAC. Its input is a one-dimensional state vector (state space dimension D). The network structure typically includes 2-3 fully connected hidden layers, each with 256 neurons, using the ReLU activation function. The design of the final output layer depends on the type of action space: for continuous actions, the output layer has 2×A nodes (A is the action dimension), representing the mean μ and standard deviation σ of each action dimension (usually log standard deviation is used to ensure positive values); for discrete actions, the output layer has K nodes (K is the total number of combinations of all discrete actions), which are transformed into a probability distribution using the Softmax function. In actual implementation, to limit the range of actions, the mean μ of continuous actions is usually compressed to the (-1, 1) interval using the tanh function, and then transformed to the physical boundary using scaling and translation. For the standard deviation, exp(log_std) is used to ensure it is positive, and a minimum standard deviation (e.g., 0.01) is set to prevent premature stagnation of exploration. During forward computation, the state vector is input into the network to obtain the parameters of the action distribution, i.e., the probability distribution of the actions. Then, the original action value a is obtained by sampling from the probability distribution of the action. raw The sampling process employs a reparameterization trick during training to facilitate gradient backpropagation, while during testing, it directly takes the mean to improve determinism. The resulting a... raw It is in the interval [-1, 1] and has not yet been mapped to a physical quantity.

[0057] In this embodiment, because the tanh activation function is used, the original action value a output by the policy network is... raw Located in the range [-1, 1], while actual physical actions, such as the regulation power of industrial interruptible loads, may range from [0, 200] kW to [-100, 100] kW (energy storage charging and discharging). Therefore, a linear mapping is needed to transform a... raw Transform to the physical domain. The formula is: a physical =a low +(a raw +1) / 2×(ahigh -a low ) Where a low and a high These represent the lower and upper bounds of the action dimension, respectively. This also applies to asymmetric intervals such as response time. For discrete actions (e.g., adjusting gears), the policy network outputs the probability of each gear. The gear index is obtained through argmax or probability sampling and then mapped to specific gear parameters (e.g., power percentage). After mapping, the physical action needs to be rounded or truncated to meet the device's step size requirements (e.g., power adjustment can only be done in 10kW steps). The final scheduling action is a multi-dimensional vector, with each dimension having a defined physical unit and value range. This action will be used to interact with the environment (i.e., issue execution) and update the network based on the reward after execution. It is worth noting that during the training phase, the mapped action still needs to have some exploration noise (e.g., Gaussian noise or OU noise) added, while no further noise is added during testing or practical application.

[0058] In this embodiment, after scaling and mapping, the scheduling actions yield target values ​​for each controllable resource. However, these target values ​​still need to be assembled into a standardized scheduling instruction format before being issued. Specifically, the scheduling instruction is a structured message body, typically including the following fields: timestamp (instruction effective time), device ID or device group ID, action type (e.g., power regulation, start / stop control), action value (e.g., -150kW indicates a reduction of 150kW), duration (e.g., 300 seconds), regulation rate limit (e.g., not exceeding 50kW / s), and instruction priority. For ILES (Industrial Interruptible Loads), the instruction also needs to include the interruption reason (e.g., peak shaving and valley filling) and the expected compensation amount. For energy storage systems, the instruction needs to specify the charging / discharging mode (constant power, constant current, or constant voltage) and the termination condition (e.g., stopping when SOC reaches 90%). When generating the instruction, the corresponding fields need to be filled according to the components in the action vector. If multiple similar devices exist, load allocation is also required, for example, distributing the total power regulation proportionally to each device. Finally, the instruction is encapsulated in JSON or binary protocol format, with cyclic redundancy check (CRC) to prevent transmission errors. This instruction is the scheduling instruction.

[0059] In this embodiment, through three sub-steps—normalization, vector concatenation, and dimension verification—it is ensured that the fused state vector retains the deep semantic information of the LSTM output, includes the current key physical measurements, and strictly conforms to the predefined state space dimension. Zero padding is performed when the dimensions do not match to avoid neural network input errors. Finally, the state vector is fed into the evolutionary SAC network to obtain scheduling instructions, forming a complete perception, fusion, and decision-making link, thereby improving the quality of the scheduling strategy.

[0060] In this embodiment, step S105 involves verifying and correcting the scheduling instruction to generate an executable scheduling instruction. This includes: verifying the scheduling instruction against industrial interruptible load constraints, including verifying whether the interruptible power is within a preset range, whether the duration of a single interruption does not exceed the upper limit, whether the number of daily interruptions does not exceed the upper limit, and whether the current time is a critical period during which interruption is not allowed; verifying the scheduling instruction against power grid safety constraints, including verifying whether the frequency deviation is within the allowable range, whether the node voltage exceeds the limit, whether the line power flow exceeds the limit, and whether the system power is balanced; if the scheduling instruction violates any constraint, the parameter value of the violated constraint in the scheduling instruction is adjusted to the nearest feasible boundary value of the violated constraint to obtain a verification instruction; and using a moving average filter to smooth and correct the verification instruction, and using the smoothed and corrected verification instruction as an executable scheduling instruction.

[0061] In this embodiment, the industrial interruptible load constraint verification mainly targets the production process limitations of ILES equipment. First, it verifies whether the interruptible power is within a preset range. Each ILES device has a minimum and maximum interruptible power. For example, some devices require an interruptible power of at least 50kW to be meaningful, and it cannot exceed 80% of its rated capacity. If the requested power is less than the minimum allowable value, it is directly set to 0 (no interruption); if it is greater than the maximum allowable value, it is clamped to the maximum allowable value. Second, it verifies the upper limit of a single interruption duration. For example, a certain refrigeration station allows a maximum of 10 minutes of continuous interruption at a time; exceeding this time may cause overheating and damage to products. The response time or duration in the instruction must be ≤ the upper limit. Third, it verifies the upper limit of daily interruptions. Each ILES device is allowed a maximum number of interruptions per day (e.g., 3 times). It is necessary to query the interruption records executed that day. If the current instruction would cause the upper limit to be exceeded, the instruction is rejected or the interruption depth is reduced. Fourth, it verifies whether the current time is a critical process period where interruption is not allowed. The production plan predefines certain time periods (e.g., 8:00-11:00 is a high-load production period) during which ILES interruptions are prohibited. If the current time falls within this window, the interrupt power in the instruction must be 0. The above verification requires real-time access to the equipment ledger database and production scheduling system, typically through API calls to obtain the latest constraint parameters and status information. Instructions that violate constraints are adjusted according to the "safety first" principle, and the adjustment record is written to the log.

[0062] In this embodiment, the grid security constraint verification aims to prevent dispatch commands from causing the distribution network to exceed limits or become unstable. First, frequency deviation is verified. The virtual power plant should maintain its frequency within the allowable range (e.g., 49.5–50.5 Hz) at the grid connection point. If the current frequency is already close to the lower limit, commands requiring increased discharge or reduced load may further lower the frequency, necessitating an appropriate reduction in the command's adjustment magnitude. This can be implemented dynamically based on the frequency-power droop curve. Second, node voltage is verified. The voltage at each critical node (especially at the feeder end where the ILES is located) should be within the range of 0.93–1.07 pu. Power flow calculations or voltage sensitivity matrices are used to predict voltage changes after command execution. If the predicted voltage exceeds the limit, the adjustment step size is reduced. Third, line power flow is verified. The current power flow of each feeder plus the increment caused by the command is checked to see if it exceeds the thermal stability limit (e.g., cable current carrying capacity). If it does, the command power is reduced or resources from other locations are called upon. Fourth, power balance verification: The difference between the total generation of the virtual power plant (PV + wind power + energy storage discharge) and the total load (ILES + other stationary loads) should equal the power exchanged with the grid, and the exchanged power cannot exceed the grid connection protocol capacity. If the command disrupts the balance, other resources need to be coordinated to compensate for the difference, such as simultaneously adjusting energy storage charging and discharging. These verifications are typically performed quickly based on real-time state estimation and sensitivity matrices, without relying on iterative power flow calculations.

[0063] In this embodiment, when a dispatch instruction is found to violate any industrial or grid constraint, the instruction cannot be simply discarded, as this would cause the virtual power plant to lose its responsiveness to the system. This embodiment employs the principle of minimum intervention, adjusting the parameter values ​​of the violated constraints in the dispatch instruction to the nearest feasible boundary value of the violated constraint. First, a violation severity function for all constraints is defined; for example, the penalty for exceeding power limits is max(0, P). req -P max Then, starting from the original command, a feasible solution is found using gradient descent or binary search, which satisfies all constraints while minimizing the modification to the original command. For single-dimensional power commands, adjustments are directly taken from the constraint boundary values ​​(such as upper or lower limits). For multi-dimensional actions (e.g., simultaneously regulating ILES and energy storage), a linear or quadratic programming problem needs to be solved, with the objective being minimization and constraints consisting of all inequality constraints. Since the action dimensions in a virtual power plant typically do not exceed 10, this optimization problem can be solved quickly. The adjusted command may not achieve the original requested regulation target, but it ensures the safety of the system.

[0064] In this embodiment, even after constraint verification and boundary adjustment, abrupt changes may still exist between adjacent moments in the instruction sequence. For example, one moment a request might be made to reduce ILES by 100kW, and the next moment a request might suddenly be made to increase it by 200kW. Such abrupt changes can cause mechanical shock to industrial equipment (such as compressors), shortening their service life. Therefore, in this embodiment, a moving average filter is used to smooth the verification instructions in the time domain. Specifically, an instruction history queue of length W (e.g., 3 or 5) is maintained. The final instruction at the current moment is equal to the weighted average of the W most recent instructions in the history queue, with weights that can decrease linearly (the more recent the instruction, the larger the weight) or be uniform. If the history is insufficient during the startup phase, the first valid value can be copied to fill the gap. It should be noted that smoothing may introduce response delay, so W should not be too large. In addition, a first-order low-pass filter can also be used to smooth the verification instructions in the time domain. The moving average filter is chosen because it is simple to implement, has no parameter drift, and does not introduce phase lead. After the smoothing correction is completed, the resulting instruction is the final executable scheduling instruction. This instruction is temporarily stored in the instruction buffer, waiting to be issued. When issuing the command, the executable scheduling instruction is encapsulated into a corresponding protocol message according to the communication protocol type of the controllable resource (such as Modbus, IEC 104, MQTT or OPC UA), and transmitted to the local controller or terminal execution unit of each controllable resource through an industrial Ethernet or 5G communication link. After issuance, the execution confirmation signal and real-time status feedback returned by each resource are received. If no confirmation is received within a preset timeout period (such as 500ms), a retransmission mechanism is triggered. After the number of retransmissions reaches the upper limit, an abnormal alarm is generated, and the unsuccessfully issued instructions are marked as pending manual intervention.

[0065] Specifically, the window length of the moving average filter is set according to the response time requirement. The smaller the window length, the faster the response; the larger the window length, the stronger the smoothing effect.

[0066] In this embodiment, a dual verification system of industrial load constraints and power grid safety constraints is constructed to comprehensively screen the compliance of dispatching instructions. It avoids the risk of violations from two core dimensions: equipment operation and power grid operation. For interruptible industrial loads, hard rules such as power, duration, number of interruptions, and production periods are verified to protect the normal production order of enterprises. For the power grid, safety indicators such as frequency, voltage, line power flow, and power balance are verified to ensure the stable operation of the power grid. In addition, for violations, the instructions are automatically corrected to the nearest feasible boundary to achieve compliance processing while preserving the original control strategy as much as possible. Finally, a moving average filter is used to smooth and optimize the corrected instructions, reducing the problems of instantaneous jumps and frequent fluctuations in instructions and preventing control actions from repeatedly impacting energy storage, load and other equipment.

[0067] The virtual power plant control method in the embodiments of the present invention has been described above. The apparatus in the embodiments of the present invention will be described below. Please refer to [link / reference]. Figure 2 The implementation methods of the virtual power plant control device in this invention include: Construction module 201 is used to construct an evolutionary flexible action and evaluation network, and to define the state space dimension and action space of the evolutionary flexible action and evaluation network; The preprocessing module 202 is used to acquire multi-source operating data in the virtual power plant and perform preprocessing to obtain a time-series input matrix; Prediction module 203 is used to input the historical time series data in the time series input matrix into a pre-trained long short-term memory network to obtain the feature vector of the running state within a preset time period; The scheduling module 204 is used to obtain device constraint parameters, perform scale normalization processing and splicing on the device constraint parameters, the feature vector and the real-time measurement data in the time series input matrix to construct a state vector that conforms to the state space dimension, and input the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions; The correction module 205 is used to verify and correct the scheduling instruction and generate an executable scheduling instruction.

[0068] In this embodiment, based on the integration of long short-term memory networks and time-series data processing technology, intelligent decision-making is carried out according to evolutionary flexible action and evaluation networks. On the one hand, by preprocessing multi-source operation data, extracting historical time-series features, and fusing real-time measurement data, the system comprehensively considers the time-series operation rules, real-time operating conditions, and equipment operation boundary constraints to fully mine operation information and ensure the comprehensiveness and accuracy of state perception. On the other hand, by using the evolutionary flexible action and evaluation network to output scheduling instructions and provide corresponding verification and correction links, the system not only leverages the dynamic scheduling advantages of intelligent algorithms in complex scenarios but also ensures that scheduling instructions are compliant and executable. This improves the intelligence level, response efficiency, and operational reliability of multi-source controllable resource scheduling in virtual power plants and adapts to the complex and ever-changing operation and control needs of virtual power plants.

[0069] above Figure 2 The virtual power plant control device in this embodiment of the invention will be described in detail from the perspective of modular functional entities. The virtual power plant control equipment in this embodiment of the invention will be described in detail from the perspective of hardware processing.

[0070] Figure 3This is a schematic diagram of the structure of a virtual power plant control device provided in an embodiment of the present invention. The device 300 can vary significantly due to different configurations or performance characteristics. It may include one or more central processing units (CPUs) 310 (e.g., one or more processors) and a memory 320, and one or more storage media 330 (e.g., one or more mass storage devices) for storing application programs 333 or data 332. The memory 320 and storage media 330 can be temporary or persistent storage. The program stored in the storage media 330 may include one or more modules (not shown), each module including a series of instruction operations on the device 300. Furthermore, the processor 310 may be configured to communicate with the storage media 330 and execute the series of instruction operations in the storage media on the device 300.

[0071] Device 300 may also include one or more power supplies 340, one or more wired or wireless network interfaces 350, one or more input / output interfaces 360, and / or one or more operating systems 331, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.

[0072] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0073] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0074] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A virtual power plant control method, characterized in that, include: Construct an evolutionary flexible action and evaluation network, and define the state space dimension and action space of the evolutionary flexible action and evaluation network; Acquire multi-source operating data within the virtual power plant and preprocess it to obtain the time-series input matrix; The historical time series data in the time series input matrix is ​​input into a pre-trained long short-term memory network to obtain the feature vector of the running state within a preset time period; Obtain device constraint parameters, perform scale normalization processing and concatenation on the device constraint parameters, the feature vector and the real-time measurement data in the time series input matrix to construct a state vector that conforms to the state space dimension, and input the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions; The scheduling instructions are verified and corrected to generate executable scheduling instructions.

2. The virtual power plant control method according to claim 1, characterized in that, The construction of the evolutionary flexible action and evaluation network, and the definition of the state space dimension and action space of the evolutionary flexible action and evaluation network, include: A network structure for constructing an evolutionary flexible action and evaluation network is provided. The network structure includes a policy network, a first value network, and a second value network. The policy network is used to output an action based on the input state vector. The first value network and the second value network are used to evaluate the expected reward of the action, respectively. The parameter update of the policy network is guided by the smaller of the expected reward output by the first value network and the second value network. The state space dimension is set according to a predetermined value; The action space is defined based on the range of values ​​for the controllable resource's adjustment power, response time, energy storage charging and discharging power, or control level, as well as its discrete or continuous attributes.

3. The virtual power plant control method according to claim 2, characterized in that, The process of constructing an evolutionary flexible action and evaluation network, defining the state space dimension and action space of the network, further includes: Initialize the parameters of the policy network, the first value network, and the second value network, and generate the initial individuals of the evolutionary population. Each individual corresponds to a set of initial parameters or hyperparameters of the policy network. After initialization, the evolutionary iteration process begins. In each round of evolutionary iteration, the policy network corresponding to each individual in the population interacts with the environment based on the input state vector and obtains the cumulative reward according to the preset reward function. The cumulative reward is used as the individual fitness to obtain the fitness evaluation result of each individual in the population. Based on the fitness evaluation results, selection, crossover, and mutation operations are performed to generate the next generation of individuals, and an elite retention strategy is adopted to retain the individuals with the highest fitness to the next generation. Repeat the above evolutionary iteration process until the preset convergence condition is met, and use the parameters of the optimal individual as the parameters of the policy network.

4. The virtual power plant control method according to claim 1, characterized in that, The process of acquiring multi-source operational data within the virtual power plant and preprocessing it to obtain a time-series input matrix includes: The system acquires historical time-series data and real-time measurement data. The historical time-series data includes photovoltaic power output, wind power output, industrial interruptible load power, grid frequency, node voltage, active power, reactive power, energy storage status of charge, time-of-use electricity price, meteorological data, historical dispatch instructions, equipment operating status, and execution feedback data. The real-time measurement data includes real-time power, real-time frequency, real-time voltage, and real-time energy storage status of charge. The historical time-series data and real-time measurement data are processed by unifying the timestamps and reconstructing the sampling frequency to obtain aligned data; Outliers in the aligned data are detected and removed using a physical threshold method to obtain cleaned data; The cleaned data is arranged in chronological order to form a time-series input matrix. Each row of the time-series input matrix corresponds to a sampling time, and each column corresponds to a data feature.

5. The virtual power plant control method according to claim 1, characterized in that, The step of inputting historical time-series data from the time-series input matrix into a pre-trained long short-term memory network to obtain a feature vector of the operating state within a preset time period includes: Extract columns belonging to historical time series data from the time series input matrix to form a historical time series submatrix; The historical time series submatrix is ​​used as the input sequence to input the pre-trained long short-term memory network to obtain the feature vector of the running state within a preset time period.

6. The virtual power plant control method according to claim 1, characterized in that, The process of acquiring device constraint parameters involves scaling and concatenating the device constraint parameters, the feature vector, and the real-time measurement data in the time-series input matrix to construct a state vector conforming to the state space dimension. This state vector is then input into an evolutionary flexible action and evaluation network to obtain scheduling instructions, including: Real-time measurement data at the current moment is extracted from the time-series input matrix to form a real-time measurement vector; Obtain the equipment constraint parameters and organize them into a constraint parameter vector. The equipment constraint parameters include the operational constraint parameters of various resources in the controllable resources. The constraint parameter vector, the feature vector, and the real-time measurement vector are scale-normalized to obtain three normalized vectors. The three normalized vectors are concatenated end to end to form a one-dimensional state vector. Check whether the dimension of the one-dimensional state vector is equal to the defined dimension of the state space; if not, pad with zeros. The zero-padding one-dimensional state vector is used as the current state vector and input into the evolutionary flexible action and evaluation network to obtain scheduling instructions.

7. The virtual power plant control method according to claim 6, characterized in that, The step of inputting the zero-padding one-dimensional state vector as the current state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions includes: The zero-padding one-dimensional state vector is used as the current state vector and input into the policy network of the evolutionary flexible action and evaluation network to obtain the probability distribution of the action. The original action value is obtained by sampling from the probability distribution of the action; The original action value is scaled or mapped according to the range of the action space to obtain the scheduling action, which includes industrial interruptible load adjustment power, response time, energy storage charging and discharging power or control level. Generate scheduling instructions based on scheduling actions.

8. The virtual power plant control method according to claim 1, characterized in that, The step of verifying and correcting the scheduling instruction to generate an executable scheduling instruction includes: The scheduling instructions are subjected to industrial interruptible load constraint verification, including verifying whether the interruptible power is within a preset range, whether the duration of a single interruption does not exceed the upper limit, whether the number of daily interruptions does not exceed the upper limit, and whether the current time is a critical process period that cannot be interrupted. The dispatching instructions are subjected to power grid safety constraint verification, including verifying whether the frequency deviation is within the allowable range, whether the node voltage exceeds the limit, whether the line power flow exceeds the limit, and whether the system power is balanced. If the scheduling instruction violates any constraint, the parameter value of the violated constraint in the scheduling instruction is adjusted to the nearest feasible boundary value of the violated constraint to obtain a verification instruction. The verification instruction is smoothed by using a moving average filter, and the smoothed verification instruction is used as an executable scheduling instruction.

9. A virtual power plant control device, characterized in that, include: A construction module is used to construct an evolutionary flexible action and evaluation network, and to define the state space dimension and action space of the evolutionary flexible action and evaluation network; The preprocessing module is used to acquire multi-source operating data within the virtual power plant and perform preprocessing to obtain a time-series input matrix; The prediction module is used to input the historical time series data in the time series input matrix into a pre-trained long short-term memory network to obtain the feature vector of the running state within a preset time period; The scheduling module is used to obtain device constraint parameters, perform scale normalization processing and concatenation on the device constraint parameters, the feature vector and the real-time measurement data in the time-series input matrix to construct a state vector that conforms to the state space dimension, and input the state vector into the evolutionary flexible action and evaluation network to obtain scheduling instructions; The correction module is used to verify and correct the scheduling instructions and generate executable scheduling instructions.

10. A virtual power plant control device, characterized in that, It includes a memory and at least one processor, wherein the memory stores computer-readable instructions; The at least one processor invokes the computer-readable instructions in the memory to perform the steps of the virtual power plant control method as described in any one of claims 1-8.