Industrial virtual power plant multi-time scale regulation and control method based on man-in-the-loop

By improving the density peak clustering algorithm and multi-timescale control method, the problems of inaccurate load clustering and scheduling disconnection in virtual power plants are solved, achieving efficient load control and anomaly response, and improving the system's adaptability.

CN121507977APending Publication Date: 2026-02-10STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511697718.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing virtual power plant load clustering algorithms fail to adapt to the non-convex cluster distribution and fuzzy boundary characteristics of industrial loads, resulting in instructions being detached from actual production, lack of organic connection between multi-timescale scheduling, low anomaly identification rate and high false alarm rate, and inability to update the control model online.

Method used

An improved density peak clustering algorithm is adopted to add a clustering correction step, construct a deep human-machine collaborative model, and combine it with a multi-timescale regulation potential mining method, including day-ahead, intraday, and real-time regulation. Anomaly thresholds and model parameters are optimized through expert feedback to construct a prediction-optimization-feedback closed loop.

Benefits of technology

It improves the accuracy and stability of load clustering, reduces the false alarm rate of anomaly detection and the risk of control command failure, and enhances the system's adaptability and dynamic scenario adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121507977A_ABST
    Figure CN121507977A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial virtual power plant multi-time scale regulation and control method based on a man-in-the-loop, and belongs to the technical field of power system optimization control, and the method comprises the following steps: S1, building an industrial virtual power plant load clustering response model; and S2, establishing a human-in-the-loop-based multi-time scale regulation potential mining method. According to the multi-time-scale regulation and control method for the industrial virtual power plant based on the human-in-the-loop, the accuracy and stability of industrial load clustering are improved, a human-machine deep coevolution mode is constructed, the adaptation capability to a dynamic scene is enhanced, meanwhile, the abnormal detection false alarm rate and the regulation and control instruction failure risk are reduced, and the regulation and control efficiency of the industrial virtual power plant is improved. And the abnormal response and self-adaptive capability of the system are further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system optimization control technology, specifically relating to a multi-timescale control method for industrial virtual power plants based on human-in-the-loop operation. Background Technology

[0002] As the energy transition accelerates, virtual power plants in industrial parks have become a crucial component in the construction of new power systems. For example, in energy-intensive industries, they can effectively mitigate load shocks and reduce energy costs and carbon emissions by aggregating resources such as photovoltaic and waste heat power generation within the plant area. In industrial park scenarios, they can integrate the energy equipment of multiple companies within the park, enabling regional energy collaborative management. In the electricity market, relying on multi-timescale control logic, virtual power plants can participate in various markets such as electricity and ancillary services, thereby achieving synergistic benefits. In new power systems, real-time anomaly detection and rapid command allocation mechanisms can mitigate the volatility of renewable energy sources such as photovoltaic and wind power, ensuring stable system operation.

[0003] However, existing technologies have the following shortcomings: traditional virtual power plants commonly use algorithms such as K-means and standard density peak clustering, which are not adapted to the characteristics of industrial loads that are "non-convex clustered and have fuzzy boundaries"; over-reliance on automation algorithms causes instructions to deviate from actual production; day-ahead, intraday, and real-time scheduling often operate in isolation without organic connection and data linkage; the recognition rate of common anomalies in industrial scenarios is low and the false alarm rate is high; and the control models are mostly statically set and cannot be adjusted and updated online according to changes in load characteristics and market rules.

[0004] Therefore, a new method is urgently needed. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-timescale control method for industrial virtual power plants based on human-in-the-loop operation. This method improves the accuracy and stability of industrial load clustering, establishes a deep human-machine co-evolution model, enhances the adaptability to dynamic scenarios, reduces the false alarm rate of anomaly detection and the risk of control command failure, and further strengthens the system's anomaly response and adaptive capabilities.

[0006] To achieve the above objectives, this invention provides a multi-timescale control method for an industrial virtual power plant based on human-in-the-loop operation, comprising the following steps: S1. Establish an industrial virtual power plant load clustering response model, clarify the core calculation rules of the traditional density peak clustering algorithm, including calculating local density through Euclidean distance, cutoff distance and judgment function of data points, distinguishing whether the local density of data points is the maximum to calculate the relative distance; add a clustering correction step before constructing the decision map in the traditional density peak clustering algorithm, and then select two clusters for comparison according to the descending order of local density of the cluster center, and output the industrial load clustering results. S2. Establish a multi-timescale control potential mining method based on human-in-the-loop. The multi-timescale control potential mining method includes: day-ahead timescale control potential mining, intraday timescale control potential mining, and real-time timescale control potential mining.

[0007] Preferably, in S1, a clustering correction step is added before constructing the decision graph using the traditional density peak clustering algorithm, specifically: Define clustering and clustering The cluster cross density is , is represented as: ; in, For clustering The distance between data points is less than the cutoff distance. The dataset, For clustering The set of data points whose distance from the intermediate data points is less than the cutoff distance; Define clustering Boundary point set , is represented as: ; ; in, For clustering A single data point in the data; for Sort the local density of the data points in ascending order. To retrieve the set before m One element, To round up, for Number of data points in the middle; Define the boundary density between clusters as follows: ; in, For clustering set of boundary points The average local density of each data point in the dataset; For clustering set of boundary points The average local density of each data point in the dataset; for and Cluster boundary density between them.

[0008] Preferably, in S1, two clusters are compared based on their local density in descending order of the cluster centers. : S101. Forming a cluster cross-density matrix based on cluster cross-density and boundary density. Cluster boundary density matrix ,in It is a cluster set The number of clusters included, and the initial number of comparisons defined. ; S102. Select two clusters for comparison in descending order of local density of cluster centers; S103, if satisfied Delete the original two cluster centers, select the data point with the highest local density in the two clusters as the new centers, merge the two clusters, return to S101 to recalculate the matrix and loop; if all clusters do not satisfy the condition after pairwise comparison. If the loop stops, the industrial load clustering results will be output.

[0009] Preferably, in S2, the exploration of the potential for day-ahead timescale regulation includes the following steps: S20101. Input market data, environmental data, and production data; after preprocessing, construct a state-space vector. : ; in, This is the day-ahead electricity price forecast curve; Forecasting FM prices; For temperature; Humidity; For production planning; The upper limit of aggregated energy consumption for adjustable resource load; This is the lower limit of aggregated energy consumption for adjustable resource loads; Control costs per unit; S20102. Initial control strategies are generated using a proximal policy optimization algorithm as the optimizer and an Actor-Critic network as the structure. The Actor network outputs a random policy distribution and samples it to obtain actions. ; Critic network evaluates the long-term expected value of the state Combined with reward function ; ; ; ; in, Total reward; For market profits; The current market electricity price; For FM market prices; To control costs; Cost per unit of demand response; Reference power; This is the absolute value of the load control quantity; To incur penalties; This is the penalty coefficient; This is the deviation threshold; The amount of electricity bid for in the energy market; To bid for the capacity of the FM market; The deviation between the actual output and the total bid amount; S20103. During strategic adjustments, experts adjust the weighting ratio between economic efficiency and security. Enthusiasm for participation in the FM market Risk aversion level for specific periods The reward function is then derived and updated using the maximum entropy inverse reinforcement learning algorithm: ; in, For the total reward function; As an economic reward; As a security reward; Incentives for participation in the FM market; This is a form of inverse reinforcement learning reward; This represents the weighting ratio between economic efficiency and safety. The strategy is then re-optimized; during manual fine-tuning, experts drag and drop to modify the strategy curve, and the system records the deviation value, which is used for Actor network supervised learning and MaxEntIRL preference learning. The deviation value is expressed as: ; in, In time At any given moment, the power deviation between expert bidding strategies and AI bidding strategies; In time At any given moment, the bidding efficiency of expert decision-making; In time At any given moment, the bidding power generated by the AI ​​algorithm.

[0010] Preferably, in S2, the exploration of intraday timescale regulation potential includes the following steps: S20201. Scroll optimization is triggered by time-driven or event-driven methods, and the latest data is integrated to construct scroll time-domain state information, represented as follows: ; in, For at a certain point in time The set of rolling time-domain state information at the location; For at a certain point in time The actual power of the load at the location; For time point future time points The updated predicted power; S20202. With the goal of maximizing revenue, the pre-scheduling instructions are solved using a model predictive control algorithm to satisfy power balance constraints, as follows: ; in, For the first The set power at any given time; For the number The load in the first Actual power at any given moment; Load dynamic characteristic constraints are expressed as follows: ; in, For the number The load in the first Minimum allowable power at any given time; For the number The load in the first Maximum permissible power at any given time; The gradeability constraint is expressed as: ; in, For the number The load in the first Actual power at any given moment; For the number The maximum allowable gradeability of the load; S20203, the expert reference case reasoning engine retrieves historical similar scenarios and locally interpretable model instruction interpretations, and performs instruction fine-tuning or rule injection; S20204. The operation data is fed into the online sequence extreme learning machine to incrementally update the load response prediction model. The human-computer interaction context is stored in the case library. In new scenarios, the model prediction control optimization convergence is accelerated through transfer learning.

[0011] Preferably, in S2, the real-time timescale regulation potential mining includes the following steps: S20301. The global optimization problem is decomposed into local subproblems by using the alternating direction multiplier method. Each load local controller works together to solve the global optimal solution and generate executable instructions. S20302. Real-time response data is collected at 1-5 second intervals, and online anomaly detection is performed using the isolated forest algorithm; the input feature vector is represented as: ; in, In time The input feature vector used for anomaly detection at each time step; In time Power deviation at any given moment; In time The response rate at any given moment; In time The absolute value of the rate of change of load power deviation at any given time; S20303, when When the dynamic threshold is exceeded, the system will alarm and push an anomaly report and historical handling solutions. Experts will then execute false alarm cancellation, abnormal load isolation, or manual command input. S20304. Dynamically adjust the anomaly threshold based on expert feedback. When experts confirm that it is a false alarm, increase the anomaly threshold under the corresponding working condition of this type of load. Fine-tune the isolation forest with anomaly samples, train the safety policy network with negative rewards from the experience replay pool, and optimize using the constrained alternating direction multiplier method.

[0012] Preferably, the parameter values ​​for strategic adjustments in S20103 are all within the specified range. : The higher the value, the more emphasis is placed on economic efficiency. The higher the value, the more actively the participant in the FM market. The higher the value, the greater the degree of risk aversion.

[0013] Preferably, in S20202, with the goal of maximizing profit, the objective function for optimizing the model predictive control algorithm is: ; in, To control costs; To incur penalties and costs.

[0014] Therefore, the present invention employs the above-mentioned human-in-the-loop industrial virtual power plant multi-timescale control method, and compared with the prior art, the technical solution of the present invention has the following beneficial effects: (1) The present invention adopts an improved density peak clustering algorithm and adds a “cluster cross density + boundary density” correction process to overcome the problem that the traditional algorithm is not adapted to the characteristics of “non-convex cluster distribution and fuzzy boundary” of industrial load, which leads to “inaccurate characterization of characteristics and deviation in overall potential assessment” when aggregating massive heterogeneous industrial resources, thereby improving the clustering accuracy and stability. (2) This invention constructs a symbiotic model of "experts determine strategy and machines calculate execution". Experts transmit preferences through "strategic sliders", and the system uses maximum entropy inverse reinforcement learning to back-infer the reward function to optimize the strategy. At the same time, it supports "rule injection" to transform experience into algorithmic hard constraints, overcomes the problem of shallow human-machine collaboration in existing technologies, which easily leads to instructions that are out of touch with reality or have low efficiency and are difficult to form a global optimization strategy, and realizes deep human-machine collaborative evolution. (3) By constructing a “prediction-optimization-feedback” closed loop, this invention overcomes the problems of disconnection in multi-time scale regulation, isolated operation of three-scale scheduling, lack of organic connection and data linkage in traditional technology, and improves the adaptability of regulation to dynamic scenarios; (4) The present invention introduces the isolated forest algorithm to detect anomalies and push the cause in real time, and optimizes the anomaly threshold by combining expert feedback; the model parameters are updated in real time by online sequence extreme learning machine within the day, overcoming the problem that traditional solutions cannot adapt to industrial load fluctuations and equipment characteristic changes, reducing false alarm rate and command failure risk, and strengthening anomaly response and adaptive capability.

[0015] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0016] Figure 1 This is a flowchart of the improved DPC algorithm in an embodiment of the human-in-the-loop industrial virtual power plant multi-timescale control method of the present invention; Figure 2 This is a flowchart illustrating the day-ahead time-scale control potential mining of an embodiment of the human-in-the-loop industrial virtual power plant multi-time-scale control method of the present invention. Figure 3 This is a flowchart illustrating the intraday timescale control potential mining of an embodiment of the human-in-the-loop industrial virtual power plant multi-timescale control method of the present invention. Figure 4 This is a flowchart illustrating the real-time time-scale control potential mining process of an embodiment of the human-in-the-loop industrial virtual power plant multi-time-scale control method of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Unless otherwise defined, the technical or scientific terms used in the present invention should have the ordinary meaning understood by those skilled in the art.

[0018] Example 1 like Figure 1 As shown, the multi-timescale control method for industrial virtual power plants based on human-in-the-loop systems of the present invention includes the following steps: S1. Construct an industrial virtual power plant load clustering response model, specifically as follows: First, we need to clarify the core calculation rules of the traditional Density Peak Clustering (DPC) algorithm, including the calculation methods for local density and relative distance. Local density is calculated using the Euclidean distance between the x-th and y-th data points, the cutoff distance, and a decision function. Relative distance requires distinguishing whether a data point's local density is maximum. If it is maximum, the maximum distance from other data points to it is taken; otherwise, the shortest distance from a data point with even higher local density to it is taken. This forms the basis for constructing the decision graph. The improved DPC algorithm adds a clustering correction step before constructing the decision graph in the traditional DPC algorithm. First, it defines the cluster cross-density and calculates the association between two clusters by statistically analyzing the set of data points in two clusters whose distance is less than the cutoff distance. Define clustering and clustering The cluster cross density is , is represented as: ; in, For clustering The distance between data points is less than the cutoff distance. The dataset, For clustering The set of data points whose distance from the intermediate data points is less than the cutoff distance; Redefining Clustering Boundary point set After sorting the local density of data points within a cluster in ascending order, the top 20% (determined by rounding up to 20% of the total number of data points in the cluster) are taken as boundary points; clustering Boundary point set Represented as: ; ; in, For clustering A single data point in the data; for Sort the local density of the data points in ascending order. To retrieve the set before m One element, To round up, for The number of data points in the cluster, i.e. The boundary point is Data points within the 20% range of areas with lower local density; Simultaneously calculate the cluster boundary density, which is the average local density of each data point in the boundary point set; define the boundary density between clusters as follows: ; in, For clustering set of boundary points The average local density of each data point in the dataset; For clustering set of boundary points The average local density of each data point in the dataset; for and Cluster boundary density between; Based on the above definitions, the clustering cross-density matrix and clustering boundary density matrix are formed. The improved clustering correction process of the DPC algorithm includes the following steps: S101, Based on cluster sets The cluster cross-density and cluster boundary density among the clusters form a cluster cross-density matrix. and cluster boundary density matrix ,in yes The number of clusters included, and the initial number of comparisons defined. ; S102. Select cluster centers in descending order of local density. and Comparing elements in, for example: assuming The 4th and 7th cluster centers are For the data points with the highest local density, first... and Compare, if not satisfied > Then, the parameters of the first and third local densities are compared, and so on, until the condition is met. > Alternatively, all clusters can be compared pairwise; S103, if satisfied > ,from Delete the original two cluster centers, and reselect the data points with the highest local density in the two clusters as the new cluster centers. middle, Merge the two clusters; return to S101, calculate the new cluster cross density matrix and cluster boundary density matrix, and continue the loop; if the conditions are still not met... > That is, the existing The clustering in the loop can no longer be merged, so stop the loop and output the clustering results; S2. Establish a multi-timescale regulation potential mining method based on human-in-the-loop, which includes day-ahead timescale regulation potential mining, intraday timescale regulation potential mining, and real-time timescale regulation potential mining, specifically as follows: S201, such as Figure 2 As shown, exploring the potential for regulation on a current timescale includes the following steps: S20101, Data input includes market data (day-ahead electricity price forecast curve) FM price forecast ), environmental data (temperature) ,humidity ), production data (production plan) Adjustable resource load aggregation energy consumption limit Adjustable resource load aggregated energy consumption lower limit Unit control cost After preprocessing, a multidimensional feature vector is constructed as the state space vector. , is represented as: ; Use Long Short-Term Memory (LSTM) networks to extract spatiotemporal correlations; S20102, Joint Optimization and Initial Policy Generation: Using the PPO algorithm as the optimizer and an Actor-Critic network as the structure, the Actor network outputs a random policy distribution (sampled from market bidding electricity / capacity), and the Critic network evaluates the long-term expected value of the state. Combined with the reward function, the initial policy is generated, specifically as follows: In an Actor network, the input state space vector The output is a random policy distribution (such as a Gaussian distribution), from which actions are sampled. That is, the amount of electricity bid for the energy market and the capacity bid for the frequency regulation market in each time period; In a Critic network, the input state space vector Output a scalar value to evaluate the long-term expected value of the state. ; Set the reward function as follows: ; ; ; ; in, Total reward; For market profits; The current market electricity price; For FM market prices; To control costs; Cost per unit of demand response; Reference power; This is the absolute value of the load control quantity; To incur penalties; This is the penalty coefficient; This is the deviation threshold; The amount of electricity bid for in the energy market; To bid for the capacity of the FM market; The deviation between the actual output and the total bid amount; Using the Proximal Policy Optimization (PPO) algorithm as the optimization algorithm and the Actor-Critic network as the network structure, combined with the aforementioned reward function, three sets of initial control policies with different styles are generated: Aggressive approach: With high returns as the core objective, it allows for higher risks and is suitable for scenarios with favorable market conditions and high equipment redundancy. Robust solution: Based on behavioral cloning technology, it uses a pre-trained network to mimic the decision-making of historical experts, seeking a balance between returns and risks, and has strong universality; Conservative approach: Prioritizes ensuring the continuity of industrial production, with profit targets being relatively secondary. Suitable for situations where production tasks are urgent and equipment operating boundaries are strictly constrained.

[0019] S20103. The system displays three strategies—an aggressive, a moderate, and a conservative approach—on the human-computer interaction interface. Experts can participate in guidance and correction through the following three methods: When adjusting strategies, experts do not directly modify the strategy curve, but rather adjust the parameters of the strategy preference module. These parameters include the weighting ratio between economic efficiency and security. (Values ​​range from 0 to 1, with larger values ​​indicating a greater emphasis on economic efficiency), FM market participation enthusiasm. (Value range 0 to 1, larger values ​​indicate greater participation in the FM market), specific time periods t Risk aversion level (Values ​​range from 0 to 1; the larger the value, the higher the degree of risk aversion.) The system interprets the expert's slider adjustments as an indirect specification of the human reward function. Using the Maximum Entropy Inverse Reinforcement Learning (MaxEntIRL) algorithm, it deduces the expert's internal reward function from the slider adjustments and updates the reward function accordingly. ; in, For the total reward function; As an economic reward; As a security reward; Incentives for participation in the FM market; This is a form of inverse reinforcement learning reward; This represents the weighting ratio between economic efficiency and safety. This triggers the PPO algorithm to quickly re-optimize and generate a new strategy that aligns with expert preferences; During manual fine-tuning, experts can directly drag and drop to modify a strategy curve, generating a new curve. The system records the deviation value, which is represented as: ; in, In time At any given moment, the power deviation between expert bidding strategies and AI bidding strategies; In time At any given moment, the bidding efficiency of expert decision-making; In time At any given moment, the bidding power generated by the AI ​​algorithm; The data is used for two purposes: first, for supervised learning, to immediately fine-tune the Actor network so that its output is closer to the expert's choice; and second, as expert demonstration input to the MaxEntIRL algorithm to learn expert decision preferences more accurately. When direct approval is granted, experts can directly select one of the three initial strategies generated by the system to quickly determine the day-ahead control strategy, which is suitable for scenarios with high acceptance of the strategy and tight time constraints. S202, such as Figure 3 As shown, exploring the potential for intraday timescale regulation includes the following steps: S20201, the rolling triggering mechanism is divided into time-driven and event-driven; among them, the time-driven mechanism is automatically triggered once every 15 minutes (one scheduling cycle); the event-driven mechanism is triggered immediately when a new ultra-short-term wind power / photovoltaic power forecast is received (updated every 15 minutes), the real-time electricity price fluctuates drastically, or a new instruction is received from the grid; Data assimilation is the integration of the latest data into the system to construct the current point in time. The rolling time-domain state information is expressed as follows: ; in, For at a certain point in time The set of rolling time-domain state information at the location; For at a certain point in time The actual power of the load at the location; For time point future time points The updated predicted power; S20202, MPC optimization solution: Within each rolling window, the objective function is to maximize the profit. The objective function is expressed as: ; in, To control costs; To incur penalties; At the same time, the power balance constraint is satisfied, which can be expressed as: ; in, For the first The set power at any given time; For the number The load in the first Actual power at any given moment; Load dynamic characteristic constraints are expressed as follows: ; in, For the number The load in the first Minimum allowable power at any given time; For the number The load in the first Maximum permissible power at any given time; The gradeability constraint is expressed as: ; in, For the number The load in the first Actual power at any given moment; For the number The maximum allowable gradeability of the load; The pre-scheduled instructions are obtained by solving the model predictive control (MPC) algorithm; S20203. The system displays pre-scheduling instructions in the form of Gantt charts and power curves. Simultaneously, it uses a Case-Based Reasoning (CBR) engine to retrieve three most similar historical scenarios and human decision-making results from the historical database for expert reference. It also utilizes locally interpretable models (such as LIME) to explain key instructions (e.g., "Reduce cooling load at 13:00 due to predicted temperature decrease"). Experts can fine-tune instructions (by dragging and dropping charts and annotating the reasons for correction, such as equipment failure) or inject rules (by entering business rules through the "Rule Editor," such as setting the power limit to 0 when equipment fails). These rules are then transformed into hard constraints for MPC optimization. S20204. In the online model update and case management phase, after the scheduler completes and confirms the instruction correction, the operational data (such as power changes, time, specific causes, etc.) will form new samples and be fed into the online Sequence Extreme Learning Machine (OS-ELM) algorithm for online incremental updates to the response prediction model of the affected load. Since the OS-ELM update is an analytical solution, the calculation speed is extremely fast (milliseconds), which hardly increases the time overhead of MPC rolling optimization, thus achieving "learning as soon as it is modified", and the corrected model can take effect immediately in the next rolling optimization cycle.

[0020] Meanwhile, the complete context of each human-computer interaction (state, instruction, correction, result) is stored as a new case in the case library. When the system encounters a new scenario, the CBR engine will search for similar cases; if the similarity is below a threshold, the system uses transfer learning to transfer the strategy from the highly similar case to the new scenario as the initial solution for MPC optimization, thereby accelerating optimization convergence and improving the quality of the initial solution.

[0021] The system forms a rapid autonomous closed loop of "execution-feedback-learning-optimization". Dispatchers only need to handle anomalies and inject knowledge, without having to do everything themselves. The intraday dispatch system has also evolved from a "static fool" that requires constant manual correction to an "intelligent partner" that can work together with dispatchers, significantly improving the ability of industrial virtual power plants to cope with real-time uncertainties. S203, such as Figure 4 As shown, the exploration of real-time timescale regulation potential includes the following steps: S20301 employs a distributed optimization strategy based on the Alternating Directional Multiplier Method (ADMM). This strategy is a computational framework applicable to distributed convex optimization problems. Through a decomposition and coordination process, the global optimization problem is broken down into multiple local subproblems. Each load's local controller (agent) communicates in real-time with neighboring agents based on its own constraints (such as power upper and lower limits, ramp rate), independently solving the local subproblems and then collaboratively obtaining the global optimal solution, quickly calculating the optimal allocation scheme for executable instructions. This approach features high distributed computing speed and high reliability, effectively avoiding the risk of single-point failure of the central controller and ensuring the stability of instruction allocation and execution. S20302. Real-time monitoring and anomaly detection includes two main components: data stream processing and online anomaly detection. In the data stream processing stage, the system collects real-time response data of each controlled load at high frequency (1-5 second interval), covering key process variables such as real-time power, set power, response deviation, and response rate, providing comprehensive data support for anomaly detection; In the online anomaly detection phase, the Isolation Forest algorithm is used as the core detection method. This algorithm is an unsupervised learning algorithm specifically designed for outlier detection. It isolates data points by constructing random binary trees, leveraging the characteristic that outliers are easier to isolate (shorter isolation paths) to identify anomalies. It has the advantages of high efficiency and simplicity, and is suitable for large-scale real-time data stream scenarios. The input feature vector during detection is: ; in, In time The input feature vector used for anomaly detection at each time step; In time Power deviation at any given moment; In time The response rate at any given moment; In time The absolute value of the rate of change of load power deviation at any given time; The output is a real-time anomaly score calculated for each data point. The closer the score is to 1, the more abnormal the load response behavior is at that moment. S20303, when When the dynamic threshold is exceeded, the system immediately highlights the alarm on the monitoring screen and pushes the alarm information to the operator's mobile device. At the same time, the interpreter is automatically triggered to analyze the key features that cause the anomaly score to soar, and presents them to the operator in a visual chart with feature importance ranking. It also generates a structured anomaly report and pushes historical similar anomalies and handling solutions from the case library to provide reference for the operator's decision-making. Combining the interpreted report with their own experience, the operator makes a decision and performs the following actions within seconds: If the alarm is confirmed to be false, click the "Confirm as Normal" button to cancel the alarm and continue executing the original command. When a genuine anomaly is confirmed, clicking the "One-Click Pause" button will immediately stop the system from sending instructions to the abnormal load and isolate it, while simultaneously activating backup adjustment resources. In complex scenarios, operators can manually input temporary power commands or adjust allocation weights to flexibly respond to special situations; S20304. Through dynamic threshold adjustment and incremental model updates, the system achieves self-optimization and adaptation, specifically as follows: Dynamic threshold adjustment is based on operator feedback of false alarms. The system automatically increases the abnormal threshold for this type of load under the corresponding operating conditions to reduce subsequent unnecessary false alarm interference. The incremental model update includes two mechanisms: supervised learning fine-tuning and reinforcement learning feedback. Supervised learning fine-tuning uses operator-confirmed anomaly samples as supervised learning samples, periodically fine-tuning the isolation forest or auxiliary supervised anomaly classifiers (such as gradient boosting trees, GBDT) to improve anomaly detection accuracy. Reinforcement learning feedback treats operator "rejection" operations as strong negative rewards and stores these signals in an experience replay pool, periodically using them to train a safety policy network. This network provides additional safety constraints to the distributed optimizer, preventing the generation of instructions that might trigger similar anomalies in the future, thus continuously improving the system's operational safety.

[0022] Therefore, the present invention adopts the above-mentioned human-in-the-loop industrial virtual power plant multi-timescale control method, which improves the accuracy and stability of industrial load clustering, constructs a deep human-machine collaborative evolution mode, enhances the adaptability to dynamic scenarios, reduces the false alarm rate of anomaly detection and the risk of control command failure, and further strengthens the system's anomaly response and adaptive capabilities.

[0023] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0024] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-timescale control method for industrial virtual power plants based on human-in-the-loop systems, characterized in that, Includes the following steps: S1. Establish an industrial virtual power plant load clustering response model, clarify the core calculation rules of the traditional density peak clustering algorithm, including calculating local density through Euclidean distance, cutoff distance and judgment function of data points, distinguishing whether the local density of data points is the maximum to calculate the relative distance; add a clustering correction step before constructing the decision map in the traditional density peak clustering algorithm, and then select two clusters for comparison according to the descending order of local density of the cluster center, and output the industrial load clustering results. S2. Establish a multi-timescale control potential mining method based on human-in-the-loop. The multi-timescale control potential mining method includes: day-ahead timescale control potential mining, intraday timescale control potential mining, and real-time timescale control potential mining.

2. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 1, characterized in that, In S1, a clustering correction step is added before constructing the decision graph in the traditional density peak clustering algorithm. Specifically: Define clustering and clustering The cluster cross density is , is represented as: ; in, For clustering The distance between data points is less than the cutoff distance. The dataset, For clustering The set of data points whose distance from the intermediate data points is less than the cutoff distance; Define clustering Boundary point set , is represented as: ; ; in, For clustering A single data point in the data; for Sort the local density of the data points in ascending order. To retrieve the set before m One element, To round up, for Number of data points in the middle; Define the boundary density between clusters as follows: ; in, For clustering set of boundary points The average local density of each data point in the dataset; For clustering set of boundary points The average local density of each data point in the dataset; for and Cluster boundary density between them.

3. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 1, characterized in that, In S1, two clusters are compared based on their local density in descending order of the cluster centers. : S101. Forming a cluster cross-density matrix based on cluster cross-density and boundary density. Cluster boundary density matrix ,in It is a cluster set The number of clusters included, and the initial number of comparisons defined. ; S102. Select two clusters for comparison in descending order of local density of cluster centers; S103, if satisfied Delete the original two cluster centers, select the data point with the highest local density in the two clusters as the new centers, merge the two clusters, return to S101 to recalculate the matrix and loop; if all clusters do not satisfy the condition after pairwise comparison. If the loop stops, the industrial load clustering results will be output.

4. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 1, characterized in that, In S2, the exploration of the potential for day-ahead timescale regulation includes the following steps: S20101. Input market data, environmental data, and production data; after preprocessing, construct a state-space vector. : ; in, This is the day-ahead electricity price forecast curve; Forecasting FM prices; For temperature; Humidity; For production planning; The upper limit of aggregated energy consumption for adjustable resource load; This is the lower limit of aggregated energy consumption for adjustable resource loads; Control costs per unit; S20102. Initial control strategies are generated using a proximal policy optimization algorithm as the optimizer and an Actor-Critic network as the structure. The Actor network outputs a random policy distribution and samples it to obtain actions. ; Critic network evaluates the long-term expected value of the state Combined with reward function ; ; ; ; in, Total reward; For market profits; The current market electricity price; For FM market prices; To control costs; Cost per unit of demand response; Reference power; This is the absolute value of the load control quantity; To incur penalties; This is the penalty coefficient; This is the deviation threshold; The amount of electricity bid for in the energy market; To bid for the capacity of the FM market; The deviation between the actual output and the total bid amount; S20103. During strategic adjustments, experts adjust the weighting ratio between economic efficiency and security. Enthusiasm for participation in the FM market Risk aversion level for specific periods The reward function is then derived and updated using the maximum entropy inverse reinforcement learning algorithm: ; in, For the total reward function; As an economic reward; As a security reward; Incentives for participation in the FM market; This is a form of inverse reinforcement learning reward; This represents the weighting ratio between economic efficiency and safety. The strategy is then re-optimized; during manual fine-tuning, experts drag and drop to modify the strategy curve, and the system records the deviation value, which is used for Actor network supervised learning and MaxEntIRL preference learning. The deviation value is expressed as: ; in, In time At any given moment, the power deviation between expert bidding strategies and AI bidding strategies; In time At any given moment, the bidding efficiency of expert decision-making; In time At any given moment, the bidding power generated by the AI ​​algorithm.

5. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 1, characterized in that, In S2, the exploration of intraday timescale manipulation potential includes the following steps: S20201. Scroll optimization is triggered by time-driven or event-driven methods, and the latest data is integrated to construct scroll time-domain state information, represented as follows: ; in, For at a certain point in time The set of rolling time-domain state information at the location; For at a certain point in time The actual power of the load at the location; For time point future time points The updated predicted power; S20202. With the goal of maximizing revenue, the pre-scheduling instructions are solved using a model predictive control algorithm to satisfy power balance constraints, as follows: ; in, For the first The set power at any given time; For the number The load in the first Actual power at any given moment; Load dynamic characteristic constraints are expressed as follows: ; in, For the number The load in the first Minimum allowable power at any given time; For the number The load in the first Maximum permissible power at any given time; The gradeability constraint is expressed as: ; in, For the number The load in the first Actual power at any given moment; For the number The maximum allowable gradeability of the load; S20203, the expert reference case reasoning engine retrieves historical similar scenarios and locally interpretable model instruction interpretations, and performs instruction fine-tuning or rule injection; S20204. Input the operation data into the online sequence extreme learning machine incremental update load response prediction model, and store the human-computer interaction context in the case library.

6. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 1, characterized in that, In S2, the exploration of real-time time-scale regulation potential includes the following steps: S20301. The global optimization problem is decomposed into local subproblems by using the alternating direction multiplier method. Each load local controller works together to solve the global optimal solution and generate executable instructions. S20302. Real-time response data is collected at 1-5 second intervals, and online anomaly detection is performed using the isolated forest algorithm; the input feature vector is represented as: ; in, In time The input feature vector used for anomaly detection at each time step; In time Power deviation at any given moment; In time The response rate at any given moment; In time The absolute value of the rate of change of load power deviation at any given time; S20303, when When the dynamic threshold is exceeded, the system will alarm and push an anomaly report and historical handling solutions. Experts will then execute false alarm cancellation, abnormal load isolation, or manual command input. S20304. Dynamically adjust the anomaly threshold based on expert feedback. When experts confirm that it is a false alarm, increase the anomaly threshold under the corresponding working condition of this type of load. Fine-tune the isolation forest with anomaly samples, train the safety policy network with negative rewards from the experience replay pool, and optimize using the constrained alternating direction multiplier method.

7. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 4, characterized in that, The parameter range for strategic adjustments in S20103 is [missing information]. : The higher the value, the more emphasis is placed on economic efficiency. The higher the value, the more actively the participant in the FM market. The higher the value, the greater the degree of risk aversion.

8. The multi-timescale control method for industrial virtual power plants based on human-in-the-loop as described in claim 5, characterized in that, In S20202, with the goal of maximizing profit, the objective function for optimizing the model predictive control algorithm is: ; in, To control costs; To incur penalties and costs.

9. A computer device, characterized in that, include: A processor configured to be coupled to memory, read and execute instructions and / or program code in the memory to perform the method as described in any one of claims 1-8.

10. A computer-readable medium, characterized in that, The computer-readable medium stores computer program code that, when executed on a computer, causes the computer to perform the method as described in any one of claims 1-8.