An air conditioner scheduling method and device based on the response of a virtual power plant
By obtaining the real-time operating parameters of the air conditioner terminal and the dynamic regulation instructions of the virtual power plant, a multi-objective optimization model is used to generate an air conditioner regulation priority sequence, and incrementally updates are solved through the feedback parameter set, the conflict between power grid requirements, user comfort and equipment energy consumption in air conditioning scheduling of virtual power plant is solved, and a scheduling strategy with high adaptability and accuracy is achieved.
Patent Information
- Application Number
- CN202510412843.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The existing virtual power plant air conditioning scheduling methods cannot respond to grid load regulation needs in real time, resulting in conflicts between user comfort and equipment energy consumption efficiency, and lack of a closed-loop mechanism for policy execution to feedback learning, resulting in a deterioration of model performance.
By obtaining the real-time operating parameter set of the air conditioning terminal and the dynamic regulation instructions of the virtual power plant, a multi-objective optimization model is used to generate the priority sequence of air conditioning regulation, and incrementally update it through the feedback parameter set to dynamically balance the grid requirements, user comfort and equipment energy consumption efficiency.
It realizes high adaptability and accuracy of the air conditioning scheduling strategy in complex scenarios, responds to emergency adjustment of grid load at minute levels, reduces the risk of scheduling failure, and improves the stability and accuracy of the scheduling strategy by continuously learning to adapt to changes in grid operation mode and user preferences.
Smart Images

Figure CN119914989B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and particularly to an air conditioner scheduling method and device based on virtual power plant response. Background Art
[0002] In the prior art, when a virtual power plant adjusts the power grid load by aggregating air conditioner terminals, static scheduling methods based on fixed rules or single-objective optimization are mostly used, and such methods have significant defects. There may be dynamic conflicts among multiple objectives such as power grid peak shaving requirements, user comfort, and equipment energy consumption efficiency. Traditional strategies using fixed priorities (such as grid priority first) lead to rigid strategies and cannot adapt to the coordinated requirements of real-time electricity price fluctuations, environmental parameter changes, and user preference migrations, easily causing a sharp drop in user satisfaction or an increase in energy consumption caused by frequent start and stop of equipment. Secondly, the static model relies on offline training of historical data, and the input features are limited to static parameters such as rated power and fixed electricity price, lacking the modeling of real-time response capabilities. When a sudden control instruction is issued, the deviation of the strategy execution results in the failure of load adjustment. In addition, the existing methods lack a closed-loop mechanism from strategy execution to feedback learning, and the problem of model performance decline cannot be solved during long-term operation. In summary, the prior art is difficult to meet the requirements of the virtual power plant for the accuracy, real-time performance, and sustainability of air conditioner terminal scheduling. Summary of the Invention
[0003] The present invention provides an air conditioner scheduling method and device based on virtual power plant response.
[0004] According to one aspect of the present invention, an air conditioner scheduling method based on virtual power plant response is provided, and the method includes:
[0005] Obtain a real-time operation parameter set of each air conditioner terminal in a target area within a preset time window, where the real-time operation parameter set includes current energy consumption characteristics, environmental temperature and humidity characteristics, and user preference setting characteristics, and the user preference setting characteristics are used to describe the priority configuration of the user for the air conditioner operation mode;
[0006] Perform feature alignment processing on the real-time operation parameter set and a dynamic control instruction issued by the virtual power plant to generate a combined input feature, where the dynamic control instruction includes a load distribution strategy of the virtual power plant for the target area and real-time electricity price fluctuation characteristics;
[0007] Input the combined input feature into a pre-trained multi-objective optimization model, and perform hierarchical weight allocation on the combined input feature through a feature fusion layer in the multi-objective optimization model to generate an air conditioner control priority sequence, where the priority sequence is used to describe the scheduling order and parameter adjustment range of each air conditioner terminal under the dynamic control instruction;
[0008] Generate a target scheduling strategy based on the priority sequence, where the target scheduling strategy includes the start / stop time configuration of each air-conditioning terminal, the adjustment range of the temperature setting threshold, and the energy consumption constraint conditions;
[0009] Send the target scheduling strategy to each air-conditioning terminal in the target area for execution, and collect the feedback parameter set after execution in real time. Use the feedback parameter set to perform incremental parameter update on the multi-objective optimization model.
[0010] According to another aspect of the present invention, there is provided an air-conditioning scheduling device based on the response of a virtual power plant, including:
[0011] A parameter acquisition module, configured to acquire the real-time operation parameter set of each air-conditioning terminal in the target area within a preset time window. The real-time operation parameter set includes the current energy consumption characteristics, environmental temperature and humidity characteristics, and user preference setting characteristics. The user preference setting characteristics are used to describe the priority configuration of the user for the air-conditioning operation mode;
[0012] A feature alignment module, configured to perform feature alignment processing on the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant to generate a combined input feature. The dynamic regulation instruction includes the load distribution strategy of the virtual power plant for the target area and the real-time electricity price fluctuation characteristics;
[0013] A priority generation module, configured to input the combined input feature into a pre-trained multi-objective optimization model, perform hierarchical weight assignment on the combined input feature through the feature fusion layer in the multi-objective optimization model, and generate an air-conditioning regulation priority sequence. The priority sequence is used to describe the scheduling order and parameter adjustment range of each air-conditioning terminal under the dynamic regulation instruction;
[0014] A strategy generation module, configured to generate a target scheduling strategy based on the priority sequence. The target scheduling strategy includes the start / stop time configuration of each air-conditioning terminal, the adjustment range of the temperature setting threshold, and the energy consumption constraint conditions;
[0015] An incremental update module, configured to send the target scheduling strategy to each air-conditioning terminal in the target area for execution, and collect the feedback parameter set after execution in real time. Use the feedback parameter set to perform incremental parameter update on the multi-objective optimization model.
[0016] The present invention has at least the following beneficial effects:
[0017] The air conditioner scheduling method based on the response of the virtual power plant provided by the present invention generates combined input features by obtaining the real-time operation parameter sets of each air conditioner terminal in the target area and combining the dynamic regulation instructions issued by the virtual power plant; performs hierarchical weight allocation on the combined input features based on a multi-objective optimization model to generate an air conditioner regulation priority sequence, and further generates a target scheduling strategy including start-stop times, temperature thresholds, and energy consumption constraints; and performs incremental update on the model through the execution feedback parameter set to form a dynamic optimization closed loop. Based on this, the real-time operation parameter set integrates the physical operation status, environmental parameters, and user preferences of the air conditioner terminal. After feature alignment with the dynamic regulation instructions of the virtual power plant, the generated combined input features cover multi-dimensional information of grid demand, equipment capabilities, and user behavior. Through hierarchical weight allocation of the above features by the multi-objective optimization model, the conflicts between grid peak shaving demand, user comfort, and equipment energy consumption efficiency can be dynamically balanced, avoiding the problem of policy bias caused by single-objective optimization in traditional methods, and significantly improving the adaptability of the scheduling strategy in complex scenarios. In addition, by performing time-segment mapping on the load adjustment target value and the energy consumption fluctuation curve in the real-time operation parameters, and integrating the gradient change characteristics of electricity price fluctuations, the generated combined input features can accurately depict the spatio-temporal correlation between grid demand and terminal operation status. Compared with traditional static strategies, this mechanism can capture the impact of electricity price fluctuations on user preferences in real time and dynamically adjust the equipment scheduling order and parameter adjustment amplitude in the priority sequence, so as to achieve minute-level response during emergency grid load adjustment and reduce the risk of scheduling failure caused by instruction delay. The multi-objective optimization model pre-learns the trade-off rules of grid stability, user satisfaction, and energy consumption efficiency from historical data. When generating the priority sequence, it not only allocates the equipment scheduling order based on the current combined input features, but also embeds the prediction information of the equipment response ability through the parameter adjustment amplitude. In this way, while meeting the hard constraints of the virtual power plant load distribution, the target scheduling strategy reserves a temperature adjustment buffer range allowed by user preferences, avoiding policy execution interruption caused by equipment response delay or sudden changes in user behavior, and achieving the unity of policy rigid requirements and elastic fault tolerance. Finally, by collecting the feedback parameter set after the strategy execution in real time and comparing the differences with the model prediction results, a multi-objective training loss value is generated to drive the incremental update of the model parameters, enabling the model to continuously learn dynamic change rules such as grid operation mode switching, equipment aging attenuation, and user preference migration, and avoiding the problem of performance decline of traditional models due to time-varying environments. At the same time, the incremental update process only performs local parameter optimization for abnormal data points, improving the adaptability of the model while retaining the stable strategy generation ability in historical training.In summary, through the synergistic effects of multi-source data fusion, dynamic feature alignment, multi-objective flexible strategy generation, and closed-loop incremental update, the present invention solves the technical problem of the difficulty in real-time collaborative optimization of grid demand, user comfort, and equipment energy consumption in air conditioner scheduling in the virtual power plant environment, and significantly improves the accuracy, response speed, and long-term stability of the scheduling strategy.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings exemplarily show embodiments and constitute a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0020] Figure 1 FIG. shows a schematic diagram of an application scenario of an air conditioner scheduling method based on virtual power plant response according to an embodiment of the present invention.
[0021] Figure 2 FIG. shows a flowchart of an air conditioner scheduling method based on virtual power plant response according to an embodiment of the present invention.
[0022] Figure 3 FIG. shows a schematic diagram of the functional module architecture of an air conditioner scheduling device based on virtual power plant response according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] In the present invention, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and do not intend to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0025] Figure 1FIG. 0 shows an application scenario of an embodiment of the present invention, and specifically provides an air conditioner scheduling system 100. The air conditioner scheduling system 100 includes one or more air conditioner terminals 101, a server 120, and one or more communication networks 110 that couple the one or more air conditioner terminals 101 to the server 120.
[0026] In an embodiment of the present invention, the server 120 may run one or more services or software applications that enable an air conditioner scheduling method based on virtual power plant response to be executed.
[0027] In Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or combinations thereof that may be executed by one or more processors.
[0028] The network 110 may be any type of network known to those skilled in the art, and it may use any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication.
[0029] The server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframes, server clusters, or any other suitable arrangement and / or combination. The server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that may be virtualized to maintain virtual storage devices of the server). In various embodiments, the server 120 may run one or more services or software applications that provide the functions described below.
[0030] The air conditioner scheduling system 100 may further include one or more databases 130. In certain embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store historical data. The databases 130 may reside in various locations. For example, the databases used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection.
[0031] Please refer to Figure 2 , the air conditioner scheduling method based on virtual power plant response provided by the embodiment of the present invention may include the following steps:
[0032] Step S100: Obtain the real-time operation parameter sets of each air-conditioning terminal in the target area within a preset time window. The real-time operation parameter sets include current energy consumption characteristics, environmental temperature and humidity characteristics, and user preference setting characteristics. The user preference setting characteristics are used to describe the priority configuration of the user for the air-conditioning operation mode.
[0033] In the target area, the preset time window is, for example, a continuous time interval set by a virtual power plant, and its length can be dynamically adjusted according to the power grid load forecasting requirements. For example, it is set to a variable period from 5 minutes to 30 minutes. The current energy consumption characteristics are collected in real time through the current sensor and power measurement module of the air-conditioning terminal, and specifically include parameters such as the compressor operating frequency, air supply wind speed, and the difference sequence between the set temperature and the actual temperature. For example, an air-conditioning terminal records that the compressor operating frequency is 45 Hz, the air supply wind speed is 3 m / s, and the difference between the set temperature of 24 °C and the actual temperature of 26 °C has not converged for 2 minutes within the time window. The environmental temperature and humidity characteristics are obtained through the temperature and humidity sensors deployed in the space where the air-conditioning terminal is located, including indoor temperature, humidity, and outdoor environmental temperature gradient change data. For example, the indoor temperature in a certain area linearly drops from 28 °C to 26 °C within the time window, and the humidity maintains a fluctuation of 60% ± 2%. The user preference setting characteristics are collected through a mobile terminal application or an air-conditioning control panel, and are specifically manifested as the temperature sensitivity coefficient set by the user, the energy-saving priority flag bit, and the allowable regulation time period. For example, user A sets the temperature sensitivity coefficient to 0.8 (in the range of 0-1, the higher the value, the lower the tolerance for temperature fluctuations), activates the energy-saving priority flag bit, and designates 14:00-16:00 every day as the allowable regulation time period. The above three types of characteristics are transmitted to the regional data center through the Internet of Things communication protocol, and after timestamp alignment processing, a real-time operation parameter set with spatio-temporal consistency is formed. The data structure of this parameter set is a multi-dimensional time series matrix, where each air-conditioning terminal corresponds to a row in the matrix, and each characteristic parameter is arranged as a column vector according to the sampling frequency within the time window. This step S100 ensures through a data synchronization mechanism that the operation states and environmental parameters of all air-conditioning terminals can be correlated and analyzed under a unified time benchmark, providing complete input conditions for subsequent dynamic regulation.
[0034] As an implementation manner, for the above step S100, obtaining the real-time operation parameter sets of each air-conditioning terminal in the target area within a preset time window may specifically include:
[0035] Step S110: Collect the operation state data of each air-conditioning terminal in the target area in multiple consecutive time segments through an Internet of Things gateway. The operation state data includes the compressor operating frequency, air supply wind speed, and the difference sequence between the set temperature and the actual temperature.
[0036] As a data mediation node between the air conditioner terminals and the central control system within the target area, the Internet of Things gateway establishes a two-way communication link with each terminal controller using the MQTT protocol or the Modbus protocol, and collects operation status data at fixed time intervals. The length of the time interval is dynamically configured according to the scheduling requirements of the virtual power plant, and the typical setting is an adjustable range from 30 seconds to 5 minutes. For example, in the scenario of urgent grid frequency modulation requirements, the time interval is shortened to 30 seconds to ensure high-frequency acquisition of the instantaneous change in the compressor operating frequency, while in the normal scheduling mode, it is extended to 5 minutes to reduce the communication load. The compressor operating frequency is obtained in real time through the frequency converter control module of the air conditioner terminal, which represents the rotational speed fluctuation of the compressor motor per unit time. For example, a certain terminal records that the operating frequency linearly rises from 45 Hz to 48 Hz and remains stable within a time interval. The air supply wind speed is measured by the Hall sensor of the built-in fan of the air conditioner, which specifically represents the air flow velocity generated by the rotation of the fan impeller. For example, a certain terminal maintains a steady-state value of 3.2 m / s ± 0.1 m / s for the air supply wind speed in multiple consecutive time intervals. The difference sequence between the set temperature and the actual temperature is obtained through periodic sampling by the temperature sensor, and the difference sequence is sorted by timestamp to form discrete time series data. For example, if a certain terminal sets the temperature to 24 °C, and the actual temperatures in three consecutive time intervals are 25.3 °C, 24.8 °C, and 24.5 °C respectively, then the difference sequence is [1.3, 0.8, 0.5]. In this step, through the multiplexing communication mechanism of the Internet of Things gateway, the operation status data of all air conditioner terminals within the target area is synchronously collected and encapsulated into a data packet in JSON format and transmitted to the edge computing node to ensure that the data collection process meets the requirements of low latency and high reliability.
[0037] Step S120: Perform outlier removal and timestamp alignment processing on the operation status data to generate a standardized set of operation parameters; among them, the timestamp alignment processing includes mapping sensor data with different sampling frequencies to the same time reference axis, and the outlier removal includes identifying compressor current values or temperature and humidity readings that exceed the preset physical thresholds.
[0038] Exemplarily, the outlier rejection process is implemented based on a preset physical threshold library, which stores the rated operating parameter ranges of various air-conditioning terminals. For example, for a certain model of air conditioner, the normal range of the compressor current is 3A - 6A. When the instantaneous current value of 8A is collected, the system determines it as abnormal data and triggers the rejection operation, and fills the missing value by linear interpolation. The abnormal detection of temperature and humidity readings adopts a dynamic threshold adjustment strategy. For example, when the reported value of the indoor temperature sensor suddenly rises by 5°C compared with the previous period and exceeds the outdoor temperature gradient change range, it is determined as a sensor failure and the data point is rejected. The timestamp alignment process is implemented through a time reference axis mapping algorithm, which uses the global clock of the Internet of Things gateway as the reference to resample data with different sampling frequencies. For example, the compressor operating frequency data is sampled at 10-second intervals, while the air supply wind speed data is sampled at 30-second intervals. During alignment processing, the two types of data are uniformly mapped to a time reference axis with a 1-minute interval, and the data at the missing time points is completed by cubic spline interpolation. The generation of the standardized operating parameter set includes two stages: data normalization and unit unification. In the normalization process, each parameter is scaled to the interval [0, 1]. For example, the compressor operating frequency of 45Hz - 55Hz is mapped to 0.45 - 0.55. In the unit unification process, parameters with different dimensions are converted to standard units. For example, the temperature difference is converted from Fahrenheit to Celsius. The processed standardized operating parameter set is a structured data matrix, whose row dimension represents the air-conditioning terminal identifier, and the column dimension contains the normalized parameter values under the timestamp sequence, ensuring the spatio-temporal consistency for subsequent analysis.
[0039] Step S130: Extract the preference setting instructions submitted by the user through the mobile terminal. The preference setting instructions include the temperature sensitivity coefficient, the energy-saving priority flag bit, and the allowable regulation time period.
[0040] Exemplarily, the user preference setting instruction is input through the interaction interface of the mobile terminal application and transmitted to the user preference database of the virtual power plant in an encrypted form. The temperature sensitivity coefficient is defined as a quantitative index of the user's tolerance to temperature fluctuations, with a value range of 0 - 1, where 0 indicates that the user can accept large temperature fluctuations, and 1 indicates that the temperature is required to be strictly maintained at the set value. For example, if user A sets the temperature sensitivity coefficient to 0.9, the system will prioritize ensuring the temperature stability of its air conditioner terminal during scheduling. The energy-saving priority flag is a Boolean parameter. When its value is true, it indicates that the user allows the virtual power plant to automatically adjust the air conditioner operation mode during peak electricity price periods to reduce energy consumption. For example, when the flag is activated, the air conditioner terminal can accept a 10% reduction in the compressor frequency. The allowed regulation time period is the time interval specified by the user for acceptable external scheduling intervention. For example, user B sets 14:00 - 16:00 every day as the allowed regulation time period, and the system shall not modify the operation parameters of its air conditioner terminal outside this time period. The extraction process of the preference setting instruction includes three links: data decryption, format conversion, and validity verification. In the decryption link, the asymmetric encryption algorithm is used to restore the user's original input. Format conversion converts the text instruction into a numerical feature vector. For example, the allowed regulation time period "14:00 - 16:00" is converted into a start timestamp 1620000000 and an end timestamp 1620007200. Validity verification detects the legality of the input parameters through a rule engine. For example, when the temperature sensitivity coefficient exceeds 1, the system automatically truncates it to 1 and generates an alarm log.
[0041] Step S140: Concatenate the features of the standardized operation parameter set and the preference setting instruction to generate a real-time operation parameter set containing spatio-temporal correlation.
[0042] Exemplarily, the feature splicing process can be implemented through a matrix expansion algorithm to align the two-dimensional matrix of the standardized operating parameter set with the three-dimensional feature vector of the user preference setting instruction in terms of space-time dimensions. Specifically, the parameter row vector corresponding to each timestamp of the standardized operating parameter set (e.g., [0.45, 3.2, 0.5] represents the compressor frequency, air supply wind speed, and temperature difference) is horizontally spliced with the preference feature vector of the corresponding user (e.g., [0.9, 1, 1620000000] represents the temperature sensitivity coefficient, energy-saving priority flag, and allowable regulation start time) to form an extended composite feature vector. The establishment of space-time correlation depends on the joint coding of timestamps and spatial positions. For example, a unique geographical location code (such as GPS coordinates) is assigned to each air-conditioning terminal in the target area, and it is used together with the timestamp as a composite index key. The real-time operating parameter set is finally stored in a space-time tensor structure, whose dimensions include four parts: the timestamp sequence, the spatial coordinates of the air-conditioning terminal, the operating state parameters, and the user preference parameters. For example, the parameter entry of a certain terminal at timestamp 1620000000 can be expressed as (1620000000, (116.404, 39.915), 0.48, 3.1, 0.6, 0.9, 1, 1620000000), where the first three values are the standardized operating parameters, and the last three values are the user preference parameters. This step enables the subsequent scheduling algorithm to simultaneously consider the real-time state of the equipment, environmental factors, and user personalized needs through the deep fusion of the feature space, laying a data foundation for generating high-precision regulation strategies.
[0043] Step S200: Perform feature alignment processing on the real-time operating parameter set and the dynamic regulation instruction issued by the virtual power plant to generate a joint input feature. The dynamic regulation instruction includes the load distribution strategy of the virtual power plant for the target area and the real-time electricity price fluctuation characteristics.
[0044] Exemplarily, the dynamic regulation instruction is a set of regulation requirements generated by the central controller of the virtual power plant. Its load distribution strategy specifically includes the total load value to be adjusted in the target area, the load change rate limit, and the load distribution ratio in different time periods. For example, the virtual power plant requires that in the next 15-minute time window, the total load in the target area be reduced by 50 kW, and the load reduction rate does not exceed 10 kW / minute, and 30% of the load reduction needs to be completed in the first 5 minutes. The real-time electricity price fluctuation characteristics are obtained through the power market trading platform, which are represented by the time-of-use electricity price curve and its gradient change rate. For example, the current electricity price in the current period is 0.5 yuan / kWh, and it is predicted that the electricity price will rise to 0.6 yuan / kWh in the next period. The feature alignment process first parses the time constraint parameters in the dynamic regulation instruction and maps and matches them with the time window of the real-time operation parameter set. For example, the initial load reduction target of 5 minutes required in the load distribution strategy is associated with the historical energy consumption data of the air-conditioning terminal in the first 5 minutes of the current time window. The device response rate data is obtained by calculating the product of the average instruction response time in the historical scheduling record of the air-conditioning terminal and the current compressor operating frequency. For example, if the historical average response delay of a certain terminal is 20 seconds and the current compressor frequency is 50 Hz, the response rate is quantified as 20 seconds × 50 Hz = 1000 response units. The feature superposition process disassembles the load adjustment target into multiple time segments and performs matrix convolution operations with the device response capabilities in the corresponding time periods to generate a load adjustment vector sequence. The real-time electricity price fluctuation characteristics are element-wise weighted concatenated with the load adjustment vector through time dimension expansion, where the weight coefficient is dynamically calculated by the power grid stability evaluation model. For example, when the power grid frequency deviation exceeds 0.1 Hz, the weight of the electricity price feature is increased to 0.7 to ensure that the economic target is subordinate to the stability requirement. The finally generated joint input feature is a high-dimensional tensor structure, and its dimensions include multiple axes such as time segment index, air-conditioning terminal identifier, load adjustment amount, electricity price weight, and environmental parameters, realizing the deep fusion of multi-source heterogeneous data.
[0045] As an implementation manner, in step S200, the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant are subjected to feature alignment processing to generate a joint input feature, which may specifically include:
[0046] Step S210: Parse the load distribution strategy in the dynamic regulation instruction, and extract the load adjustment target value related to the air-conditioning terminal and the corresponding time constraint parameters therein. The time constraint parameters include the start time of load adjustment, the maximum allowable delay time, and the phased adjustment ratio.
[0047] Exemplarily, the parsing process of the dynamic control instruction is based on the structured data message sent by the central control system of the virtual power plant. Among them, the load distribution strategy is encapsulated in JSON format, for example, and may include the total load adjustment amount, adjustment direction (increase / decrease), and constraint conditions in the target area within a specified time window. The load adjustment target value represents the power change amount that the air-conditioning terminal group needs to cooperate to complete. For example, it is required that the total load in the target area be reduced by 50 kW in the next scheduling cycle. The start time in the time constraint parameter is the trigger time point of the load adjustment operation, accurate to the millisecond level in the form of Unix timestamp. For example, the start time is set to 1620000000 (corresponding to 00:00:00 UTC on May 3, 2021). The maximum allowable delay time refers to the maximum tolerable duration from the instruction issuance to the completion of parameter adjustment by the air-conditioning terminal. For example, it is set to 120 seconds. If a certain terminal does not reach the target adjustment value within 120 seconds, it is marked as a response timeout. The phased adjustment ratio decomposes the total load adjustment target into multiple sub-goals according to time segments. For example, it is required to complete a 30% load reduction (15 kW) in the first 5 minutes, and the remaining 70% is completed in the subsequent 10 minutes. In the parsing process, a regular expression matching algorithm is used to extract key parameters, and type conversion and range verification are performed on numerical data. For example, it is verified whether the sum of the phased adjustment ratios is 100%. If an abnormality is detected, the instruction retransmission mechanism is triggered.
[0048] Step S220: Extract the current energy consumption feature subset that matches the time constraint parameter from the real-time operation parameter set. The current energy consumption feature subset includes the energy consumption fluctuation curve and the device response rate data of each air-conditioning terminal within a preset time window before the start time; among them, the device response rate data is calculated by multiplying the average time from receiving the instruction to completing parameter adjustment of the air-conditioning terminal in the historical scheduling record by the compressor operating frequency in the current operating state.
[0049] Exemplarily, the preset time window can be dynamically set according to the starting moment adjusted by the load, and its length is, for example, the historical data backtracking period. For example, the energy consumption data in the 30 minutes before the starting moment is extracted to analyze the recent fluctuation trend. The energy consumption fluctuation curve is composed of the power sampling values of the air-conditioning terminals arranged in time series. For example, the power value sequence of a certain terminal within the time window is [3.2 kW, 3.0 kW, 2.8 kW], reflecting the trend of its energy consumption decreasing over time. The calculation of the device response rate data depends on the historical scheduling record database, which stores the delay time and execution effect of each terminal's previous instruction executions. For example, in the past 10 schedules of a certain terminal, the average delay time from receiving the shutdown instruction to the complete stop of the compressor is 25 seconds, and the current compressor operating frequency is 45 Hz, then the device response rate is quantified as 25 seconds × 45 Hz = 1125 second·Hz. This value characterizes the adjustment ability of the terminal under a specific operating state, and the higher the value, the lower the response efficiency. The extraction process filters out the energy consumption data segments aligned with the current time constraint parameters through the timestamp matching algorithm. For example, when the starting moment is 1620000000, all power sampling points during the period from 1620000000 to 1620001800 (30 minutes) are extracted, and the data missing periods caused by communication interruptions are excluded.
[0050] Step S230: Split the load adjustment target value into sub-load adjustment values corresponding to multiple consecutive time segments based on the time constraint parameters, and map each time segment to the corresponding fluctuation period in the energy consumption fluctuation curve.
[0051] Exemplarily, the time segment division is determined comprehensively based on the phased adjustment ratio and the allowed maximum delay time. For example, the total load adjustment target of 50 kW is completed in two phases: the first phase reduces 15 kW within 5 minutes, and the second phase reduces 35 kW within 10 minutes. The sub-load adjustment values of each time segment are allocated to each air-conditioning terminal through the linear programming algorithm to ensure that the cumulative adjustment amount meets the phase target. When mapping to the energy consumption fluctuation curve, the sliding window matching mechanism is used to identify the fluctuation period corresponding to the time segment. For example, if a certain terminal shows periodic changes in the historical energy consumption fluctuation curve (such as a power peak every 10 minutes), then the sub-load adjustment value of the first 5 minutes in the first phase is mapped to its next expected trough period. During this process, the start and end timestamps of the time segment are aligned with the phase of the energy consumption fluctuation curve. For example, the 5 minutes in the first phase (1620000000 - 1620000300) is mapped to the descending edge period of the energy consumption fluctuation of a certain terminal (the power drops from 3.5 kW to 3.0 kW during the period from 1620000000 to 1620000300). This mapping relationship is realized by calculating the time series similarity through the dynamic time warping algorithm (DTW) to ensure the coordination of the sub-load adjustment value and the actual energy consumption change trend of the terminal.
[0052] Step S240: Superimpose the sub-load adjustment values of each time segment with the device response rate data within the corresponding fluctuation period to generate a load adjustment vector sequence containing time sensitivity.
[0053] Exemplarily, the feature superimposition process can adopt matrix convolution operation to perform element-by-element weighted fusion of the sub-load adjustment value matrix and the device response rate matrix in the time dimension. For example, the sub-load adjustment value of a certain time segment is 5 kW, and the device response rate of the corresponding terminal is 900 s·Hz. Then the superimposed feature vector contains the adjustment amount (5 kW), the response rate (900 s·Hz), and the product of the two (4500 kW·s·Hz). The time sensitivity is quantified by introducing a time decay factor. For example, time segments closer to the current moment are assigned higher weight coefficients. Specifically, for the t-th time segment, the decay factor can be set to 1 / (1 + 0.1t), so that adjustment tasks near the operation window obtain more significant decision weights. The generated load adjustment vector sequence is a three-dimensional tensor structure, and its dimensions include the time segment index, the air-conditioning terminal identifier, and the fusion feature values (adjustment amount, response rate, time sensitivity). For example, the vector of terminal A in the first-stage time segment (index 0) can be expressed as [0, A, 5 kW, 900 s·Hz, 0.91], where 0.91 is the calculation result of the decay factor.
[0054] Step S250: Perform dimension expansion and fusion on the load adjustment vector sequence and the real-time electricity price parameters in the dynamic regulation instruction. The dimension expansion and fusion include splitting the real-time electricity price parameters into gradient change segments according to the time segments and performing element-level weighted splicing with the positions of the same time index in the vector sequence to generate joint input features; among them, the weight coefficients of the weighted splicing are determined by the grid stability priority in the current operation mode of the virtual power plant.
[0055] Exemplarily, the real-time electricity price parameters can be obtained from the power market information platform and split into gradient change segments with the same time granularity as the load adjustment vector sequence according to time segments. For example, if the electricity price corresponding to a certain time segment rises from 0.5 yuan / kWh to 0.6 yuan / kWh, it is split into three ladder values of [0.5, 0.55, 0.6], and each ladder value is concatenated with the load adjustment characteristics corresponding to the time index in the vector sequence. The weight coefficients for element-wise weighted concatenation are dynamically calculated by the power grid stability assessment model, which outputs the stability priority (in the range of 0-1) based on parameters such as real-time frequency deviation and voltage volatility. For example, when the detected frequency deviation exceeds 0.1 Hz, the power grid stability priority rises to 0.8. At this time, the weight coefficient of the electricity price feature is reduced to 0.2, while the weight of the load adjustment feature rises to 0.8 to ensure that the stability goal takes precedence over the economic goal. The concatenated joint input feature is a four-dimensional tensor, including the time segment, terminal identifier, load-electricity price fusion feature, and stability weight. For example, the joint feature of terminal B at time segment 1 can be expressed as [1, B, 3kW, 800 seconds·Hz, 0.85, 0.55 yuan / kWh, 0.2], where 0.2 is the electricity price weight coefficient. This feature structure provides a complete input condition for the subsequent multi-objective optimization model to balance the power grid operation safety and economic dispatch.
[0056] Step S300: Input the joint input feature into the pre-trained multi-objective optimization model, and perform hierarchical weight allocation on the joint input feature through the feature fusion layer in the multi-objective optimization model to generate the air conditioner regulation priority sequence, which is used to describe the scheduling order and parameter adjustment range of each air conditioner terminal under the dynamic regulation instruction.
[0057] Exemplarily, the pre-trained multi-objective optimization model can be constructed using a deep reinforcement learning framework, and its feature fusion layer includes a composite structure of a multi-head attention mechanism and a gated recurrent unit. In the hierarchical weight allocation process, the importance scores of the features of each air-conditioning terminal are first calculated through the attention mechanism. For example, based on the current electricity price fluctuation characteristics, higher attention weights are assigned to the terminals with higher energy consumption. The feature fusion layer further uses a temporal convolutional network to extract the temporal dependence relationship of the joint input features. For example, it can identify the trend feature that the compressor frequency of a certain terminal continuously increases in three consecutive time segments. The generation of the air-conditioning regulation priority sequence includes two core stages: initial priority score calculation and dynamic correction. The initial score is calculated by the policy network according to the output of the feature fusion layer, specifically including the scheduling order score (in the range of 0-1, the higher the value, the more priority for scheduling) and the recommended value of the parameter adjustment range (such as the temperature setting value adjustment of ±2°C). The dynamic correction process introduces the device response delay feature and models the historical scheduling data through the gated recurrent unit. For example, if a certain terminal has an average delay of 8 seconds in the previous three schedules, the current priority score will be attenuated and compensated according to the delay variance. The urgency parameter is sent by the virtual power plant in real time. When the load adjustment demand emergency threshold exceeds the preset threshold (such as the grid frequency deviation exceeds 0.2 Hz), the terminals with a score higher than 0.8 in the priority sequence will be marked as immediate response devices, and the parameter adjustment range is restricted by the maximum allowable offset in the user preference setting feature. For example, if user B sets the maximum allowable temperature offset to 3°C, then even if the model recommends an adjustment of 4°C, the actual scheduling instruction will be automatically truncated to 3°C. The finally generated priority sequence is a structured data table, including the identification of each terminal, the scheduling time (accurate to the second level), the temperature setting value adjustment amount, and the maximum allowable energy consumption limit.
[0058] As an implementation manner, in step S300, the generation process of the priority sequence may include the following steps:
[0059] Step S310: Perform feature importance ranking on the joint input features through the attention mechanism layer in the multi-objective optimization model to generate the initial priority scores of each air-conditioning terminal.
[0060] Exemplarily, as described above, the attention mechanism layer in the multi-objective optimization model can adopt a multi-head self-attention structure. By calculating the correlation weights of different dimensions in the joint input features, the key influencing factors of each air-conditioning terminal under the dynamic regulation instruction are determined. The joint input features serve as the input tensor of the attention layer, and its dimensions include the time segment index, air-conditioning terminal identifier, load adjustment amount, real-time electricity price parameter, and power grid stability weight. For example, for the input feature vector [0.6 kW, 0.55 yuan / kWh, 0.8] of a certain air-conditioning terminal at time segment T1, the attention mechanism calculates its correlation weight with other terminals in the same time period through query-key value matching. Specifically, the query vector is generated by the linear transformation of the load adjustment amount and the real-time electricity price parameter, the key vector is composed of the device response rate and the power grid stability weight, and the attention score is normalized to the [0, 1] interval through the Softmax function. The initial priority score is generated by the dot product operation of the attention weight and the feature vector. For example, if the load adjustment amount weight of a certain terminal is 0.7 and the electricity price parameter weight is 0.3, the initial score is 0.7×0.6 kW + 0.3×0.55 yuan / kWh = 0.645. This score reflects the comprehensive contribution degree of this terminal to the power grid stability and economy within the current regulation window. The higher the score, the higher the scheduling priority.
[0061] Step S320: Input the initial priority score into the gated recurrent unit, and perform temporal correction in combination with the device response delay feature in the historical scheduling record to generate a corrected dynamic priority sequence; wherein, the device response delay feature includes the average delay time and variance data from receiving the instruction to completing the execution for each air-conditioning terminal in the historical scheduling.
[0062] The input layer of the gated recurrent unit (GRU) receives the initial priority score sequence, and its hidden state dimension is aligned with the device response delay feature in the historical scheduling record. The device response delay feature is represented in the form of a two-dimensional vector, including the historical average delay time and delay variance. For example, the average delay time of a certain terminal in the historical scheduling is 25 seconds, and the variance is 4 seconds 2, its delay feature vector is [25, 4]. During the time series correction process, the update gate unit of the GRU dynamically adjusts the influence weight of the historical delay feature according to the initial priority score of the current time segment and the hidden state of the previous moment. For example, if the current initial score is 0.8 and the hidden state of the previous moment indicates a high delay risk, the update gate coefficient is set to 0.6 and the reset gate coefficient is set to 0.4, then the corrected priority score is 0.6×0.8 + 0.4×(1 - 25 / 50) = 0.62, where the 25-second delay is converted into a 0.5 delay impact factor through normalization. The output of the dynamic priority sequence is a corrected score matrix aligned with timestamps. For example, the score sequence of terminal A within the time window [14:00 - 14:05] is [0.62, 0.71, 0.68], reflecting the result of the dynamic adjustment of its priority according to the device response ability.
[0063] Step S330: According to the load adjustment urgency parameter issued by the virtual power plant, perform threshold segmentation on the dynamic priority sequence, and mark the air-conditioning terminals with priority scores higher than the emergency threshold as immediate response devices.
[0064] Exemplarily, the load adjustment urgency parameter is issued in real time by the central control system of the virtual power plant, and its value range is [0, 1]. The larger the value, the higher the scheduling urgency. For example, when the grid frequency deviation exceeds 0.2Hz, the urgency parameter is set to 0.9. The emergency threshold is dynamically calculated through a linear interpolation algorithm, and the formula is: threshold = base threshold + urgency parameter × adjustment coefficient. Assuming the base threshold is 0.6 and the adjustment coefficient is 0.3, the threshold corresponding to the urgency parameter 0.9 is 0.6 + 0.9×0.3 = 0.87. The terminals with scores higher than 0.87 in the dynamic priority sequence are marked as immediate response devices, and their identification information and parameter adjustment instructions are preferentially added to the real-time scheduling queue. For example, if the corrected score of a terminal at time segment T2 is 0.89, exceeding the threshold 0.87, then it is marked as an immediate response device and triggers a 10% reduction in the compressor frequency. The device list after threshold segmentation is broadcast to all terminals in the target area through the message queue protocol of the edge controller to ensure the real-time synchronization of scheduling instructions.
[0065] Step S340: Gradient limit the parameter adjustment range of the immediate response device to ensure that the change in the temperature set value does not exceed the maximum allowable offset in the user preference setting feature.
[0066] Exemplarily, the gradient limit module truncates the temperature setpoint adjustment suggestion according to the maximum allowable offset in the features set by the user preference. For example, User B sets the maximum allowable temperature offset to ±2°C. If the multi-objective optimization model suggests an adjustment of -3°C, the actual adjustment amplitude is limited to -2°C. The limiting algorithm uses the projected gradient descent method to map the out-of-bounds adjustment value to the allowable range. Specifically, let the original adjustment suggestion be ΔT, and the allowable offset be ΔT max , then the corrected adjustment amount ΔT clipped = sign(ΔT) × min(|ΔT|, ΔT max ). For example, ΔT = -3°C, ΔT max = 2°C, then ΔT clipped = -2°C. The gradient limit process synchronously updates the parameter adjustment amplitude field in the priority sequence. For example, the temperature setpoint adjustment of terminal C is corrected from "-2.5°C" to "-2.0°C", and a limit flag bit is appended to the scheduling instruction. This mechanism ensures that the user comfort constraint is not violated while maintaining the maximum feasible execution effect of the power grid scheduling requirements.
[0067] Step S400: Generate a target scheduling strategy based on the priority sequence. The target scheduling strategy includes the start-stop time configuration of each air-conditioning terminal, the adjustment range of the temperature setpoint threshold, and the energy consumption constraint conditions.
[0068] The generation process of the target scheduling strategy first parses the immediate response device identifier in the priority sequence and extracts its corresponding start-stop time configuration parameters. The start-stop time configuration is generated by the timing planning algorithm of the edge controller. For example, a certain terminal needs to turn off the compressor within 30 seconds after the start of the scheduling and restart in the reduced-frequency mode after 5 minutes. The adjustment range of the temperature setpoint threshold is jointly determined according to the suggested amplitude in the priority sequence and the user preference setting features. For example, if the model suggests an adjustment of -2°C and the user allows a maximum offset of -3°C, the actual adjustment range is set to the optimal value within the interval [-2°C, -3°C]. The energy consumption constraint conditions are calculated by the linear programming algorithm to ensure that the cumulative load adjustment amount of all air-conditioning terminals in the target area meets the load distribution strategy of the virtual power plant, and the energy consumption of a single terminal does not exceed 85% of its rated power. For example, if the rated power of a certain terminal is 3kW, the maximum allowable power during its scheduling period is limited to 2.55kW. After receiving the target scheduling strategy, the edge controller converts it into a device control instruction set. The instruction set contains a sequence of drive signals with timestamps. For example, a compressor stop instruction is sent at the Unix timestamp 1620000000, and an instruction to adjust the set temperature to 26°C is sent at 1620000300. The execution timing optimization algorithm further introduces a buffer interval parameter to prevent the power grid impact caused by multiple devices starting and stopping simultaneously. For example, a 50-millisecond time misalignment trigger mechanism is set for the terminals under the same distribution circuit.
[0069] As an implementation manner, in step S400, the execution process of the target scheduling strategy may specifically include:
[0070] Step S410: Parse the air-conditioning terminal identifiers marked as immediate response devices in the priority sequence, extract the corresponding start-stop time configuration and the adjustment range of the temperature setting threshold, and generate a device control instruction set.
[0071] Exemplarily, the generation process of the target scheduling strategy takes the priority sequence as the core input, and generates an executable set of device control instructions by parsing the scheduling order and parameter adjustment range of each air-conditioning terminal. The start-stop time configuration is determined according to the scheduling time field in the priority sequence. For example, a certain terminal is assigned to turn off the compressor at timestamp 1620000300 and restart it in the reduced frequency mode at 1620000600. The adjustment range of the temperature setting threshold is jointly restricted by the recommended amplitude in the priority sequence and the maximum allowable offset in the user preference setting characteristics. For example, if the priority sequence recommends an adjustment of -2.5°C and the user allows a maximum offset of -3°C, the actual adjustment range is set to the optimal value within the interval [-2.5°C, -3°C]. The energy consumption constraint conditions are calculated by a linear programming algorithm to ensure that the cumulative load adjustment amount of all terminals in the target area meets the load distribution strategy of the virtual power plant, and at the same time, the instantaneous power of a single terminal does not exceed 85% of its rated power. For example, if the rated power of a certain terminal is 3 kW, the maximum allowable power limit during its scheduling is 2.55 kW, and it is calibrated in real time through the power monitoring module of the edge controller.
[0072] Exemplarily, the immediate response device identifier is recognized through the emergency flag bit in the priority sequence. For example, if the priority score of a certain terminal is 0.89 and the flag bit is "urgent", its identifier is added to the immediate response queue. The start-stop time configuration is extracted from the scheduling time field of the priority sequence, specifically including the compressor shutdown time, restart time, and the frequency gradual change parameter during the transition stage. For example, the shutdown time of terminal A is 1620000300, the restart time is 1620000600, and the frequency linearly decreases from 45 Hz to 40 Hz. The adjustment range of the temperature setting threshold is parsed from the parameter adjustment field of the priority sequence. For example, the temperature setting value of terminal B is adjusted from 24°C to 22°C, and the maximum allowable hysteresis is ±0.5°C. The device control instruction set is encapsulated in JSON format, including the terminal identifier, operation type (start / stop / temperature adjustment), execution timestamp, and parameter threshold. For example, {"ID":"AC-001", "Action":"SetTemp", "Time":1620000300, "Value":22}.
[0073] Step S420: Sort the device control instruction set by timestamp and send it to the edge controller in the target area. The edge controller adjusts the start and stop times of the compressor according to the buffer interval parameter in the instruction to generate a driving signal set including the execution timing sequence.
[0074] Exemplarily, the timestamp sorting algorithm can adopt ascending order to ensure that the instructions are executed according to the planned timing sequence. For example, the shutdown instruction (1620000300) of terminal A takes precedence over the temperature adjustment instruction (1620000450) of terminal B. Among them, the buffer interval parameter can be dynamically set by the virtual power plant to prevent power grid impacts caused by simultaneous start and stop of multiple devices. For example, the start and stop interval of devices under the same distribution circuit is set to 50 milliseconds. The edge controller adjusts the actual execution time according to the buffer interval. For example, the shutdown instruction of terminal A originally planned to be executed at 1620000300 is postponed to 1620000300.050 due to the buffer interval requirement. After the driving signal set is generated, it is converted into a Modbus RTU protocol frame, including the device address, function code, and data field. For example, the shutdown instruction frame of terminal A is [01 06 0001 0000 CRC].
[0075] Step S430: Real-time collect the actual operation parameters of each air-conditioning terminal after executing the driving signal set, including the compressor state switching delay duration, the execution deviation of the temperature set value, and the instantaneous energy consumption fluctuation amplitude, and generate a dynamic adjustment parameter set.
[0076] The compressor state switching delay duration records the time difference between the instruction sending and the execution completion through a high-precision clock timestamp. For example, terminal A is planned to shut down at 1620000300.050, and the actual completion time is 1620000302.100, with a delay of 2.05 seconds. The execution deviation of the temperature set value is calculated from the difference between the actual temperature and the target temperature feedback by the air-conditioning control board. For example, the set temperature is 22°C while the actual temperature stabilizes at 22.5°C, with a deviation of 0.5°C. The instantaneous energy consumption fluctuation amplitude is quantified by the standard deviation of the sampling values of the power sensor. For example, the standard deviation of the power fluctuation of terminal B within 5 minutes after adjustment is 0.3 kW. The dynamic adjustment parameter set is stored in a time series database, and the data structure includes the timestamp, terminal identifier, delay duration, temperature deviation, and power fluctuation value. For example, {"Time":1620000300, "ID":"AC-001", "Delay":2.05, "TempDev":0.5, "PowerStd":0.3}.
[0077] Step S440: Compare the dynamic adjustment parameter set with the parameter adjustment amplitude in the priority sequence to identify the abnormal device identifier whose execution deviation exceeds the allowable threshold, and generate a policy adjustment trigger event.
[0078] Optionally, the allowable thresholds can be adaptively set according to device types and user preferences. For example, the compressor delay threshold is set to 3 seconds, the temperature deviation threshold is set to 1 °C, and the power fluctuation threshold is set to 10% of the rated power. The comparison process adopts a window sliding mechanism. For example, the dynamic adjustment parameter set is checked every 5 minutes. If the delay of terminal A reaches 3.2 seconds (exceeding the 3 - second threshold), it is marked as an abnormal device. The policy adjustment trigger event includes the abnormal device identifier, deviation type, and severity level. For example, {"EventID":"E - 001", "Device":"AC - 001", "Type":"DelayExceed", "Level":"High"}, and is sent to the virtual power plant dispatching center through the message queue service.
[0079] Step S450: Call the multi - objective optimization model according to the policy adjustment trigger event to perform local recalculation on the current priority sequence, generate an updated target scheduling policy and re - issue it for execution. At the same time, mark the feedback parameter set associated with the abnormal device identifier as a high - priority sample for incremental update.
[0080] Exemplarily, the local recalculation process only updates the combined input features of the sub - region where the abnormal device is located. For example, the feedback parameter set of terminal A is injected into the input layer of the model, triggering fine - tuning of the weights in the feature fusion layer. The updated target scheduling policy is optimized by the differential evolution algorithm. For example, the priority score of terminal A is reduced from 0.89 to 0.75, and its buffer interval is extended to 100 milliseconds. The high - priority samples for incremental update are added to the training data set with a ten - fold weight. For example, the feedback parameters of terminal A are copied 10 times in the historical data pool to accelerate the model's adaptation to abnormal scenarios.
[0081] As an implementation manner, the processing process of the policy adjustment trigger event can specifically include:
[0082] Step S451: Analyze the execution deviation type corresponding to the abnormal device identifier. The execution deviation types include at least one of compressor response delay exceeding the standard, temperature set value locking failure, or instantaneous energy consumption exceeding the limit.
[0083] Exemplarily, the deviation type analysis can be realized by matching the abnormal event description field through regular expressions. For example, the event description "DelayExceed" corresponds to the compressor response delay exceeding the standard, and "TempLockFail" corresponds to the temperature set value locking failure. The determination basis for the compressor response delay exceeding the standard is that the actual delay exceeds the preset threshold. For example, the 3 - second threshold is breached. The temperature set value locking failure means that the air - conditioner control board fails to adjust the temperature according to the instruction. For example, the feedback temperature continuously deviates from the set value by more than 2 °C for 5 minutes. The instantaneous energy consumption exceeding the limit is determined by comparing the percentage of the power fluctuation amplitude with the rated power. For example, the rated power of the terminal is 3 kW, and the instantaneous power reaches 3.3 kW (exceeding the limit by 10%).
[0084] Step S452: Match the preset adjustment rule base according to the deviation type, and extract the candidate adjustment strategies associated with the current virtual power plant operation mode. The candidate adjustment strategies include equipment parameter compensation coefficients, execution timing redundancy amounts, or priority weight correction values.
[0085] Exemplarily, the adjustment rule base can be organized in a decision tree structure, and the nodes are combinations of deviation types and power grid operation modes. For example, when encountering a "DelayExceed" deviation in the "Emergency" mode, the candidate strategies include an equipment parameter compensation coefficient of 1.2 (extending the buffer interval by 20%), an increase in the execution timing redundancy amount of 50 milliseconds, and a reduction in the priority weight correction value of 0.1. The equipment parameter compensation coefficient is obtained through historical data regression analysis. For example, for every 1-second increase in delay exceeding the standard, the compensation coefficient increases by 0.1. The execution timing redundancy amount is dynamically calculated based on the variance of the device response rate, such as a variance of 4 seconds 2 corresponding to a redundancy amount of 50 milliseconds.
[0086] Step S453: Perform constraint verification on the candidate adjustment strategies and the actual operation parameters in the dynamic adjustment parameter set, eliminate the invalid strategies that violate the user preference settings or the power grid stability boundary, and generate a set of feasible strategies.
[0087] Exemplarily, the constraint verification process can use a linear programming solver to verify the feasibility of the strategies. For example, the candidate strategy recommends adjusting the temperature set value of terminal A to increase the compensation coefficient by 1.2, but the maximum offset allowed in the user preference settings is -3°C. If the original adjustment of -2.5°C becomes -3.0°C after compensation, the strategy is valid; if it is compensated to -3.1°C, it is eliminated due to exceeding the boundary. The power grid stability boundary verification simulates the power grid state after the strategy execution through a power flow calculation model. For example, if a strategy causes the load rate of the distribution line to exceed 95%, it is marked as invalid.
[0088] Step S454: Input the set of feasible strategies into the feature fusion layer of the multi-objective optimization model, and combine the real-time electricity price fluctuation feature in the current joint input feature to pre-evaluate the strategy benefits, and generate an updated candidate set of priority sequences.
[0089] Exemplarily, the feature fusion layer encodes the feasible strategies to generate strategy feature vectors. For example, the strategy "increase the buffer interval by 50 milliseconds" is encoded as [0, 50, 0, 0], and the strategy "reduce the priority weight by 0.1" is encoded as [0, 0, -0.1, 0]. The real-time electricity price fluctuation feature extracts trend information through a temporal convolutional network. For example, the electricity price in the current period increases by 0.05 yuan / kWh per hour. The benefit pre-evaluation calculates the comprehensive scores of each strategy through a value network. For example, the score of strategy A is 0.85 (power grid 0.4, user 0.3, energy consumption 0.15), and the score of strategy B is 0.78.
[0090] Step S455: The edge controller performs sandbox simulation execution on the candidate set of the priority sequence, collects the power grid load prediction data and the user satisfaction prediction score after the simulation execution, and selects the candidate strategy with the optimal comprehensive evaluation as the new target scheduling strategy for distribution and execution.
[0091] Exemplarily, the sandbox simulation environment constructs a digital twin model of the target area and runs the simulation after injecting the candidate strategy. The power grid load prediction data is generated by an improved grey model. For example, the total load is predicted to decrease by 65 kW ± 2 kW after executing Strategy A. The user satisfaction prediction score is calculated based on the transfer learning model of historical feedback data. For example, Strategy A is expected to obtain 4.2 / 5.0 points. The comprehensive evaluation uses the weighted summation method. For example, the weight of the power grid index is 0.5, the user weight is 0.3, and the energy consumption weight is 0.2. The total score of Strategy A is 0.85×0.5 + 4.2×0.3 + 0.78×0.2 = 1.923, and the total score of Strategy B is 1.845. Finally, Strategy A is selected to update the target scheduling strategy.
[0092] Step S500: The target scheduling strategy is distributed to each air-conditioning terminal in the target area for execution, and the feedback parameter set after execution is collected in real time. The multi-objective optimization model is updated incrementally using the feedback parameter set.
[0093] Exemplarily, the target scheduling strategy can be distributed to each air-conditioning terminal through industrial Internet of Things protocols (such as MQTT or OPC UA). The terminal controller generates a drive signal after parsing the instruction. The feedback parameter set includes the actual delay duration of the compressor state switch, the execution deviation of the temperature set value, and the instantaneous power fluctuation data. For example, a certain terminal plans to turn off the compressor within 30 seconds, and the actual completion time is 32 seconds, generating a 2-second delay record. The incremental parameter update adopts an online learning mechanism. First, the feedback parameter is compared with the model prediction value to generate an error distribution matrix. For example, the terminal data with an execution deviation of the temperature set value exceeding 1°C is marked as a high-weight sample. The model update process adopts a selective neuron adjustment strategy, and only the hidden layer neurons associated with the abnormal data are optimized by gradient descent. For example, the weight matrix related to the delay feature in the third-layer GRU unit. The version rollback mechanism maintains two model copies on the edge server: the main version accepts real-time updates, and the shadow version is synchronized every 5 minutes. When the scheduling accuracy rate of the main version on the validation set drops by more than 3%, the system automatically switches to the shadow version and triggers an alarm. The updated model parameters are encrypted and transmitted to each edge node through AES-256 to ensure that the data transmission process complies with the safety specifications of the power monitoring system. This step forms a closed-loop optimization mechanism, enabling the multi-objective optimization model to dynamically adapt to external disturbance factors such as equipment aging and user behavior changes, and continuously improving the accuracy and robustness of the scheduling strategy.
[0094] As an implementation manner, in step S500, the process of updating the incremental parameter may specifically include the following steps:
[0095] Step S510: Compare the feedback parameter set with the expected execution effect data for differences to generate a model prediction error distribution graph.
[0096] Exemplarily, the feedback parameter set may include the actual operation data after the air-conditioning terminal executes the scheduling strategy, specifically including the compressor state switching delay duration, the temperature set value execution deviation, and the instantaneous energy consumption fluctuation amplitude. The expected execution effect data is derived from the predicted values of the multi-objective optimization model when generating the priority sequence. For example, the model predicts that the delay duration of a certain terminal is 2 seconds, the temperature deviation is 0.3 °C, and the standard deviation of energy consumption fluctuation is 0.2 kW. The difference comparison algorithm uses tuple-level differential calculation. After aligning the actual parameters with the predicted values according to the time stamp, the differences are calculated item by item to generate an error sequence. For example, the error between the actual delay of 3 seconds and the predicted 2 seconds is +1 second, and the error between the temperature deviation of 0.5 °C and the predicted 0.3 °C is +0.2 °C. The error distribution graph is generated by the kernel density estimation (KDE) algorithm. The horizontal axis represents the range of error values, and the vertical axis represents the probability density of the occurrence of errors. For example, the delay error is concentrated in the interval [-0.5 seconds, +1.5 seconds], and the temperature deviation error is distributed in the range [-0.3 °C, +0.7 °C], forming a multi-peak distribution pattern, intuitively reflecting the prediction deviation tendency of the model under different working conditions.
[0097] Step S520: Extract the abnormal data points in the error distribution graph that exceed the tolerance interval, and trace back to the generation logic path of the corresponding scheduling strategy.
[0098] Exemplarily, the tolerance interval can be adaptively set according to the device type and scheduling objectives. For example, the compressor delay tolerance interval is [-1 second, +3 seconds], the temperature deviation tolerance interval is [-0.5 °C, +1.0 °C], and the energy consumption fluctuation tolerance interval is [-0.5 kW, +0.5 kW]. The abnormal data points are identified by the boundary detection algorithm. For example, if the delay error of a certain terminal +3.2 seconds exceeds the upper limit, it is marked as abnormal. The tracing process relies on the time stamp index of the scheduling log and the model decision tree path to locate the strategy generation link corresponding to the abnormal data. For example, the delay exceeding-standard data is associated with the 3rd layer of the GRU unit in the feature fusion layer in the log. By tracing back, it is found that the device response rate feature was not correctly weighted when the scheduling strategy was generated. The tracing result is stored in the form of a logic path graph, including the time stamp of the abnormal data, the identification of the affected model layer, and the index of the associated weight matrix.
[0099] Step S530: Perform local gradient descent optimization on the neuron connection weights in the multi-objective optimization model related to the abnormal data points.
[0100] Local gradient descent optimization focuses on specific network layers affected by abnormal data, such as the weight matrix of the second attention head in the feature fusion layer. The optimization process calculates the partial derivative of the error corresponding to the abnormal data with respect to the target weight, and the formula is ∂L / ∂W i j, where L is the loss function and W i j is the weight of the i-th row and j-th column. The learning rate is set to 1 / 10 of the global learning rate. For example, when the global learning rate is 0.001, the local learning rate is 0.0001 to prevent overfitting. For example, for abnormal data with a latency error exceeding the standard, the weight matrix W3 of the third fully connected layer of the policy network is adjusted through error backpropagation, and the update amount is ΔW3 = η·(y_true - y_pred)·x i nput, where η = 0.0001 and x i nput is the corresponding joint input feature vector. After optimization, the latency prediction error of the model for similar working conditions is reduced to within 0.8 seconds.
[0101] Step S540: Retain the historical parameter versions that perform stably on the preset validation set, and automatically roll back to the nearest stable version when the performance of the incrementally updated model drops by more than the threshold.
[0102] Exemplarily, the historical parameter versions can be stored in the version control database in the form of snapshots, and each version is attached with validation set evaluation metrics (such as grid benefit score, user satisfaction score). The validation set contains 10% of the data that did not participate in the training during the historical period. For example, 24 hours of data are randomly selected each week in the past three months. The performance drop threshold is set to a 5% relative drop in the comprehensive evaluation metric. For example, if the comprehensive score of the current version is 0.85 and drops to 0.8075 (a 5.0% drop) after incremental update, the rollback mechanism is triggered. The rollback process loads the parameter file of the stable version in the edge controller by comparing the model hash values of the latest version and the nearest stable version. For example, after the performance of version V2.1 drops due to overfitting, the system automatically rolls back to V2.0 and generates an alarm log.
[0103] Step S550: Encrypt and transmit the updated model parameters to the edge computing node for deployment.
[0104] Exemplarily, the encrypted transmission can adopt the AES-256 algorithm combined with the TLS 1.3 protocol for dual protection. The encryption key is dynamically generated and periodically updated by the central management system of the virtual power plant. After the model parameters are serialized into a binary file, they are transmitted in chunks to the edge nodes in the target area through the HTTPS protocol, and each data packet is appended with a digital signature to verify integrity. The edge node deployment process includes: 1) verifying that the hash value of the received parameter file is consistent with the record in the central system; 2) stopping the currently running model instance; 3) loading the new parameters into memory and initializing the model inference engine; 4) switching the traffic to the new model and monitoring the initial running status. For example, after receiving the encrypted parameter file, an edge node restores it to a weight matrix file through a preset decryption key and completes the hot update within 200 milliseconds to ensure the continuity of the dispatching service.
[0105] As an implementation manner, the method provided by the embodiment of the present invention further includes the pre-training process of the above multi-objective optimization model, which may specifically include the following steps:
[0106] Step S10: Obtain the operation parameter set of each air-conditioning terminal in the target area and the corresponding historical regulation instruction set of the virtual power plant within the historical period. The operation parameter set includes historical energy consumption fluctuation characteristics, environmental parameter change curves, and user setting change records.
[0107] The historical period refers to a preset time span, such as the complete dispatching records in the past 12 months, and its time window length is dynamically configured according to the model training requirements. The historical energy consumption fluctuation characteristics are collected by the power metering modules of each air-conditioning terminal in the target area, and include the fluctuation data of the compressor operating frequency, the air supply wind speed, and the difference sequence between the set temperature and the actual temperature on the historical time axis. For example, the energy consumption curve of an air-conditioning terminal shows a power peak from 14:00 to 16:00 every day during the summer historical period, and the peak power reaches 3.5 kW. The environmental parameter change curve is recorded by the temperature and humidity sensors deployed in the area where the air-conditioning terminal is located, including the indoor and outdoor temperature gradients, the humidity change rate, and the light intensity time series data. For example, the average room temperature in a certain area is 28°C ± 2°C in summer and 20°C ± 1.5°C in winter during the historical period, and the humidity fluctuation range is 45% - 65%. The user setting change records are extracted from the log database of the mobile terminal application, and specifically include the modification records of the temperature sensitivity coefficient, the energy-saving priority flag bit, and the allowable regulation time period by the user during the historical period. For example, user A adjusted the temperature sensitivity coefficient 3 times during the historical period, gradually increasing it from the initial value of 0.7 to 0.9. The operation parameter set and the historical regulation instruction set are associated and stored in the distributed database through timestamps to ensure data integrity. For example, each regulation instruction record includes the issued timestamp, the load distribution target value, and the execution status code, forming a mapping relationship with the operation parameters of the air-conditioning terminal in the same time period.
[0108] Step S20: Parse the instruction features of the historical regulation instruction set, extract the load distribution target values, electricity price constraint conditions, and grid emergency event markers contained therein, and generate a historical regulation feature set.
[0109] Exemplarily, the instruction feature parsing process can be to extract structured features from the unstructured historical regulation instruction text using regular expressions and semantic analysis algorithms. The load distribution target value represents the power adjustment amount set by the virtual power plant for each region during the historical period. For example, an instruction requires the target region to reduce the total load by 80 kW during the period from 14:00 to 15:00 on July 15, 2023. The electricity price constraint condition is parsed into the gradient change parameters of the time-of-use electricity price curve, including the peak electricity price period, valley electricity price period, and real-time floating rate. For example, in a certain historical instruction, the electricity price linearly rises from 0.5 yuan / kWh to 0.65 yuan / kWh during the period from 14:00 to 15:00, and the gradient change rate is 0.03 yuan / kWh·minute. The grid emergency event marker is realized by parsing the status code in the instruction. For example, the status code "E1" represents a frequency over-limit event, and "E2" represents a voltage sag event. After the historical regulation feature set is generated, it is stored in a three-dimensional tensor structure, and its dimensions include the time stamp, region identifier, and feature vector. For example, the feature vector of a certain region at the time stamp 1650000000 can be expressed as [80 kW, 0.5 - 0.65 yuan / kWh, E1], corresponding to the load distribution target value, electricity price constraint condition, and grid emergency event marker respectively.
[0110] Step S30: Align the operation parameter set and the historical regulation feature set in the time dimension, and intercept the parameter segments within the same time interval through a sliding window mechanism to generate a historical combined input feature set; among them, the time dimension alignment includes interpolating the air-conditioning terminal data with different sampling frequencies and the regulation instruction time stamp to ensure that the input features of each time segment contain the complete corresponding relationship between the operation parameters and the regulation instructions.
[0111] Exemplarily, the time dimension alignment algorithm can adopt bilinear interpolation to unify data with different sampling frequencies. For example, the air conditioner terminal data is sampled at 30 - second intervals, while the control instructions are issued at 5 - minute intervals. When aligning, the time stamps of the control instructions are extended to 30 - second granularity, and the instruction features at missing time points are filled by adjacent interpolation. The sliding window mechanism intercepts parameter segments of continuous time intervals with a fixed duration. For example, the window length is set to 1 hour and the step size is 30 minutes, and overlapping time series segments are extracted from historical data. The generated historical joint input feature set is a five - dimensional tensor structure, including the time window index, air conditioner terminal identifier, operating parameters, control features, and environmental parameters. For example, the segment with window index W001 contains the compressor frequency sequence [45Hz, 46Hz, 47Hz] of a certain terminal during the period from 14:00 to 15:00 on July 15, 2023, the corresponding load distribution target value of 80kW, and the room temperature data of 28℃ - 29℃ during the same period. The interpolation process ensures that each 30 - second time stamp contains a complete mapping relationship between operating parameters and control instructions, eliminating the problem of feature breakage caused by data asynchrony.
[0112] Step S40: Input the historical joint input feature set into the initialized deep reinforcement learning network. Generate a historical priority sequence through the policy network and evaluate the comprehensive benefit score of the historical priority sequence through the value network.
[0113] Exemplarily, the initialized deep reinforcement learning network consists of a policy network and a value network, and the dimension of its input layer matches that of the historical joint input feature set. The policy network adopts a gated recurrent unit (GRU) structure to process time - series features and outputs the scheduling priority scores and parameter adjustment suggestions of each air conditioner terminal within the historical time window. For example, the output for a certain terminal in time window W001 is a priority score of 0.85 (in the range of 0 - 1, the higher the value, the more priority) and a suggested temperature setting adjustment of - 1.5℃. The value network contains three parallel fully - connected layers, which calculate the grid benefit sub - score, user satisfaction sub - score, and energy consumption efficiency sub - score respectively. The grid benefit sub - score is generated by comparing the mean square error between the actual load curve and the instruction target value. For example, the error between the actual load reduction of 75kW and the target of 80kW is 5kW, corresponding to a score of 0.78. The user satisfaction sub - score is calculated based on the negative correlation between the historical feedback score and the deviation of the temperature setting value. For example, when the execution deviation of a certain terminal is 1℃, the user score drops by 20%, and the score is 0.65. The energy consumption efficiency sub - score is calculated by the dynamic time warping distance between the actual energy consumption sequence and the theoretical adjustment value, and the smaller the distance, the higher the score. The comprehensive benefit score is the weighted sum of the three sub - scores, and the initial weight values are set to [0.4, 0.3, 0.3], which are dynamically optimized during the training process.
[0114] Step S50: Collect the power grid load change data, user feedback data, and equipment energy consumption data after the actual execution of the historical priority sequence. Compare the differences between the data and the comprehensive benefit scores to generate a multi-objective training loss value.
[0115] Exemplarily, the power grid load change data is exported from the energy management system (EMS) of the virtual power plant and includes the actual power curve of the target area after historical dispatching. For example, after the execution of the priority sequence, the total regional load actually decreases by 78 kW from 14:00 to 15:00, with a deviation of 2 kW from the target value of 80 kW. The user feedback data is collected through the rating interface of the mobile application and includes the five-point satisfaction rating and the sentiment analysis results of the text evaluation. For example, user A gives a rating of 4 points after dispatching, and the sentiment analysis result is "basically satisfied". The equipment energy consumption data is transmitted back through the metering module of the air conditioner terminal, recording parameters such as the real-time power and the number of compressor starts and stops after the execution of the priority sequence. The difference comparison process uses a multi-modal data fusion algorithm to compare the actual data item by item with the comprehensive benefit scores predicted by the value network. For example, the predicted value of the power grid benefit sub-score is 0.78, and the calculated matching degree of the actual load curve is 0.75, then the difference value is 0.03. The predicted value of the user satisfaction sub-score is 0.65, and the average actual rating is 3.8 / 5.0 (converted to 0.76), with a difference value of -0.11. The predicted energy consumption efficiency sub-score is 0.70, and the calculated deviation of the actual energy consumption fluctuation from the target is 0.68, with a difference of 0.02. The above difference values constitute the initial components of the multi-objective training loss value.
[0116] Step S60: Use the gradient backpropagation algorithm to iteratively update the policy network parameters of the deep reinforcement learning network until the multi-objective training loss value reaches the convergence threshold, and obtain a pre-trained multi-objective optimization model.
[0117] Exemplarily, the Adam optimizer is used in the gradient backpropagation process, and the initial learning rate is set to 1e-4, which decays by 10% every 1000 training steps. The update direction of the policy network parameters is determined by the total gradient of the multi-objective training loss value, and the total gradient calculates the sum of the partial derivatives of each loss component with respect to the network weights through the chain rule. For example, the gradient partial derivative of the power grid stability loss component is ∂L grid / ∂W, and the partial derivative of the user satisfaction loss component is ∂L user / ∂W. After weighted summation, the policy network weights are updated. The convergence threshold is set according to the loss curve on the validation set. For example, when the loss decrease amplitude in 10 consecutive training cycles is less than 1e-5, the training is terminated. After pre-training, the multi-objective optimization model can generate a priority sequence that takes into account the power grid security, user comfort, and energy consumption efficiency. For example, on the test set, it can achieve a comprehensive performance with a load adjustment target matching degree of 92%, a user satisfaction rating of 4.2 / 5.0, and a 15% improvement in energy consumption efficiency.
[0118] As an implementation, the generation process of the multi-objective training loss value may specifically include:
[0119] Step S501: Decompose the comprehensive benefit score output by the value network into a power grid benefit sub-score, a user satisfaction sub-score, and an energy consumption efficiency sub-score. Each sub-score corresponds to three dimensions of the virtual power plant regulation target. Among them, the power grid benefit sub-score is generated by analyzing the matching degree between the historical load curve and the regulation instruction. The user satisfaction sub-score is generated by statistically analyzing the positive correlation between the historical feedback score and the execution effect of the priority sequence. The energy consumption efficiency sub-score is generated by calculating the deviation attenuation rate of the historical energy consumption fluctuation and the target adjustment value.
[0120] Exemplarily, the calculation of the power grid benefit sub-score can adopt the dynamic time warping (DTW) algorithm to measure the morphological similarity between the actual load curve and the instruction target value. For example, the target load curve in a certain period is a linear decrease of 80 kW, and the actual curve shows a stepwise decrease. The DTW distance is 12 kW·s, and the normalized score is 0.82. The user satisfaction sub-score can be quantified by the Pearson correlation coefficient. For example, if the execution deviation of the priority sequence (such as the temperature setting error) is negatively correlated with the user score (r = -0.85), the score is mapped to (1 + r) / 2 = 0.575. The energy consumption efficiency sub-score can use the exponential decay function to calculate the deviation attenuation rate, and the formula is score = exp(-λ·ΔE), where ΔE is the root mean square error between the actual energy consumption and the target value, and λ is the attenuation coefficient. For example, when ΔE = 5 kW and λ = 0.1, the score is exp(-0.5) = 0.606.
[0121] Step S502: Extract the matching degree between the actual load curve and the expected load curve in each time segment from the power grid load change data, and compare the difference between the matching degree and the power grid benefit sub-score to generate the power grid stability loss component.
[0122] Exemplarily, the matching degree calculation can adopt the piecewise integration method. For example, the load curve is divided into several intervals according to time segments, and the area coincidence degree between the actual value and the expected value in each interval is calculated. For example, in a certain 5-minute time segment, the expected load decreases by 15 kW, and the actual load decreases by 12 kW. The coincidence area is 72% of the expected curve, and the matching degree is 0.72. The power grid stability loss component is calculated by the mean square error (MSE), and the formula is L grid = Σ (matching degree - power grid benefit sub-score) 2 / N. For example, in a certain window, the matching degree sequence is [0.72, 0.85, 0.91], and the predicted power grid benefit sub-score is [0.78, 0.82, 0.88]. Then the loss component is ((0.72 - 0.78) 2 + (0.85 - 0.82) 2+(0.91 - 0.88) 2 ) / 3 = 0.0013。
[0123] Step S503: Extract the actual satisfaction scores of each user from the user feedback data, perform probability density difference analysis on the score distribution curve and the user satisfaction sub - scores, and generate the user satisfaction loss component.
[0124] Exemplarily, the probability density difference analysis uses the Kullback - Leibler (KL) divergence to measure the deviation between the predicted score distribution and the actual scores. For example, the predicted distribution of the user satisfaction sub - scores is a normal distribution with a mean of 0.7 and a variance of 0.1, and the actual score distribution has a mean of 0.76 and a variance of 0.05. The KL divergence is calculated to be 0.15. The user satisfaction loss component is defined as the weighted sum of the KL divergences, and the formula is L user = Σ w i ·D KL (P i ||Q i ), where P i is the actual distribution, Q i is the predicted distribution, and w i is the user weight coefficient. For example, the weight of high - value users is set to 1.5, and that of ordinary users is 1.0.
[0125] Step S504: Extract the actual energy consumption sequence from the device energy consumption data, calculate the trend alignment degree between the sequence fluctuation characteristics and the energy consumption efficiency sub - scores, and generate the energy consumption deviation loss component.
[0126] Exemplarily, the trend alignment degree can be measured by the cosine similarity to measure the trend consistency between the actual energy consumption sequence and the target adjustment value. For example, the actual energy consumption sequence shows a V - shaped trend of first decreasing and then increasing within a time window, while the target value is monotonically decreasing, and the cosine similarity is - 0.3. The energy consumption deviation loss component is defined as (1 - similarity) / 2, so that the loss is 0 when they are completely consistent and 1 when they are completely opposite. For example, the loss component corresponding to a similarity of - 0.3 is (1 - (- 0.3)) / 2 = 0.65.
[0127] Step S505: Input the power grid stability loss component, the user satisfaction loss component, and the energy consumption deviation loss component into the dynamic weight allocator, adjust the weight ratios of each component based on the power grid operation mode markers in the historical regulation instruction set, where the weight ratios are updated synchronously with the evaluation weights of the value network for the three dimensions; the parameter update of the dynamic weight allocator shares the back - propagation link with the gradient descent process of the policy network.
[0128] Exemplarily, the dynamic weight allocator can be implemented using an attention mechanism, whose query vector is the current power grid operation mode label (such as "normal", "emergency", "economic"), and the key-value pair is the historical weight of each loss component. For example, in the "emergency" mode, the weight of power grid stability loss is increased from 0.4 to 0.7, the weight of user satisfaction is decreased to 0.2, and the weight of energy consumption efficiency is 0.1. The weight coefficients are normalized by the Softmax function to ensure that the sum is 1. During the parameter update process, the gradient ∂L total / ∂w i of the dynamic weight allocator and the gradient ∂L total / ∂θ of the policy network jointly participate in backpropagation to achieve end-to-end joint optimization. Step S506: Input the weighted loss components into the non-linear fusion unit, calculate the total gradient value through the backpropagation path of the value network, and generate a multi-objective training loss value consistent with the optimization direction of the comprehensive benefit score.
[0129] The non-linear fusion unit consists of two layers of fully connected neural networks, and the activation function is LeakyReLU. The weighted loss components undergo non-linear transformation by the fusion unit and output the total loss value L total =α·L grid +β·L user +γ·L energy , where α, β, and γ are dynamic weight coefficients. The total gradient value is calculated by the chain rule, and the formula is ∂L total / ∂θ=∂L total / ∂L grid ·∂L grid / ∂θ+∂L total / ∂L user ·∂L user / ∂θ+∂L total / ∂L energy ·∂L energy / ∂θ. The backpropagation path shares the parameter space of the value network and the policy network to ensure that the gradient update direction is consistent with the Pareto optimal solution of the comprehensive benefit score. For example, when the weight of power grid stability loss is high, the gradient update preferentially optimizes the performance of the policy network in load tracking.
[0130] In some alternative embodiments, after the incremental parameter update in step S500, the method provided by the embodiments of the present invention may further include the following derivative steps:
[0131] Step S600: Predictive priority pre-generation is performed on air conditioner terminals in the target area that have not executed the scheduling policy based on the updated multi-objective optimization model. The pre-generation process includes: extracting the actual execution delay characteristics and energy consumption adjustment efficiency characteristics of each air conditioner terminal in the feedback parameter set to generate a device response ability profile, where the device response ability profile includes the historical average command parsing duration, the success probability of compressor state switching, and the execution accuracy of the temperature set value. The device response ability profile is dynamically associated and matched with the real-time electricity price fluctuation characteristics in the current combined input features to predict the potential priority scores of unexecuted devices in the next scheduling cycle.
[0132] The predictive priority pre-generation mechanism uses the incrementally updated multi-objective optimization model to perform forward-looking priority assessment on air conditioner terminals that have not received regulation instructions during the current scheduling cycle. Air conditioner terminals that have not executed the scheduling policy refer to the set of terminals that have not been marked as immediately responsive devices due to insufficient urgency or user preference restrictions, and their identification information is obtained through real-time screening of the device status database of the edge controller. The pre-generation process simulates scenarios by constructing combined input features for the future time window, combines historical response capabilities with real-time power grid status data, and generates potential priority scores to guide the strategy preparation for the next scheduling cycle. For example, a certain terminal is not scheduled because its current priority score of 0.72 is lower than the emergency threshold of 0.8, but it is predicted that its score may increase to 0.85 when the electricity price rises to 0.7 yuan / kWh in the next cycle, so it is included in the preparation strategy queue.
[0133] The actual execution delay feature is calculated by statistically averaging the time difference from when an instruction is issued to when it is completed in historical scheduling records. For example, a certain terminal has an average delay of 28 seconds with a standard deviation of 3 seconds in the past 50 schedules. The energy consumption adjustment efficiency feature is defined as the ratio of the actual energy consumption change rate of the terminal to the target adjustment rate. For example, if the target requires a power reduction of 1 kW within 5 minutes and the actual completion time is 5 minutes and 20 seconds, the efficiency is (1 kW / 320 seconds) = 0.003125 kW / s. The device response ability profile is stored in a structured data table, with each terminal corresponding to a row of records, including: 1) The historical average instruction parsing duration, calculated by averaging the difference between the instruction reception timestamp and the execution confirmation timestamp. For example, the average parsing duration of terminal A is 2.5 seconds; 2) The success probability of compressor state switching, which is the ratio of the number of times the compressor starts and stops according to the instruction in historical scheduling. For example, the success rate is 48 / 50 = 96%; 3) The execution accuracy of the temperature set value, defined as the root mean square error (RMSE) between the actual temperature and the target temperature. For example, the RMSE of terminal B is 0.4 °C. The profile data is dynamically updated through the time series database of the edge computing node to ensure that it reflects the latest state of the device. The dynamic association and matching algorithm adopts a fusion architecture of the attention mechanism and the time series prediction model. The real-time electricity price fluctuation feature is split into gradient segments according to the time granularity of the next scheduling period. For example, the electricity price in the next 30 minutes increases step by step from 0.6 yuan / kWh to 0.65 yuan / kWh. The historical average instruction parsing duration in the device response ability profile is convolved with the electricity price gradient change rate to generate a time sensitivity weight. For example, the parsing duration of terminal A is 28 seconds, corresponding to a time sensitivity coefficient of 0.7 during the electricity price increase stage, and the predicted increase in the priority score during the electricity price peak period is 0.7 × 0.05 yuan / kWh = 0.035. The potential priority score is calculated by the preliminary strategy branch network of the multi-objective optimization model, and the input features include the device response ability profile vector, real-time electricity price parameters, and grid frequency prediction data. For example, the input feature vector of terminal B is [28 seconds, 96%, 0.4 °C, 0.6 - 0.65 yuan / kWh, 49.98 Hz], and the output potential priority score is 0.82.
[0134] Step S700: Pre-adjust the start and stop time configurations of unexecuted devices according to the potential priority score to generate a candidate scheduling policy set including pre-buffer time slots; the pre-buffer time slots are dynamically calculated and generated based on the device response delay feature and the grid frequency fluctuation trend.
[0135] Exemplarily, the pre-adjustment algorithm can potentially use the priority score as the sorting basis to virtually allocate the start-stop time windows for unexecuted devices. For example, the potential score of terminal C is 0.78, corresponding to the time window from the 120th to the 150th second after the start moment of the next cycle. The length of the pre-buffer time slot is dynamically determined by the product of the device response delay characteristic and the grid frequency volatility. The formula is: pre-buffer time slot = historical average delay × (1 + frequency volatility). Assuming that the historical delay of terminal C is 30 seconds and the current grid frequency volatility is 0.2% (relative to 50 Hz), then the pre-buffer time slot = 30 × (1 + 0.002) = 30.06 seconds, rounded to 30 seconds. The candidate scheduling policy set contains multiple virtual time configuration schemes. For example, policy 1 sets the start moment of terminal C to 120 seconds after the start of the next cycle, and policy 2 sets it to 135 seconds, each with pre-buffer time slots of 30 seconds and 35 seconds respectively, forming an optional set of timing arrangements.
[0136] Step S800: Input the candidate scheduling policy set into the sandbox simulation environment of the edge controller, and collect the predicted grid load curve and the predicted user satisfaction value after the simulation execution.
[0137] The sandbox simulation environment constructs a virtual grid topology model and a digital twin of the air-conditioning terminal in the target area, and runs a high-precision simulation after injecting the candidate scheduling policy. The predicted grid load curve is generated by an improved ARIMA model. For example, after the simulation execution of policy 1, the total load linearly decreases from 500 kW to 470 kW in the next cycle, and for policy 2, it decreases to 465 kW. The predicted user satisfaction value is calculated based on the regression model of historical feedback data and the execution accuracy of the temperature setting value. For example, the predicted score of policy 1 is 4.1 / 5.0, and that of policy 2 is 4.3 / 5.0. The simulation results are stored as a multi-dimensional evaluation matrix, including the load change rate, the average user score, and the device energy consumption efficiency index of each policy. For example, the matrix item of policy 1 is [470 kW, 4.1, 0.85], and that of policy 2 is [465 kW, 4.3, 0.82].
[0138] Step S900: Sort the candidate policies according to the deviation degree between the predicted curve and the predicted value, select the policy with the smallest deviation degree as the preparatory policy for the next scheduling cycle, and perform timing connection with the current target scheduling policy.
[0139] Exemplarily, the deviation degree calculation can use the Mahalanobis distance to measure the multi-dimensional difference between the predicted result and the ideal target value. For example, the ideal target value is set as the Pareto front point with the maximum load reduction, the highest user score, and the optimal energy consumption efficiency. For example, the load deviation degree of policy 1 = |470 - 450| = 20 kW, the user score deviation degree = |4.1 - 5.0| = 0.9, and the corresponding deviation degrees of policy 2 are 15 kW and 0.7, and the Mahalanobis distances are √(20 2 + 0.9 2) = 20.02 and √(15 2 +0.7 2 ) = 15.02. Strategy 2 was selected as the preliminary strategy due to its smaller deviation. The time-series connection process is implemented through the scheduling time-series coordination module of the edge controller, seamlessly embedding the time window of the preliminary strategy into the current execution queue. For example, if the end time of the current strategy is 1620003600, the start time of the preliminary strategy is set to 1620003600 + buffer interval of 30 seconds = 1620003630.
[0140] As an implementation, in step S900, the process of performing time-series connection may include the following steps:
[0141] Step S901: Analyze the remaining execution time periods of each air-conditioning terminal in the current target scheduling strategy, and extract the device identifiers that have overlapping conflicts with the pre-adjusted start and stop times in the preliminary strategy.
[0142] The remaining execution time periods are dynamically calculated through the timestamp queue of the current strategy. For example, the planned shutdown period of terminal D is 1620003000 - 1620003600, and the start time of terminal E in the preliminary strategy is set to 1620003300. Then the overlapping conflict period is 1620003300 - 1620003600. The conflicting device identifiers are identified through the time interval intersection detection algorithm. For example, there are overlapping instructions between terminal D and E during 1620003300 - 1620003600, which are marked as a conflicting device pair (D, E). The conflict information is recorded as a triple data structure, including the conflicting device identifier, the start and end timestamps of the overlapping period, and the conflict type (such as start / stop conflict, power superposition conflict).
[0143] Step S902: Perform two-way sliding window matching on the execution time periods of the conflicting devices to find idle time slots in the remaining execution time periods that allow the insertion of the preliminary strategy operations; among them, the two-way sliding window matching includes forward searching for idle periods after the end time of the current strategy and backward searching for compressible periods before the start time of the current strategy.
[0144] For example, a bidirectional sliding window algorithm uses the conflict period as the center and expands the search for available time slots in both forward and backward directions. The forward search range extends from the end of the current policy to the start of the next cycle, for example, looking for a continuous idle window between 1620003600 and 1620007200. The backward search compresses the execution period of non-critical devices in the current policy, for example, by advancing the shutdown time of terminal F from 1620003200 to 1620003100 to free up a 100-second idle window. The availability of idle time slots is verified by combining power flow calculations with device response delay constraints. For example, a 100-second time slot must meet the following conditions: 1) the distribution line load factor after inserting the backup policy is ≤ 90%; 2) the device pre-buffer time slot is ≥ the historical average delay × 1.2. If both conditions are met, the time slot is marked as available.
[0145] Step S903: adjusting the pre-buffer time slot in the preparation strategy according to the length of the idle time slot, and generating a time-aligned connection strategy instruction set.
[0146] Exemplarily, the pre-buffer slot adjustment algorithm dynamically scales based on the available idle slot margin. For example, if the pre-buffer slot in the pre-policy is originally 30 seconds, and the available idle slot length is 85 seconds, the pre-buffer slot is adjusted to 30 seconds (maintaining the lower limit) to retain a 55-second safety margin. If the idle slot length is only 25 seconds, the pre-buffer slot is compressed to 25 seconds and a device response rate check is triggered. The timestamps of the connection policy instruction set are rearranged. For example, if the startup time of terminal E is adjusted from 1620003630 to 1620003550, the pre-buffer slot is reduced from 30 seconds to 20 seconds, and a new instruction frame [ID: E, Action: Start, Time: 1620003550, Buffer: 20] is generated.
[0147] Step S904: Real-time monitoring of grid load fluctuation data. When it is detected that the execution period in the connection strategy instruction set matches the grid frequency adjustment demand, a dynamic strategy switching operation is triggered. The triggering conditions of the dynamic strategy switching operation include the urgency of the grid frequency regulation instruction exceeding a preset threshold or the user comfort feedback value being lower than the safety critical line.
[0148] For example, grid frequency adjustment requirements are captured through the frequency deviation signal from the virtual power plant. For example, when the frequency drops to 49.8 Hz, the urgency parameter for triggering frequency regulation is raised to 0.9 (threshold 0.8). The user comfort safety threshold is set at a satisfaction score of 3.5 / 5.0. If the average of real-time feedback drops to 3.4, a switch is triggered. Dynamic switching operations are executed by the edge controller's immediate decision-making module. For example, if a frequency deviation exceeding 0.2 Hz is detected, the current non-critical strategy is forcibly interrupted and instructions with a priority score ≥ 0.85 in the transition strategy are immediately executed.
[0149] Step S905: Embed the switched connection strategy instruction set into the execution queue of the current target scheduling strategy according to the timestamp, and update the time constraint parameter in the input features of the multi-objective optimization model.
[0150] Exemplarily, the timestamp embedding process adopts a priority pre-emptive scheduling algorithm. For example, the shutdown instruction of terminal G in the connection strategy (priority 0.88) is inserted before the instruction with priority 0.75 in the current queue. After the execution queue is reconstructed, the time constraint parameter is synchronously updated in the input features of the multi-objective optimization model. For example, the start time of the next scheduling period is adjusted from the original 1620007200 to 1620003630, and the time window length is shortened from 3600 seconds to 3570 seconds. The time axis of the model input features is re-interpolated and aligned to ensure that the subsequent optimization process generates strategies based on the latest scheduling time constraints.
[0151] According to another aspect of the present invention, there is also provided an air conditioner scheduling device based on the response of a virtual power plant. Please refer to Figure 3 , the device 900 includes:
[0152] A parameter acquisition module 910, configured to acquire a real-time operation parameter set of each air conditioner terminal in the target area within a preset time window, where the real-time operation parameter set includes a current energy consumption feature, an ambient temperature and humidity feature, and a user preference setting feature, and the user preference setting feature is used to describe the priority configuration of the user for the air conditioner operation mode;
[0153] A feature alignment module 920, configured to perform feature alignment processing on the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant to generate combined input features, where the dynamic regulation instruction includes a load distribution strategy and a real-time electricity price fluctuation feature of the virtual power plant for the target area;
[0154] A priority generation module 930, configured to input the combined input features into a pre-trained multi-objective optimization model, and perform hierarchical weight distribution on the combined input features through a feature fusion layer in the multi-objective optimization model to generate an air conditioner regulation priority sequence, where the priority sequence is used to describe the scheduling order and parameter adjustment range of each air conditioner terminal under the dynamic regulation instruction;
[0155] A strategy generation module 940, configured to generate a target scheduling strategy based on the priority sequence, where the target scheduling strategy includes the start and stop time configurations of each air conditioner terminal, the adjustment range of the temperature setting threshold, and the energy consumption constraint condition;
[0156] An incremental update module 950, configured to issue the target scheduling strategy to each air conditioner terminal in the target area for execution, and collect the feedback parameter set after execution in real time, and use the feedback parameter set to perform incremental parameter update on the multi-objective optimization model.
[0157] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. In some embodiments, the functions or modules included in the device provided by the embodiments of the present invention can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present invention, please refer to the description of the method embodiments of the present invention for understanding.
[0158] It should be noted that in the embodiments of the present invention, if the above-mentioned air-conditioning scheduling method based on the response of a virtual power plant is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present invention are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0159] It should be understood that various forms of the processes shown above can be used, reordering, adding, or deleting steps. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. No limitations are imposed herein.
Claims
1. An air conditioner scheduling method based on the response of a virtual power plant, characterized in that The method includes: Obtaining a real-time operation parameter set of each air-conditioning terminal in the target area within a preset time window, where the real-time operation parameter set includes current energy consumption characteristics, ambient temperature and humidity characteristics, and user preference setting characteristics, and the user preference setting characteristics are used to describe the priority configuration of the user for the air-conditioning operation mode; Performing feature alignment processing on the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant to generate a combined input feature, where the dynamic regulation instruction includes a load distribution strategy of the virtual power plant for the target area and real-time electricity price fluctuation characteristics; Inputting the combined input feature into a pre-trained multi-objective optimization model, and performing hierarchical weight allocation on the combined input feature through a feature fusion layer in the multi-objective optimization model to generate an air-conditioning regulation priority sequence, where the priority sequence is used to describe the scheduling order and parameter adjustment range of each air-conditioning terminal under the dynamic regulation instruction; Generating a target scheduling strategy based on the priority sequence, where the target scheduling strategy includes start-stop time configuration, temperature setting threshold adjustment range, and energy consumption constraint conditions of each air-conditioning terminal; Issuing the target scheduling strategy to each air-conditioning terminal in the target area for execution, and collecting a feedback parameter set after execution in real time, and using the feedback parameter set to perform incremental parameter update on the multi-objective optimization model; Among them, the performing feature alignment processing on the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant to generate a combined input feature includes: Analyzing the load distribution strategy in the dynamic regulation instruction, and extracting the load adjustment target value related to the air-conditioning terminal and the corresponding time constraint parameter therein, where the time constraint parameter includes the start time of load adjustment; Extracting a current energy consumption feature subset matching the time constraint parameter from the real-time operation parameter set, where the current energy consumption feature subset includes the energy consumption fluctuation curve and equipment response rate data of each air-conditioning terminal within a preset time window before the start time; Splitting the load adjustment target value into sub-load adjustment values corresponding to multiple consecutive time segments based on the time constraint parameter, and mapping each time segment to the corresponding fluctuation period in the energy consumption fluctuation curve; Performing feature superposition on the sub-load adjustment values of each time segment and the equipment response rate data within the corresponding fluctuation period to generate a load adjustment vector sequence including time sensitivity; Performing dimension expansion and fusion on the load adjustment vector sequence and the real-time electricity price parameter in the dynamic regulation instruction, where the dimension expansion and fusion includes splitting the real-time electricity price parameter into gradient change segments according to time segments and performing element-wise weighted splicing with the positions of the same time index in the vector sequence to generate the combined input feature; Among them, the equipment response rate data is obtained by multiplying the average time from receiving the instruction to completing the parameter adjustment of the air-conditioning terminal in the historical scheduling record by the compressor operating frequency in the current operating state, and the weight coefficient of the weighted splicing is determined by the grid stability priority in the current operating mode of the virtual power plant; Among them, the execution process of the target scheduling strategy includes: Analyze the air-conditioning terminal identifiers marked as immediate response devices in the priority sequence, extract the corresponding start-stop time configuration and temperature setting threshold adjustment range, and generate a device control instruction set; Sort the device control instruction set by timestamp and send it to the edge controller in the target area. The edge controller adjusts the start and stop times of the compressor according to the buffer interval parameter in the instruction to generate a driving signal set including the execution time sequence; Collect the actual operation parameters of each air-conditioning terminal after executing the driving signal set in real time, including the compressor state switching delay duration, the execution deviation of the temperature setting value, and the instantaneous energy consumption fluctuation amplitude, and generate a dynamic adjustment parameter set; Compare the dynamic adjustment parameter set with the parameter adjustment range in the priority sequence, identify the abnormal device identifiers with execution deviations exceeding the allowable threshold, and generate a policy adjustment trigger event; According to the policy adjustment trigger event, call the multi-objective optimization model to perform local recalculation on the current priority sequence, generate an updated target scheduling policy and resend it for execution. At the same time, mark the feedback parameter set associated with the abnormal device identifier as a high-priority sample for incremental update.
2. The method according to claim 1, characterized in that, The obtaining of the real-time operation parameter set of each air-conditioning terminal in the target area within a preset time window includes: Collect the operation status data of each air-conditioning terminal in the target area in multiple consecutive time segments through the Internet of Things gateway. The operation status data includes the compressor operating frequency, the air supply speed, and the difference sequence between the set temperature and the actual temperature; Perform outlier removal and timestamp alignment processing on the operation status data to generate a standardized operation parameter set; Extract the preference setting instructions submitted by the user through the mobile terminal. The preference setting instructions include the temperature sensitivity coefficient, the energy-saving priority flag bit, and the allowable regulation time period; Perform feature splicing on the standardized operation parameter set and the preference setting instructions to generate the real-time operation parameter set including spatio-temporal correlation; Among them, the timestamp alignment processing includes mapping sensor data with different sampling frequencies to the same time reference axis, and the outlier removal includes identifying compressor current values or temperature and humidity readings that exceed the preset physical threshold.
3. The method according to claim 1, characterized in that, The pre-training process of the multi-objective optimization model includes: Obtain the operation parameter set of each air-conditioning terminal in the target area and the corresponding virtual power plant historical regulation instruction set within the historical period. The operation parameter set includes historical energy consumption fluctuation characteristics, environmental parameter change curves, and user setting change records; Perform instruction feature analysis on the historical regulation instruction set, extract the load distribution target value, electricity price constraint conditions, and grid emergency event marks included therein, and generate a historical regulation feature set; Align the operation parameter set and the historical regulation feature set in the time dimension, and intercept the parameter segments within the same time interval through the sliding window mechanism to generate a historical combined input feature set; Input the historical combined input feature set into the initialized deep reinforcement learning network, generate a historical priority sequence through the policy network, and evaluate the comprehensive benefit score of the historical priority sequence through the value network; Collect the power grid load change data, user feedback data, and equipment energy consumption data after actually executing the historical priority sequence, compare the differences between the data and the comprehensive benefit score, and generate a multi-objective training loss value; Use the gradient backpropagation algorithm to iteratively update the policy network parameters of the deep reinforcement learning network until the multi-objective training loss value reaches the convergence threshold, and obtain the pre-trained multi-objective optimization model; Among them, the time dimension alignment includes interpolating the air-conditioning terminal data with different sampling frequencies and the regulation instruction timestamp to ensure that the input features of each time segment contain the complete corresponding relationship between the operating parameters and the regulation instructions.
4. The method according to claim 3, wherein The generation process of the multi-objective training loss value includes: Decompose the comprehensive benefit score output by the value network into a power grid benefit sub-score, a user satisfaction sub-score, and an energy consumption efficiency sub-score, and each sub-score corresponds to three dimensions of the virtual power plant regulation target; Extract the matching degree between the actual load curve and the expected load curve in each time segment from the power grid load change data, compare the differences between the matching degree and the power grid benefit sub-score, and generate a power grid stability loss component; Extract the actual satisfaction scores of each user from the user feedback data, perform probability density difference analysis on the score distribution curve and the user satisfaction sub-score, and generate a user satisfaction loss component; Extract the actual energy consumption sequence from the equipment energy consumption data, calculate the trend alignment degree between the sequence fluctuation characteristics and the energy consumption efficiency sub-score, and generate an energy consumption deviation loss component; Input the power grid stability loss component, user satisfaction loss component, and energy consumption deviation loss component into the dynamic weight allocator, and adjust the weight ratio of each component based on the power grid operation mode mark in the historical regulation instruction set, where the weight ratio is updated synchronously with the evaluation weights of the value network for the three dimensions; Input the weighted loss components into the non-linear fusion unit, calculate the total gradient value through the backpropagation path of the value network, and generate a multi-objective training loss value consistent with the optimization direction of the comprehensive benefit score; Among them, the power grid benefit sub-score is generated by analyzing the matching degree between the historical load curve and the regulation instruction, the user satisfaction sub-score is generated by statistically analyzing the positive correlation between the historical feedback score and the execution effect of the priority sequence, the energy consumption efficiency sub-score is generated by calculating the deviation decay rate between the historical energy consumption fluctuation and the target adjustment value, and the parameter update of the dynamic weight allocator shares the backpropagation link with the gradient descent process of the policy network.
5. The method according to claim 1, characterized in that, The generation process of the priority sequence includes: Rank the feature importance of the joint input features through the attention mechanism layer in the multi-objective optimization model to generate the initial priority scores of each air-conditioning terminal; Input the initial priority scores into the gated recurrent unit, and perform timing correction in combination with the equipment response delay characteristics in the historical scheduling records to generate a corrected dynamic priority sequence; According to the load adjustment urgency parameter issued by the virtual power plant, perform threshold segmentation on the dynamic priority sequence, and mark the air-conditioning terminals with priority scores higher than the emergency threshold as immediately responsive devices; Perform gradient limit on the parameter adjustment range of the immediate response device to ensure that the change in the temperature set value does not exceed the maximum allowable offset in the user preference setting features.
6. The method according to claim 1, characterized in that, The processing procedure of the policy adjustment trigger event includes: Analyze the execution deviation types corresponding to the abnormal device identifiers, where the execution deviation types include at least one of compressor response delay exceeding the standard, temperature set value locking failure, or instantaneous energy consumption exceeding the limit; Match the preset adjustment rule library according to the deviation type, and extract candidate adjustment strategies associated with the current virtual power plant operation mode. The candidate adjustment strategies include device parameter compensation coefficients, execution timing redundancy amounts, or priority weight correction values; Perform constraint verification on the candidate adjustment strategies and the actual operation parameters in the dynamic adjustment parameter set, eliminate invalid strategies that violate user preference setting features or grid stability boundaries, and generate a set of feasible strategies; Input the set of feasible strategies into the feature fusion layer of the multi-objective optimization model, and combine the real-time electricity price fluctuation features in the current joint input features to pre-evaluate the strategy benefits, and generate an updated candidate set of priority sequences; Perform sandbox simulation execution on the candidate set of priority sequences through the edge controller, collect the grid load prediction data and user satisfaction prediction scores after the simulation execution, and select the candidate strategy with the best comprehensive evaluation as the new target scheduling strategy for distribution and execution.
7. The method according to claim 1, wherein The process of incremental parameter update includes: Compare the feedback parameter set with the expected execution effect data to generate a model prediction error distribution map; Extract the abnormal data points in the error distribution map that exceed the tolerance interval, and trace back to the generation logic path of the corresponding scheduling strategy; Perform local gradient descent optimization on the neuron connection weights in the multi-objective optimization model related to the abnormal data points; Retain the historical parameter version that performs stably on the preset validation set, and automatically roll back to the nearest stable version when the performance of the model after incremental update drops by more than the threshold; Encrypt and transmit the updated model parameters to the edge computing node for deployment.
8. An air conditioner scheduling device based on the response of a virtual power plant, characterized in that, Include: A parameter acquisition module for acquiring the real-time operation parameter set of each air-conditioning terminal in the target area within a preset time window. The real-time operation parameter set includes current energy consumption characteristics, environmental temperature and humidity characteristics, and user preference setting characteristics. The user preference setting characteristics are used to describe the user's priority configuration for the air-conditioning operation mode; A feature alignment module for performing feature alignment processing on the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant to generate joint input features. The dynamic regulation instruction includes the load distribution strategy of the virtual power plant for the target area and the real-time electricity price fluctuation characteristics; A priority generation module for inputting the joint input features into a pre-trained multi-objective optimization model, and performing hierarchical weight allocation on the joint input features through the feature fusion layer in the multi-objective optimization model to generate an air-conditioning regulation priority sequence. The priority sequence is used to describe the scheduling order and parameter adjustment range of each air-conditioning terminal under the dynamic regulation instruction. A policy generation module, configured to generate a target scheduling policy based on the priority sequence, where the target scheduling policy includes the start-stop time configuration of each air-conditioning terminal, the adjustment range of the temperature setting threshold, and the energy consumption constraint conditions; An incremental update module, configured to send the target scheduling policy to each air-conditioning terminal in the target area for execution, and collect the feedback parameter set after execution in real time, and use the feedback parameter set to perform incremental parameter update on the multi-objective optimization model; Among them, the feature alignment process of the real-time operation parameter set and the dynamic regulation instruction issued by the virtual power plant to generate the joint input feature includes: Analyze the load distribution policy in the dynamic regulation instruction, and extract the load adjustment target value related to the air-conditioning terminal and the corresponding time constraint parameters therein. The time constraint parameters include the start time of the load adjustment; Extract the current energy consumption feature subset matching the time constraint parameters from the real-time operation parameter set. The current energy consumption feature subset includes the energy consumption fluctuation curve and the equipment response rate data of each air-conditioning terminal within a preset time window before the start time; Based on the time constraint parameters, split the load adjustment target value into sub-load adjustment values corresponding to multiple consecutive time segments, and map each time segment to the corresponding fluctuation period in the energy consumption fluctuation curve; Overlay the sub-load adjustment values of each time segment with the equipment response rate data within the corresponding fluctuation period to generate a load adjustment vector sequence including time sensitivity; Perform dimension expansion and fusion on the load adjustment vector sequence and the real-time electricity price parameters in the dynamic regulation instruction. The dimension expansion and fusion includes splitting the real-time electricity price parameters into gradient change segments according to time segments and performing element-level weighted splicing with the positions of the same time index in the vector sequence to generate the joint input feature; Among them, the equipment response rate data is obtained by multiplying the average time from receiving the instruction to completing the parameter adjustment of the air-conditioning terminal in the historical scheduling record by the compressor operating frequency in the current operating state, and the weight coefficient of the weighted splicing is determined by the grid stability priority in the current operating mode of the virtual power plant; Among them, the execution process of the target scheduling policy includes: Analyze the air-conditioning terminal identifier marked as an immediately responsive device in the priority sequence, extract its corresponding start-stop time configuration and the adjustment range of the temperature setting threshold, and generate a device control instruction set; Sort the device control instruction set by timestamp and send it to the edge controller in the target area. The edge controller adjusts the start-stop time of the compressor according to the buffer interval parameter in the instruction to generate a driving signal set including the execution time sequence; Collect the actual operation parameters of each air-conditioning terminal after executing the driving signal set in real time, including the compressor state switching delay duration, the execution deviation of the temperature setting value, and the instantaneous energy consumption fluctuation amplitude, and generate a dynamic adjustment parameter set; Compare the dynamic adjustment parameter set with the parameter adjustment amplitude in the priority sequence, identify the abnormal device identifier with the execution deviation exceeding the allowable threshold, and generate a policy adjustment trigger event; According to the above strategy, adjust the trigger event to call the multi-objective optimization model to perform a local recalculation of the current priority sequence, generate an updated target scheduling strategy and reissue it for execution. At the same time, mark the feedback parameter set associated with the abnormal device identifier as a high-priority sample with incremental updates.
Citation Information
Patent Citations
Virtual power plant dispatching and control system and method based on central air conditioner
CN117249537A