Power dispatching system and method based on deep learning and user demand response

By using a deep learning-based power dispatching system, the problem of separating the net contribution of dispatching under external disturbances was solved, achieving accurate alignment of execution evidence and reliability of dispatching strategies, thereby improving the response efficiency and stability of power dispatching.

CN121543988BActive Publication Date: 2026-04-17湖南数界科技有限公司
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
湖南数界科技有限公司
Filing Date
2026-01-16
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

The existing power dispatching system has difficulty accurately separating the net dispatch contribution under external disturbance conditions, which leads to difficulty in aligning execution evidence, misjudging the profile of reinjected resources, and a tendency for dispatching strategies to be conservative, resulting in low response efficiency.

Method used

The power dispatching system based on deep learning and user demand response establishes a data source registry, constructs a deep learning demand response representation model, generates a controllable domain representation, constructs a set of instruction feasibility constraints, generates a set of dispatching action candidates, constructs a dispatching effect evaluator based on causal inference and counterfactual generation, and outputs a ternary result, thereby achieving the separation of execution contribution and external disturbance.

Benefits of technology

It improves the accuracy of scheduling effect evaluation, reduces the risk of misjudgment results being fed back into resource profiles, realizes the safety, controllability and closed-loop self-optimization of automated scheduling, and enhances the stable execution capability in multi-protocol and multi-object scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121543988B_ABST
    Figure CN121543988B_ABST
Patent Text Reader

Abstract

The application discloses a power dispatching system and method based on deep learning and user demand response, and relates to the technical field of power dispatching.The power dispatching system and method based on deep learning and user demand response comprises the following steps:S1, establishing a data source register and collecting a response data set, and outputting windowed dispatching data packets;S2, constructing a deep learning demand response representation model, and generating a dispatching action candidate set;S3, constructing a demand response dispatching effect evaluator, and generating a response net effect sequence;S4, constructing a global dispatching generation and cross-object collaborative arrangement mechanism, and constructing an online rollback mechanism.The application effectively improves the dispatching executability, verifiability and closed-loop self-optimization capability of controllable loads in a demand response scenario, and solves the problems that external disturbances are difficult to separate, execution evidence is difficult to align and is prone to misjudgment backfilling, the issuing link is difficult to trace and lacks rollback protection, and the dispatching tends to be conservative.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power dispatching technology, specifically to a power dispatching system and method based on deep learning and user demand response. Background Technology

[0002] With the rapid development of building energy management, IoT data acquisition, and automated control technologies, the load objects in commercial complexes are showing a trend of parallel access across multiple regions, systems, and protocols. Demand response is gradually shifting from "manual experience-based" to "platform-based, automated, and near real-time" closed-loop control. Existing technologies typically only assess response effectiveness based on the load difference before and after scheduling or simple regression baselines, lacking counterfactual baseline generation and causal identification mechanisms. This makes it difficult to accurately separate the net contribution of scheduling under disturbances such as sudden weather changes, fluctuations in customer flow, store opening and closing times, and changes in event plans, easily leading to confusion between "executed but masked by disturbances" and "unexecuted but naturally declining."

[0003] For example, invention patent CN114552658B discloses a new energy power system dispatching method that takes demand-side response into account. This method includes: considering users as dispatchable resources, establishing an optimal power system dispatching model with interruptible loads, aiming to minimize carbon emissions from thermal power units; considering wind farms as a new energy source participating in system dispatching, establishing an IGDT model considering the uncertainty of wind power output; using chance-constrained programming theory to maximize the fluctuation range of uncertain parameters, constructing a mathematical model based on probabilistic confidence gap decision theory; and using an AMPL solver to find the optimal solution for system operation. This invention achieves optimal dispatching of the power system under dual uncertainties of supply and demand, thereby ensuring the safe and stable operation of the power system.

[0004] For example, invention patent CN112234657B discloses a power system optimization scheduling method based on the combined output of new energy sources and demand response. This method includes: constructing a power dispatch optimization model based on a demand-side response model and a combined output model of multiple new energy sources; using the minimization of new energy output costs, demand-side response costs under different new energy scenarios, and rescheduling costs as the optimization objective function of the power dispatch optimization model; solving the optimization objective function to obtain a load reduction balance scheme for demand-side response under new energy scenarios, thereby controlling the access power for demand-side response. Considering the nonlinear correlation and demand response among various new energy sources in the power system, an optimized scheduling model based on new energy output fluctuations and demand-side response is obtained. An interface relationship between the nonlinear correlation of new energy sources and demand-side response is established to achieve a balance of load reduction in demand-side response under various new energy scenarios, thereby improving the level of new energy absorption.

[0005] In existing technologies, current systems only compare the load difference before and after scheduling, failing to isolate the impact of external disturbances. For example, non-scheduling factors such as a sudden large-scale event in a shopping mall leading to increased customer traffic or a sudden rise in temperature causing a natural increase in air conditioning load can mask the actual response effect. If, after scheduling to reduce load, the load actually increases due to the sudden temperature rise, the system may misjudge the resource as unreliable; conversely, if scheduling is not implemented, the load may naturally decrease due to the mall closing, leading the system to misjudge the response as effective. These misjudgments, when fed back into the resource profile, can cause subsequent demand-side response strategies for commercial complexes to become more conservative, reducing response efficiency.

[0006] Therefore, in order to address the above problems, there is an urgent need for a power dispatching system and method based on deep learning and user demand response. Summary of the Invention

[0007] Technical problems to be solved

[0008] To address the shortcomings of existing technologies, this invention provides a power dispatching system and method based on deep learning and user demand response, which solves the problems of difficulty in separating external disturbances, difficulty in aligning execution evidence leading to misjudgment and backfeeding, difficulty in tracing the distribution link, and lack of backoff guarantee, which leads to a conservative dispatching approach.

[0009] Technical solution

[0010] To achieve the above objectives, the present invention provides the following technical solution: a power dispatching system and method based on deep learning and user demand response, comprising: S1, establishing a data source registry and collecting response datasets, generating unique message identifiers, performing unified semantic encapsulation and constructing a response window index, and outputting windowed dispatching data packets; S2, constructing a deep learning demand response representation model based on the windowed dispatching data packets, generating a controllable domain representation, constructing a set of instruction feasibility constraints, and generating a candidate set of dispatching actions; S3, constructing a demand response dispatching effect evaluator based on a dual machine learning framework, generating a net response effect sequence, constructing an execution confidence judgment function, and outputting a ternary result; S4, constructing a global dispatching generation and cross-object collaborative orchestration mechanism based on the candidate set of dispatching actions, compiling the dispatching action sequence into a set of instruction transactions, and constructing an online rollback mechanism.

[0011] Furthermore, the specific process of establishing a data source registry and collecting response datasets to generate unique message identifiers is as follows: Data sources from both the field and external sources are uniformly registered to establish a data source registry; response datasets are collected, including load datasets, operating status datasets, control status datasets, and interference datasets; for each message in the acquisition link, the device's local time, the edge gateway's reception time, and the platform's entry time are simultaneously recorded, and a unique message identifier is generated; the original message archive is written to the original message archive repository and immutably saved using an append-only method, generating an immutable original message archive repository.

[0012] Furthermore, the specific process of performing unified semantic encapsulation and constructing a response window index to output windowed scheduling data packets is as follows: Input the original message archive repository and data source registry; perform structured parsing and unified semantic encapsulation on the original message to generate a standardized set of semantic data frames; retain the source pointer to the original message for each semantic data frame and write it into the resource global index key; construct a trusted time generation mechanism based on triple timestamps to output a unified analysis time; simultaneously perform quality verification on the semantic data frames to output a set of quality weights and anomaly tags; perform piecewise linear normalization on continuous numerical fields that pass the quality verification and perform standardized encapsulation; generate response event window identifiers based on trigger signals and object ranges, and construct a response window index table; output windowed scheduling data packets.

[0013] Furthermore, the specific process of constructing a deep learning demand response representation model based on windowed scheduling data packets to generate controllable domain representations is as follows: Inputting windowed scheduling data packets, constructing observation sequences and a deep learning demand response representation model; constructing a demand response representation network based on the observation sequences, learning unified latent space representations for air conditioning and lighting respectively; the demand response representation network adopts a dual-tower structure, using a temporal convolutional network to extract response features for air conditioning object sequences, and using a state event encoder and a load sequence encoder to extract response features for lighting object sequences, and introducing a cross-modal attention mechanism in the fusion layer to explicitly inject the impact of environmental and operational disturbances on load changes; constructing a multi-head prediction structure at the output of the representation network, outputting the controllable domain and uncertainty representations of the objects in the current scenario; writing the multi-head prediction outputs into a resource profile table to form a dynamic resource profile; outputting the dynamic resource profile and controllable domain representation results.

[0014] Further, the specific process of constructing the instruction feasibility constraint set and generating the scheduling action candidate set is as follows: Input the resource dynamic profile and controllable domain characterization results to construct the instruction feasibility constraint set; generate the scheduling action candidate set based on the instruction feasibility constraint set; the coupling penalty term is obtained by nonlinear coupling operation of the failure probability, delay probability, recovery rebound risk intensity and net effect uncertainty intensity in the resource dynamic profile; divide the margin characterization value by the scale parameter value and substitute it into the gate control function to obtain the feasibility pass rate; multiply all feasibility pass rates to obtain the gated product term; take the positive part of the net effect, multiply it by the net effect saturation coefficient, take the negative value and perform exponential operation, subtract the exponential operation from the fixed value to obtain the positive term, take the coupling penalty term, multiply it by the risk penalty coefficient, take the negative value and perform exponential operation to obtain the penalty term; multiply the gated product term, the positive term and the penalty term to obtain the nonlinear feasibility score value, and output the instruction feasibility constraint set and the scheduling action candidate set.

[0015] Furthermore, the specific process of constructing a demand response scheduling effect evaluator based on a dual machine learning framework and generating a net response effect sequence is as follows: Input windowed scheduling data packets and resource dynamic profiles to construct processing variables and execution evidence sufficiency indicators; perform executable segmentation on the main response window and construct a causal-aligned time axis; extract covariates from the windowed scheduling data packets and construct a temporal covariate matrix and scene label vector; construct outcome variables and perform counterfactual baseline modeling; use the dual machine learning framework to construct the demand response scheduling effect evaluator, train the counterfactual prediction model, and generate an unexecuted counterfactual baseline curve; during training... In this phase, a cross-fitting strategy is adopted for each object-level effective segment: the samples are divided into several mutually exclusive subsets, the processing model is trained on one subset to estimate the execution propensity probability, and the result model is trained on another subset to predict the load result under given covariates and execution states; the execution state is forcibly set to the non-execution state to obtain the non-execution counterfactual baseline curve, and the execution state is set to the actual execution state of the evidence chain determination to obtain the execution condition prediction curve. The net effect sequence is output based on the counterfactual baseline curve; the counterfactual baseline curve and the net effect sequence are written into the effect evaluation lineage index table, and the effect evaluation lineage index table is output.

[0016] Further, the specific process of constructing the execution confidence judgment function and outputting the ternary result is as follows: Input the effect assessment lineage index table, construct the perturbation strength vector and the execution confidence judgment function, and output the ternary result of net effect, perturbation strength, and execution confidence: Construct the execution confidence judgment function and output the execution confidence value: The evidence enhancement term is obtained by threshold shifting and scaling the execution evidence sufficiency score and then smoothing it; the penalty soft maximum aggregation term is obtained by soft maximum aggregation of the uncertainty penalty value and the conflict penalty value; the perturbation cancellation discrimination enhancement term is obtained by threshold comparison of the normalized ratio of the net effect to the perturbation and then smoothing it; Add the evidence enhancement term and the perturbation cancellation discrimination enhancement term, and substitute the sum into the control function to obtain the enhancement term; Subtract the penalty from the enhancement term. The soft maximum aggregation term is used to obtain the execution confidence value. When the execution confidence value is greater than or equal to the upper confidence threshold, a high execution confidence label is output. When the execution confidence value is greater than or equal to the lower confidence threshold and less than the upper confidence threshold, and the aggregated value of the disturbance intensity within the effective segment is greater than or equal to the disturbance intensity threshold, an execution but disturbance-canceled label is output, and the execution confidence level is marked as medium-high execution confidence. When the execution confidence value is less than the lower confidence threshold, and the operational natural decline judgment condition is met within the effective segment, an unexecuted but naturally declining label is output, and the execution confidence level is marked as low execution confidence. The ternary result of net effect, disturbance intensity, and execution confidence is written into the response effect evaluation table, and the execution confidence value is used as a gating factor to trigger a reliable update of the resource dynamic profile. The response effect evaluation table is output.

[0017] Furthermore, the specific process of constructing a global scheduling generation and cross-object collaborative orchestration mechanism based on the candidate set of scheduling actions is as follows: Input the set of instruction feasibility constraints and the candidate set of scheduling actions, the dynamic profile of resources and the representation results of the controllable domain, and the execution confidence gating parameters. Read the object index key, action type, action parameters, expected net effect range, response delay representation, recovery rebound risk and uncertainty representation from the candidate actions to construct a global scheduling generator. Perform objective constraint solving and cross-object collaborative orchestration on the candidate actions to generate action combinations and time orchestration sequences. The objective constraint solving aims to maximize the expected net effect and imposes penalties on comfort deviation, delay or failure risk, and recovery rebound. Key constraints include minimum duration, rate of change, ramp, and comfort level. The algorithm adapts to boundary conditions, partition availability, and recovery suppression. For each candidate action combination, a gain score is constructed. For each spatial region, the collaborative gain potential function of the action combination within that region is taken and substituted into the smoothing enhancement function to obtain the gain term. The gain terms corresponding to all spatial regions are summed to obtain the total collaborative gain. The uncertainty penalty potential function of the action combination within that region is taken and substituted into the smoothing enhancement function to obtain the penalty term. The penalty terms corresponding to all spatial regions are summed to obtain the total penalty. The total penalty is subtracted from the total collaborative gain to obtain the gain evaluation value. Consistency checks of the gain score values ​​are performed on the generated action combinations. A global scheduling action sequence is output, and the expected effective segment, expected duration, recovery orchestration information, and fallback priority are written for each action.

[0018] Furthermore, the specific process of compiling the scheduling action sequence into an instruction transaction set and constructing the online rollback mechanism is as follows: Input the global scheduling action sequence and the data source registry; read the control channel type, acknowledgment capability declaration, and readback field caliber corresponding to the object from the data source registry; construct the control contract and compile the scheduling action sequence into an instruction transaction set; collect proof-type acknowledgments on the edge execution side and perform execution consistency verification to form proof-type acknowledgment frames; construct an online rollback and degradation mechanism: when the communication link is unstable or acknowledgments are missing for a long time, link degradation is triggered, and the rollback reason and degradation level are written into the control transaction log to trigger the effect evaluation and reliable update closed loop of resource profile.

[0019] Furthermore, a second aspect of the present invention provides a power dispatching system based on deep learning and user demand response, applied to a power dispatching method based on deep learning and user demand response, comprising: a response data acquisition module, used to establish a data source registry and acquire response datasets, generate unique message identifiers, perform unified semantic encapsulation and construct a response window index, and output windowed dispatching data packets; a deep learning demand response representation module, used to construct a deep learning demand response representation model based on windowed dispatching data packets, generate controllable domain representations, construct a set of instruction feasibility constraints, and generate a set of dispatching action candidates; a dispatching effect evaluation module, used to construct a demand response dispatching effect evaluator based on a dual machine learning framework, generate a response net effect sequence, construct an execution confidence judgment function, and output a ternary result; and a self-optimization module, used to construct a global dispatching generation and cross-object collaborative orchestration mechanism based on the dispatching action candidate set, compile the dispatching action sequence into an instruction transaction set, and construct an online rollback mechanism.

[0020] Beneficial effects

[0021] The present invention has the following beneficial effects:

[0022] (1) This invention establishes a dynamic profile of object-level resources and a controllable domain representation model to continuously characterize adjustable boundaries, response delays, ramp-up capabilities, refusal to control tendencies, recovery rebound risks and uncertainties, providing a calculable basis for candidate action generation and risk constraints, and reducing instability caused by empirical thresholds.

[0023] (2) This invention, by constructing a set of instruction feasibility constraints and generating a set of scheduling action candidates, makes constraints such as comfort and indoor environment boundaries, equipment start-up and shutdown and minimum duration, control quantity change rate, lighting zone availability and recovery suppression explicit, so as to ensure that automated scheduling outputs feasible action combinations without triggering boundary risks.

[0024] (3) This invention constructs a scheduling effect evaluator that combines causal inference and counterfactual generation. It generates an unexecuted counterfactual baseline within the response window and calculates the net effect of the response. By combining disturbance intensity quantification and execution evidence chain, it separates the execution contribution from external disturbances, improves the accuracy of effect evaluation, and reduces the risk of misjudgment results being fed back into the resource profile.

[0025] (4) This invention adopts a control contract-driven instruction transaction compilation and idempotent delivery mechanism, combined with two-stage delivery, pre-verification limit and segmented delivery strategy, to achieve controllability, auditability and reproducibility of the automated delivery process, improve the stable execution capability in multi-protocol and multi-object scenarios, realize the safe and controllable scheduling execution and closed-loop self-optimization, and avoid long-term operation tending to be conservative.

[0026] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0027] Figure 1 This is a flowchart of the power dispatching method based on deep learning and user demand response of the present invention;

[0028] Figure 2 This is a framework diagram of the power dispatching system based on deep learning and user demand response of the present invention.

[0029] Figure 3 This is a multivariate trend comparison chart of the power dispatching system data of the present invention;

[0030] Figure 4 This is a net effect sequence aggregated view of the present invention;

[0031] Figure 5 This is a flowchart illustrating the closed-loop process of instruction transaction compilation, proof-type receipt verification, and online rollback / downgrade for this invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Please see Figures 1-5 This invention provides a technical solution: a power dispatching system and method based on deep learning and user demand response, comprising: S1, establishing a data source registry and collecting a response dataset, generating a unique message identifier, performing unified semantic encapsulation and constructing a response window index, and outputting a windowed dispatching data packet; S2, constructing a deep learning demand response representation model based on the windowed dispatching data packet, generating a controllable domain representation, constructing a set of instruction feasibility constraints, and generating a candidate set of dispatching actions; S3, constructing a demand response dispatching effect evaluator based on a dual machine learning framework, generating a net response effect sequence, constructing an execution confidence judgment function, and outputting a ternary result; S4, constructing a global dispatching generation and cross-object collaborative orchestration mechanism based on the candidate set of dispatching actions, compiling the dispatching action sequence into a set of instruction transactions, and constructing an online rollback mechanism.

[0034] Specifically, the process of establishing a data source registry and collecting response datasets to generate unique message identifiers is as follows: For the demand response scenario of air conditioning and lighting in commercial complexes, data sources on-site and external are uniformly registered to establish a data source registry; the data source registry records at least the data source identifier, data source type, access protocol and connection endpoint, collection method and scheduling strategy, field scope and unit dimension declaration, spatial area binding relationship, asset and circuit binding relationship, and parsing rule version number; the spatial area binding relationship is used to establish a stable mapping between public areas, shop zones, floor areas and power distribution circuits, the asset and circuit binding relationship is used to establish a stable mapping between chillers, fans, water pumps, terminal equipment, lighting zone controllers and metering points, and the parsing rule version number is used to identify the currently effective protocol parsing and field mapping rules of the data source, so as to facilitate replay and recalculation when the rules evolve.

[0035] The system collects response datasets, which include load datasets, operational status datasets, control status datasets, and disturbance datasets. It also collects electricity metering data from sub-meters and master meters deployed on the commercial complex's power distribution side, covering air conditioning main unit circuits, fan and water pump circuits, public area lighting circuits, and shop zone lighting circuits, to form a load dataset. Furthermore, it collects chiller start / stop status, supply and return water temperature setpoints, fan frequency, valve opening, operating mode, and fault alarm operational status data by connecting to the building automation system or air conditioning group control system, forming an operational status dataset. Finally, it collects lighting zone switch status, dimming command readback status, circuit enable status, and control failure reason codes by connecting to the lighting control gateway, forming a control status dataset. Finally, it collects outdoor temperature, humidity, irradiance, and wind speed meteorological data by connecting to meteorological services or on-site weather stations, combined with indoor environmental data collected from indoor temperature and humidity collection points, and collects store opening / closing status, activity plans, customer flow or occupancy proxy variables, and shop business zone information by connecting to the commercial complex's operation system, forming a disturbance dataset.

[0036] For each message in the acquisition link, the device's local time, the edge gateway's reception time, and the platform's entry time are recorded simultaneously, and a unique message identifier is generated. The unique message identifier is jointly constructed from the data source identifier, the source-side sequence or sampling sequence, the payload digest, and the link tag, and is used for idempotent deduplication and retransmission merging. The original message payload, protocol type, connection endpoint, data source identifier, parsing rule version number, space and asset binding reference, and triple timestamp metadata are written to the original message archive repository and stored immutably using an append-only write method, generating an immutable original message archive repository.

[0037] This implementation plan reduces event misalignment caused by out-of-order delivery and clock drift; avoids duplicate messages from contaminating profiles and evaluation results; and ensures that collected evidence is traceable, verifiable, and recalculated by appending and saving the original message immutable archive repository, providing a foundation for scheduling execution consistency verification, causal assessment lineage tracing, and system closed-loop self-optimization.

[0038] Specifically, the process of performing unified semantic encapsulation and constructing response window indexes to output windowed scheduling data packets is as follows: Input the original message archive and data source registry, read the protocol type, unit dimension declaration, ratio and phase declaration, spatial binding and asset binding relationship, and parsing rule version number of each data source from the data source registry, perform structured parsing and unified semantic encapsulation on the original message, and generate a standardized semantic data frame set: The standardized semantic data frame set includes at least: load metering semantic frames, air conditioning operation status semantic frames, lighting control status semantic frames, environmental disturbance semantic frames, and operational disturbance semantic frames.

[0039] For each semantic data frame, a source pointer pointing to the original message is retained and written into the resource global index key. The resource global index key is used to establish a consistent cross-vendor mapping between air conditioning equipment, lighting zones, metering loops, and spatial areas. The resource global index key is jointly constructed from asset binding information, spatial binding information, and unique equipment and loop numbers to ensure that the same physical object can be stably associated when reported from different data sources. A reliable time generation mechanism is built based on triple timestamps to output a unified analysis time: when the device's local time is available and drift is controllable, a drift generation mechanism is constructed based on the deviation sequence between the device time and the gateway's received time. The system performs image shifting and drift correction. When device time is missing or abnormal, the unified analysis time is estimated based on the gateway reception time and combined with the link delay profile. For records with only platform entry time, a conservative estimate is used and a low-confidence mark is added. The triple timestamp includes at least: the device local timestamp, written by the air conditioner controller, lighting controller, smart meter, or sub-metering terminal during sampling or reporting; the gateway reception timestamp, written by the edge gateway or protocol adapter when receiving the original message and completing verification before buffering; and the platform entry timestamp, written by the cloud access service when parsing and entering the database. The alignment priority of the triple timestamps is as follows: the device local timestamp is used first as the initial indication of the event occurrence time, and the gateway reception timestamp is used as the link arrival anchor point for drift correction. When the device local timestamp is missing, wrapped around, or judged as abnormal, the gateway reception timestamp is used as the benchmark for unified analysis time. When the gateway reception timestamp is unavailable or the record only comes from offline aggregation or batch import, the platform entry timestamp is used and marked as low-confidence. The determination method for whether the device's local timestamp is available and the drift is controllable is as follows: within a continuous statistical window, the difference sequence between the device's local timestamp and the gateway's received timestamp satisfies the monotonicity and upper bound constraints of fluctuation, and there is no wrap-around, jump, or long-term stagnation; when the difference sequence shows a sudden change, a reverse jump, or exceeds the drift limit, the device time is determined to be abnormal and a degradation alignment strategy is triggered.

[0040] Simultaneously, quality checks are performed on the semantic data frames, outputting quality weights and anomaly label sets. Quality checks include at least unit dimension consistency, ratio and phase consistency, state and power consistency, and item-to-summary consistency checks. Disturbances such as "store closure leading to natural decline" and "sudden weather changes leading to natural rise" are written into the disturbance labels, providing observable covariates and a basis for causal assessment. Quality weights are used to quantify the credibility of each semantic data frame. Quality weights are determined by the pass results and deviation levels of each check item: when any key check item triggers a hard anomaly label, the quality weight is lowered to a low weight range and prohibited from participating in profile updates; when only soft anomalies are triggered or deviations are slight, the quality weight is continuously decayed according to the deviation magnitude, and the quality weight and anomaly label are simultaneously written into the semantic data frame for sample weighting during training and evaluation. The anomaly tag set is a tag set that is bound one-to-one with the semantic data frame. It includes at least unit anomalies, phase ratio anomalies, state power conflict anomalies, sub-item summary inconsistency anomalies, timestamp anomalies, missing segment anomalies, and mutation anomalies. When multiple anomalies are triggered in the same frame, the anomaly tag set is merged and output according to priority, and the trigger verification item and deviation summary are retained for traceability.

[0041] For continuous numerical fields that have passed quality verification, piecewise linear normalization is performed for standardized encapsulation. For the active power, reactive power, and cumulative energy fields in the load metering semantic frame, the units are first made consistent based on the unit dimension declaration and ratio and conversion rules in the data source registry. Then, based on the quantile range or upper and lower bounds of the stable interval, piecewise linear normalization is used to generate normalized values, and the normalization method identifier and statistical window version number are written simultaneously. For the supply and return water temperature setpoints, fan frequency, and valve opening control fields in the air conditioning operation status semantic frame, a control quantity normalization mapping is constructed according to the equipment model and control range declaration, normalizing the differences in setpoint scales of different manufacturers into a unified control intensity representation. For the outdoor temperature, humidity, and irradiance fields in the environmental disturbance semantic frame, the corresponding normalization interval is selected according to the site climate interval and season label to avoid abrupt changes in feature distribution caused by seasonal changes. The normalized values ​​and original values ​​are written into the semantic data frame. Piecewise linear normalization is used to map features with different ranges and fluctuation ranges to a comparable scale. The segmentation points are adaptively determined by the historical statistical distribution: the low and high quantiles of the statistical features within the object-level historical window are used as the endpoints of the effective range, and truncation or saturation mapping is used outside the endpoints to achieve robust processing; linear mapping is used within the effective range, and the normalized values ​​are limited by a fixed saturation rule outside the endpoints to avoid stretching the scale by extreme outliers.

[0042] The unified analysis time is generated by three timestamps according to priority and credibility: When the device's local time is available and the difference sequence meets the drift controllability judgment, the unified analysis time is obtained by using the device's local time as the initial value and using the deviation profile between the device time and the gateway's received time for drift correction; when the device time is missing or is judged to be abnormal, the unified analysis time is obtained by using the gateway's received time as the benchmark and using the link delay profile for latency backtracking; when only the platform entry time is available, the unified analysis time is obtained by conservatively backtracking the platform entry time combined with the maximum link delay, and the credibility of the unified analysis time is marked as low credibility for gating and weighting.

[0043] Based on the trigger signal and object scope, a response event window identifier is generated. The control command semantic frames, execution receipt semantic frames, load metering semantic frames, and disturbance semantic frames for air conditioning and lighting are collected into the corresponding windows according to the resource global index key and the unified analysis time, and a response window index table is constructed. The response event window identifier is jointly constructed by the trigger source identifier, object scope reference, window boundary, and topology version number, which is used to strongly bind the entire process of triggering, evaluation, issuance, receipt, and effect. When the receipt has a delayed effect or partial execution, an effective sub-window reference is generated based on the effective evidence field in the receipt, and a parent-child relationship is established between the effective sub-window and the main window to support causal alignment time positioning. A windowed scheduling data packet is output.

[0044] This implementation scheme not only suppresses the contamination of profiles and models by abnormal data, but also explicitly writes clues about store closures and meteorological disturbances into labels to support counterfactual stripping. Furthermore, it uses piecewise linear normalization and coexistence encapsulation to balance cross-object comparability and interpretable backtracking. This provides verifiable, traceable, and replayable data support for controllable domain learning, action candidate generation, net effect evaluation, and execution confidence gating.

[0045] Specifically, the process of constructing a deep learning demand response representation model based on windowed scheduling data packets and generating controllable domain representations is as follows: Input the windowed scheduling data packets, read the air conditioning equipment index key and lighting zone index key from the window index, and perform time alignment and serialization on the load metering semantic frames, air conditioning operation status semantic frames, lighting control status semantic frames, environmental disturbance semantic frames, and operational disturbance semantic frames according to a unified analysis time to construct an observation sequence; group the observation sequence according to the object index key to form object-level samples, perform mask filling and credibility marking on missing segments within the same window according to quality weights to avoid replacing the true missing segments with interpolation results; embed and encode discrete state fields, use normalized values ​​as input for continuous fields and retain the original physical quantities for interpretable backtracking; construct scene label vectors for the store opening and closing status, activity plans, and customer flow proxy variables in operational disturbances, and use them together with meteorological disturbance vectors as exogenous driving inputs.

[0046] A deep learning-based demand response representation model is constructed: a demand response representation network is built based on observation sequences, and unified latent space representations are learned for air conditioning and lighting respectively. The demand response representation network adopts a dual-tower structure. Temporal convolutional networks are used to extract response features for air conditioning object sequences, and state event encoders and load sequence encoders are used to extract response features for lighting object sequences. A cross-modal attention mechanism is introduced in the fusion layer to explicitly inject the impact of environmental and operational disturbances on load changes. During training, quality weights are used as sample weight inputs to reduce the pollution of the representation space by noise and reconstructed data. Training objectives and supervision signals are derived from real execution data rather than simulated labels: control instruction semantic frames and execution receipt semantic frames in windowed scheduling data packets are used to determine processing variables and execution evidence. Net effect supervision signals are constructed using load metering semantic frames and counterfactual baseline generation results. For samples with delayed activation or partial execution, they are relabeled according to the effective sub-window based on the receipt activation evidence and power inflection point detection results to form real labels that can be used to supervise response delay, ramp-up capability, and recovery risk.

[0047] A multi-head prediction structure is constructed at the output end of the representation network to output the controllable domain and uncertainty representation of the object in the current scenario. The multi-head prediction includes at least: a load counterfactual baseline prior output head, used to predict the natural load trajectory when no scheduling is performed; an adjustable boundary output head, used to output the adjustable amplitude range and sustainability representation, and to provide monotonic mappings or piecewise response curves of control intensity and load response for air conditioning and lighting, respectively; a response dynamic output head, used to output response delay representation, ramp-up capability representation, and recovery risk representation; a failure risk output head, used to output the probability of refusal to control, the probability of delayed execution, and the probability of partial execution, and to output the attention weight of the risk source to indicate whether the risk increase is caused by equipment constraints, comfort constraints, or communication link constraints; and an uncertainty output head, used to output the confidence interval or quantile prediction results. The input feature list is organized on a unified time axis by object-level samples and includes at least: load metering features, control command features, execution evidence features, environmental disturbance features, and operational disturbance features, with missing segments appended with missing masks and quality weights as training gating inputs. To improve the realism of controllable domain learning, training samples and supervision signals are constructed. The control instruction semantic frames and execution receipt semantic frames in the windowed scheduling data packets are used to jointly determine the processing variables and execution evidence. Samples with delayed effectiveness or partial execution are relabeled according to the effective sub-window. The power inflection point detection results in the load metering semantic frames and the effectiveness evidence in the receipts are used to jointly determine the response starting point, which is used to supervise response delay and ramp-up characterization.

[0048] The training loss function employs a multi-task joint loss, including: counterfactual baseline fitting loss to constrain consistency with baseline generation results; net effect correlation loss to constrain the interpretability of predicted representations to the true net effect; classification loss for delayed, rejected, or partially executed actions to constrain consistency between risk output and relabeled feedback; and uncertainty calibration loss to constrain quantile coverage or interval coverage. Each sub-loss is weighted by quality weights, and the contribution of samples with insufficient execution confidence to the profile parameter update correlation loss is reduced.

[0049] The multi-head prediction output is written into the resource profile table to form a dynamic resource profile. The dynamic resource profile includes at least: a controllable domain parameter set, response efficiency score, response latency characteristics, recovery risk characteristics, failure risk characteristics, uncertainty representation, and the most recent update timestamp. Profile parameters are smoothly updated and version-managed over time. When a sudden change in profile parameters is detected and execution confidence is insufficient, a rollback to the previous version of the reliable profile is triggered to avoid a single abnormal window polluting the long-term profile. The dynamic resource profile and controllable domain representation results are output. Controllable domain parameters are given as adjustable amplitude ranges and sustainable time ranges in physical quantity form, and are simultaneously written into the corresponding normalized intervals for cross-object comparisons. Uncertainty representation is output as the prediction interval width or quantile bandwidth, and is bound to a unified analysis time, window identifier, and model version to support candidate action screening and gating judgment.

[0050] This implementation plan achieves integrated characterization; through smooth updates of resource profiles, version management, and low-confidence rollback mechanisms, it suppresses the pollution of long-term profiles by single abnormal windows, ultimately providing a stable, traceable, and comparable controllable domain representation foundation for instruction feasibility constraint construction, candidate action generation, and global collaborative orchestration.

[0051] Specifically, the process of constructing a set of instruction feasibility constraints and generating a candidate set of scheduling actions is as follows: Input the dynamic profile of the resource and the controllable domain representation results; read the adjustable boundary, response delay, ramp-up capability, recovery risk, and failure risk parameters from the dynamic profile of the resource; combine the object range in the window index with the current scene label to construct the set of instruction feasibility constraints; the set of instruction feasibility constraints includes at least: comfort and indoor environment boundary constraints, equipment start / stop and minimum duration constraints, control quantity change rate constraints, lighting zone service availability constraints, and recovery suppression constraints; for air conditioning objects, the controllable domain mapping is performed based on the supply and return water temperature setpoints, fan frequency, and valve opening control quantities; for lighting objects, based on the response curves and rejection probabilities of zone switching and dimming control, priority is given to action combinations that can maintain stable durations; advance time and secondary confirmation steps are added to the candidate set to avoid evaluation misalignment caused by uncertainty in the instruction's effective time.

[0052] A candidate set of scheduling actions is generated based on the set of instruction feasibility constraints. The candidate set of scheduling actions is output at the object index key granularity, including action type, action parameters, expected response start point, expected duration, recovery action and rollback action, and an execution risk score and uncertainty score are attached to each candidate action. Mutually exclusive and cooperative relationship constraints are introduced for candidate actions of air conditioning and lighting to avoid repeatedly applying mutually canceling actions to the same spatial area in the same time period. Fast simulation or incremental prediction based on counterfactual baseline is performed on candidate actions to estimate the impact of the actions on the load trajectory and output the expected net effect range.

[0053] For each candidate action, a nonlinear feasibility score is constructed. The coupling penalty term is obtained by nonlinearly coupling the failure probability, delay probability, recovery rebound risk intensity, and net effect uncertainty intensity from the resource dynamic profile. The margin characterization value is divided by the scale parameter value, and then substituted into the gate control function to obtain the feasibility pass rate. All feasibility pass rates are multiplied to obtain the gate control product term. The positive part of the net effect is multiplied by the net effect saturation coefficient, and the negative value is exponentially calculated. The fixed value is subtracted from the exponential calculation to obtain the positive term. The coupling penalty term is multiplied by the risk penalty coefficient, and the negative value is exponentially calculated to obtain the penalty term. The gate control product term, the positive term, and the penalty term are multiplied to obtain the nonlinear feasibility score value, which is used to collaboratively characterize the constraint feasibility gate, net effect benefit saturation, failure and delay risk penalty, and recovery rebound and uncertainty coupling penalty. The specific calculation formula for the nonlinear feasibility score value is as follows:

[0054] ;

[0055] In the formula, Indicates the first The nonlinear feasibility score of each candidate action is used to rank and prune the candidate actions. This represents the gate function, used to continuously map various constraint margins to the feasibility pass rate in the interval [0,1]. Indicates the candidate action in the 1st... The margin characterization value on the class constraint, which includes at least comfort, indoor environment boundary margin, equipment start-up and minimum duration margin, control change rate margin, lighting zone service availability margin and recovery inhibition margin, is calculated by substituting candidate action parameters into the corresponding constraint judgment model, and is used to quantify the remaining safety space of the action distance constraint boundary; The scale parameter value is obtained by statistically analyzing the distribution of class margins and combining it with the success and failure labels of actions. It is used to map different constraint margins to comparable scales. This represents the positive net effect, obtained by truncating the expected net effect of candidate actions, and is used to prevent actions with negative net effects from being mistakenly selected under the influence of other items. This represents the expected net effect, obtained through rapid simulation or incremental prediction based on a counterfactual baseline. It is used to drive the saturation growth of the benefit term, so that the benefit of the action decreases after reaching the margin and suppresses over-adjustment. This represents the risk penalty coefficient, which is obtained by optimizing the consistency between the ranking of candidate action scores and the actual execution results of success, delay, or rejection. The value range is a real number greater than zero, and it is used to adjust the suppression strength of risk items on the total score. The net effect saturation coefficient is obtained by fitting and cross-validating the relationship between net effect and over-action risk. It is a real number with a value greater than zero and is used to control the rate at which the benefit term increases as the net effect increases from small to large. The denotes the coupling penalty term, which is obtained by nonlinearly coupling the failure probability, delay probability, recovery rebound risk intensity, and net effect uncertainty intensity in the resource dynamic profile. It is used to multiplicatively suppress candidate actions when the risk increases and is accompanied by a rebound or increased uncertainty.

[0056] When any key constraint margin is negative, the gating term causes the score value to decay rapidly to prevent infeasible instructions from entering the automated issuance chain. Once the net effect reaches a certain level, the benefit term exhibits saturation growth to suppress excessive actions. When the risk of failure and delay increases, accompanied by increased rebound risk and uncertainty, the penalty term is coupled and amplified non-linearly to select executable and recoverable candidate actions. The output consists of a set of instruction feasibility constraints and a set of scheduling action candidates.

[0057] As shown in Table 1, the nonlinear feasibility scoring table for candidate actions is as follows: Candidate action ACTION-001, constraint type: comfort; margin characterization value: 8.2; scale parameter value: 10; coupling penalty term: 2.5; risk penalty coefficient: 0.3; nonlinear feasibility score: 0.0927. Candidate action ACTION-001, constraint type: indoor environmental boundary margin; margin characterization value: 6.5; scale parameter value: 8; coupling penalty term, risk penalty coefficient, and nonlinear feasibility score are consistent with other constraint types. Candidate action ACTION-002, constraint type: equipment start-stop margin; margin characterization value: 7.1; scale parameter value: 12; coupling penalty term: 3.1; risk penalty coefficient: 0.3; nonlinear feasibility score: 0.0546. Candidate action ACTION-002, constraint type: control quantity rate margin; margin characterization value: 6; scale parameter value: 9; coupling penalty term, risk penalty coefficient, and nonlinear feasibility score are consistent with other constraint types. Candidate action identifier ACTION-003, constraint type is lighting service availability margin: margin characterization value is 6.9, representing the remaining safety space of this action under the lighting service availability constraint; scale parameter value is 7, coupling penalty term is 2.2, risk penalty coefficient is 0.3; nonlinear feasibility score is 0.1243. Candidate action identifier ACTION-003, constraint type is recovery inhibition margin: margin characterization value is 7.2, scale parameter value is 8, coupling penalty term, risk penalty coefficient and nonlinear feasibility score are consistent with other constraint types.

[0058] Table 1. Nonlinear Feasibility Scoring Table for Candidate Actions

[0059]

[0060] like Figure 3The multivariate trend comparison chart of power dispatching system data shows that the power dispatching system compares the trends of key variables on the constraint feasibility side and the dispatching net effect side at the same time. The margin characterization curve reflects the comprehensive safety margin of objects such as air conditioning and lighting under constraints such as comfort and indoor environment boundaries, equipment start-up and shutdown and minimum duration, control quantity change rate, zonal service availability and recovery suppression. The curve fluctuations indicate that the adjustable space shrinks or expands with changes in the scenario. The scale parameter curve is used to scale-calibrate margin and risk indicators of different calibers and dimensions, ensuring stable input to the gating and scoring models under a unified discrimination scale, thereby avoiding discrimination oversensitivity or undersensitivity. The gating function result curve maps the "margin-scale" combination to the feasibility pass rate, used for candidate action admission, action combination arrangement, and... The system uses a continuous buffering approach to stabilize feasibility assessments and avoids hard switching due to short-term noise. The positive net effect curve corresponds to the positive net contribution obtained after removing external disturbances from the counterfactual baseline. It shows a phased plateau and step changes in some sections, indicating that the net effect is not a simple linear response to the margin, but is determined by the combination of actions, cross-object collaboration, and uncertainty, provided that the gating is allowed. When the net effect plateau moves up in the later stage, it usually means that the feasibility is stronger, the conflict is lower, and the prediction is more verifiable. If there is sufficient margin but no increase in net effect, it often points to problems such as combination conflict, recovery rebound, or high uncertainty. If there is a significant net effect but the gating is low for a long time, it suggests that the evidence chain is insufficient or the scale calibration and gating sensitivity need to be adjusted.

[0061] In this implementation plan, the candidate action set is converged into a set of actions that are more feasible, recoverable, and verifiable. It can quickly locate abnormal patterns with sufficient margin but poor net effect and significant net increase effect but low gating, thereby driving the adaptive tuning of gating sensitivity, scale parameters and risk penalty, and improving the overall stability, executability and long-term reliability of candidate generation.

[0062] Specifically, the process of constructing a demand response scheduling effect evaluator based on a dual machine learning framework and generating a net response effect sequence is as follows: Input windowed scheduling data packets and dynamic resource profiles, as well as effective sub-window references in the window index; extract the execution evidence chain from control instruction semantic frames and execution receipt semantic frames; construct processing variables and execution evidence sufficiency indicators; using the object index key as the granularity, use the receipt code, refusal control reason code, setpoint readback status, key remote signaling status transitions, power inflection point detection results, and time consistency as the execution evidence chain; define the execution status as a processing variable, which includes at least executed, refusal control, delayed execution, partial execution, and unknown states. Delayed execution is determined jointly by the effective time evidence in the receipt and the power inflection point time; partial execution is determined by the confirmation of a subset of actions and the consistency check that power changes only cover a portion of the object range; unknown states are marked as insufficient evidence and included in a low-weight sample set.

[0063] The minimum criteria for determining the chain of evidence are as follows: if the receipt code is confirmed and the setpoint readback is consistent, and a power inflection point or key remote signaling transition is detected within the effective sub-window to meet time consistency, then it is determined to be executed; if the receipt code is rejected or the rejection reason code is valid, and the setpoint readback is inconsistent, then it is determined to be rejected; if the receipt provides an effective time and the power inflection point time has a consistent lag relative to the issuance time and falls within the effective sub-window, then it is determined to be delayed execution; if only some objects in the same window have confirmed receipts or only some objects have inflection points and the coverage after spatial binding aggregation is insufficient, then it is determined to be partially executed; if the above evidence items are missing or conflict with each other, then it is determined to be an unknown state and the sample weight is reduced. The minimum statistical caliber for the execution evidence sufficiency index is: the evidence sufficiency score is obtained by summarizing the evidence sub-item hit and time consistency residuals, and the four types of sub-items—receipt consistency, readback consistency, remote signaling consistency, and inflection point consistency—are counted separately in the evidence sub-item count.

[0064] The system performs executable segmentation on the response main window to construct a causal alignment timeline. For each object index key, the causal alignment start point is preferentially used based on the receipt effective time. When the receipt only confirms receipt but not effectiveness, the remote signaling transition time is used as the effectiveness proxy. When evidence is still insufficient, the response inflection point obtained from the change point detection in the load metering semantic frame is used as the effectiveness proxy, and the evidence sufficiency score is reduced. The segmented effective fragments are written into the object-level effective fragment table to remove ineffective segments from the execution effect estimation, avoiding miscalculation of external disturbances as execution effects. The minimum statistical scope of the causal alignment timeline is: outputting a triplet of effective start point, effective end point, and effective duration for each object, and recording the effective start point source type as one of three categories: receipt effectiveness, remote signaling transition, or inflection point proxy. When using inflection point proxy, a proxy confidence flag is output simultaneously, and the evidence sufficiency score is reduced for gating and weighting.

[0065] Covariates are extracted from windowed scheduling data packets to construct a temporal covariate matrix and scene label vector. The temporal covariate matrix includes at least: outdoor temperature and humidity, irradiance and trend characteristics in the environmental disturbance semantic frame, and indoor temperature and humidity and deviation characteristics; store opening and closing status, activity plan status, customer flow or occupancy proxy variables and their change characteristics in the operational disturbance semantic frame; historical load pattern characteristics, component proportion characteristics, and load fluctuation characteristics in the load metering semantic frame; operating mode, start and stop status, set point, and valve and fan execution status in the air conditioning operation status semantic frame; and zoned scenes, switch and dimming status, and availability status in the lighting control status semantic frame. Covariates are aligned according to a unified analysis time and missing masks are retained. Reconstruction tags are written to the reconstructed segments and quality weights are introduced as sample weights to avoid systematic bias caused by retransmission and out-of-order repair. The minimum statistical scope of the result variables is: object-level power sequence or spatial region aggregated power sequence as the result variables, and the power increment sequence relative to the window start point can be output to weaken static differences. The aggregation relationship between object level and region level is generated through a spatial binding relationship table and written into the lineage index.

[0066] Construct outcome variables and perform counterfactual baseline modeling; the outcome variables preferably adopt the load sequence after aggregation at the object level or region, or the load increment sequence relative to the window start point can be used to reduce static differences; output object level and spatial region level for air conditioning and lighting respectively, and perform traceable aggregation of object level and region level through spatial binding relationship.

[0067] A dual machine learning framework is employed to construct a demand response scheduling effect evaluator, train a counterfactual prediction model, and generate an unexecuted counterfactual baseline curve. During the training phase, a cross-fitting strategy is used for each object-level effective segment: the samples are divided into several mutually exclusive subsets. A processing model is trained on one subset to estimate the execution propensity probability, while an outcome model is trained on another subset to predict the load outcome under given covariates and execution states. The processing model preferably uses a temporal attention network of a sequence classification network to output the propensity probability, and extreme propensity probabilities are truncated and smoothed to control variance. The outcome model preferably uses a temporal convolutional network or a self-attention encoder to encode the temporal covariate matrix and scene label vector, and then fuses them with the execution state to output a conditional load prediction sequence, while also outputting quantiles or interval widths as uncertainty representations. During the evaluation phase, the execution state is forcibly set to an unexecuted state to obtain the unexecuted counterfactual baseline curve, and the execution state is set to the actual execution state determined by the chain of evidence to obtain the execution conditional prediction curve, thus achieving the isolation of external disturbances under the same covariate conditions.

[0068] The minimum statistical approach for crossfitting is as follows: after grouping the samples by window identifier or object index key, divide them into mutually exclusive folds, train the processing model and the result model in each fold, and then generate the propensity probability and conditional prediction sequence to avoid the same sample participating in both training and evaluation, which would lead to bias; at the same time, truncate samples with propensity probabilities that are too close to zero or one and record the truncation mark to control variance and trigger a review.

[0069] The net effect sequence is output based on the counterfactual baseline curve. For each object index key, the response net effect value is calculated within the effective segment to obtain the net effect sequence, which is then aggregated by object, spatial region, and complex dimensions to form object-level net effect, region-level net effect, and complex-level net effect, respectively. For partially executed objects, the net effect is calculated segment by effective segment and then summarized within a window to avoid treating partial execution as the whole invalid. For delayed execution samples, the net effect of the segment before effective execution is fixed at zero and marked with a delay label to avoid including natural fluctuations before effective execution in the execution contribution. The counterfactual baseline curve and net effect sequence are written into the effect evaluation lineage index table. The lineage index includes at least: window identifier, object index key, covariate snapshot summary, processing model version, result model version, propensity probability interval, evidence sufficiency score, and effective segment boundary, to ensure the verification and traceability of the evaluation results, and the effect evaluation lineage index table is output. The error metric between the unexecuted counterfactual baseline curve and the executed condition prediction curve includes at least the following: calculating the baseline prediction error on the unexecuted sample set or high-confidence unexecuted segments, and calculating the executed condition prediction error on high-confidence executed segments; the error metric outputs at least the absolute error aggregation and peak-valley deviation aggregation, and writes the error summary into the lineage index to measure the stability and repeatability of the model in different scenarios.

[0070] like Figure 4The aggregated view of the net effect sequence shows an overlay of the "counterfactual baseline load without scheduling" (dashed line) and the "actual observed load" (solid line), with the net effect sequence formed by the difference between the two displayed on the right-hand coordinate system, allowing for simultaneous observation of the corresponding relationships within the same view. Multiple vertical marker lines in the figure align and annotate key moments, including the command issuance time, the start and end of partial execution, ensuring that the start point, duration, and decay / recovery segments of the net effect are strictly aligned temporally with the execution evidence chain. Within the command activation window, the net effect is aggregated using bar and fill methods, distinguishing between the "effective execution" net effect intervals and those "insufficient execution or offset by disturbances." Net effect range: When the actual load curve deviates stably, continuously, and in line with the effective window from the counterfactual baseline, the view marks the net effect segment as a valid contribution to support the reliable update of the resource profile; when the net effect fluctuates significantly within the effective window, has insufficient duration, or is more consistent with some execution intervals, it is marked as an inadequate execution / offset segment, indicating that the profile update weight should be reduced or the window should be put into review; at the window level, the net effect is upgraded from a "numerical sequence" to a "traceable evidence alignment result", which can explain why the net effect is significant in some periods and why the net effect is not significant in some periods, thus providing a visual basis for execution confidence judgment and closed-loop optimization.

[0071] This implementation plan removes disturbances and obtains the net effect sequence without affecting the pricing strategy; it ensures that the evaluation results are traceable, verifiable, and comparable. This allows for a stable distinction between effective contribution segments and segments with insufficient execution or disturbance offsetting, thus providing a reliable basis for performance confidence determination, reliable resource profile updates, and subsequent self-optimization of scheduling strategies.

[0072] Specifically, the process of constructing the execution confidence judgment function and outputting the ternary result is as follows: input the effect evaluation lineage index table, construct the disturbance intensity vector and the execution confidence judgment function, and output the ternary result of net effect, disturbance intensity and execution confidence: the disturbance intensity vector includes at least: meteorological disturbance intensity, constructed from the changing trends, abrupt changes and consistency of changes of outdoor temperature, humidity and irradiance; operational disturbance intensity, constructed from store opening and closing switching, activity plan state transitions, and abrupt changes of passenger flow and occupancy proxy variables; internal operating condition disturbance intensity, constructed from indoor temperature and humidity deviations, abrupt changes in terminal operating state and frequency of lighting scene switching; normalize and encapsulate the disturbance intensity vector and retain the disturbance type label.

[0073] The execution confidence function is constructed and the execution confidence value is output: the evidence enhancement term is obtained by threshold shifting and scaling the execution evidence sufficiency score, followed by smoothing enhancement; the penalty soft maximum aggregation term is obtained by soft maximum aggregation of the uncertainty penalty value and the conflict penalty value; the perturbation cancellation discrimination enhancement term is obtained by threshold comparison of the normalized ratio of the net effect to the perturbation, followed by smoothing enhancement; the evidence enhancement term and the perturbation cancellation discrimination enhancement term are added together, and the sum is substituted into the control function to obtain the enhancement term; the execution confidence value is obtained by subtracting the penalty soft maximum aggregation term from the enhancement term; the specific calculation formula for the execution confidence value is as follows:

[0074] ;

[0075] In the formula, This represents the execution confidence value, used to output high, medium, and low execution confidence and as a gating factor for updating the resource profile; The evidence enhancement item is obtained by performing threshold shifting and scaling normalization on the sufficiency of evidence score, followed by smoothing enhancement. It is used to continuously map the difference between "insufficient evidence" and "sufficient evidence" into confidence increments. This represents the penalty soft maximum aggregation term, which is obtained by soft maximum aggregation of uncertainty penalty value and conflict penalty value. It is used to achieve a nonlinear suppression mechanism that can significantly reduce confidence by increasing any penalty term. The perturbation cancellation discrimination enhancement term is obtained by thresholding the normalized ratio of the net effect to the perturbation and then smoothing it. It is used to characterize whether the net effect still has an interpretable contribution under strong perturbation background. The execution evidence sufficiency score is obtained by consistency verification and gating fusion of evidence sub-items such as the consistency of receipt code, the credibility of the reason code for refusal of control, the consistency of set point readback, the consistency of key remote signaling transition, and the consistency of power inflection point and effective time. It is used to quantify the strength of executed evidence. The threshold for evidence is determined by statistical analysis of evidence scores and adaptive selection based on manual sampling labels. It is used to define the boundary between evidence meeting the standard and evidence being insufficient. The evidence scale parameter is determined through cross-validation that yields the best consistency between the confidence output and the labels of execution, non-execution, or delayed execution. It is a real number with a value greater than zero and is used to control the sensitivity of evidence items to the confidence level. It represents the positive part of the net effect sequence, obtained by the difference between the executed conditional prediction curve and the unexecuted counterfactual baseline curve, and is used to avoid falsely raising confidence levels for negative net effects or reverse fluctuations under strong disturbances; The time series representing disturbance intensity is constructed by fusing meteorological disturbance intensity, operational disturbance intensity and internal operating condition disturbance intensity and then normalizing and encapsulating it. It is used to characterize the potential driving intensity of non-scheduling factors on load changes within a window. The effective segment indication sequence is generated by combining the effective time evidence, setpoint readback consistency evidence, and key remote signaling transition evidence in the execution receipt semantic frame. It is used to aggregate the net effect and disturbance intensity only within the effective segment. The stability factor is determined adaptively by the lower quantile of the statistical disturbance intensity or the minimum effective disturbance scale. Its value range is a real number greater than zero, which is used to avoid the ratio being unstable due to extremely small disturbance intensity. The baseline threshold representing the net effect relative to the perturbation is adaptively determined by statistically analyzing the ratio distribution of samples that have been executed but whose perturbations have been offset. This threshold is used to support the identification of offset labels and suppress misjudgments caused by natural variations. The uncertainty penalty value is obtained by statistical analysis of quantile bandwidth, confidence interval width, or prediction residual fluctuation. It is used to characterize the credibility of counterfactual estimation and adjust the strength of uncertainty suppression of confidence. The conflict penalty value is represented by the residual of the evidence pointing to the net effect direction and the temporal consistency, as well as the probability snapshot and the evidence chain conflict penalty. It is used to characterize the situation where the evidence and the net effect contradict each other.

[0076] When the execution confidence value is greater than or equal to the upper confidence threshold, a high execution confidence label is output; when the execution confidence value is greater than or equal to the lower confidence threshold but less than the upper confidence threshold, and the aggregated value of the disturbance intensity within the effective segment is greater than or equal to the disturbance intensity threshold, an execution but disturbance-canceled label is output, and the execution confidence level is marked as medium-high execution confidence; when the execution confidence value is less than the lower confidence threshold, and the operational natural decline judgment condition is met within the effective segment, an unexecuted but naturally declining label is output, and the execution confidence level is marked as low execution confidence; the operational natural decline judgment condition is: the store opening / closing status indicator within the effective segment is in the closed state, or the activity plan status transitions from the in progress state to the ended state, and the aggregated value of the operational natural decline indicator is greater than or equal to the operational natural decline threshold; the operational natural decline indicator is jointly constructed and normalized by the store closing status indicator, the activity end indicator, and the decline magnitude of the customer flow or occupancy proxy variable. When the acknowledgment indicates delayed execution and there is a consistent lag between the net effect inflection point and the subsequent delivery, a delayed execution label is output, and segmented confidence is output in units of effective segments. The tendency probability snapshot is used as a reference for execution predictability. When the tendency probability is extreme and conflicts with the evidence chain, the conflict degree is increased to avoid over-inference on extreme samples. The ternary result of net effect, disturbance intensity, and execution confidence is written into the response effect evaluation table, and the execution confidence value is used as a gating factor to trigger the reliable update of the resource dynamic profile. When the execution confidence value is insufficient, it is only recorded without updating or a conservative update strategy is adopted, and the window is included in the sample pool to be reviewed. For samples with high execution confidence labels, time smoothing update and version management are adopted. When the profile parameters change abruptly and the disturbance intensity is abnormal, a rollback to the previous version of the reliable profile is triggered to avoid a single abnormal window polluting the long-term profile. At the same time, the execution confidence value and disturbance type statistics are written into the profile lineage field to provide a basis for risk parameters in the generation of scheduling candidates. The response effect evaluation table is output.

[0077] In this implementation scheme, the update logic has definite boundaries rather than vague descriptions; the ternary results are written into the response effect evaluation table and the execution confidence is used as the profile update gating to avoid the external disturbances being mistakenly fed back as unreliable resources; the profile drift and the strategy tending to be conservative caused by misjudgment feedback are suppressed, forming a reliable closed loop of sustainable self-optimization.

[0078] Specifically, the process of constructing a global scheduling generation and cross-object collaborative orchestration mechanism based on the candidate set of scheduling actions is as follows: Input the instruction feasibility constraint set and the candidate set of scheduling actions, the resource dynamic profile and controllable domain representation results, and the execution confidence gating parameters. Read the object index key, action type, action parameters, expected net effect range, response delay representation, recovery rebound risk, and uncertainty representation from the candidate actions to construct a global scheduling generator. Perform target constraint solving and cross-object collaborative orchestration on the candidate actions to generate action combinations and time orchestration sequences. The global scheduling generator uses spatial region binding as the orchestration unit, grouping air conditioning objects and lighting objects within the same spatial region into collaborative clusters. Based on the relationships within the collaborative clusters... The controllable domain mapping and recovery suppression constraints establish a primary-secondary orchestration rule: first generate the air conditioning action baseline and lock the expected effective segment and duration, then generate lighting action fine-tuning in the remaining controllable domain, or orchestrate in reverse order according to business priority in business scenarios where lighting is the primary focus; introduce mutual exclusion and coupling constraints between different collaborative clusters to avoid applying mutually canceling actions to adjacent spatial areas at the same time and suppress cross disturbances caused by personnel flow; during the action selection process, sort the nonlinear feasibility scores of candidate actions and truncate them in combination with execution confidence gating. When the execution confidence value corresponding to a candidate action is lower than the gating threshold, reduce the priority of entering the combination or mark it as a candidate to be reviewed.

[0079] The objective constraint solution is defined as follows: within the scheduling window, the goal is to maximize the expected net effect aggregation value, while penalizing indoor environmental boundary approach, execution delay and failure risk, and recovery rebound risk, and satisfying hard constraints of minimum duration, rate of change, ramp-up capability, partition business availability, and recovery suppression; the collaborative gain potential function is defined as follows: robustly aggregating the execution confidence and expected net effect positive part of actions within the collaborative cluster and deducting recovery rebound risk to obtain the realizable contribution strength within the region; the penalty potential function is defined as follows: aggregating the directional conflict degree and uncertainty strength of the net effect of actions within the collaborative cluster to obtain the in-region cancellation and unprovable strength; the smoothing enhancement function is defined as a logarithmic exponential smoothing mapping, used to maintain continuous score changes to stabilize ranking and pruning when gain increases and penalties increase.

[0080] The objective constraint solution aims to maximize the expected net effect and imposes penalties on comfort deviation, delay or failure risk, and recovery rebound. Key constraints include minimum duration, rate of change, ramp-up, comfort boundary, zonal availability, and recovery inhibition. A gain score is constructed for each candidate action combination. For each spatial region, the cooperative gain potential function of the action combination within that region is taken and substituted into the smoothing enhancement function to obtain the gain term. The gain terms corresponding to all spatial regions are summed to obtain the cooperative gain sum. The uncertainty penalty potential function of the action combination within that region is taken and substituted into the smoothing enhancement function to obtain the penalty term. The penalty terms corresponding to all spatial regions are summed to obtain the penalty sum. The gain evaluation value is obtained by subtracting the penalty sum from the cooperative gain sum, which is used to screen and rank action combinations. The specific calculation formula for the gain score is as follows:

[0081] ;

[0082] In the formula, Indicates a combination of actions The gain score is used to optimize candidate combinations during the objective constraint solution process; It represents a set of spatial regions, obtained by reading the spatial region binding relationship in the data source registry and the object index key-spatial region identifier mapping, and is used to determine the orchestration granularity and cooperative cluster division boundary when generating the global schedule; It represents a spatial region or collaborative cluster, obtained by merging air conditioning and lighting objects bound to the same spatial region, and is used to implement cross-object collaborative orchestration at the region level; The cooperative gain potential function of the action combination within the spatial region is obtained by extracting the execution confidence, positive part of expected net effect and recovery rebound risk factors from the subset of actions falling into the spatial region and performing robust aggregation and nonlinear transformation. It is used to drive the global scheduling generator to prioritize action combinations with more realizable gains within the region. It represents the uncertainty penalty potential function of action combinations within a spatial region. It is obtained by aggregating the conflict of net effect directions between actions within the region and the uncertainty intensity of the prediction of net effect of actions, and then performing a nonlinear transformation. It is used in the goal constraint solution to automatically reduce action combinations with obvious offsetting risks or high uncertainties and trigger action replacement and rearrangement. This represents a smoothing enhancement function that rapidly increases the combined score when feasibility and synergistic gains are satisfied, and rapidly decreases it when offsetting and uncertainty increase.

[0083] The system performs consistency checks on the gain score of the generated action combinations; it checks whether the rate of change and minimum duration constraints of the air conditioning action combinations are simultaneously satisfied on all objects, and introduces an upper bound constraint on the recovery intensity during the recovery phase: when the predicted recovery intensity exceeds the upper bound, segmented recovery actions are automatically inserted to smooth the recovery; it checks the availability constraints of the lighting action combinations for the zoning services: the baseline state is forcibly retained for critical and emergency zones, and a batch activation strategy is adopted for adjustable zones to reduce visual mutations and the probability of rejection; it checks the non-cancellation constraint in the same area for cross-object combinations: when the air conditioning and lighting actions cause the overall net effect to cancel each other out in the same spatial area, the actions are automatically rearranged or replaced with low-conflict candidates; it outputs a global scheduling action sequence and writes the expected effective segment, expected duration, recovery orchestration information and rollback priority for each action.

[0084] This implementation plan reduces evaluation misalignment and profile contamination caused by low-quality actions entering the combination; it solidifies the consistency of change rate and minimum duration across all objects, retention and batch activation of key lighting zones, non-cancellation rearrangement within the same zone, and the upper bound of recovery and replenishment intensity and segmented recovery insertion mechanism into factory checks, ensuring that the generated global scheduling action sequence can be stably activated, smoothly restored, and has a clear rollback priority, thereby improving the stability, cross-object collaboration, and long-term reliability of automated deployment.

[0085] Specifically, the process of compiling the scheduling action sequence into a set of instruction transactions and constructing the online rollback mechanism is as follows: Figure 5 The flowchart shown illustrates the closed-loop process of instruction transaction compilation, proof-type receipt verification, and online rollback / degradation. It inputs a global scheduling action sequence and a data source registry. From the data source registry, it reads the control channel type, receipt capability declaration, and readback field specifications corresponding to the object. A control contract is constructed, and the scheduling action sequence is compiled into an instruction transaction set. The control contract includes at least action semantics, preconditions, expected state transitions, timeout policies, rollback actions, and recovery actions. For each instruction transaction, an instruction transaction identifier and idempotent key are generated, along with the object index key, action type, action parameters, expected effective segment, expected duration, and recovery action. Actions and rollbacks are written into the instruction transaction header; a two-stage delivery strategy is adopted for lighting zones: first, readback confirmation of availability and current scene status is performed, and actions are issued only after the readback meets the preconditions, and the consistency of the readback is confirmed; a precondition verification and limit issuance strategy is adopted for air conditioning objects: verify the equipment protection interval, operating mode and alarm status, generate setpoint increments according to the controllable domain limit and issue them, and use segmented issuance to meet the change rate constraint; idempotent delivery and retry strategy is enabled for all instruction transactions. When duplicate delivery is detected, it is merged by idempotent key and only the effective result is retained to avoid fluctuations caused by duplicate actions.

[0086] On the edge execution side, evidence-type receipts are collected and execution consistency is checked to form evidence-type receipt frames. The edge actuator collects the device confirmation code, control rejection reason code, setpoint readback value, key remote signaling transition and power inflection point detection results for each instruction transaction, and writes them into the evidence-type receipt frame according to a unified analysis time. Consistency checks are performed on the receipt evidence: when the device confirms but the setpoint readback is inconsistent, it is marked as suspected ineffective; when the setpoint readback is consistent but the power inflection point is missing, it is marked as suspected disturbance cancellation or unobservable effect; when the power inflection point exists but is consistent with the effective segment, it is marked as delayed execution. The consistency check results are written back to the platform and bound to the window identifier and the effective sub-window reference. Establish an online rollback and degradation mechanism: When the probability of execution failure increases or the reason codes for rejection of control appear in a concentrated manner, trigger a rollback: terminate subsequent undelivered instruction transactions and issue a rollback action to restore the object to the previous version's safety setpoint or preset baseline state; when indoor environmental indicators approach the boundary, trigger a protective rollback: prioritize rollback of the setpoint increment for air conditioning objects and extend the verification waiting time, and degrade lighting objects to execute only on non-critical partitions; when the communication link is unstable or the receipt is missing for a long time, trigger link degradation, switch to conservative actions and increase the strength of pre-confirmation, and if necessary, suspend automated issuance and enter a pending review state; write the rollback reason, degradation level and recovery completion mark to the control transaction log to trigger the closed loop of effect evaluation and reliable update of resource profile.

[0087] This implementation plan achieves deduplication and merging in cross-retry and duplicate delivery scenarios, avoiding load fluctuations and evaluation misalignments caused by repeated actions; it ensures that the system can still maintain safe, stable and interpretable execution in uncertain, disturbed or abnormal link scenarios, and provides a reliable closed-loop basis for effect evaluation and trustworthy updates of resource profiles.

[0088] Specifically, the second aspect of the present invention provides a power dispatching system based on deep learning and user demand response, applied to a power dispatching method based on deep learning and user demand response, comprising: a response data acquisition module, used to establish a data source registry and acquire response datasets, generate unique message identifiers, perform unified semantic encapsulation and construct a response window index, and output windowed dispatching data packets; the semantic encapsulation records triple timestamps and generates a unified analysis time; and the semantic frames are subjected to quality verification and output quality weights and anomaly labels. The deep learning demand response representation module is used to build a deep learning demand response representation model based on windowed scheduling data packets. It uses real instructions and receipts as supervision sources to generate controllable domain representations, construct instruction feasibility constraint sets, and generate a candidate set of scheduling actions. The scheduling effect evaluation module is used to build a demand response scheduling effect evaluator based on a dual machine learning framework. It constructs an execution evidence chain and evidence sufficiency score to form a causal aligned time axis. It generates a net response effect sequence, constructs an execution confidence judgment function, and outputs a ternary result. The self-optimization module is used to build a global scheduling generation and cross-object collaborative orchestration mechanism based on the candidate set of scheduling actions. It uses spatial regions as collaborative clusters and performs primary-to-secondary orchestration, combined with execution confidence gating pruning. It compiles the scheduling action sequence into an instruction transaction set and constructs an online rollback mechanism.

[0089] This implementation plan enhances the executability and cross-object collaboration of action combinations. It can also trigger online rollback and degradation when there are link anomalies, uncontrolled concentrations, or boundary approach, ultimately achieving a closed-loop scheduling effect that is executable, verifiable, rollback-able, and self-optimizable, thereby improving scheduling stability, long-term reliability, and automation implementation capabilities.

[0090] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0091] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A power dispatching method based on deep learning and user demand response, characterized in that, Includes the following steps: S1, establish a data source registry and collect response datasets, generate unique message identifiers, perform unified semantic encapsulation and build response window indexes, and output windowed scheduling data packets; S2, Based on windowed scheduling data packets, construct a deep learning demand response representation model, generate a controllable domain representation, construct an instruction feasibility constraint set, and generate a candidate set of scheduling actions; S3, based on a dual machine learning framework, constructs a demand response scheduling effect evaluator, generates a net response effect sequence, constructs an execution confidence judgment function, and outputs a ternary result; S4, based on the candidate set of scheduling actions, construct a global scheduling generation and cross-object collaborative orchestration mechanism, compile the sequence of scheduling actions into a set of instruction transactions, and construct an online rollback mechanism; The specific process of establishing a data source registry, collecting response datasets, and generating unique message identifiers is as follows: Unify the registration of on-site and external data sources and establish a data source registry; Collect response datasets, which include: load datasets, operating status datasets, control status datasets, and interference datasets; record the device local time, edge gateway reception time, and platform entry time for each message in the acquisition link, and generate a unique message identifier; write the original message archive and save it immutably using an append write method, generating an immutable original message archive. The specific process of constructing a deep learning demand response representation model based on windowed data packet scheduling and generating a controllable domain representation is as follows: Input windowed scheduling data packets to construct observation sequences and a deep learning demand response representation model: Based on the observation sequences, a demand response representation network is constructed, and unified latent space representations are learned for air conditioning and lighting respectively. The demand response representation network adopts a dual-tower structure. For the air conditioning object sequence, a temporal convolutional network is used to extract response features, and for the lighting object sequence, a state event encoder and a load sequence encoder are used to extract response features. A cross-modal attention mechanism is introduced in the fusion layer to explicitly inject the impact of environmental disturbances and operational disturbances on load changes. A multi-head prediction structure is constructed at the output of the representation network to output the controllable domain and uncertainty representation of the object in the current scenario. The multi-head prediction output is written into a resource profile table to form a dynamic resource profile. The dynamic resource profile and controllable domain representation results are output. The specific process of constructing the instruction feasibility constraint set and generating the scheduling action candidate set is as follows: Input the dynamic profile of the input resources and the characterization results of the controllable domain to construct a set of instruction feasibility constraints; generate a candidate set of scheduling actions based on the set of instruction feasibility constraints; the coupling penalty term is obtained by nonlinear coupling operation of the failure probability, delay probability, recovery rebound risk intensity and net effect uncertainty intensity in the dynamic profile of the resources; divide the margin characterization value by the scale parameter value and substitute it into the gate control function to obtain the feasibility pass rate; multiply all feasibility pass rates to obtain the gated product term; take the positive part of the net effect, multiply it by the net effect saturation coefficient, take the negative value and perform exponential operation, subtract the exponential operation from the fixed value to obtain the positive term, take the coupling penalty term, multiply it by the risk penalty coefficient, take the negative value and perform exponential operation to obtain the penalty term; multiply the gated product term, the positive term and the penalty term to obtain the nonlinear feasibility score value, and output the set of instruction feasibility constraints and the candidate set of scheduling actions; The specific process of constructing a demand response scheduling effect evaluator based on a dual machine learning framework and generating a net response effect sequence is as follows: Input windowed scheduling data packets and resource dynamic profiles to construct processing variables and execution evidence sufficiency indicators; use executable segmentation of the response main window to construct a causal-aligned time axis; extract covariates from the windowed scheduling data packets and construct a temporal covariate matrix and scene label vector; construct outcome variables and perform counterfactual baseline modeling; A dual machine learning framework is used to construct a demand response scheduling effect evaluator, train a counterfactual prediction model, and generate an unexecuted counterfactual baseline curve. During the training phase, a cross-fitting strategy is adopted for each object-level effective segment: the samples are divided into several mutually exclusive subsets. On one subset, a processing model is trained to estimate the execution propensity probability, and on another subset, an outcome model is trained to predict the load outcome under given covariates and execution states. The execution state is forcibly set to an unexecuted state to obtain the unexecuted counterfactual baseline curve, and the execution state is set to the actual execution state determined by the chain of evidence to obtain the execution condition prediction curve. The net effect sequence is output based on the counterfactual baseline curve. The counterfactual baseline curve and the net effect sequence are written into the effect evaluation lineage index table, and the effect evaluation lineage index table is output. The specific process of constructing and executing the confidence judgment function and outputting the ternary result is as follows: Input the lineage index table for effect evaluation, construct the perturbation intensity vector and the execution confidence judgment function, and output the ternary result of net effect, perturbation intensity, and execution confidence: Construct the execution confidence judgment function and output the execution confidence value: The evidence enhancement term is obtained by threshold shifting and scaling the execution evidence sufficiency score and then smoothing it; the penalty soft maximum aggregation term is obtained by soft maximum aggregation of the uncertainty penalty value and the conflict penalty value; the perturbation cancellation discrimination enhancement term is obtained by threshold comparison of the normalized ratio of net effect to perturbation and then smoothing it; Add the evidence enhancement term and the perturbation cancellation discrimination enhancement term, and substitute the sum into the control function to obtain the enhancement term; Subtract the penalty soft maximum aggregation term from the enhancement term to obtain the execution confidence value; When the execution confidence value is greater than or equal to the upper confidence threshold, output a high execution confidence label; when the execution confidence value is greater than or equal to the lower confidence threshold but less than the upper confidence threshold, and the aggregated value of the disturbance intensity within the effective segment is greater than or equal to the disturbance intensity threshold, output an executed but disturbance-canceled label, and mark the execution confidence level as medium-high execution confidence; when the execution confidence value is less than the lower confidence threshold, and the operational natural decline judgment condition is met within the effective segment, output an unexecuted but naturally declining label, and mark the execution confidence level as low execution confidence; write the ternary result of net effect, disturbance intensity, and execution confidence into the response effect evaluation table, and use the execution confidence value as a gating factor to trigger a reliable update of the resource dynamic profile; output the response effect evaluation table.

2. The power dispatching method based on deep learning and user demand response according to claim 1, characterized in that: The specific process of performing unified semantic encapsulation, constructing a response window index, and outputting windowed scheduling data packets is as follows: Input the original message archive repository and data source registry, perform structured parsing and unified semantic encapsulation on the original message, and generate a standardized set of semantic data frames: retain the source pointer to the original message for each semantic data frame and write it into the resource global index key, build a reliable time generation mechanism based on triple timestamp, output a unified analysis time, and perform quality verification on the semantic data frames to output a set of quality weights and anomaly tags. Perform piecewise linear normalization on continuous numerical fields that have passed quality verification, and then encapsulate them in a standardized manner; Generate response event window identifiers based on trigger signals and object scope, and construct a response window index table; Output windowed scheduling data packets.

3. The power dispatching method based on deep learning and user demand response according to claim 1, characterized in that: The specific process of constructing a global scheduling generation and cross-object collaborative orchestration mechanism based on the scheduling action candidate set is as follows: Input the set of feasibility constraints and candidate sets of scheduling actions, the dynamic profile of resources and the representation results of controllable domains, as well as the execution confidence gating parameters. Read the object index key, action type, action parameters, expected net effect range, response delay representation, recovery rebound risk and uncertainty representation from the candidate actions to construct a global scheduler. The program performs objective constraint solving and cross-object collaborative orchestration on candidate actions to generate action combinations and time-based orchestration sequences. The objective constraint solving aims to maximize the expected net effect and imposes penalties on comfort deviation, delay or failure risk, and recovery rebound. Key constraints include minimum duration, rate of change, ramp, comfort boundary, zonal availability, and recovery inhibition. For each candidate action combination, a gain score is constructed. For each spatial region, the cooperative gain potential function of the action combination within the region is taken and substituted into the smoothing enhancement function to obtain the gain term. The gain terms corresponding to all spatial regions are added together to obtain the cooperative gain sum. The uncertainty penalty potential function of the action combination within the region is taken and substituted into the smoothing enhancement function to obtain the penalty term. The penalty terms corresponding to all spatial regions are added together to obtain the penalty sum. Subtract the total penalty from the total collaborative gain to obtain the gain evaluation value, and perform a consistency check on the gain score value for the generated action combination; Output the global scheduling action sequence, and write the expected effective segment, expected duration, recovery orchestration information and rollback priority for each action.

4. The power dispatching method based on deep learning and user demand response according to claim 1, characterized in that: The specific process of compiling the scheduling action sequence into an instruction transaction set and constructing the online rollback mechanism is as follows: Input the global scheduling action sequence and the data source registry. Read the control channel type, acknowledgment capability declaration, and readback field caliber corresponding to the object from the data source registry. Construct a control contract and compile the scheduling action sequence into an instruction transaction set. Collect proof-type acknowledgments on the edge execution side and perform execution consistency verification to form proof-type acknowledgment frames. Construct an online rollback and degradation mechanism: when the communication link is unstable or acknowledgments are missing for a long time, trigger link degradation and write the rollback reason and degradation level into the control transaction log to trigger the effect evaluation and reliable update closed loop of resource profile.

5. A power dispatching system based on deep learning and user demand response, employing the power dispatching method based on deep learning and user demand response as described in any one of claims 1-4, characterized in that, include: The response data acquisition module is used to establish a data source registry and collect response datasets, generate unique message identifiers, perform unified semantic encapsulation and build a response window index, and output windowed scheduling data packets. The deep learning demand response representation module is used to build a deep learning demand response representation model based on windowed scheduling data packets, generate a controllable domain representation, construct a set of instruction feasibility constraints, and generate a set of scheduling action candidates. The scheduling effect evaluation module is used to build a demand response scheduling effect evaluator based on a dual machine learning framework, generate a net response effect sequence, construct an execution confidence judgment function, and output a ternary result. The self-optimization module is used to build a global scheduling generation and cross-object collaborative orchestration mechanism based on the candidate set of scheduling actions, compile the sequence of scheduling actions into a set of instruction transactions, and build an online rollback mechanism.

Citation Information

Patent Citations

  • Power System Optimization Dispatch Method Based on Combined Output of New Energy Sources and Demand Response

    CN112234657B

  • A dispatching method for renewable energy power system considering demand-side response

    CN114552658B

  • Demand response type virtual power plant aggregation scheduling method and system

    CN116523199A

  • Integrated energy power dispatching system and method thereof

    CN116976627A