Method and system for optimizing heating and power consumption based on neural network
Patent Information
- Application Number
- CN202610974029.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
针对现有技术的不足,本发明提供了基于神经网络的暖通能耗优化调度方法及系统,解决了现有暖通能耗优化调度中神经网络输出动作缺乏设备侧可执行性校核,导致调度指令易出现动作透支、反馈空转和联锁阻断且难以稳定落地的问题
(1)基于神经网络的暖通能耗优化调度方法及系统,通过将相邻控制指令编号对应的调度指令字段转换为调度动作令牌,并基于执行反馈数据计算令牌透支量和回补后令牌透支量,使神经网络输出的设定值变化转化为反映设备承接能力的动作单元,从而提高暖通能耗优化调度指令与真实设备执行状态之间的一致性。
Smart Images

Figure CN122813341A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy consumption scheduling technology, specifically to a method and system for optimizing HVAC energy consumption scheduling based on neural networks. Background Technology
[0002] With the development of building automation systems, building energy consumption monitoring platforms, and artificial intelligence scheduling technologies, the operation and control of HVAC systems are gradually shifting from fixed-time start / stop, manual parameter setting, and single-point threshold adjustment to continuous scheduling based on multi-source operating data, load forecasting results, and intelligent optimization models. In large public buildings, commercial complexes, and industrial parks, there are significant energy consumption coupling relationships among chillers, chilled water pumps, cooling water pumps, cooling tower fans, air handling units, and terminal valves. Changes in the setpoints of a single device often propagate to other devices through the water system, air system, and indoor thermal environment. Therefore, how to utilize neural networks, deep reinforcement learning, and building energy management models for collaborative optimization of HVAC systems has become an important technological direction in the field of building energy conservation control.
[0003] For example, application CN120278848B provides an energy storage management system for building energy conservation. The system includes a data acquisition module, a data processing module, a load forecasting module, an energy storage scheduling module, and an intelligent scheduling control module. In the load forecasting module, the system constructs a symmetrical sliding window fluctuation perception CNN model to improve the timeliness and sensitivity of heat load forecasting. In the energy storage scheduling module, the system combines opponent disturbance modeling, a two-way adversarial update mechanism, and subgradient cumulative correction technology to improve the Actor-Critic model, forming an adversarial dual-track Actor-Critic model, which enhances the robustness and adaptability of phase change energy storage scheduling strategy in complex dynamic environments.
[0004] For example, application CN120975528A discloses a microgrid-HVAC coordinated optimization method based on deep reinforcement learning, which relates to the field of building energy system management and scheduling. The method includes the following steps: 1. Constructing a refined model of the HVAC system and microgrid system considering population comfort and building energy costs; 2. For the refined model constructed in step 1, using a two-layer optimization control strategy to construct a microgrid energy management system framework integrating the HVAC system; 3. Based on step 2, constructing a multi-energy coupling model of the microgrid-HVAC system, transforming the time-series optimization problem into a Markov Decision Process (MDP) and performing mathematical representation; 4. Training the agent using an improved Priority Experience Replay Deep Deterministic Policy Gradient Algorithm (PER-DDPG).
[0005] However, existing HVAC energy consumption optimization and scheduling technologies focus more on load forecast accuracy, energy consumption objective function construction, and scheduling strategy optimization. They typically rely on neural networks or reinforcement learning models to directly output temperature setpoints, chilled water supply temperature setpoints, pump frequency setpoints, cooling tower fan frequency setpoints, or chiller start / stop commands. They lack a process-oriented assessment of whether the scheduling actions can actually be implemented by valves, pumps, chillers, and fans on the real equipment side. Especially when there are short-term load fluctuations or the model continuously outputs setpoints with large amplitudes, the equipment side may be affected by valve execution lag, pump frequency ramp-up, compressor loading valve response limitations, chiller interlock protection, and cooling tower fan switching constraints. This can lead to low-energy-consumption actions generated by the model being feasible in simulation but exhibiting overdraft, feedback idling, reverse cancellation, and interlock blocking in real HVAC systems.
[0006] Therefore, in order to address the above problems, there is an urgent need for a method and system for optimizing HVAC energy consumption scheduling based on neural networks. Summary of the Invention
[0007] Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method and system for optimizing HVAC energy consumption scheduling based on neural networks. This solves the problem that existing HVAC energy consumption optimization scheduling methods lack equipment-side executable verification of neural network output actions, which leads to scheduling commands being prone to action overdraft, feedback idling, and interlocking blockage, and are difficult to implement stably.
[0008] Technical solution To achieve the above objectives, the present invention provides the following technical solution: a neural network-based HVAC energy consumption optimization scheduling method, comprising: S1, collecting HVAC energy consumption scheduling monitoring data, performing time alignment, abnormal record removal, and missing data compensation on the HVAC energy consumption scheduling monitoring data, and outputting preprocessed HVAC energy consumption scheduling monitoring data; S2, based on the preprocessed HVAC energy consumption scheduling monitoring data, converting the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens, and generating token overdraft, token replenishment relationship marker value, and replenished token based on the execution feedback data corresponding to the scheduling action tokens. S3, Based on the scheduling action token, token overdraft amount, token overdraft amount after replenishment, and token replenishment relationship marker value, generate an overdraft embedding vector, and combine the overdraft embedding vector to construct a neural network action overdraft token graph, generating prohibited token paths and replenishable token paths; S4, Based on the current HVAC state sequence and the neural network action overdraft token graph, generate candidate action token groups, perform prohibited path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token group, convert the target scheduling action token group into HVAC energy consumption optimization scheduling instructions and issue them, and output the HVAC energy consumption optimization scheduling results.
[0009] Further, the specific steps for collecting HVAC energy consumption scheduling and monitoring data, including execution time alignment, abnormal record removal, and missing data compensation, are as follows: Collect HVAC energy consumption scheduling and monitoring data, which includes the air conditioning system number, equipment number, control command number, data acquisition timestamp, outdoor dry-bulb temperature, indoor dry-bulb temperature, supply air temperature, chilled water supply temperature, chiller unit status flag value, compressor load rate value, compressor load valve opening value, chilled water pump frequency value, cooling water pump frequency value, cooling tower fan frequency value, terminal valve opening value, air conditioning unit fan frequency value, chiller unit power value, chilled water pump power value, cooling water pump power value, and cooling tower fan... The data includes: power output, indoor temperature setpoint for the dispatch command area, chilled water supply temperature setpoint, chilled water pump frequency setpoint, cooling water pump frequency setpoint, cooling tower fan frequency setpoint, chiller unit start / stop flag, dispatch command issuance timestamp, interlock protection trigger flag, and control command execution result flag. Using the building automation system's main controller timestamp as the primary time axis, a nearest neighbor timestamp matching algorithm is employed to perform time alignment processing on the HVAC energy consumption dispatch monitoring data. A local outlier factor algorithm is used to identify and remove abnormal records in the HVAC energy consumption dispatch monitoring data, combined with piecewise linear interpolation for missing data compensation, resulting in the output of preprocessed HVAC energy consumption dispatch monitoring data.
[0010] Furthermore, based on the preprocessed HVAC energy consumption scheduling and monitoring data, the specific steps for converting the scheduling instruction fields corresponding to adjacent control instruction numbers into scheduling action tokens are as follows: Read the preprocessed HVAC energy consumption scheduling and monitoring data, and extract the following values from adjacent control instructions according to the air conditioning system number and control instruction number: indoor temperature setpoint of the scheduling instruction area, chilled water supply temperature setpoint of the scheduling instruction, chilled water pump frequency setpoint of the scheduling instruction, cooling water pump frequency setpoint of the scheduling instruction, cooling tower fan frequency setpoint of the scheduling instruction, and chiller unit start / stop flag value of the scheduling instruction area; calculate the difference between the same scheduling instruction fields in the subsequent control instruction number and the previous control instruction number to obtain the area... The system calculates the action difference for the zone's indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. For each action difference, the corresponding scheduling instruction field, action direction marker value, and action amplitude value are written into a scheduling action token to generate a scheduling action token group. The scheduling instruction field includes the zone's indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. The action direction marker value is determined by the positive or negative direction of the action difference or the direction of start / stop status change. The action amplitude value is determined by the absolute value of the action difference or the result of the start / stop status change.
[0011] Further, the specific steps for generating the token overdraft, token replenishment relationship flag value, and replenished token overdraft based on the execution feedback data corresponding to the scheduling action token are as follows: Match the execution feedback data according to the scheduling instruction fields of the scheduling action token, where the regional indoor temperature corresponds to the terminal valve opening value and the air conditioning unit fan frequency value, the chilled water supply temperature corresponds to the compressor load rate value and the compressor load valve opening value, the chilled water pump frequency corresponds to the chilled water pump frequency value, the cooling water pump frequency corresponds to the cooling water pump frequency value, the cooling tower fan frequency corresponds to the cooling tower fan frequency value, and the chiller unit start / stop corresponds to the chiller unit status flag value; read the matched execution feedback data from the N consecutive sampling points after the timestamp of the scheduling instruction issuance. According to the algorithm, the maximum change amplitude formed by the execution feedback data along the corresponding action direction is determined as the actual acceptance amount of the token; the actual acceptance amount of the token is obtained by subtracting the action amplitude value in the scheduling action token; when the token overdraft amount is greater than zero, the corresponding scheduling action token is marked as an overdraft token; when the token overdraft amount is equal to zero, the corresponding scheduling action token is marked as an acceptance token; when the actual acceptance amount of the token is greater than zero and the matched target response data remains unchanged within N consecutive sampling points, the corresponding scheduling action token is marked as a feedback idle token; among these, the target response data is matched according to the scheduling instruction field recorded in the scheduling action token, and the indoor temperature setpoint in the scheduling instruction area is matched with the indoor dry-bulb temperature value. The dispatch command matches the chilled water supply temperature setpoint with the chilled water supply temperature value; the dispatch command matches the chilled water pump frequency setpoint, the dispatch command matches the cooling water pump frequency setpoint, and the dispatch command matches the cooling tower fan frequency setpoint with the supply air temperature value and the indoor air dry-bulb temperature value; the dispatch command matches the chiller unit start / stop flag value with the chilled water supply temperature value; it reads the overdraft token, token overdraft amount, and action direction flag value from the previous control command number and the execution feedback data from the next control command number; when the change direction of the execution feedback data corresponding to the same dispatch command field in the next control command number is consistent with the action direction flag value of the overdraft token, it generates a same-equipment compensation relationship and subtracts the token overdraft amount from the corresponding execution feedback data. The token overdraft amount is obtained by subtracting the change magnitude of the target response data from the token overdraft amount. When the execution feedback data corresponding to different scheduling instruction fields in the subsequent control instruction number changes, and causes the corresponding target response data to change, a cross-device overdraft relationship is generated, and the token overdraft amount is subtracted from the change magnitude of the corresponding target response data to obtain the token overdraft amount after replenishment. When the change direction of the execution feedback data corresponding to the same scheduling instruction field in the subsequent control instruction number is opposite to the action direction flag value of the overdraft token, a reverse cancellation relationship is generated, and the reverse change magnitude of the corresponding execution feedback data is added to the token overdraft amount to obtain the token overdraft amount after replenishment. The same-device overdraft relationship, cross-device overdraft relationship, and reverse cancellation relationship are written into the token overdraft relationship flag value.
[0012] Further, the specific steps for generating the overdraft embedding vector based on the scheduling action token, token overdraft amount, token overdraft amount after replenishment, and token replenishment relationship marker value are as follows: The scheduling action tokens corresponding to M consecutive control command numbers under the same air conditioning system number, ending with the current control command number, are arranged into a token state sequence according to the timestamp of the scheduling command issuance; wherein, the scheduling action tokens under each control command number are arranged in the order of regional indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop; when the actual number of tokens in the token state sequence is less than 6M, an invalid padding token is written at the beginning of the sequence; the action direction is... The marker values, overdraft token marker values, feedback idle token marker values, token replenishment relationship marker values, interlocking protection trigger marker values, and control command execution result marker values are converted into discrete codes. The action amplitude value, token overdraft amount, and replenished token overdraft amount are converted into standardized values to form a token state input vector. The token state input vector is input into the overdraft embedding neural network, which is then time-encoded through a gated recurrent unit network to output the overdraft embedding vector corresponding to each scheduling action token. During the training of the overdraft embedding neural network, the error between the replenished overdraft prediction value and the replenished token overdraft amount, and the error between the forbidden path prediction marker value and the forbidden token path marker are used as training losses.
[0013] Furthermore, the specific steps for constructing a neural network action overdraft token graph by combining overdraft embedding vectors and generating prohibited token paths and recoverable token paths are as follows: Using scheduling action tokens as graph nodes, token recovery relationship marker values as graph edges, and overdraft embedding vectors as graph node features, construct a neural network action overdraft token graph; when the interlocking protection trigger marker value is in a triggered state or the control command execution result marker value is in a failed state, write the corresponding scheduling action token into the blocking node; starting from the graph node corresponding to each scheduling action token, search along the graph edges, write the graph path where the overdraft amount of the recovered token is greater than the overdraft tolerance threshold and reaches the blocking node into the prohibited token path, and write the graph path where the overdraft amount of the recovered token is less than or equal to the overdraft tolerance threshold and does not reach the blocking node into the recoverable token path, outputting the neural network action overdraft token graph data.
[0014] Furthermore, the specific steps for generating candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token graph are as follows: Extract preprocessed HVAC energy consumption scheduling monitoring data from the current sampling point according to the air conditioning system number, and assemble the current HVAC state sequence according to the data acquisition timestamp; input the current HVAC state sequence into a gated recurrent unit network to obtain the current state latent vector; concatenate the current state latent vector with the overdraft embedding vector and input it into the graph attention layer. The graph attention layer resets the attention weights of the graph nodes corresponding to the forbidden token paths to zero, and weights the current state latent vector according to the attention weights of the graph nodes corresponding to the recoverable token paths. Update and generate an overdraft mask state vector; input the overdraft mask state vector into a multilayer perceptron to output candidate action token sets; map the candidate action token sets to the neural network action overdraft token graph; when the graph path corresponding to the candidate action token set has overlapping nodes or overlapping edges with the forbidden token path, search along the graph edges for the replenishable token path that has the same action direction label value as the candidate action token set and has the smallest token overdraft amount after replenishment, and replace the candidate action token set with the corresponding replenishable action token set; when the graph path corresponding to the candidate action token set does not have overlapping nodes or overlapping edges with the forbidden token path, retain the candidate action token set.
[0015] Further, the specific steps for performing forbidden path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token groups are as follows: The retained candidate action token groups and the replacement replenishable action token groups are combined with the current HVAC state sequence to form a prediction input vector, and the prediction input vector is input into the multilayer perceptron to output the system's predicted total power value corresponding to each action token group; the replenished token overdraft amount, feedback idle token flag value, and blocking node connection result corresponding to each action token group are read, and action token groups that meet any of the following conditions are eliminated: feedback idle token flag value, connection to a blocking node, or replenished token overdraft amount exceeding the overdraft tolerance threshold; if there are remaining action token groups after elimination, the action token group with the smallest system predicted total power value is selected as the target scheduling action token group; if there are no remaining action token groups after elimination, the scheduling instruction field, action direction flag value, and action amplitude value corresponding to the current control instruction number are written into the holding action token group, and the holding action token group is used as the target scheduling action token group.
[0016] Further, the specific steps for converting the target scheduling action token group into HVAC energy consumption optimization scheduling instructions and issuing them, and outputting the HVAC energy consumption optimization scheduling results are as follows: Generate HVAC energy consumption optimization scheduling instructions according to the scheduling instruction field, action direction marker value, and action amplitude value in the target scheduling action token group; for the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, and cooling tower fan frequency, according to the action direction marker value, match the scheduling instruction setting value in the current control instruction number that is the same as the scheduling instruction field of the target scheduling action token with the corresponding action... The amplitude values are added or subtracted to obtain the corresponding target setpoint. For chiller unit start-up and shutdown, the start-up and shutdown flag values of the scheduling command corresponding to the current control command number are updated according to the action direction flag value to obtain the target chiller unit start-up and shutdown flag value. The target setpoints and target chiller unit start-up and shutdown flag values are combined to generate HVAC energy consumption optimization scheduling commands. New control command numbers are assigned to the HVAC energy consumption optimization scheduling commands and sent to the main controller of the building automation system. The system receives and writes back the execution feedback data bound to the new control command number and outputs the HVAC energy consumption optimization scheduling results.
[0017] The second aspect of this invention provides a neural network-based HVAC energy consumption optimization scheduling system, comprising: an HVAC data alignment processing module, an action overdraft token generation module, a neural overdraft map construction module, and a mask scheduling instruction output module, wherein: the HVAC data alignment processing module is used to collect HVAC energy consumption scheduling monitoring data, perform execution time alignment, abnormal record removal, and missing data compensation processing on the HVAC energy consumption scheduling monitoring data, and output preprocessed HVAC energy consumption scheduling monitoring data; the action overdraft token generation module is used to convert the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens based on the preprocessed HVAC energy consumption scheduling monitoring data, and generate token overdrafts based on the execution feedback data corresponding to the scheduling action tokens. The system includes: a token overdraft mapping module, a neural network action overdraft token mapping module, a mask scheduling instruction output module, and a mask scheduling instruction output module. The mask scheduling instruction output module generates candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token mapping module. It performs forbidden path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token group. The target scheduling action token group is then converted into an HVAC energy consumption optimization scheduling instruction and issued, outputting the HVAC energy consumption optimization scheduling result.
[0018] Beneficial effects The present invention has the following beneficial effects: (1) A method and system for optimizing HVAC energy consumption scheduling based on neural networks converts the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens, and calculates the token overdraft amount and the token overdraft amount after replenishment based on the execution feedback data, so that the change of the set value output by the neural network is transformed into an action unit reflecting the equipment's capacity, thereby improving the consistency between the HVAC energy consumption optimization scheduling instructions and the actual equipment execution status.
[0019] (2) A method and system for optimizing HVAC energy consumption based on neural networks, by constructing same-equipment replenishment relationship, cross-equipment replenishment relationship and reverse offset relationship, writes the delay acceptance, equipment collaborative compensation and reverse weakening process of the scheduling action in the subsequent control cycle into the token replenishment relationship flag value, which can distinguish between replenishable actions and feedback idle actions, and improve the accuracy of identifying execution lag and coupled response state.
[0020] (3) A method and system for optimizing HVAC energy consumption scheduling based on neural networks. By inputting the token state sequence into the overdraft embedding neural network and using the overdraft embedding vector as the graph node feature to construct the neural network action overdraft token graph, the scheduling action, action overdraft, replenishment relationship and blocking state can be uniformly mapped into the graph structure, thereby identifying prohibited token paths and replenishable token paths.
[0021] (4) A method and system for optimizing HVAC energy consumption scheduling based on neural networks, by performing prohibited path replacement and power sorting on candidate action token groups, firstly eliminate or replace action token groups with feedback idling, blocked connection or overdraft without replenishment, and then select the target scheduling action token group with the lowest predicted total power value from the executable candidate set, so that the scheduling result takes into account low energy consumption, equipment safety boundary and command landing stability. Attached Figure Description
[0022] Figure 1 The flowchart shows a neural network-based HVAC energy consumption optimization scheduling method. Figure 2 The diagram shows the structure of a neural network-based HVAC energy consumption optimization scheduling system. Figure 3 A schematic diagram of constructing a neural network action overdraft token graph; Figure 4 Replace the executable domain and forbidden path with the action overdraft token. Detailed Implementation
[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] Please see Figures 1-4 This invention provides a technical solution: a neural network-based HVAC energy consumption optimization scheduling method, comprising: S1, collecting HVAC energy consumption scheduling monitoring data, performing time alignment, abnormal record removal and missing data compensation processing on the HVAC energy consumption scheduling monitoring data, and outputting preprocessed HVAC energy consumption scheduling monitoring data; S2, based on the preprocessed HVAC energy consumption scheduling monitoring data, converting the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens, and generating token overdraft, token replenishment relationship marker value and replenished token overdraft based on the execution feedback data corresponding to the scheduling action tokens; S3: Generate an overdraft embedding vector based on the scheduling action token, token overdraft amount, token overdraft amount after replenishment, and token replenishment relationship marker value. Combine the overdraft embedding vector to construct a neural network action overdraft token graph, generating prohibited token paths and replenishable token paths. S4: Generate candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token graph. Perform prohibited path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token group. Convert the target scheduling action token group into HVAC energy consumption optimization scheduling instructions and issue them, outputting the HVAC energy consumption optimization scheduling results.
[0025] Specifically, the steps for collecting HVAC energy consumption scheduling and monitoring data, including time alignment, abnormal record removal, and missing data compensation, are as follows: Collecting HVAC energy consumption scheduling and monitoring data includes the air conditioning system number, equipment number, control command number, data acquisition timestamp, outdoor dry-bulb temperature, indoor dry-bulb temperature, supply air temperature, chilled water supply temperature, chiller unit status flag value, compressor load rate value, compressor load valve opening value, chilled water pump frequency value, cooling water pump frequency value, cooling tower fan frequency value, terminal valve opening value, and air conditioning unit fan frequency. The following parameters are included: power rating, chiller unit power rating, chilled water pump power rating, cooling water pump power rating, cooling tower fan power rating, indoor temperature setpoint for the dispatch command area, chilled water supply temperature setpoint for the dispatch command, chilled water pump frequency setpoint for the dispatch command, cooling water pump frequency setpoint for the dispatch command, cooling tower fan frequency setpoint for the dispatch command, chiller unit start / stop flag value for the dispatch command, timestamp of dispatch command issuance, interlock protection trigger flag value, and control command execution result flag value. The air conditioning system number is obtained from the control group number of the air conditioning unit, chiller unit, chilled water pump, cooling water pump, and cooling tower fan in the building automation equipment ledger. Equipment numbers are read from the unique equipment codes in the building automation point table; control command numbers are generated sequentially by the building automation main controller each time a scheduling command is issued; data acquisition timestamps are uniformly synchronized and written by the building automation main controller; outdoor dry-bulb temperature values are collected by outdoor temperature sensors (in degrees Celsius); indoor dry-bulb temperature values are collected by target area temperature sensors (in degrees Celsius); supply air temperature values are collected by air conditioning unit supply air duct temperature sensors (in degrees Celsius); chilled water supply temperature values are collected by chilled water supply pipe temperature sensors (in degrees Celsius); chiller unit status flag values are obtained from chiller unit start / stop feedback. Contact readings: The compressor load rate value is read from the chiller unit controller, in percentage; the compressor load valve opening value is read from the compressor load valve position feedback signal, in percentage; the chilled water pump frequency value, cooling water pump frequency value, cooling tower fan frequency value, and air conditioning unit fan frequency value are respectively read from the corresponding frequency converter operating frequency feedback signal, in Hertz; the terminal valve opening value is read from the terminal valve actuator position feedback signal, in percentage; the chiller unit power value, chilled water pump power value, cooling water pump power value, and cooling tower fan power value are respectively read from the corresponding energy meter real-time power register, in kilowatts.The indoor temperature setpoint, chilled water supply temperature setpoint, chilled water pump frequency setpoint, cooling water pump frequency setpoint, cooling tower fan frequency setpoint, and chiller unit start / stop flag value for the dispatch command area are read from the dispatch command cache table. The dispatch command issuance timestamp is generated by the building automation main controller when the command is written. The interlock protection trigger flag value is read from the feedback signal of the equipment interlock protection circuit. The control command execution result flag value is generated by the building automation main controller based on the feedback arrival status, feedback change status, and interlock protection trigger status after the command is issued. The chiller unit status flag value is represented by an operating flag, a shutdown flag, and a fault shutdown flag. The chiller unit start / stop flag for the dispatch command is... The values are represented by start, stop, and hold flags. The interlocking protection trigger flag values are represented by non-triggered and triggered states. When the equipment interlocking protection circuit feedback signal is at a valid level, a triggered state is written; when the equipment interlocking protection circuit feedback signal is at an invalid level, a non-triggered state is written. The control command execution result flag values are represented by execution success, execution failure, execution timeout, and partial execution. When, within the feedback sampling window after the dispatch command is issued, the execution feedback data of the target equipment arrives and its change direction is consistent with the corresponding action direction flag value of the dispatch command field, and the interlocking protection trigger flag value is in the non-triggered state, an execution success state is written; when the interlocking protection trigger flag value is in the triggered state, the target... When any condition in the fault status returned by the equipment feedback signal is met, an execution failure status is written; when no feedback data from the target equipment is received before the feedback sampling window ends, an execution timeout status is written; when only some of the execution feedback data corresponding to multiple scheduling instruction fields corresponding to the same control instruction number has changed, a partial execution status is written; in subsequent blocking node determination, both the execution failure status and the execution timeout status are treated as failure statuses, and scheduling action tokens that have not changed in the partial execution status are treated as failure statuses; using the building automation main controller timestamp as the main time axis, the nearest neighbor timestamp matching algorithm is used to align the execution time of the HVAC energy consumption scheduling monitoring data, specifically according to the air conditioning system number. Each sampling record is grouped by device number, and a unified time index is established using the scheduling instruction issuance timestamp and data acquisition timestamp. The sampling period of the building automation main controller is 10 to 60 seconds. The main time axis sampling points are generated at equal intervals according to the sampling period of the building automation main controller. At each main time axis sampling point, the sampling record with the smallest absolute time difference is found, and the sampling record with the absolute time difference not greater than the time alignment allowable deviation is written into the corresponding sampling point. The time alignment allowable deviation is determined based on the sampling period of the building automation main controller and the transmission delay of the field control network, and the value range is 1 to 5 seconds. When there are multiple candidate records with the same data name at the same sampling point, the record with the data acquisition timestamp closest to the main time axis sampling point is selected.When two sampling records with the same data name exist at the same main time axis sampling point and have the same absolute time difference, the sampling record whose data acquisition timestamp is no later than the main time axis sampling point is selected first. When both sampling records are earlier than the main time axis sampling point and have the same absolute time difference, or when both sampling records are later than the main time axis sampling point and have the same absolute time difference, the sampling record whose device number is bound to the current control command number is selected. When the absolute time difference of the sampling records is greater than the allowable time alignment deviation, no forced matching is performed, and the corresponding data name is written to the missing data marker to avoid mismatch across sampling periods, which could lead to misalignment of execution feedback data and scheduling command data. Local outlier factor calculation is used. The method identifies and removes abnormal records in HVAC energy consumption scheduling and monitoring data. Specifically, it establishes sliding verification windows according to air conditioning system number, equipment number, and data name. Outdoor dry-bulb temperature, indoor dry-bulb temperature, supply air temperature, chilled water supply temperature, compressor load rate, compressor load valve opening, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, terminal valve opening, air conditioning unit fan frequency, chiller unit power, chilled water pump power, cooling water pump power, and cooling tower fan power are treated as continuous numerical verification objects. The local outlier factor algorithm uses Euclidean distance to calculate the continuous numerical verification objects within the sliding verification window. The sample distance is determined by the sliding validation window length, which is 12 to 36 sampling points, and the neighborhood number k, which is 5 to 10, and the neighborhood number k is less than the number of valid sampling points within the sliding validation window. When the number of valid sampling points within the sliding validation window is less than the neighborhood number k, the local outlier calculation is not performed, and the corresponding sampling record is retained for verification in the next sampling window. For temperature data, frequency data, opening degree data, and power data, independent sliding validation windows are established according to the corresponding data names, and data from different physical units are not mixed to calculate the local outlier value. After meeting the neighborhood number k requirement, the reachability distance and local reachability density between each sampling record and adjacent sampling records within the sliding validation window are calculated. The local outlier value is obtained based on the local reachability density ratio. When the local outlier value is greater than the outlier determination threshold, and the corresponding sampling record exceeds at least one of the sensor range boundary, equipment operation boundary, and control command boundary, the corresponding sampling record is marked as an abnormal record and removed. The outlier determination threshold is determined by the local outlier distribution of historical normal operation data, and the value range is 1.5 to 2.5. For the chiller unit status flag value, the chiller unit start / stop flag value of the dispatch command, the interlock protection trigger flag value, and the control command execution result flag value, consistency verification is performed according to the status flag value set. When the status flag does not belong to the legal value set, the corresponding sampling record is marked as an abnormal record and removed.Missing data compensation is performed using piecewise linear interpolation. Specifically, the system number, equipment number, and data name are used to locate the two valid sampling points before and after the missing data. When the duration of the missing data is no greater than the upper limit of the missing data compensation time, the linear change slope is calculated based on the sampling time difference and numerical difference between the two valid sampling points. The compensation value is then calculated according to the sampling point position on the main time axis. The upper limit of the missing data compensation time is determined based on the sampling cycle of the building automation main controller and the response inertia of the HVAC equipment, ranging from 2 to 6 sampling cycles. Piecewise linear interpolation is only used for outdoor dry-bulb temperature, indoor dry-bulb temperature, supply air temperature, chilled water supply temperature, compressor load rate, compressor load valve opening, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, terminal valve opening, air conditioning unit fan frequency, chiller unit power, chilled water pump power, cooling water pump power, and cooling tower fan power. For chiller unit status flag values and scheduling command values... For water turbine start / stop flags, interlock protection trigger flags, and control command execution result flags, piecewise linear interpolation is not used for missing data compensation. When the duration of the missing data is no greater than the upper limit of the missing data compensation time, and the status flags of the two valid sampling points before and after the missing data are consistent, the consistent status flags are used for maintenance compensation. When the status flags of the two valid sampling points before and after the missing data are inconsistent, the corresponding data name is retained as the missing status, and it is not used as the action direction flag value and the basis for determining the blocking node when generating subsequent scheduling action tokens, so as to avoid continuous processing of start / stop status, interlock status, and execution result status. After completing time alignment, abnormal record removal, and missing data compensation, the air conditioning system number, equipment number, control command number, data acquisition timestamp, execution feedback data, target response data, power data, scheduling command data, interlock protection trigger flag value, and control command execution result flag value corresponding to each sampling point are written into the same time sequence record, and the preprocessed HVAC energy consumption scheduling monitoring data is output.
[0026] In this implementation plan, by unifying the data acquisition sources, time base, continuous numerical data, discrete state markers, power data, and scheduling command data in the HVAC energy consumption scheduling and monitoring data under the time axis of the building automation main controller, the corresponding relationship between command changes, equipment feedback, target response, energy consumption results, and interlocking status can be obtained simultaneously when generating scheduling action tokens. This avoids the calculation deviations of token overdraft, token overdraft after replenishment, and feedback idle token marker values caused by sampling time misalignment, abnormal record mixing, state marker distortion, and short-term missing data. This improves the data credibility of the neural network action overdraft token map construction and the feasibility of HVAC energy consumption optimization scheduling results.
[0027] Specifically, based on the preprocessed HVAC energy consumption scheduling and monitoring data, the specific steps for converting the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens are as follows: Read the preprocessed HVAC energy consumption scheduling and monitoring data, and extract the following values from adjacent control instructions according to the air conditioning system number and control instruction number: indoor temperature setpoint of the scheduling instruction area, chilled water supply temperature setpoint of the scheduling instruction, chilled water pump frequency setpoint of the scheduling instruction, cooling water pump frequency setpoint of the scheduling instruction, cooling tower fan frequency setpoint of the scheduling instruction, and chiller unit start / stop flag value of the scheduling instruction; First, the preprocessed HVAC energy consumption scheduling and monitoring data is grouped according to the air conditioning system number, and then sorted in ascending order according to the scheduling instruction issuance timestamp corresponding to the control instruction number, so that two consecutive control instructions under the same air conditioning system number have a continuous temporal relationship; a unique correspondence is established between the control instruction number and the scheduling instruction issuance timestamp under the same air conditioning system number; when there are multiple records for the same control instruction number, the record with the latest scheduling instruction issuance timestamp and whose control instruction execution result flag value is not in the revocation state is retained; when the same scheduling instruction issuance timestamp exists ... When a stamp corresponds to multiple control instruction numbers, the order in which the building automation main controller writes the instructions to the scheduling instruction cache table is used to determine the sequence. Control instruction numbers marked as canceled, resent overwritten, or invalid are not included in the matching of adjacent control instructions. Two adjacent control instructions must simultaneously meet the following conditions: the air conditioning system number is consistent, the scheduling instruction timestamps are consecutive, and the latter control instruction number has not been canceled or resent overwritten. This avoids duplicate calculation of action differences caused by resentment, cancellation, and concurrent instructions. When the air conditioning system number is consistent in two adjacent control instructions, and the scheduling instruction timestamp corresponding to the latter control instruction number is later than the scheduling instruction timestamp corresponding to the former control instruction number, the former control instruction number is used as the action base instruction, and the latter control instruction number is used as the action update instruction. The differences in the same scheduling instruction fields in the latter and former control instruction numbers are calculated to obtain the action difference for the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop.Specifically, the regional indoor temperature action difference is obtained by subtracting the regional indoor temperature setpoint from the scheduling command setpoint in the previous control command number from the regional indoor temperature setpoint in the subsequent control command number; the chilled water supply temperature action difference is obtained by subtracting the chilled water supply temperature setpoint from the scheduling command setpoint in the previous control command number from the chilled water supply temperature setpoint in the subsequent control command number; the chilled water pump frequency action difference is obtained by subtracting the chilled water pump frequency setpoint from the scheduling command setpoint in the previous control command number from the chilled water pump frequency setpoint in the subsequent control command number; and the cooling water pump frequency action difference is obtained by subtracting the chilled water pump frequency setpoint from the cooling water pump frequency setpoint in the previous control command number from the chilled water pump frequency setpoint in the subsequent control command number. The frequency setting value of the cooling water pump in the dispatch command of a control command number is obtained. The frequency action difference of the cooling tower fan is obtained by subtracting the frequency setting value of the cooling tower fan in the dispatch command of the previous control command number from the frequency setting value of the cooling tower fan in the dispatch command of the next control command number. The start-stop action difference of the chiller unit is obtained by the state change result of the chiller unit start-stop flag value in the dispatch command of the next control command number and the chiller unit start-stop flag value in the dispatch command of the previous control command number. For the regional indoor temperature action difference and the chilled water supply temperature action difference, a positive difference is determined as the heating action direction, a negative difference is determined as the cooling action direction, and a zero difference is determined as the holding action direction. For the chilled water supply temperature action difference, a positive difference is determined as the heating action direction, a negative difference is determined as the cooling action direction, and a zero difference is determined as the holding action direction. For the frequency difference of chilled water pumps, cooling water pumps, and cooling tower fans, positive differences are defined as the frequency increase direction, negative differences as the frequency decrease direction, and zero differences as the hold direction. For the chiller unit start-stop action difference, changing the stop flag to an operating flag determines the start direction, changing the operating flag to a stop flag determines the stop direction, and no change in flags determines the hold direction. The action direction flag values use a fixed set of values consisting of temperature increase, temperature decrease, frequency increase, frequency decrease, start, stop, and hold flags. Among them, the temperature increase and temperature decrease flags are only used for the regional indoor temperature action difference and chilled water pump frequency difference. The chilled water supply temperature action difference, frequency increase flag and frequency decrease flag are only used for the frequency action difference of chilled water pump, cooling water pump and cooling tower fan. The start flag and stop flag are only used for the start and stop action difference of chiller unit. The hold flag is used for scheduling action tokens where the action difference is zero or the start and stop status has not changed. When the direction of change of the subsequent execution feedback data is consistent with the action direction flag value of the overdraft token, the comparison is made according to the same type of direction flag under the same scheduling instruction field, and the action direction flag value is not compared across scheduling instruction fields. The scheduling instruction field, action direction flag value and action amplitude value corresponding to each action difference are written into the scheduling action token to generate a scheduling action token group.The dispatch instruction fields include regional indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. The action direction marker value is determined by the positive or negative direction of the action difference or the direction of start / stop status change, and the action amplitude value is determined by the absolute value of the action difference or the result of the start / stop status change. The action amplitude values corresponding to regional indoor temperature and chilled water supply temperature are retained in degrees Celsius, while the action amplitude values corresponding to chilled water pump frequency, cooling water pump frequency, and cooling tower fan frequency are retained in Hertz. The action amplitude values corresponding to chiller unit start / stop are... The action amplitude value is represented by the start / stop state change result; the action amplitude value retains its original physical meaning in the scheduling action token and is used to record the actual adjustment amount of the corresponding scheduling instruction field; before calculating the actual token acceptance amount, token overdraft amount, and token overdraft amount after replenishment, action amplitude scaling benchmarks are established according to the scheduling instruction fields, and a robust normalization processing method based on the median and interquartile range is used to convert the regional indoor temperature action difference, chilled water supply temperature action difference, chilled water pump frequency action difference, cooling water pump frequency action difference, and cooling tower fan frequency action difference into... Dimensionless standardized values for action amplitude; the difference between chiller unit start-up and stop actions is not subtracted from the difference between temperature and frequency actions; the difference between chiller unit start-up and stop actions uses state change encoding to participate in the token state input vector concatenation, and the start-up and stop actions are determined by the chiller unit state flag value; the air conditioning system number, control command number, dispatch command issuance timestamp, and dispatch command field are synchronously written into the dispatch action token, so that each dispatch action token can be traced back to the specific control command and the specific dispatch object; when multiple dispatch actions are generated within the same adjacent control command number... When creating tokens, the scheduling action tokens are written into the scheduling instruction field in a fixed order: regional indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. By converting continuous setpoint differences into scheduling action tokens, the continuous setpoint changes output by the neural network can be decomposed into the smallest action unit with scheduling instruction fields, action direction marker values, and action amplitude values. This provides a unified input for subsequent calculations of the actual token acceptance, token overdraft, and token overdraft after replenishment based on execution feedback data.
[0028] In this implementation scheme, by unifying the changes in the set values of different scheduling instruction fields into scheduling action tokens with scheduling instruction fields, action direction marker values, and action amplitude values, a comparable, traceable, and calculable action expression format is formed between regional indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start-up and shutdown. This avoids the problems of inconsistent field meanings, unclear action directions, and misaligned action amplitudes caused by the direct mixing of multiple types of set values output by the neural network in subsequent judgments. This improves the input consistency for calculating the actual token acceptance amount, token overdraft amount, and token overdraft amount after replenishment, and provides a stable data foundation for the neural network action overdraft token map to accurately identify prohibited token paths and replenishable token paths.
[0029] Specifically, the steps for generating the token overdraft, token replenishment relationship flag value, and replenished token overdraft based on the execution feedback data corresponding to the scheduling action token are as follows: Match the execution feedback data according to the scheduling instruction field of the scheduling action token. Specifically, the regional indoor temperature corresponds to the terminal valve opening value and the air conditioning unit fan frequency value; the chilled water supply temperature corresponds to the compressor load rate value and the compressor load valve opening value; the chilled water pump frequency corresponds to the chilled water pump frequency value; the cooling water pump frequency corresponds to the cooling water pump frequency value; the cooling tower fan frequency corresponds to the cooling tower fan frequency value; and the chiller unit start / stop corresponds to the chiller unit status flag value. The execution feedback data is used to determine whether the scheduling action has been accepted by the actual equipment execution end. The target response data is used to determine whether changes at the execution end are transmitted to the air-side temperature result and the water-side temperature result. When the area indoor temperature corresponds to the terminal valve opening value and the air conditioning unit fan frequency value, the terminal valve opening value and the air conditioning unit fan frequency value are used as execution feedback data to participate in the calculation of the actual token acceptance amount. The indoor air dry-bulb temperature value is used as target response data to participate in the feedback idling token mark value and cross-device compensation relationship determination, but is not directly used as the actual token acceptance amount for the area indoor temperature action. When the chilled water supply temperature corresponds to the compressor load rate value and the compressor load valve opening value, the compressor load rate value and the compressor load valve opening value are used as execution feedback data to participate in the calculation of the actual token acceptance amount. The target response data is used to participate in the feedback idle token flag value and cross-device replenishment relationship determination. When matching execution feedback data, the scheduling instruction issuance timestamp, scheduling instruction field, action direction flag value, and action amplitude value in the scheduling action token are read first. The execution feedback data of the sampling point corresponding to the scheduling instruction issuance timestamp is used as the feedback baseline data. Then, the corresponding execution feedback data is read according to the scheduling instruction field. When one scheduling instruction field corresponds to two execution feedback data, the change amplitude of the two execution feedback data relative to the feedback baseline data is calculated respectively. Based on the linear calibration coefficient between the change amount of the scheduling instruction field and the change amount of the execution feedback data in the historical scheduling instruction sample, the change amplitude of the execution feedback data is determined. The values are converted to the same dimension as the action amplitude values. Specifically, the terminal valve opening value and air conditioning unit fan frequency value corresponding to the regional indoor temperature are converted to the regional indoor temperature acceptance amplitude, the compressor load rate value and compressor load valve opening value corresponding to the chilled water supply temperature are converted to the chilled water supply temperature acceptance amplitude, the chilled water pump frequency value, cooling water pump frequency value and cooling tower fan frequency value are directly used as the acceptance amplitude based on the Hertz change amplitude, and the chiller unit status flag value is used as the acceptance amplitude based on the start-stop status change result. Before calculating the actual acceptance amount of the token, the action amplitude value and acceptance amplitude are converted to the dimensionless action amplitude standardized value and dimensionless acceptance amplitude standardized value under the same scheduling instruction field, respectively.Among them, the regional indoor temperature action difference, chilled water supply temperature action difference, chilled water pump frequency action difference, cooling water pump frequency action difference, and cooling tower fan frequency action difference are robustly standardized according to the median and interquartile range of the historical action amplitude of the corresponding dispatch instruction field; the terminal valve opening value, air conditioning unit fan frequency value, compressor load rate value, and compressor load valve opening value are first converted into the acceptance amplitude under the same dispatch instruction field through the corresponding linear calibration coefficient, and then robustly standardized according to the same scaling benchmark; the actual acceptance amount of the token is represented by the dimensionless acceptance amplitude standardized value, and the token overdraft amount is represented by the dimensionless action amplitude standardized value minus the dimensionless value. The standardized value of the acceptance amplitude is obtained. The start-up and shutdown actions of the chiller unit are judged by state change coding, without continuous subtraction from temperature and frequency actions. Matching execution feedback data is read from N consecutive sampling points after the dispatch command's timestamp. The maximum change amplitude formed by the execution feedback data along the corresponding action direction is determined as the actual acceptance amount of the token. Here, N is determined based on the building automation main controller's sampling period and the HVAC equipment's execution delay, ranging from 3 to 12 sampling points. The maximum change amplitude along the corresponding action direction is the maximum acceptance amplitude in the same direction relative to the feedback reference data within N consecutive sampling points. When the dispatch command field corresponds to two... When processing feedback data, the larger of the two dimensionless standardized values of the acceptance amplitude corresponding to the two acceptance amplitudes is taken as the actual acceptance amount of the token. This ensures that the actual acceptance amount of the token can represent the maximum observable acceptance level of the scheduling action on the real equipment side. The actual acceptance amount of the token is obtained by subtracting the dimensionless standardized value of the action amplitude in the scheduling action token. When the token overdraft amount is less than zero, it is corrected to zero to avoid negative overdraft caused by equipment feedback overshoot. When the token overdraft amount is greater than the overdraft tolerance threshold, the corresponding scheduling action token is marked as an overdraft token. When the token overdraft amount is less than or equal to the overdraft tolerance threshold, the corresponding scheduling action token is marked as an acceptance token. When the actual number of tokens received is greater than zero and the matched target response data remains unchanged within N consecutive sampling points, the corresponding scheduling action token is marked as a feedback idle token. Among them, the target response data is matched according to the scheduling instruction fields recorded in the scheduling action token. The scheduling instruction area indoor temperature setting value is matched with the indoor air dry bulb temperature value, the scheduling instruction chilled water supply temperature setting value is matched with the chilled water supply temperature value, the scheduling instruction chilled water pump frequency setting value, the scheduling instruction cooling water pump frequency setting value and the scheduling instruction cooling tower fan frequency setting value are matched with the air supply temperature value and the indoor air dry bulb temperature value, and the scheduling instruction chiller unit start / stop flag value is matched with the chilled water supply temperature value.The scheduling commands for chilled water pump frequency settings, cooling water pump frequency settings, and cooling tower fan frequency settings are all matched with the supply air temperature and indoor air dry-bulb temperature values. This is because the chilled water pump frequency affects the chilled water distribution flow rate, the cooling water pump frequency affects the condenser-side heat exchange flow rate, and the cooling tower fan frequency affects the cooling water heat dissipation capacity. Changes in these three types of execution feedback data are transmitted to the indoor air dry-bulb temperature value through the chilled water supply temperature value and the supply air temperature value. When determining whether the target response data has changed, the N1st to N2th sampling points after the timestamp of the scheduling command issuance are used as the target response lag window, with N1 ranging from 1 to 3 and N2 ranging from 6 to 18. The supply air temperature value and indoor air dry-bulb temperature value are read within the target response lag window. The maximum unidirectional change in the dry bulb temperature value is used to avoid misinterpreting the normal lag before the command is transmitted to the air side as feedback idling; the target response data remains unchanged within N consecutive sampling points, meaning that the change in the target response data relative to the sampling point corresponding to the timestamp of the scheduling command is less than the target response change threshold. The target response change threshold is determined by the corresponding sensor resolution and the upper limit of the sampling noise of the building automation main controller. The target response change thresholds corresponding to the indoor air dry bulb temperature value, supply air temperature value, and chilled water supply temperature value are taken as 0.1 degrees Celsius to 0.3 degrees Celsius; the overdraft token, token overdraft amount, action direction marker value in the previous control command number and the execution feedback data in the next control command number are read; when the next control command When the direction of change of the execution feedback data corresponding to the same scheduling instruction field in the numbering is consistent with the action direction marker value of the overdraft token, a replenishment relationship with the same device is generated, and the token overdraft amount is subtracted from the same-direction change amplitude of the corresponding execution feedback data to obtain the replenished token overdraft amount; wherein, the same-direction change amplitude is calculated as the same-direction increment of the execution feedback data in the sampling window corresponding to the next control instruction number relative to the end sampling point of the previous control instruction number, and converted into the replenishment amplitude of the same field as the token overdraft amount according to the same linear calibration coefficient, and then converted into a dimensionless replenishment amplitude according to the same scaling reference; when the execution feedback data corresponding to different scheduling instruction fields in the next control instruction number changes, and satisfies the coupling relationship and target response hysteresis recorded in the device coupling table, the replenishment relationship is generated. When the direction of change of the target response data in the subsequent window corresponds to the action direction of the previous overdraft token, and the magnitude of the change of the target response data is greater than the target response change threshold, a cross-device compensation relationship is generated. Among them, the device coupling table is established by the building automation point table and the HVAC system process relationship, recording the target response transmission relationship between the chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, compressor load rate, terminal valve opening value, and air conditioning unit fan frequency value. The magnitude of change of the target response data is first robustly standardized according to the historical median and interquartile range of the corresponding target response data to obtain the dimensionless response contribution value. Then, the dimensionless response contribution value is used as the cross-device compensation magnitude and deducted from the token overdraft amount to obtain the token overdraft amount after compensation.When the direction of change of the execution feedback data corresponding to the same scheduling instruction field in the subsequent control instruction number is opposite to the action direction flag value of the overdraft token, a reverse cancellation relationship is generated, and the reverse change amplitude of the corresponding execution feedback data is added to the token overdraft amount to obtain the replenished token overdraft amount. The reverse change amplitude is calculated as the reverse increment of the execution feedback data within the sampling window corresponding to the subsequent control instruction number relative to the end sampling point of the previous control instruction number, converted to the cancellation amplitude of the same field as the token overdraft amount using the same linear calibration coefficient, and then converted to a dimensionless cancellation amplitude using the same scaling reference. When the calculated replenished token overdraft amount is less than zero, the replenished token overdraft amount is corrected to zero, so that the replenished token overdraft amount only represents the remaining amount of action that has not yet been received by the device feedback and target response. When the replenished token overdraft amount is less than or equal to the overdraft tolerance threshold, After replenishment, the token overdraft amount is returned to zero. The overdraft tolerance threshold is determined based on the sensor resolution corresponding to the scheduling instruction field, the minimum adjustment step size of the controller setpoint, and the robust normalization back calculation error. This threshold is used to eliminate the influence of sensor noise, sampling error, and floating-point calculation error on the overdraft clearing determination. The replenishment relationship within the same device, the replenishment relationship across devices, and the reverse cancellation relationship are written into the token replenishment relationship marker value. The token replenishment relationship marker value includes the previous control instruction number, the next control instruction number, the scheduling action token number, the scheduling instruction field, the action direction marker value, the relationship type, the replenishment amplitude, and the token overdraft amount after replenishment. This ensures that when constructing the neural network action overdraft token graph, the scheduling action token can be used as the graph node, and the token replenishment relationship marker value as the graph edge, while maintaining consistency in the calculation link between the token overdraft amount, the token overdraft amount after replenishment, and the feedback idle token marker value.
[0030] In this implementation scheme, the relationship between scheduling action tokens, execution feedback data, target response data, and subsequent control commands is transformed into token overdraft, feedback idle token flag value, token replenishment relationship flag value, and replenished token overdraft. This allows the executability of HVAC scheduling actions to no longer rely solely on the difference in a single command, but to reflect the combined impact of equipment feedback delay, target response lag, cross-equipment compensation, and reverse cancellation on the action acceptance result. This avoids misjudging actions that have generated feedback but have not formed an effective temperature response as executable actions, and also avoids directly judging actions that can be completed within the subsequent sampling window as failed actions. This improves the accuracy of the neural network action overdraft token map in identifying prohibited token paths, replenishable token paths, and target scheduling action token groups.
[0031] Specifically, the steps for generating an overdraft embedding vector based on the scheduling action token, token overdraft amount, replenished token overdraft amount, and token replenishment relationship flag value are as follows: Take the scheduling action token, action direction flag value, action amplitude value, token overdraft amount, replenished token overdraft amount, overdraft token flag value, feedback idling token flag value, token replenishment relationship flag value, interlocking protection trigger flag value, and control command execution result flag value corresponding to M consecutive control command numbers ending with the current control command number under the same air conditioning system number, and arrange them into a token state sequence according to the scheduling command issuance timestamp order, where M is 2 to 8; First, the scheduling action tokens are grouped according to the air conditioning system number, and then grouped according to the current control command number... The timestamp of the corresponding scheduling instruction is used as the end time of the sequence. The scheduling action tokens corresponding to the preceding M consecutive control instruction numbers are extracted, ensuring the token state sequence covers the overdraft continuation state and the recovery resolution state prior to the current scheduling action. For multiple scheduling action tokens under the same control instruction number, they are written into the token state sequence in a fixed order: regional indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. The maximum number of tokens in the token state sequence is 6M. When the actual number of tokens is less than 6M, an invalid padding token is written at the beginning of the token state sequence, and an invalid bit flag is added to the invalid padding token, ensuring the token state generated in different scheduling cycles remains consistent. The sequence has a consistent input length; each scheduling action token corresponds to a token state unit, which includes the scheduling instruction field code, action direction flag value, action amplitude value, token overdraft amount, token overdraft amount after replenishment, overdraft token flag value, feedback idle token flag value, token replenishment relationship flag value, interlock protection trigger flag value, control instruction execution result flag value, and invalid bit flag; among them, the scheduling instruction field code uses a one-hot encoding method to represent the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop; the action direction flag value uses discrete encoding corresponding to heating up, cooling down, frequency increasing, frequency decreasing, start, stop, and hold, allowing... The token replenishment relationship flag value is represented by discrete encoding corresponding to replenishment relationship with the same device, replenishment relationship across devices, reverse cancellation relationship, and no replenishment relationship. The overdraft token flag value, feedback idle token flag value, interlock protection trigger flag value, control command execution result flag value, and invalid bit flag value are represented by binary encoding. For the action amplitude value, token overdraft amount, and replenished token overdraft amount, numerical scaling benchmarks are first established according to the scheduling command field, and then the dimensionless processing method based on the median and interquartile range is used to obtain the action amplitude standardized value, token overdraft standardized value, and replenished overdraft standardized value, so that the Celsius, Hertz, percentage, and start / stop state change results can be input into the same time-series coding network.One-hot encoding, discrete encoding, binary encoding, normalized action amplitude value, normalized token overdraft value, and normalized overdraft value after replenishment are concatenated to form the token state input vector. When concatenating the token state input vector, a fixed-length concatenation is performed in the following order: scheduling instruction field encoding, action direction flag value encoding, token replenishment relationship flag value encoding, overdraft token flag value encoding, feedback idle token flag value encoding, interlocking protection trigger flag value encoding, control instruction execution result flag value encoding, invalid bit flag encoding, normalized action amplitude value, normalized token overdraft value, and normalized overdraft value after replenishment. The one-hot encoding uses... To maintain the category boundaries between different scheduling instruction fields, discrete encoding is used to represent the category status of action direction marker values and token replenishment relationship marker values; binary encoding is used to represent the trigger status of overdraft token marker values, feedback idle token marker values, interlocking protection trigger marker values, control instruction execution result marker values, and invalid bit marker values; and standardized numerical values are used to represent the continuous change intensity of action amplitude values, token overdraft amounts, and replenished token overdraft amounts. Discrete marker features and continuous numerical features are written into the same token state input vector through a fixed splicing order, enabling the gated recurrent unit network to read the input from each token state unit. The input dimension remains consistent. The token state input vector is input into the overdraft embedding neural network, and the token state input vector is temporally encoded by the gated recurrent unit network to output the overdraft embedding vector corresponding to each scheduling action token. The overdraft embedding neural network includes a token input layer, a gated recurrent unit network, and an embedding output layer. The token input layer receives the token state input vector, and the gated recurrent unit network reads the token state input vector in the order of the scheduling instruction issuance timestamp. At each token state unit, it calculates the update gate, reset gate, and candidate hidden vector based on the current token state input vector and the previous token state hidden vector. The update gate is used to retain the overdraft information that has not yet been replenished in the preceding scheduling action; the reset gate is used to weaken the historical impact that has been resolved by the same device replenishment relationship, cross-device replenishment relationship and reverse cancellation relationship; the candidate latent vector is used to write the action amplitude, token overdraft amount, token overdraft amount after replenishment, feedback idle state and blocking state features of the current scheduling action token; the embedding output layer maps the token state latent vector at each time step to a fixed-dimensional overdraft embedding vector, and the overdraft embedding vector simultaneously retains the action intensity feature, overdraft remaining feature, replenishment relationship feature, feedback idle feature and blocking risk feature of the current scheduling action token;When training the overdraft embedding neural network, the overdraft amount of the replenished tokens in the historical scheduling samples is used as the supervision label for the predicted overdraft amount after replenishment. The forbidden token path markers determined from the historical scheduling samples based on the blocking node connection results, feedback idle token marker values, and the overdraft amount of the replenished tokens are used as the supervision label for the predicted forbidden path marker values. The errors between the predicted overdraft amount and the replenished tokens, and between the predicted forbidden path markers and the forbidden token path markers, are used as the training loss. This ensures that the overdraft embedding vector not only represents the temporal changes of scheduling actions but also the risks of incomplete overdraft clearing and forbidden path risks. After training, the overdraft embedding vectors are used as graph node features of the neural network's action overdraft token graph for subsequent identification of forbidden token paths and replenishable token paths.
[0032] In this implementation, by converting the token state sequence into an overdraft embedding vector, the action direction marker, action amplitude marker, token overdraft amount, token overdraft amount after replenishment, feedback idle token marker, token replenishment relationship marker, interlocking protection trigger marker, and control command execution result marker in the scheduling action token form a unified neural network representation. This can preserve the overdraft continuation relationship, replenishment resolution relationship, and blocking risk relationship between continuous control commands, avoiding the loss of historical overdraft impact due to judging the executability of scheduling actions based solely on action data from a single sampling point. This improves the integrity of the graph node features in the neural network action overdraft token graph and enhances the stability of identifying prohibited token paths and replenishable token paths.
[0033] Specifically, the steps for constructing a neural network action overdraft token graph by combining overdraft embedding vectors and generating forbidden token paths and replenishable token paths are as follows: Figure 3As shown, a neural network action overdraft token graph is constructed using scheduling action tokens as graph nodes, token replenishment relationship markers as graph edges, and overdraft embedding vectors as graph node features. First, a unique graph node number is assigned to each scheduling action token. This number is generated by concatenating the air conditioning system number, control command number, scheduling command field, and scheduling command issuance timestamp, ensuring that each graph node can be traced back to its corresponding scheduling action token. The overdraft embedding vector is then written into the graph node feature field of the corresponding graph node, along with the action direction marker, action amplitude, token overdraft amount, replenished token overdraft amount, overdraft token marker, and feedback idle token marker, allowing the graph node to simultaneously retain the neural network temporal encoding structure. Results and executability determination data; A directed graph edge is established based on the token replenishment relationship marker value. The starting point of the graph edge is the graph node corresponding to the overdraft token in the previous control instruction number, and the ending point of the graph edge is the graph node corresponding to the scheduling action token in the subsequent control instruction number that forms any of the following relationships: same-device replenishment relationship, cross-device replenishment relationship, or reverse cancellation relationship. Specifically, the graph edge corresponding to the same-device replenishment relationship indicates that the same scheduling instruction field continues to inherit the previous overdraft token in the subsequent control instruction number; the graph edge corresponding to the cross-device replenishment relationship indicates that different scheduling instruction fields compensate for the previous overdraft token through changes in target response data; and the graph edge corresponding to the reverse cancellation relationship indicates that the reverse execution feedback data in the subsequent control instruction number increases. Add the unaccepted remaining amount of the previous overdraft token; when the token replenishment relationship flag value is written into the graph edge, the same-device replenishment relationship, cross-device replenishment relationship, and reverse cancellation relationship correspond to three types of directed edges respectively; the direction of the directed edge corresponding to the same-device replenishment relationship is from the overdraft token in the previous control instruction number to the accept token in the same scheduling instruction field in the next control instruction number; the direction of the directed edge corresponding to the cross-device replenishment relationship is from the overdraft token in the previous control instruction number to the accept token in the next control instruction number that causes the same target response data change; the direction of the directed edge corresponding to the reverse cancellation relationship is from the overdraft token in the previous control instruction number to the cancellation token in the next control instruction number that has the opposite direction of the execution feedback data change; graph edge attribute The parameters include relationship type, compensation range, reverse offset range, token overdraft after compensation, target response data change range, and graph edge weight. The graph edge weights corresponding to same-device compensation relationships and cross-device compensation relationships are represented by the ratio of compensation contribution value to action range value, with the compensation contribution value being the difference between token overdraft and token overdraft after compensation. The graph edge weights corresponding to reverse offset relationships are represented by the ratio of reverse offset range to action range value, and are written to a reverse weight flag to distinguish between overdraft mitigation and overdraft amplification effects. When either the interlocking protection trigger flag value is in the trigger state or the control command execution result flag value is in the failure state, the corresponding scheduling action token is written to the blocking node.The graph nodes corresponding to the blocking nodes retain trigger status markers, failure status markers, and corresponding control command numbers, enabling the blocking nodes to characterize the definite risk states of the scheduling action on the real equipment side, such as interlocking protection, execution failure, or execution interruption. Starting from the graph node corresponding to each scheduling action token, a search is performed along the graph edges. Graph paths where the overdraft amount of the replenished token is greater than the overdraft tolerance threshold and reaches the blocking node are written into the prohibited token path. Graph paths where the overdraft amount of the replenished token is less than or equal to the overdraft tolerance threshold and does not reach the blocking node are written into the replenishable token path. The neural network action overdraft token graph data is output. The overdraft tolerance threshold is adjusted based on the sensor resolution of the scheduling command field corresponding to the action amplitude value and the minimum adjustment step size of the controller setpoint. The robust standardization back-calculation error is determined; when the scheduling instruction field is the regional indoor temperature or chilled water supply temperature, the overdraft tolerance threshold is 0.1℃ to 0.3℃; when the scheduling instruction field is the chilled water pump frequency, cooling water pump frequency, or cooling tower fan frequency, the overdraft tolerance threshold is 0.1Hz to 0.5Hz; when the scheduling instruction field is the chiller unit start / stop, the overdraft tolerance threshold is 0; when searching along the graph edges, the graph node corresponding to each overdraft token is used as the starting node, and a depth-first search strategy is used to traverse the directed graph edges in the direction of increasing control instruction number; the maximum path length of a single graph path is L, where L is 2 to 6, and the search path length stops expanding when it reaches L; the visited graph node number is recorded within the same graph path, and the next graph node is used as the next graph node. If the node number already exists in the set of visited graph node numbers, stop the branch search and discard the graph path that forms a loop; stop the branch search when a blocking node, a graph node with a valid feedback idle token flag, a graph node whose token overdraft amount after replenishment is less than or equal to the overdraft tolerance threshold, or when there is no next directed graph edge; after the search, determine the forbidden token path and the replenishable token path based on the replenished token overdraft amount of the terminal graph node, the connection result of the blocking node, and the feedback idle token flag value; when the feedback idle token flag value of any graph node in the graph path is valid, write the graph path into the forbidden token path to avoid executing scheduling actions where the feedback data changes but the target response data does not change and enters the subsequent candidate action selection; when the graph When a path reaches a blocking node, the path is written into the prohibited token path to prevent the scheduling actions corresponding to the interlock protection trigger state and the control command execution failure state from being reissued. When the overdraft amount of the token corresponding to the end node of the path is less than or equal to the overdraft tolerance threshold, and the path does not pass through the blocking node and the node corresponding to the feedback idle token, the path is written into the recoverable token path. When multiple recoverable token paths exist for the same node, the paths are retained in the order of overdraft amount after replenishment from small to large, path length from short to long, and total power value predicted by the system from small to large, so that subsequent candidate action token groups can prioritize matching recoverable token paths with sufficient overdraft resolution, shorter equipment connection links, and lower power results.The neural network action overdraft token graph data includes a set of graph nodes, a set of graph edges, graph node features, graph edge weights, blocking nodes, forbidden token paths, and recoverable token paths. This provides graph structure input for subsequent generation of candidate action token groups based on the current HVAC state sequence, execution of forbidden path replacement, and power ranking.
[0034] In this implementation scheme, the overdraft embedding vector, token replenishment relationship marker value, blocking node, feedback idle token marker value, and replenished token overdraft amount are uniformly organized into a neural network action overdraft token graph. This enables the continuous tracking of the connection links, risk transmission links, and overdraft resolution links between scheduling action tokens. It avoids candidate action token groups from entering the scheduling screening based solely on single-point power prediction results. It can preemptively exclude prohibited token paths with connected blocking nodes, feedback idleness, and unresolved overdrafts, while retaining replenishable token paths with real equipment support. This improves the safety, stability, and executability of subsequent prohibited path replacement, target scheduling action token group selection, and HVAC energy consumption optimization scheduling command issuance.
[0035] Specifically, the steps for generating candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token graph are as follows: Extract preprocessed HVAC energy consumption scheduling monitoring data from the current sampling point according to the air conditioning system number, and assemble the current HVAC state sequence according to the data acquisition timestamp; wherein, the current sampling point is determined by the timestamp of the building automation main controller, and the current HVAC state sequence is composed of the data from each sampling point within the continuous sampling window before the current sampling point and the data from the current sampling point in ascending order of data acquisition timestamps. The length of the continuous sampling window is determined based on the response lag time of the chilled water supply temperature, supply air temperature, and indoor air dry-bulb temperature, ranging from 6 to 24 sampling points; when the number of valid sampling points within the continuous sampling window is insufficient, a zero-filling state is written at the beginning of the sequence, and an invalid bit flag is written simultaneously, enabling the gated loop unit network to distinguish between real sampled data and filled data; the current HVAC state sequence includes the outdoor air dry-bulb temperature, indoor air dry-bulb temperature, supply air temperature, chilled water supply temperature, and chilled water temperature within the continuous sampling window before the current sampling point. The system includes the following parameters: chiller unit status flag value, compressor load rate value, compressor load valve opening value, chilled water pump frequency value, cooling water pump frequency value, cooling tower fan frequency value, terminal valve opening value, air conditioning unit fan frequency value, chiller unit power value, chilled water pump power value, cooling water pump power value, cooling tower fan power value, interlock protection trigger flag value, and control command execution result flag value. Continuous numerical data in the current HVAC status sequence are dimensionless using robust normalization based on median and interquartile range. The flag values, interlock protection trigger flag values, and control command execution result flag values are processed using binary encoding to obtain the current state input sequence. This current HVAC state sequence is then input into the gated loop unit network to obtain the current state latent vector. Before being input into the gated loop unit network, the continuous numerical data of the current HVAC state sequence is first dimensionless, and the chiller unit state flag values, interlock protection trigger flag values, and control command execution result flag values are binary encoded. The encoded sequence serves as the actual input data for the gated loop unit network. The gated loop unit network reads the current state input sequence according to the data acquisition timestamp sequence. It retains the slowly changing states of indoor air dry-bulb temperature, supply air temperature, and chilled water supply temperature by updating the gates, and weakens historical states already received by equipment feedback by resetting the gates. This ensures that the current state latent vector simultaneously represents the current heat load state, equipment operating state, electrical power state, and command execution state. The current state latent vector is then concatenated with the overdraft embedding vector and input into the graph attention layer.Specifically, the current state latent vector is copied to each graph node in the neural network action overdraft token graph, resulting in a state copy vector matching the number of graph nodes. Then, the state copy vector corresponding to each graph node is concatenated with the overdraft embedding vector of the same graph node along the feature dimension to form a node fusion feature vector. The graph attention layer uses this node fusion feature vector as input to the graph nodes, enabling the global HVAC operation state to match the overdraft features of each scheduling action token along the same node dimension. The graph attention layer resets the attention weight of the graph node corresponding to the forbidden token path to zero and, based on the replenishable token... The attention weights of the graph nodes corresponding to the path are used to weight and update the current state latent vector, generating an overdraft mask state vector. Specifically, a path mask matrix is generated based on forbidden token paths and recoverable token paths. The mask values for graph nodes and edges corresponding to forbidden token paths are written to zero, while the mask values for graph nodes and edges corresponding to recoverable token paths are written to one. When the same graph node belongs to both forbidden token paths and recoverable token paths, it is processed according to the forbidden token path first, and the attention matching value of that graph node is set to invalid. After completing the forbidden mask processing, normalization is re-performed only for graph nodes with a mask value of one. The remaining attention weights are summed to one. When there are no valid graph nodes in the replenishable token path, candidate action generation is stopped, and the action token group is output. The graph attention layer uses the current state latent vector as the query vector and the overdraft embedding vector of each graph node in the neural network action overdraft token graph as the key vector and value vector to calculate the attention matching value between the current HVAC state sequence and each scheduling action token. When a graph node belongs to the forbidden token path, the corresponding attention matching value is set to an invalid value, so that the connection blocking node, the token mark value containing feedback idle tokens, and the overdraft amount of the replenished token are greater than the overdraft tolerance. The action path of the threshold does not participate in the generation of candidate actions; when the graph node belongs to the replenishable token path, the corresponding attention matching value is retained, and the retained attention matching value is normalized. Then, the overdraft embedding vector is weighted and summed using the normalized attention weight to obtain the executable action context vector; the executable action context vector is concatenated with the current state latent vector and input into the linear mapping layer to obtain the overdraft mask state vector, so that the candidate action generation process is shielded by the forbidden token path and guided by the replenishable token path; the overdraft mask state vector is input into the multilayer perceptron to output the candidate action token group;The multilayer sensor includes a field classification output terminal, a direction classification output terminal, and an amplitude regression output terminal. The field classification output terminal outputs candidate scheduling command fields from the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. The direction classification output terminal outputs action direction marker values from the heating marker, cooling marker, frequency increase marker, frequency decrease marker, start marker, stop marker, and hold marker. The amplitude regression output terminal outputs candidate action amplitude values and performs boundary mapping according to the setpoint boundary, equipment operating boundary, and minimum adjustment step size corresponding to the scheduling command field. Scheduling command field and action direction marker value combinations that do not meet the valid action constraints are eliminated. After elimination, the action tokens with the highest predicted probabilities are retained to form a candidate action token group. The output layer of the multilayer perceptron corresponds to candidate actions for regional indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. The output includes a scheduling instruction field, an action direction marker value, and an action amplitude value. For regional indoor temperature and chilled water supply temperature, the action amplitude value is in degrees Celsius; for chilled water pump frequency, cooling water pump frequency, and cooling tower fan frequency, the action amplitude value is in Hertz; for chiller unit start / stop, the action amplitude value is represented by the start / stop state change result. The candidate action token group is mapped to a neural network action overdraft command. In the token graph analysis, when the graph path corresponding to the candidate action token group has overlapping nodes or edges with the forbidden token path, robust standardization is performed based on the median and interquartile range of the historical token overdraft amounts corresponding to different scheduling instruction fields before comparing the overdraft amounts after replenishment for different scheduling instruction fields, resulting in a dimensionless overdraft remaining score. For the overdraft amounts after replenishment corresponding to the start-up and shutdown of chiller units, the start-up and shutdown success markers are converted into a dimensionless overdraft remaining score. When selecting paths, the minimum dimensionless overdraft remaining score is used as the first sorting criterion, followed by the shortest path length and the disconnection node connection result being unconnected as subsequent sorting criteria, avoiding direct comparison of results based on Celsius, Hertz, percentage, and start-up / shutdown status changes. This can lead to path selection bias. The search along the graph edges finds the recoverable token path that matches the action direction marker value of the candidate action token group and has the smallest dimensionless overdraft remaining score. The candidate action token group is then replaced with the corresponding recoverable action token group. The graph path corresponding to the candidate action token group is obtained by matching the scheduling instruction field, action direction marker value, and action amplitude value in the candidate action token group with the scheduling instruction field, action direction marker value, and action amplitude value in the graph nodes. Overlapping nodes refer to graph node numbers that match the node numbers in the forbidden token path, and overlapping edges refer to token recovery relationship marker values that match the graph edge relationships in the forbidden token path.When searching along the graph edges, execution follows the ascending direction of the control command number. The path with the smallest recoverable token score—where the action direction marker value is consistent, the feedback idle token marker value is invalid, there are no connected blocking nodes, and the remaining score is the smallest—is selected as the replacement path. During replacement, the scheduling command field of the candidate action token group is retained, the action amplitude value is updated to the action amplitude value corresponding to the graph node at the end of the replacement path, and the action direction marker value is updated to the action direction marker value corresponding to the graph node at the end of the replacement path, forming a recoverable action token group. When the graph path corresponding to the candidate action token group and the forbidden token path do not have overlapping nodes or overlapping edges, the candidate action token group is retained.
[0036] In this implementation scheme, by masking and fusing the current HVAC state sequence with the neural network action overdraft token graph, the generation process of candidate action token groups is simultaneously constrained by the current heat load state, equipment operating state, electrical power state, command execution state, and recoverable token paths. This avoids the neural network outputting low-power action results that fall into forbidden token paths based solely on the current state latent vector. After shielding forbidden token paths and introducing recoverable token paths at the graph attention layer, candidate action token groups can preferentially approach action areas that have completed overdraft resolution and are not connected to blocking nodes. This reduces the risk impact of feedback idle token flag values, interlock protection trigger flag values, and control command execution result flag values on scheduling landing, thereby improving the safety, stability, and equipment acceptability of the target scheduling action token group generation process.
[0037] Specifically, the steps for performing forbidden path replacement and power sorting on candidate action token groups to obtain target scheduling action token groups are as follows: The retained candidate action token groups and the replacement-obtainable replenishable action token groups are combined with the current HVAC state sequence to form a prediction input vector. This prediction input vector is then input into a multilayer perceptron, which outputs the predicted total power value of the system corresponding to each action token group. The prediction input vector consists of the outdoor dry-bulb temperature, indoor dry-bulb temperature, supply air temperature, chilled water supply temperature, chiller status flag, compressor load rate, compressor loading valve opening, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and terminal valve opening values from the current HVAC state sequence. The values of air conditioning unit fan frequency, chiller unit power, chilled water pump power, cooling water pump power, and cooling tower fan power, along with the scheduling instruction field, action direction marker value, and action amplitude value in the action token group, are concatenated to form a data structure. Before concatenation, the outdoor air dry-bulb temperature, indoor air dry-bulb temperature, supply air temperature, chilled water supply temperature, compressor load rate, compressor load valve opening, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, terminal valve opening, air conditioning unit fan frequency, chiller unit power, chilled water pump power, cooling water pump power, and cooling tower fan power are processed using a robust normalization method based on the median and interquartile range. Dimensionality is achieved by using discrete encoding for the chiller unit status flag values, scheduling instruction fields, and action direction flag values. Action amplitude values are scaled according to the units corresponding to the scheduling instruction fields, enabling data from different physical units to serve as input to the same multilayer perceptron. The multilayer perceptron includes an input layer, a hidden layer, and an output layer. The input layer receives the predicted input vector. The hidden layer extracts the power response relationship between the current HVAC state sequence and the action token group using a nonlinear activation function. The output layer outputs the system's predicted total electrical power value, which is obtained by summing the predicted electrical power values of the chiller unit, chilled water pump, cooling water pump, and cooling tower fan. The system reads the corresponding feedback values for each action token group. The token overdraft amount, feedback idle token flag value, and blocking node connection result are used to eliminate action token groups that meet any of the following conditions: feedback idle token flag value, connection to a blocking node, or token overdraft amount exceeding the overdraft tolerance threshold after replenishment. The feedback idle token flag value is used to exclude action token groups where the execution feedback data has changed but the target response data has not changed effectively. The blocking node connection result is used to exclude action token groups associated with interlock protection trigger status and control command execution failure status. The token overdraft amount after replenishment is used to exclude action token groups that have not yet been accepted by the actual equipment side. When there are remaining action token groups after elimination, the action token group with the smallest predicted total power value of the system is selected as the target scheduling action token group.When multiple action token groups with the same predicted total power value exist in the remaining action token groups, the target scheduling action token group is determined in the following order: smaller token overdraft after replenishment, invalid feedback idle token flag, disconnection node connection result is not connected, and smaller action amplitude value. This ensures that the target scheduling action token group simultaneously satisfies the replenishable path constraint, equipment acceptance constraint, and minimum power constraint. If no remaining action token group exists after removal, the indoor temperature setpoint of the scheduling command area corresponding to the current control command number, the chilled water supply temperature setpoint of the scheduling command, the chilled water pump frequency setpoint of the scheduling command, and the scheduling command... The system sets the cooling water pump frequency setting, the cooling tower fan frequency setting, and the chiller unit start / stop flag value. It then writes the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop value into the scheduling command field of the hold action token group. The system also writes the action direction flag value corresponding to each scheduling command field into the hold field, and writes the action amplitude value corresponding to each scheduling command field into zero. This hold action token group is then used as the target scheduling action token group, ensuring that even when all candidate action token groups are eliminated, the target scheduling action token group can still be output without increasing equipment movement disturbance.
[0038] In this implementation scheme, by introducing a joint screening of the overdraft amount of the replenished token, the feedback idle token flag value, the blocking node connection result, and the predicted total power value before the candidate action token group enters the final issuance, the low power candidate actions must first meet the actual equipment acceptance conditions, the target response validity conditions, and the interlocking safety conditions. This avoids selecting action token groups that still have overdrafts that have not been cleared, feedback idle, or blocking connections simply because the power prediction is low. As a result, the target scheduling action token group retains the energy consumption optimization direction and can match the actual execution boundary of the HVAC equipment, thereby improving the stability and executability of the HVAC energy consumption optimization scheduling command after issuance.
[0039] Table 1. Action Overdraft Token Path Replacement and Power Filtering Data Table
[0040] In this embodiment, eight scheduling action tokens generated by an air conditioning system within a continuous scheduling cycle are used as an example to illustrate the process of replacing prohibited paths and selecting target scheduling actions in the neural network action overdraft token graph. The system calculates the token overdraft amount after replenishment based on the execution feedback data corresponding to the scheduling action token, and determines whether each scheduling action token can be accepted by the actual HVAC equipment by combining the feedback idle token flag value and the blocking node connection result. For scheduling action tokens that enter prohibited token paths, a replenishable token path is searched along the neural network action overdraft token graph, and the prohibited token is replaced with a replenishable token that satisfies the action direction constraint and has a zero token overdraft amount after replenishment. Subsequently, the executable candidate set is sorted according to the system's predicted total power value to determine the target scheduling action token, and the target scheduling action token is converted into an HVAC energy consumption optimization scheduling instruction.
[0041] In Table 1, the overdraft amount of the token after replenishment is used to characterize the remaining amount of the scheduling action that has not been accepted by the equipment side after processing by the same equipment replenishment relationship, cross-equipment replenishment relationship, or reverse offset relationship; the feedback idle token flag value is used to characterize the state that the execution feedback data has changed but the target response data has not changed effectively; the blocking node connection result is used to characterize whether the scheduling action token is connected to the blocking node corresponding to the interlock protection trigger or the control command execution failure. As can be seen from Table 1, the overdraft amount of the token after replenishment of T1 is still greater than the overdraft tolerance threshold. Although it is not connected to the blocking node, it does not meet the conditions for the replenishable token path, so it is classified as a token to be replenished; although the total power values predicted by the system for T2, T3, and T4 are 418.6kW, 414.2kW, and 421.5kW, respectively, they all have overdraft amounts of the token after replenishment that are greater than zero or are connected to the blocking node, so they are classified as prohibited token paths; T2, T3, and T4 are replaced by T6, T7, and T8, respectively. T5, T6, T7, and T8 were not connected to any blocking nodes, did not trigger feedback idling, and had a token overdraft of 0.00 after replenishment. Among them, T8's system predicted total power value was 417.9kW, which was the lowest in the executable candidate set. Therefore, T8 was determined as the target scheduling action token.
[0042] like Figure 4As shown in the figure, the horizontal axis represents the token overdraft amount after replenishment, and the vertical axis represents the total predicted power value of the system. In the figure, T1 is the token to be replenished. After replenishment, the token overdraft amount is still greater than the overdraft tolerance threshold, which does not meet the conditions for replenishable token paths. Therefore, it does not participate in the selection of target scheduling action token groups. T2, T3, and T4 are located in the area where the token overdraft amount is greater than zero after replenishment and are marked as prohibited tokens. This indicates that although the corresponding candidate actions have a low total predicted power value of the system, there is still a risk of action overdraft, feedback idling, or blocking node connections. They are not suitable for direct issuance as HVAC energy consumption optimization scheduling instructions. T5, T6, T7, and T8 are located in the executable area where the token overdraft amount is zero after replenishment and are marked as replenishable tokens or target tokens. This indicates that the corresponding actions have been received and corrected through token replenishment relationships. The purple arrows indicate the direction of forbidden path replacement, where T2 is replaced by T6, T3 by T7, and T4 by T8, illustrating the forbidden path replacement process of candidate action token groups based on the neural network action overdraft token graph of this invention. The green trajectory line represents the power trajectory of the replenishable token. Among the executable candidate set, T8 corresponds to the lowest predicted total electrical power value of the system, and is therefore identified as the target scheduling action token. Figure 4 This invention demonstrates that it does not directly select the lowest power action output by the neural network, but rather determines the HVAC energy consumption optimization scheduling result based on the system's predicted total power value after eliminating forbidden token paths and completing the replaceable replacement.
[0043] Specifically, the steps for converting the target scheduling action token group into HVAC energy consumption optimization scheduling instructions and issuing them, and outputting the HVAC energy consumption optimization scheduling results are as follows: Generate HVAC energy consumption optimization scheduling instructions according to the scheduling instruction fields, action direction marker values, and action amplitude values in the target scheduling action token group; First, read each target scheduling action token in the target scheduling action token group, extract the scheduling instruction fields, action direction marker values, and action amplitude values corresponding to the target scheduling action token, and then read the indoor temperature setpoint of the scheduling instruction area, the chilled water supply temperature setpoint, the chilled water pump frequency setpoint, and the scheduling instruction value corresponding to the current control instruction number according to the air conditioning system number and the current control instruction number. The command cooling water pump frequency setpoint, the dispatch command cooling tower fan frequency setpoint, and the dispatch command chiller unit start / stop flag value serve as the command conversion reference. For the zone indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, and cooling tower fan frequency, based on the action direction flag value, the dispatch command setpoint with the same dispatch command field as the target dispatch action token in the current control command number is added to or subtracted from the corresponding action amplitude value to obtain the corresponding target setpoint. Among them, when the action direction flag value is the heating action direction, the dispatch command zone indoor temperature setpoint corresponding to the current control command number is added to the action amplitude value corresponding to the zone indoor temperature to obtain the target zone indoor temperature setpoint. When the action direction marker is a cooling action direction, the target area indoor temperature setting value is obtained by subtracting the corresponding action amplitude value from the current control command's area indoor temperature setting value. When the action direction marker is a heating action direction, the target chilled water supply temperature setting value is obtained by adding the current control command's area chilled water supply temperature setting value to the corresponding action amplitude value. When the action direction marker is a cooling action direction, the target chilled water supply temperature setting value is obtained by subtracting the corresponding action amplitude value from the current control command's area chilled water supply temperature setting value. When the frequency increase action direction is selected, the frequency setting values of the chilled water pump, cooling water pump, and cooling tower fan corresponding to the current control command number are added to the corresponding action amplitude values to obtain the target chilled water pump frequency setting value, target cooling water pump frequency setting value, and target cooling tower fan frequency setting value. When the action direction marker value is the frequency decrease action direction, the corresponding action amplitude values are subtracted from the frequency setting values of the chilled water pump, cooling water pump, and cooling tower fan corresponding to the current control command number to obtain the target chilled water pump frequency setting value, target cooling water pump frequency setting value, and target cooling tower fan frequency setting value.When the action direction marker is "maintain action direction", the scheduling instruction setting value of the corresponding scheduling instruction field in the current control instruction number is retained, and the retained setting value is used as the corresponding target setting value. After generating the target setting value, boundary limiting processing, minimum adjustment step size processing, and rate of change constraint processing are performed on the target area indoor temperature setting value, target chilled water supply temperature setting value, target chilled water pump frequency setting value, target cooling water pump frequency setting value, and target cooling tower fan frequency setting value. Among them, the boundary limiting processing is based on the upper limit and lower limit values that can be issued recorded in the building automation point table. When the target setting value is higher than the upper limit value that can be issued, it is written as the upper limit value that can be issued; when the target setting value is lower than the lower limit value that can be issued, it is written as the lower limit value that can be issued. The minimum adjustment step size processing is based on the minimum resolution of the setpoints recorded in the building automation point table. The target area indoor temperature setpoint and the target chilled water supply temperature setpoint are corrected to integer multiples of the corresponding temperature setting step size. The target chilled water pump frequency setpoint, the target cooling water pump frequency setpoint, and the target cooling tower fan frequency setpoint are also corrected to integer multiples of the corresponding frequency setting step size. The rate of change constraint processing is based on the allowable single change and the allowable change per unit time. When the change of the target setpoint relative to the setpoint corresponding to the current control command number exceeds the allowable single change, the target setpoint is truncated according to the allowable single change. When the rate of change of the target setpoint relative to the timestamp of the previous scheduling command exceeds the allowable change per unit time... When optimizing the quantity, the target setpoint is corrected according to the allowable change per unit time to ensure that the HVAC energy consumption optimization scheduling command meets the receiving range of the building automation main controller and the execution range of the actual equipment. For chiller unit start-up and shutdown, the chiller unit start-up and shutdown flag value corresponding to the current control command number is updated according to the action direction flag value to obtain the target chiller unit start-up and shutdown flag value. Among them, when the action direction flag value is the start action direction, the chiller unit status flag value, interlock protection trigger flag value, control command execution result flag value, chilled water supply temperature value, indoor air dry bulb temperature value, and current chiller unit start-up and shutdown duration are read first. Only when the chiller unit status flag value is the shutdown flag value, the interlock protection trigger flag value is the non-triggered state, and the control command execution result flag value is read, will the target chiller unit start-up and shutdown flag value be changed. The target chiller unit start / stop flag value will only be written as running flag when the result flag value is not in the execution failure state, the current chiller unit shutdown duration is not less than the minimum shutdown interval, the chilled water supply temperature is higher than the target chilled water supply temperature setting value, and the indoor air dry bulb temperature is higher than the target area indoor temperature setting value; when the action direction flag value is the shutdown action direction, the target chiller unit start / stop flag value will only be written as shutdown flag when the chiller unit status flag value is running flag, the interlock protection trigger flag value is not triggered, the current chiller unit operation duration is not less than the minimum operation interval, the chilled water supply temperature is not higher than the target chilled water supply temperature setting value, and the indoor air dry bulb temperature is not higher than the target area indoor temperature setting value.When any start / stop safety verification condition is not met, the start / stop flag value of the target chiller unit is kept at the start / stop flag value of the chiller unit corresponding to the current control command number, and the start / stop safety verification failure flag is written into the HVAC energy consumption optimization scheduling command; when the action direction flag value is to maintain the action direction, the start / stop flag value of the chiller unit corresponding to the current control command number is kept unchanged; the target setpoints, the target chiller unit start / stop flag values, and the start / stop safety verification failure flags are combined to generate the HVAC energy consumption optimization scheduling command, a new control command number is assigned to the HVAC energy consumption optimization scheduling command, and it is sent to the main controller of the building automation system. The system receives and writes back the execution feedback data bound to the new control command number, and outputs the HVAC energy consumption optimization scheduling result. The new control command number is generated sequentially according to the air conditioning system number and the dispatch command issuance timestamp, and is written into the dispatch command cache table along with the target area indoor temperature setpoint, target chilled water supply temperature setpoint, target chilled water pump frequency setpoint, target cooling water pump frequency setpoint, target cooling tower fan frequency setpoint, target chiller unit start / stop flag value, and dispatch command issuance timestamp. The building automation main controller writes the HVAC energy consumption optimization dispatch command into the chiller unit controller, chilled water pump inverter, cooling water pump inverter, cooling tower fan inverter, air conditioning unit controller, and terminal valve actuator under the corresponding air conditioning system number according to the new control command number. In the feedback sampling window after the dispatch command issuance timestamp, the new control command is read. The command number is bound to the following values: chiller unit status flag, compressor load rate, compressor load valve opening, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, terminal valve opening, air conditioning unit fan frequency, chiller unit power, chilled water pump power, cooling water pump power, cooling tower fan power, interlock protection trigger flag, and control command execution result flag. The read results are then bound to the new control command number and written back. The measured total power value is obtained by summing the chiller unit power, chilled water pump power, cooling water pump power, and cooling tower fan power values. When the control command execution result flag is in a successful execution state and the interlock protection trigger flag is in a non-triggered state, the new command number is... The control command number, target area indoor temperature setpoint, target chilled water supply temperature setpoint, target chilled water pump frequency setpoint, target cooling water pump frequency setpoint, target cooling tower fan frequency setpoint, target chiller unit start / stop flag value, execution feedback data, system predicted total power value, and measured total power value are written into the HVAC energy consumption optimization scheduling result. When the control command execution result flag is in the execution failure state, or the interlock protection trigger flag is in the trigger state, the new control command number, target scheduling action token group, blocking node connection result, execution feedback data, system predicted total power value, and measured total power value are written into the HVAC energy consumption optimization scheduling result as data input for the next round of neural network action overdraft token graph update.
[0044] In this implementation plan, by converting the target scheduling action token set into HVAC energy consumption optimization scheduling instructions with new control instruction numbers, and by forming a closed-loop binding of the target setpoint, target chiller start / stop flag value, execution feedback data, interlock protection trigger flag value, and control instruction execution result flag value, the scheduling result can be smoothly implemented from the neural network output action into an instruction format that the building automation main controller can recognize, avoiding the target scheduling action token set remaining at the calculation result level. At the same time, by feedback writing back to incorporate the actual equipment execution status into the subsequent neural network action overdraft token map update process, the deviation between the candidate action token set and the actual HVAC equipment's capacity can be continuously corrected, thereby improving the continuous traceability, execution verifiability, and long-term stability of the HVAC energy consumption optimization scheduling results.
[0045] like Figure 2 As shown, the second aspect of the present invention provides a neural network-based HVAC energy consumption optimization scheduling system, comprising: an HVAC data alignment processing module, an action overdraft token generation module, a neural network overdraft map construction module, and a mask scheduling instruction output module, wherein: the HVAC data alignment processing module is used to collect HVAC energy consumption scheduling monitoring data, perform execution time alignment, abnormal record removal, and missing data compensation processing on the HVAC energy consumption scheduling monitoring data, and output preprocessed HVAC energy consumption scheduling monitoring data; the action overdraft token generation module is used to convert the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens based on the preprocessed HVAC energy consumption scheduling monitoring data, and generate a token based on the execution feedback data corresponding to the scheduling action tokens. The system includes: a token overdraft mapping module, a neural network action overdraft token mapping module, a mask scheduling instruction output module, and a mask scheduling instruction output module. The mask scheduling instruction output module generates candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token mapping module. It performs forbidden path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token group. The target scheduling action token group is then converted into an HVAC energy consumption optimization scheduling instruction and issued, outputting the HVAC energy consumption optimization scheduling result.
[0046] In this implementation plan, a continuous processing link is established by combining HVAC energy consumption scheduling monitoring data, scheduling action tokens, token overdraft amount, token replenishment relationship marker value, overdraft embedding vector, neural network action overdraft token map, and HVAC energy consumption optimization scheduling results. This ensures that the candidate action token group output by the neural network is simultaneously constrained by the actual equipment capacity, historical overdraft resolution status, prohibited token path, and replenishable token path before being issued. This prevents low-power candidate actions from directly bypassing the interlock protection trigger marker value, control command execution result marker value, and feedback idling token marker value to enter the actual control process, thereby improving the execution reliability of HVAC energy consumption optimization scheduling commands, the stability of energy consumption optimization, and the ability to be implemented in real scenarios.
[0047] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0048] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A neural network-based HVAC energy consumption optimization scheduling method, characterized in that, Includes the following steps: S1 collects HVAC energy consumption scheduling and monitoring data, performs time alignment, abnormal record removal and missing data compensation on the HVAC energy consumption scheduling and monitoring data, and outputs the pre-processed HVAC energy consumption scheduling and monitoring data. S2, based on the preprocessed HVAC energy consumption scheduling and monitoring data, convert the scheduling instruction field corresponding to the adjacent control instruction number into a scheduling action token, and generate the token overdraft amount, token replenishment relationship flag value and replenished token overdraft amount according to the execution feedback data corresponding to the scheduling action token; S3 generates an overdraft embedding vector based on the scheduling action token, token overdraft amount, token overdraft amount after replenishment, and token replenishment relationship label value. It then constructs a neural network action overdraft token graph by combining the overdraft embedding vector, generating prohibited token paths and replenishable token paths. S4. Based on the current HVAC state sequence and the neural network action overdraft token map, generate candidate action token groups, perform forbidden path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token group, convert the target scheduling action token group into HVAC energy consumption optimization scheduling instructions and issue them, and output the HVAC energy consumption optimization scheduling results.
2. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 1, characterized in that: The specific steps for collecting HVAC energy consumption scheduling and monitoring data, including time alignment, abnormal record removal, and missing data compensation, are as follows: Collect HVAC energy consumption dispatch monitoring data, which includes air conditioning system number, equipment number, control command number, data acquisition timestamp, outdoor dry bulb temperature, indoor dry bulb temperature, supply air temperature, chilled water supply temperature, chiller unit status flag value, compressor load rate value, compressor load valve opening value, chilled water pump frequency value, cooling water pump frequency value, cooling tower fan frequency value, terminal valve opening value, air conditioning unit fan frequency value, chiller unit power value, chilled water pump power value, cooling water pump power value, cooling tower fan power value, dispatch command area indoor temperature setpoint, dispatch command chilled water supply temperature setpoint, dispatch command chilled water pump frequency setpoint, dispatch command cooling water pump frequency setpoint, dispatch command cooling tower fan frequency setpoint, dispatch command chiller unit start / stop flag value, dispatch command issuance timestamp, interlock protection trigger flag value, and control command execution result flag value. Using the timestamp of the main controller of the building automation system as the main time axis, the nearest neighbor timestamp matching algorithm is used to perform time alignment processing on the HVAC energy consumption scheduling and monitoring data. The local outlier factor algorithm is used to identify and remove abnormal records in the HVAC energy consumption scheduling and monitoring data, and piecewise linear interpolation is used for missing data compensation to output preprocessed HVAC energy consumption scheduling and monitoring data.
3. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 2, characterized in that: The specific steps for converting the scheduling instruction field corresponding to adjacent control instruction numbers into scheduling action tokens based on the preprocessed HVAC energy consumption scheduling and monitoring data are as follows: Read the preprocessed HVAC energy consumption dispatch monitoring data, and extract the following values from adjacent control commands according to the air conditioning system number and control command number: indoor temperature setpoint for the dispatch command area, chilled water supply temperature setpoint, chilled water pump frequency setpoint, cooling water pump frequency setpoint, cooling tower fan frequency setpoint, and chiller unit start / stop flag value. Calculate the differences between the same dispatch command fields in the subsequent control command number and the preceding control command number to obtain the area indoor temperature action difference, chilled water supply temperature action difference, and chilled water pump frequency action difference. The system calculates the frequency difference of the cooling water pump, the frequency difference of the cooling tower fan, and the start / stop difference of the chiller unit. It then writes the corresponding scheduling instruction field, action direction marker value, and action amplitude value into a scheduling action token to generate a scheduling action token group. The scheduling instruction field includes the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. The action direction marker value is determined by the positive or negative direction of the action difference or the direction of start / stop status change. The action amplitude value is determined by the absolute value of the action difference or the result of the start / stop status change.
4. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 3, characterized in that: The specific steps for generating the token overdraft amount, token replenishment relationship flag value, and replenished token overdraft amount based on the execution feedback data corresponding to the scheduling action token are as follows: The execution feedback data is matched with the scheduling instruction fields of the scheduling action token. Among them, the regional indoor temperature corresponds to the terminal valve opening value and the air conditioning unit fan frequency value, the chilled water supply temperature corresponds to the compressor load rate value and the compressor load valve opening value, the chilled water pump frequency corresponds to the chilled water pump frequency value, the cooling water pump frequency corresponds to the cooling water pump frequency value, the cooling tower fan frequency corresponds to the cooling tower fan frequency value, and the chiller unit start / stop corresponds to the chiller unit status flag value. The matched execution feedback data is read from N consecutive sampling points after the timestamp of the scheduling instruction, and the maximum change amplitude formed by the execution feedback data along the corresponding action direction is determined as the actual token acceptance amount. Subtract the actual amount received by the token from the action amplitude value in the scheduling action token to obtain the token overdraft amount; when the token overdraft amount is greater than zero, the corresponding scheduling action token is marked as an overdraft token; when the token overdraft amount is equal to zero, the corresponding scheduling action token is marked as an acceptance token; when the actual amount received by the token is greater than zero and the matched target response data remains unchanged within N consecutive sampling points, the corresponding scheduling action token is marked as a feedback idle token; among them, the target response data is matched according to the scheduling instruction fields recorded in the scheduling action token, the scheduling instruction area indoor temperature setpoint is matched with the indoor air dry bulb temperature value, the scheduling instruction chilled water supply temperature setpoint is matched with the chilled water supply temperature value, the scheduling instruction chilled water pump frequency setpoint, the scheduling instruction cooling water pump frequency setpoint and the scheduling instruction cooling tower fan frequency setpoint are matched with the supply air temperature value and the indoor air dry bulb temperature value, and the scheduling instruction chiller unit start / stop flag value is matched with the chilled water supply temperature value; Read the overdraft token, overdraft amount, and action direction marker value from the previous control instruction number, and the execution feedback data from the next control instruction number; when the change direction of the execution feedback data corresponding to the same scheduling instruction field in the next control instruction number is consistent with the action direction marker value of the overdraft token, generate a same-device compensation relationship, and subtract the same-direction change amplitude of the corresponding execution feedback data from the overdraft amount to obtain the compensated overdraft amount; when the execution feedback data corresponding to different scheduling instruction fields in the next control instruction number changes, and causes a change in the corresponding target response data, generate a cross-device compensation relationship, and subtract the change amplitude of the corresponding target response data from the overdraft amount to obtain the compensated overdraft amount; when the change direction of the execution feedback data corresponding to the same scheduling instruction field in the next control instruction number is opposite to the action direction marker value of the overdraft token, generate a reverse cancellation relationship, and add the reverse change amplitude of the corresponding execution feedback data to the overdraft amount to obtain the compensated overdraft amount. Write the same-device compensation relationship, cross-device compensation relationship, and reverse offset relationship into the token compensation relationship flag value.
5. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 4, characterized in that: The specific steps for generating the overdraft embedding vector based on the scheduling action token, token overdraft amount, token overdraft amount after replenishment, and token replenishment relationship marker value are as follows: The scheduling action tokens corresponding to M consecutive control command numbers under the same air conditioning system number, ending with the current control command number, are arranged into a token state sequence according to the timestamp of the scheduling command issuance. Within each control command number, the scheduling action tokens are arranged in the order of zone indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, cooling tower fan frequency, and chiller unit start / stop. When the actual number of tokens in the token state sequence is less than 6M, an invalid filler token is written at the beginning of the sequence. The action direction flag, overdraft token flag, feedback idling token flag, and token return / replenishment flag are set. The system flag value, interlock protection trigger flag value, and control command execution result flag value are converted into discrete codes. The action amplitude value, token overdraft amount, and replenished token overdraft amount are converted into standardized values to form a token state input vector. The token state input vector is input into the overdraft embedding neural network and time-series encoded through a gated recurrent unit network to output the overdraft embedding vector corresponding to each scheduling action token. During the training of the overdraft embedding neural network, the error between the replenished overdraft prediction value and the replenished token overdraft amount, and the error between the forbidden path prediction flag value and the forbidden token path flag are used as training losses.
6. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 5, characterized in that: The specific steps for constructing a neural network action overdraft token graph by combining overdraft embedding vectors and generating forbidden token paths and replenishable token paths are as follows: Using scheduling action tokens as graph nodes, token replenishment relationship markers as graph edges, and overdraft embedding vectors as graph node features, a neural network action overdraft token graph is constructed. When either the interlocking protection trigger marker value is in the trigger state or the control command execution result marker value is in the failure state, the corresponding scheduling action token is written to the blocking node. Starting from the graph node corresponding to each scheduling action token, search along the graph edges. Write the graph path where the token overdraft amount after replenishment is greater than the overdraft tolerance threshold and reaches the blocking node into the forbidden token path. Write the graph path where the token overdraft amount after replenishment is less than or equal to the overdraft tolerance threshold and does not reach the blocking node into the replenishable token path. Output the neural network action overdraft token graph data.
7. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 6, characterized in that: The specific steps for generating candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token graph are as follows: Preprocessed HVAC energy consumption scheduling and monitoring data of the current sampling point are extracted according to the air conditioning system number, and the current HVAC state sequence is formed according to the data acquisition timestamp; the current HVAC state sequence is input into the gated cyclic unit network to obtain the current state hidden vector; The current state latent vector and the overdraft embedding vector are concatenated and input into the graph attention layer. The graph attention layer resets the attention weight of the graph node corresponding to the forbidden token path to zero, and updates the current state latent vector with weights according to the attention weight of the graph node corresponding to the replenishable token path to generate the overdraft mask state vector. The overdraft mask state vector is input into a multilayer perceptron, which outputs candidate action token sets. The candidate action token sets are mapped to the neural network action overdraft token graph. When the graph path corresponding to the candidate action token set has overlapping nodes or overlapping edges with the forbidden token path, the graph path that matches the action direction label value of the candidate action token set and has the smallest overdraft amount after being filled is searched along the graph edge. The candidate action token set is then replaced with the corresponding fillable action token set. When the graph path corresponding to the candidate action token set does not have overlapping nodes or overlapping edges with the forbidden token path, the candidate action token set is retained.
8. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 7, characterized in that: The specific steps for performing forbidden path replacement and power sorting on the candidate action token group to obtain the target scheduling action token group are as follows: The retained candidate action token groups and the replacement replenishable action token groups are combined with the current HVAC state sequence to form a prediction input vector. The prediction input vector is then input into the multilayer perceptron, and the system prediction total power value corresponding to each action token group is output. The replenished token overdraft amount, feedback idle token flag value, and blocking node connection result corresponding to each action token group are read. Action token groups that meet any of the following conditions are eliminated: feedback idle token flag value, connection to blocking node, or replenished token overdraft amount exceeding the overdraft tolerance threshold. If there are remaining action token groups after elimination, select the action token group with the smallest predicted total power value from the remaining action token groups as the target scheduling action token group; if there are no remaining action token groups after elimination, write the scheduling instruction field, action direction flag value and action amplitude value corresponding to the current control instruction number into the holding action token group, and use the holding action token group as the target scheduling action token group.
9. The HVAC energy consumption optimization scheduling method based on neural networks according to claim 8, characterized in that: The specific steps for converting the target scheduling action token group into a HVAC energy consumption optimization scheduling instruction and issuing it, and outputting the HVAC energy consumption optimization scheduling result are as follows: HVAC energy consumption optimization scheduling instructions are generated according to the scheduling instruction field, action direction flag value, and action amplitude value in the target scheduling action token group; for the area indoor temperature, chilled water supply temperature, chilled water pump frequency, cooling water pump frequency, and cooling tower fan frequency, the scheduling instruction setting value in the current control instruction number that is the same as the scheduling instruction field of the target scheduling action token is added to or subtracted from the corresponding action amplitude value according to the action direction flag value to obtain the corresponding target setting value; for chiller unit start-up and shutdown, the scheduling instruction chiller unit start-up and shutdown flag value corresponding to the current control instruction number is updated according to the action direction flag value to obtain the target chiller unit start-up and shutdown flag value; The target setpoints and target chiller start / stop flags are combined to generate HVAC energy consumption optimization scheduling instructions. New control instruction numbers are assigned to the HVAC energy consumption optimization scheduling instructions and sent to the main controller of the building automation system. The system receives and writes back the execution feedback data bound to the new control instruction number and outputs the HVAC energy consumption optimization scheduling results.
10. A neural network-based HVAC energy consumption optimization and scheduling system, characterized in that, include: The system includes a HVAC data alignment processing module, an action overdraft token generation module, a neural overdraft atlas construction module, and a mask scheduling instruction output module, among which: The HVAC data alignment processing module is used to collect HVAC energy consumption scheduling and monitoring data, perform time alignment, abnormal record removal and missing data compensation processing on the HVAC energy consumption scheduling and monitoring data, and output preprocessed HVAC energy consumption scheduling and monitoring data. The action overdraft token generation module is used to convert the scheduling instruction field corresponding to the adjacent control instruction number into a scheduling action token based on the preprocessed HVAC energy consumption scheduling monitoring data, and to generate the token overdraft amount, token replenishment relationship flag value and replenished token overdraft amount according to the execution feedback data corresponding to the scheduling action token. The neural overdraft map construction module is used to generate overdraft embedding vectors based on scheduling action tokens, token overdraft amount, token overdraft amount after replenishment, and token replenishment relationship marker values, and to construct a neural network action overdraft token map by combining the overdraft embedding vectors, generating prohibited token paths and replenishable token paths. The mask scheduling instruction output module is used to generate candidate action token groups based on the current HVAC state sequence and the neural network action overdraft token map, perform forbidden path replacement and power sorting on the candidate action token groups to obtain the target scheduling action token group, convert the target scheduling action token group into HVAC energy consumption optimization scheduling instructions and issue them, and output the HVAC energy consumption optimization scheduling results.
Citation Information
Patent Citations
An energy storage management system for building energy conservation
CN120278848B
Micro-grid-heating ventilation air conditioner coordinated optimization method based on deep reinforcement learning
CN120975528A