An air conditioning unit control method, equipment and medium
By constructing a control method for air conditioning units and utilizing multi-source time-series data and dynamic control strategies, the problem of traditional air conditioning systems being unable to adapt to complex environments has been solved, achieving improved energy efficiency and stability while ensuring indoor comfort.
Patent Information
- Application Number
- CN202510180930.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Traditional air conditioning system control methods are difficult to fully adapt to complex and ever-changing indoor environmental factors, resulting in energy waste and poor indoor comfort. They also lack effective use of historical data and cannot achieve efficient energy-saving control.
By collecting multi-source time-series data from air conditioning units, and using a preset time window slicing algorithm and noise injection algorithm to process the preprocessed multi-source time-series data, an enhanced dataset is generated. A Mamba time-series prediction model is constructed, and the enhanced dataset is processed based on the Mamba time-series prediction model. A comprehensive reward function and a noise injection algorithm are defined to generate a dynamic control strategy. Combined with the SAC algorithm of the maximum entropy framework and the dual-Q network architecture, the equipment control command with the lowest energy consumption is output.
This improved the air conditioning system, enhancing energy efficiency and operational stability while ensuring indoor environmental comfort, thus achieving efficient energy utilization.
Smart Images

Figure CN119860584B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of air conditioning unit control, and in particular to an air conditioning unit control method, equipment and medium. Background Technology
[0002] With the continuous advancement of modern building technology and people's increasing demands for indoor environmental comfort, air conditioning systems, as key equipment for regulating indoor temperature, humidity, and air quality, have seen their energy consumption account for over 30% of total building energy consumption. Traditional air conditioning system control methods are mostly based on simple sensor feedback and preset rules, such as starting or stopping heating / cooling based on the deviation between the temperature detected by the temperature sensor and the set value. However, this method has significant limitations. On the one hand, the indoor environment is a complex dynamic system influenced by multiple factors, including changes in outdoor weather, indoor occupant activity, and air humidity. Simple rules are insufficient to fully adapt to these complex and ever-changing environmental factors, often leading to energy waste and poor indoor comfort. On the other hand, traditional control methods lack effective utilization of historical data, making it impossible to proactively predict the operating energy consumption of the air conditioning system and achieve efficient and energy-saving control.
[0003] Therefore, how to improve the energy efficiency of air conditioning unit systems while ensuring indoor environmental comfort has become an urgent technical problem to be solved. Summary of the Invention
[0004] This application provides an air conditioning unit control method, equipment, and medium to solve the following technical problem: how to improve the energy efficiency of the air conditioning unit system while ensuring indoor environmental comfort.
[0005] In a first aspect, embodiments of this application provide an air conditioning unit control method, applied to the control system of an air conditioning unit. The method includes: collecting multi-source time-series data of the air conditioning unit; wherein the multi-source time-series data includes environmental data, equipment operation data, and energy consumption data, and the environmental data includes internal environmental data and external environmental data; processing the preprocessed multi-source time-series data based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset; constructing a Mamba time-series prediction model and processing the enhanced dataset based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operation parameters of the air conditioning unit in a first time period; defining a comprehensive reward function and processing the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value; processing the original comprehensive reward value based on a reward clustering algorithm to output a normalized comprehensive reward signal; processing the augmented data based on a preset maximum entropy framework SAC algorithm and combined with a preset dual-Q network architecture to generate a dynamic control strategy; the augmented data includes environmental data, equipment operation data, energy consumption data, and predicted energy consumption value; and outputting equipment control commands under the lowest energy consumption based on the dynamic control strategy.
[0006] In one implementation of this application, the collection of multi-source time-series data from an air conditioning unit specifically includes: constructing an edge server and multiple sensor networks; wherein the edge server is connected to multiple sensor networks; and monitoring the temperature and humidity of the service area of the air conditioning unit based on the sensor networks, monitoring the operating data and energy consumption data of the air conditioning unit to generate multi-source time-series data.
[0007] In one implementation of this application, preprocessed multi-source time-series data is processed based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset. Specifically, this includes: cleaning, denoising, and standardizing the multi-source time-series data to generate preprocessed multi-source time-series data; slicing the preprocessed multi-source time-series data according to a preset time window size to generate multiple time segments; injecting a preset level of random noise into the multiple time segments; and combining the processed time segments to generate an enhanced dataset.
[0008] In one implementation of this application, a Mamba time-series prediction model is constructed, and an augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period. Specifically, this includes: constructing a Mamba time-series prediction model based on a preset Transformer architecture; inputting the augmented dataset into the Mamba time-series prediction model; and analyzing the time-series data based on the linear layer, convolutional layer, gating mechanism layer, and state-space model (SSM) layer inside the Mamba time-series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period.
[0009] In one implementation of this application, a comprehensive reward function is defined, and real-time energy consumption values are processed based on the comprehensive reward function to output an original comprehensive reward value. Specifically, the comprehensive reward function is defined as a weighted sum of energy efficiency and indoor comfort; wherein, energy efficiency is a negative reward, and energy efficiency is inversely proportional to the predicted energy consumption value, and indoor comfort is a positive reward, and the deviation of indoor temperature from the set value is inversely proportional to indoor comfort; the real-time energy consumption value and indoor temperature data are input into the comprehensive reward function to calculate the original comprehensive reward value.
[0010] In one implementation of this application, the original comprehensive reward value is processed based on a reward clustering algorithm to output a normalized comprehensive reward signal. Specifically, this includes: calculating the moving average of the original comprehensive reward value; wherein the initial moving average is zero; adjusting the update rate of the moving average based on a dynamic step size parameter; wherein the initial value of the dynamic step size parameter is a non-zero positive number and decays exponentially with the training process; and subtracting the current moving average from the original comprehensive reward value to generate a normalized comprehensive reward signal centered on the zero mean.
[0011] In one implementation of this application, a SAC algorithm based on a preset maximum entropy framework is used, combined with a preset dual-Q network architecture to process augmented data to generate a dynamic control strategy. The augmented data includes environmental data, equipment operation data, energy consumption data, and predicted energy consumption values. Specifically, this includes: constructing a SAC algorithm based on a maximum entropy framework and constructing a dual-Q network architecture; wherein the dual-Q network architecture includes two independent Q networks; when updating the SAC algorithm, the output of the network with the smaller estimated value among the two independent Q networks is selected as the target value to update the SAC algorithm; the augmented signal is input to the SAC algorithm to generate the dynamic control strategy. In another implementation of this application, the method further includes: analyzing the equipment parameter adjustment range in the dynamic control strategy; monitoring the equipment operating status and energy consumption deviation in real time through a closed-loop feedback mechanism; if the actual energy consumption deviates from the predicted energy consumption value by more than a preset threshold, then adjusting the Mamba time-series prediction model.
[0012] Secondly, embodiments of this application also provide an air conditioning unit control device, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: collect multi-source time-series data of the air conditioning unit; wherein the multi-source time-series data includes environmental data, equipment operation data, and energy consumption data, the environmental data including internal environmental data and external environmental data; process the preprocessed multi-source time-series data based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset; and construct a Mamb... A time-series prediction model is used, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period. A comprehensive reward function is defined, and the real-time energy consumption value is processed based on the comprehensive reward function to output the original comprehensive reward value. The original comprehensive reward value is processed based on the reward clustering algorithm to output the normalized comprehensive reward signal. The SAC algorithm based on the preset maximum entropy framework is combined with the preset dual-Q network architecture to process the augmented data to generate a dynamic control strategy. The augmented data includes environmental data, equipment operating data, energy consumption data, and energy consumption prediction value. Based on the dynamic control strategy, the equipment control command under the lowest energy consumption is output.
[0013] Thirdly, embodiments of this application also provide a non-volatile computer storage medium for air conditioning unit control, storing computer-executable instructions, characterized in that the computer-executable instructions are configured as follows: collecting multi-source time-series data of the air conditioning unit; wherein the multi-source time-series data includes environmental data, equipment operation data, and energy consumption data, and the environmental data includes internal environmental data and external environmental data; processing the preprocessed multi-source time-series data based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset; constructing a Mamba time-series prediction model and processing the enhanced dataset based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operation parameters of the air conditioning unit in a first time period; defining a comprehensive reward function and processing the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value; processing the original comprehensive reward value based on a reward clustering algorithm to output a normalized comprehensive reward signal; processing the augmented data based on a preset maximum entropy framework SAC algorithm and combined with a preset dual-Q network architecture to generate a dynamic control strategy; the augmented data includes environmental data, equipment operation data, energy consumption data, and predicted energy consumption values; and outputting equipment control instructions at the lowest energy consumption based on the dynamic control strategy.
[0014] This application provides an air conditioning unit control method, device, and medium. By collecting multi-source time-series data from the air conditioning unit, the operating status of the air conditioning system can be understood. A preset time window slicing algorithm and noise injection algorithm are used to enhance the preprocessed data, improving its adaptability to complex operating conditions and random fluctuations, generating a high-quality enhanced dataset. Based on the constructed Mamba time-series prediction model, the enhanced dataset is analyzed to predict the energy consumption and equipment operating parameters of the air conditioning unit in the first time period, providing a basis for the formulation of control strategies. By defining a comprehensive reward function and combining it with a reward clustering algorithm, real-time energy consumption values are processed, and a normalized comprehensive reward signal is output, which can improve the stability and convergence performance of the reinforcement learning algorithm to a certain extent. Based on the SAC algorithm of the maximum entropy framework and a preset dual-Q network architecture, dynamic control strategies suitable for different operating conditions are generated, and the lowest energy consumption equipment control commands are output while ensuring indoor comfort. This effectively improves the operating energy efficiency, stability, and comfort of the air conditioning system, achieving efficient energy utilization. Attached Figure Description
[0015] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0016] Figure 1 A flowchart of an air conditioning unit control method provided in an embodiment of this application;
[0017] Figure 2 This is a schematic diagram of the internal structure of an air conditioning unit control device provided in an embodiment of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] This application provides an air conditioning unit control method, equipment, and medium to solve the following technical problem: how to improve the energy efficiency of the air conditioning unit system while ensuring indoor environmental comfort.
[0020] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0021] Figure 1This is a control flowchart for an air conditioning unit provided as an embodiment of this application. Figure 1 As shown in the figure, an air conditioning unit control method provided in this application embodiment specifically includes the following steps:
[0022] Step 1: Collect multi-source time-series data of the air conditioning unit; the multi-source time-series data includes environmental data, equipment operation data and energy consumption data, and the environmental data includes internal environmental data and external environmental data.
[0023] Internal environmental data refers to data such as temperature and humidity within the service area of the air conditioning unit that directly affect indoor comfort. This data is monitored in real time by sensors installed within the service area.
[0024] External environmental data refers to external environmental factors that affect the operating efficiency of air conditioning units, such as outdoor temperature, humidity, wind speed, and solar radiation.
[0025] Equipment operation data refers to the operating status parameters of various components of the air conditioning unit (such as compressor, fan, condenser), such as speed, current and voltage.
[0026] Energy consumption data refers to the energy consumption of air conditioning units during operation, such as electrical energy consumption. This data is collected through electricity metering devices (such as electricity meters).
[0027] Step 11: Build an edge server and multiple sensor networks; wherein the edge server is connected to multiple sensor networks.
[0028] An edge server is a computing device located at the edge of a network and connected to multiple sensor networks.
[0029] Sensor networks include temperature sensors, humidity sensors, wind speed and direction sensors, solar radiation sensors, equipment operation status sensors, and electricity metering equipment.
[0030] Temperature sensor: Used to measure the temperature within the service area and the external ambient temperature.
[0031] Humidity sensor: Used to measure humidity within the service area and the humidity of the external environment.
[0032] Wind speed and direction sensor: used to measure wind speed and direction in the external environment.
[0033] Solar radiation sensor: Used to measure the intensity of solar radiation.
[0034] Equipment operating status sensors: such as current sensors and voltage sensors, are used to monitor the operating status parameters of various components of the air conditioning unit.
[0035] Electricity metering equipment, such as smart meters, is used to measure the electricity consumption of air conditioning units.
[0036] Step 12: Based on sensor network, monitor the temperature and humidity of the area served by the air conditioning unit, monitor the operation data and energy consumption data of the air conditioning unit, and generate multi-source time-series data.
[0037] Temperature and humidity data within the service area are monitored in real time by installing temperature and humidity sensors within the service area of the air conditioning unit.
[0038] Multiple temperature sensors are evenly deployed within the service area, and multiple humidity sensors are deployed in coordination.
[0039] The operating status parameters of the air conditioning unit are monitored in real time by using operating status sensors (such as current sensors and voltage sensors) installed on various components of the air conditioning unit.
[0040] The energy consumption of the air conditioning unit can be monitored in real time by installing an energy metering device (such as a smart meter) at the power input terminal of the air conditioning unit.
[0041] The collected energy consumption data is transmitted to the edge server for storage in real time.
[0042] Step 2: Process the preprocessed multi-source time series data based on the preset time window slicing algorithm and noise injection algorithm to generate an enhanced dataset.
[0043] The time window slicing algorithm is a commonly used time series data processing technique that aims to slice continuous time series data according to a preset time window size, thereby generating multiple time segments.
[0044] The time window refers to the length of the time period used for slicing. In this invention, the size of the time window is preset according to the operating characteristics and data features of the air conditioning unit.
[0045] To fully utilize the information in the data, a sliding window technique can be used. This involves sliding a time window across the data, moving it by a fixed step each time, thus generating multiple overlapping time segments. This method helps capture long-term dependencies and dynamic changes in the data.
[0046] Noise injection is a technique that enhances the robustness of a model by injecting random noise into the data. In time series data, noise injection helps simulate measurement errors and random fluctuations in the real environment, improving the model's tolerance to noise and its prediction accuracy.
[0047] Step 21: Clean, denoise, and standardize the multi-source time series data to generate preprocessed multi-source time series data.
[0048] Existing cleaning algorithms will be used, which will not be elaborated here.
[0049] Step 22: Slice the preprocessed multi-source time series data according to the preset time window size to generate multiple time segments.
[0050] The choice of time window size needs to be comprehensively considered based on the operating characteristics of the air conditioning unit and the data features. An excessively large time window may smooth out local features and dynamic changes in the data, while an excessively small time window may fail to fully capture long-term dependencies in the data.
[0051] Therefore, cross-validation is used to compare and evaluate different time window sizes, and the optimal time window size is selected to determine the time window size.
[0052] The preprocessed multi-source time series data is sliced according to a determined time window size to generate multiple time segments.
[0053] Step 23: Inject a preset level of random noise into multiple time segments, and combine the processed time segments to generate an enhanced dataset.
[0054] A preset level of random noise is injected into each time segment to generate a noisy time segment.
[0055] The noisy time segments are combined in the order of the original data to generate a complete augmented dataset.
[0056] In a specific example, large data center A has multiple air conditioning units responsible for providing a stable temperature and humidity environment for server equipment. To improve the energy efficiency and stability of the air conditioning units, the data center decided to adopt the control method and equipment of this invention.
[0057] Multiple sensors (such as temperature sensors, humidity sensors, and current sensors) are installed in the service area of the air conditioning unit to monitor and collect multi-source time-series data of the air conditioning unit in real time.
[0058] The collected multi-source time-series data were cleaned, denoised, and standardized to generate preprocessed multi-source time-series data. Specifically, interpolation was used to handle missing values, low-pass filtering was used to remove noise, and Z-score standardization was used for data standardization.
[0059] The preprocessed multi-source time-series data is sliced according to a preset time window size (e.g., 0.5 hours) to generate multiple time segments. A sliding window technique is used to generate overlapping time segments to fully utilize the information in the data.
[0060] Inject a preset level of random noise (such as 10% of the data standard deviation) into each time segment to generate noisy time segments.
[0061] The noisy time segments are combined in the order of the original data to generate a complete augmented dataset. The augmented dataset is then stored in cloud storage for subsequent model training and testing.
[0062] Step 3: Construct a Mamba time-series forecasting model and process the augmented dataset based on the Mamba time-series forecasting model to output the predicted energy consumption of the air conditioning unit and the equipment operating parameters in the first time period.
[0063] Step 31: Build a Mamba time series prediction model based on the preset Transformer architecture.
[0064] The Mamba time series forecasting model is a selective state-space model that incorporates the Transformer architecture and is specifically designed for processing time series data.
[0065] Step 32: Input the augmented dataset into the Mamba time series prediction model, and analyze the time series data based on the linear layer, convolutional layer, gating mechanism layer and state space model (SSM) layer inside the Mamba time series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period.
[0066] Step 4: Define the comprehensive reward function and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value.
[0067] Step 41: Define the comprehensive reward function as a weighted sum of energy efficiency and indoor comfort; where energy efficiency is a negative reward, and energy efficiency is inversely proportional to the predicted energy consumption value, and indoor comfort is a positive reward, and the deviation of indoor temperature from the set value is inversely proportional to indoor comfort.
[0068] Energy efficiency is a crucial indicator for measuring the energy effectiveness of an air conditioning system. It is the ratio of energy consumed to cooling / heating capacity provided per unit time. In the comprehensive reward function, energy efficiency is set as a negative reward, meaning that the higher the energy efficiency (the lower the energy consumption), the larger the reward value (the smaller the negative value). Energy efficiency is inversely proportional to the predicted energy consumption value; that is, the lower the predicted energy consumption value, the higher the energy efficiency, and the larger the reward value.
[0069] Indoor comfort is an important indicator for measuring indoor environmental quality, and it is a comprehensive effect of multiple parameters such as indoor temperature, humidity, and airflow. In the comprehensive reward function, indoor comfort is set as a positive reward, meaning that the higher the indoor comfort, the greater the reward value. The deviation of indoor temperature from the set value is inversely proportional to indoor comfort; that is, the closer the indoor temperature is to the set value, the higher the indoor comfort and the greater the reward value.
[0070] To balance energy efficiency and indoor comfort, this invention uses a weighted sum to combine both into a comprehensive reward function. The weighting coefficients represent the relative importance of each objective and can be adjusted according to actual needs.
[0071] Step 42: Input the real-time energy consumption value and indoor temperature data into the comprehensive reward function to calculate the original comprehensive reward value.
[0072] By substituting real-time energy consumption and indoor temperature data into the mathematical expression of the comprehensive reward function, the original comprehensive reward value can be calculated.
[0073] Step 5: Process the original comprehensive reward value based on the reward clustering algorithm to output a normalized comprehensive reward signal.
[0074] Step 51: Calculate the moving average of the original total reward value; where the initial moving average is zero.
[0075] A moving average is a method used to smooth out data fluctuations by calculating the average value of data over a period of time to reflect the overall trend of the data.
[0076] The original comprehensive reward value may fluctuate due to factors such as environmental noise and model uncertainty, which can affect the model's learning efficiency and stability. Therefore, a moving average is used to smooth the original comprehensive reward value.
[0077] The moving average is a method of calculating the average value of data within a fixed-size window. As the window moves, new data enters and older data leaves, resulting in a series of averages. These averages reflect the overall trend of the data over a period of time, effectively reducing the impact of data fluctuations on model learning.
[0078] During the initialization phase, the moving average is zero. As the model interacts with the environment, the initial total reward value is continuously generated, and the moving average is gradually updated to reflect the reward trend under the current environmental conditions.
[0079] Step 52: Adjust the update rate of the moving average based on the dynamic step size parameter; wherein the initial value of the dynamic step size parameter is a non-zero positive number, and decays exponentially with the training process.
[0080] The update rate of the moving average determines its sensitivity to fluctuations in the original total reward value. An update rate that is too fast may prevent the moving average from effectively smoothing data fluctuations, while an update rate that is too slow may prevent the moving average from reflecting changes in environmental conditions in a timely manner. Therefore, this invention uses a dynamic step size parameter to adjust the update rate of the moving average.
[0081] The dynamic step size parameter is a time-varying coefficient used to control the update rate of the moving average. It starts as a non-zero positive number and decays exponentially as training progresses. This means that in the early stages of training, the moving average updates quickly, rapidly reflecting changes in the environment; as training continues, the update rate gradually slows down, making the moving average more stable.
[0082] Exponential decay is a commonly used method for adjusting parameter values, which controls the rate of change of the parameter value by introducing a decay factor. In this invention, the dynamic step size parameter decays exponentially with the training process to achieve dynamic adjustment of the moving average update rate.
[0083] Step 53: Subtract the current moving average from the original comprehensive reward value to generate a normalized comprehensive reward signal centered at the zero mean.
[0084] The moving average, after smoothing and dynamic step size parameter adjustment, can reflect the overall trend of the original comprehensive reward value. To generate a more stable reward signal, this invention subtracts the current moving average from the original comprehensive reward value to obtain a normalized comprehensive reward signal centered at zero mean.
[0085] The normalized composite reward signal is the difference between the original composite reward value and its moving average, representing the degree of deviation of the current reward value from the overall trend. Since the moving average reflects the overall trend of the reward value, the normalized composite reward signal can eliminate the influence of environmental noise and model uncertainty on the reward value, allowing the model to focus more on learning effective control strategies.
[0086] For example, Commercial Complex B has multiple air conditioning units responsible for providing stable temperature and humidity environments for different areas. To improve the energy efficiency and intelligence of these air conditioning units, the commercial complex decided to adopt the control method and equipment of this invention.
[0087] The sliding window size is set to 10 (i.e., N=10), and the initial sliding average is zero. As the model interacts with the environment, the original comprehensive reward value is continuously generated, and the sliding average is gradually updated. For example, in the first 10 time steps, the sliding average is the average of the original comprehensive reward value at each time step; from the 11th time step onwards, the sliding average is dynamically updated according to the formula.
[0088] The initial dynamic step size parameter is set to 0.1, and the decay factor is set to 0.99. The update rate is adjusted using the current dynamic step size parameter each time the moving average is updated. As training progresses, the dynamic step size parameter gradually decreases, and the update rate of the moving average also slows down.
[0089] The normalized comprehensive reward signal is obtained by subtracting the current moving average from the original comprehensive reward value. For example, assuming the current original comprehensive reward value is 1.5 and the moving average is 1.2, the normalized comprehensive reward signal R is 1.5 − 1.2 = 0.3.
[0090] Step 6: Based on the preset maximum entropy framework, the SAC algorithm is combined with the preset dual-Q network architecture to process the extended data to generate a dynamic control strategy; the extended data includes environmental data, equipment operation data, energy consumption data and energy consumption prediction values.
[0091] Step 61: Construct the SAC algorithm of the maximum entropy framework and construct the dual-Q network architecture; wherein, the dual-Q network architecture includes two independent Q networks.
[0092] SAC (Soft Actor-Critic) is a reinforcement learning algorithm based on the principle of maximum entropy. It balances exploration and exploitation by maximizing cumulative reward and policy entropy. In SAC, the policy objective is not only to maximize cumulative reward but also to maximize policy entropy, thereby enabling the model to explore more of the state-action space and avoid premature convergence to local optima.
[0093] The dual-Q network architecture is used to mitigate the problem of Q-value overestimation. In this architecture, two independent Q-networks are used, each estimating the Q-value of a state-action pair. During policy updates, the output of the Q-network with the smaller estimate is selected as the target value for updating, thus avoiding policy bias caused by overestimation from a single Q-network. This architecture improves the stability and performance of the algorithm.
[0094] Step 62: When updating the SAC algorithm, select the output of the network with the smaller estimate from the two independent Q networks as the target value to update the SAC algorithm.
[0095] In each training iteration, for each state-action pair, its Q-value estimate is computed using both the Q1 and Q2 networks. Then, the smaller of the two Q-values is selected as the target Q-value to update the value network in the SAC algorithm.
[0096] Step 63: Input the extended signal into the SAC algorithm to generate a dynamic control strategy.
[0097] After the SAC algorithm is trained, an augmented signal is input into the SAC algorithm, which then generates a corresponding control strategy based on the augmented signal. This control strategy is dynamic and can be adjusted in real time according to changes in the augmented signal to achieve optimal control performance.
[0098] For example, the SAC algorithm can adjust the control strategy of air conditioning equipment in real time based on indoor environmental data such as temperature and humidity, as well as the operating status and energy consumption data of the equipment, in order to achieve the goal of energy saving and consumption reduction.
[0099] Step 7: Output equipment control commands with the lowest energy consumption based on the dynamic control strategy.
[0100] The dynamic control strategy, generated by the SAC algorithm based on a normalized comprehensive reward signal, aims to achieve efficient operation of the air conditioning unit with minimal energy consumption. This strategy includes parameter adjustment suggestions for various components within the air conditioning unit, such as the frequencies of the chilled water pump and cooling pump, the frequency of the cooling tower fan, and the outlet water temperatures of the main unit and cooling tower.
[0101] After analyzing the equipment parameter adjustment range in the dynamic control strategy, specific equipment control commands are output based on the strategy recommendations. These commands will directly affect each piece of equipment in the air conditioning unit to achieve efficient operation with minimal energy consumption.
[0102] Based on the parameter adjustment suggestions and ranges in the dynamic control strategy, specific equipment control commands are generated. These commands include adjusting the frequency of chilled water pumps and cooling pumps, changing the frequency of cooling tower fans, and adjusting the outlet water temperature of the main unit and cooling tower.
[0103] The generated control commands are sent to the air conditioning unit's control system, which then executes the corresponding equipment adjustment operations. During this process, the control system monitors the equipment's operating status and energy consumption in real time.
[0104] The method also includes: analyzing the adjustment range of equipment parameters in the dynamic control strategy; monitoring the equipment operating status and energy consumption deviation in real time through a closed-loop feedback mechanism; and adjusting the Mamba time-series prediction model if the actual energy consumption deviates from the predicted energy consumption value by more than a preset threshold.
[0105] In one example: parameter adjustment suggestions for each device are extracted from the dynamic control strategy, and a reasonable adjustment range is determined for each parameter based on the device's physical limitations and historical operating data. For instance, the frequency adjustment range for the chilled pump is set between 30Hz and 60Hz.
[0106] Based on the parameter adjustment suggestions and ranges in the dynamic control strategy, specific equipment control commands are generated and sent to the air conditioning unit's control system for execution. For example, the frequency of the chilled water pump can be adjusted to 45Hz to reduce energy consumption.
[0107] The system collects real-time operating status and energy consumption data of the air conditioning unit using sensors and monitoring equipment, and feeds this data back to the SAC algorithm and the Mamba time-series prediction model. It calculates the deviation between actual and predicted energy consumption and determines whether the deviation exceeds a preset threshold (e.g., 5%).
[0108] When an energy consumption deviation exceeding a preset threshold is detected (e.g., actual energy consumption is more than 5% higher than predicted energy consumption), the Mamba time-series prediction model is adjusted. Adjustment methods include retraining the model and updating model parameters. By adjusting the model, its accuracy and adaptability in predicting the operating status of air conditioning units are improved, and the dynamic control strategy generated by the SAC algorithm is optimized to achieve efficient and stable operation with minimal energy consumption.
[0109] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide an air conditioning unit control device, the structure of which is as follows: Figure 2 As shown.
[0110] Figure 2 This is a schematic diagram of the internal structure of an air conditioning unit control device provided in an embodiment of this application. Figure 2 As shown, the device includes:
[0111] At least one processor 201;
[0112] And a memory 202 that is communicatively connected to at least one processor;
[0113] The memory 202 stores instructions executable by at least one processor, which are executed by at least one processor 201 to enable at least one processor 201 to:
[0114] Multi-source time-series data of the air conditioning unit is collected, including environmental data, equipment operation data, and energy consumption data. The environmental data includes both internal and external environmental data. The pre-processed multi-source time-series data is processed using a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset. A Mamba time-series prediction model is constructed, and the enhanced dataset is processed based on this model to output the predicted energy consumption and equipment operation parameters of the air conditioning unit within the first time period. A comprehensive reward function is defined, and real-time energy consumption values are processed based on this function to output the original comprehensive reward value. The original comprehensive reward value is processed using a reward clustering algorithm to output a normalized comprehensive reward signal. The SAC algorithm based on a preset maximum entropy framework, combined with a preset dual-Q network architecture, is used to process the augmented data to generate a dynamic control strategy. The augmented data includes environmental data, equipment operation data, energy consumption data, and predicted energy consumption values. Based on the dynamic control strategy, equipment control commands at the lowest energy consumption are output.
[0115] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium for air conditioning unit control, storing computer-executable instructions, wherein the computer-executable instructions are configured as follows:
[0116] Multi-source time-series data of the air conditioning unit is collected, including environmental data, equipment operation data, and energy consumption data. The environmental data includes both internal and external environmental data. The pre-processed multi-source time-series data is processed using a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset. A Mamba time-series prediction model is constructed, and the enhanced dataset is processed based on this model to output the predicted energy consumption and equipment operation parameters of the air conditioning unit within the first time period. A comprehensive reward function is defined, and real-time energy consumption values are processed based on this function to output the original comprehensive reward value. The original comprehensive reward value is processed using a reward clustering algorithm to output a normalized comprehensive reward signal. The SAC algorithm based on a preset maximum entropy framework, combined with a preset dual-Q network architecture, is used to process the augmented data to generate a dynamic control strategy. The augmented data includes environmental data, equipment operation data, energy consumption data, and predicted energy consumption values. Based on the dynamic control strategy, equipment control commands at the lowest energy consumption are output.
[0117] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0118] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.
[0119] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0120] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0121] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0123] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0124] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0125] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0126] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An air conditioning unit control method, applied to the control system of an air conditioning unit, characterized in that, The method includes: Collect multi-source time-series data of the air conditioning unit; wherein, the multi-source time-series data includes environmental data, equipment operation data and energy consumption data, and the environmental data includes internal environmental data and external environmental data; The preprocessed multi-source time-series data is processed based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset; A Mamba time-series prediction model is constructed, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period. Define a comprehensive reward function, and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value; The original comprehensive reward value is processed based on the reward clustering algorithm to output a normalized comprehensive reward signal; The SAC algorithm, based on a preset maximum entropy framework, is combined with a preset dual-Q network architecture to process augmented data in order to generate a dynamic control strategy; the augmented data includes environmental data, equipment operation data, energy consumption data, and energy consumption prediction values. Based on the dynamic control strategy, output equipment control commands with the lowest energy consumption. A Mamba time-series prediction model is constructed, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption and equipment operating parameters of the air conditioning unit in the first time period, specifically including: A Mamba time series prediction model is built based on a pre-defined Transformer architecture; The augmented dataset is input into the Mamba time series prediction model, and the time series data is analyzed based on the linear layer, convolutional layer, gating mechanism layer and state space model (SSM) layer inside the Mamba time series prediction model to output the predicted energy consumption value of the air conditioning unit and the equipment operating parameters in the first time period. Define a comprehensive reward function, and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value, specifically including: The comprehensive reward function is defined as a weighted sum of energy efficiency and indoor comfort; wherein, energy efficiency is a negative reward, and energy efficiency is inversely proportional to the predicted energy consumption value, and indoor comfort is a positive reward, and the deviation of indoor temperature from the set value is inversely proportional to indoor comfort. The real-time energy consumption value and indoor temperature data are input into the comprehensive reward function to calculate the original comprehensive reward value. The original comprehensive reward value is processed based on a reward clustering algorithm to output a normalized comprehensive reward signal, specifically including: Calculate the moving average of the original total reward value; where the initial moving average is zero; The update rate of the moving average is adjusted based on the dynamic step size parameter; wherein the initial value of the dynamic step size parameter is a non-zero positive number, and decays exponentially with the training process; Subtract the current moving average from the original comprehensive reward value to generate a normalized comprehensive reward signal centered on the zero mean. The SAC algorithm, based on a pre-defined maximum entropy framework, is combined with a pre-defined dual-Q network architecture to process augmented data, thereby generating a dynamic control strategy. The augmented data includes environmental data, equipment operation data, energy consumption data, and energy consumption prediction values, specifically including: The SAC algorithm with a maximum entropy framework is constructed, and a dual-Q network architecture is built; wherein, the dual-Q network architecture includes two independent Q networks; When updating the SAC algorithm, the output of the network with the smaller estimate among the two independent Q networks is selected as the target value to update the SAC algorithm. The expanded data is input into the SAC algorithm to generate a dynamic control strategy.
2. The air conditioning unit control method according to claim 1, characterized in that, The collection of multi-source time-series data from the air conditioning unit specifically includes: An edge server and multiple sensor networks are constructed; wherein the edge server is connected to the multiple sensor networks. Based on the sensor network, the temperature and humidity of the area served by the air conditioning unit are monitored, as well as the operating data and energy consumption data of the air conditioning unit, to generate multi-source time-series data.
3. The air conditioning unit control method according to claim 1, characterized in that, The preprocessed multi-source time-series data is processed based on a preset time window slicing algorithm and a noise injection algorithm to generate an augmented dataset, specifically including: The multi-source time-series data is cleaned, denoised, and standardized to generate preprocessed multi-source time-series data. The preprocessed multi-source time series data is sliced according to a preset time window size to generate multiple time segments; A preset level of random noise is injected into multiple time segments, and the processed time segments are combined to generate an enhanced dataset.
4. The air conditioning unit control method according to claim 1, characterized in that, The method further includes: Analyze the adjustment range of equipment parameters in dynamic control strategies; The system monitors the equipment's operating status and energy consumption deviation in real time through a closed-loop feedback mechanism. If the actual energy consumption deviates from the predicted energy consumption value by more than a preset threshold, the Mamba time series prediction model is adjusted.
5. An air conditioning unit control device, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Collect multi-source time-series data of the air conditioning unit; wherein, the multi-source time-series data includes environmental data, equipment operation data and energy consumption data, and the environmental data includes internal environmental data and external environmental data; The preprocessed multi-source time-series data is processed based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset; A Mamba time-series prediction model is constructed, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period. Define a comprehensive reward function, and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value; The original comprehensive reward value is processed based on the reward clustering algorithm to output a normalized comprehensive reward signal; The SAC algorithm, based on a preset maximum entropy framework, is combined with a preset dual-Q network architecture to process augmented data in order to generate a dynamic control strategy; the augmented data includes environmental data, equipment operation data, energy consumption data, and energy consumption prediction values. Based on the dynamic control strategy, output equipment control commands with the lowest energy consumption. A Mamba time-series prediction model is constructed, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption and equipment operating parameters of the air conditioning unit in the first time period, specifically including: A Mamba time series prediction model is built based on a pre-defined Transformer architecture; The augmented dataset is input into the Mamba time series prediction model, and the time series data is analyzed based on the linear layer, convolutional layer, gating mechanism layer and state space model (SSM) layer inside the Mamba time series prediction model to output the predicted energy consumption value of the air conditioning unit and the equipment operating parameters in the first time period. Define a comprehensive reward function, and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value, specifically including: The comprehensive reward function is defined as a weighted sum of energy efficiency and indoor comfort; wherein, energy efficiency is a negative reward, and energy efficiency is inversely proportional to the predicted energy consumption value, and indoor comfort is a positive reward, and the deviation of indoor temperature from the set value is inversely proportional to indoor comfort. The real-time energy consumption value and indoor temperature data are input into the comprehensive reward function to calculate the original comprehensive reward value. The original comprehensive reward value is processed based on a reward clustering algorithm to output a normalized comprehensive reward signal, specifically including: Calculate the moving average of the original total reward value; where the initial moving average is zero; The update rate of the moving average is adjusted based on the dynamic step size parameter; wherein the initial value of the dynamic step size parameter is a non-zero positive number, and decays exponentially with the training process; Subtract the current moving average from the original comprehensive reward value to generate a normalized comprehensive reward signal centered on the zero mean. The SAC algorithm, based on a pre-defined maximum entropy framework, is combined with a pre-defined dual-Q network architecture to process augmented data, thereby generating a dynamic control strategy. The augmented data includes environmental data, equipment operation data, energy consumption data, and energy consumption prediction values, specifically including: The SAC algorithm with a maximum entropy framework is constructed, and a dual-Q network architecture is built; wherein, the dual-Q network architecture includes two independent Q networks; When updating the SAC algorithm, the output of the network with the smaller estimate among the two independent Q networks is selected as the target value to update the SAC algorithm. The expanded data is input into the SAC algorithm to generate a dynamic control strategy.
6. A non-volatile computer storage medium for air conditioning unit control, storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: Collect multi-source time-series data of the air conditioning unit; wherein, the multi-source time-series data includes environmental data, equipment operation data and energy consumption data, and the environmental data includes internal environmental data and external environmental data; The preprocessed multi-source time-series data is processed based on a preset time window slicing algorithm and a noise injection algorithm to generate an enhanced dataset; A Mamba time-series prediction model is constructed, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption value and equipment operating parameters of the air conditioning unit in the first time period. Define a comprehensive reward function, and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value; The original comprehensive reward value is processed based on the reward clustering algorithm to output a normalized comprehensive reward signal; The SAC algorithm, based on a preset maximum entropy framework, is combined with a preset dual-Q network architecture to process augmented data in order to generate a dynamic control strategy; the augmented data includes environmental data, equipment operation data, energy consumption data, and energy consumption prediction values. Based on the dynamic control strategy, output equipment control commands with the lowest energy consumption. A Mamba time-series prediction model is constructed, and the augmented dataset is processed based on the Mamba time-series prediction model to output the predicted energy consumption and equipment operating parameters of the air conditioning unit in the first time period, specifically including: A Mamba time series prediction model is built based on a pre-defined Transformer architecture; The augmented dataset is input into the Mamba time series prediction model, and the time series data is analyzed based on the linear layer, convolutional layer, gating mechanism layer and state space model (SSM) layer inside the Mamba time series prediction model to output the predicted energy consumption value of the air conditioning unit and the equipment operating parameters in the first time period. Define a comprehensive reward function, and process the real-time energy consumption value based on the comprehensive reward function to output the original comprehensive reward value, specifically including: The comprehensive reward function is defined as a weighted sum of energy efficiency and indoor comfort; wherein, energy efficiency is a negative reward, and energy efficiency is inversely proportional to the predicted energy consumption value, and indoor comfort is a positive reward, and the deviation of indoor temperature from the set value is inversely proportional to indoor comfort. The real-time energy consumption value and indoor temperature data are input into the comprehensive reward function to calculate the original comprehensive reward value. The original comprehensive reward value is processed based on a reward clustering algorithm to output a normalized comprehensive reward signal, specifically including: Calculate the moving average of the original total reward value; where the initial moving average is zero; The update rate of the moving average is adjusted based on the dynamic step size parameter; wherein the initial value of the dynamic step size parameter is a non-zero positive number, and decays exponentially with the training process; Subtract the current moving average from the original comprehensive reward value to generate a normalized comprehensive reward signal centered on the zero mean. The SAC algorithm, based on a pre-defined maximum entropy framework, is combined with a pre-defined dual-Q network architecture to process augmented data, thereby generating a dynamic control strategy. The augmented data includes environmental data, equipment operation data, energy consumption data, and energy consumption prediction values, specifically including: The SAC algorithm with a maximum entropy framework is constructed, and a dual-Q network architecture is built; wherein, the dual-Q network architecture includes two independent Q networks; When updating the SAC algorithm, the output of the network with the smaller estimate among the two independent Q networks is selected as the target value to update the SAC algorithm. The expanded data is input into the SAC algorithm to generate a dynamic control strategy.
Citation Information
Patent Citations
Smart air conditioner control method and system
CN111795484A
Building air conditioner energy consumption prediction method based on indoor temperature optimization control
CN116045443A