Intelligent curtain wall energy-saving control method and system based on deep learning
By employing a regional and hierarchical deep learning control strategy, the problems of high energy consumption and poor comfort in traditional curtain wall control methods have been solved. This has enabled refined energy-saving control and adaptive adjustment of the curtain wall system, thereby improving building energy efficiency and user experience.
Patent Information
- Application Number
- CN202510034054.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-01-09
AI Technical Summary
Traditional curtain wall control methods cannot dynamically adjust according to environmental changes and user needs, resulting in high building energy consumption and difficulty in ensuring indoor comfort. Existing intelligent control systems are unable to effectively process multi-source heterogeneous data and lack adaptive learning capabilities.
A regional and hierarchical control strategy is adopted. Temperature, light, and energy consumption data are collected and processed using deep learning technology. An energy consumption feature vector and pattern feature database are established, and a deep reinforcement learning model is trained to achieve adaptive droop coefficient adjustment and multi-objective optimization control. Combined with an online update mechanism, the control accuracy and flexibility are improved.
It achieves refined control of the curtain wall system, balances energy-saving effects with indoor comfort, has the ability to continuously learn and adapt to dynamic changes in the building, and improves data quality and control accuracy.
Smart Images

Figure CN119846969B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of deep learning, in particular to an intelligent curtain wall energy-saving control method and system based on deep learning. BACKGROUND
[0002] As an important part of building envelope, the energy consumption control of curtain wall system directly affects the overall energy utilization efficiency of the building. Traditional curtain wall control methods mainly rely on simple threshold control or fixed strategy control, which cannot dynamically adjust according to environmental changes and user needs, resulting in high building energy consumption and difficulty in ensuring indoor environmental comfort.
[0003] Although existing intelligent curtain wall control systems have introduced some intelligent algorithms, there are still many problems in practical application: the control system cannot effectively process multi-source heterogeneous data such as temperature, illumination and energy consumption in the curtain wall system, and cannot accurately capture the system operation characteristics; secondly, the traditional control strategy cannot balance the multi-objective optimization problem of energy-saving effect and indoor comfort, and often loses one to gain the other; thirdly, the control system lacks self-adaptive learning ability and is difficult to cope with dynamic changes in building use. SUMMARY
[0004] The main purpose of the present application is to provide an intelligent curtain wall energy-saving control method and system based on deep learning, which adopts a regional and hierarchical control strategy, fully considers the influence of building orientation and floor height, and realizes fine control management.
[0005] To achieve the above purpose, the present application provides an intelligent curtain wall energy-saving control method based on deep learning, comprising the following steps:
[0006] Collecting temperature data, illumination data and energy consumption data of each level area of the curtain wall system to obtain a sensor data matrix;
[0007] Calculating the curtain wall energy consumption residual value according to the sensor data matrix, and performing time sequence segmentation and K-means clustering on the curtain wall energy consumption residual value to obtain an energy consumption feature vector and an energy consumption mode feature database;
[0008] Based on the energy consumption feature vector and the energy consumption mode feature database, setting state space, action space and reward function, and training an initial deep reinforcement learning model;
[0009] Based on the initial deep reinforcement learning model, performing adaptive droop coefficient adjustment on the curtain wall control area, and adjusting the shading coefficient, ventilation opening degree and hollow layer temperature to obtain the energy-saving control parameters of each area;
[0010] Perform multi-objective optimization control according to the energy-saving control parameters of each region, and update the initial deep reinforcement learning model online to obtain a target deep reinforcement learning model.
[0011] The application further provides an intelligent curtain wall energy-saving control system based on deep learning, which comprises:
[0012] The acquisition module is configured to acquire temperature data, illumination data and energy consumption data of each hierarchical region of the curtain wall system to obtain a sensor data matrix.
[0013] The calculation module is configured to calculate a curtain wall energy consumption residual value according to the sensor data matrix, and perform time sequence segmentation and K-means clustering on the curtain wall energy consumption residual value to obtain an energy consumption feature vector and an energy consumption mode feature database.
[0014] The training module is configured to set a state space, an action space and a reward function based on the energy consumption feature vector and the energy consumption mode feature database, and train an initial deep reinforcement learning model.
[0015] The adjustment module is configured to perform adaptive droop coefficient adjustment on the curtain wall control region based on the initial deep reinforcement learning model, and adjust a sunshade coefficient, a ventilation opening degree and a hollow layer temperature to obtain energy-saving control parameters of each region.
[0016] The update module is configured to perform multi-objective optimization control according to the energy-saving control parameters of each region, and update the initial deep reinforcement learning model online to obtain a target deep reinforcement learning model.
[0017] The application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method according to any one of the above embodiments when executing the computer program.
[0018] The application further provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the method according to any one of the above embodiments when executed by a processor.
[0019] In summary, the technical scheme provided by the present application effectively solves the processing problem of multi-source heterogeneous data such as temperature, illumination and energy consumption by establishing a multi-level data acquisition network and a standardized processing flow, and improves the data quality and availability. The control strategy based on deep reinforcement learning does not need to establish an accurate mathematical model, and realizes intelligent control of the curtain wall system through online learning and experience accumulation. The introduction of the adaptive droop coefficient adjustment mechanism enables the control system to dynamically adjust the control parameters according to the characteristics of different regions, improving the accuracy and flexibility of the control. Through the multi-objective optimization algorithm, balanced control of energy saving effect and indoor comfort is realized, which not only guarantees the building energy efficiency, but also meets the user experience demand. The online updating mechanism based on the incremental learning method enables the control system to have the ability of continuous learning and optimization, and can continuously adapt to the dynamic changes in the building use process. The regional and hierarchical control strategy fully considers the influence of building orientation and floor height, and realizes more refined control management. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 is a step schematic diagram of the intelligent curtain wall energy-saving control method based on deep learning in an embodiment of the present application;
[0021] Figure 2 is a structure block diagram of the intelligent curtain wall energy-saving control system based on deep learning in an embodiment of the present application;
[0022] Figure 3 is a structure schematic block diagram of the computer device of an embodiment of the present application.
[0023] The implementation of the object of the present application, functional characteristics and advantages will be further described with reference to the drawings. DETAILED DESCRIPTION
[0024] In order to make the object, technical scheme and advantages of the present application more clear, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0025] With reference to Figure 1 The present embodiment provides an intelligent curtain wall energy-saving control method based on deep learning, which comprises the following steps:
[0026] S1, collecting temperature data, illumination data and energy consumption data of each level region of the curtain wall system to obtain a sensor data matrix;
[0027] The curtain wall system is hierarchically divided into regions, and monitoring points are set on the inner surface, outer surface and hollow layer of the curtain wall by comprehensively considering the orientation, floor height and environmental characteristics of the curtain wall, forming a systematic monitoring point layout scheme. The layout scheme not only needs to cover different levels of regions, but also needs to ensure that the monitoring points are evenly distributed and can fully capture the key environmental parameters of the curtain wall system in operation. Based on the monitoring point layout scheme, appropriate sensors are installed at each monitoring point to construct a sensor network layout. Temperature sensors, light sensors and energy consumption sensors are configured for each monitoring point to ensure that each sensor can accurately capture the corresponding physical quantities, such as temperature, light intensity and energy consumption data, etc. After the sensor network layout is completed, the temperature data of each level of the system is collected by calling the data acquisition module. The collection range of temperature data covers outdoor temperature, indoor temperature and curtain wall surface temperature, ensuring that the thermodynamic characteristics of different regions are effectively recorded. After collection, to eliminate data errors caused by different sensor accuracy, range and installation environment, the collected temperature data is standardized to obtain standardized temperature data. At the same time, under the same sensor network layout, light data is collected by the data acquisition module. The content of light data includes outdoor light intensity, indoor light intensity and glass daylight coefficient, which can reflect the daylighting performance and light-heat conversion efficiency of the curtain wall system. Similarly, the collected light data is standardized to form standardized light data. On this basis, the energy consumption data is collected by the energy consumption sensors arranged in the system, including shading coefficient, regional energy consumption, air conditioning load and other important energy consumption parameters related to the operation of the curtain wall. To ensure the applicability of energy consumption data in subsequent analysis, it is standardized to obtain standardized energy consumption data. The standardized temperature data, light data and energy consumption data are input into the data preprocessing module for processing. The preprocessing module detects outliers and removes abnormal data points caused by sensor failure or external interference to improve data accuracy. The cleaned data set is stored as a preprocessed data set. After the preprocessed data set is generated, it is structured according to time series and region number to construct a three-dimensional data array, where the time dimension records the data changes at different time points, the space dimension identifies the physical location of the sensor, and the parameter dimension stores the standardized values of temperature, light and energy consumption data. In order to meet the input requirements of the deep learning model, the three-dimensional data array is processed by tensor transformation to generate a sensor data matrix.
[0028] S2, calculating curtain wall energy consumption residual values according to the sensor data matrix, and performing time series segmentation and K-means clustering on the curtain wall energy consumption residual values to obtain energy consumption feature vectors and an energy consumption mode feature database;
[0029] Specifically, principal component analysis is performed on the sensor data matrix to extract the energy consumption characteristic principal components in the data. Principal component analysis maps high-dimensional sensor data into a lower-dimensional feature space by removing redundant information from the data, retaining the core features in the curtain wall system that best explain the changes in energy consumption. An energy consumption benchmark model is established based on the energy consumption characteristic principal components. The model uses a three-layer neural network structure to predict the energy consumption of the curtain wall by fitting actual energy consumption data. The input of the neural network is the energy consumption characteristic principal components, and the output is the energy consumption prediction value, while the hidden layer in the middle captures the complex energy consumption change relationship through a nonlinear activation function. In order to optimize the fitting performance of the neural network, the backpropagation algorithm is used to continuously adjust the network parameters, and the loss function is minimized to ensure that the model output is close to the actual energy consumption. The energy consumption benchmark model that has been fully trained can accurately predict the energy consumption of the curtain wall under given conditions. The actual energy consumption data in the sensor data matrix is subtracted from the energy consumption value predicted by the energy consumption benchmark model to obtain the difference calculation result. The difference calculation result is subjected to a moving average process to obtain more representative curtain wall energy consumption residual values. The moving average process eliminates the influence of sudden abnormal changes, so that the residual values can more stably reflect the energy consumption deviation of the curtain wall system. Based on its time series characteristics, the curtain wall energy consumption residual values are segmented, and local energy consumption deviation information is extracted by time period. In order to reveal the potential periodicity and frequency domain characteristics of the time series data, fast Fourier transform is used to perform frequency domain analysis on each segment of data. The frequency domain features reveal the dominant frequency and change pattern of the energy consumption deviation in different time periods. Due to the high dimensionality of the frequency domain features, dimensionality reduction processing is performed to extract the most discriminative time series features, thereby reducing computational complexity and improving clustering accuracy. The time series features are input into the improved K-means clustering model, and the optimal number of cluster centers is dynamically determined by the silhouette coefficient and Davies-Bouldin index to ensure that the clustering results can reflect the actual structural characteristics of the data. In the clustering process, the K-means algorithm optimizes the compactness of the data points within the cluster and the separation degree of the data points between the clusters through multiple iterations to obtain the energy consumption mode clustering result, reflecting the typical energy consumption modes of the curtain wall system under different operating conditions. Based on the energy consumption mode clustering result, the mean vector, variance vector, peak coefficient, and skewness coefficient of each cluster are calculated. These features describe the central tendency and dispersion of each energy consumption mode, and reflect the symmetry and extreme value characteristics of its data distribution. Through feature fusion, the above statistical features are combined into a complete energy consumption feature vector. Normalization is performed on the energy consumption feature vector to scale each feature to a uniform numerical range, obtaining a compressed feature vector. The compressed feature vector is associated with the timestamp and spatial location information to form an energy consumption mode feature database.
[0030] S3, based on the energy consumption feature vector and the energy consumption mode feature database, setting the state space, action space and reward function, training the initial deep reinforcement learning model;
[0031] It should be noted that the energy consumption feature vector and energy consumption pattern feature database are combined by integrating indoor temperature parameters, outdoor temperature parameters, light intensity parameters, energy consumption parameters, and time parameters into a unified feature set to construct a state vector, which is then organized into a state space representation matrix. Based on the state space representation matrix, controllable variables of the curtain wall are extracted. A first range for the shading coefficient, a second range for the ventilation opening, and a third range for the insulated layer temperature are set to generate an action space parameter set. These parameters represent adjustable options in the curtain wall energy-saving control process, and their design needs to comprehensively consider the physical limitations of the actual hardware equipment and the operational constraints of the curtain wall system. Based on the indoor temperature parameters, light intensity parameters, and energy consumption parameters in the state space representation matrix, the reward function is defined as an objective function, aiming to balance energy consumption optimization, indoor comfort, and lighting conditions. The reward function expression is R = -w1 × energy consumption value - w2 × |actual room temperature - set room temperature| - w3 × |actual illuminance - set illuminance|, where w1, w2, and w3 are weighting coefficients, respectively measuring the importance of energy consumption, room temperature deviation, and lighting deviation. This reward function penalizes behaviors that deviate from the ideal state by using negative values, prompting the reinforcement learning model to optimize energy consumption while meeting the temperature and illuminance requirements of the indoor environment. A deep Q-network is constructed using the state space representation matrix and the action space parameter set. This network adopts the Q-learning method in deep reinforcement learning, and its architecture includes an input layer, three hidden layers, and an output layer. The number of neurons in the hidden layers is set to 256, 128, and 64, respectively. The layered design can capture the complex features in the high-dimensional state space while ensuring computational efficiency. The input of the deep Q-network is a state vector, and the output is the Q-value of the corresponding action. The Q-value represents the long-term cumulative reward obtained by choosing a certain action in a specific state. During the training of the deep Q-network, key training parameters such as the size of the experience replay buffer, the discount factor γ, the exploration factor ϵ, and the learning rate α are set. Among them, the experience replay buffer is used to store past state transition samples to break the temporal correlation between samples; the discount factor γ controls the importance of long-term rewards; the exploration factor ϵ balances the exploration and utilization of the model; and the learning rate α determines the update magnitude of the model parameters. By inputting the state space representation matrix into the deep Q-network, the network calculates the Q-value for each action. By combining a reward function model, an immediate reward value is calculated, and a temporal difference algorithm is used to compute the target Q-value, thus defining a loss function. The loss function measures the difference between the Q-value predicted by the deep Q-network and the target Q-value. The Adam optimization algorithm is used to minimize the loss function, and the network parameters are adjusted through backpropagation. Continuous optimization of the network parameters allows the deep Q-network to gradually approach the optimal Q-value function, thereby learning the optimal control policy. During training, the model gradually improves the effectiveness of its policy by repeatedly calculating Q-values and updating network weights.After each training phase, the trained deep Q-network is applied to the validation dataset for performance evaluation, calculating the average reward value and monitoring performance metrics on the validation set. When the average reward value reaches convergence and the validation performance remains stable, the initial deep reinforcement learning model is considered successfully trained. Through this model, the curtain wall control system can achieve intelligent adjustment based on deep learning, optimizing energy consumption and improving indoor environmental comfort.
[0032] S4, based on the initial deep reinforcement learning model, performs adaptive sag coefficient adjustment on the curtain wall control area, and adjusts the shading coefficient, ventilation opening and hollow layer temperature to obtain energy-saving control parameters for each area;
[0033] Specifically, the Q-value output of the initial deep reinforcement learning model is mapped to the actual building space. Based on the orientation characteristics of the curtain wall, the control area is divided into four directions: east, south, west, and north. Simultaneously, it is further divided into low, middle, and high zones based on floor height, constructing a control unit area matrix. Each cell in this matrix represents a specific spatial location. A droop control model is then constructed based on the control unit area matrix. By establishing a linear mapping between the actual energy consumption and standard energy consumption of each control unit, the corresponding droop characteristic coefficient set is calculated. The droop characteristic coefficients describe the energy consumption deviation characteristics of each control unit area. To adapt to changes in different time periods and environmental conditions, the droop characteristic coefficient set is adaptively adjusted, forming an adaptive droop coefficient matrix. This matrix can reflect the energy-saving needs of each control unit area in real time and serves as the core parameter input for subsequent calculations. Based on the adaptive droop coefficient matrix, shading control parameters are calculated, taking into account the solar altitude angle. and azimuth By establishing a formula for calculating the angle of shading louvers To determine the shading angle of the curtain wall system under different solar incidence conditions. Indicates the angle of incidence of the sun; The shading angle represents the angle between the sun and the curtain wall, describing the relationship between sunlight and the curtain wall direction. Based on the shading angle and the indoor illuminance requirement, a shading control matrix is generated to ensure that dynamically adjusting the angle of the shading louvers reduces the heat gain from solar radiation while meeting indoor illuminance comfort requirements. A ventilation control model is established based on the shading control matrix to regulate the ventilation effect in the curtain wall area. The ventilation volume is calculated based on the indoor-outdoor temperature difference and wind speed. By establishing the relationship between airflow and vent opening, a ventilation opening matrix is generated. The ventilation opening matrix is input into the heat transfer model, and the temperature distribution in the hollow layer is calculated by solving the heat transfer differential equation, generating a temperature control matrix. To achieve overall optimization, the shading control matrix, ventilation opening matrix, and temperature control matrix are combined to construct a multi-objective optimization function. The optimization function takes the form of... ,in The energy consumption per unit area target value is used to reflect energy saving efficiency; The human comfort index value is used to evaluate the comfort of indoor temperature and illumination; The control stability index value is used to measure the stability of the adjustment process. The balance factor controls the weight distribution among the three targets to ensure that the optimization result takes into account energy consumption, comfort, and stability. The above multi-objective optimization function is input into the particle swarm optimization algorithm, and the energy-saving control parameters of each region are obtained through iterative optimization calculation of the particle swarm algorithm. The particle swarm optimization algorithm simulates the cooperative behavior of the particle swarm in the search space, gradually approaches the global optimal solution, and thus generates optimal shading coefficients, ventilation openings, and hollow layer temperature adjustment parameters for each control unit region.
[0034] S5, according to the energy-saving control parameters of each region, performing multi-objective optimization control, and online updating the initial deep reinforcement learning model to obtain the target deep reinforcement learning model.
[0035] The energy-saving control parameters of each region are introduced into the control execution unit of the curtain wall. The control region is subdivided into specific execution units according to the orientation and floor height of the curtain wall. For the specific needs of each region, the angle of the sunshade louver is adjusted, the ventilation opening is adjusted, and the temperature set value is modified to generate a regional control execution matrix. This matrix directly corresponds to the specific control actions of each region, ensuring that the control strategy can be distributed and implemented in space and time. According to the regional control execution matrix, the time interval for data collection is set, and the real-time environmental parameter data is collected using the sensor network installed in the curtain wall system. The temperature sensor collects indoor and outdoor temperature change information, and the light sensor captures indoor and outdoor light intensity fluctuations. The energy consumption sensor monitors the actual energy consumption data of different regions. These data are integrated into an environmental parameter set, which describes the current system operating state and external environmental conditions. To ensure data consistency and availability, the environmental parameter set is normalized to convert the data into a standardized parameter matrix. The standardized parameter matrix is input into the initial deep reinforcement learning model. The model extracts the current system's feature state through state space mapping and selects the optimal adjustment action in the action space. At the same time, based on the pre-set reward function, the immediate reward value of the current action is calculated, and the optimization control vector of each region is output. These vectors contain specific sunshade adjustment, ventilation adjustment, and temperature adjustment, representing the model's optimal adjustment scheme for the current operating state. To apply the optimization control vector to actual control, it is decoded and mapped. Through decoding, the normalized control amount is restored to the actual sunshade louver angle value, ventilation opening value, and temperature set value. The decoded control amount is distributed according to the specific control region to generate the execution control instruction. After generating the execution control instruction, the system's execution mechanism is driven to operate, including adjusting the sunshade louver motor angle to change the sunshade state, adjusting the ventilation window opening to optimize the ventilation amount, and controlling the temperature of the hollow layer to adjust the heat exchange. These actions are independently executed in different regions, and the execution state data is collected through sensors to obtain real-time system responses. These response data are used not only to verify the execution effect of the control instruction but also to provide feedback information for the online update of the model. Based on the system response data, energy-saving indicators and comfort indicators are calculated. The energy-saving indicator evaluates the overall energy efficiency of the curtain wall system, and the comfort indicator measures the impact of indoor temperature and light on the human body. Through incremental learning, these indicators are fed back to the initial deep reinforcement learning model to gradually update the model parameters. The incremental learning method can optimize the strategy based on new data without damaging the original model performance, making the model adapt to changing environmental conditions and operational requirements. Through this process, the initial deep reinforcement learning model evolves into the target deep reinforcement learning model. The target model has higher control accuracy and adaptability and can achieve long-term stable operation in complex practical application scenarios, thereby improving the energy-saving effect and user comfort of the curtain wall system.
[0036] Cumulative energy consumption value, peak-valley energy consumption ratio value, and energy consumption standard deviation of unit area are calculated by extracting energy consumption data from system response data and performing time series statistics on 24-hour operation status of each control area. These statistics respectively reflect the energy consumption intensity, fluctuation amplitude, and uniformity of energy consumption distribution of the curtain wall system. By performing weighted average on these indicators, an energy-saving performance indicator matrix is generated. At the same time, indoor environmental parameters are extracted from system response data to comprehensively evaluate comfort performance. By combining physical parameters such as indoor temperature, humidity, air flow speed, and radiant temperature, the predicted mean vote (PMV) value is calculated, which is a comprehensive indicator reflecting human thermal comfort. The PMV value measures comfort by quantifying the deviation between the environment and human thermal balance. By calculating the uniformity of illumination and glare value, visual comfort is evaluated, and the two indicators respectively describe the uniformity of light distribution and the potential degree of eye irritation in the light environment. Based on these calculation results, a comfort performance indicator matrix is generated, quantifying the comfort effect of the curtain wall system under the operation of different areas. According to the energy-saving performance indicator matrix and the comfort performance indicator matrix, a model performance evaluation function is constructed to comprehensively evaluate the actual performance of the initial deep reinforcement learning model. The performance evaluation function reasonably allocates the weights of energy-saving and comfort indicators to generate a global performance score. When the output value of the model performance evaluation function is lower than the preset threshold, it means that the current model's control strategy cannot meet the system optimization goal, triggering the model update mechanism and generating an update trigger instruction. After the update mechanism is started, the system response data within the continuous 24 hours is sampled to construct model training data containing state transition sequence, action sequence, and reward sequence. These data reflect the dynamic behavior of the control system in actual operation and are an important basis for enhancing the adaptability of the model. Online feature extraction is performed on the sampled data to extract dynamic feature vectors. These feature vectors capture the transition patterns between different states of the system, the influence range of actions, and the regularity of reward changes. The extracted dynamic feature vectors are input into the initial deep reinforcement learning model to perform incremental training. The incremental training method updates the model parameters step by step to enable the model to absorb the operation rules reflected by the new data while maintaining the stability of the original model performance. In this process, the loss function of the model is dynamically adjusted to minimize the deviation between the new and old strategies while ensuring the model's rapid adaptation to new environmental conditions. After multiple rounds of training and optimization, the target deep reinforcement learning model is finally formed, which has higher control accuracy and flexibility and can continuously optimize the energy-saving performance and user comfort of the curtain wall system in a changing environment.
[0037] In one example, temperature data, illumination data, and energy consumption data of each level area of the curtain wall system are collected to obtain a sensor data matrix, including:
[0038] The curtain wall system is hierarchically regionally divided, monitoring points are set on the inner surface, outer surface and hollow layer according to the curtain wall orientation and floor height, and a monitoring point layout scheme is obtained;
[0039] Based on the monitoring point layout scheme, temperature sensors, light sensors and energy consumption sensors are installed and configured for each monitoring point, and a sensor network layout is obtained;
[0040] According to the sensor network layout, the data acquisition module is called to collect temperature data, including outdoor temperature, indoor temperature and curtain wall surface temperature, and standardization processing is performed on the temperature data to obtain standardized temperature data;
[0041] Based on the sensor network layout, the data acquisition module is called to collect light data, including outdoor light intensity, indoor light intensity and glass lighting coefficient, and standardization processing is performed on the light data to obtain standardized light data;
[0042] Through the energy consumption sensors in the sensor network layout, energy consumption data is collected, including shading coefficient, regional energy consumption and air conditioning load, and standardization processing is performed on the energy consumption data to obtain standardized energy consumption data;
[0043] The standardized temperature data, standardized light data and standardized energy consumption data are input into the data preprocessing module for outlier detection and deletion, and a preprocessed data set is obtained;
[0044] The preprocessed data set is constructed into a three-dimensional data array according to time sequence and region number, including time dimension, space dimension and parameter dimension, and a sensor data matrix is obtained through tensor transformation.
[0045] In this example, according to the specific structure of the building and the distribution characteristics of the curtain wall, the curtain wall system is hierarchically regionally divided, and grouped according to the orientation of the curtain wall (such as east, south, west, north) and the floor height (such as low, middle, high). In each region, according to the physical structure of the curtain wall, monitoring points are set on the inner surface, outer surface and hollow layer to capture the characteristics of heat transfer, light change and energy consumption data, forming a complete monitoring point layout scheme. According to the monitoring point layout scheme, sensors are installed and configured for each monitoring point, including temperature sensors, light sensors and energy consumption sensors. Temperature sensors are used to monitor indoor temperature ), outdoor temperature ) and curtain wall surface temperature ). Light sensors record outdoor light intensity ), indoor light intensity ) and glass lighting coefficient , where the glass lighting coefficient is calculated by the formula Calculations are performed to quantify the efficiency of light energy transmission through the glass. An energy sensor is used to collect the shading coefficient ( ). ), regional energy consumption ) and air conditioning load ( These data reflect the shading system's regulation efficiency, overall energy consumption level, and cooling or heating load. After the sensor network layout is completed, the data acquisition module collects and processes various parameters. Temperature data is collected, including... , and Because different sensors have different measurement ranges or resolutions, the temperature data is standardized. The standardization formula is expressed as:
[0046] ;
[0047] in, This represents the raw temperature data. This is the mean of this type of data. Standard deviation, This is standardized temperature data. Standardization eliminates scale differences in the temperature data, making it suitable for further analysis. Illumination data is collected based on the sensor network layout, including... , and Similarly, the illumination data is standardized using the following formula:
[0048] ;
[0049] in, This represents the original lighting data. and These represent the mean and standard deviation of the illumination data, respectively. Standardized illumination data reflects the relative illumination characteristics of different regions. Energy consumption data collected by energy consumption sensors includes... , and It also needs to be standardized. Its formula is similar to the standardization method for temperature and light intensity:
[0050] ;
[0051] in, This is the raw energy consumption data. and respectively. The normalized temperature data, illumination data, and energy consumption data are input into a data preprocessing module for outlier detection and cleaning. By setting upper and lower threshold values (e.g., according to a 95% confidence interval of historical data), outliers caused by sensor failure or environmental interference are removed, and a preprocessed data set is obtained. The preprocessed data set is constructed into a three-dimensional data array according to time sequence and region number. The three-dimensional data array includes time dimension (e.g., time points within 24 hours per day), space dimension (e.g., monitoring point positions of different floors and orientations), and parameter dimension (e.g., 、 、 ). Through tensor transformation operations, the three-dimensional data array is converted into a unified sensor data matrix to describe the running state of the entire curtain wall system. For example, the sensor data matrix at a specific time is represented by the formula:
[0052] ;
[0053] wherein, 、 and represent the temperature, illumination, and energy consumption standardized data, respectively, at time , on the floor, in the region. Through the above steps, the sensor data matrix is finally obtained.
[0054] In one example, the curtain wall energy consumption residual value is calculated according to the sensor data matrix, and the curtain wall energy consumption residual value is time-series segmented and K-means clustered to obtain an energy consumption feature vector and an energy consumption mode feature database, including:
[0055] Performing principal component analysis on the sensor data matrix to obtain energy consumption feature principal components;
[0056] Establishing an energy consumption benchmark model based on the energy consumption feature principal components, fitting the energy consumption data using a three-layer neural network structure, optimizing the network parameters through a backpropagation algorithm, and obtaining energy consumption prediction values;
[0057] Calculating the difference between the actual energy consumption data in the sensor data matrix and the energy consumption prediction values to obtain a difference calculation result, and performing a moving average process on the difference calculation result to obtain a curtain wall energy consumption residual value;
[0058] Segmenting the curtain wall energy consumption residual value, extracting the frequency domain features of each segment of data through a fast Fourier transform, and performing dimensionality reduction processing on the frequency domain features to obtain a time series feature sequence;
[0059] The time series of features is input into the improved K-means clustering model, the optimal number of cluster centers is determined by using the silhouette coefficient and the Davies-Bouldin index, the iterative clustering calculation is performed, and the energy consumption mode clustering result is obtained;
[0060] According to the energy consumption mode clustering result, the mean vector, variance vector, peak coefficient and skewness coefficient of each cluster are calculated, and the energy consumption feature vector is obtained through feature fusion;
[0061] The energy consumption feature vector is normalized to obtain the compressed feature vector, and the compressed feature vector is associated with the timestamp and spatial position information to obtain the energy consumption mode feature database.
[0062] In this example, principal component analysis is performed on the energy consumption related features using the sensor data matrix to reduce the complexity of high-dimensional data while retaining the key information that best explains the energy consumption changes. The sensor data matrix contains multiple variable dimensions, such as temperature data , illumination data and energy consumption data . Assuming is a matrix, where is the time series length, is the number of parameters, by calculating the covariance matrix and decomposing its eigenvalues and eigenvectors, the principal component matrix is obtained. Each column of the principal component matrix is a linear combination, for example:
[0063] ;
[0064] where, is the th principal component, is the linear combination coefficient, representing the contribution weight of the th variable to the th principal component. Based on the energy consumption feature principal component, an energy consumption benchmark model is established. This model uses a three-layer neural network structure to fit the nonlinear relationship between actual energy consumption data and principal component features. The number of input layer nodes of the network is equal to the number of principal components, the number of neurons in the hidden layer is set to and , and the output layer is the energy consumption prediction value . The prediction process of the neural network is represented as:
[0065] ;
[0066] where, and are weight matrices, and is the bias vector, is the activation function (such as ReLU), is the linear transformation of the output layer. By minimizing the loss function and updating the parameters and using the backpropagation algorithm, a neural network model capable of accurately predicting energy consumption is obtained. The energy consumption prediction value is subtracted from the actual energy consumption data to obtain the difference sequence :
[0067] ;
[0068] To eliminate the interference of short-term fluctuations on analysis, a moving average process is performed on the difference sequence to generate smooth curtain energy residual values :
[0069] ;
[0070] wherein, is the size of the sliding window, is the time point. The energy residual values are segmented by time sequence, and a fast Fourier transform is performed on each segment of data to extract frequency domain features. The fast Fourier transform converts the time domain signal into a frequency domain signal, and its formula is:
[0071] ;
[0072] wherein, is the frequency domain signal, is the frequency, is the data length. The frequency domain features reveal the distribution of energy residual values at different frequency components, and through dimension reduction techniques (such as principal component analysis or feature screening), the most discriminant time series features are extracted. The time series feature sequence is input into the improved K-means clustering model, and the optimal number of cluster centers is determined by the silhouette coefficient and the Davies-Bouldin index. The formula for calculating the silhouette coefficient is:
[0073] ;
[0074] wherein, is the average distance from the sample to other points in the cluster, is the average distance from the sample to the nearest neighbor cluster. And the Davies-Bouldin index is:
[0075] ;
[0076] wherein, and are diameters of the clusters and , respectively, is the distance between cluster centers. The energy consumption pattern clustering results are obtained by iterative clustering calculation. According to the clustering results, the mean vector , the variance vector , the peak coefficient and the skewness coefficient of each cluster are calculated:
[0077] ;
[0078] wherein, is the data sample within the cluster, is the sample quantity. Through feature fusion, these statistical indicators are combined into a complete energy consumption feature vector. The energy consumption feature vector is normalized to eliminate the scale effect, and the normalization formula is:
[0079] ;
[0080] The normalized compressed feature vector is associated with the time stamp and spatial location information to form an energy consumption pattern feature database. For example, for a certain cluster center, its feature vector , time and location are recorded. The database is used for energy consumption pattern analysis and optimization control decision.
[0081] In one example, based on the energy consumption feature vector and the energy consumption pattern feature database, the state space, the action space and the reward function are set, and the initial deep reinforcement learning model is trained, including:
[0082] The energy consumption feature vector and the energy consumption pattern feature database are combined to construct a state vector containing indoor temperature parameters, outdoor temperature parameters, light intensity parameters, energy consumption parameters and time parameters, and a state space representation matrix is obtained;
[0083] According to the state space representation matrix, the curtain control variable is extracted, the first value range of the shading coefficient, the second value range of the ventilation opening degree and the third value range of the hollow layer temperature are set, and an action space parameter set is obtained;
[0084] According to the indoor temperature parameter in the state space representation matrix, the illumination intensity parameter in the state space representation matrix and the energy consumption parameter in the state space representation matrix, a reward function R=-w1x energy consumption value-w2x|actual room temperature-set room temperature|-w3x|actual illumination-set illumination| is constructed, to obtain a reward function model, wherein R represents the reward function, w1, w2 and w3 represent weight coefficients;
[0085] Based on the state space representation matrix and the action space parameter set, a deep Q network is constructed, which includes an input layer, three hidden layers and an output layer, and the number of neurons in the hidden layers is 256, 128 and 64 respectively;
[0086] The experience replay buffer size, the discount factor γ, the exploration factor ε and the learning rate α are set for the deep Q network, to obtain the deep Q network training parameters;
[0087] The state space representation matrix is input into the deep Q network, the Q value is calculated through the deep Q network, the immediate reward is calculated by using the reward function model, and the target Q value is calculated based on the time difference algorithm, to obtain a loss function value;
[0088] According to the loss function value, a back propagation operation is performed, the network parameters are updated by using the Adam optimization algorithm, a trained deep Q network is obtained, and the trained deep Q network is used for performance evaluation on a verification data set, when the average reward value converges and the performance index on the verification set is stable, an initial deep reinforcement learning model is obtained.
[0089] In this example, the energy consumption feature vector and the energy consumption mode feature database are combined, the indoor temperature parameter , the outdoor temperature parameter , the illumination intensity parameter , the energy consumption parameter and the time parameter are integrated into a unified state description, and a state vector is constructed. For a specific time , the state vector is expressed as:
[0090] ;
[0091] All the state vectors at different times are combined to form a state space representation matrix , wherein each row represents the state of a time, and the columns represent different parameter dimensions. After the state space is defined, the control variables of the curtain system are extracted according to the state space representation matrix, including the shading coefficient , the ventilation opening degree and the hollow layer temperature These variables control the behavior of the shading system, ventilation system and heating system respectively. To ensure the value range of the control variables is reasonable, a specific range is set for each variable, i.e. the first value range of the shading coefficient, the second value range of the ventilation opening degree and the third value range of the hollow layer temperature, to obtain the action space parameter set , i.e.
[0092] ;
[0093] Based on the state space representation matrix and the energy consumption parameter , the indoor temperature parameter and the light intensity parameter , a reward function model is constructed . The reward function punishes the undesirable energy consumption, temperature deviation and light deviation in the form of negative value, and its specific expression is:
[0094] ;
[0095] wherein , and are weight coefficients, respectively representing the relative importance of energy consumption, temperature deviation and light deviation to the overall optimization goal; and are the set room temperature and light target values. Based on the state space representation matrix and the action space parameter set , a deep Q network is constructed, whose structure contains an input layer, three hidden layers and an output layer. The number of nodes of the input layer is equal to the dimension of the state vector, for example, it contains , , , , and time a total of six nodes. The hidden layer is set to three layers, and the number of neurons is 256, 128 and 64 respectively. The design of reducing layer by layer helps to capture high-dimensional features in the state space and effectively compress them. The number of nodes of the output layer is equal to the dimension of the action space, and the output value of each node corresponds to the Q value of an action. In order to train the deep Q network, the experience replay buffer size , the discount factor , the exploration factor and the learning rate are set. Among them, the buffer size is used to store the past state transition samples; the discount factor determines the weight of future rewards; the exploration factor is used to control the balance between random exploration and greedy selection; the learning rate Step size of model parameter update is determined. During training, state space representation matrix Input deep Q network, network according to current state Q value of all possible actions is calculated Instantaneous reward Through reward function model, target Q value is calculated according to time difference algorithm:
[0096] ;
[0097] Based on predicted Q value and target Q value, loss function is defined:
[0098] ;
[0099] Wherein, Sample batch size. Use back propagation algorithm to optimize loss function, use Adam optimization algorithm to adjust parameters of deep Q network, including weight and bias. Model parameters are updated during training, and finally loss function converges. The trained deep Q network is evaluated on the validation dataset, and the average reward value And other performance indicators (such as mean square error) are calculated to judge the effect of the model. When the average reward value Converge and the index on the validation set is stable, the initial deep reinforcement learning model is obtained.
[0100] In an example, based on the initial deep reinforcement learning model, adaptive sag coefficient adjustment is performed on the curtain wall control area, and the shading coefficient, ventilation opening degree and hollow layer temperature are adjusted to obtain the energy saving control parameters of each area, including:
[0101] Map the Q value output result of the initial deep reinforcement learning model according to the building space attribute, divide the control area into eastward control area, southward control area, westward control area and northward control area according to the curtain wall orientation, and divide the control area into low area control area, middle area control area and high area control area according to the floor height, to obtain the control unit area matrix;
[0102] Based on the control unit area matrix, establish the sag control model, establish the linear mapping relationship between the actual energy consumption of each control unit and the standard energy consumption, obtain the sag characteristic coefficient set, and adjust the sag characteristic coefficient set adaptively to obtain the adaptive sag coefficient matrix;
[0103] Based on the adaptive sag coefficient matrix, the shading coefficient is calculated, the calculation formula of shading louver angle αs is established according to the solar elevation angle θs and the azimuth angle φs, that is, αs=arctan(tanθs×cosφs), wherein θs represents the solar incident angle, φs represents the solar azimuth angle, and the shading control matrix is obtained combined with the indoor illuminance demand value.
[0104] According to the shading control matrix, a ventilation control model is established, the ventilation volume is calculated based on the indoor and outdoor temperature difference and wind speed, the ventilation opening degree matrix is obtained, and the ventilation opening degree matrix is input into the heat transfer model, the hollow layer temperature distribution is calculated by solving the heat transfer differential equation, and the temperature control matrix is obtained;
[0105] The shading control matrix, the ventilation opening degree matrix and the temperature control matrix are combined to construct a multi-objective optimization function F = λ1 × E1 + λ2 × E2 + λ3 × E3, wherein E1 represents the unit area energy consumption target value, E2 represents the human comfort index value, and E3 represents the control stability index value, and λ1, λ2 and λ3 represent balance factors;
[0106] The multi-objective optimization function is input into the particle swarm optimization algorithm, and the energy-saving control parameters of each region are calculated by iterative optimization.
[0107] In this example, the Q value output result of the initial deep reinforcement learning model is mapped according to the spatial properties of the building, combined with the curtain wall characteristics of the building, divided into east, south, west and north control regions according to the orientation, and subdivided into low, middle and high regions according to the floor height. Through the two-dimensional division, the control unit region matrix is formed , and each element in the matrix corresponds to the control unit of the th orientation and the th floor height. On this basis, a droop control model is constructed to dynamically adjust the energy consumption level of each control unit. For each control unit, the difference between the actual energy consumption and the standard energy consumption is calculated, and a linear mapping relationship is established between the two, and the droop characteristic coefficient is defined as:
[0108] ;
[0109] wherein describes the deviation of each unit relative to the standard energy consumption. By continuously collecting new energy consumption data, the adaptive droop characteristic coefficient is adjusted, and an adaptive droop coefficient matrix is generated, and the element of the adaptive droop coefficient matrix represents the droop characteristic of the th orientation and the th floor. Based on the adaptive droop coefficient matrix , the shading coefficient is calculated. Combined with the solar elevation angle and the azimuth angle , the adjustment angle of the shading louver is derived through geometric relationship, and the formula is:
[0110] ;
[0111] wherein, denotes the incident angle of sunlight; is the azimuth angle of sunlight relative to the curtain wall. Combined with the indoor illuminance requirement value , the shading coefficient is adjusted to meet the illuminance requirement, and a shading control matrix is finally formed. According to the shading control matrix , a ventilation control model is established to calculate the opening of the ventilation opening . The calculation formula of the ventilation volume is:
[0112] ;
[0113] wherein, is the air volume coefficient, denotes the wind speed, is the opening of the ventilation opening. Combined with the indoor and outdoor temperature difference , the opening of the ventilation opening is adjusted so that the temperature change of the indoor environment meets the comfort requirement, and a ventilation opening matrix is formed. The ventilation opening matrix is input into the heat transfer model, and the hollow layer temperature distribution is calculated by solving the heat transfer differential equation:
[0114] ;
[0115] wherein, is the thermal diffusion coefficient, is the spatial position, is the time. The solved temperature distribution is used to generate a temperature control matrix . The shading control matrix , the ventilation opening matrix and the temperature control matrix are combined to construct a multi-objective optimization function , which is in the form of:
[0116] ;
[0117] wherein, denotes the unit area energy consumption target value, which is calculated as:
[0118] ;
[0119] denotes the human comfort index value, which comprehensively considers the temperature, humidity and illumination conditions and is represented by the predicted mean vote (PMV). The calculation formula of PMV is:
[0120] ;
[0121] wherein is the metabolic rate, is the thermal load. represents the control stability index value, used to constrain the frequent changes of control actions. The multi-objective optimization function is iteratively optimized by the particle swarm optimization algorithm, and the initial particle position is a random value of the control parameter, and the particle update rule is:
[0122] ;
[0123] ;
[0124] wherein, is the particle velocity, is the particle position, is the historical optimal position of the particle, is the global optimal position, , , are inertia weight and learning factor respectively, , is a random number. Through multiple iterations of optimization, the energy-saving control parameters of each region are finally obtained, including the optimized shading coefficient, ventilation opening and hollow layer temperature value.
[0125] In one example, multi-objective optimization control is performed according to the energy-saving control parameters of each region, and the initial deep reinforcement learning model is updated online to obtain a target deep reinforcement learning model, including:
[0126] The energy-saving control parameters of each region are imported into a control execution unit, and the shading louver angle adjustment, ventilation opening adjustment and temperature set value adjustment are performed according to the control region corresponding to the curtain wall orientation and floor height, to obtain a regional control execution matrix;
[0127] The data collection time interval is set according to the regional control execution matrix, the indoor and outdoor temperature data are collected by a temperature sensor, the indoor and outdoor light intensity data are collected by a light sensor, and the regional energy consumption data are collected by an energy consumption sensor, to obtain an environmental parameter set;
[0128] The environmental parameter set is normalized to obtain a standardized parameter matrix, and the standardized parameter matrix is input into the initial deep reinforcement learning model, and after state space mapping, action space selection and reward function calculation, the shading adjustment amount, ventilation adjustment amount and temperature adjustment amount of each region are output, to obtain an optimization control vector;
[0129] The optimized control vector is decoded and mapped to restore the normalized control quantities to actual values of the sunshade louver angle, ventilation opening degree and temperature set value, and is distributed according to the control area to obtain an execution control instruction;
[0130] According to the execution control instruction, the execution mechanism is driven to operate, the sunshade louver motor angle, ventilation window opening degree and hollow layer temperature are adjusted, and execution state data is collected to obtain system response data;
[0131] Based on the system response data, energy saving indicators and comfort indicators are calculated, and an incremental learning method is used to update the initial deep reinforcement learning model online to obtain a target deep reinforcement learning model.
[0132] In this example, the energy saving control parameters of each area are imported into the control execution unit, and the building is partitioned according to the specific orientation (such as east, south, west and north) and floor height (low, middle and high) of the curtain wall. Each control area corresponds to an execution unit, and its control parameters include the sunshade louver angle , ventilation opening degree and temperature set value . These parameters are organized into a regional control execution matrix , and the element of the matrix represents the control parameter set of the th orientation and the th floor. After the control parameter distribution is completed, the time interval for data collection is set according to the regional control execution matrix , for example, once every 15 minutes, the indoor temperature and outdoor temperature are recorded by temperature sensors, the indoor light intensity and outdoor light intensity are collected by light sensors, and the area energy consumption is measured by energy consumption sensors. These data are integrated into an environmental parameter set to describe the current state of the curtain wall system. The environmental parameter set is normalized, and the normalization formula is:
[0133] ;
[0134] wherein is the original data, and are the minimum and maximum values of the parameter, respectively. Through normalization, a standardized parameter matrix is obtained, so that all data values are within the range [0, 1], which is convenient for subsequent input into the deep reinforcement learning model. The standardized parameter matrix The input is fed into an initial deep reinforcement learning model, which extracts features from the current system state through state space mapping to obtain a state vector , which includes elements including , , , , , and a time parameter . The deep reinforcement learning model selects the optimal action in the action space A, and calculates the Q value under the current state . The immediate reward is calculated through a reward function model :
[0135] ;
[0136] wherein , , are weight coefficients representing the influence of energy consumption, room temperature deviation, and light deviation on the total reward, respectively, and are the set room temperature and light target values. The model outputs an optimized control vector . The optimized control vector needs to be decoded and mapped to restore the normalized control values to actual values. For example, the decoding formula for the angle of the sunshade louver is:
[0137] ;
[0138] wherein and are the minimum and maximum values of the louver angle, respectively. Similarly, the actual values of the ventilation opening and temperature set value are calculated through the corresponding decoding formula. After decoding, the control parameters of each region are assigned to the corresponding execution unit to generate the final execution control instruction. The execution control instruction is passed to the hardware layer to drive the sunshade louver motor to adjust the angle , adjust the ventilation window opening , and set the hollow layer temperature . At the same time, real-time execution state data such as the actual adjusted louver angle, ventilation volume, and hollow layer temperature are collected through sensors to generate system response data. Based on the system response data, energy saving indicators and comfort indicators are calculated. The energy saving indicator represents the energy consumption per unit area, and the formula is:
[0139] ;
[0140] While the comfort indicator is represented by the predicted mean vote (PMV), and the specific formula is:
[0141] ;
[0142] wherein is the metabolic rate, is the human heat load. Incremental learning is performed on the initial deep reinforcement learning model in combination with the system response data. The incremental learning method samples new state transition sequences to construct training data, and the model is optimized online. In the updating process, the target Q value is calculated by the time difference formula:
[0143]
[0144] and the model parameters are updated by minimizing the loss function . The optimization algorithm uses the Adam optimizer, and finally forms the optimized target deep reinforcement learning model, realizing a dynamic and adaptive energy-saving control strategy.
[0145] In one example, the energy-saving indicators and comfort indicators are calculated based on the system response data, and the incremental learning method is used to update the initial deep reinforcement learning model online to obtain the target deep reinforcement learning model, including:
[0146] The energy consumption data in the system response data is time series statistical, and the cumulative energy consumption per unit area of each control area within 24 hours, the peak-valley energy consumption ratio and the energy consumption standard deviation are calculated. The energy-saving performance index matrix is obtained by weighted average;
[0147] The indoor environment parameters in the system response data are evaluated and calculated, the predicted mean vote value PMV is calculated based on temperature, humidity, air flow velocity and radiation temperature, and the visual comfort is calculated based on the uniformity of illumination and glare value, and the comfort performance index matrix is obtained;
[0148] According to the energy-saving performance index matrix and the comfort performance index matrix, the model performance evaluation function is constructed, and the performance of the initial deep reinforcement learning model is evaluated. When the output value of the model performance evaluation function is less than the preset threshold value, the model updating mechanism is activated, and the update trigger instruction is obtained;
[0149] Based on the update trigger instruction, the system response data of 24 consecutive hours is sampled to construct the model training data containing state transition sequences, action sequences and reward sequences;
[0150] The model training data is online feature extracted to obtain a dynamic feature vector, and the dynamic feature vector is input into the initial deep reinforcement learning model for incremental training to obtain the target deep reinforcement learning model.
[0151] In this example, the energy consumption data of each control area is extracted from the system response data, and time series statistical analysis is performed. The cumulative energy consumption per unit area of each control area within 24 hours is calculated , formula is:
[0152] ;
[0153] wherein, is the control area The actual energy consumption value at time , is the area of the region. Calculate the peak-to-valley energy consumption ratio of each control area :
[0154] ;
[0155] This index reflects the volatility of energy consumption. Calculate the energy consumption standard deviation of each control area , indicating the degree of dispersion of energy consumption:
[0156] ;
[0157] wherein, is the average energy consumption of the region , is the number of time points. After weighted average of the above three indexes, the energy saving performance index matrix is constructed, the element of which corresponds to the th region and the th performance dimension. At the same time, the indoor environment parameters are extracted from the system response data to evaluate the comfort performance index. Through temperature , humidity , air flow velocity and radiant temperature , the predicted mean vote (PMV) is calculated, and the formula is:
[0158] ;
[0159] wherein, is the metabolic rate, indicates the human heat load, which is calculated by comprehensive factors such as radiant heat transfer, convective heat transfer and evaporative heat dissipation. The visual comfort is evaluated by the uniformity of illumination and the glare value , which are calculated as follows:
[0160] ;
[0161] ;
[0162] wherein and are the minimum and maximum illuminance, is the brightness of the th light source, is the average luminance, is the number of light sources. Based on PMV and visual comfort, a comfort performance index matrix is constructed . By combining the energy-saving performance index matrix and the comfort performance index matrix , a model performance evaluation function is defined:
[0163] ;
[0164] wherein, and are the target values of energy saving and comfort respectively, and are the weights. When the performance evaluation function is less than a preset threshold, a model updating mechanism is triggered. Based on the updating trigger instruction, a state transition sequence is generated by sampling from the system response data of continuous 24 hours, wherein, represents the current state, including environmental parameters; is an action vector, including shading, ventilation and temperature adjustment; is the immediate reward; is the next state. After constructing the action sequence and the reward sequence, the input data is subjected to online feature extraction to extract a dynamic feature vector :
[0165] ;
[0166] wherein is a feature extraction function for capturing time correlation. The dynamic feature vector is input into an initial deep reinforcement learning model for incremental training. The target Q value is updated by a time difference algorithm:
[0167] ;
[0168] wherein is a discount factor. A loss function is defined:
[0169] ;
[0170] The loss function is minimized using the Adam optimization algorithm to update the model parameters. After multiple rounds of iterative training, the model gradually learns the latest environmental change rule, and finally forms a target deep reinforcement learning model.
[0171] With reference to Figure 2 , the embodiment provides an intelligent curtain energy-saving control system based on deep learning, comprising:
[0172] The collection module 1 is used for collecting temperature data, illumination data and energy consumption data of each hierarchical region of the curtain wall system to obtain a sensor data matrix;
[0173] The calculation module 2 is used for calculating a curtain wall energy consumption residual value according to the sensor data matrix, and performing time sequence segmentation and K-means clustering on the curtain wall energy consumption residual value to obtain an energy consumption feature vector and an energy consumption mode feature database;
[0174] The training module 3 is used for setting a state space, an action space and a reward function based on the energy consumption feature vector and the energy consumption mode feature database, and training an initial deep reinforcement learning model;
[0175] The adjustment module 4 is used for performing adaptive droop coefficient adjustment on the curtain wall control region based on the initial deep reinforcement learning model, and adjusting a sunshade coefficient, a ventilation opening degree and a hollow layer temperature to obtain regional energy-saving control parameters;
[0176] The update module 5 is used for performing multi-objective optimization control according to the regional energy-saving control parameters, and performing online update on the initial deep reinforcement learning model to obtain a target deep reinforcement learning model.
[0177] In the embodiment, the specific implementation of each unit in the system embodiment is described above, and will not be repeated here.
[0178] Referring to Figure 3 , the embodiment of the present application also provides a computer device, which can be a server, and the internal structure thereof can be as shown in Figure 3 . The computer device comprises a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer device is used for providing calculation and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used for storing corresponding data in the embodiment. The network interface of the computer device is used for communicating with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.
[0179] Those skilled in the art can understand Figure 3 that the structure shown in the figure is only a block diagram of part of the structure related to the present application scheme, and does not constitute a limitation on the computer device to which the present application scheme is applied.
[0180] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the method.
[0181] It is understood by those skilled in the art that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium provided by the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM, etc.
[0182] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, devices, articles or methods including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, devices, articles or methods. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0183] The above description is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, as described in the specification and drawings of the present application, are also included in the patent protection scope of the present application.
Claims
1. A deep learning-based intelligent curtain wall energy-saving control method, characterized in that, The method comprises the following steps: temperature data, illumination data and energy consumption data of each hierarchical region of the curtain wall system are collected to obtain a sensor data matrix; a curtain wall energy consumption residual value is calculated according to the sensor data matrix, and the curtain wall energy consumption residual value is time-series segmented and K-means clustered to obtain an energy consumption feature vector and an energy consumption mode feature database; a state space, an action space and a reward function are set based on the energy consumption feature vector and the energy consumption mode feature database, and an initial deep reinforcement learning model is trained; based on the initial deep reinforcement learning model, adaptive droop coefficient adjustment is performed on the curtain wall control region, and the shading coefficient, the ventilation opening degree and the hollow layer temperature are adjusted to obtain regional energy-saving control parameters; specifically, the Q value output result of the initial deep reinforcement learning model is mapped according to the building space attribute, the control region is divided into an east-facing control region, a south-facing control region, a west-facing control region and a north-facing control region according to the curtain wall orientation, and the control region is divided into a low-zone control region, a middle-zone control region and a high-zone control region according to the floor height to obtain a control unit region matrix; a droop control model is established based on the control unit region matrix, a linear mapping relationship between the actual energy consumption of each control unit and the standard energy consumption is established to obtain a droop characteristic coefficient set, and the adaptive droop coefficient matrix is obtained by adaptively adjusting the droop characteristic coefficient set; based on the adaptive droop coefficient matrix, the shading coefficient is calculated, a calculation formula of the shading louver angle αs is established according to the solar elevation angle θs and the azimuth angle φs, αs=arctan(tanθs×cosφs), wherein θs represents the solar incident angle and φs represents the solar azimuth angle, and the shading control matrix is obtained in combination with the indoor illumination demand value; a ventilation control model is established according to the shading control matrix, the ventilation quantity is calculated based on the indoor and outdoor temperature difference and the wind speed to obtain a ventilation opening degree matrix, and the ventilation opening degree matrix is input into a heat transfer model to calculate the hollow layer temperature distribution by solving the heat transfer differential equation to obtain a temperature control matrix; the shading control matrix, the ventilation opening degree matrix and the temperature control matrix are combined to construct a multi-objective optimization function F=λ1×E1+λ2×E2+λ3×E3, wherein E1 represents the unit area energy consumption target value, E2 represents the human comfort index value, E3 represents the control stability index value, and λ1, λ2 and λ3 represent balance factors; the multi-objective optimization function is input into a particle swarm optimization algorithm, and the regional energy-saving control parameters are calculated by iterative optimization; multi-objective optimization control is performed according to the regional energy-saving control parameters, and the initial deep reinforcement learning model is updated online to obtain a target deep reinforcement learning model. 2.The deep learning-based intelligent curtain wall energy-saving control method of claim 1, wherein, The method comprises the following steps: the curtain wall system is divided into hierarchical regions, monitoring points are set on the inner surface, the outer surface and the hollow layer according to the curtain wall orientation and the floor height to obtain a monitoring point layout scheme; Based on the monitoring point layout scheme, a temperature sensor, a light sensor, and an energy consumption sensor are installed and configured for each monitoring point to obtain a sensor network layout; According to the sensor network layout, a data acquisition module is called to acquire temperature data, including outdoor temperature, indoor temperature, and curtain surface temperature, and standardization processing is performed on the temperature data to obtain standardized temperature data; Based on the sensor network layout, a data acquisition module is called to acquire light data, including outdoor light intensity, indoor light intensity, and glass light coefficient, and standardization processing is performed on the light data to obtain standardized light data; Energy consumption data is collected through the energy consumption sensor in the sensor network layout, including shading coefficient, regional energy consumption, and air conditioning load, and standardization processing is performed on the energy consumption data to obtain standardized energy consumption data; The standardized temperature data, the standardized light data, and the standardized energy consumption data are input into a data preprocessing module for outlier detection and deletion to obtain a preprocessed data set; The preprocessed data set is constructed into a three-dimensional data array according to time sequence and region number, including time dimension, space dimension, and parameter dimension, and a sensor data matrix is obtained through tensor transformation. 3.The deep learning-based intelligent curtain wall energy-saving control method of claim 2, wherein, The sensor data matrix is used to calculate curtain energy consumption residual values, and the curtain energy consumption residual values are time-series segmented and K-means clustered to obtain energy consumption feature vectors and an energy consumption mode feature database, including: Principal component analysis is performed on the sensor data matrix to obtain energy consumption feature principal components; An energy consumption benchmark model is established based on the energy consumption feature principal components, a three-layer neural network structure is used to fit energy consumption data, network parameters are optimized through a back propagation algorithm, and energy consumption prediction values are obtained; Actual energy consumption data in the sensor data matrix is subtracted from the energy consumption prediction values to obtain a difference calculation result, and a moving average processing is performed on the difference calculation result to obtain curtain energy consumption residual values; The curtain energy consumption residual values are segmented, frequency domain features of each segment of data are extracted through fast Fourier transform, and the frequency domain features are dimensionally reduced to obtain time sequence features; The time sequence features are input into an improved K-means clustering model, the optimal number of clustering centers is determined using the silhouette coefficient and Davies-Bouldin index, iterative clustering calculation is performed, and energy consumption mode clustering results are obtained; According to the energy consumption mode clustering results, the mean vector, variance vector, peak coefficient, and skewness coefficient of each cluster are calculated, and energy consumption feature vectors are obtained through feature fusion; The energy consumption feature vectors are normalized to obtain compressed feature vectors, and the compressed feature vectors are associated with time stamps and spatial location information to obtain an energy consumption mode feature database. 4.The deep learning-based intelligent curtain wall energy-saving control method of claim 3, wherein, Based on the energy consumption feature vectors and the energy consumption mode feature database, a state space, an action space, and a reward function are set, an initial deep reinforcement learning model is trained, including: The energy consumption feature vector and the energy consumption mode feature database are combined to construct a state vector containing indoor temperature parameters, outdoor temperature parameters, light intensity parameters, energy consumption parameters and time parameters, and a state space representation matrix is obtained; According to the state space representation matrix, curtain control variables are extracted, and a first value range of the shading coefficient, a second value range of the ventilation opening degree, and a third value range of the hollow layer temperature are set, to obtain an action space parameter set; According to the indoor temperature parameters in the state space representation matrix, the light intensity parameters in the state space representation matrix, and the energy consumption parameters in the state space representation matrix, a reward function R=-w1×energy consumption value-w2×|actual room temperature-set room temperature|-w3×|actual illumination-set illumination| is constructed, to obtain a reward function model, wherein R represents the reward function, w1, w2, and w3 represent weight coefficients; Based on the state space representation matrix and the action space parameter set, a deep Q network is constructed, which contains an input layer, three hidden layers and an output layer, and the number of neurons in the hidden layers is 256, 128 and 64 respectively; The experience replay buffer size, the discount factor γ, the exploration factor ε and the learning rate α are set for the deep Q network, to obtain deep Q network training parameters; The state space representation matrix is input into the deep Q network, the Q value is calculated through the deep Q network, the immediate reward is calculated using the reward function model, and the target Q value is calculated based on the time difference algorithm, to obtain a loss function value; According to the loss function value, a back propagation operation is performed, the network parameters are updated using the Adam optimization algorithm, a trained deep Q network is obtained, and the trained deep Q network is evaluated on a validation data set. When the average reward value converges and the performance index on the validation set is stable, an initial deep reinforcement learning model is obtained. 5.The deep learning based intelligent curtain wall energy saving control method according to claim 1, characterized in that, According to the energy-saving control parameters of each region, multi-objective optimization control is performed, and the initial deep reinforcement learning model is updated online to obtain a target deep reinforcement learning model, which includes: The energy-saving control parameters of each region are imported into a control execution unit, and according to the control region corresponding to the curtain orientation and floor height, the angle adjustment of the sunshade louver, the ventilation opening degree adjustment and the temperature set value adjustment are performed to obtain a regional control execution matrix; According to the regional control execution matrix, a data collection time interval is set, indoor and outdoor temperature data are collected through a temperature sensor, indoor and outdoor light intensity data are collected through a light sensor, and regional energy consumption data are collected through an energy consumption sensor, to obtain an environmental parameter set; The environmental parameter set is normalized to obtain a standardized parameter matrix, and the standardized parameter matrix is input into the initial deep reinforcement learning model. After state space mapping, action space selection and reward function calculation, the optimization control vector is output, including the shading adjustment amount, the ventilation adjustment amount and the temperature adjustment amount of each region. The optimization control vector is decoded and mapped to restore the normalized control quantity to actual shading louver angle value, ventilation opening value and temperature set value, and is distributed according to the control area to obtain an execution control instruction; According to the execution control instruction, the execution mechanism is driven to operate to adjust the shading louver motor angle, the ventilation window opening and the hollow layer temperature, and the execution state data is collected to obtain system response data; Based on the system response data, energy saving indicators and comfort indicators are calculated, and the initial deep reinforcement learning model is updated online by using an incremental learning method to obtain a target deep reinforcement learning model. 6.The deep learning based intelligent curtain wall energy saving control method according to claim 5, characterized in that, The calculation of the energy saving indicators and the comfort indicators based on the system response data, and the online updating of the initial deep reinforcement learning model by using the incremental learning method to obtain the target deep reinforcement learning model, include: The energy consumption data in the system response data is time series statistical, the cumulative energy consumption value per unit area, the peak-valley energy consumption ratio and the energy consumption standard deviation of each control area within 24 hours are calculated, and the energy saving performance indicator matrix is obtained by weighted average; The indoor environment parameters in the system response data are evaluated and calculated, the predicted mean vote PMV is calculated based on temperature, humidity, air flow velocity and radiation temperature, and the visual comfort is calculated based on illumination uniformity and glare value, and the comfort performance indicator matrix is obtained; According to the energy saving performance indicator matrix and the comfort performance indicator matrix, a model performance evaluation function is constructed, and the performance of the initial deep reinforcement learning model is evaluated, and when the output value of the model performance evaluation function is less than a preset threshold, the model updating mechanism is activated to obtain an update trigger instruction; Based on the update trigger instruction, the system response data of continuous 24 hours is sampled to construct model training data containing state transition sequence, action sequence and reward sequence; The model training data is online feature extracted to obtain a dynamic feature vector, and the dynamic feature vector is input into the initial deep reinforcement learning model for incremental training to obtain a target deep reinforcement learning model.
7. A deep learning-based intelligent curtain wall energy-saving control system, characterized in that, The system for implementing the steps of the method of any one of claims 1 to 6 comprises: A collection module for collecting temperature data, illumination data and energy consumption data of each hierarchical region of the curtain wall system to obtain a sensor data matrix; A calculation module for calculating a curtain wall energy consumption residual value according to the sensor data matrix, and time series segmenting and K-means clustering the curtain wall energy consumption residual value to obtain an energy consumption feature vector and an energy consumption mode feature database; A training module for setting a state space, an action space and a reward function based on the energy consumption feature vector and the energy consumption mode feature database, and training an initial deep reinforcement learning model; An adjustment module for adjusting the adaptive droop coefficient of the curtain wall control region based on the initial deep reinforcement learning model, and adjusting the shading coefficient, the ventilation opening and the hollow layer temperature to obtain the energy saving control parameters of each region. An updating module is configured to perform multi-objective optimization control according to the energy-saving control parameters of the regions, and update the initial deep reinforcement learning model online to obtain a target deep reinforcement learning model.
8. A computer device comprising a memory and a processor, the memory having stored therein a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 6.
9. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Self-adaptive deep learning optimization energy-saving control algorithm for central air-conditioning cold supply system
CN114543273A
Zero-carbon building optimization design method based on deep reinforcement learning
CN114692265A