HVAC load mode prediction method based on time sequence decomposition and ensemble learning
Through time-series decomposition and integrated learning methods, the trend and seasonal components of HVAC load data are extracted, and cluster analysis and feature screening are solved, and the problem of inaccurate classification of HVAC load modes in the existing technology is achieved, achieving higher classification accuracy and model generalization ability.
Patent Information
- Application Number
- CN202510638098.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The prior art is difficult to accurately identify different operating modes in the HVAC load mode classification, resulting in large limitations in energy consumption optimization. The main problems are the lack of feature engineering, uncombined timing decomposition and lack of integrated learning.
The method based on time-series decomposition and ensemble learning is adopted, including steps: obtaining the original HVAC load data for preprocessing, performing time-series decomposition and extracting trends and seasonal components, obtaining the optimal cluster number through clustering analysis, performing feature screening and multi-classifier integrated learning model prediction.
It significantly improves the accuracy and stability of load mode classification, enhances the generalization ability of prediction models, and provides a high-reliability load mode recognition basis for building energy-saving control.
Smart Images

Figure CN120184949A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of HVAC load pattern prediction, and particularly to an HVAC load pattern prediction method based on time series decomposition and ensemble learning. Background Art
[0002] During the operation and management of building environments such as subway stations and rail transit hubs, the energy consumption of the Heating, Ventilation, and Air Conditioning (HVAC) system accounts for a relatively large proportion. How to reasonably classify HVAC load patterns to formulate more accurate energy-saving control strategies is an important topic in current research on energy conservation in intelligent buildings and rail transit. The load of subway stations has obvious time-dependence and periodic characteristics, and is affected by various factors such as passenger flow, environmental temperature and humidity. Traditional load pattern classification methods are difficult to accurately identify different operation modes, resulting in great limitations in energy consumption optimization.
[0003] There has been some research on load pattern classification in the prior art, and the main methods include classification methods based on clustering analysis and machine learning. For example, unsupervised clustering methods such as K-means and DBSCAN are used to directly classify load data. Traditional methods directly cluster the original load data, resulting in insufficient utilization of features such as time periodicity and trend, and poor stability of classification results. Classification methods based on machine learning have been applied in load pattern recognition, but these methods often do not combine time series decomposition and feature engineering, only use the original load data for classification, do not fully consider the time series characteristics of load data, and at the same time, the generalization ability of a single model is limited and is easily affected by changes in data distribution.
[0004] Although the above methods have achieved the classification of HVAC load patterns to a certain extent, there are still the following deficiencies: (1) Lack of feature engineering. Most methods directly use the original load data for clustering or classification, without fully extracting the long-term trend and periodic characteristics of the load, resulting in unstable classification effects; (2) Do not combine time series decomposition. Traditional clustering methods are difficult to accurately capture the trend and periodic characteristics of the load; (3) Lack of ensemble learning. Most existing classification methods are based on a single model, resulting in poor adaptability and insufficient generalization ability under different data conditions. It can be seen that most of the existing methods focus on a certain aspect of load pattern classification, rather than having a comprehensive and systematic framework.
[0005] Therefore, there is an urgent need for an HVAC load pattern prediction method based on time series decomposition and ensemble learning, which can more accurately identify different operation modes, improve the stability and generalization ability of the prediction model, and provide more reliable basic data support for subsequent energy-saving optimization. Summary of the Invention
[0006] To solve the above technical problems, on the one hand, the present invention provides an HVAC load pattern prediction method based on time series decomposition and ensemble learning, including the following steps: Step S1: Obtain the original HVAC load data, preprocess the original HVAC load data to generate preprocessed load time series data; Step S2: Based on the preprocessed load time series data, perform time series decomposition to extract the trend component and seasonal component, and generate an enhanced data set; Step S3: Based on the enhanced data set, obtain the optimal number of clusters through clustering analysis to obtain HVAC load pattern classification labels; Step S4: Based on the HVAC load pattern classification labels and the original HVAC load data, screen through correlation analysis to obtain a feature set; Step S5: According to the feature set, perform prediction through a multi-classifier ensemble learning model and combine a voting mechanism to obtain a load pattern prediction result.
[0007] Further, in step S2, based on the preprocessed load time series data, perform time series decomposition to extract the trend component and seasonal component, and generate an enhanced data set, specifically including: S21: Decompose the preprocessed load time series data into a trend component, a seasonal component, and a residual component by using the STL method; S22: Combine the trend component and the seasonal component with the original HVAC load data to form the enhanced data set.
[0008] Further, the calculation formula of the STL method is: ; In the formula, Y t represents the given load data, T t represents the trend component, S t represents the seasonal component, and R t represents the residual component.
[0009] Further, in step S3, based on the enhanced data set, obtain the optimal number of clusters through clustering analysis to obtain HVAC load pattern classification labels, including: S31: Determine the range of the number of clusters based on the data volume of the enhanced data set; S32: According to the enhanced data set, perform clustering analysis with each number of clusters within the range of the number of clusters respectively; S33: Calculate the evaluation indexes of the clustering results corresponding to each number of clusters respectively; S34: According to the evaluation indexes of the clustering results corresponding to each number of clusters, use the normalized clustering evaluation index fusion method to determine the optimal number of clusters; S35: Perform clustering analysis on the enhanced dataset with the optimal number of clusters to obtain the HVAC load pattern classification labels.
[0010] Further, in S34, according to the evaluation metrics of the clustering results corresponding to each number of clusters, the normalized clustering evaluation metric fusion method is used to determine the optimal number of clusters, including: S341: Obtain the evaluation metrics of the clustering results corresponding to each number of clusters, where the evaluation metrics include the silhouette coefficient, Gap Statistic value, DBI value, and CH value; S342: Perform normalization processing on the evaluation metrics of the clustering results corresponding to each number of clusters respectively to obtain the normalized scores of the evaluation metrics for each number of clusters; S343: Select the number of clusters with the highest normalized score of the evaluation metric as the optimal number of clusters.
[0011] Further, step S4, based on the HVAC load pattern classification labels and the original HVAC load data, through correlation analysis screening, obtain the feature set, including: Based on the HVAC load pattern classification labels and the original HVAC load data, perform continuous variable correlation analysis to obtain a preliminary feature set; Perform discrete variable correlation analysis according to the preliminary feature set to obtain the feature set.
[0012] Further, before performing the continuous variable correlation analysis, it also includes: For the time information in the original HVAC load data, use sine and cosine functions to perform periodic encoding on the time information to construct feature variables reflecting time periodicity; For the week information in the original HVAC load data, perform discrete encoding to generate 7-dimensional one-hot encoding features; For the passenger flow data in the original HVAC load data, perform non-linear transformation to eliminate scale differences; Perform correlation analysis screening on the feature variables reflecting time periodicity, 7-dimensional one-hot encoding features, non-linearly transformed passenger flow data, and other data of the original HVAC load data.
[0013] Further, in step S5, based on the feature set, through the prediction of the multi-classifier ensemble learning model and combined with the voting mechanism, obtain the load pattern prediction result, including: Based on the feature set, respectively obtain the HVAC load pattern classification labels predicted by each classifier through decision tree, random forest, and XGBoost; Through a hard voting mechanism, the prediction results of each classifier are aggregated, and the HVAC load pattern classification label with the highest frequency is selected as the final prediction result.
[0014] Furthermore, when training each classifier, the grid search cross-validation method is used to optimize the hyperparameters of each classifier. The embodiments of the present invention have the following technical effects: The HVAC load pattern prediction method based on time series decomposition and ensemble learning provided by the present invention. First, by preprocessing the original HVAC load data, the data time interval is unified and the interference of outliers and missing values is eliminated, providing a high-quality time series data basis for subsequent analysis. Secondly, based on time series decomposition, the trend component and seasonal component are extracted, enhancing the in-depth mining of the long-term change law and periodic characteristics of the load, making up for the defect that traditional methods do not fully capture the time series characteristics, and providing a more discriminative enhanced data set for clustering analysis. When generating classification labels through clustering analysis, a dynamic determination strategy for the optimal number of clusters is particularly introduced, solving the problem of pattern division deviation caused by fixed or empirical setting of the number of clusters in traditional methods, significantly improving the objectivity and accuracy of load pattern classification, and providing a more reliable classification label basis for subsequent feature screening and model training. Combining multi-dimensional feature extraction to screen the feature set strongly related to the classification label, avoiding the interference of redundant features on the model. Finally, a multi-classifier ensemble learning model combined with a voting mechanism is used for prediction. Through the complementary advantages and result fusion of different models, the stability and generalization ability of the prediction results are enhanced. This solution provides a highly reliable load pattern recognition basis for building energy-saving control through the coordination of time series feature enhancement, adaptive selection of the optimal number of clusters, and the ensemble learning framework, solving the technical problems of unstable classification effect, strong clustering subjectivity, and poor model adaptability in traditional methods. Description of the Drawings
[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 It is a step relationship diagram of the HVAC load pattern prediction method based on time series decomposition and ensemble learning provided by the embodiments of the present invention; Figure 2 It is a flowchart of the HVAC load pattern prediction method based on time series decomposition and ensemble learning provided by the embodiments of the present invention. Specific Embodiments
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0018] By comprehensively applying technical means such as time series decomposition, ensemble learning, multi-level feature extraction, periodic transformation, and feature enhancement, the present invention can achieve accurate and stable results in subway HVAC load pattern classification. The introduction of each technical means is to make up for the deficiencies of traditional methods. By decomposing data, integrating multiple learning models, considering time dependence, and enhancing the comprehensiveness of feature data, accurate identification and efficient classification of load patterns are achieved. Ultimately, this provides a solid foundation for the energy-saving optimization of the subway HVAC system and has strong applicability and generalization ability.
[0019] From Figure 1 and Figure 2 it can be seen that on the one hand, the present invention provides an HVAC load pattern prediction method based on time series decomposition and ensemble learning, including the following steps: To solve the above technical problems, on the one hand, the present invention provides an HVAC load pattern prediction method based on time series decomposition and ensemble learning, including the following steps: Step S1: Obtain the original HVAC load data, preprocess the original HVAC load data to generate preprocessed load time series data; Collect relevant time series data as the original HVAC load data, including but not limited to data such as load, passenger flow, outdoor air temperature, outdoor air humidity, date, etc. Each group of data corresponds to a timestamp, indicating the variation law of the data over time.
[0020] The original HVAC load data may be obtained from different sources. When collecting data from different sources, the time intervals may not be uniform, resulting in different time intervals for the load data obtained from each source. In irregularly sampled data, it is necessary to convert data with different time intervals into data with the same time interval. Resample the time series data to convert irregular time interval data into time series data with the same interval. The following methods can be adopted: (1) Forward filling: Fill the missing values with the values of the previous moment.
[0021] (2) Linear interpolation: Interpolate the missing values according to the variation trend between time points to obtain a smooth time series.
[0022] (3) Mean value interpolation: Fill the missing values according to the average value of adjacent data points.
[0023] In addition, it is also necessary to clean the data to complete data preprocessing. This includes filling missing values using a simple moving average method, removing outlier data, etc.
[0024] The moving average method fills in missing values by calculating the average value of the data within a time window that rolls before and after the missing value. Its effectiveness depends on the choice of window size. A smaller window is sensitive to data changes but is easily interfered by noise; a larger window helps to smooth the data but may delay the capture of information. Determine the window size according to the data change rate. Use a cross-validation method that simulates missing values to evaluate the filling effect under different window sizes. The specific method is as follows: Randomly delete a part of the data from the original complete data as "simulated missing", and use the moving average method with different window sizes \(n\) to fill it, and then calculate the root mean square error (RMSE) and coefficient of variation (CV) between the filling result and the original data to comprehensively measure the accuracy and stability of the filling.
[0025] After supplementing the missing values, it may also include removing outliers. Exemplarily, other methods such as the interquartile range rule and local outlier factor can be used. Removing outliers provides a high-quality time series data basis for subsequent analysis.
[0026] Step S2: Perform time series decomposition on the preprocessed load time series data to extract the trend component and seasonal component, and generate an enhanced dataset; In some embodiments, in step S2, performing time series decomposition on the preprocessed load time series data to extract the trend component and seasonal component, and generate an enhanced dataset, specifically includes: S21: Decompose the preprocessed load time series data into a trend component, a seasonal component, and a residual component using the STL method; The STL method is a technique for decomposing time series through the locally weighted regression (LOESS) method, which can decompose any input time series Y t and output it as the sum of three parts: the trend component T t 、the seasonal component S t and the residual component R t , respectively extracting the long-term evolution trend, periodic operation law, and random perturbation component, effectively reducing the interference between different components. In some embodiments, the calculation formula of the STL method is: ; In the formula, Y t represents the given load data, T t represents the trend component, S t represents the seasonal component, and R t represents the residual component.
[0027] The trend component Tt Represents the long-term change pattern in time series data, reflecting the changing trend of the data over time. It is obtained by applying Loess smoothing to the data after removing the seasonal component, as shown in the following formula: ; Seasonal component S t Represents the periodic fluctuations in the time series, reflecting the repeating pattern of the data within the period. It is obtained by applying Loess smoothing to the data after removing the trend component. The seasonal cycle length is represented by P, as shown in the following formula: ; Residual component R t Represents the random fluctuations remaining in the data after removing the trend and seasonal components, usually reflecting noise, abnormal data, or unpredictable factors, as shown in the following formula: ; Seasonal and Trend decomposition using Loess (STL, a seasonal and trend decomposition method based on locally weighted regression) is a key step in mining the internal laws of subway HVAC load data. In the formula, the trend component Characterizes the slow changes in the load over time, such as the baseline energy consumption offset caused by changes in the thermal performance of the building envelope; the calculation uses a locally weighted regression algorithm to smoothly extract the trend characteristics by fitting a low-order polynomial function within a sliding window. The seasonal component Reflects the repeating fluctuations within a fixed period, such as the sharp increase in cooling demand during peak morning and evening hours; during the calculation, the detrended data is segmented by the cycle length first, and the mean value of the same-phase points in each segment is taken to form a reference seasonal pattern, and then the high-frequency noise is eliminated through low-pass filtering. The residual component Contains random interferences such as measurement errors and sudden events. As a by-product of the decomposition process, it can be used for data quality assessment. Larger residuals often indicate abnormal events or special patterns not captured by the model. This formula establishes an additive decomposition model for the load data, providing a theoretical basis for subsequent component analysis and feature engineering.
[0028] This decomposition method breaks through the limitation that the traditional Fourier transform is only applicable to stationary signals and can effectively handle the non-linear and non-stationary characteristics of building loads. The explicit separation of the trend and seasonal components provides structured features for subsequent analysis. For example, the rate of change of equipment energy efficiency can be evaluated through the slope of the trend component, and the intensity of the periodic fluctuations can be quantified through the amplitude of the seasonal component, making the classification of load patterns more physically interpretable.
[0029] By performing time series decomposition on complex fluctuations, the long-term trend, seasonality, and residuals are extracted to enhance the data representation ability. The trend component reflects the impact of long-term factors such as building equipment aging and operation strategy adjustment on the load. The seasonal component captures the repetitive fluctuation patterns within fixed cycles such as daily and weekly. The residual component retains random interference and unmodeled factors. It effectively analyzes the inherent periodic characteristics (such as daily, weekly, and seasonal variations) and long-term change trends of the load data. Through the decomposed components, the load patterns can be analyzed more precisely, avoiding mixing periodic and trend changes into the noise. Clustering of the long-term trend can identify the long-term operation patterns of the subway HVAC system, while clustering of the seasonal component can identify the seasonal changes in daily operations. In this way, the potential connections between different load patterns can be deeply explored, improving the classification accuracy.
[0030] S22: Merge the trend component and the seasonal component with the original HVAC load data to form the enhanced dataset.
[0031] The decomposed trend component T t , seasonal component S t are fused with the preprocessed load time series data to construct an enhanced dataset, providing multi-dimensional feature inputs for clustering analysis, breaking through the limitations of traditional methods that only rely on the original data, and improving the prediction accuracy.
[0032] Step S3: Based on the enhanced dataset, by obtaining the optimal number of clusters, perform clustering analysis to obtain the HVAC load pattern classification labels; In some embodiments, in the step S3, based on the enhanced dataset, by obtaining the optimal number of clusters, perform clustering analysis to obtain the HVAC load pattern classification labels, including: S31: Determine the range of the number of clusters based on the data volume of the enhanced dataset; Exemplarily, if the data volume of the enhanced dataset is M groups, then the range of the number of clusters is set to [2, M / 5], and the integers within this range are taken as the number of clusters. According to the actual situation, the range of the number of clusters can be appropriately expanded or reduced.
[0033] S32: According to the enhanced dataset, perform clustering analysis respectively with each number of clusters within the range of the number of clusters; S33: Calculate the evaluation indicators of the clustering results corresponding to each number of clusters respectively; S34: According to the evaluation indicators of the clustering results corresponding to each number of clusters, use the normalized clustering evaluation index fusion method to determine the optimal number of clusters; In some embodiments, in the S34, according to the evaluation indicators of the clustering results corresponding to each number of clusters, use the normalized clustering evaluation index fusion method to determine the optimal number of clusters, including: S341: Obtain the evaluation metrics for the clustering results corresponding to each number of clusters. The evaluation metrics include the silhouette coefficient, Gap Statistic value, DBI value, and CH value. The silhouette coefficient is an indicator for evaluating the clustering quality of each sample, taking into account the compactness within the cluster and the separation between clusters. The value range of the silhouette coefficient is from -1 to 1, and the larger the value, the better the clustering effect. By calculating the silhouette coefficients under different K values, the K value corresponding to the maximum silhouette coefficient value is selected as the first number of clusters. Calculate the similarity difference between each sample and the same-class cluster and the nearest different-class cluster, and the maximum value of the overall mean corresponds to the optimal balance of classification compactness and separation. The present invention directly applies the existing silhouette coefficient calculation method for calculation to obtain the silhouette coefficients of the clustering results corresponding to each number of clusters.
[0034] The Gap Statistic is used to evaluate the rationality of the number of clusters. It judges whether the number of clusters is optimal by comparing the clustering effect of the clustering result with that of random data. If the clustering effect is significantly better than that of random data, it indicates that the current number of clusters is appropriate. The larger the Gap Statistic value, the better the number of clusters. First, generate multiple random data sets and calculate the sum of squared errors of these data sets under different numbers of clusters. The present invention directly applies the existing Gap Statistic value calculation method for calculation to obtain the Gap Statistic values of the clustering results corresponding to each number of clusters.
[0035] The Davies - Bouldin index (DBI) is used to evaluate the compactness and separation of clusters. The smaller the DBI value, the more compact the clustering and the greater the separation between clusters, and the better the clustering effect. Select the K value with the smallest DBI value as the third number of clusters. Evaluate the classification quality by calculating the ratio of the inter - cluster distance to the intra - cluster diameter. The smaller the value, the higher the inter - cluster separation and the better the intra - cluster compactness. DBI measures the clustering effect by calculating the similarity of each pair of clusters. The present invention directly applies the existing DBI value calculation method for calculation to obtain the DBI values of the clustering results corresponding to each number of clusters.
[0036] The Calinski - Harabasz index (CH index, also known as the variance ratio criterion) evaluates the quality of clustering by comparing the compactness within the cluster and the separation between clusters. The larger the CH index, the better the clustering result. Measure the classification effect by the ratio of the inter - cluster dispersion to the intra - cluster dispersion. The larger the value, the clearer the clustering structure. The present invention directly applies the existing CH value calculation method for calculation to obtain the CH values of the clustering results corresponding to each number of clusters.
[0037] S342: Normalize the evaluation metrics for the clustering results corresponding to each number of clusters respectively to obtain the normalized scores of the evaluation metrics for each number of clusters. Normalize the evaluation metrics corresponding to each number of clusters in the clustering results respectively, and normalize each metric to the range of [0, 1] in the following way.
[0038] For maximization metrics, i.e., evaluation metrics where a larger value indicates a better clustering result: silhouette coefficient, Gap Statistic value, and CH value, when normalizing the scores of these three evaluation metrics corresponding to each number of clusters in the clustering results, use the following formula: ; In the formula, represents the normalized score of this metric when the number of clusters is K, with a range of [0, 1], and M K represents the original calculated value of a certain evaluation metric when the number of clusters is K, min(M) represents the minimum value of a metric for all K values, and max(M) represents the maximum value of a metric for all K values.
[0039] Taking the silhouette coefficient as an example, use the above formula to calculate the silhouette coefficient corresponding to each number of clusters in the clustering results. represents the normalized score of the silhouette coefficient when the number of clusters is K, corresponding to in the subsequent formula, and M K represents the original calculated value of the silhouette coefficient when the number of clusters is K, min(M) represents the minimum value of the silhouette coefficient for all K values, and max(M) represents the maximum value of the silhouette coefficient for all K values.
[0040] Similarly, calculate the normalized score of the Gap Statistic value and the normalized score of the CH value .
[0041] For minimization metrics, i.e., evaluation metrics where a smaller value indicates a better clustering result, namely the DBI value, use the following formula to calculate.
[0042] ; represents the normalized score of the DBI value when the number of clusters is K, with a range of [0, 1], and DBI K represents the original calculated value of the DBI value when the number of clusters is K, min(DBI) represents the minimum value of the DBI value for all K values, and max(DBI) represents the maximum value of the DBI value for all K values.
[0043] Then, according to the normalized scores of each evaluation metric, use the following formula to calculate the normalized score of the evaluation metric for each number of clusters: ; S K represents the normalized score of the evaluation metric when the number of clusters is K.
[0044] S343: Select the number of clusters with the highest normalized score of the evaluation index as the optimal number of clusters. This method of fusing normalized clustering evaluation indexes ensures that the load pattern division not only conforms to the internal structure of the data but also meets the requirements of engineering applications for the stability of classification results.
[0045] S35: Perform clustering analysis on the enhanced data set with the optimal number of clusters to obtain the HVAC load pattern classification labels.
[0046] Exemplarily, first standardize and normalize the trend component and seasonal component in the enhanced data set to eliminate the influence of different feature scales.
[0047] Standardization is to make the data follow the standard normal distribution by subtracting the mean from the data and dividing by the standard deviation. Normalization is to scale the data to the interval [0, 1].
[0048] Exemplarily, select a suitable clustering method (such as K-Means, K-Medoids, GMM, etc.) to cluster the standardized and normalized trend component and seasonal component data. The clustering result will assign a cluster label to each data point, indicating the category it belongs to. During the clustering process, according to the similarity of the data, the load data with similar patterns are grouped into one category, and the corresponding clustering labels are generated as the HVAC load pattern classification labels. Exemplarily, the conforming patterns are divided into Pattern 1, Pattern 2, …, or peak mode, flat peak mode, etc.
[0049] Step S4: Based on the HVAC load pattern classification labels and the original HVAC load data, obtain a feature set through correlation analysis screening; Use the obtained HVAC load pattern classification labels as new features and add them to the original HVAC load data to form a new data set for correlation analysis.
[0050] In some embodiments, step S4, based on the HVAC load pattern classification labels and the original HVAC load data, obtain a feature set through correlation analysis screening, including: In some embodiments, before performing the continuous variable correlation analysis, it further includes: S4A: For the time information in the original HVAC load data, use sine and cosine functions to perform periodic encoding on the time information to construct feature variables reflecting time periodicity; the feature variables reflecting time periodicity retain the periodic characteristics of time, making it easier for the model to identify the time dependence of the load pattern.
[0051] S4B: Perform discrete encoding on the day-of-week information in the original HVAC load data to generate 7-dimensional one-hot encoded features; perform discrete encoding on the day-of-week information in the original HVAC load data to generate 7-dimensional one-hot encoded features; the original day-of-week data (represented by numbers from 1 to 7) may mislead the model into thinking that the relationship between day 1 and day 2 is linear or sequential, while in fact they are just different categories. Turning each day of the week into an independent feature eliminates this false "sequential relationship".
[0052] S4C: Perform a non-linear transformation on the passenger flow data in the original HVAC load data to eliminate scale differences; perform a Log transformation on the number of inbound passengers and the number of outbound passengers to prevent the large difference in the order of magnitude between the passenger flow data and other data from affecting the analysis results, improve the classification accuracy, and thus enhance the model's learning ability regarding the relationships between different features.
[0053] S4D: Conduct a correlation analysis and screening on the feature variables reflecting time periodicity, the 7-dimensional one-hot encoded features, the non-linearly transformed passenger flow data, and other data of the original HVAC load data.
[0054] Multi-dimensional feature extraction aims to construct a system of influencing factors that comprehensively represents the load pattern. In the time feature transformation, the hour information is mapped to a continuous periodic variable through sine-cosine transformation, avoiding feature mutations caused by treating 00:00 and 23:59 as discrete points; the day-of-week information uses one-hot encoding to generate a vector composed of 0 / 1, retaining the independent influence of each working day. The passenger flow data undergoes Log transformation to prevent the large difference in the order of magnitude from affecting the analysis results.
[0055] S41: Based on the HVAC load pattern classification labels and the original HVAC load data, conduct a correlation analysis of continuous variables to obtain a preliminary feature set; The correlation analysis of continuous variables uses the method of analysis of variance to test whether there are significant mean differences among continuous variables under different categories. If the mean differences among different categories are large, it indicates a strong relationship between this continuous variable and the category output variable. When the input variable is continuous and the output variable is categorical, one-way analysis of variance (One-Way ANOVA) can be used. Through ANOVA, the differences between each continuous input variable and the output category are tested, thereby obtaining the correlation between various category data in the original HVAC load data and the HVAC load pattern classification labels, that is, screening out the factors that have an impact on the HVAC load pattern classification labels for subsequent prediction calculations, avoiding inputting too many irrelevant factors and reducing the prediction efficiency. In the present invention, when p ≤ significance level (such as 0.05), a preliminary feature set is obtained by screening from the original HVAC load data.
[0056] S42: Perform discrete variable correlation analysis based on the preliminary selected feature set to obtain the feature set.
[0057] Similarly, perform discrete variable correlation analysis based on the preliminary selected feature set. Exemplarily, select the Chi-Square Test to evaluate the correlation between various types of data and the HVAC load pattern classification label in the preliminary selected feature set, and also take p ≤ significance level (such as 0.05).
[0058] Merge the preliminary selected feature set screened by discrete variable correlation analysis with the HVAC load pattern classification label to generate a feature set.
[0059] Step S5: Based on the feature set, through a multi-classifier ensemble learning model prediction, combined with a voting mechanism, obtain the load pattern prediction result.
[0060] In some embodiments, in step S5, based on the feature set, through a multi-classifier ensemble learning model prediction, combined with a voting mechanism, obtaining the load pattern prediction result includes: S51: Based on the feature set, respectively obtain the HVAC load pattern classification labels predicted by each classifier through a decision tree, a random forest, and XGBoost. The ensemble learning framework improves the prediction robustness through heterogeneous model complementarity. The decision tree uses the information gain ratio to select splitting features and constructs an intuitive load pattern discrimination rule, but is sensitive to noise data and prone to overfitting. The random forest constructs multiple decision trees through Bootstrap sampling, introduces feature randomness to reduce variance, but may ignore subtle local patterns. XGBoost adopts the gradient boosting strategy to iteratively optimize the loss function, controls the model complexity through a regularization term, is good at capturing non-linear relationships but takes a long time to train.
[0061] In some embodiments, when training each classifier, use the grid search cross-validation method to perform hyperparameter tuning on each classifier. S52: Through a hard voting mechanism, aggregate the prediction results of each classifier, and select the HVAC load pattern classification label with the highest frequency as the final prediction result.
[0062] The hard voting mechanism requires each base classifier to make predictions independently, and the final result selects the category with the most votes. This democratic decision-making mechanism effectively suppresses the influence of misjudgments of individual models. For example, when two classifiers determine "peak mode" and one determines "flat peak mode", the system finally outputs "peak mode". During the training stage, stratified sampling is used to ensure the balance of samples of each category and avoid the voting process being biased towards high-frequency categories.
[0063] Ensemble learning conducts voting selection by combining multiple base classifiers, improving the robustness and generalization ability of the model. Each classifier can analyze the load pattern from different perspectives, and the voting mechanism can reduce the error of individual classifiers by aggregating the judgments of different models. The model integration strategy gives full play to the advantages of each algorithm. Decision trees provide highly interpretable rules, random forests ensure generalization performance, and XGBoost captures complex non-linear relationships. This combination enables the system to handle both linearly separable basic patterns and identify complex working conditions with multi-factor coupling, significantly enhancing the adaptability in different building scenarios. It improves the robustness and generalization ability of the model, can be applied in different subway stations and different operating conditions, and improves the applicability of the method.
[0064] Exemplarily, the construction and training of the ensemble learning model are as follows: Features with higher feature importance indicate that they contribute more to the model prediction, and these features are usually preferentially selected during feature selection. Train each classifier to predict the class labels to which different features belong. By integrating different classifiers, the classification accuracy and generalization ability are improved.
[0065] (1) Data partitioning Partition the dataset into a training set and a test set. For example, use 70% of the data for training and 30% for testing to ensure that the training and test datasets are representative, thereby effectively avoiding overfitting.
[0066] (2) Model construction and training Step A: Hyperparameter tuning Adjust the hyperparameters for each classifier. The specific hyperparameters to be adjusted include: Decision tree: Adjust the maximum depth of the tree and the minimum number of samples for splitting at each node.
[0067] Random forest: Adjust the number of decision trees and the maximum number of features used for each tree.
[0068] XGBoost: Adjust hyperparameters such as the learning rate, maximum tree depth, number of trees, etc., and explicitly set the evaluation metrics.
[0069] The Grid Search with Cross-Validation method is adopted. This method exhaustively enumerates the preset hyperparameter combinations, conducts cross-validation evaluation under each combination, and selects the parameter configuration with the optimal comprehensive performance as the setting of the final model. The performance evaluation of each parameter combination is achieved through K-fold cross-validation, that is, the dataset is divided into Z subsets. In each iteration, (Z - 1) subsets are used for training, and the remaining one subset is used for validation. After Z cycles, the average evaluation index is taken as the final score of this parameter combination. This method can effectively reduce the accidental influence caused by different data partitions and improve the stability and reliability of parameter selection. Finally, the optimal parameter configuration selected by this method is used in the subsequent model training and integration process.
[0070] Step B: Model Training In some embodiments, when training each classifier, the Grid Search with Cross-Validation method is used to optimize the hyperparameters of each classifier respectively.
[0071] The training set is trained using three classification models: Decision Tree, Random Forest, and XGBoost respectively. Each classifier model adopts different learning strategies to capture the patterns in the data: Decision Tree: By recursively splitting the feature space, the best splitting feature is selected using information gain or Gini index.
[0072] Random Forest: By integrating multiple decision trees and voting, the stability of the model is improved.
[0073] XGBoost: Adopts the gradient boosting tree method to gradually optimize the model to improve the prediction accuracy.
[0074] (3) Ensemble Learning and Classification Prediction The ensemble learning method is used to improve the model performance. An ensemble voting classifier is constructed by integrating each classifier (Decision Tree, Random Forest, and XGBoost) to complete the overall construction of the ensemble learning model. Integration is carried out during the training process to improve the prediction stability and accuracy of the final model. The voting classifier adopts the hard voting method, makes a voting decision based on the prediction results of each classifier, and selects the HVAC load pattern with the highest frequency as the final prediction result.
[0075] The ensemble learning model is trained and predicted on the test set. Multiple metrics are used to evaluate the performance of the ensemble learning model.
[0076] Accuracy: The proportion of samples correctly predicted by the model in the total samples.
[0077] Precision: The proportion of samples actually being positive among the samples predicted as positive.
[0078] Recall: The proportion of samples that are actually positive and are correctly predicted as positive.
[0079] F1-Score: The harmonic mean of precision and recall, comprehensively considering the performance of the model on positive samples.
[0080] Cross-validation: Use cross-validation to evaluate the ensemble learning model, ensure the stable performance of the model on different data subsets, and further avoid overfitting.
[0081] The present invention constructs a systematic framework of "time series decomposition - clustering analysis - feature engineering - ensemble prediction". First, by preprocessing the original HVAC load data, unifying the data time interval and eliminating the interference of outliers and missing values, it provides a high-quality time series data basis for subsequent analysis; second, based on time series decomposition, the trend component and seasonal component are extracted, enhancing the in-depth mining of the long-term change law and periodic characteristics of the load, making up for the defect that traditional methods do not fully capture the time series characteristics, and providing a more discriminative enhanced data set for clustering analysis; further, when generating classification labels through clustering analysis, a dynamic determination strategy for the optimal number of clusters is particularly introduced, solving the problem of pattern division deviation caused by fixed or empirical setting of the number of clusters in traditional methods, significantly improving the objectivity and accuracy of load pattern classification, and providing a more reliable classification label basis for subsequent feature screening and model training; combining multi-dimensional feature extraction to screen the feature set strongly related to the classification label, avoiding the interference of redundant features on the model; finally, a multi-classifier ensemble learning model combined with a voting mechanism is used for prediction. Through the complementary advantages and result fusion of different models, the stability and generalization ability of the prediction results are enhanced. This solution, through the coordination of time series feature enhancement, adaptive selection of the optimal number of clusters and the ensemble learning framework, provides a highly reliable load pattern recognition basis for building energy-saving control, and solves the technical problems of unstable classification effect, strong clustering subjectivity and poor model adaptability of traditional methods.
[0082] On the other hand, the present invention also provides an HVAC load pattern prediction system based on time series decomposition and ensemble learning, which executes the HVAC load pattern prediction method based on time series decomposition and ensemble learning described in any one of the above, including: an original data acquisition module, a data enhancement module, an HVAC load pattern classification label determination module, a feature set generation module, and an output module; The original data acquisition module is used to acquire the original HVAC load data, preprocess the original HVAC load data, and generate the preprocessed load time series data; the data enhancement module is connected to the original data acquisition module and is used to perform time series decomposition based on the preprocessed load time series data, extract the trend component and the seasonal component, and generate an enhanced data set; the HVAC load pattern classification label determination module is connected to the data enhancement module and is used to perform clustering analysis based on the enhanced data set by obtaining the optimal number of clusters to obtain the HVAC load pattern classification label; the feature set generation module is connected to the HVAC load pattern classification label determination module and the original data acquisition module and is used to obtain a feature set through correlation analysis screening based on the HVAC load pattern classification label and the original HVAC load data; the output module is connected to the feature set generation module and is used to obtain the load pattern prediction result through multi-classifier ensemble learning model prediction and in combination with a voting mechanism based on the feature set.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. HVAC load pattern prediction method based on time series decomposition and ensemble learning, characterized in that: The following steps are involved: Step S1: obtaining original HVAC load data, preprocessing the original HVAC load data, and generating preprocessed load time series data; Step S2: performing time series decomposition based on the preprocessed load time series data, extracting trend components and seasonal components, and generating an enhanced data set; Step S3: Based on the enhanced data set, performing cluster analysis by obtaining an optimal number of clusters to obtain a classification label of the HVAC load mode; Step S4: Based on the HVAC load mode classification label and the original HVAC load data, a feature set is obtained by screening through correlation analysis; Step S5: Based on the feature set, a multi-classifier ensemble learning model is used for prediction, combined with a voting mechanism, to obtain a load pattern prediction result.
2. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 1 is characterized in that: In the step S2, time series decomposition is performed based on the preprocessed load time series data, trend components and seasonal components are extracted, and an enhanced data set is generated, which specifically includes: S21: decomposing the preprocessed load time series data into trend components, seasonal components and residual components using the STL method; S22: merging the trend component and the seasonal component with the original HVAC load data to form the enhanced data set.
3. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 2 is characterized in that: The calculation formula of the STL method is: ; Where Y t Represents the given load data, T t represents the trend component, S t represents the seasonal component, R t represents the residual component.
4. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 1 is characterized in that: In step S3, based on the enhanced data set, cluster analysis is performed by obtaining the optimal number of clusters to obtain the HVAC load mode classification label, including: S31: Determine a range of cluster numbers based on the data volume of the enhanced data set; S32: performing cluster analysis with each cluster number within the cluster number range according to the enhanced data set; S33: Calculate the evaluation index of the clustering result corresponding to each cluster number respectively; S34: according to the evaluation index of the clustering result corresponding to each cluster number, a normalized clustering evaluation index fusion method is adopted to determine the optimal cluster number; S35: Performing cluster analysis on the enhanced data set with the optimal number of clusters to obtain the HVAC load mode classification label.
5. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 4 is characterized in that: In S34, according to the evaluation index of the clustering result corresponding to each clustering number, a normalized clustering evaluation index fusion method is used to determine the optimal clustering number, including: S341: Obtaining evaluation indicators of clustering results corresponding to each cluster number, wherein the evaluation indicators include silhouette coefficient, GapStatistic value, DBI value and CH value; S342: normalizing the evaluation index of the clustering result corresponding to each cluster number to obtain a normalized score of the evaluation index of each cluster number; S343: Select the cluster number with the highest normalized score of the evaluation index as the optimal cluster number.
6. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 1, characterized in that: Step S4, based on the HVAC load mode classification label and the original HVAC load data, a feature set is obtained by correlation analysis screening, including: Based on the HVAC load mode classification label and the original HVAC load data, a continuous variable correlation analysis is performed to obtain a preliminary feature set; A discrete variable correlation analysis is performed based on the initially selected feature set to obtain the feature set.
7. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 6, characterized in that: Before the continuous variable correlation analysis is performed, the method further includes: For the time information in the original HVAC load data, sine and cosine functions are used to periodically encode the time information to construct a characteristic variable reflecting the time periodicity; Perform discrete encoding on the week information in the original HVAC load data to generate a 7-dimensional one-hot encoding feature; Performing nonlinear transformation on the passenger flow data in the original HVAC load data to eliminate scale differences; The characteristic variables reflecting the time periodicity, the 7-dimensional unique hot encoding features, the nonlinearly transformed passenger flow data, and other data of the original HVAC load data are subjected to correlation analysis and screening.
8. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 1, characterized in that: In step S5, based on the feature set, a multi-classifier ensemble learning model is used for prediction, combined with a voting mechanism, to obtain a load pattern prediction result, including: According to the feature set, the classification labels of the HVAC load patterns predicted by each classifier are obtained by using decision tree, random forest and XGBoost respectively; Through the hard voting mechanism, the prediction results of each classifier are aggregated, and the HVAC load mode classification label with the highest frequency is selected as the final prediction result.
9. The HVAC load pattern prediction method based on time series decomposition and ensemble learning according to claim 8, characterized in that: When training each classifier, the grid search cross-validation method is used to tune the hyperparameters of each classifier.
Citation Information
Patent Citations
Wave power generation typical scene generation method based on evaluation indexes
CN112308412A
Regional medium-term load prediction method and device based on clustering electric quantity curve decomposition
CN113449933A
Power distribution network medium-term load decomposition-set prediction method and system
CN115062864A
Electric vehicle charging load prediction method based on empirical mode decomposition
CN117993603A
Engine group fault mode recognition method and device based on density clustering-support vector machine and multi-moment prediction weighting and electronic equipment
CN119312181A