Data feature engineering processing method for improving building cooling and heating load prediction precision
Through the MRMR-PCA feature engineering processing method, the problem of feature redundancy in the building load prediction model is solved, and the prediction accuracy and efficiency are improved.
Patent Information
- Application Number
- CN202311509472.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
In the building load prediction model based on machine learning, there are many input features and high redundancy, resulting in low prediction accuracy and efficiency.
The feature engineering processing method based on MRMR-PCA is adopted to select subsets of features with strong correlation and low redundancy with the target variables through the MRMR algorithm, and the PCA algorithm is used to convert these features into comprehensive feature sets that are independent and orthogonal to each other, thereby reducing the mutual influence of redundant information.
Effectively filter out important variables that are meaningful for prediction, reduce the number of input variables, and improve the accuracy and efficiency of building load prediction.
Smart Images

Figure CN120011773A_ABST
Abstract
Description
Technical Field
[0001] The new patent of this invention relates to the field of building HVAC system prediction technology, especially a data feature engineering processing method to improve the accuracy of building cooling and heating load prediction based on machine learning. Background Art
[0002] According to statistics, in 2020, the total energy consumption of the whole process of buildings in China was 2.27 billion tce, and the total carbon emissions were 5.08 billion tco2, accounting for 45.5% of the total energy consumption and 50.9% of the total carbon emissions in China, respectively. Among them, the energy consumption in the building operation stage was 1.06 billion tce, and the carbon emissions were 2.82 billion tco2, accounting for 21.3% of the total energy consumption and 28.2% of the total carbon emissions in China, respectively. Therefore, improving building energy efficiency and reducing building operation energy consumption can save huge energy and reduce a lot of carbon emissions. Due to changes in factors such as internal and external disturbances, the demand for building cooling and heating loads will also change continuously. It is necessary to adjust the operation strategy to meet the thermal comfort demand while maintaining the high-efficiency operation of the unit. Due to the high delay and large hysteresis characteristics of the building HVAC system, the air-conditioning system that relies solely on negative feedback regulation may not be able to guarantee indoor thermal comfort and energy-saving operation of equipment. Therefore, accurate building load prediction is of great significance to the efficient operation of the building HVAC system and to reduce energy consumption.
[0003] At present, there are two main methods for building load prediction: load prediction based on physical methods and prediction methods based on data-driven methods. The load prediction calculation based on physical methods is mainly based on the principle of unstable heat transfer, using building and environmental information, such as external climate conditions, building thermal parameters, power density of indoor electrical equipment, personnel density, indoor and outdoor shading and human behavior, and HVAC equipment operation mode as input, to accurately calculate the energy consumption of all building components step by step. The prediction method based on machine learning obtains the statistical law between input and output through a large amount of historical data. In today's booming artificial intelligence, the calculation amount of load obtained by using energy consumption simulation software is large and time-consuming. In subsequent actual projects, it cannot be well integrated with intelligent control systems to achieve the purpose of real-time regulation. Therefore, machine learning methods are currently more used to predict building loads.
[0004] A typical HVAC system is a complex nonlinear network containing hundreds of features, reflecting all aspects of the system. In addition to being affected by static parameters such as the building's shape and structure, thermal parameters of the envelope structure, and thermal characteristics of objects inside the building, it is also mainly affected by dynamic parameters such as outdoor meteorological parameters, the number of indoor people and human behavior, lighting power, and equipment power. At the same time, due to the influence of thermal hysteresis, the historical data of these parameters will also have an impact on load forecasting. When training a machine learning model, the input feature set determines the upper limit of the prediction performance of machine learning. When there is a correlation between input features, information redundancy will occur. The redundant parameters copy most or all of the information contained in one or more other parameters, resulting in large deviations and fluctuations in the prediction results. This redundancy of data features is more obvious in building load forecasting. For example, there is a strong correlation between outdoor temperature and solar radiation. In some buildings, an increase in the occupancy rate of people in the room will inevitably lead to an increase in equipment power. Therefore, the reasonable selection of input features is crucial to improving the accuracy and efficiency of building load forecasting. Summary of the invention
[0005] In order to solve the problem of many input features and high redundancy in the building load prediction model based on machine learning, the purpose of the present invention is to provide a feature engineering data processing method for improving the generalization ability and prediction accuracy of building cooling and heating load prediction, which can effectively screen out important variables that are meaningful for prediction from the many input features that affect the building cooling and heating load, while reducing the redundancy between variables, effectively improving the prediction accuracy of the model.
[0006] To achieve the above object, the present invention proposes a feature engineering processing method based on MRMR-PCA, which uses the Max-Relevance Min-Redundancy (MRMR) algorithm to select a feature subset with strong correlation with the target variable and low redundancy with other feature variables, and further uses the Principal Component Analysis (PCA) algorithm to convert the feature subset selected by the MRMR algorithm into a comprehensive feature set that is independent and orthogonal to each other, mainly including the following contents:
[0007] S1. Construct a building cooling and heating load prediction data set: collect building cooling and heating load data, meteorological data, and building internal disturbance data;
[0008] S2. Data preprocessing: The collected data may be abnormal or missing. The IQR method is used to identify and process data outliers, and the interpolation method is used to fill in missing values and outliers, and the characteristic variables are standardized.
[0009] S3. Feature selection: Use the MRMR algorithm to calculate the maximum correlation-minimum redundancy of the variables, and select the feature subset that has a strong correlation with the target variable and a low redundancy with other feature variables;
[0010] S4. Feature extraction: PCA algorithm is used to transform the feature subsets selected by MRMR algorithm into independent and orthogonal comprehensive feature sets, further reducing the mutual influence of redundant information. The data processed by MRMR-PCA is the final feature variable input into the building load prediction model.
[0011] It can be seen from the above technical scheme that the beneficial effects of the present invention are: first, load data, meteorological data and building internal disturbance data are collected by sensors and smart energy systems to construct a building cooling and heating load prediction data set; second, outliers are identified by the IQR method, and interpolation is used to fill in outliers and missing values, and then the cleaned data is standardized; third, the MRMR algorithm is used to calculate the maximum correlation-minimum redundancy of the variables, and a feature subset with strong correlation with the target variable and low redundancy with other feature variables is selected, which can better solve the interference and redundancy problems between multiple variables in load forecasting; fourth, the PCA algorithm is used to convert the feature subsets screened by the MRMR algorithm into independent and orthogonal comprehensive feature sets, thereby further reducing the mutual influence of redundant information, reducing the number of input variables, and improving prediction accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 Flowchart of the feature engineering method for MRMR-PAC data DETAILED DESCRIPTION
[0013] A data feature engineering processing method based on MRMR-PAC uses the MRMR algorithm to select a feature subset with strong correlation with the target variable and low redundancy with other feature variables, and further uses the PCA algorithm to convert the feature subset selected by the MRMR algorithm into a comprehensive feature set that is independent and orthogonal to each other. Specifically, the following steps are included:
[0014] S1. Construct a building cooling and heating load prediction data set: collect building cooling and heating load data, meteorological data, and building internal disturbance data;
[0015] S2. Data preprocessing: The collected data may be abnormal or missing. The IQR method is used to identify and process data outliers, and the interpolation method is used to fill in missing values and outliers, and the characteristic variables are standardized.
[0016] S3. Feature selection: Use the MRMR algorithm to calculate the maximum correlation-minimum redundancy of the variables, and select the feature subset with strong correlation with the target variable and low redundancy with other feature variables;
[0017] S4. Feature extraction: PCA algorithm is used to transform the feature subsets selected by MRMR algorithm into independent and orthogonal comprehensive feature sets, further reducing the mutual influence of redundant information. The data processed by MRMR-PCA is the final feature variable input into the building load prediction model.
[0018] Further, the specific data of step S1 include: the building cooling and heating load data include the water supply flow and water supply temperature difference at the thermal inlet, the meteorological data include the outdoor temperature, dew point temperature, relative humidity, wind speed, wind direction, horizontal plane total radiation, normal direction solar radiation, horizontal plane scattered radiation, building internal disturbance data include the number of people in the room, lighting power and equipment power, these data can be obtained by installing corresponding sensors and smart energy systems;
[0019] Further, the IQR method described in step S2 determines whether the data is an outlier based on the quartile range of the number, and the specific steps are as follows:
[0020] S21. Calculate the first quartile (Q1), the second quartile (median), and the third quartile (Q3) of the data;
[0021] S22. Calculate IQR = Q3 - Q1;
[0022] S23. Calculate the lower limit lower = Q1-k*IQR and the upper limit upper = Q3+k*IQR, k can be 1.5 or 3, when k = 3 represents an extreme outlier, when k = 1.5 represents a moderate outlier;
[0023] S24. If the data value is less than lower or greater than upper, it is considered an outlier;
[0024] The interpolation method in step S2 to fill the missing values refers to using the mean of two adjacent moments to replace the missing values;
[0025] S25. Standardize the variables according to formula (1):
[0026]
[0027] Where x is the original value of a feature, x * is the eigenvalue after standardization, μ is the average value of the feature in all samples, and σ is the standard deviation of the feature in all samples;
[0028] Further, in step S3, the maximum correlation-minimum redundancy of the variables is calculated using the MRMR algorithm, and a feature subset with a strong correlation with the target variable and a low redundancy with other feature variables is selected. The specific steps are as follows:
[0029] S31. In the MRMR algorithm, correlation and redundancy are measured by mutual information, that is, the degree of reduction in the uncertainty (measured by information entropy) of the output variable Y after knowing a variable X. The calculation formula for the mutual information of two continuous variables is shown in formula (2):
[0030]
[0031] In the formula, I(X, Y) is the mutual information of variables X and Y, P(x, y) is the joint probability density function of X and Y, and the marginal probability density distribution function of P(x) and P(y). The value range of the calculation result of mutual information is [0, 1]. The larger its value, the greater the correlation between the two variables.
[0032] According to the maximum correlation criterion, I(X i , Y) in an appropriate order to search for the best feature variable related to the output variable Y, and use formula (3) to calculate its correlation size:
[0033]
[0034] S32. In order to eliminate the redundant information between the selected variables in the feature set and select mutually exclusive features, the minimum redundancy criterion is calculated using formula (4):
[0035]
[0038] S33. Integrate formula (4) and (5) by addition to find the set S with maximum correlation and minimum redundancy:
[0039] Max φ(F), φ(D, R)=DR (5)
[0040] S34. Use the incremental search method to obtain an approximate optimal solution: that is, assuming that S p-1 feature subsets, and need to be in the remaining XS p-1 The pth feature is selected from the features, and feature selection is performed by maximizing φ(D, R), that is, maximizing formula (6):
[0041]
[0042] Furthermore, in step S4, the PCA algorithm is used to transform the feature subsets screened by the MRMR algorithm into independent and orthogonal comprehensive feature sets, so as to further reduce the mutual influence of redundant information. The specific steps are as follows:
[0043] S41. Assume that p feature variables are selected after the MRMR algorithm, the feature matrix is represented by X, and the variables are standardized according to formula (7) to obtain the standardized matrix B:
[0044]
[0045] In the formula,
[0046] S42. Solve the covariance of the standardized matrix according to formulas (8) and (9):
[0047]
[0048]
[0049] S43. Calculate the eigenvalues of the covariance matrix R and the corresponding orthogonalized eigenvectors:
[0050] According to the calculation formula of the correlation coefficient matrix R, we can get r ij =r ji According to linear algebra knowledge, the correlation coefficient matrix R is a real symmetric matrix. There must be a standard orthogonal matrix U and a diagonal matrix Λ, such that:
[0051]
[0052] where λ1, λ2, ..., λ p are the eigenvalues of the correlation coefficient matrix R, and their corresponding unit eigenvectors are U1, U2, ..., U p ;
[0053] S44. Perform linear combination of unit eigenvectors according to formula (11) to form new eigenvariables:
[0054]
[0055] In the formula, F i is the i-th characteristic variable after spatial transformation, i = 1, 2, ..., p. From the knowledge of linear algebra, we know that the new characteristic variables are orthogonal and uncorrelated with each other.
[0056] S45. Calculate the variance of the new feature variable according to formula (12):
[0057]
[0058] From step S43, we know that Var(F i )=λ i , that is, the variance of the new feature variable is the eigenvalue of the covariance matrix. For this reason, it is assumed that the order of the eigenvalues of the covariance matrix is λ1≥λ2≥λ3…≥λ p ≥0, then the variance order of the new characteristic variable composed of the eigenvectors corresponding to its eigenvalues is also Var(F1)≥Var(F2)≥Var(F3)…≥Var(F p)≥0. In this case, F1 is called the first principal component, F2 is called the second principal component, ..., F p It is called the pth principal component.
[0059] S46. Calculate the principal component contribution rate and accumulation according to formulas (13) and (14) to determine how many principal components to retain:
[0060]
[0061]
[0062] In the formula, F i is the contribution rate of the i-th principal component, η i is the cumulative contribution rate of the first i principal components.
[0063] Generally, the eigenvalues λ1, λ2, ..., λ1 whose cumulative contribution rate reaches 85% to 95% are selected. m The corresponding 1st, 2nd, ..., m (m≤p) variables are taken as the principal components, which are the final data feature engineering processing results.
[0064] In order to better understand the present invention, the present invention is verified by distance as follows:
[0065] The original data set is composed of the hourly heat load data of a building measured from January 25, 2023 to February 25, 2023, the hourly meteorological data released by the local meteorological station, and the actually measured hourly number of people in the room. The model input parameters are considered: meteorological data include outdoor temperature, dew point temperature, relative humidity, wind speed, wind direction, horizontal radiation, direct radiation, and scattered radiation. For dynamic parameters, the historical data of the prediction time and the previous 24 hours are considered. For the historical heat load data, the time dimension of 72 hours before the prediction time is considered, totaling 297 input feature variables.
[0066] Data preprocessing is performed for each type of data: using the IQR method to identify outliers, interpolation to fill in missing values and replace anomalies, and standardizing the cleaned data.
[0067] The maximum correlation-minimum redundancy of 297 variables was calculated using the MRMR algorithm, and then 41 characteristic variables were screened out using the incremental search method.
[0068] The PCA algorithm is used to further transform the 41 features and variables screened by the MRMR algorithm into 41 principal components, and the cumulative contribution rate of each component is calculated. The principal components with a cumulative contribution rate of 95% are selected - the first 31 principal components as the final input features.
[0068] The input features selected by the MRMR-PCA algorithm are used as the input of the building heat load prediction model. Three common machine learning algorithms: support vector machine, BP neural network and ridge regression are used to predict the building heat load, and compared and verified with three feature engineering algorithms: Pearson coefficient method, MRMR and PCA.
[0067] CV (RMSE) (Coefficient of Variation of the Root-Mean-Square Error) is used as the prediction effect evaluation index, which is defined as shown in formula (15):
[0068]
[0069] In the formula, y i represents the actual measured value, represents the simulated value, and n represents the number of samples. MBE reflects the changing trend of the overall deviation. A positive result indicates that the result is overestimated, and a negative result indicates that the result is underestimated. CV (RMSE) is used to measure the error between the actual data and the simulated value. The smaller the CV (RMSE) value, the smaller the error, and the closer the prediction result is to the true value.
[0070] Table 1 is a comparison between the data feature engineering processing method based on MRMR-PCA proposed in the present invention and three data feature engineering processing methods, namely, the Pearson coefficient method, MRMR, and PCA. It can be seen that the method proposed in the present invention obtains the least variables and the smallest fitting error, and can provide efficient input feature variable data for the subsequent building heat load prediction model.
[0071] Table 1
[0072]
Claims
1. A data feature engineering processing method for improving the accuracy of building cooling and heating load prediction based on machine learning, characterized in that: The method comprises the following steps in order: S1. Construct a building cooling and heating load prediction data set: collect building cooling and heating load data, meteorological data, and building internal disturbance data; S2. Data preprocessing: The collected data may be abnormal or missing. The IQR method is used to identify and process data outliers, and the interpolation method is used to fill in missing values and outliers, and the characteristic variables are standardized. S3. Feature selection: Use the MRMR algorithm to calculate the maximum correlation-minimum redundancy of the variables, and select the feature subset that has a strong correlation with the target variable and a low redundancy with other feature variables; S4. Feature extraction: PCA algorithm is used to transform the feature subsets selected by MRMR algorithm into independent and orthogonal comprehensive feature sets, further reducing the mutual influence of redundant information. The data processed by MRMR-PCA is the characteristic variable finally input into the building load prediction model.
2. According to claim 1, a data feature engineering processing method for improving the accuracy of building cooling and heating load prediction based on machine learning is characterized in that: The specific data in step S1 include: the building's cooling and heating load data include the water supply flow and water supply temperature difference at the thermal inlet; the meteorological data include outdoor temperature, dew point temperature, relative humidity, wind speed, wind direction, total horizontal radiation, normal direction solar radiation, horizontal scattered radiation; the building's internal disturbance data include the number of people in the room, lighting power and equipment power. These data can be obtained by installing corresponding sensors and smart energy systems.
3. According to claim 1, a data feature engineering processing method for improving the accuracy of building cooling and heating load prediction based on machine learning is characterized in that: The IQR method described in step S2 determines whether the data is an outlier based on the quartile range of the data. The specific steps are as follows: S21. Calculate the first quartile (Q1), the second quartile (median), and the third quartile (Q3) of the data; S22. Calculate IQR = Q3 - Q1; S23. Calculate the lower limit lower = Q1-k*IQR and the upper limit upper = Q3+k*IQR, k can be 1.5 or 3, when k = 3 represents an extreme outlier, when k = 1.5 represents a moderate outlier; S24. If the data value is less than lower or greater than upper, it is considered an outlier; The interpolation method in step S2 to fill the missing values refers to using the mean of two adjacent moments to replace the missing values; S25. Standardize the variables according to formula (1): In the formula x is the original value of a feature, x * is the eigenvalue after standardization, μ is the average value of the feature in all samples, and σ is the standard deviation of the feature in all samples.
4. According to claim 1, a data feature engineering processing method for improving the accuracy of building cooling and heating load prediction based on machine learning is characterized in that: In step S3, the maximum correlation-minimum redundancy of the variables is calculated using the MRMR algorithm, and a feature subset with a strong correlation with the target variable and a low redundancy with other feature variables is selected. The specific steps are as follows: S31. In the MRMR algorithm, correlation and redundancy are measured by mutual information, that is, the degree of reduction in the uncertainty (measured by information entropy) of the output variable Y after knowing a variable X. The calculation formula for the mutual information of two continuous variables is shown in formula (2): In the formula, I(X, Y) is the mutual information of variables X and Y, P(x, y) is the joint probability density function of X and Y, and the marginal probability density distribution function of P(x) and P(y). The value range of the calculation result of mutual information is [0, 1]. The larger its value, the greater the correlation between the two variables. According to the maximum correlation criterion, I(X i , Y) in an appropriate order to search for the best feature variable related to the output variable Y, and use formula (3) to calculate its correlation size: S32. In order to eliminate the redundant information between the selected variables in the feature set and select mutually exclusive features, the minimum redundancy criterion is calculated using formula (4): S33. Integrate formula (4) and (5) by addition to find the set S with maximum correlation and minimum redundancy: Maxφ(F),φ(D,R)=DR (5) S34. Use the incremental search method to obtain the approximate optimal solution, that is, assuming that S has been obtained p-1 feature subsets, it is necessary to p-1 The pth feature is selected from the features, and feature selection is performed by maximizing φ(D, R), that is, maximizing:
5. According to claim 1, a data feature engineering processing method for improving the accuracy of building cooling and heating load prediction based on machine learning is characterized in that: In step S4, the PCA algorithm is used to transform the feature subsets screened by the MRMR algorithm into independent and orthogonal comprehensive feature sets to further reduce the mutual influence of redundant information. The specific steps are as follows: S41. Assume that p feature variables are selected after the MRMR algorithm, the feature matrix is represented by X, and the variables are standardized according to formula (7) to obtain the standardized matrix B: In the formula, S42. Solve the covariance of the standardized matrix according to formulas (8) and (9): S43. Calculate the eigenvalues of the covariance matrix R and the corresponding orthogonalized eigenvectors: According to the calculation formula of the correlation coefficient matrix R, we can get r ij =r ji According to linear algebra knowledge, the correlation coefficient matrix R is a real symmetric matrix. There must be a standard orthogonal matrix U and a diagonal matrix Λ, such that: where λ1, λ2, ..., λ p are the eigenvalues of the correlation coefficient matrix R, and their corresponding unit eigenvectors are U1, U2, ..., U p ; S44. Perform linear combination of unit eigenvectors according to formula (11) to form new eigenvariables: In the formula, F i is the i-th characteristic variable after spatial transformation, i = 1, 2, ..., p. From the knowledge of linear algebra, we know that the new characteristic variables are orthogonal and uncorrelated with each other. S45. Calculate the variance of the new feature variable according to formula (12): From step S43, we know that Var(F i )=λ i , that is, the variance of the new feature variable is the eigenvalue of the covariance matrix. For this reason, it is assumed that the order of the eigenvalues of the covariance matrix is λ1≥λ2≥λ3…≥λ p ≥0, then the variance order of the new characteristic variable composed of the eigenvectors corresponding to its eigenvalues is also Var(F1)≥Var(F2)≥Var(F3)…≥Var(F p )≥0. In this case, F1 is called the first principal component, F2 is called the second principal component, ..., F p It is called the pth principal component. S46. Calculate the principal component contribution rate and accumulation according to formulas (13) and (14) to determine how many principal components to retain: Where F i is the contribution rate of the i-th principal component, η i is the cumulative contribution rate of the first i principal components. Generally, the eigenvalues λ1, λ2, ..., λ1 whose cumulative contribution rate reaches 85% to 95% are selected. m The corresponding 1st, 2nd, ..., m (m≤p) variables are taken as principal components, which is the final data feature engineering processing result.