An air conditioner load prediction method based on eigenvalue selection

By employing an eigenvalue selection method that combines mathematical mutual information with physical redundancy, the problem of inaccurate eigenvalue selection in air conditioning load forecasting is solved, thereby improving forecast accuracy and reducing average error, and achieving more efficient air conditioning load forecasting.

CN119760385BActive Publication Date: 2026-04-17ECONOMIC TECH RES INST OF STATE GRID HENAN ELECTRIC POWER +3
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ECONOMIC TECH RES INST OF STATE GRID HENAN ELECTRIC POWER
Filing Date
2024-12-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing air conditioning load forecasting methods suffer from insufficient prediction accuracy in feature value selection, especially when dealing with dynamic features, making it difficult to accurately select the feature values ​​for model input.

Method used

An eigenvalue selection method based on the synergy of mathematical mutual information and physical redundancy is adopted. Through data preprocessing, correlation analysis, mutual information calculation and model training, the optimal eigenvalue is selected for air conditioning load prediction. This includes data correction, standardization, normalization, Person correlation coefficient evaluation, Shannon information entropy theory calculation and long short-term memory network model construction.

Benefits of technology

It improves the accuracy of the air conditioning load prediction model, increases the correlation coefficient by 5.62%, and reduces the average error by 16% to 22%, while keeping the calculation cost unchanged, and has good practical application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119760385B_ABST
    Figure CN119760385B_ABST
Patent Text Reader

Abstract

This invention provides an air conditioning load forecasting method based on eigenvalue selection, belonging to the fields of machine learning and building energy conservation technology. This eigenvalue-based air conditioning load forecasting method selects eigenvalues ​​through the coordinated integration of mathematical mutual information and physical redundancy, and evaluates the selected eigenvalues ​​using a posterior method of the prediction model, accurately determining the optimal input eigenvalues ​​for the model. This eigenvalue-based air conditioning load forecasting method can quantitatively capture the dynamic characteristics affecting real-time air conditioning load values, ensuring high prediction accuracy. It can improve the accuracy of the air conditioning load forecasting model by 5.62% in the correlation coefficient and reduce the average error by approximately 16%–22%, while essentially without increasing computational costs, demonstrating significant practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of machine learning and building energy conservation technology, and specifically relates to an air conditioning load prediction method based on feature value selection. Background Technology

[0002] Accurate prediction of building air conditioning load is crucial for optimizing the operation of air conditioning systems and improving energy management efficiency. In recent years, air conditioning load prediction based on data-driven and machine learning has become an increasingly popular research topic. Related research mainly focuses on three aspects: (1) research on the prediction model itself; (2) comparative research on optimizing the model structure and parameters using optimization algorithms; and (3) research on the impact of eigenvalue selection methods on model performance. Although many scholars have achieved considerable research results in these areas, much work still needs to be further explored and improved, especially in the research on eigenvalue selection methods to improve the accuracy of model prediction, where there is still much room for improvement.

[0003] This air conditioning load forecasting method, based on black-box modeling, extracts features from historical operating data of the actual system and learns the mapping relationship between the physical quantity to be predicted (load) and other physical quantities that affect it. This completes model training and enables the prediction of future loads. Since these physical quantities affecting the load have dynamic characteristics that change over time, only by accurately selecting the feature values ​​input to the model can these dynamic characteristics be quantitatively captured, ensuring high prediction accuracy. Summary of the Invention

[0004] This invention is made to solve the above-mentioned problems, and aims to provide an air conditioning load prediction method based on feature value selection.

[0005] This invention provides an air conditioning load prediction method based on eigenvalue selection, used to predict air conditioning load based on historical load values ​​L of the air conditioning system. T-m And the corresponding indoor and outdoor characteristic prediction of the current load value L of the air conditioning system T T represents the current time, m is the time difference, and Tm represents the historical time. Based on the characteristic of eigenvalue selection using a combination of mathematical mutual information and physical redundancy, the method includes the following steps: S10, given a fixed building envelope, select the historical load value L for a set time period of the air conditioning system. T-m The data and their corresponding indoor and outdoor characteristics were used as a sample set. After correcting and imputing outliers and missing values ​​in the sample set, the data in the sample set were standardized and normalized. S20, Person correlation coefficient was used to evaluate the historical load value L. T-m Compared with indoor and outdoor characteristics and the current load value L TThe correlation is analyzed to determine its impact, thereby eliminating data that violates physical principles and obtaining the dataset; S30, based on Shannon's information entropy theory, the historical load value L in the dataset is calculated. T-m The indoor and outdoor characteristics are respectively related to the current load value L T After sorting the mutual information values, select the historical load values ​​L corresponding to several mutual information values ​​greater than the set value. T-m S40: Using indoor and / or outdoor features as feature values, a candidate pool is formed; S50: According to a set time period, the dataset is divided into a training set and a test set, and an air conditioning load prediction model is constructed accordingly; S60: Considering mathematical mutual information, the 1 to k feature values ​​with the largest mutual information values ​​in the candidate pool are selected sequentially as input feature values ​​to train and predict the air conditioning load prediction model, thus obtaining the k corresponding to the air conditioning load prediction model with the highest prediction accuracy, which is the specific selection of feature values ​​based on mathematical mutual information; S70: Considering physical redundancy, the 1 to w feature values ​​with the largest mutual information values ​​remaining in the candidate pool and different from those selected in step S50 are selected sequentially as input feature values, thereby training and predicting the air conditioning load prediction model with the highest prediction accuracy obtained in step S50, thus obtaining the w corresponding to the air conditioning load prediction model with the highest prediction accuracy, which is the specific selection of feature values ​​based on physical redundancy; S81: Based on the optimal air conditioning load prediction model obtained by training the k+w feature values ​​finally determined in steps S50 to S60, the current load value of the air conditioning system is predicted by inputting historical load values ​​and / or indoor and outdoor features into it.

[0006] The air conditioning load forecasting method based on eigenvalue selection provided by this invention may also have the following features: In step S10, the indoor and outdoor features include indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity; the building where the air conditioning system is located is underground; and in step S20, the sign of the calculated Person correlation coefficient r indicates the historical load value L. T-m Compared with indoor and outdoor characteristics and the current load value L T They are positively or negatively correlated.

[0007] The air conditioning load forecasting method based on feature value selection provided by the present invention may also have the following feature: wherein step S10 includes the following sub-steps:

[0008] S11, Given a fixed building envelope, select the historical load value L for the set time period of the air conditioning system. T-m The data, along with their corresponding indoor and outdoor features, form the sample set; S12, using a level processing method, continuous outliers in the sample set are corrected:

[0009] |Y(d,t)-Y(d,t-1)|>α(t),

[0010] |Y(d,t)-Y(d,t+1)|>β(t),

[0011]

[0012] α(t) and β(t) represent the threshold values ​​of the deviation between the historical load value at time t and the adjacent historical load values, respectively. Y(d,t) represents the historical load value at time t on day d. Y(d,t-1) and Y(d,t+1) represent the historical load values ​​at time t-1 and time t+1 on day d, respectively.

[0013] S13, using a vertical processing method, based on the similar characteristics of historical load values ​​at the same time on similar days, outliers are replaced with data from nearby dates at the same time:

[0014] |Y(d,t)-a(t)|>r(t),

[0015]

[0016] a(t) represents the average historical load value of similar days at time t, and r(t) represents the threshold between the historical load value at time t and the average value;

[0017] S14. For imputation of a single missing value, the average interpolation method is used, which is to impute the missing value by calculating the average of the two data points before and after the missing value. For imputation of two or more consecutive missing values, the cubic spline interpolation method is used: Given that the function f(x) has n+1 distinct nodes on the interval [a,b], a=x0 <x1<…<x i <… <x n The function value at point b is y. i =f(x) i The constructor s(x) satisfies the condition in [x... i ,x i The highest power of the polynomial in [+1] is 3, and s(x) and its first and second derivatives are continuous on [a,b]. The calculation formula is: c, d, e, and f are polynomial coefficients;

[0018] S15, normalizes the sample set, scaling the overall data to the range [0,1], reducing the order-of-magnitude difference between different data types and eliminating outliers: Y Normalization Y represents the normalized data. i Y represents the i-th data point. max Y min These represent the maximum and minimum values ​​of all data, respectively.

[0019] S16 standardizes the sample set, transforming the data to a normal distribution with a mean of 0 and a standard deviation of 1, thus better mapping the relationship between the physical quantity to be predicted and the input features. Y Standardization This represents the standardized data, where μ and σ represent the mean and variance of the data, respectively.

[0020] The air conditioning load forecasting method based on feature value selection provided by this invention may also have the following feature: wherein step S20 includes the following sub-steps: S21, using the Person correlation coefficient r to evaluate the historical load value and indoor and outdoor characteristics, respectively, and the current load value L. T Correlation: N represents the number of feature samples, α j and β j Let r represent the characteristic variables between the j-th historical load value and the influencing factors, respectively. The value range of r is [-1, 1]. r in the range of [-1, 0] indicates that the two variables are linearly negatively correlated, and r in the range of [0, 1] indicates that the two variables are linearly positively correlated. |r| tending to 1 indicates that the correlation between the two variables is relatively high, and |r| tending to 0 indicates that the correlation is relatively low. S22, calculate the correlation coefficient r between several historical load values ​​and / or indoor and outdoor characteristics at past historical moments and the current load value, and remove the historical load values ​​and / or indoor and outdoor characteristics in the sample set corresponding to the correlation coefficient r of the phenomena that violate physical principles.

[0021] The air conditioning load prediction method based on feature value selection provided by this invention may also have the following feature: wherein step S30 includes the following sub-steps: S31, based on Shannon's information entropy theory, using the mutual information calculation formula to calculate the mutual information values ​​of historical load values ​​and indoor and outdoor features in the dataset with respect to the current load value: G represents the number of information events, u and v represent the u-th and v-th information events respectively, P(X u ,Y v ) represents the Xth u ,Y v The joint probability of the occurrence of _x_ information events, P(X_1) u ) and P(Y v ) respectively represent the Xth u and Y v The probability of an information event occurring is used to quantitatively describe the correlation between two event sets using mutual information values. The larger the mutual information value, the greater the degree of interdependence between the two variables and the stronger the correlation. S32: The mutual information values ​​calculated in step S31 are merged and arranged in descending order. The historical load values ​​and / or indoor and outdoor characteristics corresponding to all mutual information values ​​not less than the set value are selected as feature values ​​and arranged in descending order to form a candidate pool.

[0022] The air conditioning load prediction method based on feature value selection provided by the present invention may also have the following feature: in step S32, the set value is 0.185.

[0023] The air conditioning load prediction method based on feature value selection provided by the present invention may also have the following feature: in step S40, the air conditioning load prediction model is constructed by a long short-term memory network or a BP neural network.

[0024] The air conditioning load forecasting method based on eigenvalue selection provided by this invention may also have the following feature: In steps S50 to S60, the forecasting accuracy is calculated as follows: correlation coefficient Mean Absolute Error Root mean square error Mean absolute percentage error For predicted values The covariance with the actual value γ, And Var[γ] are the predicted values. The variance of the actual value γ, where n is the prediction duration. and γ z ξ represents the predicted value and the actual value at time z, respectively. The smaller the value of ξ, the lower the prediction accuracy. The larger the values ​​of MAE, RMSE and MAPE, the lower the prediction accuracy.

[0025] The air conditioning load prediction method based on feature value selection provided by the present invention may also have the following features: In steps S50 to S60, when adding similar feature values ​​as input feature values ​​for training and prediction of the air conditioning load prediction model, and the prediction accuracy decreases, the addition is stopped; otherwise, the addition continues until the prediction accuracy begins to decrease. When adding similar feature values ​​as input feature values ​​for training and prediction of the air conditioning load prediction model until the prediction accuracy begins to decrease, different types of feature values ​​are added as input feature values ​​for training and prediction of the air conditioning load prediction model. If the prediction accuracy improves, the addition continues; otherwise, the addition stops.

[0026] The role and effect of invention

[0027] According to the air conditioning load forecasting method based on eigenvalue selection of the present invention, the eigenvalue selection method based on the synergy of mathematical mutual information and physical redundancy accurately selects the eigenvalues ​​of the model input.

[0028] Therefore, the air conditioning load prediction method based on eigenvalue selection of the present invention can quantitatively capture the dynamic characteristics that affect the real-time load value of air conditioning, ensuring that the model has high prediction accuracy. It can improve the accuracy of the air conditioning load prediction model by 5.62% in the correlation coefficient and reduce the average error by about 16% to 22%, while basically not increasing the calculation cost, and has good practical application value. Attached Figure Description

[0029] Figure 1 This is a flowchart of an air conditioning load prediction method based on feature value selection according to an embodiment of the present invention;

[0030] Figure 2 This refers to the hourly historical air conditioning load values ​​in the sample set after data preprocessing in step S10 of the embodiment of the present invention.

[0031] Figure 3 The historical load value L is calculated in step S22 of the embodiment of the present invention. T-m Indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity over the past 24 hours (m = 0–24) compared to the current load value L. T The curve showing the change in the Person correlation coefficient r between them;

[0032] Figure 4 This is the candidate pool of feature values ​​obtained by merging mutual information values ​​and arranging them in descending order in the embodiments of the present invention;

[0033] Figure 5 This is a training set sample illustration of a feature value in step S50 of an embodiment of the present invention;

[0034] Figure 6 This is a diagram illustrating the training set samples of the two feature values ​​in step S50 of an embodiment of the present invention;

[0035] Figure 7 This illustrates the impact of the number of feature values ​​on prediction accuracy in step S50 of an embodiment of the present invention.

[0036] Figure 8 These are the air conditioning load prediction results under six scenarios in step S50 of an embodiment of the present invention;

[0037] Figure 9 This is a box plot of the relative errors of the prediction results under six scenarios in step S50 of an embodiment of the present invention.

[0038] Figure 10 This is a comparison of the accuracy and performance evaluation index of the air conditioning load prediction model when the feature value type is historical load value and outdoor air temperature in step S60 of an embodiment of the present invention.

[0039] Figure 11 This is a comparison of the accuracy and performance evaluation index of the air conditioning load prediction model when the feature value types are historical load value, outdoor air temperature and indoor air temperature in step S60 of the embodiment of the present invention.

[0040] Figure 12 This is a comparison of the accuracy and performance evaluation indicators of the air conditioning load prediction model when the feature values ​​in step S60 of the embodiment of the present invention are historical load values, outdoor air temperature, indoor air temperature and outdoor air relative humidity. Detailed Implementation

[0041] To make the technical means, creative features, objectives and effects of the present invention easy to understand, the following embodiments, in conjunction with the accompanying drawings, specifically illustrate an air conditioning load prediction method and its application based on feature value selection.

[0042] <Example>

[0043] Figure 1 This is a flowchart of an air conditioning load prediction method based on feature value selection according to an embodiment of the present invention.

[0044] like Figure 1 As shown, this embodiment provides an air conditioning load forecasting method based on eigenvalue selection. This eigenvalue selection method, based on the synergy of mathematical mutual information and physical redundancy, is used to predict the air conditioning load based on the historical load value L of the air conditioning system. T-m And the corresponding indoor and outdoor characteristic prediction of the current load value L of the air conditioning system T T represents the current time, m is the time difference, and Tm represents the historical time. The process includes the following steps S10 to S60:

[0045] S10, given a defined enclosure structure, obtain a sample set and preprocess the data therein, including the following sub-steps S11 to S16:

[0046] S11. The actual operating data of an underground building air conditioning system in a hot summer and warm winter region from 0:00 on May 1, 2020 to 23:00 on August 31, 2020 was selected as the sample set. The air conditioning system has a total of 2 chiller units.

[0047] The sample set contains the historical load values ​​L of two chiller units for a set time period of the air conditioning system. T-m Data on its corresponding indoor and outdoor characteristics.

[0048] The factors influencing building air conditioning load include two categories: building envelope and indoor / outdoor characteristics. Once the building envelope is determined, its dynamic impact on air conditioning load is relatively small; while indoor / outdoor characteristics, including indoor air temperature and humidity, indoor occupants, equipment and lighting, outdoor air temperature and humidity, outdoor wind speed, and solar radiation, often have a significant impact on the dynamic changes of air conditioning load.

[0049] In this embodiment, the building is located underground, therefore solar radiation and outdoor wind speed parameters have a relatively small impact on the building's air conditioning load. Therefore, the indoor and outdoor characteristics selected in this embodiment are: indoor air temperature at four sampling points within the underground building, indoor relative humidity at four sampling points within the underground building, outdoor air temperature at the entrance of the underground building, and outdoor relative humidity at the entrance of the underground building. In the above sample set, the data sampling interval is 2 minutes, totaling 88,560 sets of data.

[0050] When an air conditioning system is running, under natural conditions, due to factors such as sensor accuracy, systematic errors in measuring instruments and data collected by personnel, and sudden instability in operating conditions, outliers inevitably appear in the sample set. If these outliers are directly used to train a prediction model, it will inevitably affect the model's prediction accuracy and lead to inaccurate results. Therefore, the data in the sample set needs to be preprocessed through the following steps S12 to S14.

[0051] S12, using a leveling method, corrects consecutive outliers in the sample set:

[0052] |Y(d,t)-Y(d,t-1)|>α(t), |Y(d,t)-Y(d,t+1)|>β(t),

[0053] Where α(t) and β(t) represent the threshold values ​​of the deviation between the historical load value at time t and the adjacent historical load values, respectively; Y(d,t) represents the historical load value at time t on day d; and Y(d,t-1) and Y(d,t+1) represent the historical load values ​​at time t-1 and time t+1 on day d, respectively.

[0054] S13, using a vertical processing method, based on the similar characteristics of historical load values ​​at the same time on similar days, outliers are replaced with data from nearby dates at the same time: |Y(d,t)-a(t)|>r(t),

[0055]

[0056] Where a(t) represents the average historical load value of similar days at time t, and r(t) represents the threshold between the historical load value at time t and the average value.

[0057] S14, imputing missing values ​​in the sample set:

[0058] (1) For filling in a single missing value, the average interpolation method is used to fill in the missing value by calculating the average of the two data before and after the missing value.

[0059] (2) For imputation of two or more consecutive missing values, cubic spline interpolation is used:

[0060] Given that the function f(x) has n+1 distinct nodes on the interval [a,b], and a = x0 <x1<…<x i <… <x n The function value at point b is y. i =f(x) i The constructor s(x) satisfies the condition in [x... i ,x i The highest power of the polynomial in [+1] is 3, and s(x) and its first and second derivatives are continuous on [a,b]. The calculation formula is: Where c, d, e, and f are polynomial coefficients.

[0061] S15, normalizes the sample set, scaling the overall data to the range [0,1], reducing the order-of-magnitude difference between different data types and eliminating outliers:

[0062] Among them, Y Normalization Y represents the normalized data. i Y represents the i-th data point. max Y min These represent the maximum and minimum values ​​of all data, respectively.

[0063] S16 standardizes the sample set, transforming the data to a normal distribution with a mean of 0 and a standard deviation of 1, thus better mapping the relationship between the physical quantity to be predicted and the input features.

[0064] Among them, Y Standardization This represents the standardized data, where μ and σ represent the mean and variance of the data, respectively.

[0065] Figure 2 It refers to the hourly historical air conditioning load values ​​in the sample set after data preprocessing in step S10 of the embodiment of the present invention.

[0066] like Figure 2 As shown, after the relevant data processing in steps S11 to S16, the data in the sample set has continuity and can be used for black-box prediction of air conditioning load.

[0067] S20, Remove data from the sample set that violate physical principles to obtain the dataset, including the following sub-steps S21 to S22:

[0068] S21, Person correlation coefficient r is used to evaluate historical load values ​​and indoor / outdoor characteristics, and current load value L, respectively. T Correlation:

[0069] Where N represents the number of feature samples, α j and β j Let r represent the characteristic variables between the j-th historical load value and the influencing factors, respectively. The value range of r is [-1, 1]. r in the range of [-1, 0] indicates that the two variables are linearly negatively correlated, r in the range of [0, 1] indicates that the two variables are linearly positively correlated, |r| approaching 1 indicates that the correlation between the two variables is high, and |r| approaching 0 indicates that the correlation is low.

[0070] S22, due to thermal inertia and time delay, the historical load value L T-m and indoor and outdoor characteristics (indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity) at different historical times for the current air conditioning load (current load value L). T The degree of influence varies, and SPSS Statistics 21 software was used to calculate five characteristic physical quantities (historical load value L) respectively. T-m Indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity over the past 24 hours (m = 0–24) compared with the current load value L. T The Person correlation coefficient r between them is as follows: Figure 3 As shown. Among them, Figure 3 The historical load value L is calculated in step S22 of the embodiment of the present invention. T-m Indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity over the past 24 hours (m = 0–24) compared to the current load value L. T The Person correlation coefficient r curve between them.

[0071] Based on the physical relationship between indoor air conditioning load and indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity, the correlation rule determination condition is as follows: During the summer air conditioning season, the indoor air temperature, outdoor air temperature, and current load value L... T There is a positive correlation between indoor relative humidity, outdoor relative humidity and current load value L. T They are negatively correlated, and r is a negative value.

[0072] The outdoor air temperature (r) value is negative from T-9 to T-17, the outdoor relative humidity (r) value is positive from T-4 to T-20, and the indoor relative humidity (r) value is positive from T to T-24. All these violate the rule-based judgment conditions. Therefore, these characteristic physical quantities in the original dataset for these time periods are not relevant to the current air conditioning load (current load value L). T The original data exhibited phenomena that violated physical principles, therefore these original data were discarded.

[0073] The specific data removed are: outdoor air temperature from historical time T-9 to T-17, outdoor relative humidity from historical time T-4 to T-20, and indoor relative humidity from historical time T to T-24.

[0074] After removing the above data from the sample set, the dataset is obtained.

[0075] S30, Select the dataset that can be used to determine the current load value L T Historical load value L that has a significant impact T-m And / or indoor and outdoor features are used as feature values ​​to form a candidate pool, including the following sub-steps S31 to S32:

[0076] S31, based on Shannon's information entropy theory, uses the mutual information calculation formula to calculate the mutual information values ​​of historical load values ​​and indoor / outdoor features in the dataset with respect to the current load value:

[0077]

[0078] Where G represents the number of information events, u and v represent the u-th and v-th information events respectively, and P(X u ,Y v ) represents the Xth u ,Y v The joint probability of the occurrence of _x_ information events, P(X_1) u ) and P(Y v ) respectively represent the Xth u and Y v The probability of an event occurring is used to quantitatively describe the correlation between two sets of events using mutual information values. The larger the mutual information value, the greater the interdependence between the two variables and the stronger the correlation.

[0079] Quantitative calculation of four characteristic physical quantities (historical load values ​​L of m = 1 to 24) over the past 24 hours T-m The indoor air temperature from time T to T-24, the outdoor air temperature from time T to T-8 and from time T-18 to T-24, and the outdoor relative humidity from time T to T-3 and from time T-21 to T-24 are compared with the current load value L. T The mutual information value.

[0080] Using the Python programming language, the continuous function in the entropy_estimators self-editing library was called to calculate the mutual information value. The results of the mutual information value are shown in Table 1 below.

[0081] Table 1 (mutual information values ​​of historical load values, indoor air temperature, outdoor air temperature, and outdoor relative humidity over the past 24 hours with respect to the current load value)

[0082]

[0083] S32. Since the mutual information value reflects the degree of interdependence between two variables, the calculation conditions of all mutual information values ​​in Table 1 in step S31 are the same, and they are comparable and consistent with each other. Therefore, all mutual information values ​​in Table 1 reflect the correlation between historical load values ​​and indoor and outdoor characteristics at different times in the past 24 hours and the current load value. The larger the mutual information value, the greater its influence on the current load value, so it is used as the input feature value of the air conditioning load prediction model.

[0084] Figure 4 This is the candidate pool of feature values ​​obtained by merging mutual information values ​​and arranging them in descending order in the embodiments of the present invention.

[0085] like Figure 4 As shown, all mutual information values ​​in Table 1 of step S31 are merged and sorted in descending order. All historical load values ​​with mutual information values ​​greater than or equal to 0.185 and indoor and outdoor characteristics (indoor air temperature, outdoor air temperature and outdoor air relative humidity) are taken as feature values, thus forming a candidate pool of input feature values ​​for the building's air conditioning load prediction model.

[0086] like Figure 4 As shown, the horizontal axis represents all feature values ​​from left to right, and their influence on air conditioning load decreases from left to right. Therefore, all feature values ​​in the candidate pool have reason to be selected as input feature values ​​for the air conditioning load prediction model. However, if... Figure 4 Using all feature values ​​to train the model not only greatly increases the cost of parameter acquisition and the computational cost of the prediction model in the actual system, but also, due to the overfitting characteristics of machine learning models, may lead to low prediction accuracy in the end.

[0087] Therefore, selecting the optimal eigenvalues ​​is crucial for improving the accuracy of air conditioning load prediction models. This embodiment employs the following steps S40–S60 to select the optimal eigenvalues.

[0088] S40, using hours as the time granularity, divide the original dataset obtained in step S20 into a training set and a test set, using the data from 0:00 on May 1st to 23:00 on August 29th as the training set; and the data from 0:00 on August 30th to 23:00 on August 31st as the test set.

[0089] An air conditioning load prediction model (LSTM model) was constructed using a Long Short-Term Memory (LSTM) network. The relevant parameter settings for the model, obtained from multiple simulation experiments, are shown in Table 2 below. The LSTM model algorithm was written using MATLAB R2019a programming language, calling the Deep Learning library in Simulink.

[0090] Table 2 (Parameter Settings for Prediction Model)

[0091]

[0092] The model prediction evaluation indicators used in this embodiment for the air conditioning load prediction model are: correlation coefficient ξ, mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE), and the calculation formulas are as follows:

[0093] Correlation coefficient

[0094] Mean Absolute Error

[0095] Root mean square error

[0096] Mean absolute percentage error

[0097] in, For predicted values The covariance with the actual value γ, And Var[γ] are the predicted values. The variance of the actual value γ, where n is the prediction duration. and γ z ξ represents the predicted value and the actual value at time z, respectively. The smaller the value of ξ, the lower the prediction accuracy. The larger the values ​​of MAE, RMSE and MAPE, the lower the prediction accuracy.

[0098] S50, based on mathematical mutual information, feature values ​​are selected. The 1 to k feature values ​​with the largest mutual information values ​​in the candidate pool are selected sequentially as input feature values ​​for training and prediction of the air conditioning load prediction model. The specific selection of feature values ​​based on mathematical mutual information is obtained by obtaining the k corresponding to the air conditioning load prediction model with the highest prediction accuracy. The specific method of this step is as follows:

[0099] The air conditioning load prediction model constructed in step S40 is used to... Figure 4 Simulation experiments were conducted on various feature value selection methods for the top 1, top 2, top 3, top 4, top 5, and top 6 mutual information values ​​in the candidate pool. A total of 6 experimental scenarios were set up, as shown in Table 3 below.

[0100] Table 3 (6 experimental scenario settings)

[0101]

[0102] Taking scenario 1 as an example, the first LSTM network model is constructed, and the relevant parameter settings are shown in Table 2 in step S40. First, Figure 3 The leftmost quantity in the candidate pool, i.e., the historical air conditioning load L at time T-1. T-1 As input feature values ​​(referred to as "a feature value"), a dataset is constructed to train and predict the model.

[0103] Figure 5 This is a training set sample illustration of a feature value in step S50 of an embodiment of the present invention.

[0104] like Figure 5 As shown, the training set data is input into the model with a matrix dimension of 2904×1 for training. After the iterative calculation converges, the test set data is tested with a matrix dimension of 48×1, and the corresponding accuracy index is calculated.

[0105] Similarly, in scenario 2, a second LSTM network model is constructed, with the input feature value being... Figure 3 The leftmost two values ​​in the candidate pool.

[0106] Figure 6 This is a diagram illustrating the training set samples of two feature values ​​in step S50 of an embodiment of the present invention.

[0107] like Figure 6 As shown, the method for scenario 2 is similar, but it requires handling the missing values ​​of the data matrices in the training and test sets separately. Figure 6 (The historical moments in the yellow section) will be supplemented.

[0108] Based on the selected feature values ​​in Table 3, an experiment was conducted using an LSTM model to predict the air conditioning load of the underground building. The selected features were... Figure 4 The first 6 characteristic physical quantities from left to right (L) T-1 L T-2 L T-3 L T-4 L T-5 L T-6 ).

[0109] Figure 7This illustrates the impact of the number of feature values ​​on prediction accuracy in step S50 of an embodiment of the present invention. Figure 8 These are the air conditioning load prediction results under six scenarios in step S50 of an embodiment of the present invention; Figure 9 This is a box plot of the relative errors of the prediction results under six scenarios in step S50 of an embodiment of the present invention.

[0110] like Figures 7-9 As shown, before the number of eigenvalues ​​is increased to four, the prediction evaluation indicators of each error decrease continuously and the prediction accuracy gradually improves; however, after the number of eigenvalues ​​exceeds four, if the number of eigenvalues ​​is further increased, the evaluation indicators of each error will increase with the increase of the eigenvalues, and the prediction accuracy will begin to decrease.

[0111] This is because too few feature values ​​(k) will prevent the model from fully learning the input features, leading to underfitting during training and thus lower prediction accuracy. Conversely, too many feature values ​​will increase the maximum and minimum redundancy of parameters, creating competition between multiple sets of input parameters of the same category. Historical loads with low mutual information values, after being trained on the model, will affect historical loads with high mutual information values, causing a decrease in the prediction accuracy of the model that originally had quantitative feature values ​​as input.

[0112] Taking the LSTM model as an example, when there are too many input feature values, the model becomes more complex, the training time increases, and the cost of data collection increases. The evaluation index and computation time τ of the load prediction model in six scenarios are shown in Table 4 below.

[0113] Table 4 (Evaluation metrics and calculation time τ for load forecasting models under 6 scenarios)

[0114]

[0115] like Figure 7 , Figure 8 , Figure 9 As shown in Table 4, for the air conditioning load prediction model, the number of eigenvalues ​​selected based on mutual information ranking is not necessarily better the more there are; rather, there exists an optimal number of eigenvalues, k. Exceeding this number may negatively impact the model's predictive performance. This embodiment, through simulation experiments, found that the optimal number of eigenvalues ​​is four (i.e., k = 4, and the top four eigenvalues ​​in terms of mutual information are L...). T-1 L T-2 L T-3 L T-4 ), and the fourth-ranked eigenvalue L T-4 The mutual information value between the model and the predicted load is 0.363. Furthermore, choosing the optimal number of eigenvalues ​​k=4 did not increase the model's prediction time or computational cost.

[0116] from Figure 4 Looking at the mutual information values ​​of the candidate pool, if we only consider the mathematical mutual information values, the ones at the top are all historical load values ​​(black bars).

[0117] As shown in step S50, selecting the first four feature values ​​already achieves the best prediction accuracy. If the number is increased to the first five (when k=5), the prediction accuracy decreases. Furthermore, for those features... Figure 3 Indoor air temperature, outdoor air temperature, or indoor relative humidity (red, yellow, or green bars) ranked lower in the model are generally considered unsuitable as input features due to their smaller mutual information values.

[0118] However, the current load value of air conditioning in actual underground buildings is related to the load of people, indoor lighting and equipment heating, as well as the heat storage of the underground soil layer. Therefore, the current load value of air conditioning is also affected by indoor and outdoor air parameters. Thus, from the perspective of load generation mechanism, indoor and outdoor temperature and humidity also affect the building air conditioning load.

[0119] Therefore, in Figure 4 In the candidate pool, it is also necessary to select other categories of physical quantities (indoor air temperature, outdoor air temperature, or outdoor relative humidity) that may affect the current load value of the air conditioner and are not historical load values ​​as feature values, while taking into account their mutual information numerical ranking. Therefore, the following step S60 is performed:

[0120] S60, under the physical redundancy coordination, the remaining 1 to w feature values ​​with the largest mutual information values ​​in the candidate pool, which are different from those selected in step S50, are selected sequentially as input feature values. This allows the air conditioning load prediction model with the highest prediction accuracy obtained in step S50 to be trained and predicted again, resulting in w corresponding to the air conditioning load prediction model with the highest prediction accuracy (better than step S50). This yields the specific selection of feature values ​​based on physical redundancy coordination. The specific method for this step is as follows:

[0121] As shown in Table 5, based on the four optimal feature values ​​selected in step S50, the outdoor air temperature C, which has the highest mutual information value among the other categories of feature physical quantities (outdoor air temperature) of non-historical load values, is ranked first. T,o (o represents outdoors, T represents the current time, and C represents the temperature) as the fifth characteristic value ( Figure 4 The first red column appears in the image; the selected feature values ​​are still evaluated using the posterior method of the LSTM model until the number of feature values ​​for the optimal outdoor air temperature is obtained.

[0122] Continue to add other categories of characteristic physical quantities (indoor air temperature) as characteristic values. Figure 4The first yellow bar to appear in the table); and so on, after selecting the optimal characteristic value (indoor air temperature), add outdoor relative humidity as a characteristic value. Figure 4 The number of eigenvalues ​​with the largest mutual information value that should be selected in this step (the first green bar that appears in the model) until the prediction accuracy of the air conditioning load prediction model reaches its maximum can be obtained.

[0123] Table 5 (Types and Number of Different Category Feature Values ​​Added in Step S60)

[0124]

[0125] In Table 5 above, L represents the air conditioning load, T represents the current time, Tm represents the historical time (m = 1 to 24), C represents the temperature, o represents the outdoor temperature, q represents the indoor temperature, and RH represents the relative humidity of the air.

[0126] In this step, we still use the posterior method of the LSTM model to conduct simulation experiments on the various scenarios in Table 5.

[0127] Figure 10 This is a comparison of the accuracy and performance evaluation index of the air conditioning load prediction model when the feature value type is historical load value and outdoor air temperature in step S60 of an embodiment of the present invention. Figure 11 This is a comparison of the accuracy and performance evaluation index of the air conditioning load prediction model when the feature value types are historical load value, outdoor air temperature and indoor air temperature in step S60 of the embodiment of the present invention. Figure 12 This is a comparison of the accuracy and performance evaluation indicators of the air conditioning load prediction model when the feature values ​​in step S60 of the embodiment of the present invention are historical load values, outdoor air temperature, indoor air temperature and outdoor air relative humidity.

[0128] like Figures 10-12 The accuracy comparison and performance evaluation indicators of air conditioning load prediction models under different eigenvalue selections are shown below:

[0129] (1) When the feature values ​​are selected as the top 5 historical load values ​​sorted by mutual information, the model prediction accuracy begins to decrease. However, if physical redundancy is considered while taking into account the mathematical mutual information sorting, C, which is only ranked 13th in the total mutual information sorting but has the first ranking among feature physical quantities such as outdoor temperature, which affects the current load value, can be selected. T,o Choosing C as the fifth feature value reduces various error indices in the model prediction and further improves the model accuracy; however, if C, which is ranked second among physical quantities such as outdoor temperature, is also selected... T-1,o Choosing it as the sixth feature value reduces the model's prediction accuracy.

[0130] (2) Next, C, which ranks first in other categories of characteristic physical quantities such as indoor temperature, although it is only ranked 16th in the total mutual information value ranking, is ranked first. T-18,q As the sixth feature value, it can also further improve the final prediction accuracy of the model. However, adding such physical quantities would not be conducive to the final prediction accuracy of the model.

[0131] (3) However, if we rank RH, the first among the fourth category of characteristic physical quantities, outdoor relative humidity, T-24,o As the seventh feature value, it does not lead to a further improvement in model prediction accuracy. This is because although RH T-24,o It belongs to a different category of physical quantity than the previous ones, but it is ranked too low in the total mutual information value and has too little correlation with the predicted load.

[0132] Therefore, while considering mathematical quantification indicators, selecting parameters with relatively low overall ranking but high mutual information values ​​among other physical quantities as feature values ​​helps improve the model's final prediction accuracy. This can be explained by the physical redundancy of the model's input feature values. Choosing outdoor temperatures, which have lower mutual information values ​​but are different from historical load physical quantities, as feature values ​​avoids maximum and minimum redundancy between similar physical quantities. Furthermore, the model can extract features from other physical quantities to more fully explore the mapping relationship between historical data and predicted quantities, thus helping to improve the model's prediction accuracy. However, when a second outdoor temperature is added, the model's prediction accuracy shows a downward trend, indicating that adding too many outdoor temperatures of the same type increases the competition among them, making it difficult for the model to effectively learn relevant features, leading to a decrease in prediction accuracy.

[0133] However, while considering physical redundancy, it is also necessary to take into account the overall ranking of characteristic physical quantities in the correlation quantification index. In this embodiment, the outdoor air relative humidity (RH) is... T-24,o A lower overall ranking indicates a lower correlation between the data and the current load value to be predicted. Therefore, it cannot provide more valuable data features for the prediction model. If selected as a feature value, it will actually reduce the prediction accuracy.

[0134] like Figures 10-12 The comparison of accuracy and performance evaluation indicators of air conditioning load prediction models under different eigenvalue selections is shown. When selecting... Figure 4When the mutual information ranking of the candidate pool consists of six feature values—historical load at times T-1, T-2, T-3, and T-4, outdoor air temperature at time T, and indoor air temperature at time T-18—ranked as 1st to 4th, 13th, and 16th respectively (i.e., k=4 in step S50 and w=2 in step S60), the accuracy evaluation indicators of the LSTM air conditioning load prediction model are as follows: correlation coefficient 0.94, mean absolute error 15.75kW, root mean square error 22.03kW, and mean absolute percentage error 3.05%. Compared with the model accuracy evaluation indicators when only the historical load at times T-1, T-2, T-3, and T-4, which rank in the top four overall, are selected as feature values, the correlation coefficient is improved by 5.62%, the mean absolute error is reduced by 16.13%, the root mean square error is reduced by 22.76%, and the mean absolute percentage error is reduced by 16.44%. The slightly extended computation time hardly increases the computational cost and is negligible compared to the improvement in model prediction accuracy.

[0135] S70, based on the k+w feature values ​​finally determined in steps S50 to S60 (specifically, in this embodiment...) Figure 4 The optimal air conditioning load prediction model is trained by using the six feature values ​​(historical load at times T-1, T-2, T-3, and T-4, outdoor air temperature at time T, and indoor air temperature at time T-18) in the candidate pool with mutual information sorted as 1-4, 13, and 16 respectively. By inputting historical load values ​​and / or indoor and outdoor features into it, the current load value of the air conditioning system can be predicted.

[0136] The role and effect of the embodiments

[0137] According to the air conditioning load prediction method based on feature value selection involved in this embodiment, the optimal input feature value of the model is accurately determined because the feature value selection method based on the coordination of mathematical mutual information and physical redundancy is used, and the selected feature value is evaluated by the posterior method of LSTM model.

[0138] Therefore, the air conditioning load prediction method based on eigenvalue selection of the present invention can quantitatively capture the dynamic characteristics that affect the real-time load value of air conditioning, ensuring that the model has high prediction accuracy. It can improve the accuracy of the air conditioning load prediction model by 5.62% in the correlation coefficient and reduce the average error by about 16% to 22%, while basically not increasing the calculation cost, and has good practical application value.

[0139] Those skilled in the art should understand that this invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to this invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. An air conditioning load prediction method based on eigenvalue selection, for predicting a current load value L T-m of an air conditioning system according to historical load values L T of the air conditioning system and corresponding indoor and outdoor characteristics, wherein T represents a current time, m represents a time difference, and T-m represents a historical time. Its features are, The eigenvalue selection method based on the synergistic optimization of mathematical mutual information and physical redundancy includes the following steps: S10, given a fixed building envelope, select the historical load value L for the set time period of the air conditioning system. T-m The data of the indoor and outdoor features and their corresponding data are used as a sample set. After correcting and supplementing the outliers and missing values ​​in the sample set, the data in the sample set are standardized and normalized. S20, The Person correlation coefficient is used to evaluate the historical load value L. T-m The indoor and outdoor characteristics and the current load value L are respectively T The correlation is analyzed to determine its impact, thereby eliminating data that violates physical principles and obtaining the dataset; S30, Based on Shannon's information entropy theory, calculate the historical load value L in the dataset. T-m The indoor and outdoor characteristics are respectively related to the current load value L. T After sorting the mutual information values, the historical load values ​​L corresponding to several mutual information values ​​greater than a set value of 0.185 are selected. T-m And / or the indoor and outdoor features are used as feature values ​​to form a candidate pool; S40, according to the set time period, the dataset is divided into a training set and a test set, and an air conditioning load prediction model is constructed based on this. The air conditioning load prediction model is constructed through a long short-term memory network or a BP neural network. S50, considering mathematical mutual information, sequentially select the 1 to k feature values ​​with the largest mutual information values ​​in the candidate pool as input feature values ​​to train and predict the air conditioning load prediction model, and obtain the k corresponding to the air conditioning load prediction model with the highest prediction accuracy, that is, obtain the specific selection of the feature value based on mathematical mutual information. S60, considering physical redundancy, sequentially select the 1 to w feature values ​​with the largest mutual information values ​​remaining in the candidate pool that are different from the types selected in step S50 as input feature values, thereby training and predicting the air conditioning load prediction model with the highest prediction accuracy obtained in step S50, and obtaining w corresponding to the air conditioning load prediction model with the highest prediction accuracy, that is, obtaining the specific selection of the feature values ​​based on physical redundancy. S70: Based on the optimal air conditioning load prediction model trained using the k+w feature values ​​finally determined in steps S50-S60, the current load value of the air conditioning system is predicted by inputting historical load values ​​and / or indoor and outdoor features into the model. In steps S50-S60, the prediction accuracy is calculated as follows: correlation coefficient Mean Absolute Error Root mean square error Mean absolute percentage error COV(ŷ,y) is the predicted value. Compared with actual value The covariance, Var[ ] and Var[ [These are the predicted values] Compared with actual value The variance, where n is the prediction duration. and Let Z be the predicted value and the actual value at time z, respectively. The smaller the value, the lower the prediction accuracy; the larger the values ​​of MAE, RMSE, and MAPE, the lower the prediction accuracy. When adding similar feature values ​​as input feature values ​​for training and prediction of the air conditioning load prediction model, if the prediction accuracy decreases, then the addition should be stopped; otherwise, the addition should continue until the prediction accuracy begins to decrease. When the air conditioning load prediction model is trained and predicted by adding similar feature values ​​as input feature values ​​until the prediction accuracy begins to decrease, the air conditioning load prediction model is trained and predicted by adding different feature values ​​as input feature values. If the prediction accuracy improves, the addition continues; otherwise, the addition stops.

2. The air conditioning load forecasting method based on eigenvalue selection according to claim 1, characterized in that: in, In step S10, the indoor and outdoor characteristics include indoor air temperature, indoor relative humidity, outdoor air temperature, and outdoor relative humidity. The air conditioning system is located in an underground building. In step S20, the sign of the calculated Person correlation coefficient r indicates the historical load value L. T-m The indoor and outdoor characteristics and the current load value L are respectively T They are positively or negatively correlated.

3. The air conditioning load forecasting method based on eigenvalue selection according to claim 1 or 2, characterized in that: in, Step S10 includes the following sub-steps: S11, Given a fixed building envelope, select the historical load value L for the set time period of the air conditioning system. T-m The data of the indoor and outdoor features and their corresponding characteristics are used as the sample set; S12, using a leveling method, correct the continuous outliers in the sample set: , , , α(t) and β(t) represent the threshold values ​​of the deviation between the historical load value at time t and the adjacent historical load values, respectively. Y(d,t) represents the historical load value at time t on day d. Y(d,t-1) and Y(d,t+1) represent the historical load values ​​at time t-1 and time t+1 on day d, respectively. S13, using a vertical processing method, based on the similar characteristics of historical load values ​​at the same time on similar days, outliers are replaced with data from nearby dates at the same time: , , a(t) represents the average historical load value of similar days at time t, and r(t) represents the threshold between the historical load value at time t and the average value; S14. For imputation of a single missing value, the average interpolation method is used, which is to impute the missing value by calculating the average of the two data points before and after the missing value. For imputation of two or more consecutive missing values, cubic spline interpolation is used: Given that the function f(x) has n+1 distinct nodes on the interval [a,b], and a=x0 <x1<…<x i <… <x n The function value at point b is y. i =f(x i The constructor s(x) satisfies the condition in [x... i ,x i The highest power of the polynomial in [+1] is 3, and s(x) and its first and second derivatives are continuous on [a,b]. The calculation formula is: , c, d, e, and f are polynomial coefficients; S15, normalizes the sample set, scaling the overall data to the range [0,1], reducing the order-of-magnitude difference between different data types and eliminating outliers: , Y Normalization Y represents the normalized data. i Y represents the i-th data point. max Y min These represent the maximum and minimum values ​​of all data, respectively. S16 standardizes the sample set, transforming the data to a normal distribution with a mean of 0 and a standard deviation of 1, thus better mapping the relationship between the physical quantity to be predicted and the input features. , Y Standardization This represents the standardized data, where μ and σ represent the mean and variance of the data, respectively.

4. The air conditioning load forecasting method based on eigenvalue selection according to claim 1 or 2, characterized in that: in, Step S20 includes the following sub-steps: S21, Person correlation coefficient r is used to evaluate the historical load value and the indoor and outdoor characteristics, respectively, and the current load value L. T Correlation: , N represents the number of feature samples, α j and β j Let r represent the characteristic variables between the j-th historical load value and the influencing factors, respectively. The value range of r is [-1, 1]. r in the range of [-1, 0] indicates that the two variables are linearly negatively correlated, r in the range of [0, 1] indicates that the two variables are linearly positively correlated, |r| approaching 1 indicates that the correlation between the two variables is high, and |r| approaching 0 indicates that the correlation is low. S22, calculate the correlation coefficient r between the historical load values ​​and / or the indoor and outdoor characteristics at past historical moments and the current load values ​​respectively, and remove the historical load values ​​and / or the indoor and outdoor characteristics in the sample set corresponding to the calculated correlation coefficient r of phenomena that violate physical principles.

5. The air conditioning load forecasting method based on eigenvalue selection according to claim 1 or 2, characterized in that: in, Step S30 includes the following sub-steps: S31, Based on Shannon's information entropy theory, the mutual information values ​​of the historical load value and the indoor / outdoor features in the dataset with respect to the current load value are calculated using the mutual information calculation formula: , G represents the number of information events, u and v represent the u-th and v-th information events respectively, and P(X) u ,Y v ) represents the Xth u ,Y v The joint probability of the occurrence of _x_ information events, P(X_1) u ) and P(Y v ) represent the Xth u and Y v The probability of an information event occurring is used to quantitatively describe the correlation between two sets of events using mutual information values. The larger the mutual information value, the greater the degree of interdependence between the two variables and the stronger the correlation. S32, after merging the mutual information values ​​calculated in step S31, arrange them in descending order, and select the historical load value and / or the indoor and outdoor characteristics corresponding to all mutual information values ​​not less than a set value as feature values ​​and arrange them in descending order to form a candidate pool.

Citation Information

Patent Citations

  • Building cooling load prediction method based on improved FENN and related device

    CN116384570A

  • Clustering and transfer learning-based proxy electricity purchasing user load prediction method and system

    CN118554424A