Intelligent regulation and control method for distributed electric boiler heat storage heating system

The heating power prediction model constructed through the random forest algorithm solves the problem of inaccurate boiler power regulation in the distributed electric boiler heat storage heating mode, realizes accurate prediction of heating power, improves energy utilization and reduces operating costs.

CN120450319APending Publication Date: 2025-08-08LIAONING UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510534775.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing heating power prediction methods are difficult to cope with the complex environmental factors and user needs in the distributed electric boiler heat storage heating mode, resulting in inaccurate boiler power regulation and affecting energy utilization.

Method used

The heating power prediction model is constructed using a random forest algorithm. By processing fuzzy meteorological information, the effective feature input vector is determined, and combined with Pearson's correlation coefficient analysis and grid search optimization model parameters, the accurate prediction of heating power is achieved.

Benefits of technology

It improves the proximity between the total production and total consumption of heat resources in the heating season, improves energy utilization, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450319A_ABST
    Figure CN120450319A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent regulation and control method for a distributed electric boiler heat storage and supply system, and belongs to the field of energy management and prediction. According to the method, firstly, historical heating data and related influence factors such as air temperature, weather and room temperature are collected and arranged, then a random forest model is constructed, and the relation between each feature and the heating power is learned through training data. And then, predicting the future short-term heating power by using the trained model. The method not only can process a non-linear relationship, but also has relatively high robustness, and can adapt to different climate changes and user requirements. In order to verify the effectiveness of the method, evaluation and analysis are performed through a plurality of indexes, and the result shows that the method has relatively high performance in prediction precision and calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of heating energy consumption management, specifically a heating power prediction model and solution based on the random forest algorithm. This system utilizes machine learning techniques, specifically the random forest algorithm used in ensemble learning, to accurately analyze environmental data to effectively predict heating power to meet user heating needs. Background Art

[0002] Accurately predicting heating power is crucial for modern building heating systems. This impacts not only indoor comfort but also energy consumption and operating costs. Traditional heating power prediction methods often rely on empirical formulas and linear regression models, which struggle to adapt to changing environmental factors and complex user needs. Furthermore, with climate change and rising building energy efficiency standards, the applicability of these traditional methods is declining.

[0003] With the development of the heating industry, the heating model has also changed from traditional coal-fired centralized heating to a distributed electric boiler heat storage heating model. This model takes advantage of tiered electricity prices and hot water storage tanks to improve energy utilization. Specifically, this model uses a two-level pipeline network. The first-level pipeline network forms a primary loop from the electric boiler to the hot water storage tank; the second-level pipeline network forms a secondary loop from the hot water storage tank to the heating user. The heating company mainly wants to take advantage of the tiered electricity price. During the off-peak period, the electric boiler is turned on to heat the hot water storage tank, storing the heat in the hot water storage tank, and then use the hot water storage tank for heating during the day or other periods of electricity price. This heating model is more complex, and the precise power control of the electric boiler is more challenging.

[0004] Due to this shift in heating models, research on boiler power for distributed electric boiler thermal storage heating requires analysis of specific heating scenarios before considering the model's generalizability. Certain heating scenarios also suffer from low-quality historical datasets, such as insufficiently detailed meteorological information or overly vague descriptions, which presents challenges in determining model eigenvectors. Currently, neither industry nor academia has developed a more effective and precise boiler power control solution for this new heating model. Summary of the Invention

[0005] To address the challenge of accurately predicting boiler power in existing new heating systems, this paper proposes an intelligent control method for distributed electric boiler thermal storage heating systems. This solution consists of two main components: determining effective feature input vectors when meteorological information is ambiguous, and accurately predicting daily underground heating power at heating sites. Ultimately, this solution ensures that total heat generation throughout the heating season is approximately equal to total heat consumption, achieving optimal energy utilization.

[0006] The present invention is achieved through the following technical solutions:

[0007] An intelligent control method for a distributed electric boiler thermal storage heating system, characterized by the following steps:

[0008] Step 1) Processing fuzzy meteorological information: Construct the correlation between fuzzy meteorological information and illumination, and establish the meteorological information matrix W ij * .

[0009] First, obtain all different meteorological information in the dataset, split the meteorological information in each column according to “~” to obtain a meteorological array W and remove duplicate elements;

[0010] Then, the obtained meteorological array W is sorted according to the light intensity to obtain W 1 , W 1 The element values of are replaced by integers, the first element is set to 1 and the subsequent elements are incremented by 1, thus obtaining the weather vector W 2 =[1,2,3,......,n]; In order to maintain generality, a meteorological matrix W* is established to describe the meteorological information containing “~”; where,

[0011] Finally, traverse the meteorological information of the original data set and replace the corresponding meteorological information with W ij * The value of .

[0012] Step 2) Analyze the correlation between each feature and heating power, and determine the eigenvector θ* = [outdoor maximum temperature, outdoor minimum temperature, outdoor average temperature, weather information, yesterday's power on, indoor temperature at 9:00, indoor temperature at 14:00], where the weather information will be represented by W ij * Replace it with the specific value in .

[0013] 2.1) A pre-selected eigenvector θ = [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, weather information, date, indoor temperature at 9:00 AM, indoor temperature at 2:00 PM] is used to analyze the correlation between each eigenvalue in the eigenvector θ and the heating power of the heating site using the Pearson correlation coefficient;

[0014] Since the meteorological information in the dataset contains fuzzy information, the meteorological information in the pre-selected feature vector θ will be represented by W ij * The specific value in is used to replace it to facilitate model training; when the weather information in the dataset is sunny to cloudy, first, the weather information "sunny" is replaced by the W obtained from the dataset. 2 Determine its location i; secondly, the weather information "cloudy" is based on W 2 Determine its position j; then obtain the specific value W in the meteorological matrix W*ij * To replace the vague weather information "sunny ~ cloudy".

[0015] The Pearson correlation coefficient is a statistic used to measure the strength of the linear relationship between two variables. It is represented by the symbol r and the formula is:

[0016]

[0017] in: are the sample means of X and Y respectively, the numerator is the covariance, and the denominator is the product of the standard deviations;

[0018] The Pearson correlation coefficient between power on and indoor temperature at 9:00 AM is -0.12, the Pearson correlation coefficient between power on and indoor temperature at 2:00 PM is -0.11, the Pearson correlation coefficient between power on and date is 0.07, the Pearson correlation coefficient between power on and maximum outdoor temperature is -0.60, the Pearson correlation coefficient between power on and minimum outdoor temperature is -0.66, the Pearson correlation coefficient between power on and average outdoor temperature is -0.66, the Pearson correlation coefficient between power on and weather information is -0.15, and the Pearson correlation coefficient between power on and yesterday's power on is 0.89.

[0019] 2.2) Select parameters with high correlation coefficients and determine the eigenvector as θ* = [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, meteorological information, power on yesterday, indoor temperature at 9:00, indoor temperature at 14:00].

[0020] Step 3) Use Random Forest to Build a Prediction Model: The feature matrix determined in Step 2 is θ*, and the data to be predicted is R = [power on, total power]. θ* is used as the input matrix, and R is the output matrix. A machine learning model is fitted, using a random forest regressor and grid search cross-validation to optimize model parameters. The hyperparameter combination is RandomForestRegressor(max_depth = 10, min_samples_leaf = 3, max_features = log2, min_samples_split = 15, n_estimators = 300).

[0021] 3.1) Use θ* as the input matrix and R as the output matrix to fit the machine learning model;

[0022] Using random forest regressor and grid search cross validation to optimize model parameters;

[0023] A parameter grid is defined, which includes different hyperparameter combinations: the number of decision trees n_estimators, the maximum number of features max_features, the maximum depth max_depth, the minimum number of samples to be split at each node min_samples_split, and the minimum number of samples for leaf nodes min_samples_leaf; the number of decision trees selected is 50, 100, 200, 250 and 300, and for the maximum number of features, the total number of features 'auto', the square root of the number of features 'sqrt' and the logarithm of the number of features 'log2' are considered; the maximum depth is None or set to 5, 10, 20, 30 to control the depth of the tree; the minimum number of sample splits is set to 2, 5, 10 or 15, which determines the minimum number of samples required for each node to split; the minimum number of samples for leaf nodes, to control the growth of the tree, is set to 1, 2, 3 and 4;

[0024] 3.2) Use GridSearchCV to create a grid search object, which will perform 5-fold cross-validation on all possible hyperparameter combinations to evaluate the performance of each set of parameters. The best hyperparameter combination obtained on the training set is RandomForestRegressor(max_depth=10,min_samples_leaf=3,max_features=log2,min_samples_split=15,n_estimators=300).

[0025] Step 4) Based on tomorrow's weather forecast, obtain the corresponding meteorological information and temperature information, input them into the prediction model obtained in step 3), and output the start-up power of the electric boiler tomorrow; the predicted value of the start-up power of the electric boiler obtained is used as a reference for heating power regulation.

[0026] Beneficial effects of the present invention: The present invention can accurately and effectively predict heating power in a distributed electric boiler heat storage heating mode through the above method. Compared with the traditional heating power calculation method, the present invention empowers the heating industry with the help of machine learning technology, and the resulting electric boiler start-up power is closer to the actual value. From the perspective of the entire heating season, the present invention makes the total demand for heat energy in the entire heating system closer to the total supply, thereby improving energy utilization, thereby achieving the purpose of reducing costs and increasing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Dataset infographic;

[0028] Figure 2 Weather matrix chart;

[0029] Figure 3 Correlation analysis diagram;

[0030] Figure 4 Model structure diagram;

[0031] Figure 5 Model hyperparameter grid plot;

[0032] Figure 6 Evaluation metrics graph. DETAILED DESCRIPTION

[0033] Step 1: Data Collection:

[0034] Meteorological data: According to research by domestic and foreign scholars, the main factors that affect heating power in meteorological information are wind speed, light, outdoor temperature, etc. These factors play a key role in heating power prediction. For example, when the wind speed is high, the building may need to increase heating due to accelerated heat loss. When the light is strong, the amount of solar radiation entering the room through the building increases, causing the indoor temperature to rise; at the same time, when the light is strong, the ground and air absorb more solar radiation, and the heat loss of the building will also decrease. Therefore, light intensity is particularly critical in heating power prediction. For this reason, it is necessary to comprehensively consider the impact of these external factors on heating predictions. Due to the low quality of historical meteorological data sets, only fuzzy meteorological information can be obtained, such as Figure 1 shown.

[0035] Indoor Temperature: Heating power is also related to the indoor temperature set by the end user. End users cannot adjust their own indoor temperature settings; the heating company monitors and controls the indoor temperature remotely. Indoor temperature data should be collected in a closed space with minimal human interference.

[0036] Outdoor temperature: Outdoor temperature has a very important influence on the level of heating power. Figure 3 The correlation analysis in

[15] can also show that the higher the outdoor temperature, the lower the corresponding heating power, and the lower the outdoor temperature, the higher the corresponding heating power. The outdoor temperatures collected by this invention are the highest temperature, lowest temperature, and average temperature of the day, and the outdoor temperature data are all obtained from the weather forecast of the day.

[0037] Yesterday's power: This method assumes that the outdoor environmental conditions of heating users in the past two days are similar, and therefore the heating power in the past two days should also be similar. Therefore, the power on the previous day is also used as a reference value to predict the future heating power. The idea is derived from the principle of time locality in computer memory, which is also derived from Figure 3 The correlation analysis was confirmed.

[0038] Finally, the collected features and historical heating power, daily total electricity consumption and other data are formed into a data set for heating forecasting.

[0039] Since the meteorological information in the dataset is too fuzzy to be quantified, such as sunny or sunny to cloudy, it is also difficult to establish a correlation between it and heating power. Therefore, it is necessary to use a certain method to construct the correlation between fuzzy meteorological information and wind speed and light intensity. The method proposed in this invention is to establish a meteorological information matrix. This method only considers the correlation between it and light intensity, such as Figure 2 shown.

[0040] First, obtain all different meteorological information in the dataset, split the meteorological information in each column according to “~” to obtain a meteorological array W and remove duplicate elements. Then, sort the obtained meteorological array W according to the light intensity to obtain W 1 , W 1 The element values of are replaced by integers, the first element is set to 1 and the subsequent elements are incremented by 1, thus obtaining the weather vector W 2 =[1,2,3,......,n]. In order to keep generality, we further establish a meteorological matrix W* to better describe meteorological information such as sunny to cloudy. Finally, traverse the meteorological information of the original data set and replace the corresponding meteorological information with W ij * The value of .

[0041] Step 2: Data preprocessing:

[0042] The collected data sets are cleaned to remove missing values and outliers to ensure the quality of the data sets. Since the historical meteorological information is too vague, it is difficult for the model to process it, which has a serious impact on the prediction of heating power. Therefore, the above-mentioned method is used for the meteorological information column in the data set, such as Figure 2 The meteorological matrix shown is processed to facilitate subsequent training of the model.

[0043] The pre-selected eigenvector θ = [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, meteorological information, date, indoor temperature at 9:00, indoor temperature at 14:00] is used to analyze the correlation between each eigenvalue in the eigenvector θ and the heating power of the heating site through the Pearson correlation coefficient.

[0044] Since the meteorological information in the dataset is too vague, the meteorological information in the pre-selected feature vector θ will be represented by W ij * For example, sunny to cloudy, first, the weather information "sunny" is replaced by the specific value W obtained from the dataset. 2 Determine its location i; secondly, the weather information "cloudy" is based on W 2 Determine its position j; then obtain the specific value W in the meteorological matrix W* ij* To replace the vague weather information "sunny ~ cloudy".

[0045] The Pearson correlation coefficient is a statistic used to measure the strength of the linear relationship between two variables, usually represented by the symbol r. Formula: in: are the sample means of X and Y respectively, the numerator is the covariance, and the denominator is the product of the standard deviations.

[0046] The correlation coefficients between each characteristic value and the heating start power are as follows: Figure 3 As shown in the figure, the Pearson correlation coefficients between power consumption and indoor temperature at 9:00 AM are -0.12, -0.11, 0.07, -0.60, -0.66, and -0.66, respectively. The Pearson correlation coefficients between power consumption and weather information are -0.15, and 0.89, respectively. The correlation analysis above shows that the correlation between date and power consumption is the lowest, even below 0.1. Therefore, date is removed from the eigenvector θ. Drawing on the principle of temporal locality in computers, it is believed that heating forecasting also follows this principle: the temperature and heating power between two consecutive days are generally similar, and so are the heating power consumption between two consecutive days. Therefore, yesterday's power consumption is added to the eigenvector θ. In the correlation coefficient analysis, the Pearson correlation coefficient between the startup power and yesterday's startup power is 0.89, which also confirms the above conjecture.

[0047] The input feature vector θ* obtained based on the above analysis is [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, meteorological information, power on yesterday, indoor temperature at 9:00 a.m., indoor temperature at 2:00 p.m.], which is then normalized to scale the feature values to the range of [0, 1] to improve model training efficiency.

[0048] Divide the dataset: Divide the normalized dataset into a training set and a test set in a ratio of 7:3, which are used for model training and testing respectively.

[0049] Step 3: Model training:

[0050] After step 2, the feature matrix is determined to be θ*, and the data to be predicted is R = [on power, total power]. θ* is used as the input matrix, R is the output matrix, and the machine learning model is used for fitting.

[0051] This paper innovatively applies the Random Forest algorithm in ensemble learning to predict heating power in a distributed electric boiler thermal storage heating system. This algorithm, a decision tree-based ensemble learning model, is widely used for classification and regression tasks. Due to its good prediction accuracy and resistance to overfitting, it has been widely used in various fields. However, it has not yet been applied to the new distributed electric boiler thermal storage heating system.

[0052] Model training setup: In this invention, we used random forest regressor and grid search cross-validation to optimize model parameters and thus improve the predictive performance of the model.

[0053] First, a parameter grid is defined, which includes different hyperparameter combinations, such as the number of decision trees (n_estimators), the maximum number of features (max_features), the maximum depth (max_depth), the minimum number of samples to be split at each node (min_samples_split), and the minimum number of samples for leaf nodes (min_samples_leaf). Specifically, 50, 100, 200, 250, and 300 decision trees were selected, and for the maximum number of features, 'auto' (the total number of features), 'sqrt' (the square root of the number of features), and 'log2' (the logarithm of the number of features) were considered; the maximum depth can be None or set to 5, 10, 20, 30 to control the depth of the tree; and the minimum number of sample splits is set to 2, 5, 10, or 15, which determines the minimum number of samples required for each node to be split; the minimum number of samples for leaf nodes, to control the growth of the tree, is set to 1, 2, 3, and 4. Figure 5 shown.

[0054] Next, a grid search object was created using GridSearchCV. This object will perform 5-fold cross validation (cv=5) on all possible hyperparameter combinations to evaluate the performance of each set of parameters. The best hyperparameter combination obtained on the training set is RandomForestRegressor (max_depth=10, min_samples_leaf=3, max_features=log2, min_samples_split=15, n_estimators=300). The optimal random forest model is used to predict heating power. The model structure diagram is shown in the figure below. Figure 4 shown.

[0055] Step 4: Heating power prediction:

[0056] The data related to tomorrow's heating supply is input using the input feature vector θ* determined above. After normalizing θ*, the trained random forest model is used to predict the heating load. The model's output provides a reference for the heating company to determine the power consumption of the electric boiler tomorrow.

[0057] This invention is based on a distributed electric boiler heat storage heating model. This heating model consists of a two-tiered network consisting of electric boilers, heat storage tanks, and heating users. The boilers are operated during the off-peak hours in the region. According to research, the off-peak hours in Anshan are between 10 p.m. and 5 a.m.

[0058] To determine tomorrow's electric boiler power, we first obtain the corresponding meteorological and temperature information based on tomorrow's weather forecast. Next, we calculate the boiler power and indoor temperature setpoints for that day and construct the model's input matrix θ* = [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, meteorological information, yesterday's power, indoor temperature at 9:00 AM, indoor temperature at 2:00 PM]. This constructed θ* is then used as input to our trained random forest model to predict the short-term electric boiler power. Finally, the predicted electric boiler power serves as a reference for heating power control.

[0059] Step 5: Result Analysis:

[0060] The predicted results were compared with the actual heating power, and the mean square error (MSE), mean absolute error (MAE), and determination coefficient R were calculated. 2 and other indicators to evaluate the accuracy of the model.

[0061] The trained model is evaluated on the test set, and the evaluation indicators are MSE, RMSE, MAE, and R2.

[0062] (1) Mean Squared Error (MSE)

[0063] Definition: The average of the squares of the differences between predicted and actual values.

[0064] formula:

[0065] (2) Root Mean Squared Error (RMSE)

[0066] Definition: The square root of the mean square error, which can be used to raise the units of the error to the same as the target value.

[0067] formula:

[0068] (3) Mean Absolute Error (MAE)

[0069] Definition: The average of the absolute values of the differences between the predicted and actual values.

[0070] formula:

[0071] (4) Coefficient of determination (R 2 Score

[0072] Definition: An indicator that reflects the model's ability to explain data changes, with a value range of 0 to 1.

[0073] formula:

[0074] The results after model evaluation are as follows Figure 6 As shown:

[0075] Mean square error: 33866.564397213515

[0076] Root mean square error: 184.02870536199922

[0077] Mean absolute error: 97.58220738513441

[0078] R2 coefficient of determination: 0.8203924531990887

[0079] According to the values of the above evaluation indicators, it can be seen that the random forest algorithm has a good performance in predicting heating power in the distributed electric boiler thermal storage heating mode.

[0080] Example 1:

[0081] In order to test the performance of the method of the present invention, a total of 936 heating data from 2018 to 2023 were collected in the school heating scenario, including the outdoor maximum temperature, outdoor minimum temperature, outdoor average temperature, meteorological information, daily total electricity consumption, daily boiler power, indoor temperature at 9:00, indoor temperature at 14:00, and date. The code in the method of the present invention was implemented in Python language and run on a computer with the following hardware information: CPU Core TM i7-10750H CPU @ 2.60GHz, graphics card: NVIDIA GeForce GTX 1660Ti, memory: 8GB, operating environment is Windows 11 64-bit system. Specific data set information is as follows Figure 1 shown.

[0082] Data preprocessing:

[0083] Before model training, proper data preprocessing is essential. First, missing values need to be handled. For missing items that may exist in the original data, this paper uses a forward filling method, which fills the missing value with the previous valid value to ensure data continuity and integrity. Furthermore, feature selection is required to ensure that only features that are strongly correlated with the target variable are used when training the model.

[0084] In conjunction with the steps required for data preprocessing, the dataset must first be divided into training and test sets. In this study, 70% of the data was used as the training set, and the remaining 30% was used as the test set for validation. To enhance the model's generalization capabilities, feature normalization was also performed during the preprocessing phase. The goal is to ensure that all feature values are distributed within a similar range, thereby minimizing the impact of different features on model training results.

[0085] Model building

[0086] The method proposed in this study is based on the Random Forest Regressor model, an ensemble learning method that can handle large feature sets with high accuracy and robustness. The model construction process is implemented using the Python programming language and its related libraries, such as scikit-learn.

[0087] In model setting, the selection of hyperparameters for random forest is crucial. After considering the number of data samples and characteristics, we set several representative hyperparameter combinations, including:

[0088] n_estimators: The number of trees to build. We chose 50, 100, 200, 300, and 500 trees. Generally, the greater the number of trees, the better the stability and accuracy of the model, but it also increases the computational cost.

[0089] max_features: used to control the random feature selection of each tree. "log2", "auto" and "sqrt" are used as parameter settings to find the best feature combination.

[0090] max_depth: Limits the maximum depth of the tree to prevent overfitting. It can be set to None (unlimited) and specific values of 5, 10, 20, and 30.

[0091] min_samples_split: controls the minimum number of samples required for a node to split to avoid excessive model complexity, and is set to 2, 5, 10, and 15.

[0092] min_samples_leaf: defines the minimum number of samples for leaf nodes to control the growth of the tree, and is set to 1, 2, 3, and 4.

[0093] In order to optimize the above hyperparameter combinations, a grid search method was used for cross-validation to evaluate the performance of the model under various combinations and finally select the best performing parameters.

[0094] Model training and evaluation:

[0095] After the model training is completed, its performance on the test set needs to be verified. To evaluate the accuracy of the model, we use the root mean square error (RMSE) and the coefficient of determination (R 2 ) as the evaluation metric. The model predicts the boiler power demand of the test set and compares it with the actual value. Through RMSE, we can intuitively understand the deviation between the model prediction value and the actual value, while R 2 It can be used to measure the model's ability to explain data variability.

[0096] Experimental results analysis:

[0097] After evaluating the model performance on the test dataset, the results show that the random forest model shows good accuracy in heating power prediction. Root mean square error (RMSE): 184.02870536199922; R 2 Coefficient of determination: 0.8203924531990887; the model shows a low root mean square error, indicating that the model's prediction results are relatively close to the true value; at the same time, the coefficient of determination R 2 A value close to 1 indicates that the model can fully account for data variability. This demonstrates the feasibility of the proposed random forest method in practical applications. The model can flexibly adjust the predicted boiler power based on historical heating data, depending on varying climate conditions and indoor and outdoor temperature variations.

Claims

1. An intelligent control method for a distributed electric boiler thermal storage heating system, characterized in that: Here are the steps: Step 1) Processing fuzzy meteorological information: Construct the correlation between fuzzy meteorological information and light intensity, and establish the meteorological information matrix W ij * ; Step 2) Analyze the correlation between each feature and heating power, and determine the eigenvector θ* = [outdoor maximum temperature, outdoor minimum temperature, outdoor average temperature, weather information, yesterday's power on, indoor temperature at 9:00, indoor temperature at 14:00], where the weather information will be represented by W ij * Replace the specific value in ; Step 3) Use random forest to build a prediction model: After step 2, the feature matrix is determined to be θ*, and the data to be predicted is R = [power on, total power]. θ* is used as the input matrix and R is the output matrix. Fit the model using a machine learning model, using a random forest regressor and grid search cross-validation to optimize the model parameters. The hyperparameter combination is RandomForestRegressor(max_depth=10,min_samples_leaf=3,max_features=log2,min_samples_split=15,n_estimators=300). Step 4) Based on tomorrow's weather forecast, obtain the corresponding meteorological information and temperature information, input them into the prediction model obtained in step 3), and output the start-up power of the electric boiler tomorrow; the predicted value of the start-up power of the electric boiler obtained is used as a reference for heating power regulation.

2. The intelligent control method for a distributed electric boiler thermal storage heating system according to claim 1, characterized in that: In the step 1), the specific method is: First, obtain all different meteorological information in the dataset, split the meteorological information in each column according to "~" to obtain a meteorological array W for each column and remove duplicate elements; Then, the obtained meteorological array W is sorted according to the light intensity to obtain W 1 , W 1 The element values of are replaced by integers, the first element is set to 1 and the subsequent elements are incremented by 1, thus obtaining the weather vector W 2 =[1,2,3,......,n]; In order to maintain generality, a meteorological matrix W* is established to describe the meteorological information containing "~"; where, Finally, traverse the meteorological information of the original data set and replace the corresponding meteorological information with W ij * value.

3. The intelligent control method for a distributed electric boiler thermal storage heating system according to claim 2, characterized in that: In step 2), the method for determining the eigenvector θ* is: 2.1) A pre-selected eigenvector θ = [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, weather information, date, indoor temperature at 9:00 AM, indoor temperature at 2:00 PM] is used to analyze the correlation between each eigenvalue in the eigenvector θ and the heating power of the heating site using the Pearson correlation coefficient; Since the meteorological information in the dataset contains fuzzy information, the meteorological information in the pre-selected feature vector θ will be represented by W ij * The specific values in are replaced to facilitate model training; The Pearson correlation coefficient is a statistic used to measure the strength of the linear relationship between two variables. It is represented by the symbol r and the formula is: in: are the sample means of X and Y respectively, the numerator is the covariance, and the denominator is the product of the standard deviations; The Pearson correlation coefficient between power on and indoor temperature at 9:00 AM is -0.12, the Pearson correlation coefficient between power on and indoor temperature at 2:00 PM is -0.11, the Pearson correlation coefficient between power on and date is 0.07, the Pearson correlation coefficient between power on and maximum outdoor temperature is -0.60, the Pearson correlation coefficient between power on and minimum outdoor temperature is -0.66, the Pearson correlation coefficient between power on and average outdoor temperature is -0.66, the Pearson correlation coefficient between power on and weather information is -0.15, and the Pearson correlation coefficient between power on and yesterday's power on is 0.

89. 2.2) Select parameters with high correlation coefficients and determine the eigenvector as θ* = [maximum outdoor temperature, minimum outdoor temperature, average outdoor temperature, meteorological information, power on yesterday, indoor temperature at 9:00, indoor temperature at 14:00].

4. The intelligent control method for a distributed electric boiler thermal storage heating system according to claim 3, characterized in that: In the step 2), when the weather information in the data set is sunny to cloudy, first, the weather information "sunny" is calculated based on the W obtained from the data set. 2 Determine its location i; secondly, the weather information "cloudy" is based on W 2 Determine its position j; then obtain the specific value W in the meteorological matrix W* ij * To replace the vague weather information "sunny ~ cloudy".

5. The intelligent control method for a distributed electric boiler thermal storage heating system according to claim 1, characterized in that: In the step 3), the specific method is: 3.1) Use θ* as the input matrix and R as the output matrix to fit the machine learning model; Using random forest regressor and grid search cross validation to optimize model parameters; A parameter grid is defined, which includes different hyperparameter combinations: the number of decision trees n_estimators, the maximum number of features max_features, the maximum depth max_depth, the minimum number of samples to be split at each node min_samples_split, and the minimum number of samples for leaf nodes min_samples_leaf; the number of decision trees selected is 50, 100, 200, 250 and 300, and for the maximum number of features, the total number of features 'auto', the square root of the number of features 'sqrt' and the logarithm of the number of features 'log2' are considered; the maximum depth is None or set to 5, 10, 20, 30 to control the depth of the tree; the minimum number of sample splits is set to 2, 5, 10 or 15, which determines the minimum number of samples required for each node to split; the minimum number of samples for leaf nodes, to control the growth of the tree, is set to 1, 2, 3 and 4; 3.2) Use GridSearchCV to create a grid search object, which will perform 5-fold cross-validation on all possible hyperparameter combinations to evaluate the performance of each set of parameters. The best hyperparameter combination obtained on the training set is RandomForestRegressor(max_depth=10,min_samples_leaf=3,max_features=log2,min_samples_split=15,n_estimators=300).