Heat supply pipe network prediction method, system and equipment based on statistics and medium

Through a statistical method, the Spearman correlation coefficient is used to screen features and construct a gradient boosting tree model, which solves the complexity and computing resource consumption problems in heating pipeline network modeling and prediction, and realizes efficient, accurate prediction and dynamic adjustment of the heating system.

CN120671905APending Publication Date: 2025-09-19QINGDAO ENERGY TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510766895.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing heating network modeling and prediction methods have problems such as high data complexity, large consumption of computing resources, poor adaptability and insufficient utilization of real-time data, making it difficult to establish accurate and efficient prediction models under complex conditions.

Method used

A statistical method is used to screen features through the Spearman correlation coefficient, and a gradient boosting tree model is constructed. Combined with data preprocessing and model optimization, efficient prediction of the heating network is achieved.

Benefits of technology

It reduces model complexity, improves computational efficiency, ensures efficient operation and accurate prediction under large-scale data sets, and dynamically adjusts heating parameters through real-time monitoring, thereby improving the operating efficiency of the heating system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671905A_ABST
    Figure CN120671905A_ABST
Patent Text Reader

Abstract

The invention relates to a heat supply pipe network prediction method based on statistics, and the method comprises the steps: obtaining related data from a heat supply pipe network system, and the related data comprise meteorological data, user demand data, pipe network operation state data and other factors which may affect a heat supply load; a reference time period is set for each characteristic value at different acquisition time points, and average value calculation is performed on sampling data of all the characteristics in the time period to obtain a representative value in the reference time period, so that data deviation caused by asynchronous acquisition time sequences is eliminated; the uniformity and consistency of the characteristic values in the time dimension are ensured; checking and filling missing values: filling missing data points in the original data by adopting an interpolation method, and calculating and filling specific numerical values of the missing data on the basis of numerical value change trends of previous and later time periods or characteristic values with relatively high correlation, so that the continuity and the integrity of the whole time sequence are ensured; and the preprocessed data is stored in a database, so that subsequent analysis and modeling are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of data science and artificial intelligence, and in particular to intelligent modeling and prediction technologies in heating engineering. Background Art

[0002] Heating networks are a vital component of urban infrastructure, responsible for transporting heat from heat sources to individual users. With the acceleration of urbanization and the growth of energy demand, the efficient operation of heating systems has become increasingly important. However, heating network data is complex and diverse, involving multiple factors such as meteorological conditions, user needs, and network structure. This results in high data dimensionality and significant feature redundancy. Developing accurate and efficient prediction models under these complex conditions is a major challenge.

[0003] In recent years, with the development of big data and artificial intelligence technologies, machine learning and deep learning methods have been widely used in the modeling and prediction of heating networks. While these methods have improved prediction accuracy to a certain extent, they also increase the complexity of the models, and many problems still exist in practical applications.

[0004] At present, the main technologies for modeling and prediction of heating pipeline networks are divided into three categories: modeling methods based on physical theory, modeling methods based on data prediction, and modeling methods that combine the two.

[0005] Among them, the modeling method based on physical theory is based on the physical characteristics of the heating pipeline network. This method is highly complex and usually requires complex and large-scale differential calculations; it has low data dependence and makes less use of real-time data, which may not capture the dynamic changes in the actual system; it has poor adaptability and needs to readjust or re-establish the physical model in the face of changes in system parameters. The modeling method based on data prediction mainly relies on historical operating data and statistical learning algorithms. This method has high requirements for data quality and requires a large amount of high-quality historical data to train the model; data selection is difficult and requires accurate selection of relevant features for model establishment and prediction. The modeling method that combines the two attempts to combine physical principles with data analysis. This method is more complex and needs to process more variables and complex calculation processes; it consumes a lot of resources and requires the ability to simultaneously possess physical modeling and data analysis, and has high requirements for hardware and computing resources. Summary of the Invention

[0006] To solve the above problems, a statistically based heating network prediction method is provided. The technical solutions provided are as follows, including:

[0007] Obtain relevant data from the heating network system, including meteorological data, user demand data (historical heat consumption), network operation status data (pressure, flow), and other factors that may affect the heating load;

[0008] Set a benchmark time period for each feature value collected at different time points, calculate the average value of all feature sampling data within this time period, and obtain the representative value within this benchmark time period. This eliminates data deviations caused by asynchronous collection timing and ensures the uniformity and consistency of each feature value in the time dimension.

[0009] Check and fill missing values: For missing data points in the original data, interpolation is used to fill them. Based on the numerical change trend of the previous and next time periods or the characteristic values ​​with high correlation, the specific values ​​of the missing data are calculated and filled in, thereby ensuring the continuity and integrity of the entire time series;

[0010] Storing preprocessed data in a structured database for subsequent analysis and modeling;

[0011] Calculate the Spearman correlation coefficient between each feature; the calculation formula is:

[0012]

[0013] in:

[0014] Rx, Ry are the positions of x and y respectively, Respectively represent the average rank;

[0015] The Spearman correlation coefficient reflects the strength and direction of the monotonic relationship between two variables, and its value range is between -1 and +1;

[0016] Save the calculated correlation coefficient table into the database;

[0017] Select the target feature and find the correlation coefficient between the target feature and other features;

[0018] According to the set threshold, features with higher absolute values ​​of correlation coefficients are screened out, and features with little impact on prediction or irrelevant features are removed, thereby reducing the data dimension;

[0019] The selected feature set is used as input variables to train the gradient boosting tree (GTB) model.

[0020] Based on the above technical solution, the specific steps for optimizing the gradient boosting tree (GTB) model are as follows:

[0021] Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function

[0022] For each step m = 1 to M (total number of iterations);

[0023] Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value:

[0024]

[0025] Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), such as minimizing the squared error loss:

[0026]

[0027] where h(x i ) is the true value of the point;

[0028] The output of the new predictor is added to the overall prediction with a learning rate η:

[0029] f m (x) = f m-1 (x)+ηh m (x)

[0030] After M iterations, the final prediction function is obtained:

[0031]

[0032] Use the training data set to fit the initial model and obtain preliminary prediction results;

[0033] The mean square error (MSE), mean absolute error (MAE) and R square coefficient (R 2 ) and other indicators to evaluate the prediction effect in regression tasks;

[0034] Compare the predicted results with the actual measurement results. If the difference between the two exceeds the set threshold, an alarm will be triggered and the data for that period will be added to the error log.

[0035] The trained model is applied to the load forecast of the actual heating system; through real-time monitoring and data analysis, accurate prediction of the heat load is achieved, and the heating parameters are dynamically adjusted according to the prediction results.

[0036] Based on the above technical solutions, we gradually optimize the loss function to improve the model performance. The expression for constructing the model is as follows:

[0037] Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function;

[0038] For each step m = 1 to M (total number of iterations);

[0039] Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value:

[0040]

[0041] Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), such as minimizing the squared error loss:

[0042]

[0043] where h(x i ) is the true value of the point

[0044] The output of the new predictor is added to the overall prediction with a learning rate η:

[0045] f m (x) = f m-1 (x)+ηh m (x)

[0046] After M iterations, the final prediction function is obtained:

[0047]

[0048] Based on the above technical solutions, the trained model is applied to load forecasting of actual heating systems. Through real-time monitoring and data analysis, accurate prediction of heat load is achieved, and heating parameters are dynamically adjusted based on the prediction results.

[0049] The predicted results are compared with the actual measurement results. If the difference between the two exceeds the set threshold, an alarm is triggered and the data for that period is added to the error log.

[0050] On the basis of the above technical solution, during the operation, the obtained measurement data is continuously incorporated into the database, and after a certain amount of data has been accumulated, the model is rebuilt to adapt to the working status of the heating network;

[0051] The statistically based heating network prediction system provides the following technical solutions, including:

[0052] Data acquisition module, used to collect historical heating data of the heating network, historical meteorological data and historical data of each node;

[0053] The data preprocessing module is used to set a reference time period for each feature value collected at different time points, calculate the average value of the sampled data of all features within the time period to obtain the representative value within the reference time period; and use interpolation to supplement missing data points;

[0054] Feature extraction module, used to extract key features with high correlation with the target by calculating the Spearman correlation coefficient;

[0055] Model building module, used to build GTB model using key features;

[0056] The model new connection module is used to train the GTB model using historical data to optimize the model parameters;

[0057] The optimized alarm module is used to use the GTB model to make real-time predictions of future heating loads and compare the predicted results with the actual measurement results. If the difference between the two exceeds the set threshold, an alarm will be issued and the data for that period will be added to the error log.

[0058] Based on the above technical solution, historical heating data refers to the daily heating supply of the heating station;

[0059] Historical meteorological data refers to the external ambient temperature;

[0060] The historical data of each node refers to the temperature measured by the sensors at each node in the heating network;

[0061] Data preprocessing means taking two hours as a benchmark unit and using the average value of all data within these two hours as the representative value of this benchmark unit to ensure the time synchronization and consistency of the data;

[0062] Interpolation is used to fill in missing data points to ensure time continuity.

[0063] An electronic device provides the following technical solution, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the statistical-based heating network prediction method according to claim 1.

[0064] A computer-readable storage medium provides the following technical solution, which stores a computer program, and when the computer program is executed by a processor, implements the statistical-based heating network prediction method as described in claim 1.

[0065] Beneficial effects:

[0066] 1. Use the Spearman correlation coefficient to evaluate the strength of the relationship between each feature and the target variable (heat load demand). This not only screens out highly correlated features but also identifies nonlinear relationships, helping to reduce model complexity.

[0067] Reduce computational costs: By removing irrelevant or weakly relevant features, the computational effort during model training and prediction is reduced, improving efficiency.

[0068] 2. Choose gradient boosting tree models such as XGBoost and LightGBM, which can handle large amounts of data and optimize performance by automatically adjusting parameters. Suitable for large-scale data sets: This technology is particularly suitable for large-scale data sets commonly found in heating systems, ensuring efficient operation and accurate prediction in big data environments.

[0069] 3. Integrate the trained model into the real-time monitoring platform of the heating system and dynamically adjust the heating parameters based on the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 This is a flow chart of the heating pipe network prediction method of the present invention.

[0071] Figure 2 For the present invention, Figure 1 The prediction method is introduced into the gradient tree to form a flow chart of the heating network prediction system. DETAILED DESCRIPTION

[0072] Example 1.

[0073] This embodiment provides a method for predicting a heating network using statistics.

[0074] Obtain relevant data from the heating network system, including meteorological data, user demand data (historical heat consumption), network operation status data (pressure, flow), and other factors that may affect the heating load.

[0075] Set a benchmark time period (for example, two hours) for each feature value collected at different time points, calculate the average value of the sampled data of all features within this time period, and obtain the representative value within this benchmark time period. This can eliminate data deviations caused by asynchronous collection timing and ensure the uniformity and consistency of each feature value in the time dimension.

[0076] Check and fill missing values: Missing data points in the original data are filled using interpolation. Specifically, this method calculates and fills missing values ​​based on the numerical trend of the preceding and following time periods or highly correlated feature values, thus ensuring the continuity and integrity of the entire time series.

[0077] The preprocessed data is stored in a structured database to facilitate subsequent analysis and modeling.

[0078] Calculate the Spearman correlation coefficient between each feature. The calculation formula is:

[0079]

[0080] in:

[0081] Rx, Ry are the positions of x and y respectively, Represents the average ranking

[0082] The Spearman correlation coefficient reflects the strength and direction of the monotonic relationship between two variables, and its value range is between -1 and +1.

[0083] Save the calculated correlation coefficient table into the database;

[0084] Select the target feature and find the correlation coefficient between the target feature and other features;

[0085] According to the set threshold, features with higher absolute values ​​of correlation coefficients are screened out, and features with little impact on prediction or irrelevant features are removed, thereby reducing the data dimension;

[0086] The selected feature set is used as input variables to train the gradient boosting tree (GTB) model.

[0087] The specific steps for optimizing the Gradient Boosting Tree (GTB) model are:

[0088] Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function

[0089] For each step m = 1 to M (total number of iterations)

[0090] Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value:

[0091]

[0092] Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), for example, minimizing the squared error loss:

[0093]

[0094] where h(x i ) is the true value of the point

[0095] The output of the new predictor is added to the overall prediction with a learning rate η:

[0096] f m (x) = f m-1 (x)+ηh m (x)

[0097] After M iterations, the final prediction function is obtained:

[0098]

[0099] The initial model is fitted using the training data set to obtain preliminary prediction results.

[0100] The mean square error (MSE), mean absolute error (MAE) and R square coefficient (R 2 ) and other indicators to evaluate the prediction effect in regression tasks

[0101] Compare the predicted results with the actual measured results. If the difference between the two exceeds the set threshold, an alarm will be triggered and the data for that period will be added to the error log.

[0102] The trained model is applied to load forecasting in actual heating systems. Through real-time monitoring and data analysis, accurate predictions of heat loads are achieved, and heating parameters are dynamically adjusted based on the prediction results.

[0103] By calculating the Spearman correlation coefficient between the target and other features, the features with a correlation coefficient greater than the set threshold are selected as relevant features to participate in the model establishment. The formula for calculating the Spearman correlation coefficient is:

[0104]

[0105] Among them: Rx, Ry are the positions of x and y respectively, Represent the average ranking respectively.

[0106] Gradually optimize the loss function to improve model performance. The expression for building the model is as follows:

[0107] Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function

[0108] For each step m = 1 to M (total number of iterations)

[0109] Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value:

[0110]

[0111] Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), for example, minimizing the squared error loss:

[0112]

[0113] where h(x i ) is the true value of the point

[0114] The output of the new predictor is added to the overall prediction with a learning rate η:

[0115] f m (x) = f m-1 (x)+ηh m (x)

[0116] After M iterations, the final prediction function is obtained:

[0117]

[0118] Example 2.

[0119] The present invention proposes a GTB model for a heating network based on Spearman correlation coefficient feature selection, including:

[0120] Obtain relevant data from the heating network system, including meteorological data, user demand data (historical heat consumption), network operation status data (pressure, flow), and other factors that may affect the heating load.

[0121] Set a benchmark time period (for example, two hours) for each feature value collected at different time points, calculate the average value of the sampled data of all features within this time period, and obtain the representative value within this benchmark time period. This can eliminate data deviations caused by asynchronous collection timing and ensure the uniformity and consistency of each feature value in the time dimension.

[0122] Check and fill missing values: Missing data points in the original data are filled using interpolation. Specifically, this method calculates and fills missing values ​​based on the numerical trend of the preceding and following time periods or highly correlated feature values, thus ensuring the continuity and integrity of the entire time series.

[0123] The preprocessed data is stored in a structured database to facilitate subsequent analysis and modeling.

[0124] Calculate the Spearman correlation coefficient between each feature. The calculation formula is:

[0125]

[0126] in:

[0127] Rx, Ry are the positions of x and y respectively, Represents the average ranking

[0128] The Spearman correlation coefficient reflects the strength and direction of the monotonic relationship between two variables, and its value range is between -1 and +1.

[0129] Save the calculated correlation coefficient table to the database

[0130] Select the target feature and find the correlation coefficient between the target feature and other features

[0131] According to the set threshold, features with higher absolute values ​​of correlation coefficients are screened out, and features with little impact on prediction or irrelevant features are removed, thereby reducing the data dimension.

[0132] The selected feature set is used as input variables to train the gradient boosting tree (GTB) model.

[0133] Gradient boosting tree is a powerful ensemble learning method that improves model performance by gradually optimizing the loss function. The specific steps are:

[0134] Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function

[0135] For each step m = 1 to M (total number of iterations)

[0136] Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value:

[0137]

[0138] Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), for example, minimizing the squared error loss:

[0139]

[0140] where h(x i ) is the true value of the point

[0141] The output of the new predictor is added to the overall prediction with a learning rate η:

[0142] f m (x) = f m-1 (x)+ηh m (x)

[0143] After M iterations, the final prediction function is obtained:

[0144]

[0145] The initial model is fitted using the training data set to obtain preliminary prediction results.

[0146] The mean square error (MSE), mean absolute error (MAE) and R square coefficient (R 2 ) and other indicators to evaluate the prediction effect in regression tasks

[0147] Compare the predicted results with the actual measured results. If the difference between the two exceeds the set threshold, an alarm will be triggered and the data for that period will be added to the error log.

[0148] The trained model is applied to load forecasting in actual heating systems. Through real-time monitoring and data analysis, accurate predictions of heat loads are achieved, and heating parameters are dynamically adjusted based on the prediction results.

[0149] Example 3.

[0150] This embodiment further provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the statistical-based heating network prediction method according to claim 1.

[0151] Example 4.

[0152] This embodiment further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the statistical-based heating network prediction method as claimed in claim 1.

[0153] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0154] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A statistically based heat supply network prediction method, characterized in that: Includes: Obtain relevant data from the heating network system, including meteorological data, user demand data, network operation status data, and other factors that may affect the heating load; Set a benchmark time period for each feature value collected at different time points, calculate the average value of all feature sampling data within this time period, and obtain the representative value within this benchmark time period. This eliminates data deviations caused by asynchronous collection timing and ensures the uniformity and consistency of each feature value in the time dimension. Check and fill missing values: For missing data points in the original data, interpolation is used to fill them. Based on the numerical change trend of the previous and next time periods or the characteristic values ​​with high correlation, the specific values ​​of the missing data are calculated and filled in, thereby ensuring the continuity and integrity of the entire time series; Storing preprocessed data in a structured database for subsequent analysis and modeling; Calculate the Spearman correlation coefficient between each feature; The calculation formula is: in: Rx, Ry are the positions of x and y respectively, Respectively represent the average rank; The Spearman correlation coefficient reflects the strength and direction of the monotonic relationship between two variables, and its value range is between -1 and +1; Save the calculated correlation coefficient table into the database; Select the target feature and find the correlation coefficient between the target feature and other features; According to the set threshold, features with higher absolute values ​​of correlation coefficients are screened out, and features with little impact on prediction or irrelevant features are removed, thereby reducing the data dimension; The selected feature set is used as input variables to train the gradient boosting tree (GTB) model.

2. The statistical-based heat supply network prediction method according to claim 1, characterized in that: The specific steps for optimizing the Gradient Boosting Tree (GTB) model are: Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function; For each step m = 1 to M (total number of iterations); Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value: Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), for example, minimizing the squared error loss: where h(x i ) is the true value of the point; The output of the new predictor is added to the overall prediction with a learning rate η: f m (x)=f m-1 (x)+ηh m (x) After M iterations, the final prediction function is obtained: Use the training data set to fit the initial model and obtain preliminary prediction results; The mean square error (MSE), mean absolute error (MAE) and R square coefficient (R 2 ) and other indicators to evaluate the prediction effect in regression tasks; Compare the predicted results with the actual measurement results. If the difference between the two exceeds the set threshold, an alarm will be triggered and the data for that period will be added to the error log. The trained model is applied to the load forecast of the actual heating system; through real-time monitoring and data analysis, accurate prediction of the heat load is achieved, and the heating parameters are dynamically adjusted according to the prediction results.

3. The statistical-based heat supply network prediction method according to claim 2, characterized in that: Gradually optimize the loss function to improve model performance. The expression for building the model is as follows: Let the initial prediction function be a constant f0(x)=c, where c is determined by minimizing the loss function; For each step m = 1 to M (total number of iterations); Calculate the residual of the current model, which is the result of the derivative of the loss function with respect to the current predicted value: Fit a weak predictor, using the residual r mi As the new target variable, train a weak prediction tree h m (x), for example, minimizing the squared error loss: where h(x i ) is the true value of the point The output of the new predictor is added to the overall prediction with a learning rate η: f m (x)=f m-1 (x)+ηh m (x) After M iterations, the final prediction function is obtained:

4. The statistical-based heat supply network prediction method according to claim 3, characterized in that: Apply the trained model to load forecasting of actual heating systems; achieve accurate prediction of heat load through real-time monitoring and data analysis, and dynamically adjust heating parameters based on the prediction results; The predicted results are compared with the actual measurement results. If the difference between the two exceeds the set threshold, an alarm is triggered and the data for that period is added to the error log.

5. The statistical-based heat supply network prediction method according to claim 4, characterized in that: During operation, the obtained measurement data are continuously incorporated into the database, and after a certain amount of data has been accumulated, the model is rebuilt to adapt to the working status of the heating network.

6. A statistically based heat supply network prediction system, using the method of claim 1, characterized in that: Includes: Data acquisition module, used to collect historical heating data of the heating network, historical meteorological data and historical data of each node; The data preprocessing module is used to set a reference time period for each feature value collected at different time points, calculate the average value of the sampled data of all features within the time period to obtain the representative value within the reference time period; and use interpolation to supplement missing data points; Feature extraction module, used to extract key features with high correlation with the target by calculating the Spearman correlation coefficient; Model building module, used to build GTB model using key features; The model new connection module is used to train the GTB model using historical data to optimize the model parameters; The optimized alarm module is used to use the GTB model to make real-time predictions of future heating loads and compare the predicted results with the actual measurement results. If the difference between the two exceeds the set threshold, an alarm will be issued and the data for that period will be added to the error log.

7. The system according to claim 6, wherein: Historical heating data refers to the daily heating supply of the heating station; Historical meteorological data refers to the external ambient temperature; The historical data of each node refers to the temperature measured by the sensors at each node in the heating network; Data preprocessing means taking two hours as a benchmark unit and using the average value of all data within these two hours as the representative value of this benchmark unit to ensure the time synchronization and consistency of the data; Interpolation is used to fill in missing data points to ensure time continuity.

8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the statistical-based heating network prediction method according to claim 1.

9. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed by a processor, implements the statistical-based heating network prediction method as claimed in claim 1.