Natural gas consumption prediction model and method based on gas power generation

By integrating multi-source data through the random forest algorithm, a natural gas consumption prediction model was constructed, which solved the problems of cross-industry data barriers and high-dimensional nonlinear data processing, achieved high-precision natural gas consumption prediction, supported the stability of natural gas dispatch and power production, and improved enterprise profits.

CN120996262APending Publication Date: 2025-11-21SHANGHAI SHENNENG FENGXIAN THERMAL POWER CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511096073.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, natural gas consumption forecasting faces challenges such as cross-industry data barriers, a lack of successful experience in cross-domain forecasting models, and the difficulty of traditional algorithms in handling high-dimensional nonlinear data. These issues result in insufficient forecasting accuracy and an inability to effectively support the needs of natural gas dispatching and power production.

Method used

The random forest algorithm is used to integrate 11 key feature factors, including historical gas load data, meteorological data, LNG supply dynamics, and hydropower generation, to build a prediction model. Through data cleaning, feature selection, and model training, the model parameters are optimized to achieve accurate prediction of high-dimensional nonlinear data.

Benefits of technology

The model achieved a coefficient of determination of 0.989 on the training set and an R² of 0.932 on the test set, with an actual test bias of only -2.00%, significantly improving prediction accuracy. This supports the scientific scheduling of natural gas and the stable supply of electricity, thereby increasing the trading revenue of gas-fired power generation companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996262A_ABST
    Figure CN120996262A_ABST
Patent Text Reader

Abstract

The invention discloses a natural gas consumption prediction model and method based on gas power generation, and the method comprises the steps: selecting characteristic factors, such as gas load historical data, meteorological data, economic indexes, LNG supply dynamics and hydropower station generating capacity, collecting data in a certain time period by taking a day as a unit, constructing a multi-dimensional characteristic data set, obtaining original data, and carrying out the prediction of the natural gas consumption based on the original data. Collected original data are preprocessed, 11 key characteristic factors such as gas load historical data, meteorological data, LNG supply dynamics and hydropower station generating capacity are integrated through a random forest algorithm, the model training set decision coefficient (R2) reaches 0.989, the test set R2 reaches 0.932, the actual test deviation rate in 2023-2024 is only-2.00%, the average deviation rate is 1.93%, and the model training set decision coefficient (R2) reaches 0.932. The method is obviously superior to a traditional method in high-dimensional and nonlinear data processing capacity, the accurate prediction result can support scientific scheduling of natural gas, supply pressure caused by too large peak-valley difference of natural gas in Shanghai city is relieved, and stable supply of natural gas for power generation is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of natural gas consumption prediction, and particularly relates to a natural gas consumption prediction model and method based on gas power generation. BACKGROUND

[0002] With the transformation of energy structure and the promotion of new power system construction, natural gas, as a clean and efficient energy, is increasingly widely used in the field of power generation, especially in grid peak shaving. Taking Shanghai as an example, the natural gas consumption continues to grow, and the average daily gas consumption increases from 20 million cubic meters in 2020 to 30 million cubic meters in 2024, and the load peak-valley difference is significant, such as the peak-valley difference of 30 million cubic meters in some time period in January 2024. The natural gas for power generation accounts for one third of the total gas consumption, and is the core field of ensuring energy supply security.

[0003] Accurate prediction of the amount of natural gas for power generation is of great significance to both gas and electricity:

[0004] At the level of natural gas dispatching: scientific dispatching needs to be achieved through prediction to balance supply and demand fluctuations and avoid unstable supply caused by excessive peak-valley difference;

[0005] At the level of power production and transaction: gas power generation enterprises need to develop production plans based on the prediction and provide a basis for future power spot transaction to improve the income.

[0006] However, the existing prediction faces three major challenges: cross-industry data barrier: gas and electricity industry data are not connected, making it difficult to obtain complete batch operation information; there is a lack of successful experience for reference in the research and development of energy industry cross-domain prediction model; traditional algorithms such as ARIMA and linear regression are difficult to handle high-dimensional and nonlinear gas load data, and the prediction accuracy is insufficient, so we need to provide a natural gas consumption prediction model and method based on gas power generation. SUMMARY

[0007] The purpose of the application is to provide a natural gas consumption prediction model and method based on gas power generation, which integrates 11 key characteristic factors such as gas load historical data, meteorological data, LNG supply dynamics, and hydropower generation capacity through a random forest algorithm, and the determination coefficient (R 2 ) of the model training set is 0.989, and the R 2The actual test deviation rate is only-2.00% in 2023-2024, and the average deviation is 1.93%, which is significantly better than the processing ability of the traditional method for high-dimensional and nonlinear data, and the accurate prediction result can support natural gas scientific scheduling, relieve the supply pressure caused by the large peak-valley difference of natural gas in Shanghai, and ensure the stable supply of natural gas for power generation, so as to solve the problem of the cross-industry data barrier in the prior art: the data of gas and electricity industries are not connected, it is difficult to obtain complete batch operation information, and there is lack of successful experience for reference in the research and development of cross-field prediction models in the energy industry; and the traditional algorithms such as ARIMA and linear regression are difficult to process high-dimensional and nonlinear gas load data, and the prediction accuracy is insufficient.

[0008] To achieve the above object, the following technical scheme is adopted in the present application: A natural gas consumption prediction method based on gas power generation, comprising the following steps:

[0009] Selecting gas load historical data, meteorological data, economic indicators, LNG supply dynamics and hydropower station power generation and other characteristic factors, collecting data in a certain time period in units of days, constructing a multi-dimensional feature data set, and obtaining original data;

[0010] The collected original data is preprocessed, including data cleaning to remove outliers and missing values, analyzing the correlation between each characteristic factor and the target parameter by fitting, calculating the determination coefficient R2, and sorting according to the size of R2 to select key characteristic factors that significantly affect the target parameter;

[0011] A prediction model is constructed using a random forest algorithm, the data set corresponding to the selected key characteristic factors is divided into a training set and a test set, the training set is used to train the model, the model performance is optimized by adjusting the model parameters, and real-time incremental data flow is simulated to obtain a trained model;

[0012] The performance of the trained model is evaluated using the test set, and indicators such as mean square error (MSE) and determination coefficient (R2) are calculated, and if the R2 of the model on the training set and the test set is high and the results are similar;

[0013] The model is tested for a long time in the actual scene, the deviation rate of the predicted value and the actual value is calculated, the reasons for the large deviation in the test are analyzed and corresponding measures are taken, and the model is iteratively improved.

[0014] Preferably, the gas load historical data includes daily actual gas consumption, daily actual gas supply, and the difference between gas consumption and gas supply in Shanghai; the LNG supply dynamics includes daily gas supply of Yangshan LNG, daily inventory of Yangshan LNG, and daily gas supply from the first West-East Gas Transmission Pipeline to Shanghai; and the hydropower station power generation includes daily power generation of Xiangjiaba Hydropower Station and daily power generation of Three Gorges Hydropower Station.

[0015] Preferably, the meteorological data includes daily maximum temperature, daily minimum temperature, temperature difference, wherein the temperature difference is calculated by the formula:

[0016] Temperature difference = daily maximum temperature - daily minimum temperature.

[0017] Preferably, the certain period of time is from January 1, 2020 to December 31, 2022, a total of 1196 groups of data are constructed; the division ratio of the training set and the test set is 8:2, that is, 80% of the data is used for model training, and 20% of the data is used for testing.

[0018] Preferably, the parameters of the random forest algorithm include: n_estimators=55, max_depth=20, min_samples_split=3, min_samples_leaf=1, max_features='sqrt'.

[0019] Preferably, in the model performance evaluation, the coefficient of determination R 2 is calculated by the formula:

[0020]

[0021] Wherein, y_true is the actual amount of natural gas for power generation, y_pred is the model prediction value, is the average value of the actual value.

[0022] Preferably, the time range of the long-term test is from January 1, 2023 to December 31, 2024, and the deviation rate calculation formula is:

[0023] Deviation rate = (predicted natural gas consumption for power generation - actual natural gas consumption for power generation) / actual natural gas consumption for power generation x 100%;

[0024] When the absolute value of the deviation rate exceeds 10%, the model is iteratively optimized by adding real-time data of the previous 3 months to the training set.

[0025] The natural gas consumption prediction model based on gas-fired power generation based on the method, the model comprises:

[0026] Multi-source data access module: used for accessing multi-source data such as gas load historical data, meteorological data, economic indicators, LNG supply dynamics and hydropower station power generation, supporting access to actual values of the previous day at 09:00 and planned values of the next day at 20:00;

[0027] Feature preprocessing module: performs data cleaning, feature fitting and R 2 calculation, and outputs 11 key feature factors;

[0028] Random forest algorithm core module: based on the parameters, an algorithm model is constructed to support real-time incremental data stream simulation training;

[0029] Model evaluation module: calculate the MSE and R of the training set and the test set 2 , and output an evaluation report;

[0030] Prediction result output interface: output the daily natural gas consumption prediction value and the deviation rate, support the power spot trading system and the natural gas pipeline network scheduling platform.

[0031] An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the natural gas consumption prediction method based on gas power generation when executing the program.

[0032] A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the natural gas consumption prediction method based on gas power generation.

[0033] The technical effects and advantages of the present application: the natural gas consumption prediction model and method based on gas power generation proposed by the present application has the following advantages compared with the prior art:

[0034] The present application integrates 11 key characteristic factors such as gas load historical data, meteorological data, LNG supply dynamics, and hydropower generation capacity through random forest algorithm, the determination coefficient (R2) of the model training set reaches 0.989, the R2 of the test set reaches 0.932, the actual test deviation rate in 2023-2024 is only -2.00%, and the average deviation is 1.93%, which is significantly better than the processing ability of traditional methods for high-dimensional and nonlinear data, and the accurate prediction result can support natural gas scientific scheduling, relieve the supply pressure caused by the large peak-valley difference of natural gas in Shanghai, and ensure the stable supply of natural gas for power generation;

[0035] The model can convert the natural gas quantity into electricity quantity, help the gas power generation enterprise to accurately predict the power generation scale of the next day or period, and formulate production and maintenance plan. At the same time, it provides reliable day-ahead quotation basis for future power spot trading, helps the enterprise to obtain favorable electricity price in trading, and improves the income. In addition, the load prediction ability of the model for the gas turbine unit for grid peak shaving can enhance its flexible peak shaving response efficiency in the new power system, the present application successfully applies the random forest algorithm in the cross-field prediction of the energy industry for the first time, overcomes the limitations of traditional methods, provides a referenceable technical path for similar cross-industry load prediction problems, especially provides a practical example for the processing of high-dimensional and nonlinear energy data, and has strong popularization value.

[0036] Other features and advantages of the present application will be set forth in the description that follows, and in part will be apparent from the description, or can be learned by practice of the application. The purposes and other advantages of the application will be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 A natural gas daily gas consumption trend graph of Shanghai in January 2024 of the present application;

[0038] Figure 2 A multi-data scatter plot initially established by the present application;

[0039] Figure 3 A natural gas amount graph for power generation of the present application;

[0040] Figure 4 A daily natural gas amount graph for power generation of the present application;

[0041] Figure 5 A training set and test set simulation prediction situation graph of the present application;

[0042] Figure 6 A step flow schematic diagram of the present application. DETAILED DESCRIPTION

[0043] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments. The specific embodiments described herein are only used to explain the present application, and are not used to limit the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0044] The present application provides a natural gas consumption prediction method based on gas power generation as shown in Figures 1-6 The present application provides a natural gas consumption prediction method based on gas power generation as shown in

[0045] Selecting gas load historical data, meteorological data, economic indicators, LNG supply dynamics, and power generation of hydropower stations as characteristic factors, collecting data of a certain time period in units of days, constructing a multi-dimensional characteristic data set, and obtaining original data;

[0046] The collected original data is preprocessed, including data cleaning to remove outliers and missing values, analyzing the correlation between each characteristic factor and the target parameter through fitting and other methods, calculating the determination coefficient R2, sorting according to the size of R2, and screening out key characteristic factors that significantly affect the target parameter;

[0047] The random forest algorithm is used to construct a prediction model, the data set corresponding to the screened key characteristic factors is divided into a training set and a test set, the training set is used to train the model, the model performance is optimized by adjusting the model parameters, and a real-time incremental data stream is simulated to obtain the trained model;

[0048] The trained model is evaluated by using the test set, and the mean square error (MSE) and the determination coefficient (R2) are calculated. If the R2 of the model on the training set and the test set is high and the results are similar;

[0049] In the actual scene, the model is tested for a long time, the deviation rate of the predicted value and the actual value is calculated, and for the case that the deviation is large in the test, the reasons are analyzed and corresponding measures are taken to improve the model.

[0050] The gas load historical data includes the actual gas consumption of Shanghai city per day, the actual gas supply per day, the difference between the gas consumption and the gas supply; the LNG supply dynamics includes the daily gas supply of Yangshan LNG, the daily inventory of Yangshan LNG, the daily gas supply from the first West-East Gas Transmission to Shanghai; the power generation of the hydropower station includes the daily power generation of Xiangjiaba Hydropower Station and the daily power generation of Three Gorges Hydropower Station;

[0051] Specifically, the data is from the natural gas pipeline network company, in units of ten thousand cubic meters, and the actual gas consumption data of the previous day is provided at 09:00 every day, reflecting the actual consumption scale of natural gas in Shanghai city, which is the core basic data for measuring gas load. Its change is directly related to the fluctuation of the amount of natural gas for power generation, for example, from January 14 to 23, 2024, the peak-valley difference of natural gas consumption in Shanghai city reached 30 million cubic meters, which provides a key basis for the model to capture the load fluctuation characteristics; the difference between the actual gas consumption and the actual gas supply of Shanghai city per day is calculated (difference = gas consumption - gas supply), in units of ten thousand cubic meters, which directly reflects the gap or surplus of daily natural gas supply and demand. When the difference is positive, it means that the gas consumption exceeds the gas supply, and there may be a risk of supply shortage, which needs to be balanced by adjusting the amount of natural gas for power generation; when the difference is negative, the gas supply is excessive, which can provide space for adjustment of the amount of natural gas for power generation. This index provides a quantitative basis for the model to judge the tightness of supply and demand and optimize the prediction results. The amount of natural gas supplied by Yangshan LNG receiving station to Shanghai city per day, in units of ten thousand cubic meters, is one of the important sources of natural gas supply in Shanghai city. The fluctuation of its supply directly affects the availability of natural gas for power generation, for example, the increase of Yangshan LNG gas supply during the winter heating period can support the increase of natural gas for power generation. This data provides input for the model to capture the supporting role of LNG supply on natural gas for power generation.

[0052] The meteorological data includes daily maximum temperature, daily minimum temperature, and temperature difference, wherein the temperature difference calculation formula is:

[0053] Temperature difference = Daily maximum temperature - Daily minimum temperature

[0054] Specifically, the daily maximum temperature data is from the China Meteorological Administration, with units of °C, and is broadcast twice a day at 08:00 and 18:00. The model uses the morning broadcast value. The maximum temperature directly affects social electricity demand (such as the summer high temperature leading to a sharp increase in air conditioning load), and indirectly affects the amount of gas for power generation - when the temperature is too high, the power grid electricity load rises, and the gas generator set may need to increase output to meet the peak shaving demand, resulting in an increase in the amount of natural gas for power generation; conversely, under moderate temperature, the electricity load is stable, and the amount of natural gas for power generation is relatively stable. For example, in Shanghai, when the summer maximum temperature exceeds 35°C, the gas power load often increases significantly, and this data provides a key input for the model to capture the correlation between temperature and gas for power generation;

[0055] Also provided by the Central Meteorological Observatory, with units of °C, broadcast simultaneously with the maximum temperature, reflecting the lower limit of the daily temperature. The minimum temperature has a significant impact on gas demand during the winter heating period: when the minimum temperature is too low (such as below 0°C), residential and industrial heating gas consumption increases, which may compete with the amount of gas for power generation, resulting in a reduction in the amount of gas for power generation; when the minimum temperature is higher, the heating gas demand decreases, and the amount of gas for power generation can be more adequately supplied. The model optimizes the prediction logic of the amount of gas for power generation in different seasons (especially in winter) through the minimum temperature data.

[0056] Also provided by the Central Meteorological Observatory, with units of °C, broadcast simultaneously with the maximum temperature, reflecting the lower limit of the daily temperature. The minimum temperature has a significant impact on gas demand during the winter heating period: when the minimum temperature is too low (such as below 0°C), residential and industrial heating gas consumption increases, which may compete with the amount of gas for power generation, resulting in a reduction in the amount of gas for power generation; when the minimum temperature is higher, the heating gas demand decreases, and the amount of gas for power generation can be more adequately supplied. The model optimizes the prediction logic of the amount of gas for power generation in different seasons (especially in winter) through the minimum temperature data

[0057] The certain period of time is from January 1, 2020 to December 31, 2022, a total of 1196 groups of data; the division ratio of the training set and the test set is 8:2, that is, 80% of the data is used for model training, and 20% of the data is used for testing;

[0058] Specifically, the data time period is selected based on the selection of January 1, 2020 to December 31, 2022 as the time range of the basic data set, mainly based on the following considerations:

[0059] Data integrity and representativeness: This time period covers three complete natural years, including meteorological characteristics of all four seasons (such as high summer temperatures and severe winter cold), holiday electricity fluctuations (such as the Spring Festival and National Day), energy policy adjustments (such as the early stages of new power system construction), and other factors that can fully reflect the periodic and sudden changes in the amount of natural gas used for power generation. For example, the extreme high temperatures in the summer of 2021 led to a surge in electricity demand, and the amount of gas used for gas-fired power generation broke a historical peak. During the winter of 2022, heating and power generation gas demand overlapped, forming a typical load fluctuation case, providing ample samples for the model to learn prediction logic in complex scenarios.

[0060] Sample size adaptability: The 3-year data forms 1196 daily samples, which not only meets the sample size requirements of the random forest algorithm for high-dimensional feature modeling (to avoid overfitting due to insufficient samples), but also captures the lag effect of load through time series continuity (such as the impact of the previous day's gas consumption on the current day), providing data support for the model to explore feature correlation.

[0061] Training set and test set division logic: The training set (about 957 data) and the test set (about 239 data) are divided in an 8:2 ratio, with the specific logic as follows:

[0062] Training set function: 80% of the samples are used for model parameter learning, and the random forest algorithm iteratively optimizes key parameters such as the number of trees (n_estimators = 55) and the maximum depth (max_depth = 20) to enable the model to fully fit the nonlinear relationship between features and target parameters (natural gas used for power generation). For example, the summer high-temperature period data in 2020-2022 is included in the training set, which can help the model learn the correlation between "high temperature, electricity load, and gas-fired power generation gas."

[0063] Test set function: 20% of the samples are used to verify the model's generalization ability, and the test set samples are continuous in time but not overlapping with the training set, simulating the actual application scenario of "predicting the future from the past." The mean square error (MSE = 13836.55) and the determination coefficient (R 2 = 0.932) calculated from the test set can objectively evaluate the prediction accuracy of the model on unseen data, ensuring the stability of the model in actual deployment.

[0064] Division method: Time series segmentation is used instead of random sampling to avoid learning future information (such as predicting the past with future data) due to the disruption of the time sequence, ensuring the scientific nature of the division and consistency with actual application scenarios.

[0065] The parameters of the random forest algorithm include: n_estimators=55, max_depth=20, min_samples_split=3, min_samples_leaf=1, and max_features='sqrt';

[0066] Specifically, n_estimators=55

[0067] The number of decision trees in the model is 55. This parameter reduces the overfitting risk of a single decision tree by integrating the prediction results of multiple decision trees (using voting or averaging), and improves the stability of the model. After multiple experiments, when the number of decision trees is 55, the performance of the model on the training set and the test set (training set R 2 = 0.989, test set R 2 = 0.932) reaches a balance - less than 55, the model is not fully learned, and the prediction accuracy is insufficient; more than 55, the model complexity increases, the calculation cost rises but the accuracy improves limitedly, so 55 is the optimal choice considering accuracy and efficiency.

[0068] max_depth=20

[0069] The maximum depth of each decision tree is limited to 20 layers. Shallow depth (such as less than 10) will lead to underfitting of the model, which cannot capture the complex relationship between features and target parameters (natural gas consumption for power generation); too deep depth (such as greater than 30) may cause the decision tree to overfit the noise in the training data (such as abnormal gas consumption data), reducing the generalization ability. When set to 20, the model can fully learn the deep rules of multi-source features (such as the interaction between weather data and LNG supply dynamics), while avoiding overfitting and ensuring high accuracy on the test set.

[0070] min_samples_split=3

[0071] The minimum number of samples required for node splitting in the decision tree is 3. That is, when the number of samples in a node is less than 3, it will not be split and directly output the result as a leaf node. This parameter prevents the model from over-splitting small sample nodes and reduces the risk of overfitting. For the 1196 groups of data in this project, the threshold of 3 can effectively filter noise samples (such as single-day abnormal low gas consumption data), allowing the decision tree to focus on statistically significant feature rules (such as seasonal gas consumption trends).

[0072] min_samples_leaf=1

[0073] The leaf node (the final node of the decision tree) contains at least 1 sample. This parameter controls the minimum output unit granularity of the decision tree. When set to 1, it allows the model to capture fine-grained feature differences (such as different gas consumption fluctuations corresponding to different temperature intervals), while the max_depth=20 limit prevents overfitting due to excessive leaf nodes, balancing the preservation of detailed information and the maintenance of generalization ability.

[0074] max_features='sqrt'

[0075] This parameter represents the square root of the total number of features randomly selected as candidate features for each node split. For example, for the 11 key features in this project, 4 features will be randomly selected for evaluation each time the split is made (sqrt(11) ≈ 3.316, rounded up to 4), and the optimal split method is selected. This parameter reduces the correlation between decision trees by introducing randomness, enhancing the diversity of ensemble learning, and ultimately improving the overall predictive performance of the model, especially in scenarios where multiple sources of features (such as gas load, weather, LNG supply, etc.) interact significantly.

[0076] In the model performance evaluation, the coefficient of determination R 2 is calculated as follows:

[0077]

[0078] where y_true is the actual amount of natural gas for power generation, y_pred is the model prediction, and is the average of the actual values.

[0079] Specifically, the core significance of R 2

[0080] The coefficient of determination R 2 is a key indicator that measures the degree of fit between the model's predicted values and the actual values, with a value range of [0, 1]. The closer R 2 is to 1, the stronger the model's ability to explain the target parameter (natural gas for power generation), and the smaller the deviation between the predicted value and the actual value. In this model, the training set R 2 = 0.989 and the test set R 2 = 0.932, indicating that the model can explain 98.9% of the training data fluctuations and 93.2% of the test data fluctuations, with a significantly better fitting effect than traditional linear models (which usually have R 2 < 0.8).

[0081] Formula parameter analysis

[0082] ​y_true: refers to the actual amount of natural gas used for power generation in Shanghai per day (unit: ten thousand cubic meters), sourced from the actual measurement records of the natural gas pipeline network company, and is the benchmark true value for model prediction. For example, the actual amount of gas used for power generation on January 1, 2023 was 381.48 ten thousand cubic meters, which is used as y_true for calculation.

[0083] y_pred: refers to the predicted value of the amount of natural gas used for power generation per day output by the model (unit: ten thousand cubic meters), calculated by the random forest algorithm based on 11 key features such as historical gas load data and weather data. For example, the model predicted value on January 1, 2023 was 442.23 ten thousand cubic meters, which is used as y_pred for calculation.

[0084] refers to the average value of the actual amount of natural gas used for power generation (unit: ten thousand cubic meters), obtained by taking the arithmetic mean of all y_true in the training set or test set, representing the overall average level of the target parameter. For example, the average value of the actual amount of gas used for power generation in the training set from 2020 to 2022 was about 800 ten thousand cubic meters, reflecting the average gas consumption during that period.

[0085] Formula calculation logic

[0086] Numerator Σ(y_true-y_pred) 2 : represents the sum of the squared differences between the predicted value and the actual value, measuring the total error of the model prediction. The smaller the value, the smaller the deviation between the predicted value and the actual value.

[0087] Denominator represents the sum of the squared differences between the actual value and the average value, measuring the fluctuation range of the actual data itself (i.e. total deviation).

[0088] The overall formula quantifies the proportion of the model's explanation of the actual data fluctuations in the form of "1-(prediction error / total deviation)". For example, if the prediction error accounts for only 1.1% of the total deviation, then R 2 = 0.989, indicating that the model has almost captured all the effective fluctuation rules.

[0089] Application value in this model

[0090] The calculation result of this formula directly supports the judgment of the model's practicality: when the R 2 of the training set and the test set are both high and the difference is small (such as the difference of this model being only 0.057), it indicates that the model has not overfit (overfitting the training data and failing to adapt to new data) and has good generalization ability. For example, the R 2 ​= 0.932 indicates that the model can still maintain high-precision prediction on new data not involved in training, providing a reliable basis for natural gas scheduling and power generation planning in actual scenarios.

[0091] The long-term test time range is from January 1, 2023 to December 31, 2024, and the deviation rate calculation formula is:

[0092] Deviation rate = (predicted natural gas consumption for power generation - actual natural gas consumption for power generation) / actual natural gas consumption for power generation x 100%;

[0093] When the absolute value of the deviation rate exceeds 10%, the model is iteratively optimized by adding real-time data of the past three months to the training set;

[0094] Specifically, January 1, 2023 to December 31, 2024 is selected as the actual scenario test period, covering two complete natural years, a total of 730 daily data. This time period is based on the following considerations:

[0095] Scenario coverage integrity: includes typical scenarios such as seasonal weather changes (e.g. high temperature in summer 2023, cold wave in winter 2024), holiday electricity fluctuations (e.g. Spring Festival, National Day), energy policy adjustments (e.g. expansion of inter-provincial green electricity trading in Shanghai power grid in 2024), etc. The adaptability of the model in real environment can be fully verified. For example, in March 2024, due to the surge in green electricity trading, the natural gas consumption for power generation dropped to about 2 million cubic meters. Such extreme cases provide key samples for model robustness testing.

[0096] Time sequence connection with training data: the test period follows the training data (2020-2022) to form a continuous time sequence, which can simulate the actual application scenario of the model "real-time prediction after going online", avoid prediction deviation caused by data gaps, and ensure the authenticity of the evaluation results.

[0097] Deviation rate calculation formula and quantitative significance

[0098] The deviation rate formula is: deviation rate = (predicted natural gas consumption for power generation - actual natural gas consumption for power generation) / actual natural gas consumption for power generation x 100%, where:

[0099] Predicted natural gas consumption for power generation: daily power generation gas consumption prediction value (unit: million cubic meters) output by the model, calculated based on real-time access of gas load, weather, LNG supply, etc.

[0100] Actual natural gas consumption for power generation: daily actual measurement value (unit: million cubic meters) provided by the natural gas pipeline network company, serving as the benchmark true value for deviation calculation.

[0101] The formula quantifies the relative deviation of the predicted value from the actual value, for example: on January 1, 2023, the predicted value is 4.4223 million cubic meters, the actual value is 3.8148 million cubic meters, the deviation rate is 15.92%, reflecting the magnitude of the predicted value higher than the actual value.

[0102] Overall test effect and abnormal situation definition

[0103] The overall performance of the model during the test period is stable: the total predicted natural gas consumption for power generation is 702576 million cubic meters, the actual consumption is 716893 million cubic meters, the overall deviation rate is -2.00%, the average deviation rate is 1.93%, indicating that the model can meet the practical needs under normal circumstances.

[0104] When the absolute value of the deviation rate exceeds 10%, it is defined as "abnormal situation", such situations are mostly caused by sudden factors, for example: in March 2024, due to the surge in green electricity transactions, the single-day deviation rate appeared 117.37% (predicted value much higher than actual value) and -32.58% (predicted value much lower than actual value) extreme fluctuations, such data, although belonging to "noise", but not historically rare, therefore not excluded from the test set, but as a trigger condition for model optimization.

[0105] Iterative optimization mechanism for abnormal situations

[0106] When the absolute value of the deviation rate exceeds 10%, the model is iterated by "adding 3 months of real-time data to the training set", the specific logic is as follows:

[0107] Data supplement range: select complete data (including gas load, weather, LNG supply, hydropower generation, etc.) of the previous 3 months before the abnormal situation occurs, for example, after the abnormal situation in March 2024, supplement the data from January 1, 2023 to March 1, 2024, increase the sample size of the training set by 365 groups, and improve the model's learning ability for extreme scenarios;

[0108] Optimization effect verification: after supplementing the data, the prediction deviation of the model for similar green electricity transaction surges and other scenarios is significantly reduced, and there is no case where the absolute value of the deviation rate exceeds 10% from April to December 2024, indicating that this mechanism can effectively enhance the model's adaptability to sudden factors.

[0109] The natural gas consumption prediction model based on gas-fired power generation based on the method, the model comprises:

[0110] Multi-source data access module: used for accessing multi-source data such as gas load historical data, weather data, economic indicators, LNG supply dynamics, and hydropower generation, supporting daily 09:00 access to the previous day's actual value, 20:00 access to the next day's planned value;

[0111] Feature preprocessing module: perform data cleaning, feature fitting and R 2 Calculate and output 11 key characteristic factors;

[0112] Random forest algorithm core module: build algorithm model based on the parameters, support real-time incremental data stream simulation training;

[0113] Model evaluation module: calculate the MSE and R 2 of the training set and the test set, and output the evaluation report;

[0114] Prediction result output interface: output daily natural gas consumption prediction value and deviation rate, support docking power spot trading system and natural gas pipeline network scheduling platform.

[0115] An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the program to implement the steps of the natural gas consumption prediction method based on gas-fired power generation according to any one of claims 1 to 7.

[0116] A non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executable by a processor to implement the steps of the natural gas consumption prediction method based on gas-fired power generation according to any one of claims 1 to 7.

[0117] Working principle: select gas load historical data, meteorological data, economic indicators, LNG supply dynamics, and power generation of hydropower stations, etc. characteristic factors, collect data of a certain period of time in units of days, construct a multi-dimensional feature data set, and obtain original data; preprocess the collected original data, including data cleaning to remove outliers and missing values, analyze the correlation between each characteristic factor and the target parameter through fitting and other methods, calculate the coefficient of determination R2, sort according to the size of R2, and select the key characteristic factors that have a significant impact on the target parameter; use the random forest algorithm to build a prediction model, divide the data set corresponding to the selected key characteristic factors into a training set and a test set, use the training set to train the model, optimize the model performance by adjusting the model parameters, and simulate real-time incremental data flow at the same time, to obtain a trained model; evaluate the performance of the trained model using the test set, calculate the mean square error (MSE) and the coefficient of determination (R2), etc. If the R2 of the model on the training set and the test set is high and the results are similar; test the model for a long time in the actual scene, calculate the deviation rate of the predicted value and the actual value, analyze the reasons for the large deviation in the test and take corresponding measures, and improve the model iteratively.

[0118] In addition, the text unit creating unit, the image generating unit, the high-quality image training unit and the model optimization processing unit described above are also used to realize other functions of the natural gas consumption prediction method based on gas power generation when they are executed, which will not be described here.

[0119] In addition, the present application also provides a terminal device, the natural gas consumption prediction method based on gas power generation involved in the present embodiment is mainly applied in the terminal device, which can be a PC, a portable computer, a mobile terminal or other devices with display and processing functions.

[0120] Specifically, the terminal device can include a processor (for example, a CPU), a communication bus, a user interface, a network interface and a memory. The communication bus is used to realize the connection and communication between these components; the user interface can include a display screen (Display) and an input unit such as a keyboard (Keyboard); the network interface can optionally include a standard wired interface and a wireless interface (such as a WI-FI interface); the memory can be a high-speed RAM memory or a stable memory (non-volatile memory) such as a disk memory, and the memory can optionally be a storage device independent of the aforementioned processor.

[0121] The memory stores a readable storage medium, and the readable storage medium stores a natural gas consumption prediction program; the processor can call the natural gas consumption prediction program stored in the memory and execute the natural gas consumption prediction method based on gas power generation provided by the present embodiment.

[0122] It can be understood that the readable storage medium can be a tangible device that maintains and stores instructions for use by an instruction execution device. The computer readable storage medium can be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any appropriate combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or a concave and convex structure in a groove, and any appropriate combination of the above. The computer readable storage medium used herein is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (for example, an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0123] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0124] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computing / processing device, partly on the user's computing / processing device, as a stand-alone software package, partly on the user's computing / processing device and partly on a remote computing / processing device or entirely on the remote computing / processing device or server. In the latter scenario, the remote computing / processing device can be connected to the user's computing / processing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing / processing device, for example, through the Internet using an Internet Service Provider. In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0125] Finally, it should be noted that the above-described embodiments are merely possible examples of implementing the present disclosure, and thus are not used to limit the present disclosure. In the above embodiments, although the present disclosure has been described in detail with reference to the foregoing embodiments, the technical solutions recorded in the foregoing embodiments can be modified or replaced by other technical solutions with similar technical features, or some technical features can be modified or replaced by other technical features with similar technical effects, without departing from the spirit and principle of the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for predicting a natural gas consumption based on gas-fired power generation, characterized by, The method comprises the following steps: Selecting gas load historical data, meteorological data, economic indicators, LNG supply dynamics, and hydropower station power generation, etc. characteristic factors, collecting data of a certain time period on a daily basis, constructing a multi-dimensional feature data set, and obtaining original data; The collected original data is pre-processed, including data cleaning to remove outliers and missing values, analyzing the correlation between each characteristic factor and the target parameter by fitting method, calculating the determination coefficient R 2 , according to the size of R 2 , screening out the key characteristic factors which have significant influence on the target parameter; Using a random forest algorithm to construct a prediction model, dividing the data set corresponding to the screened key characteristic factors into a training set and a test set, training the model using the training set, optimizing the model performance by adjusting the model parameters, and simultaneously simulating real-time incremental data flow to obtain a trained model; The trained model is evaluated for performance using the test set, and mean square error (MSE) and coefficient of determination (R 2 ) and other indicators are calculated. If the R 2 of the model on the training set and the test set is high and the results are similar; In the actual scene, the model is tested for a long time, the deviation rate of the predicted value and the actual value is calculated, the reasons for the large deviation in the test are analyzed and corresponding measures are taken, and the model is iteratively improved.

2. The method for predicting the natural gas consumption based on the gas power generation according to claim 1, characterized in that: The gas load historical data includes the actual gas consumption of Shanghai per day, the actual gas supply per day, and the difference between the gas consumption and the gas supply; the LNG supply dynamics includes the daily gas supply of Yangshan LNG, the daily inventory of Yangshan LNG, and the daily gas supply from the West-East Gas Transmission Line 1 to Shanghai; and the hydropower station power generation includes the daily power generation of Xiangjiaba Hydropower Station and the daily power generation of Three Gorges Hydropower Station.

3. The method for predicting the natural gas consumption based on the gas power generation according to claim 1, characterized in that: The meteorological data includes the daily maximum temperature and the daily minimum temperature, and the temperature difference is calculated by the formula: Temperature difference = daily maximum temperature - daily minimum temperature.

4. The method for predicting the natural gas consumption based on the gas power generation according to claim 1, characterized in that: The certain time period is from January 1, 2020 to December 31, 2022, and a total of 1196 groups of data are constructed; the division ratio of the training set and the test set is 8:2, i.e. 80% of the data is used for model training and 20% of the data is used for testing.

5. The method of claim 1, wherein: The parameters of the random forest algorithm include: n_estimators = 55, max_depth = 20, min_samples_split = 3, min_samples_leaf = 1, and max_features ='sqrt'.

6. The method of claim 1, wherein: In the model performance evaluation, the coefficient of determination R 2 The calculation formula is: where y_true is the actual amount of natural gas used for power generation, and y_pred is the model prediction value, is the average of the actual values.

7. The method of claim 1, wherein: The time range of the long-time test is from January 1, 2023 to December 31, 2024, and the deviation rate is calculated by the formula: Deviation rate = (predicted natural gas consumption for power generation - actual natural gas consumption for power generation) / actual natural gas consumption for power generation x 100%; When the absolute value of the deviation rate exceeds 10%, the model is iteratively optimized by adding real-time data of the last 3 months to the training set.

8. A gas consumption prediction model for gas-fired power generation based on the method of any one of claims 1 to 7, characterized in that, The model comprises: A multi-source data access module for accessing gas load historical data, meteorological data, economic indicators, LNG supply dynamics, and hydropower station power generation, etc. multi-source data, supporting the access of actual values of the previous day at 09:00 and the access of planned values of the next day at 20:00; Feature pre-processing module: perform data cleaning, feature fitting and R 2 Compute, output 11 key feature factors; A random forest algorithm core module for constructing an algorithm model based on the parameters and supporting real-time incremental data flow simulation training; Model evaluation module: Calculate the MSE and R of the training set and the test set 2 , output evaluation report; A prediction result output interface for outputting daily predicted natural gas consumption for power generation and deviation rate, supporting the connection with the power spot trading system and the natural gas pipeline network dispatching platform.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the natural gas consumption prediction method based on gas power generation according to any one of claims 1 to 7 when executing the program.

10. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, which is executed by a processor, implements the steps of a method for predicting the consumption of natural gas based on gas-fired electricity generation according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Natural gas load prediction method based on LSTM recurrent neural network

    CN110852496A

  • Photovoltaic power generation capacity short-term prediction method based on optimized random forest

    CN118735055A

  • Urban natural gas consumption prediction method based on XGB-PSO-LSTM model

    CN119671328A