Method and system for predicting yield of oil and gas well

By combining the models of random forests and long and short-term memory neural networks in oil and gas well yield prediction, key features are screened and non-stationary characteristics are captured, problems that are difficult to predict by traditional algorithms are solved, and higher prediction accuracy and adaptability are achieved.

CN120197098AInactive Publication Date: 2025-06-24SHAANXI DIYUAN YUNCHUANG TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510318587.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Oil and gas well output forecast faces complex influencing factors and uncertainties, traditional machine learning algorithms are difficult to effectively fit and predict, and the model is insufficiently interpretable and adaptable.

Method used

A random forest algorithm is used to build a preliminary yield prediction model, and key features are screened through feature importance analysis. If the prediction error exceeds the threshold, a long and short-term memory neural network is introduced to capture non-stationary characteristics, optimize the prediction results, and dynamically update the model through time series cross-validation and incremental training mechanism.

Benefits of technology

It improves the accuracy and adaptability of oil and gas well production forecasts, can effectively deal with complex geological conditions and production characteristics, and provides a reliable basis for oil and gas field production decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197098A_ABST
    Figure CN120197098A_ABST
Patent Text Reader

Abstract

The invention discloses an oil and gas well yield prediction method and system, and the method comprises the steps: obtaining geological feature data, development history data and process parameter data in an oil and gas field database, employing an interpolation method to process missing values for the data, employing an outlier detection algorithm to process abnormal values, and obtaining a cleaned data set; constructing a yield prediction model by adopting a random forest algorithm, and screening characteristic variables which have obvious influence on the yield through characteristic importance analysis to obtain a preliminary prediction result; and dividing the cleaned data set into a plurality of subsets according to a time sequence by adopting a time sequence cross validation method, training and verifying the yield prediction model in sequence, and evaluating the prediction performance of the yield prediction model in different time periods to obtain a dynamic verification result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to a method and system for predicting oil and gas well production. Background Art

[0002] Predicting oil and gas well production is a complex process involving multiple influencing factors and uncertainties. When establishing a machine learning model, it is first necessary to determine appropriate input features. Although the geological feature data and development history data of oil and gas fields can serve as important references, the quality and integrity of these data may be problematic. For example, there may be errors in the collection and interpretation of geological data, and historical data may be missing or inconsistent. In addition, other factors that may affect production, such as the technological parameters of oil and gas wells and production management measures, also need to be considered. Selecting appropriate features and performing necessary data preprocessing and feature engineering are the keys to establishing a high-quality machine learning model.

[0003] When selecting machine learning algorithms, although algorithms such as support vector machines and artificial neural networks have achieved good results in many fields, their applicability in predicting oil and gas well production still needs to be further verified. The problem of predicting oil and gas well production may have characteristics such as non-linearity, non-stationarity, and multi-scale, and traditional machine learning algorithms may be difficult to fit and predict well. In addition, the interpretability of the model is also an issue that needs attention because decision-makers need to understand the key factors affecting production and their action mechanisms. In practical applications, it may be necessary to explore multiple machine learning algorithms and conduct systematic comparisons and optimizations to obtain the best prediction performance.

[0004] Model validation and optimization are important links in the development of machine learning models. Although common validation methods such as cross-validation and holdout method can evaluate the performance of the model, they may not fully consider the dynamic change characteristics of the oil and gas well production process. Over time, the geological conditions, production status, etc. of oil and gas wells may change, resulting in differences in the distribution of historical data and future data. Therefore, it is necessary to explore more suitable model validation methods, such as time series cross-validation, etc., to better evaluate the performance of the model in practical applications. At the same time, a perfect model update mechanism also needs to be established to regularly retrain and optimize the model using new production data to adapt to the changes in the production conditions of oil and gas wells. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for predicting oil and gas well production.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A method for predicting oil and gas well production, comprising:

[0008] Obtain geological feature data, development history data, and process parameter data in the oil and gas field database;

[0009] Extract feature variables related to production according to the geological feature data, development history data, and process parameter data in the oil and gas field database;

[0010] Construct a production prediction model using the random forest algorithm, and through feature importance analysis, screen out the feature variables that have a significant impact on production to obtain a preliminary prediction result;

[0011] If the error of the preliminary prediction result exceeds the preset threshold, use a long short-term memory neural network to model the time series data, take the output result of the random forest model and the original features as inputs, capture the non-stationary characteristics and multi-scale change trends of production, and obtain an optimized prediction result;

[0012] Evaluate the prediction performance of the production prediction model in different time periods according to the optimized prediction result to obtain a dynamic verification result;

[0013] If the dynamic verification result shows that the performance of the production prediction model has declined, obtain new samples from the latest production data, perform incremental training on the production prediction model, update the parameters of the production prediction model, and obtain a prediction model adapted to the latest production situation.

[0014] Preferably, interpolation method is used to process missing values in the geological feature data, development history data, and process parameter data in the oil and gas field database, and an outlier detection algorithm is used to process outliers to obtain a cleaned data set.

[0015] Preferably, according to the cleaned data set, extract feature variables related to production. The feature variables include formation pressure, permeability, wellbore pressure, and production time, and perform dimensionality reduction on the feature variables through principal component analysis to obtain a low-dimensional representation of the feature variables.

[0016] Preferably, use the time series cross-validation method to divide the cleaned data set into multiple subsets in chronological order, train and verify the production prediction model in turn, evaluate the prediction performance of the production prediction model in different time periods, and obtain a dynamic verification result.

[0017] The present invention also provides a prediction system for the production of oil and gas wells, including:

[0018] A data acquisition and cleaning module for obtaining geological feature data, development history data, and process parameter data in the oil and gas field database;

[0019] A feature extraction and dimensionality reduction module for extracting feature variables related to production according to the geological feature data, development history data, and process parameter data in the oil and gas field database;

[0020] The production prediction model construction module is used to construct a production prediction model using the random forest algorithm, screen out the feature variables that have a significant impact on production through feature importance analysis, and obtain a preliminary prediction result;

[0021] The model optimization module is used to, if the error of the preliminary prediction result exceeds a preset threshold, model the time series data using a long short-term memory neural network, take the output result of the random forest model and the original features as inputs, capture the non-stationary characteristics and multi-scale change trends of production, and obtain an optimized prediction result;

[0022] The dynamic verification module is used to evaluate the prediction performance of the production prediction model in different time periods according to the optimized prediction result, and obtain a dynamic verification result;

[0023] The model update module is used to, if the dynamic verification result shows that the performance of the production prediction model has declined, obtain new samples from the latest production data, perform incremental training on the production prediction model, update the parameters of the production prediction model, and obtain a prediction model adapted to the latest production situation.

[0024] Preferably, the data acquisition and cleaning module uses the interpolation method to process missing values in the geological feature data, development history data, and process parameter data in the oil and gas field database, and uses an outlier detection algorithm to process outliers, obtaining a cleaned data set.

[0025] Preferably, the feature extraction and dimensionality reduction module extracts the feature variables related to production according to the cleaned data set. The feature variables include formation pressure, permeability, wellbore pressure, and production time, and performs dimensionality reduction on the feature variables through the principal component analysis method to obtain a low-dimensional representation of the feature variables.

[0026] Preferably, the dynamic verification module uses the time series cross-validation method to divide the cleaned data set into multiple subsets in chronological order, train and verify the production prediction model in turn, evaluate the prediction performance of the production prediction model in different time periods, and obtain a dynamic verification result.

[0027] The present invention first obtains geological features, development history, and process parameter data from a database, and performs data cleaning through interpolation and outlier detection algorithms. Then, it extracts feature variables related to production, such as formation pressure, permeability, etc., and uses principal component analysis for dimensionality reduction. Next, a random forest algorithm is used to construct a preliminary prediction model, and key variables are screened through feature importance analysis. If the prediction error is too large, a long short-term memory neural network is introduced to capture non-stationary characteristics and optimize the prediction results. The present invention also uses time series cross-validation to evaluate the model performance and maintains the adaptability of the model to the latest production status through incremental training. This method combines multiple algorithms, can effectively handle the complex geological conditions and production characteristics of oil and gas fields, improve the accuracy and adaptability of production prediction, and provide a reliable basis for oil and gas field production decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of the prediction method for the production of oil and gas wells of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] To further understand the content of the present invention, the present invention will be described in detail with reference to the drawings and embodiments. The following further describes the present application with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. Additionally, it should be noted that for the sake of description, only parts related to the invention are shown in the drawings.

[0030] Embodiment 1:

[0031] As Figure 1 shown, the embodiment of the present invention provides a prediction method for the production of oil and gas wells, including:

[0032] S101. Obtain geological feature data, development history data, and process parameter data in the oil and gas field database, process missing values for the data using interpolation, and process outliers using an outlier detection algorithm to obtain a cleaned data set.

[0033] Obtain geological feature data, development history data, and process parameter data in the oil and gas field database. Preprocess the obtained original data, and use the least squares method to interpolate and estimate the missing values to obtain a completed dataset. For the completed dataset, use a distance-based outlier detection algorithm to calculate the distance between each data point and other data points. If the distance exceeds a preset threshold, it is determined as an outlier and marked as an abnormal value. According to the outlier detection results, correct the abnormal values. Use the median substitution method to replace the abnormal values with the median of the corresponding dimension to obtain a processed dataset. For the processed dataset, use the Z-score standardization method for data normalization to map data with different dimensions to the same scale and eliminate the influence of dimensions to obtain a standardized dataset. According to the standardized dataset, use the Pearson correlation coefficient method to calculate the correlation between different attributes. If the correlation coefficient exceeds a preset threshold, it is determined that there is a strong correlation between the two attributes. According to the correlation analysis results between attributes, construct an oil and gas field data feature association graph, with attributes as nodes and correlations as edges, to visually display the association relationships between various attributes. Based on the association graph, use the decision tree algorithm to establish an oil and gas field data classification model, comprehensively consider multi-dimensional attributes such as geological features, development history, and process parameters, classify and predict the oil and gas field data, and provide data support for subsequent oil and gas field development decisions.

[0034] Furthermore, the oil and gas field database contains rich information on geological characteristics, development history, and process parameters. Taking a certain oilfield as an example, the geological characteristic data may include reservoir depth, porosity, permeability, etc.; the development history data includes cumulative oil production, water cut change, etc.; the process parameter data includes injection pressure, oil production method, etc. These original data may have missing values, such as the missing permeability data of a certain well. Using the least squares method for interpolation estimation, the best curve can be fitted based on the known data points to calculate the missing permeability value. After data completion, outlier detection is required to identify outliers. For example, the water cut data in a certain block is generally between 60% - 80%, but individual data points suddenly show an abnormally high value of 95%. Using a distance-based algorithm, calculate the Euclidean distance between each data point and other points. If the distance of a certain point is significantly greater than other points, it may be an outlier. Replacing the detected outliers with the median of this dimension can effectively reduce the impact of outliers on subsequent analysis. It is very necessary to perform standardization processing after data cleaning. Taking reservoir depth and porosity as an example, the former has the unit of meters and the value may be between 1000 - 3000; the latter is a percentage and is usually between 5% - 30%. The numerical differences between these two indicators are huge, which is not conducive to subsequent modeling analysis. Using the Z-score standardization method, map the data with different dimensions to a distribution with a mean of 0 and a standard deviation of 1 to eliminate the influence of dimensions and make different indicators comparable. The standardized dataset can be used to analyze the correlation between different attributes. For example, calculate the Pearson correlation coefficient between reservoir depth and porosity. If the coefficient is close to -1, it indicates a strong negative correlation between the two, that is, as the depth increases, the porosity gradually decreases. This correlation analysis helps to understand the geological characteristics of the oil and gas reservoir. Based on the results of correlation analysis, an association map of oil and gas field data characteristics can be constructed. In the map, attributes such as reservoir depth and porosity are used as nodes, and the correlation coefficient is used as the edge connecting the nodes to visually display the association strength between each attribute. This visualization method helps geological engineers quickly grasp the overall characteristics of the oil and gas reservoir. Finally, using the decision tree algorithm to establish a classification model can comprehensively consider multi-dimensional attributes to classify and predict oil and gas fields. For example, according to reservoir characteristics, development history, and process parameters, predict whether a certain block is suitable for water flooding development. The advantage of the decision tree model lies in its strong interpretability, which can clearly show the influence degree of each factor on the development decision and provide data support for the formulation of oil and gas field development plans. Through the above series of data processing and analysis steps, valuable information can be extracted from the original oil and gas field data, providing a scientific basis for oil and gas field exploration and development decisions, improving development efficiency and economic benefits. This data-driven method is gradually becoming an important development trend in the oil and gas industry.

[0035] S102. Extract the feature variables related to production from the cleaned data set. The feature variables include formation pressure, permeability, wellbore pressure, and production time. Perform dimensionality reduction on the feature variables through the principal component analysis method to obtain the low-dimensional representation of the feature variables.

[0036] Based on the cleaned data set, obtain the feature variables related to production, including formation pressure, permeability, wellbore pressure, production time, etc. For the obtained feature variables, perform dimensionality reduction processing using the principal component analysis method. By calculating the covariance matrix of the feature variables, obtain the eigenvalues and eigenvectors of the principal components. According to the eigenvalue magnitudes, select the eigenvectors corresponding to the top k largest eigenvalues as the principal components to obtain the low-dimensional representation form of the feature variables. Use the projection coefficients of the feature variables on the principal components as new features to replace the original high-dimensional feature variables, realizing the dimensionality reduction and compression of the data. Based on the compressed low-dimensional feature data, construct a production prediction model and use machine learning algorithms such as support vector machines or random forests for training and optimization. Apply the trained production prediction model to new well data. According to the input feature variables such as formation pressure, permeability, wellbore pressure, and production time, predict the production level of the well. Evaluate and analyze the prediction results, calculate the prediction error and determination coefficient of the model, judge the prediction performance of the model, and optimize and adjust the model as needed to improve the accuracy of production prediction.

[0037] Furthermore, in oil and gas field development, production prediction is a crucial task. To improve the prediction accuracy, it is first necessary to obtain the characteristic variables related to production. Taking a certain oilfield as an example, formation pressure, permeability, wellbore pressure, and production time are selected as the main characteristics. These characteristics reflect the geological conditions and production status of the reservoir and have a direct impact on production. However, high-dimensional characteristics may contain redundant information, affecting the model performance. Therefore, the principal component analysis method is used for dimensionality reduction. By calculating the covariance matrix of the characteristic variables, the eigenvalues and eigenvectors of the principal components are obtained. Assuming the calculation results show that the cumulative contribution rate of the first three principal components reaches 95%, these three principal components can be selected as the new characteristics. Projecting the original characteristics into the principal component space, the compressed low-dimensional characteristic data is obtained. This dimensionality reduction not only reduces the data volume but also retains the main information of the original data, which is beneficial to improving the efficiency and accuracy of subsequent modeling. Next, the compressed characteristic data is used to construct a production prediction model. Support vector machine (SVM) and random forest are two commonly used machine learning algorithms. SVM is suitable for dealing with non-linear relationships, while random forest is good at handling complex feature interactions. In this example, the random forest algorithm is selected because it can handle various types of features and is insensitive to outliers. The dataset is divided into a training set and a test set at a ratio of 8:2. The random forest model is constructed using the training set, and the model performance is optimized by adjusting hyperparameters such as the number of trees and the maximum depth. Assuming that after optimization, a random forest model with 100 decision trees and a maximum depth of 10 is obtained. Applying the optimized model to the test set can evaluate the prediction performance of the model. Assuming that the root mean square error (RMSE) of the model on the test set is 100 barrels per day and the coefficient of determination (R 2 ) is 0.85. This indicates that the model has good prediction ability but there is still room for improvement. To further improve the model performance, more relevant features such as reservoir temperature and water cut can be considered. At the same time, other algorithms such as gradient boosting trees can be tried, or the prediction results of multiple models can be integrated. Through continuous optimization and adjustment, a more accurate and reliable production prediction model can be finally obtained, providing strong support for oil and gas field development decisions.

[0038] S103. Construct a production prediction model using the random forest algorithm, and through feature importance analysis, screen the characteristic variables that have a significant impact on production to obtain the preliminary prediction results.

[0039] Obtain historical production data and related characteristic variable data, and preprocess the data, including operations such as missing value handling, outlier handling, and data standardization, to obtain a preprocessed dataset. Use the random forest algorithm to construct a production prediction model, optimize the model hyperparameters through methods such as grid search, and train to obtain the optimal random forest model. Use feature importance analysis methods, such as metrics like mean impurity reduction or mean precision reduction, to calculate the importance scores of each characteristic variable, sort according to the importance scores, and screen out the key characteristic variables that have a significant impact on production. Based on the screened key characteristic variables, retrain the random forest model to obtain a simplified production prediction model. Apply the simplified random forest model to the test set data to predict the production in a future period of time and obtain preliminary prediction results. Use methods such as cross-validation to evaluate the prediction performance of the simplified random forest model, measure the prediction accuracy of the model through metrics such as mean squared error and mean absolute error, and determine whether the model meets the requirements of practical applications. If the prediction performance of the model is not ideal enough, consider introducing other machine learning algorithms such as support vector machines or neural networks, and further improve the accuracy of production prediction through methods such as ensemble learning to finally obtain reliable production prediction results.

[0040] Furthermore, obtaining historical production data and related characteristic variable data is the basis for constructing a prediction model. For example, for the production prediction of a certain oilfield, monthly production data for the past 5 years, as well as corresponding characteristic variables such as formation pressure, permeability, and wellbore pressure, can be collected. The data preprocessing step is crucial. Missing values can be processed by interpolation methods. For example, if production data is missing for a certain month, it can be filled with the average value of the previous and next months. Outlier handling can use the boxplot method. Data that exceeds 1.5 times the interquartile range is regarded as an outlier and can be removed or corrected. Data standardization can use the min-max normalization method to unify each characteristic variable into the 0-1 interval and eliminate the influence of dimensions. The random forest algorithm is an ensemble learning method that classifies or regresses by constructing multiple decision trees and taking the majority vote. In the production prediction model, hyperparameters such as the number of trees, the maximum depth of each tree, and the minimum number of samples per node can be set. Grid search is a commonly used hyperparameter optimization method. For example, the number of trees can be set from 100 to 500 with a step size of 100; the maximum depth can be set from 5 to 15 with a step size of 2; the minimum number of samples can be set from 2 to 10 with a step size of 2. By traversing all parameter combinations, the parameter combination with the best performance on the validation set is selected as the optimal hyperparameter. Feature importance analysis helps identify the factors that have the most significant impact on production. The mean decrease in impurity method measures importance based on the degree of impurity reduction caused by the split of a feature in a decision tree. For example, if the mean decrease in impurity value of formation pressure in the decision tree is the highest, it indicates that it makes the greatest contribution to production prediction. The mean decrease in accuracy method evaluates importance by randomly shuffling the values of a feature and observing the degree of decrease in model accuracy. If the model accuracy drops the most after shuffling the wellbore pressure, it means that the wellbore pressure is a key characteristic variable. Retraining the random forest model based on the selected key characteristic variables can simplify the model structure and improve computational efficiency. For example, if it is found that formation pressure, permeability, and production time are the three most important features, a new random forest model can be constructed using only these three variables. This simplification can not only reduce the model complexity but also potentially improve the generalization ability and reduce the risk of overfitting. Applying the simplified random forest model to the test set data can predict the production for a future period. For example, using the data from the past 5 years to train the model and then predicting the monthly production for the next 6 months. This rolling prediction method can timely capture the production change trend and provide support for production decisions. Cross-validation is an effective method for evaluating model performance. 5-fold cross-validation can be used, dividing the dataset into 5 parts and taking turns using 4 parts as the training set and 1 part as the validation set. By calculating metrics such as the mean squared error and mean absolute error, the performance of the model on different data subsets can be comprehensively evaluated. For example, if the mean absolute error of the model in cross-validation is lower than 5% of the historical production, the model can be considered to meet the requirements of practical applications. If the performance of the random forest model is not good, other algorithms can be considered. Support vector machines are suitable for handling nonlinear relationships and can capture the complex associations between production and characteristic variables.Neural networks are good at processing large-scale data and can discover potential patterns in massive production data. Through ensemble learning methods, such as voting or stacking, the advantages of multiple models can be combined to further improve prediction accuracy. For example, the prediction results of random forests, support vector machines, and neural networks can be weighted averaged to obtain the final output prediction value, thereby achieving more reliable output prediction.

[0041] The simplified random forest model was applied to the test set data to predict the output in the future and obtain preliminary prediction results.

[0042] Obtain the simplified random forest model and test set data, and apply the model to the test set data. According to the model prediction results, obtain the preliminary forecast value of the output in the future period. Use the time series analysis method to further process the output forecast value to eliminate the influence of outliers. Through seasonal decomposition and trend analysis, judge the change trend and periodic characteristics of the output forecast value. If the output forecast value shows a clear upward or downward trend, determine the rate of change of future output based on the trend slope. If the output forecast value shows a significant seasonal cycle, use the periodic factor to correct the output in each future time period. Comprehensively consider the trend and periodic characteristics of the output, obtain the output forecast results for each future time period, and provide a basis for the formulation of production plans.

[0043] S104. If the error of the preliminary prediction result exceeds a preset threshold, a long short-term memory neural network is used to model the time series data, and the output result and original features of the random forest model are used as input to capture the non-stationary characteristics and multi-scale change trends of the output, so as to obtain an optimized prediction result.

[0044] Obtain the time series data related to the yield, preprocess the data, and extract the original features that reflect the yield change trend. Use the random forest model to make a preliminary prediction on the preprocessed data to obtain the preliminary prediction result. Determine whether the error of the preliminary prediction result exceeds the preset threshold. If it exceeds the preset threshold, execute the next step; otherwise, output the preliminary prediction result as the final prediction result. In view of the non-stationary characteristics and multi-scale change trends of the yield, construct a long short-term memory neural network model, and use the output results and original features of the random forest model as input. Model the time series data through the long short-term memory neural network model to capture the non-stationary characteristics and multi-scale change trends of the yield and obtain the optimized prediction result. According to the optimized prediction results and the preliminary prediction results, determine the final yield prediction value, and output the final prediction result. Evaluate and analyze the prediction results. If the prediction effect does not meet expectations, return to the fourth step, adjust the parameters of the long short-term memory neural network model, re-model and predict until a satisfactory prediction effect is obtained.

[0045] Model the time series data through a long short-term memory neural network model to capture the non-stationary characteristics and multi-scale change trends of production, and obtain optimized prediction results.

[0046] Obtain the production time series data within a certain time range, sort it according to the timestamps of the production data to form a time series data set. Preprocess the time series data, including missing value processing, outlier processing, and data normalization, to obtain the preprocessed time series data. According to the preprocessed time series data, use the sliding window method to construct training samples and labels. Each training sample contains historical data with a certain time step, and the corresponding label is the production value at the next moment. Construct a long short-term memory neural network model, set the input layer, several LSTM hidden layers, and a fully connected output layer, and select appropriate loss functions and optimization algorithms. Input the constructed training samples and labels into the long short-term memory neural network model for training, update the model parameters through the backpropagation algorithm to obtain a trained prediction model. Use the trained prediction model to predict new time series data, and obtain the production prediction values for a future period through the forward propagation process of the model. Post-process the prediction results, including data denormalization and outlier processing, to obtain the final production prediction results for guiding decisions such as production planning and resource scheduling.

[0047] Furthermore, obtaining production time series data is the basis of the prediction model. For example, the daily production data of a steel plant, including the date and the corresponding steel production. Data preprocessing is crucial for the model's performance. For missing value handling, interpolation methods can be used. For example, if data is missing on a certain day, it can be filled with the average of the previous and next two days. For outlier handling, the box plot method can be used, considering data beyond 1.5 times the interquartile range as outliers and replacing them with nearby values. For normalization, the maximum-minimum method can be adopted to map the data to the 0-1 interval, which helps improve the model's convergence speed. Constructing samples using a sliding window is a common method for time series prediction. Assuming a prediction of the production in the next 7 days, the window size can be set to 30 days. Each sample contains 30 days of historical data, and the corresponding label is the production on the 31st day. This can make full use of historical information and capture the long-term dependencies of the time series. Long Short-Term Memory Neural Network (LSTM) is suitable for processing time series data. The model structure can include an input layer, two LSTM hidden layers, and a fully connected output layer. The LSTM layer can learn long-term dependencies and effectively capture the production change trend. The mean squared error (MSE) can be selected as the loss function, and the Adam optimization algorithm can be used to balance the convergence speed and stability. During the model training process, the batch gradient descent method can be adopted. Each batch contains 64 samples, and the model parameters are optimized through multiple rounds of iteration. To prevent overfitting, the dropout technique can be used by adding a dropout layer after the LSTM layer with a dropout rate set to 0.2. At the same time, an early stopping strategy can be adopted, stopping the training when the validation set loss does not decrease for 5 consecutive rounds to avoid overfitting the training data. In the prediction stage, the data of the most recent 30 days is input into the model to obtain the predicted values of the production in the next 7 days. In the post-processing step, first, inverse normalization is performed to map the prediction results back to the original data range. Then, outlier handling is carried out. For example, if the predicted value is negative, it can be adjusted to 0 or the historical minimum value. The final prediction results can be used for production plan formulation. For example, if the prediction shows an upward trend in production in the next week, the enterprise can correspondingly increase the raw material procurement volume and adjust the personnel schedule to ensure that the production capacity meets the demand. If the prediction shows large fluctuations in production, the inventory strategy can be adjusted in advance to balance the supply and demand relationship. Through production prediction, the enterprise can optimize resource allocation, improve production efficiency, and reduce operating costs. This prediction method comprehensively considers the data characteristics and model advantages. Data preprocessing ensures the input quality, constructing samples using a sliding window makes full use of historical information, and the LSTM model can effectively capture the long-term and short-term dependencies of the time series. Through model training, prediction, and post-processing, reliable production prediction results are finally obtained, providing strong support for enterprise decision-making.

[0048] S105. Adopt the time series cross-validation method to divide the cleaned data set into multiple subsets in chronological order, train and validate the production prediction model in turn, evaluate the prediction performance of the production prediction model in different time periods, and obtain the dynamic verification result.

[0049] According to the time series characteristics, the sliding time window method is adopted to divide the cleaned historical data set into a training set and a validation set in chronological order. The training set is used to train the yield prediction model, and the validation set is used to evaluate the model performance. For each subset of the divided training set, a yield prediction model is constructed. Machine learning algorithms such as linear regression, support vector machine, and neural network can be selected, and the model hyperparameters are optimized through methods such as grid search to train the optimal model. The trained yield prediction model is used to predict the validation set subset, and the prediction performance of the model in this time period is evaluated through evaluation indicators such as mean square error and mean absolute error. The sliding time window method is adopted to continuously update the training set and the validation set, and steps 2-3 are repeated to obtain the prediction performance evaluation results of the model in different time periods. According to the prediction performance of the model in each time period, a trend chart of the performance changing with time is drawn to analyze the dynamic generalization ability and stability of the model. If the prediction performance of the model is poor in some time periods, the influencing factors of the yield are analyzed according to the data characteristics of this time period, additional features are introduced, the data preprocessing method is optimized, and the model is retrained. The overall effect of the yield prediction model is evaluated by synthesizing the prediction performance of each time period, and the model with the best and stable performance is selected as the final production application model to guide the yield prediction and production plan formulation in the future for a period of time.

[0050] Furthermore, the processing of time - series data is crucial for yield prediction. The sliding - time - window method can effectively capture the dynamic characteristics of data. By continuously updating the training set and validation set, the performance of the model at different time periods can be evaluated. For example, for a manufacturing enterprise, a monthly - based sliding window can be selected. Each time, the data of the past 12 months is used as the training set, and the next 3 months are used as the validation set. This can reflect seasonal variations and short - term trends. When constructing a yield - prediction model, choosing the appropriate algorithm is crucial. Linear regression is suitable for simple linear relationships, while support vector machines can handle non - linear relationships. Neural networks perform well in dealing with complex patterns. For example, for crop - yield prediction affected by multiple factors, a multi - layer perceptron model can be tried. The input layer includes features such as temperature, precipitation, and soil conditions, and the hidden layer captures the complex interactions of these factors. Optimizing the model hyperparameters is crucial for improving prediction performance. Grid search is a commonly used method, which finds the optimal solution by exhaustively searching all possible parameter combinations. For example, for a support vector machine model, grid search can be performed on the kernel function type (such as linear, polynomial, radial basis) and the regularization parameter C. This helps to find a balance between model complexity and generalization ability. When evaluating model performance, different metrics reflect different aspects of the prediction. The mean squared error is more sensitive to large errors and is suitable for evaluating the model's ability to handle outliers. The mean absolute error gives the average magnitude of the prediction error and is easier to interpret. For yield prediction, the mean percentage error can also be considered, which intuitively reflects the relative accuracy of the prediction. By plotting the trend of model performance over time, the dynamic characteristics of the model can be analyzed in depth. For example, if it is found that the model performs poorly in the third quarter of each year, this may imply the existence of uncaught seasonal factors. At this time, seasonal indicators can be considered or seasonal adjustment techniques can be adopted to improve the model. For the time periods with poor performance, it is crucial to deeply analyze the factors affecting yield. For example, for an electronics manufacturer, it may be found that the traditional model's prediction effect is poor before and after the release of new products. At this time, the product - life - cycle stage can be introduced as a new feature, or separate sub - models can be constructed for different product types to improve prediction accuracy. The finally selected production - application model should perform stably in all time periods. This not only ensures the reliability of the prediction but also provides a solid foundation for production - plan formulation. For example, a model with a quarterly prediction error of less than 5% in the past two years is more suitable for guiding long - term production decisions compared to a model with a smaller overall error but larger fluctuations.

[0051] S106. If the dynamic verification result shows that the performance of the yield - prediction model has declined, obtain new samples from the latest production data, perform incremental training on the yield - prediction model, update the parameters of the yield - prediction model, and obtain a prediction model adapted to the latest production situation.

[0052] Obtain the prediction results of the production prediction model on the latest production data, compare the prediction results with the actual production, calculate the prediction error, and obtain the dynamic verification results of the model. If the dynamic verification results show a decline in the performance of the production prediction model, randomly extract a certain number of samples from the latest production data as new training data. Combine the newly extracted sample data with the original training data, and preprocess the combined data, including data cleaning, feature selection, and feature engineering operations, to obtain a dataset suitable for model training. Adopt incremental learning algorithms, such as online learning or transfer learning, and use the new training data to perform incremental training on the existing production prediction model, update the parameters and weights of the model, and make the model adapt to the latest production situation. During the incremental training process, evaluate the performance of the model through methods such as cross-validation, select the optimal model parameters and hyperparameters, avoid overfitting or underfitting problems, and improve the generalization ability of the model. Apply the incrementally trained production prediction model to the latest production data, obtain the prediction results, compare them with the actual production, calculate the prediction error, and evaluate whether the prediction performance of the model has been improved. If the performance of the incrementally trained model meets the requirements, use it as the new production prediction model for subsequent production prediction tasks; otherwise, continue to iterate steps 2-6 until a prediction model with satisfactory performance is obtained.

[0053] Furthermore, dynamic verification of the output prediction model is a key step to ensure the continued effectiveness of the model. By comparing the prediction results of the latest production data with the actual output, the prediction error can be calculated to evaluate the current performance of the model. For example, a steel plant uses a machine learning model to predict daily output. The prediction error in the last week suddenly increased from an average of 3% to 7%, indicating that the model performance may have declined. In order to cope with the decline in model performance, new training samples need to be extracted from the latest production data. This approach enables the model to adapt to changes in the production environment. For example, the above steel plant can randomly extract 500 data from the production records of the last month as new training samples, which contain the latest production trends and features. After merging the new and old data, comprehensive preprocessing is required. This includes removing outliers, processing missing data, selecting key features, etc. In the example of the steel plant, it may be necessary to eliminate abnormal output records caused by equipment failure and introduce new features such as raw material price volatility index to improve the prediction accuracy of the model. The application of incremental learning algorithms enables the model to effectively absorb the information of new data. Online learning algorithms such as stochastic gradient descent (SGD) can process new data one by one and continuously update model parameters. Transfer learning, on the other hand, can quickly adapt to new production environments by leveraging models trained on other similar tasks. For example, a steel plant can use a model trained on a similar-sized plant as a starting point and fine-tune it to suit its own production characteristics. In the incremental training process, cross-validation is an effective tool for evaluating model performance. By dividing the dataset into multiple subsets and taking turns as validation sets, the generalization ability of the model can be fully evaluated. For the case of the steel plant, time series cross-validation can be used, using data from the last few months as the validation set to ensure that the model can accurately predict future production. After incremental training, the model performance needs to be verified again on the latest data. If the steel plant's model reduces the prediction error to less than 4% in the new week, the incremental training can be considered successful. This process may require multiple iterations until the model performance reaches the expected level. Through this dynamic update and verification method, the output forecasting model can continuously adapt to changes in the production environment and provide accurate decision support for enterprises. This not only improves production efficiency, but also helps enterprises better cope with market fluctuations and optimize resource allocation.

[0054] Embodiment 2:

[0055] The embodiment of the present invention provides a system for predicting oil and gas well production, including:

[0056] The data acquisition and cleaning module is used to obtain geological feature data, development history data and process parameter data in the oil and gas field database, use interpolation method to process missing values, use outlier detection algorithm to process abnormal values, and obtain a cleaned data set;

[0057] A feature extraction and dimensionality reduction module, which is used to extract feature variables related to production according to the cleaned data set. The feature variables include formation pressure, permeability, wellbore pressure, and production time, and perform dimensionality reduction on the feature variables through the principal component analysis method to obtain a low-dimensional representation of the feature variables;

[0058] A production prediction model construction module, which is used to construct a production prediction model using the random forest algorithm, screen out the feature variables that have a significant impact on production through feature importance analysis, and obtain a preliminary prediction result;

[0059] A model optimization module, which is used to, if the error of the preliminary prediction result exceeds a preset threshold, use a long short-term memory neural network to model time series data, take the output result of the random forest model and the original features as inputs, capture the non-stationary characteristics and multi-scale change trends of production, and obtain an optimized prediction result;

[0060] A dynamic verification module, which is used to adopt the time series cross-validation method, divide the cleaned data set into multiple subsets in chronological order, train and verify the production prediction model in turn, evaluate the prediction performance of the production prediction model in different time periods, and obtain a dynamic verification result;

[0061] A model update module, which is used to, if the dynamic verification result shows that the performance of the production prediction model has declined, obtain new samples from the latest production data, perform incremental training on the production prediction model, update the parameters of the production prediction model, and obtain a prediction model adapted to the latest production situation.

[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting oil and gas well production, characterized in that: include: Obtain geological characteristic data, development history data and process parameter data from the oil and gas field database; Extract characteristic variables related to production based on geological characteristic data, development history data and process parameter data in the oil and gas field database; The random forest algorithm was used to build a yield prediction model. The feature importance analysis was performed to screen the feature variables that had a significant impact on yield and obtain preliminary prediction results. If the error of the preliminary prediction result exceeds a preset threshold, a long short-term memory neural network is used to model the time series data, and the output result and original features of the random forest model are used as input to capture the non-stationary characteristics and multi-scale change trends of the output, so as to obtain an optimized prediction result; Evaluate the prediction performance of the production prediction model in different time periods according to the optimized prediction results to obtain dynamic verification results; If the dynamic verification result shows that the performance of the production prediction model has declined, new samples are obtained from the latest production data, incremental training is performed on the production prediction model, and the parameters of the production prediction model are updated to obtain a prediction model that adapts to the latest production conditions.

2. The method for predicting oil and gas well production according to claim 1, characterized in that: The missing values ​​of geological characteristic data, development history data and process parameter data in the oil and gas field database are processed by interpolation method, and the outlier detection algorithm is used to process abnormal values ​​to obtain the cleaned data set.

3. The method for predicting oil and gas well production according to claim 2, characterized in that: According to the cleaned data set, characteristic variables related to production are extracted, and the characteristic variables include formation pressure, permeability, wellbore pressure and production time. The characteristic variables are reduced in dimension by principal component analysis to obtain a low-dimensional representation of the characteristic variables.

4. The method for predicting oil and gas well production according to claim 3, characterized in that: The cleaned data set is divided into multiple subsets in chronological order by using a time series cross-validation method, the yield prediction model is trained and validated in turn, the prediction performance of the yield prediction model in different time periods is evaluated, and a dynamic validation result is obtained.

5. A prediction system for oil and gas well production, characterized in that: include: Data acquisition and cleaning module, used to obtain geological feature data, development history data and process parameter data from oil and gas field database; Feature extraction and dimension reduction module, used to extract characteristic variables related to production based on geological feature data, development history data and process parameter data in the oil and gas field database; The yield prediction model building module is used to build a yield prediction model using the random forest algorithm, screen the characteristic variables that have a significant impact on the yield through feature importance analysis, and obtain preliminary prediction results; A model optimization module is used to use a long short-term memory neural network to model the time series data if the error of the preliminary prediction result exceeds a preset threshold, and use the output result and original features of the random forest model as input to capture the non-stationary characteristics and multi-scale change trends of the output to obtain an optimized prediction result; A dynamic verification module, used to evaluate the prediction performance of the yield prediction model in different time periods according to the optimized prediction results, and obtain dynamic verification results; The model updating module is used to obtain new samples from the latest production data, perform incremental training on the production prediction model, update the parameters of the production prediction model, and obtain a prediction model that adapts to the latest production conditions if the dynamic verification result shows that the performance of the production prediction model has declined.

6. The oil and gas well production prediction system according to claim 5, characterized in that: The data acquisition and cleaning module uses the interpolation method to process missing values ​​of geological feature data, development history data and process parameter data in the oil and gas field database, and uses the outlier detection algorithm to process abnormal values ​​to obtain a cleaned data set.

7. The oil and gas well production prediction system according to claim 6, characterized in that: The feature extraction and dimensionality reduction module extracts characteristic variables related to production based on the cleaned data set, wherein the characteristic variables include formation pressure, permeability, wellbore pressure and production time, and reduces the dimension of the characteristic variables through principal component analysis to obtain a low-dimensional representation of the characteristic variables.

8. The oil and gas well production prediction system according to claim 7, characterized in that: The dynamic verification module adopts a time series cross-validation method to divide the cleaned data set into multiple subsets in chronological order, trains and verifies the yield prediction model in turn, evaluates the prediction performance of the yield prediction model in different time periods, and obtains dynamic verification results.

Citation Information

Patent Citations

  • Electric power material protocol inventory prediction method based on RF-LSTM

    CN117932252A

  • Oil and gas well yield prediction method

    CN118195092A

  • Shale gas yield prediction method based on fracturing construction parameter clustering

    CN118485182A

  • Intelligent supply chain and logistics optimization management method and system

    CN119250687A

  • Oil field yield prediction method and system based on digital twinborn technology

    CN119377571A

Cited By

  • Method and device for intelligently predicting post-fracturing productivity of multi-layer combined producing well

    CN122022024A