Method for predicting pipe detection period of injection-production well
By collecting production and status data from injection and production wells, and using random forest and extreme gradient boosting tree algorithms to screen features and construct decision trees, the problems of subjectivity and narrow applicability in injection and production well inspection and management cycle prediction are solved, and high-precision inspection and management cycle prediction is achieved, which reduces costs and improves the scientific nature of production decisions.
Patent Information
- Application Number
- CN202510838317.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-09-26
AI Technical Summary
The existing methods for determining the pipe inspection cycle for injection and production wells are highly subjective, have a narrow scope of application, and do not take all factors into consideration. They are difficult to achieve accurate predictions under complex working conditions, resulting in premature or late inspections and increased production costs and risks.
By collecting production and status data from multiple injection and production wells, using random forest and extreme gradient boosting tree algorithms for feature screening and model training, a prediction model for the inspection cycle of injection and production wells was constructed. Real-time predictions were performed by combining engineering data with machine learning algorithms, screening the main controlling factors and constructing a decision tree, and optimizing hyperparameters to improve prediction accuracy.
It achieves high-precision prediction of the pipe inspection cycle of injection and production wells, is applicable to most environments, and has an accuracy rate of over 90%. It reduces unnecessary operating costs, improves the scientific nature and economic benefits of production decisions, and supports intelligent oilfield management.
Smart Images

Figure CN120705507A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for predicting a pipe inspection cycle of an injection-production well, and belongs to the field of petroleum and natural gas industry. Background Art
[0002] In the oil and gas industry, pipe inspections in injection and production wells are crucial for ensuring efficient and stable production in oil and gas fields. Properly determining the inspection cycle is crucial for maintaining normal well operation, reducing production costs, and improving economic benefits. However, with the deepening of oil and gas field development, tubing corrosion in domestic oil and gas wells is becoming increasingly prominent, especially under conditions of alternating gas and water injection. As gas and water injection volumes increase, scaling and blockage become increasingly severe, significantly increasing the risk of tubing corrosion. These factors place higher demands on the precise prediction of the inspection cycle for injection and production wells. Accurate prediction of the inspection cycle is crucial. Premature inspections can lead to unnecessary operating costs, while delayed inspections can cause tubing failure, resulting in significant losses such as reduced oil production efficiency and increased well repair costs, and can even lead to serious production accidents.
[0003] Currently, traditional methods for determining the inspection cycle for injection and production wells rely primarily on empirical judgment or fixed time intervals, which present significant limitations. Empirical methods rely too heavily on subjective judgment, leading to significant discrepancies in assessment results among different technicians and failing to fully account for differences in various factors related to injection and production wells, such as production factors and state factors. Fixed-period methods, on the other hand, can easily lead to excessive or insufficient inspections. Luo Zhengshan et al. published a paper in the Journal of Safety and Environment titled "Prediction of Corrosion Rates in Gas Storage Injection and Production Pipelines Based on IAOA-KELM." They developed a model combining the Archimedean optimization algorithm with a nuclear extreme learning machine to improve the accuracy of corrosion rate predictions. However, the modeling process did not consider factors such as the number of days in production (well opening), the maximum H2S content during the period, the total gas injection volume, and the number of days the well was shut down. Feng Fuping et al. published a paper titled "Prediction of Corrosion Rate and Service Life of CO2 Injection Well Pipeline" in the journal Thermal Processing Technology. Through simulation experiments, they established a corrosion rate prediction model for Q125, 3Cr, and 13Cr pipes, and predicted the safe service life of Q125 casing. However, the factors considered were only temperature, CO2 partial pressure, and corrosion time, and the experimental research included limited pipes. This method faced many constraints in actual engineering applications, resulting in a relatively narrow scope of application. Liu Guozhen et al. published a paper titled "Prediction and Safety Control of the Ultimate Life of Heavy Oil Water-Drive Injection and Production Pipeline" in the journal Petroleum Machinery. They established a method for predicting the corrosion rate of injection and production well casing based on CO2 corrosion prediction and O2 corrosion prediction, as well as a method for calculating the strength of the pipe considering corrosion. On this basis, they derived a method for predicting the ultimate life of casing. However, this method was primarily targeted at heavy oil water-drive injection and production wells, and did not consider factors such as the maximum H2S content, production (well opening) days, and well shut-in days, resulting in a limited scope of application.
[0004] In summary, the traditional methods in the existing injection and production well inspection cycle determination technology have the limitations of subjectivity or fixed cycles. The existing models or methods have defects such as incomplete considerations and narrow scope of application, making it difficult to accurately adapt to the inspection cycle prediction needs under complex working conditions. Summary of the Invention
[0005] The present invention aims to provide a high-precision method for predicting the pipe inspection cycle of injection and production wells, effectively solving many problems existing in the existing methods for determining the pipe inspection cycle, comprehensively improving the production decision-making level and economic benefits of the oil and gas industry, and ensuring production safety.
[0006] Specifically, the implementation steps of the present invention are as follows:
[0007] Step 1: Collect the actual inspection cycle Y (days) and related production data and status data of multiple injection and production wells, and assign values to the related status data; the collected related production data include total gas injection volume X1 (10,000 cubic meters), total water injection volume X2 (10,000 cubic meters), production (well opening) days X3 (days), well shut-in days X4 (days), actual gas injection days X5 (days), maximum H2S content during the period X 11 (mg / m 3 ); The collected relevant status data include nitrogen generation mode X6, anti-corrosion measures X7, production status X8, whether it is sealed X9, oil pipe status X 10 , where the production status X8 includes the moisture content α (%) and whether it is acidified X 82 ;
[0008] The assignment rules for the relevant status data are as follows:
[0009] Nitrogen production mode X6: the value of centralized nitrogen production is 1, and the value of decentralized nitrogen production is 1.5;
[0010] Anti-corrosion measures X7: coating anti-corrosion is assigned a value of 0.4, lining anti-corrosion is assigned a value of 0.5, corrosion inhibitor anti-corrosion is assigned a value of 0.9, and no anti-corrosion measures are assigned a value of 1;
[0011] The value of production status X8 is calculated as follows:
[0012] X8=X 81 ×X 82 (1)
[0013] Where: X8 is the value assigned to the production status, dimensionless; X 81 is the value assigned to the moisture content, dimensionless. When the moisture content α is less than 50%, X 81 Assigned value 1, when the moisture content α ≥ 50% X 81 Assign a value of 1.5; X 82 is the value assigned to the acidification situation, dimensionless, no acidification X 82 Assign a value of 1, acidify X 82Assign a value of 1.1;
[0014] Whether sealed X9: The value of the column with seal is 1, and the value of the column without seal is 1.3;
[0015] Oil pipe status X 10 : If the status of the oil pipeline is new, the value is 1; if the status of the oil pipeline is old, the value is 1.2;
[0016] Construct the data set X as formula (2):
[0017]
[0018] Where: X (1,n) Indicates the total gas injection volume of the nth well; X (2,n) represents the total water injection volume of the nth well; X (3,n) Indicates the production (opening) days of the nth well; X (4,n) represents the number of days of well shut-in for the nth well; X (5,n) represents the actual gas injection days of the nth well; X (6,n) Indicates the nitrogen production mode value of the nth well; X (7,n) represents the anti-corrosion measure value of the nth well; X (8,n) represents the production status value of the nth well; X (9,n) Indicates the sealed status value of the nth well; X (10,n) Indicates the status value of the tubing of the nth well; X (11,n) Indicates the maximum H2S content value during the nth well; Y n Indicates the actual inspection cycle value of the nth well;
[0019] Step 2: Preprocess the dataset X. Perform data cleaning on the dataset X listed in formula (2), delete duplicate records, identify outliers through box plots, and then delete, correct, or process them separately according to the situation. The sample set S is obtained as shown in formula (3):
[0020]
[0021] Where: S represents the sample set; X′ (1,n) represents the pre-processed value of the total gas injection volume of the nth well; X′ (2,n) represents the pre-processed value of the total water injection volume of the nth well; X′ (3,n) represents the pre-processed value of the number of days the n-th well has been in production (opening); X′ (4,n) represents the pre-processed value of the number of days of soaking of the nth well; X′ (5,n) represents the pre-processed value of the actual gas injection days of the nth well; X′ (6,n) represents the pre-processing value of the nitrogen production mode of the nth well; X′ (7,n) represents the pretreatment value of the anti-corrosion measures for the nth well; X′ (8,n)represents the pre-processed value of the production status of the nth well; X′ (9,n) The pre-processing value indicating whether the nth well is sealed; X′ (10,n) The pre-processed value of the tubing status of the nth well; X′ (11,n) represents the pre-processed value of the maximum H2S content during the nth well; Y′ n Indicates the pre-processing value of the actual inspection cycle of the nth well;
[0022] Step 3: Screening the main controlling factors. First, Python is used to generate the actual inspection cycle feature heat map according to the sample set S as shown in formula (3), and then correlation analysis and multicollinearity elimination are performed. Then, the features of the actual inspection cycle of the injection and production wells after preprocessing are sorted using the random forest and extreme gradient boosting tree after hyperparameter optimization. Secondly, feature selection is performed based on the forward selection method to obtain the features selected by the random forest and the features selected by the extreme gradient boosting tree. All features selected by the random forest are counted as being selected once, and all features selected by the extreme gradient boosting tree are counted as being selected once. Finally, the frequencies of the selected features are calculated, and the features selected by both the random forest and the extreme gradient boosting tree, i.e., the features with a frequency of "2", are selected as the main controlling factors affecting the actual inspection cycle. The specific steps are as follows:
[0023] Step 3-1: Use Python to generate the actual inspection cycle feature heat map based on the sample set S as shown in formula (3), and then perform correlation analysis and multicollinearity elimination to exclude features with high correlation, that is, features with a correlation coefficient greater than 0.7;
[0024] Step 3-2: Define a parameter tuning space containing multiple hyperparameters, including the number of trees n_estimators, the maximum depth of the tree max_depth, the minimum number of split samples min_samples_split, and the minimum number of leaf node samples min_sample_leaf;
[0025] Step 3-3: Adjust hyperparameters and use the learning curve results to determine the number of trees n_estimators. Use the GridSearchCV class in the scikit-learn library to search for other hyperparameters and use 5-fold cross-validation to obtain the minimum generalization error of the random forest model and output the optimal hyperparameter combination for the model.
[0026] Step 3-4: Use the RandomForestClassifier class in the Python scikit-learn library to train the sample set S as shown in formula (3). After the training is completed, the importance score of each feature is obtained through the feature_importances attribute of the random forest model, and the size relationship of each feature importance score is compared to rank the features;
[0027] Step 3-5: Define a parameter space containing multiple hyperparameters, including the learning rate learning_rate, the number of trees n_estimators, the maximum depth of the tree max_depth, and the number of leaf nodes num_leaves;
[0028] Steps 3-6: Adjust hyperparameters and use the learning curve results to determine the number of trees n_estimators. Use the GridSearchCV class in the scikit-learn library to search for other hyperparameters and use 5-fold cross-validation to obtain the minimum generalization error of the extreme gradient boosting tree model and output the optimal hyperparameter combination for the model.
[0029] Step 3-7: Use the LGBMRegressor class in the Python lightgbm library to train the sample set S as shown in formula (3). After the training is completed, the importance score of each feature is obtained through the feature_importances attribute of the extreme gradient boosting tree model, and the size relationship of each feature importance score is compared to perform feature sorting;
[0030] Step 3-8: Construct a feature set D 0:k As shown in formula (4), the first feature set does not contain any features at the beginning;
[0031] D 0:k =[x 0:k ,f(x 0:k )](4)
[0032] Where: D 0:k Represents a dataset consisting of features from 0th to kth, after the importance scores of the features are sorted from high to low, x 0:k represents the feature combination consisting of the 0th to kth features, f(x 0:k ) represents the feature combination x 0:k The value of the evaluation index MSE;
[0033] Step 3-9: Based on the results of random forest feature sorting, starting from the most important feature, a new feature is added in each iteration. MSE is used as the evaluation index to evaluate the impact of the updated feature set on the performance of the random forest model. When all factors are selected, the iteration ends. According to formula (4), the change of MSE under different combination features is considered, and the feature combination with the smallest MSE is selected to obtain the features selected by random forest.
[0034] Step 3-10: Based on the results of the extreme gradient boosting tree feature sorting, starting from the most important feature, a new feature is added in each iteration, and the MSE is used as the evaluation index to evaluate the impact of the updated feature set on the performance of the extreme gradient boosting tree model. When all factors are selected, the iteration ends. According to formula (4), the change of MSE under different combination features is considered, and the feature combination with the smallest MSE is selected to obtain the features selected by the extreme gradient boosting tree.
[0035] Step 3-11: All features selected by random forest are counted once, and all features selected by extreme gradient boosting tree are counted once;
[0036] Step 3-12: Calculate the frequencies of the selected features and select the features selected by both the random forest and extreme gradient boosting tree, i.e., the features with a frequency of 2, as the main control factors affecting the actual inspection and management cycle;
[0037] Step 4: Use extreme gradient boosting trees to build a prediction model for the inspection cycle of injection and production wells and adjust hyperparameters. A greedy algorithm is used to construct the decision tree, and split gains are calculated for the first- and second-order derivatives of the residuals, while also considering regularization terms. The adjusted hyperparameters include the number of trees, n_estimators, and the maximum tree depth, max_depth. The number of trees, n_estimators, is determined using the learning curve results. The maximum tree depth, max_depth, is then searched using the GridSearchCV class in the scikit-learn library. Finally, 5-fold cross-validation is used.
[0038] Step 5: Use the pre-processed sample data of the actual inspection cycle of the injection and production wells and their main controlling factors to train the injection and production well inspection cycle prediction model;
[0039] Step 6: Calculate the evaluation index of the injection and production well inspection cycle prediction model, using the mean square error (MSE), mean absolute percentage error (MAPE), and coefficient of determination (R) 2 Conduct performance evaluations;
[0040] Step 7: Obtain data on the main controlling factors of the pipe inspection cycle of in-service injection and production wells;
[0041] Step 8: Use the trained injection-production well inspection cycle prediction model to predict the inspection cycle of the injection-production wells in service.
[0042] Furthermore, the specific steps of step 4 are as follows:
[0043] Step 4-1: Establish the objective function as shown in formula (5):
[0044]
[0045] Where: Obj is the objective function; is the loss function; Y i ′ is the pre-processing value of the ith actual inspection cycle; is the i-th predicted value; Ω(w) is the regularization term; w is the model weight;
[0046] Step 4-2: Taylor second-order expansion, extreme gradient boosting tree model uses feature splitting to continuously generate new trees, each tree corresponds to a new function, and then the loss function Perform Taylor second-order expansion, that is, the model focuses on the first-order and second-order derivatives of the current model in each iteration and fits the residual of the previous prediction until the convergence condition is met;
[0047] Step 4-3: Construct a decision tree. The extreme gradient boosting tree uses a greedy algorithm to construct a decision tree and calculates the split gain of the first-order and second-order derivatives of the residual. At the same time, considering the regularization term, the split that makes the objective function change the most before and after the split is selected.
[0048] Step 4-4: Define a parameter tuning space, where the hyperparameters include the number of trees n_estimators and the maximum depth of the tree max_depth;
[0049] Steps 4-5: Adjust hyperparameters. First, use the learning curve results to determine the number of trees n_estimators. Then, use the GridSearchCV class in the scikit-learn library to search for the maximum depth max_depth of the tree. Finally, use 5-fold cross-validation to ensure the reliability of the results.
[0050] The beneficial effects of the present invention are:
[0051] ① Breaking through the drawbacks of traditional empirical judgment and fixed-cycle methods, fully considering the production data and status of injection and production wells, combining engineering data with machine learning algorithms to achieve real-time predictions, providing support for the rational formulation of inspection and management operation cycles;
[0052] ② It is not limited to specific oil and gas field types, tubing materials, or single operating conditions. It is applicable to injection and production wells in most environments and has strong universality, providing an effective solution for predicting the inspection cycle of various injection and production wells.
[0053] ③ With an accuracy rate of over 90%, the system can compare the inspection cycle prediction results with the service time of the pipe string to directly derive the remaining service time, providing reliable data for production decision-making and enhancing the scientific nature of decision-making;
[0054] ④ Compared with traditional methods, it can accurately grasp the actual factors of the project, effectively avoid premature or late inspection and management, significantly reduce unnecessary operating costs, improve economic benefits, and make inspection and management operations more scientific in cost control;
[0055] ⑤ Provide the oil and gas industry with a new approach to predicting inspection and management cycles, from optimizing production decisions to improving economic benefits, build a solid technical foundation for intelligent oilfield management in multiple dimensions, and promote the upgrading of the industry's operation and maintenance model. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0057] Figure 1 This is a flow chart of a method for predicting the pipe inspection cycle of an injection-production well provided by the present invention;
[0058] Figure 2 This is a flow chart of screening the main control factors of the pipe inspection cycle of injection and production wells provided by the present invention;
[0059] Figure 3 This is a characteristic heat map of an actual inspection cycle provided by an embodiment of the present invention;
[0060] Figure 4 This is a graph showing the training results of a pipe inspection cycle prediction model for injection and production wells provided by an embodiment of the present invention;
[0061] Figure 5 This is a graph showing the prediction results of the pipe inspection cycle of an in-service injection-production well provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] The present invention will be described in detail below with reference to the accompanying drawings and specific implementation plans of the embodiments of the present invention.
[0063] The embodiment of the present invention discloses a method for predicting the pipe inspection cycle of an injection-production well. The specific implementation process is as follows:
[0064] Step 1: Collect the actual inspection cycle Y (days) and related production data and status data of multiple injection and production wells, and assign values to the related status data; the collected related production data include total gas injection volume X1 (10,000 cubic meters), total water injection volume X2 (10,000 cubic meters), production (well opening) days X3 (days), well shut-in days X4 (days), actual gas injection days X5 (days), maximum H2S content during the period X 11 (mg / m 3 ); The collected relevant status data include nitrogen generation mode X6, anti-corrosion measures X7, production status X8, whether it is sealed X9, oil pipe status X 10 , where the production status X8 includes the moisture content α (%) and whether it is acidified X 82 ;
[0065] The assignment rules for the relevant status data are as follows:
[0066] Nitrogen production mode X6: the value of centralized nitrogen production is 1, and the value of decentralized nitrogen production is 1.5;
[0067] Anti-corrosion measures X7: coating anti-corrosion is assigned a value of 0.4, lining anti-corrosion is assigned a value of 0.5, corrosion inhibitor anti-corrosion is assigned a value of 0.9, and no anti-corrosion measures are assigned a value of 1;
[0068] The value of production status X8 is calculated as follows:
[0069] X8=X 81 ×X 82 (1)
[0070] Where: X8 is the value assigned to the production status, dimensionless; X 81 is the value assigned to the moisture content, dimensionless. When the moisture content α is less than 50%, X 81 Assigned value 1, when the moisture content α ≥ 50% X 81 Assign a value of 1.5; X 82 is the value assigned to the acidification situation, dimensionless, no acidification X 82 Assign a value of 1, acidify X 82 Assign a value of 1.1;
[0071] Whether sealed X9: The value of the column with seal is 1, and the value of the column without seal is 1.3;
[0072] Oil pipe status X 10 : If the status of the oil pipeline is new, the value is 1; if the status of the oil pipeline is old, the value is 1.2;
[0073] The actual inspection and management cycle Y (days) and the related production data and status data assignment results are shown in Table 1;
[0074] Table 1 Actual inspection cycle of injection and production wells and related production data and status data
[0075]
[0076]
[0077]
[0078] Construct the data set X as formula (2):
[0079]
[0080] Step 2: Preprocess the dataset X. Perform data cleaning on the dataset X listed in formula (2), delete duplicate records, identify outliers through box plots, and then delete, correct, or process them separately according to the situation. The sample set S is obtained as shown in formula (3):
[0081]
[0082] Step 3: Screening the main controlling factors. First, Python is used to generate the actual inspection cycle feature heat map according to the sample set S as shown in formula (3), and then correlation analysis and multicollinearity elimination are performed. Then, the features of the actual inspection cycle of the injection and production wells after preprocessing are sorted using the random forest and extreme gradient boosting tree after hyperparameter optimization. Secondly, feature selection is performed based on the forward selection method to obtain the features selected by the random forest and the features selected by the extreme gradient boosting tree. All features selected by the random forest are counted as being selected once, and all features selected by the extreme gradient boosting tree are counted as being selected once. Finally, the frequencies of the selected features are calculated, and the features selected by both the random forest and the extreme gradient boosting tree, i.e., the features with a frequency of "2", are selected as the main controlling factors affecting the actual inspection cycle. The specific steps are as follows:
[0083] Step 3-1: Use Python and generate the actual inspection cycle characteristic heat map according to the sample set S as formula (3) Figure 3 As shown, correlation analysis and multicollinearity elimination were then performed to exclude features with high correlation, that is, features with correlation coefficients greater than 0.7. Since the correlation between each feature was not greater than 0.7, no feature was removed;
[0084] Step 3-2: Define a parameter tuning space containing multiple hyperparameters, including the number of trees n_estimators, the maximum depth of the tree max_depth, the minimum number of split samples min_samples_split, and the minimum number of leaf node samples min_sample_leaf;
[0085] Step 3-3: Adjust hyperparameters and use the learning curve results to determine the number of trees n_estimators. Use the GridSearchCV class in the scikit-learn library to search for other hyperparameters and use 5-fold cross-validation to obtain the minimum generalization error of the random forest model and output the optimal hyperparameter combination for the model. The parameter search results are shown in Table 2.
[0086] Table 2 Random forest model parameters
[0087] Parameter name Parameter default values Search range Parameter search results n_estimators 100 [50,1000] 230 max_depth None [2,50] 4 min_samples_split 2 [2,10] 8 min_sample_leaf 1 [1,6] 1
[0088] Step 3-4: Use the RandomForestClassifier class in the Python scikit-learn library to train the sample set S as shown in formula (3). After the training is completed, the importance score of each feature is obtained through the feature_importances attribute of the random forest model. As shown in Table 3, the feature importance ranking of the random forest model is: production (well opening) days X3 (days) > total water injection volume X2 (10,000 cubic meters) > production status X8 > maximum H2S content during the period X 11 (mg / m 3 )>Total gas injection volume X1 (10,000 cubic meters)>Non-well days X4 (days)>Actual gas injection days X5 (days)>Nitrogen generation mode X6>Whether sealed X9>Tubing status X 10 >Anti-corrosion measures X7;
[0089] Table 3 Random Forest Model Feature Importance Score Table
[0090]
[0091] Step 3-5: Define a parameter space containing multiple hyperparameters, including the learning rate learning_rate, the number of trees n_estimators, the maximum depth of the tree max_depth, and the number of leaf nodes num_leaves;
[0092] Steps 3-6: Adjust hyperparameters. Use the learning curve results to determine the number of trees, n_estimators. Use the GridSearchCV class in the scikit-learn library to search for other hyperparameters and perform 5-fold cross-validation to obtain the minimum generalization error of the extreme gradient boosting tree model and output the optimal hyperparameter combination for the model. The results of the parameter search are shown in Table 4.
[0093] Table 4. Parameters of the extreme gradient boosting tree model for feature importance scoring
[0094] Parameter name Parameter default values Search range Parameter search results learning_rate 0.3 [0.01,0.3] 0.13 n_estimators 100 [10,500] 100 max_depth 6 [1,10] 2 min_child_weight 1 [1,6] 1
[0095] Step 3-7: Use the LGBMRegressor class in the lightgbm library of Python to train the sample set S as shown in formula (3). After the training is completed, the importance score of each feature is obtained through the feature_importances attribute of the extreme gradient boosting tree model. As shown in Table 5, the feature importance ranking of the extreme gradient boosting tree model is: production (well opening) days X3 (days) > total water injection volume X2 (10,000 cubic meters) > nitrogen production mode X6 > total gas injection volume X1 (10,000 cubic meters) > well shut-in days X4 (days) > maximum H2S content during the period X 11 (mg / m 3)>Production status X8>Actual gas injection days X5 (days)>Whether it is sealed X9>Oil pipe status X 10 >Anti-corrosion measures X7;
[0096] Table 5. Feature importance scores of extreme gradient boosting tree model
[0097]
[0098] Step 3-8: Construct a feature set D 0:k As shown in formula (4), the first feature set does not contain any features at the beginning;
[0099] D 0:k =[x 0:k ,f(x 0:k )](4)
[0100] Where: D 0:k Represents a dataset consisting of features from 0th to kth, after the importance scores of the features are sorted from high to low, x 0:k represents the feature combination consisting of the 0th to kth features, f(x 0:k ) represents the feature combination x 0:k The value of the evaluation index MSE;
[0101] Step 3-9: Based on the results of random forest feature sorting, starting from the most important feature, a new feature is added in each iteration. MSE is used as the evaluation index to evaluate the impact of the updated feature set on the performance of the random forest model. When all factors are selected, the iteration ends. According to formula (4), the changes in MSE under different combination features are considered, and the feature combination with the smallest MSE is selected. The features selected by random forest are: production (well opening) days X3 (days), total water injection volume X2 (million cubic meters), production status X8, and maximum H2S content during the period X 11 (mg / m 3 ), total gas injection volume X1 (10,000 cubic meters), number of days of well shut-in X4 (days), actual gas injection days X5 (days);
[0102] Step 3-10: Based on the results of the extreme gradient boosting tree feature sorting, starting from the most important feature, a new feature is added in each iteration, and the MSE is used as the evaluation index to evaluate the impact of the updated feature set on the performance of the extreme gradient boosting tree model. When all factors are selected, the iteration ends. According to formula (4), the changes in MSE under different combination features are considered, and the feature combination with the smallest MSE is selected. The features selected by the extreme gradient boosting tree are: production (well opening) days X3 (days), total water injection volume X2 (10,000 cubic meters), nitrogen production mode X6, total gas injection volume X1 (10,000 cubic meters), well shut-in days X4 (days), maximum H2S content during the period X 11 (mg / m 3), production status X8, actual gas injection days X5 (days);
[0103] Step 3-11: All features selected by random forest are counted once, and all features selected by extreme gradient boosting tree are counted once;
[0104] Step 3-12: Calculate the frequency of the selected features, as shown in Table 6. Select the features selected in both the random forest and extreme gradient boosting trees, that is, the features with "frequency = 2" as the main control factors affecting the actual inspection and management cycle, specifically: production (well opening) days x3 (days), total water injection volume x2 (million cubic meters), production status x8, maximum H2S content during the period x 11 (mg / m 3 ), total gas injection volume X1 (10,000 cubic meters), number of days of well shut-in X4 (days), actual number of days of gas injection X5 (days).
[0105] Table 6 Frequency table of main control factors
[0106] feature Frequency feature Frequency <![CDATA[Production (Well Open) Days X3 (days)]]> 2 <![CDATA[Total gas injection volume X1 (10,000 m³)]]> 2 <![CDATA[Total water injection volume X2 (10,000 m³)]]> 2 <![CDATA[Soaking well days X4 (days)]]> 2 <![CDATA[Production status X8]]> 2 <![CDATA[Actual gas injection days X5 (days)]]> 2 <![CDATA[Maximum H2S content X during 11 (mg / m 3 )]]> 2 <![CDATA[Nitrogen generation mode X6]]> 1
[0107] Step 4: Use extreme gradient boosting trees to build a prediction model for the inspection cycle of injection and production wells and adjust hyperparameters. A greedy algorithm is used to build the decision tree, and split gains are calculated on the first-order and second-order derivatives of the residuals, while also considering the regularization term. The adjusted hyperparameters include the number of trees, n_estimators, and the maximum tree depth, max_depth. First, the number of trees, n_estimators, is determined using the learning curve results. Then, the maximum tree depth, max_depth, is searched using the GridSearchCV class in the scikit-learn library. Finally, 5-fold cross-validation is used.
[0108] Furthermore, step 4 includes the following steps:
[0109] Step 4-1: Establish the objective function as shown in formula (5):
[0110]
[0111] Where: Obj is the objective function; is the loss function; Y i ′ is the pre-processing value of the ith actual inspection cycle; is the i-th predicted value; Ω(w) is the regularization term; w is the model weight;
[0112] Step 4-2: Taylor second-order expansion, extreme gradient boosting tree model uses feature splitting to continuously generate new trees, each tree corresponds to a new function, and then the loss function Perform Taylor second-order expansion, that is, the model focuses on the first-order and second-order derivatives of the current model in each iteration and fits the residual of the previous prediction until the convergence condition is met;
[0113] Step 4-3: Construct a decision tree. The extreme gradient boosting tree uses a greedy algorithm to construct a decision tree and calculates the split gain of the first-order and second-order derivatives of the residual. At the same time, considering the regularization term, the split that makes the objective function change the most before and after the split is selected.
[0114] Step 4-4: Define a parameter tuning space, where the hyperparameters include the number of trees n_estimators and the maximum depth of the tree max_depth;
[0115] Steps 4-5: Adjust hyperparameters. First, use the learning curve results to determine the number of trees n_estimators. Then, use the GridSearchCV class in the scikit-learn library to search for the maximum depth max_depth of the tree. The results are shown in Table 7. Finally, use 5-fold cross-validation to ensure the reliability of the results.
[0116] Table 7 Parameters of the injection-production well inspection cycle prediction model
[0117] Parameter name Parameter default values Search range Parameter search results n_estimators 100 [10,500] 150 max_depth 6 [1,10] 1
[0118] Step 5: Use the pre-processed sample data of the actual inspection cycle of injection and production wells and their main controlling factors to train the injection and production well inspection cycle prediction model.
[0119] Furthermore, the step 5 includes the following steps:
[0120] Step 5-1: Divide the sample set S as shown in formula (3), randomly select 90% of the sample set as the training set, and the remaining 10% of the sample set as the test set;
[0121] Step 5-2: The pre-processed sample data of the actual inspection cycle of injection and production wells and their seven main control factors are input into the injection and production well inspection cycle prediction model after parameter adjustment for training. The results are as follows: Figure 4 As shown in FIG, the trained injection-production well inspection cycle prediction model is obtained.
[0122] Step 6: Calculate the evaluation index of the injection and production well inspection cycle prediction model, using the mean square error (MSE), mean absolute percentage error (MAPE), and coefficient of determination (R) 2 Conduct performance evaluation.
[0123] Furthermore, step 6 includes the following steps:
[0124] Step 6-1: Calculate and output the mean square error (MSE) of the injection and production well inspection cycle prediction model on the training set and test set as formula (6), the mean absolute percentage error (MAPE) as formula (7), and the coefficient of determination (R) 2 As shown in formula (8), the mean square error (MSE), mean absolute percentage error (MAPE) and determination coefficient (R) of the injection and production well inspection cycle prediction model on the training set can be obtained: 2 The mean square error (MSE), mean absolute percentage error (MAPE), and coefficient of determination (R) on the test set are 5589.6562, 6.63%, and 0.9559, respectively. 2 They are 10430.2166, 11.99% and 0.9236 respectively, indicating that the model performs well and the inspection cycle prediction accuracy can reach 92.36%;
[0125]
[0126] Where: Y i Indicates the actual inspection cycle value of the i-th period; represents the i-th predicted value; represents the mean of the actual inspection cycle; n represents the number of predicted data; R 2 The value range is [0,1]. The closer the value is to 1, the better the model fitting effect is. The values of MSE and MAPE are ≥ 0. The smaller the MSE and MAPE values are, the higher the model prediction accuracy is.
[0127] Step 7: Obtain data on the main controlling factors of the pipe inspection cycle of in-service injection and production wells.
[0128] Furthermore, the step 7 includes the following steps:
[0129] Step 7-1: Obtain the data of the main control factors of the inspection cycle of 41 in-service wells, including: production (well opening) days x3 (days), total water injection volume x2 (million cubic meters), production status x8, maximum H2S content during the period x 11 (mg / m 3 ), total gas injection volume X1 (10,000 cubic meters), number of days of well shut-in X4 (days), actual number of days of gas injection X5 (days).
[0130] Step 8: Use the trained injection and production well inspection cycle prediction model to predict the inspection cycle of the injection and production wells in service. The inspection cycle prediction results of 41 in-service wells are as follows: Figure 5 shown.
[0131] In summary, the present invention proposes a method for predicting the pipe inspection cycle of injection and production wells. By making full use of historical engineering data and real-time production data and status data, combined with machine learning algorithms, it successfully breaks through the limitations of traditional experience judgment and fixed cycle methods, and realizes accurate prediction of the pipe inspection cycle of injection and production wells. Subsequently, the pipe inspection cycle prediction results of the in-service injection and production wells can be compared with the service time of the tubing to obtain the remaining service time of the tubing. Knowing the remaining service time of the tubing in a timely manner helps to identify potential problems such as corrosion and damage in advance, reduce production stoppages or sudden repairs caused by tubing failures, and help formulate more scientific inspection plans. This method minimizes operating costs and resource waste, effectively avoids tubing failures or production interruptions caused by premature or late inspections, provides strong technical support for intelligent oilfield management, and provides new solutions for improving production efficiency and safety in the oil and gas industry. It has broad application prospects and market value.
[0132] The specific embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments. Those skilled in the art can make various changes based on the actual situation on site without departing from the purpose of the present invention.
Claims
1. A method for predicting the pipe inspection cycle of an injection-production well, characterized in that: The following steps are involved: Step 1: Collect the actual inspection cycle Y (days) and related production data and status data of multiple injection and production wells, and assign values to the related status data; the collected related production data include total gas injection volume X1 (10,000 cubic meters), total water injection volume X2 (10,000 cubic meters), production (well opening) days X3 (days), well shut-in days X4 (days), actual gas injection days X5 (days), maximum H2S content during the period X 11 (mg / m 3 ); The collected relevant status data include nitrogen generation mode X6, anti-corrosion measures X7, production status X8, whether it is sealed X9, oil pipe status X 10 , where the production status X8 includes the moisture content α (%) and whether it is acidified X 82 ; The assignment rules for the relevant status data are as follows: Nitrogen production mode X6: the value of centralized nitrogen production is 1, and the value of decentralized nitrogen production is 1.5; Anti-corrosion measures X7: coating anti-corrosion is assigned a value of 0.4, lining anti-corrosion is assigned a value of 0.5, corrosion inhibitor anti-corrosion is assigned a value of 0.9, and no anti-corrosion measures are assigned a value of 1; The value of production status X8 is calculated as follows: X8=X 81 ×X 82 (1) Where: X8 is the value assigned to the production status, dimensionless; X 81 is the value assigned to the moisture content, dimensionless. When the moisture content α is less than 50%, X 81 Assigned value 1, when the moisture content α ≥ 50% X 81 Assign a value of 1.5; X 82 is the value assigned to the acidification situation, dimensionless, no acidification X 82 Assign a value of 1, acidify X 82 Assign a value of 1.1; Whether sealed X9: If the column is sealed, the value is 1; if the column is not sealed, the value is 1.3; Oil pipe status X 10 : If the status of the oil pipeline is new, the value is 1; if the status of the oil pipeline is old, the value is 1.2; Construct the data set X as formula (2): Where: X (1,n) Indicates the total gas injection volume of the nth well; X (2,n) represents the total water injection volume of the nth well; X (3,n) Indicates the production (opening) days of the nth well; X (4,n) represents the number of days of well shut-in for the nth well; X (5,n) represents the actual gas injection days value of the nth well; X (6,n) Indicates the nitrogen production mode value of the nth well; X (7,n) represents the anti-corrosion measure value of the nth well; X (8,n) represents the production status value of the nth well; X (9,n) Indicates the sealed status value of the nth well; X (10,n) Indicates the status value of the tubing of the nth well; X (11,n) Indicates the maximum H2S content value during the nth well; Y n Indicates the actual inspection cycle value of the nth well; Step 2: Preprocess the dataset X. Perform data cleaning on the dataset X listed in formula (2), delete duplicate records, identify outliers through box plots, and then delete, correct, or process them separately according to the situation. The sample set S is obtained as shown in formula (3): Where: S represents the sample set; X′ (1,n) represents the pre-processed value of the total gas injection volume of the nth well; X′ (2,n) represents the pre-processed value of the total water injection volume of the nth well; X′ (3,n) represents the pre-processed value of the number of days the n-th well has been in production (opening); X′ (4,n) represents the pre-processed value of the number of days of soaking of the nth well; X′ (5,n) represents the pre-processed value of the actual gas injection days of the nth well; X′ (6,n) represents the pre-processing value of the nitrogen production mode of the nth well; X′ (7,n) represents the pretreatment value of the anti-corrosion measures for the nth well; X′ (8,n) represents the pre-processed value of the production status of the nth well; X′ (9,n) The pre-processing value indicating whether the nth well is sealed; X′ (10,n) The pre-processed value of the tubing status of the nth well; X′ (11,n) represents the pre-processed value of the maximum H2S content during the nth well; Y′ n Indicates the pre-processing value of the actual inspection cycle of the nth well; Step 3: Screening the main controlling factors. First, Python is used to generate a heat map of the actual inspection and maintenance cycle characteristics based on the sample set S as shown in formula (3), and then correlation analysis and multicollinearity elimination are performed. Then, the features of the actual inspection and maintenance cycle of the injection and production wells after preprocessing are sorted using the random forest and extreme gradient boosting tree after hyperparameter optimization. Secondly, feature selection is performed based on the forward selection method to obtain the features selected by the random forest and the features selected by the extreme gradient boosting tree. All features selected by the random forest are counted as being selected once, and all features selected by the extreme gradient boosting tree are counted as being selected once. Finally, the frequencies of the selected features are calculated, and the features selected by both the random forest and the extreme gradient boosting tree, i.e., the features with "frequency = 2", are selected as the main controlling factors affecting the actual inspection and maintenance cycle. The specific steps are as follows: Step 3-1: Use Python to generate the actual inspection cycle feature heat map based on the sample set S as shown in formula (3), and then perform correlation analysis and multicollinearity elimination to exclude features with high correlation, that is, features with a correlation coefficient greater than 0.7; Step 3-2: Define a parameter tuning space containing multiple hyperparameters, including the number of trees n_estimators, the maximum depth of the tree max_depth, the minimum number of split samples min_samples_split, and the minimum number of leaf node samples min_sample_leaf; Step 3-3: Adjust hyperparameters and use the learning curve results to determine the number of trees n_estimators. Use the GridSearchCV class in the scikit-learn library to search for other hyperparameters and use 5-fold cross-validation to obtain the minimum generalization error of the random forest model and output the optimal hyperparameter combination for the model. Step 3-4: Use the RandomForestClassifier class in the Python scikit-learn library to train the sample set S as shown in formula (3). After the training is completed, the importance score of each feature is obtained through the feature_importances attribute of the random forest model, and the size relationship of each feature importance score is compared to rank the features; Step 3-5: Define a parameter space containing multiple hyperparameters, including the learning rate learning_rate, the number of trees n_estimators, the maximum depth of the tree max_depth, and the number of leaf nodes num_leaves; Steps 3-6: Adjust hyperparameters and use the learning curve results to determine the number of trees n_estimators. Use the GridSearchCV class in the scikit-learn library to search for other hyperparameters and use 5-fold cross-validation to obtain the minimum generalization error of the extreme gradient boosting tree model and output the optimal hyperparameter combination for the model. Step 3-7: Use the LGBMRegressor class in the Python lightgbm library to train the sample set S as shown in formula (3). After the training is completed, the importance score of each feature is obtained through the feature_importances attribute of the extreme gradient boosting tree model, and the size relationship of each feature importance score is compared to perform feature sorting; Step 3-8: Construct a feature set D 0:k As shown in formula (4), the first feature set does not contain any features at the beginning; D 0:k =[x 0:k ,f(x 0:k )] (4) Where: D 0:k Represents a dataset consisting of features from 0th to kth, after the importance scores of the features are sorted from high to low, x 0:k represents the feature combination consisting of the 0th to kth features, f(x 0:k ) represents the feature combination x 0:k The value of the evaluation index MSE; Step 3-9: Based on the results of random forest feature sorting, starting from the most important feature, a new feature is added in each iteration. MSE is used as the evaluation index to evaluate the impact of the updated feature set on the performance of the random forest model. When all factors are selected, the iteration ends. According to formula (4), the change of MSE under different combination features is considered, and the feature combination with the smallest MSE is selected to obtain the features selected by random forest. Step 3-10: Based on the results of the extreme gradient boosting tree feature sorting, starting from the most important feature, a new feature is added in each iteration, and the MSE is used as the evaluation index to evaluate the impact of the updated feature set on the performance of the extreme gradient boosting tree model. When all factors are selected, the iteration ends. According to formula (4), the change of MSE under different combination features is considered, and the feature combination with the smallest MSE is selected to obtain the features selected by the extreme gradient boosting tree. Step 3-11: All features selected by random forest are counted once, and all features selected by extreme gradient boosting tree are counted once; Step 3-12: Calculate the frequencies of the selected features and select the features selected by both the random forest and extreme gradient boosting tree, i.e., those with a frequency of 2, as the main control factors affecting the actual inspection and management cycle; Step 4: Use extreme gradient boosting trees to build a prediction model for the inspection cycle of injection and production wells and adjust hyperparameters. A greedy algorithm is used to construct the decision tree, and split gains are calculated for the first- and second-order derivatives of the residuals, while also considering regularization terms. The adjusted hyperparameters include the number of trees, n_estimators, and the maximum tree depth, max_depth. The number of trees, n_estimators, is determined using the learning curve results. The maximum tree depth, max_depth, is then searched using the GridSearchCV class in the scikit-learn library. Finally, 5-fold cross-validation is used. Step 5: Use the pre-processed sample data of the actual inspection cycle of the injection and production wells and their main controlling factors to train the injection and production well inspection cycle prediction model; Step 6: Calculate the evaluation index of the injection and production well inspection cycle prediction model, using the mean square error (MSE), mean absolute percentage error (MAPE), and determination coefficient (R) 2 Conduct performance evaluations; Step 7: Obtain data on the main controlling factors of the pipe inspection cycle of in-service injection and production wells; Step 8: Use the trained injection-production well inspection cycle prediction model to predict the inspection cycle of the injection-production wells in service.
2. A method for predicting the pipe inspection cycle of an injection-production well according to claim 1, characterized in that: The step 4 specifically includes: Step 4-1: Establishing the objective function as shown in formula (5): Where: Obj is the objective function; is the loss function; Y′ i is the pre-processing value of the ith actual inspection cycle; is the i-th predicted value; Ω(w) is the regularization term; w is the model weight; Step 4-2: Taylor second-order expansion, the extreme gradient boosting tree model uses feature splitting to continuously generate new trees, each tree corresponds to a new function, and then the loss function Perform Taylor second-order expansion, that is, the model focuses on the first-order and second-order derivatives of the current model in each iteration and fits the residual of the previous prediction until the convergence condition is met; Step 4-3: Construct a decision tree. The extreme gradient boosting tree uses a greedy algorithm to construct a decision tree and calculates the split gain of the first-order and second-order derivatives of the residual. At the same time, considering the regularization term, the split that makes the objective function change the most before and after the split is selected. Step 4-4: Define a parameter tuning space, where the hyperparameters include the number of trees n_estimators and the maximum depth of the tree max_depth; Steps 4-5: Adjust hyperparameters. First, use the learning curve results to determine the number of trees n_estimators. Then, use the GridSearchCV class in the scikit-learn library to search for the maximum depth max_depth of the tree. Finally, use 5-fold cross-validation to ensure the reliability of the results.
3. The method for predicting the pipe inspection cycle of an injection-production well according to claim 1, characterized in that: The step 5 specifically includes: Step 5-1: Divide the sample set S as shown in formula (3), randomly select 90% of the sample set as the training set, and the remaining 10% of the sample set as the test set; Step 5-2: Input the pre-processed sample data of the actual inspection cycle of the injection-production well and its main controlling factors into the injection-production well inspection cycle prediction model after parameter adjustment for training to obtain a trained injection-production well inspection cycle prediction model.
4. A method for predicting the pipe inspection cycle of an injection-production well according to claim 1, characterized in that: The step 6 specifically includes: Use the mean square error MSE as formula (6), the mean absolute percentage error MAPE as formula (7), and the coefficient of determination R 2 The prediction performance of the injection-production well inspection cycle prediction model is evaluated as shown in formula (8) to obtain the evaluation results; Where: Y i Indicates the actual inspection cycle value of the i-th period; represents the i-th predicted value; represents the mean of the actual inspection cycle; n represents the number of predicted data; R 2 The value range is [0,1]. The closer the value is to 1, the better the model fitting effect is. The values of MSE and MAPE are ≥ 0. The smaller the MSE and MAPE values are, the higher the model prediction accuracy is.