Prejudgment method and device for classification of fracture pressure and yield of tight oil reservoir
By constructing a machine learning-based fracture pressure and production prediction model, the problem of inaccurate fracture pressure calculation in the hydraulic fracturing design of tight oil wells was solved, and accurate production prediction and development strategy optimization of tight oil wells were achieved.
Patent Information
- Application Number
- CN202510817014.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In the existing technology of tight oil well hydraulic fracturing design, the fracture pressure calculation is inaccurate, resulting in construction failure or accidents, and the production cannot be accurately analyzed, and there is a lack of effective prediction methods.
By collecting tight oil horizontal well data and using a variety of machine learning algorithms to build fracture pressure prediction models and production prediction models, the main controlling factors are screened out, a prediction method for tight oil reservoir fracture pressure and production classification is established, and a development strategy is generated.
It achieves accurate prediction of the fracture pressure and production of tight oil wells, provides reasonable development strategies, increases production and reduces production costs, and adapts to the complex geological conditions of different blocks.
Smart Images

Figure CN120649871A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of oil reservoir exploration, and in particular to a method and device for predicting the fracture pressure and production classification of a tight oil reservoir. Background Art
[0002] This section is intended to provide a background or context for embodiments of the present invention. No description herein is admitted to be prior art by virtue of its inclusion in this section.
[0003] When designing hydraulic fracturing for tight oil wells, a common challenge is the occurrence of abnormal fracture pressure values. These abnormal fracture pressures often lead to abnormal operation pressures, which can lead to failure or accidents during the operation. Therefore, fracture pressure is a crucial parameter in hydraulic fracturing design and a crucial factor in optimizing operation parameters. Due to the complex geological structure of different blocks, traditional fracture pressure calculation formulas are no longer sufficient. Although regional fitting formulas have been developed based on these formulas, their accuracy and rationality are significantly limited. Furthermore, inaccurate fracture pressure measurements make it impossible to accurately analyze and calculate tight oil production.
[0004] In summary, there is an urgent need for a technical solution that can overcome the shortcomings of existing technologies, accurately predict the fracture pressure and production classification of tight oil reservoirs, and provide strong technical support for the development of tight oil wells. Summary of the Invention
[0005] To address the problems of the prior art, the present invention proposes a method and device for predicting the breakdown pressure and production classification of tight oil reservoirs. By combining theory and practice with the actual geological data and production data of the target block, the present invention establishes a mathematical model for predicting the breakdown pressure of tight oil wells. This model can be used to predict the breakdown pressure, thereby guiding hydraulic fracturing design and determining whether the breakdown pressure can break open each cluster. Based on this, a machine learning prediction method for the classification of tight oil production is established to verify its developability. The overall solution establishes a complete set of prediction models to provide guidance for the subsequent development of tight oil wells, which is of great significance for increasing the production of tight oil wells, reducing production costs, and promoting technological innovation in actual production.
[0006] In a first aspect of an embodiment of the present invention, a method for predicting tight oil reservoir fracture pressure and production classification is proposed, the method comprising:
[0007] Collect tight oil horizontal well data;
[0008] Extracting factors affecting the breakdown pressure of the tight oil horizontal well from the tight oil horizontal well data, inputting the factors affecting the breakdown pressure of the tight oil horizontal well into a tight oil horizontal well breakdown pressure prediction model to obtain the breakdown pressure of the tight oil horizontal well; wherein the tight oil horizontal well breakdown pressure prediction model is trained according to the following method: constructing multiple initial models respectively using multiple machine learning algorithms, constructing a first training set and a first test set using sample data of factors affecting the breakdown pressure of the tight oil horizontal well, training the multiple initial models respectively using the first training set, and testing them using the first test set, selecting a model based on the test results, and obtaining the tight oil horizontal well breakdown pressure prediction model;
[0009] Extracting factors affecting tight oil production from the tight oil horizontal well data, inputting the factors affecting tight oil production into a tight oil production prediction model to obtain tight oil production; wherein the tight oil production prediction model is trained according to the following method: constructing multiple initial models using multiple prediction algorithms, constructing a second training set and a second test set using sample data of factors affecting tight oil production, training the multiple initial models using the second training set, and testing them using the second test set, selecting a model based on the test results, and obtaining a tight oil horizontal well fracture pressure prediction model;
[0010] The tight oil horizontal wells are classified according to the tight oil production to obtain a tight oil production classification result, and a tight oil horizontal well development strategy is generated according to the tight oil horizontal well fracture pressure and the tight oil production classification result.
[0011] In a second aspect of an embodiment of the present invention, a device for predicting tight oil reservoir fracture pressure and production classification is provided, the device comprising:
[0012] Data acquisition module, used to collect tight oil horizontal well data;
[0013] A tight oil horizontal well burst pressure prediction module is used to extract factors affecting the burst pressure of the tight oil horizontal well from the tight oil horizontal well data, input the factors affecting the burst pressure of the tight oil horizontal well into a tight oil horizontal well burst pressure prediction model, and obtain the tight oil horizontal well burst pressure; wherein the tight oil horizontal well burst pressure prediction model is trained according to the following method: multiple initial models are constructed respectively using multiple machine learning algorithms, a first training set and a first test set are constructed using sample data of factors affecting the burst pressure of the tight oil horizontal well, the multiple initial models are trained respectively using the first training set, and the models are tested using the first test set, and a model is selected according to the test results to obtain the tight oil horizontal well burst pressure prediction model;
[0014] a tight oil production prediction module, configured to extract factors influencing tight oil production from the tight oil horizontal well data, input the factors influencing tight oil production into a tight oil production prediction model, and obtain tight oil production; wherein the tight oil production prediction model is trained according to the following method: constructing multiple initial models using multiple prediction algorithms, constructing a second training set and a second test set using sample data of factors influencing tight oil production, training the multiple initial models using the second training set, and testing them using the second test set, selecting a model based on the test results, and obtaining a tight oil horizontal well fracture pressure prediction model;
[0015] The development strategy generation module is used to classify the tight oil horizontal wells according to the tight oil production, obtain the tight oil production classification results, and generate the tight oil horizontal well development strategy according to the tight oil horizontal well fracture pressure and the tight oil production classification results.
[0016] In a third aspect of an embodiment of the present invention, a computer device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a method for predicting the fracture pressure and production classification of tight oil reservoirs when executing the computer program.
[0017] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a method for predicting the fracture pressure and production classification of tight oil reservoirs is implemented.
[0018] In a fifth aspect of an embodiment of the present invention, a computer program product is proposed. The computer program product includes a computer program. When the computer program is executed by a processor, a method for predicting the fracture pressure and production classification of tight oil reservoirs is implemented.
[0019] The method and device for predicting the fracture pressure and production classification of tight oil reservoirs proposed in the present invention screen the main controlling factors affecting the fracture pressure and the main controlling factors of the production classification, construct a fracture pressure prediction model based on multiple machine learning algorithms, select the model that best suits the complex geological conditions of different blocks, and further use multiple prediction algorithms to construct a production prediction model to achieve production classification prediction. It can effectively distinguish between high-yield wells and low-yield wells, provide a reasonable development strategy for tight oil horizontal wells, and provide accurate technical support for fracturing design and benefit evaluation of tight oil development. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 The figure is a flow chart of a method for predicting the fracture pressure and production classification of tight oil reservoirs according to an embodiment of the present invention.
[0022] Figure 2 It is a schematic diagram of a data preprocessing process according to an embodiment of the present invention.
[0023] Figure 3 1 is a flow chart of extracting factors influencing the fracture pressure of a tight oil horizontal well according to an embodiment of the present invention.
[0024] Figure 4 2 is a schematic diagram of random forest importance ranking results when processing factors affecting the fracture pressure of tight oil horizontal wells according to an embodiment of the present invention.
[0025] Figure 5 4 is a schematic diagram of correlation analysis when processing factors affecting tight oil production according to an embodiment of the present invention.
[0026] Figure 6 1 is a schematic diagram of a training process for a tight oil horizontal well fracture pressure prediction model according to an embodiment of the present invention.
[0027] Figure 7 3 is a schematic diagram comparing the actual value and the predicted value of GA-BP according to an embodiment of the present invention.
[0028] Figure 8 3 is a schematic diagram comparing the actual value and the predicted value of SVR according to an embodiment of the present invention.
[0029] Figure 9 3 is a schematic diagram comparing the true value and the predicted value of XGBoost according to an embodiment of the present invention.
[0030] Figure 10 2 is a schematic diagram for comparing prediction results according to an embodiment of the present invention.
[0031] Figure 11 It is a flowchart of extracting factors affecting tight oil production from the tight oil horizontal well data according to an embodiment of the present invention.
[0032] Figure 12 2. It is a schematic diagram of the random forest importance ranking results when processing factors affecting tight oil production according to one embodiment of the present invention.
[0033] Figure 134 is a schematic diagram of correlation analysis when processing factors affecting tight oil production according to an embodiment of the present invention.
[0034] Figure 14 1 is a schematic diagram of the training process of the tight oil production prediction model according to an embodiment of the present invention.
[0035] Figure 15 FIG. 4 is a schematic diagram of SVC parameter optimization according to an embodiment of the present invention.
[0036] Figure 16 FIG. 4 is a schematic diagram of RF parameter optimization according to an embodiment of the present invention.
[0037] Figure 17 1 is a schematic diagram of AdaBoost parameter optimization according to an embodiment of the present invention.
[0038] Figure 18 2 is a schematic diagram of XGBoost parameter optimization according to an embodiment of the present invention.
[0039] Figure 19 2 is a schematic diagram of an SVC confusion matrix according to an embodiment of the present invention.
[0040] Figure 20 FIG. 4 is a schematic diagram of an RF confusion matrix according to an embodiment of the present invention.
[0041] Figure 21 Schematic diagram of an AdaBoost confusion matrix according to an embodiment of the present invention.
[0042] Figure 22 Schematic diagram of the XGBoost confusion matrix according to an embodiment of the present invention.
[0043] Figure 23 It is a technical route schematic diagram of an embodiment of the present invention.
[0044] Figure 24 2 is a schematic diagram of the architecture of a device for predicting tight oil reservoir fracture pressure and production classification according to an embodiment of the present invention.
[0045] Figure 25 It is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0047] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.
[0048] According to an embodiment of the present invention, a method and device for predicting the fracture pressure and production classification of tight oil reservoirs are proposed, which relate to the field of oil reservoir exploration technology.
[0049] The principles and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0050] Figure 1 FIG. 1 is a flow chart of a method for predicting the fracture pressure and production classification of tight oil reservoirs according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0051] S101, collecting tight oil horizontal well data;
[0052] S102, extracting factors affecting the breakdown pressure of the tight oil horizontal well from the tight oil horizontal well data, inputting the factors affecting the breakdown pressure of the tight oil horizontal well into a tight oil horizontal well breakdown pressure prediction model to obtain the breakdown pressure of the tight oil horizontal well; wherein the tight oil horizontal well breakdown pressure prediction model is trained according to the following method: constructing multiple initial models using multiple machine learning algorithms, constructing a first training set and a first test set using sample data of factors affecting the breakdown pressure of the tight oil horizontal well, training the multiple initial models using the first training set, and testing them using the first test set, selecting a model based on the test results, and obtaining the tight oil horizontal well breakdown pressure prediction model;
[0053] S103, extracting factors affecting tight oil production from the tight oil horizontal well data, inputting the factors affecting tight oil production into a tight oil production prediction model to obtain tight oil production; wherein the tight oil production prediction model is trained according to the following method: constructing multiple initial models using multiple prediction algorithms, constructing a second training set and a second test set using sample data of factors affecting tight oil production, training the multiple initial models using the second training set, and testing them using the second test set, selecting a model based on the test results, and obtaining a tight oil horizontal well fracture pressure prediction model;
[0054] S104, classifying the tight oil horizontal wells according to the tight oil production to obtain a tight oil production classification result, and generating a tight oil horizontal well development strategy according to the tight oil horizontal well fracture pressure and the tight oil production classification result.
[0055] The method for predicting the fracture pressure and production classification of tight oil reservoirs proposed in the present invention screens the main controlling factors affecting the fracture pressure and the main controlling factors of the production classification, constructs a fracture pressure prediction model based on multiple machine learning algorithms, selects the model that best suits the complex geological conditions of different blocks, and further uses multiple prediction algorithms to construct a production prediction model to achieve production classification prediction. It can effectively distinguish high-yield wells from low-yield wells, provide a reasonable development strategy for tight oil horizontal wells, and provide accurate technical support for the fracturing design and benefit evaluation of tight oil development.
[0056] In order to explain the above prediction method of tight oil reservoir fracture pressure and production classification more clearly, each step is described in detail below.
[0057] In one embodiment, S101 , data of tight oil horizontal wells are collected.
[0058] Due to the existence of anomalies and missing data, data preprocessing is required. For details, refer to Figure 2 After collecting the tight oil horizontal well data, the method further includes:
[0059] S201, pre-processing the tight oil horizontal well data, filling in missing data, removing outliers and filling in missing data;
[0060] S202, based on the pre-processed tight oil horizontal well data, preliminarily select the factors that affect the fracture pressure and production of the tight oil horizontal well. The preliminarily selection method is as follows: preliminary selection of factors through manual investigation, analysis and induction; and analysis and organization of the logging, fracturing and production data of different blocks of the tight oil horizontal well to eliminate factors not involved in the block and factors with abnormal data.
[0061] In practical applications, fracture pressure prediction and production classification are based on a wealth of characteristic data. However, due to incompleteness, inconsistency, and significant missing data, direct data mining is often impossible or produces poor results. To improve data mining quality, data preprocessing is necessary. Research on fracture pressure prediction is conducted on a cluster basis, while production classification is conducted on a well basis. For example, by collecting and organizing field data on tight oil in a target area, 102 clusters of well logging data and 295 well production data samples were obtained.
[0062] Missing values are generated when measurements are not taken during construction, or when measurements are taken but not recorded. Methods for addressing missing values primarily include filling in missing values and rational summarization. Abnormalities in production data are caused by human factors, equipment, and environmental factors, such as weather, mechanical failure, or improper management. These factors are unpredictable and have a significant impact on tight gas fracture pressure. Based on analysis of production patterns, abnormal values caused by uncertainties in the production process are not considered and are removed and filled in as missing data.
[0063] Before feature extraction, we conducted a preliminary selection of factors influencing the breakdown pressure and production classification of tight oil horizontal wells. This was primarily accomplished through two approaches: 1. Literature research: By reviewing extensive literature and analyzing and summarizing the research of numerous scholars, we initially selected some parameters. 2. We analyzed and organized logging, fracturing, and production data from different blocks of tight oil horizontal wells. We eliminated factors not present in some blocks and factors with data anomalies.
[0064] In one embodiment, S102, factors affecting the breakdown pressure of the tight oil horizontal well are extracted from the tight oil horizontal well data, and the factors affecting the breakdown pressure of the tight oil horizontal well are input into a tight oil horizontal well breakdown pressure prediction model to obtain the breakdown pressure of the tight oil horizontal well; wherein, the tight oil horizontal well breakdown pressure prediction model is trained according to the following method: multiple initial models are constructed respectively through multiple machine learning algorithms, a first training set and a first test set are constructed using sample data of factors affecting the breakdown pressure of the tight oil horizontal well, the multiple initial models are trained respectively using the first training set, and the first test set is used to test, and a model is selected according to the test results to obtain the tight oil horizontal well breakdown pressure prediction model.
[0065] refer to Figure 3 The specific process of extracting factors affecting the fracture pressure of tight oil horizontal wells from the tight oil horizontal well data is as follows:
[0066] S301, for the tight oil horizontal well data, using a random forest importance algorithm and a grey relational analysis algorithm to obtain all factors affecting the breakdown pressure of the tight oil horizontal well and perform a comprehensive ranking of the influence;
[0067] S302, eliminating factors that do not meet the correlation standard with the breakdown pressure of the tight oil horizontal well through stepwise regression analysis, and obtaining the main controlling factors affecting the breakdown pressure of the tight oil horizontal well.
[0068] Among them, the main controlling factors affecting the fracture pressure of tight oil horizontal wells include at least: acoustic wave delay, density, Poisson's ratio, brittleness index, depth, minimum principal stress, natural gamma and maximum principal stress.
[0069] Random Forest is an ensemble learning algorithm based on a decision tree of a classification and regression tree. The underlying predictor is a decision tree of a classification and regression tree (CRAT). Random Forest is suitable for processing high-dimensional datasets with highly correlated features because it is insensitive to collinearity. Random Forest determines the importance of different variables by randomly replacing variables in sample data and then calculating the resulting decrease in prediction accuracy. The value used to determine variable importance is described by the mean square error, which represents the increase in model prediction accuracy.
[0070] In grey relational analysis, grey represents that the information obtained about an object or system may not be completely available or clear, while relation represents that there is a certain connection between objects or systems. Since grey relational analysis quantifies comparable factors, the degree of correlation can be obtained through calculation. The specific steps are:
[0071] 1. Determine the analysis sequence. The reference sequence can be used to obtain characteristics related to the target entity, while the comparison sequence can be used to obtain characteristics related to the reference entity. Therefore, it is necessary to first determine the reference sequence and the comparison sequence. The reference sequence refers to the sequence in all data sequences that reflects the system behavior characteristics.
[0072] 2. Dimensionless variables. To eliminate the problem of incorrect conclusions due to different meanings and very similar units among various factors in the indicator system, it is necessary to dimensionlessly process the data to maintain indicator consistency. Initialization and averaging can be used to eliminate dimensionless variables.
[0073] 3. Calculate the correlation coefficient and degree of correlation. The correlation coefficient represents the degree of correlation between the reference sequence and the comparison sequence at different times (at various points on the curve). The degree of correlation is calculated based on the correlation coefficient.
[0074] 4. Sort the correlations by size.
[0075] Taking a target block as an example, 16 factors influencing fracture pressure were summarized based on 102 clusters of data from the target block. These 16 factors were then ranked by importance using a random forest method. The results of the random forest importance ranking of all factors are shown in Table 1.
[0076] Table 1 Comprehensive ranking results of factors affecting the fracture pressure of tight oil horizontal wells
[0077]
[0078] Figure 4This is a schematic diagram of the random forest importance ranking results when processing factors affecting the fracture pressure of tight oil horizontal wells according to an embodiment of the present invention. Figure 4 The random forest importance ranking results shown in the figure show that the eight factors with the highest scores for tight oil fracture pressure obtained by random forest analysis are acoustic transit time, Poisson's ratio, density, depth, brittleness index, natural gamma ray, minimum principal stress, Young's modulus, and wellhead pressure. The remaining eight factors (Young's modulus, wellhead pressure, displacement, porosity, wellbore caliper 2, wellbore caliper 3, wellbore caliper 1, and resistivity) have lower importance scores.
[0079] The data is processed dimensionlessly and standardized using the minimization method. The correlation analysis of these 16 factors is carried out in combination with the grey correlation analysis (e.g. Figure 5 (As shown in the figure, which is a schematic diagram of correlation analysis). The correlation of fracture pressure is as follows: acoustic wave transit time, density, brittleness index, Poisson's ratio, depth, minimum principal stress, porosity, natural gamma, maximum principal stress, displacement, Young's modulus, wellbore caliper 2, resistivity, wellbore caliper 1, wellhead pressure, and wellbore caliper 3.
[0080] Combining the results of random forest importance ranking and grey correlation analysis correlation ranking, a comprehensive ranking of the two methods was obtained through statistical analysis, and the main controlling factors affecting the fracture pressure of tight oil were obtained. The specific ranking results are shown in Table 1.
[0081] Referring to Table 1, we can see that the 10 factors of acoustic wave time difference, density, Poisson's ratio, brittleness index, depth, minimum principal stress, natural gamma, maximum principal stress, porosity, and Young's modulus rank high. Figure 5 , is a schematic diagram of the correlation analysis of factors affecting the fracture pressure of tight oil horizontal wells in one embodiment of the present invention. At the same time, the relative degree trend line analysis of the two models (such as Figure 4 , Figure 5 ), the scores and rankings of the six factors, namely, displacement, well diameter 2, wellhead pressure, well diameter 1, resistivity, and well diameter 3, were relatively small, so they were eliminated.
[0082] The analysis revealed that the remaining 10 geological factors had a significant impact on the fracture pressure of tight oil, so a secondary screening was performed using stepwise regression analysis. Since there were 10 factors in the regression model, there might be factors that were not significantly correlated with the fracture pressure. Therefore, a backward stepwise regression method was used to select the main controlling factors, construct a full variable model, and gradually identify insignificant factors until all factors were significantly correlated with the fracture pressure. The full variable model is:
[0083] Y=a1X1+a2X2+...+a 10 X 10 +X0
[0084] Where Y represents the burst pressure, X iThe specific meaning is shown in Table 2. i Represents the coefficient, i=1,2,...,10.
[0085] Table 2 Definition of relevant variables
[0086]
[0087] After three backward regressions, porosity and Young's modulus did not meet the criteria and were removed from the model. After final screening, eight factors with significant impact on tight oil fracture pressure were obtained: acoustic time difference, density, Poisson's ratio, brittleness index, depth, minimum principal stress, natural gamma, and maximum principal stress.
[0088] refer to Figure 6 The training method of the tight oil horizontal well fracture pressure prediction model includes:
[0089] S601, constructing corresponding initial models respectively through XGBoost algorithm, GA-BP neural network, support vector regression algorithm (SVR) and partition fitting algorithm (conventional method);
[0090] S602 , after training and testing multiple initial models respectively, corresponding prediction accuracy rates are obtained, and the model with the highest prediction accuracy rate is selected as the tight oil horizontal well fracture pressure prediction model.
[0091] Specifically, the partition fitting algorithm assumes that the rock has no permeability and applies the Terzaghi effective stress to propose the first formula for calculating the rock fracture pressure when vertical fractures occur in the formation under open hole completion conditions:
[0092] p b =3σ h -σ H +σ f -p o ;
[0093] p b represents the rock fracture pressure; σ h represents the minimum horizontal principal stress of the formation; σ H represents the maximum horizontal principal stress of the formation; σ f Indicates the uniaxial tensile stress intensity of the formation rock; p o Indicates the pore pressure of the formation rock. If it is an impermeable rock, the pore pressure p o is 0.
[0094] In actual application and production, the traditional fracture pressure calculation formula can no longer meet the needs. In view of complex geological structures and other conditions, a polynomial fitting formula is performed on the basis of the traditional fracture and fracturing calculation formula to obtain fitting formulas for different blocks.
[0095] The Support Vector Regression (SVR) algorithm uses a regression-type support vector machine, which is an important model for processing regression problems. SVR differs from some traditional regression models in the following ways, as shown in Table 3.
[0096] Table 3 Differences between SVR model and traditional regression model
[0097]
[0098] The GA-BP neural network is a BP neural network algorithm based on the genetic algorithm. This algorithm is a random search method derived from the genetic and evolutionary mechanisms of nature. The algorithm first encodes and assigns values to individuals and calculates their fitness. Then, based on the fitness function, the best individuals are selected. All selected individuals are then subjected to a series of operations, including selection, crossover, and mutation. Individuals with good fitness are retained, while those with poor fitness are discarded. New individuals inherit the good genes of the previous generation. This process is repeated until the required number of iterations is met or the number of iterations is reached.
[0099] The difference between the GA-BP neural network and the BP neural network is that the initial weights and thresholds of the GA-BP neural network are optimized using genetic algorithms such as selection, crossover, and mutation to obtain the optimal weights and thresholds. The most important thing for the GA-BP neural network model to predict coalbed methane production is to optimize the network weights and thresholds.
[0100] The XGBoost algorithm is a member of the Boost family of algorithms. The fundamental principle of Boost is to construct a highly accurate strong classifier from multiple simple weak classifiers. XGBoost is primarily an improvement on the Gradient Boosted Decision Tree (GBDT). Besides some differences in engineering implementation and problem solving, the primary difference between the two lies in the objective function. The boosting concept combines individual learners to create dependencies, enabling efficient construction of boosted trees and parallel execution. The XGBoost algorithm boasts advantages such as fast computation, high efficiency and accuracy, high stability, and strong generalization. Its key concept is to learn a new function by adding trees and fitting the residuals of the final prediction. Sample scores are then obtained, and the final prediction score is then summed by adding the scores of each tree. During the training process, when building a tree, a key issue is finding the optimal split point for leaf nodes. XGBoost supports two node splitting methods: greedy and approximate.
[0101] In order to accurately and intuitively evaluate the performance of the three models, the following three evaluation indicators are used: 1. Determination coefficient R 2, the closer the value is to 1, the better the prediction accuracy of the model; 2. Mean relative error MAPE, its value range is [0,+∞), the closer it is to 0, the better the model performance is. If the value exceeds 100%, it means that the model is a poor model; 3. Mean square error RMSE, the smaller the value of the RMSE indicator, the better the prediction performance of the model, and it can also indicate the degree of discreteness of a data set.
[0102] The fracture pressure data come from three blocks in a certain region: Block 2, Block 18, and Block 131, totaling 102 clusters (as shown in Table 4). Due to the differences in geological conditions and other factors within each block, directly using the traditional fracture pressure calculation formula would result in significant errors. Therefore, based on the classical fracture pressure calculation formula, 80% of the sample data from each block was selected and multivariate linear regression was used to obtain the fitting formula for these three blocks (as shown in Table 5). The remaining 20% of the sample data was used to verify the accuracy of the fitting formula. The verification and accuracy rates for Blocks 2, 18, and 131 are shown in Table 6.
[0103] Table 4 Data of some wells in the block
[0104]
[0105]
[0106] Table 5 Polynomial fitting formula
[0107] Block Fitting formula Block 2 <![CDATA[P b =-74.68σ h +72.90s H +0.80P0-666.30]]> Block 18 <![CDATA[P b =20.79σ h -19.47s H +0.80P0+201.45]]> Block 131 <![CDATA[P b =0.20σ h +0.23σ H -3.17P0+193.87]]>
[0108] Table 6 Block prediction results
[0109] Block 2 Block 18 Block 131 average Fitting accuracy 92.71% 90.16% 94.98% 92.62% Prediction accuracy 90.50% 82.87% 87.72% 87.03%
[0110] As shown in Table 5, the fitting formulas for Block 2, Block 18, and Block 131 differ, reflecting differences in geological factors and other factors within each block. Furthermore, clusters within these three blocks were validated using the fitted formulas. The best validation accuracy was achieved in Block 2, at 90.50%. The average validation accuracy across the three blocks was 87.03%.
[0111] Through the study of the main controlling factors of tight gas well fracture pressure, the eight parameters of acoustic time difference, density, Poisson's ratio, brittleness index, depth, minimum principal stress, natural gamma, and maximum principal stress were used as input samples. The fracture pressure was used as the output sample. Three models (XGBoost, GA-BP, SVR) were established for all samples, and the prediction effects of the three models were compared and analyzed. Among the complete 102 clusters of samples selected, 80% of the random samples (as shown in Table 7) were used as the training set, and the remaining 20% of the samples were used as the test set (as shown in Table 8) to verify the consistency of the prediction model. Next, these three methods were applied respectively.
[0112] Table 7 Training dataset
[0113]
[0114] Table 8 Test dataset
[0115]
[0116]
[0117] The model was trained and simulated using deep learning software. The evolutionary generations in the GA-BP neural network were set to 50, the population size was set to 10, the crossover probability was set to 0.2, the mutation probability was set to 0.1, the number of hidden layer nodes was set to 15, the maximum number of iterations was set to 1000, the error target was set to 0.001, and the Logsigmod activation function was selected from the input layer to the hidden layer and from the hidden layer to the output layer.
[0118] The GA-BP neural network model was used to train 81 clusters of training set samples and predict the rupture pressure of 21 clusters of test set. The comparison between the actual value and the predicted value is as follows: Figure 7 As shown in Figure 2, the GA-BP prediction accuracy is 84.35%. The prediction error index RMSE is 18.31MPa, the index MAPE is 19.78%, and R 2 The value of is 0.84.
[0119] The model was trained and simulated using programming software. The RBF function was selected as the kernel function in the SVR algorithm, and the kernel coefficient was set to "auto" by default. After five-fold cross-validation, the optimal penalty coefficient C for the model was 5.
[0120] SVR is used to train 81 clusters of training set samples and predict the rupture pressure of 21 clusters of test set. The comparison between the true value and the predicted value is as follows: Figure 8 As shown in the figure, the SVR prediction accuracy is 79.89%. The prediction error indicator RMSE is 20.51MPa, the indicator MAPE is 25.39%, and the R2 value is 0.74.
[0121] The model was trained and simulated using programming software. The learning rate in the XGBoost algorithm was set to 0.1, the maximum depth of the tree was set to 6, the sample sampling rate was set to 0.5, and the minimum weight sum was set to 1.
[0122] The XGBoost algorithm is used to train the 81 clusters of training set samples and predict the rupture pressure of the 21 clusters of the test set. Figure 9As shown in the figure, the XGBoost prediction accuracy is 92.05%. The prediction error indicator RMSE is 8.79MPa, the indicator MAPE is 8.65%, and the R2 value is 0.88.
[0123] The following discussion and comparative analysis of different methods, different numbers of samples, and different blocks are conducted to verify the universality of the tight oil fracture pressure prediction method based on the XGBoost algorithm.
[0124] In order to verify the superiority of the tight oil fracture pressure prediction method established in the present invention and the rationality of the main controlling factors, the main controlling factors were used as input and the XGBoost algorithm, GA-BP neural network, SVR and fitting formula were used for comparison to obtain the prediction results of different methods; Figure 10 The following is a schematic diagram for comparing the prediction results.
[0125] The comparison results show that the XGBoost algorithm is better than the GA-BP neural network, SVR and fitting formula.
[0126] Due to differences in geological conditions across different regions, different fitting formulas are used for different regions, which are not universally applicable. However, machine learning algorithms such as the XGBoost algorithm learn from data across all regions, demonstrating strong learning capabilities, universal applicability, and higher prediction accuracy for a variety of complex geological conditions.
[0127] The number of training samples affects prediction accuracy. Generally speaking, the more training samples, the better. However, in practical applications, this number must be determined based on different modeling methods. Therefore, the optimal number of samples was selected by analyzing the impact of different numbers of training samples on the accuracy of the tight oil fracture pressure prediction model established using the XGBoost algorithm. The results of the comparison of different sample numbers are shown in Table 9.
[0128] Table 9 The impact of different training samples on prediction accuracy
[0129] Number of training samples Accuracy / % 66 87.25 75 85.63 81 92.05 87 89.58
[0130] Through analysis of Table 9 combined with the previous research results, it can be seen that when the number of training samples is selected as 81, the tight oil fracture pressure prediction model established by the XGBoost algorithm has the highest accuracy.
[0131] Block 2 is located on the northwest margin of a basin, bounded by Block 131 to the north and Block 18 to the southwest. Wells in Block 131, Block 2, and Block 18 are drilled at depths of approximately 3,100 meters, 3,300 meters, and 3,800 meters, respectively. Formation parameters vary significantly with increasing depth. The entire block has a buried depth of 2,812 to 4,230 meters, oil saturation of 45% to 65%, porosity of 7.7% to 11.8%, overburden permeability of 0.02 to 0.45 mD, gravel particle size of 2 to 35 mm, lack of natural fractures, biaxial stress difference of 10 to 22 MPa, Young's modulus of 19 to 25 GPa, and Poisson's ratio of 0.2 to 0.25.
[0132] After studying the fitting formulas for Blocks 2, 18, and 131, we developed a tight oil fracture pressure prediction method based on the XGBoost algorithm and integrated the data from these three blocks. The prediction results of the four methods are shown in Table 10.
[0133] Table 10 Prediction results of blocks 2, 18, and 131
[0134] Block Block 2 Block 18 Block 131 XGBoost (comprehensive three blocks) Prediction accuracy 90.50% 82.87% 87.72% 92.05%
[0135] Table 10 shows that the XGBoost algorithm performs well in predicting the overall fracture pressure data for these three blocks, significantly reducing the associated workload. Therefore, the proposed tight oil fracture pressure prediction method is not only applicable to all three blocks but also paves the way for further application of machine learning methods to predict fracture pressure.
[0136] Factors influencing tight oil breakdown pressure were ranked by importance using random forest analysis and by correlation analysis using grey correlation analysis. The results of these two methods were combined to form a comprehensive ranking. Based on this ranking, stepwise regression was used to further select the optimal controlling factors. A machine learning prediction model for tight oil breakdown pressure was established based on these controlling factors. This model was then compared with a partition fitting formula.
[0137] In one embodiment, S103, factors affecting tight oil production are extracted from the tight oil horizontal well data, and the factors affecting tight oil production are input into a tight oil production prediction model to obtain tight oil production; wherein, the tight oil production prediction model is trained according to the following method: multiple initial models are constructed respectively through multiple prediction algorithms, a second training set and a second test set are constructed using sample data of factors affecting tight oil production, the multiple initial models are trained respectively using the second training set, and the second test set is used to test, and a model is selected according to the test results to obtain a tight oil horizontal well fracture pressure prediction model.
[0138] refer to Figure 11The specific process of extracting factors affecting tight oil production from the tight oil horizontal well data is as follows:
[0139] S1101, for the tight oil horizontal well data, using a random forest importance algorithm and a grey relational analysis algorithm to obtain all factors affecting tight oil production and perform a comprehensive ranking of importance and relevance;
[0140] S1102: By using stepwise regression analysis, factors that are not correlated with tight oil production are eliminated to obtain the main controlling factors affecting tight oil production.
[0141] Among them, the main controlling factors affecting tight oil production include at least: cluster spacing, total sand volume, permeability, number of fractures, average operation displacement, total liquid volume, average pump stop pressure, pre-fluid ratio, porosity and oil saturation.
[0142] Taking a block as an example, 15 factors affecting tight oil production were summarized based on the data of 295 wells in the block. Then, the importance of these 15 factors was ranked using random forest. The results of random forest importance ranking of all factors are as follows: Figure 12 shown.
[0143] Figure 12 Schematic diagram of the random forest importance ranking results when processing factors affecting tight oil production according to an embodiment of the present invention. Figure 12 The 10 factors with the highest scores for tight oil production obtained by random forest are total sand volume, permeability, cluster spacing, total liquid volume, number of fractures, average operation displacement, average pump stop pressure, pre-fluid ratio, oil saturation, and porosity. The remaining 5 factors have lower importance scores. The data is then dimensionlessly processed and standardized using the minimization and maximization methods. The correlation analysis of these 15 factors is then performed in combination with grey correlation analysis (e.g. Figure 13 As shown, Figure 13 Figure 2 is a schematic diagram illustrating a correlation analysis for factors affecting tight oil production according to an embodiment of the present invention. The correlations for tight oil production are as follows: cluster spacing, number of fractures, average operation displacement, total sand volume, average pump-off pressure, permeability, total fluid volume, pad fluid ratio, porosity, oil saturation, reservoir encounter rate, horizontal section length, average sand ratio, reservoir thickness, and wellbore depth.
[0144] Combining the importance ranking results of random forest and the correlation ranking results of grey correlation analysis, a comprehensive ranking of the two methods was obtained through statistical collation, and the main controlling factors affecting tight oil production were obtained. The specific ranking results are shown in Table 11.
[0145] Table 11 Comprehensive ranking results of factors affecting tight oil production
[0146]
[0147] Referring to Table 11, we can see that the 12 factors, cluster spacing, total sand volume, permeability, number of fractures, average construction displacement, total liquid volume, average pump stop pressure, pre-fluid ratio, porosity, oil saturation, oil layer conversion rate, and average sand ratio, rank high. At the same time, the relative degree trend line analysis of the two models (such as Figure 12 , Figure 13 ), the three factors of oil layer thickness, horizontal section length and wellbore depth had lower scores and rankings, so they were eliminated.
[0148] The analysis revealed that the remaining 12 geological factors had a significant impact on tight oil production, and a secondary screening was performed using stepwise regression analysis. Since there were 12 factors in the regression model, there might be factors that were not significantly correlated with production. Therefore, a backward stepwise regression method was used to select the main controlling factors, construct a full-variable model, and gradually identify insignificant factors until all factors were significantly correlated with fracture pressure. The full-variable model is:
[0149] Z=β1Y1+β2Y2+...+β 11 Y 11 +β 12 Y 12 +Y0
[0150] Among them, Z represents the burst pressure, Y i The specific meaning is shown in Table 12, β i Represents the coefficient, i=1,2,...,12.
[0151] Table 12 Definition of relevant variables
[0152]
[0153] After three backward regressions, the oil layer encounter rate and average sand ratio did not meet the criteria and were removed from the model. After final screening, 10 factors with significant impact on tight oil production were obtained: cluster spacing, total sand volume, permeability, number of fractures, average operation displacement, total liquid volume, average pump-off pressure, front fluid ratio, porosity, and oil saturation.
[0154] refer to Figure 14 The training method of the tight oil production prediction model includes:
[0155] S1401, constructing corresponding initial models using support vector machine classification algorithm (SVM), random forest classification algorithm, AdaBoos algorithm, and XGBoost algorithm;
[0156] S1402: After training and testing multiple initial models, the corresponding prediction accuracy, precision, and recall rates are obtained, and the model with the highest average of the three is selected as the tight oil production prediction model; wherein, the accuracy rate is the proportion of samples correctly predicted by the model to the total number of samples; the precision rate is the proportion of samples predicted by the model to be positive that are actually positive; and the recall rate is the proportion of samples that are actually positive that are correctly predicted by the model to be positive.
[0157] Yield prediction is a hot research topic in oil and gas field development. Accurately predicting high-yield and low-yield tight oil wells can provide effective guidance for scientific decision-making and contribute to the efficient and sustainable development of tight oil. This paper uses a binary classification method to develop a prediction method for high-yield and low-yield tight oil wells. The results of various prediction methods are compared and analyzed.
[0158] Support vector machines (SVMs) are a commonly used machine learning algorithm in data mining. Their superior generalization capabilities prevent overfitting and the local optimality problems common in traditional machine learning. Furthermore, the algorithm's complexity is calculated based on the number of support vectors, thus avoiding the curse of dimensionality. Support vector machines can be used for both classification (SVC) and regression (SVR). SVC (Support Vector Machine Classification Algorithm) is particularly effective in solving binary classification problems.
[0159] The Random Forest (RF) classification algorithm is an ensemble learning algorithm based on binary decision trees. Since the process of generating decision trees is independent, Random Forest is convenient for parallel computing when processing large sample data sets. In the classification problem of high-dimensional data, it has obvious advantages such as fast speed, high accuracy, and good stability. Its effectiveness has been verified in a large number of applications. In addition, the difference in sample subsets will reduce the correlation between each decision tree, thereby ensuring that the Random Forest algorithm has a strong generalization ability. The Random Forest algorithm has very wide applications in the fields of classification, regression, and feature screening. The application of Random Forest in feature screening is described in the above embodiment of the present invention.
[0160] The AdaBoost algorithm (Adaptive Boosting) is an ensemble learning algorithm that achieves high accuracy by integrating multiple weak classifiers with accuracy greater than random guessing (50%) to make joint decisions. Compared to various classifier training algorithms, the AdaBoost algorithm also has the function of feature screening. In addition, the classification model trained by the AdaBoost algorithm has low complexity, good real-time computational performance, strict algorithm convergence, and is not prone to overfitting. Both experimental and theoretical research have demonstrated the AdaBoost algorithm's strong generalization ability.
[0161] XGBoost (Extreme Gradient Boosting) is a boosting ensemble algorithm based on CART regression trees. XGBoost has no specific data requirements. Regardless of the size or distribution of the data sample, XGBoost demonstrates advantages such as fast computation, input data invariance, and high prediction accuracy. Therefore, it has been widely used and acclaimed in fields such as machine learning and data mining. Unlike traditional decision tree-based ensemble algorithms, XGBoost incorporates regularization terms such as tree depth and leaf node weights into its cost function. This not only controls the complexity of the prediction model but also prevents overfitting. Furthermore, XGBoost uses a second-order Taylor expansion to approximate the cost function, making the objective function approximation closer to the actual value, thereby achieving higher prediction accuracy. The objective function can determine the quality of the tree structure. In theory, all possible tree structures can be enumerated to find the optimal spanning tree, but this is difficult to achieve in practice. Therefore, XGBoost optimizes one layer of the tree at a time, using the above method to find the optimal split point, thereby gradually optimizing the optimal tree structure.
[0162] For the purpose of distinguishing and subsequent classification and optimization research based on the field data of tight oil wells in a certain region, 295 tight oil well samples were divided into two types: high-yield wells and low-yield wells according to their average daily oil production. The specific classification criteria and the number of samples in each category are shown in Table 13.
[0163] Table 13 Production classification criteria
[0164] category Sample type Production range (t / d) Label Sample size Average output (t / d) 1 Low-yield wells (0,15] 0 131 8.25 2 High-yield wells (15,61] 1 164 24.92
[0165] For the 295 tight oil well data samples, 20 samples were randomly selected from the 163 samples belonging to the high-yield well category, and 15 samples were randomly selected from the 131 samples belonging to the low-yield well category. These 35 data samples formed the test dataset for classification and yield prediction research. The remaining 260 samples served as the training dataset for building the various models. Tables 14 and 15 show partial data from the training and test sets before normalization. All samples were normalized to their maximum values before formal calculations.
[0166] Table 14 Training dataset
[0167]
[0168] Table 15 Test dataset
[0169]
[0170] The kernel function of the SVC algorithm (Support Vector Machine Classification Algorithm) is RBF, so the kernel parameter sigma and penalty factor C need to be determined. The RF algorithm needs to determine the number of classification trees and leaves. The AdaBoost algorithm needs to determine the number of iterations. The XGBoost algorithm needs to determine the depth of the tree and the gamma parameter. In order to make each algorithm reach the best state as much as possible, the important parameters are optimized. Figures 15 to 18 , which are schematic diagrams of SVC parameter optimization, RF parameter optimization, AdaBoost parameter optimization, and XGBoost parameter optimization, respectively. The results of optimizing one or two of the most important parameters for each of the four algorithms were obtained. The SVC, RF, and AdaBoost algorithms were implemented using built-in functions in deep learning software; the XGBoost algorithm was implemented using the third-party library xgboost imported into programming software. The subsample parameter, eta parameter, and iteration number were set to 0.7, 0.1, and 500, respectively. The remaining parameters not to be optimized were left as they were. Based on the optimization results, the optimal parameter values were determined, as shown in Table 16.
[0171] Table 16 Optimal values of parameters for each classification algorithm
[0172] Serial number algorithm software method Optimal parameter values 1 SVC Deep learning software svmtrain() sigma is 20, C is 5 2 RF Deep learning software TreeBagger() The number of leaves is 100 and the number of trees is 28 3 AdaBoost Deep learning software fitensemble() The number of iterations is 670 4 XGBoost Programming software xgboost.train() Gamma is 1, and the depth of the tree is 6
[0173] For the binary classification problem of high-yield and low-yield tight oil wells, let the number of correctly classified high-yield well samples in the test set be TP, the number of incorrectly classified high-yield well samples be FP, the number of incorrectly classified low-yield well samples be FN, and the number of correctly classified low-yield well samples be TN. These constitute a second-order confusion matrix by row. Accuracy, precision, recall, and F1 value are used to evaluate the specific effectiveness of the prediction method for high-yield and low-yield tight oil wells. The calculation formula is as follows:
[0174]
[0175] For the 260 tight oil well samples in the training set, prediction methods for tight oil high-yield wells and low-yield wells based on the SVC algorithm, RF algorithm, AdaBoost algorithm and XGBoost algorithm were established according to the optimal parameter values. The 35 samples in the test set were predicted, and the results are as follows: Figures 19 to 22The following diagrams show the SVC confusion matrix, the RF confusion matrix, the AdaBoost confusion matrix, and the XGBoost confusion matrix. Classification task 1 (class1) and classification task 0 (class 0) represent the two categories in the classification task, while model prediction result 1 (output 1) and model prediction result 0 (output 0) represent the predicted categories. Comparing the number of misclassifications (numbers where the predicted category differs from the actual category) in the four confusion matrices reveals that XGBoost performs best in this classification task, with the fewest misclassified samples. RF comes in second, while SVC and AdaBoost have relatively more misclassifications. The confusion matrix can be used to determine which algorithm is most suitable for tight oil production classification, providing a quantitative basis for model selection. Table 17 shows the evaluation results of the prediction methods.
[0176] Table 17 Evaluation results of prediction methods
[0177]
[0178]
[0179] In Table 17, the accuracy rate represents the ratio of all correctly predicted samples to the total number of test set samples; the precision rate represents the ratio of the number of samples correctly predicted as high-yield wells to the total number of samples predicted as high-yield wells; the recall rate represents the ratio of the number of samples correctly predicted as high-yield wells to the total number of high-yield well samples in the test set; and the F1 value is the harmonic mean of the precision rate and the recall rate. Figures 19 to 22 The confusion matrix is calculated, and the larger the value, the better the prediction method.
[0180] As shown in Table 17, among the four prediction methods for high-yield and low-yield tight oil wells, the RF algorithm and the XGBoost algorithm achieved the highest accuracy, both at 88.57%. The RF algorithm achieved the highest accuracy at 94.44%, but the XGBoost algorithm also achieved an accuracy of 90.00%, with the highest recall rate at 90.00 and the highest F1 value at 0.90. Overall, considering all the evaluation indicators, the prediction method based on the SVC algorithm performed the worst overall. The prediction method based on the AdaBoost algorithm had too few adjustable parameters, making it difficult to effectively improve accuracy through parameter adjustment. Therefore, the prediction effect of this method was not outstanding. The prediction methods based on the RF algorithm and the XGBoost algorithm each had their own advantages. In comparison, the prediction method based on the XGBoost algorithm achieved better overall results. Furthermore, it has more parameter settings and a more complete and complex theoretical basis, and can achieve further improvement in prediction effect through more intensive parameter adjustment.
[0181] For the prediction of high-yield and low-yield tight oil wells, the present invention describes the SVC classification algorithm, RF classification algorithm, AdaBoost classification algorithm, and XGBoost classification algorithm respectively, and forms a prediction method for high-yield and low-yield tight oil wells based on the results of the main controlling factors, and implements the program through deep learning software and programming software. By integrating the evaluation indicators of multiple classification problems and combining the advantages and disadvantages of the algorithm itself, the effects of each prediction method are analyzed and evaluated in detail. The prediction method for high-yield and low-yield tight oil wells based on XGBoost can meet the current accuracy requirements and has a broader application prospect.
[0182] The present invention will be described below with reference to a specific embodiment. Figure 23 , is a technical route diagram of a specific embodiment of the present invention. Figure 23 As shown in the figure, the main research contents were determined through the analysis of the main controlling factors of tight oil fracture pressure and production classification, the investigation, analysis and summary of tight oil fracture pressure prediction and production classification methods. The specific process is as follows:
[0183] S2301, collect tight oil field data.
[0184] S2302, random forest method, grayscale correlation analysis.
[0185] By utilizing the differences in tight oil field data and a large number of surveys, a preliminary selection of factors was achieved. On this basis, a comprehensive study of the factors affecting tight oil fracture pressure and production classification was conducted in combination with random forest importance ranking analysis and grey correlation analysis.
[0186] S2303, stepwise regression, main controlling factors.
[0187] The main controlling factors affecting tight oil fracture pressure and production classification were selected through stepwise regression analysis.
[0188] Specifically, by analyzing geological and fracturing data from tight oil wells in the region and reviewing relevant literature, we initially identified factors influencing the breakdown pressure and production of tight oil wells. Based on this, we used random forest analysis and grey correlation analysis to analyze importance and correlation, respectively. The results were comprehensively ranked, and stepwise regression analysis was used to optimize the factors. The resulting selection of key controlling factors was more reasonable and laid the foundation for predicting breakdown pressure and classifying production in tight oil wells.
[0189] S2304, rupture pressure prediction model.
[0190] Based on field data from tight oil reservoirs in target areas, we studied cluster-based fracture pressure prediction to address the question of whether fractures can be broken. Based on traditional fracture pressure formulas, we developed polynomial formulas for fracture pressure suitable for different reservoirs. Furthermore, we studied the key factors controlling tight oil fracture pressure and combined various classic machine learning algorithms to develop a machine learning-based tight oil fracture pressure prediction model. We compared and analyzed the prediction results to select the optimal model.
[0191] Specifically, using field data from tight oil fields in the target area, the team studied fracture prediction on a cluster-by-cluster basis to address the question of whether fracture pressure could be achieved. By optimizing eight key controlling factors, they established a machine learning model for predicting fracture pressure in tight oil wells. The XGBoost algorithm achieved the highest prediction accuracy of 92.05%, while the GA-BP neural network and SVR algorithms achieved 84.35% and 79.89%, respectively. All three methods achieved higher prediction accuracy than the 87.03% average accuracy achieved across different regions using a modified polynomial fitting formula based on traditional fracture pressure methods. The advantage of machine learning algorithms like XGBoost over polynomial fitting formulas lies not only in their higher prediction accuracy but also in their ability to learn from and adapt to complex geological environments.
[0192] S2305, tight oil production prediction model.
[0193] Based on field data from tight oil reservoirs in the target area, a well-by-well classification study was conducted. The development benefits of tight oil wells were analyzed based on the wells that had already been opened. Based on the key factors controlling tight oil production, various classic classification algorithms and ensemble learning classification algorithms were combined to study methods for predicting high- and low-yield tight oil wells. The application results were compared and analyzed to select the optimal model.
[0194] Specifically, a tight oil production classification study was conducted on a well-by-well basis, based on field data from the target region. The development benefits of tight oil wells, based on wells already opened by pressure, were analyzed. A prediction method for high- and low-yield tight oil wells was developed based on 10 empirically selected tight oil production classification factors. The highest accuracy of each method achieved 94.44% across a test set of 35 wells.
[0195] S2306, Tight oil well development strategy.
[0196] Based on the prediction results, a reasonable development strategy is provided for the tight oil wells in the target block. For example, wells with more suitable fracture pressure and production rates are selected for development first.
[0197] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and drawings, this does not require or imply that these operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0198] After introducing the method of the exemplary embodiment of the present invention, next, reference is made to Figure 24 A device for predicting fracture pressure and production classification of tight oil reservoirs according to an exemplary embodiment of the present invention is introduced.
[0199] The implementation of the device for predicting tight oil reservoir fracture pressure and production classification can be referenced to the implementation of the aforementioned method, and any repetitions will not be repeated. The terms "module" or "unit" used below may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0200] Based on the same inventive concept, the present invention also proposes a prediction device for classification of tight oil reservoir fracture pressure and production, such as Figure 24 As shown, the device includes:
[0201] Data acquisition module 2410, for collecting tight oil horizontal well data;
[0202] The tight oil horizontal well breakdown pressure prediction module 2420 is configured to extract factors influencing the breakdown pressure of the tight oil horizontal well from the tight oil horizontal well data, input the factors influencing the breakdown pressure of the tight oil horizontal well into a tight oil horizontal well breakdown pressure prediction model, and obtain the breakdown pressure of the tight oil horizontal well; wherein the tight oil horizontal well breakdown pressure prediction model is trained according to the following method: constructing multiple initial models using multiple machine learning algorithms, constructing a first training set and a first test set using sample data of factors influencing the breakdown pressure of the tight oil horizontal well, training the multiple initial models using the first training set, and testing them using the first test set, selecting a model based on the test results, and obtaining the tight oil horizontal well breakdown pressure prediction model;
[0203] The tight oil production prediction module 2430 is configured to extract factors influencing tight oil production from the tight oil horizontal well data, input the factors influencing tight oil production into a tight oil production prediction model, and obtain tight oil production. The tight oil production prediction model is trained according to the following method: constructing multiple initial models using multiple prediction algorithms, constructing a second training set and a second test set using sample data of factors influencing tight oil production, training the multiple initial models using the second training set, and testing them using the second test set. A model is selected based on the test results to obtain a tight oil horizontal well fracture pressure prediction model.
[0204] The development strategy generation module 2440 is used to classify the tight oil horizontal wells according to the tight oil production, obtain the tight oil production classification results, and generate the tight oil horizontal well development strategy according to the tight oil horizontal well fracture pressure and the tight oil production classification results.
[0205] In one embodiment, after collecting tight oil horizontal well data, the data acquisition module 2410 is further configured to:
[0206] Preprocessing the tight oil horizontal well data, filling in missing data, removing outliers and filling in missing data;
[0207] Based on the pre-processed tight oil horizontal well data, the factors affecting the fracture pressure and production of tight oil horizontal wells were preliminarily selected. The preliminary selection method was as follows: factors were preliminarily selected through manual investigation, analysis and induction; factors not involved in the blocks and factors with abnormal data were eliminated by analyzing and collating the logging, fracturing and production data of different blocks of tight oil horizontal wells.
[0208] In one embodiment, the tight oil horizontal well breakdown pressure prediction module 2420 extracts factors affecting the breakdown pressure of the tight oil horizontal well from the tight oil horizontal well data, including:
[0209] For the tight oil horizontal well data, all factors affecting the fracture pressure of the tight oil horizontal well are obtained by using the random forest importance algorithm and the grey relational analysis algorithm, and the influence is comprehensively ranked;
[0210] Through stepwise regression analysis, factors with unsatisfactory correlation with the breakdown pressure of tight oil horizontal wells were eliminated, and the main controlling factors affecting the breakdown pressure of tight oil horizontal wells were obtained.
[0211] In one embodiment, the main controlling factors affecting the fracture pressure of tight oil horizontal wells include at least: acoustic wave transit time, density, Poisson's ratio, brittleness index, depth, minimum principal stress, natural gamma and maximum principal stress.
[0212] In one embodiment, the method for training the tight oil horizontal well fracture pressure prediction model includes:
[0213] The corresponding initial models are constructed using the XGBoost algorithm, GA-BP neural network, support vector regression algorithm, and partition fitting algorithm;
[0214] After training and testing multiple initial models, the corresponding prediction accuracy was obtained, and the model with the highest prediction accuracy was selected as the fracture pressure prediction model for tight oil horizontal wells.
[0215] In one embodiment, the tight oil production prediction module 2430 extracts factors affecting tight oil production from the tight oil horizontal well data, including:
[0216] For the tight oil horizontal well data, all factors affecting tight oil production are obtained by using a random forest importance algorithm and a grey relational analysis algorithm, and are comprehensively ranked by importance and relevance;
[0217] Through stepwise regression analysis, factors with unsatisfactory correlation with tight oil production were eliminated to obtain the main controlling factors affecting tight oil production.
[0218] In one embodiment, the main controlling factors affecting tight oil production include at least: cluster spacing, total sand volume, permeability, number of fractures, average operation displacement, total liquid volume, average pump-off pressure, pre-fluid ratio, porosity and oil saturation.
[0219] In one embodiment, the training method of the tight oil production prediction model includes:
[0220] Use the support vector machine classification algorithm, random forest classification algorithm, AdaBoos algorithm, and XGBoost algorithm to build corresponding initial models respectively;
[0221] After training and testing multiple initial models respectively, the corresponding prediction accuracy, precision and recall rates are obtained, and the model with the highest average of the three is selected as the tight oil production prediction model; wherein, the accuracy rate is the proportion of samples correctly predicted by the model to the total samples; the precision rate is the proportion of samples predicted by the model to be positive that are actually positive; the recall rate is the proportion of samples that are actually positive that are correctly predicted by the model to be positive.
[0222] It should be noted that while the detailed description above mentions several modules of the apparatus for predicting tight oil reservoir fracture pressure and production classification, this division is merely exemplary and not mandatory. In practice, according to embodiments of the present invention, the features and functions of two or more modules described above may be embodied in a single module. Conversely, the features and functions of a single module described above may be further divided and embodied by multiple modules.
[0223] Based on the above invention concept, Figure 25 As shown, the present invention also proposes a computer device 2500, including a memory 2510, a processor 2520, and a computer program 2530 stored in the memory 2510 and executable on the processor 2520. When the processor 2520 executes the computer program 2530, the aforementioned method for predicting the fracture pressure and production classification of tight oil reservoirs is implemented.
[0224] Based on the aforementioned inventive concept, the present invention proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned method for predicting tight oil reservoir fracture pressure and production classification.
[0225] Based on the aforementioned inventive concept, the present invention proposes a computer program product, which includes a computer program. When the computer program is executed by a processor, a method for predicting the fracture pressure and production classification of tight oil reservoirs is implemented.
[0226] The method and device for predicting the fracture pressure and production classification of tight oil reservoirs proposed in the present invention screen the main controlling factors affecting the fracture pressure and the main controlling factors of the production classification, construct a fracture pressure prediction model based on multiple machine learning algorithms, select the model that best suits the complex geological conditions of different blocks, and further use multiple prediction algorithms to construct a production prediction model to achieve production classification prediction. It can effectively distinguish between high-yield wells and low-yield wells, provide a reasonable development strategy for tight oil horizontal wells, and provide accurate technical support for fracturing design and benefit evaluation of tight oil development.
[0227] The acquisition, storage, use, and processing of data in the technical solution of this application comply with relevant laws and regulations.
[0228] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0229] The present invention is described with reference to flowcharts and / or block diagrams of methods and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0230] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0231] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0232] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A method for predicting the fracture pressure and production classification of tight oil reservoirs, characterized in that: The method includes: Collect tight oil horizontal well data; Extracting factors affecting the breakdown pressure of the tight oil horizontal well from the tight oil horizontal well data, inputting the factors affecting the breakdown pressure of the tight oil horizontal well into a tight oil horizontal well breakdown pressure prediction model to obtain the breakdown pressure of the tight oil horizontal well; wherein the tight oil horizontal well breakdown pressure prediction model is trained according to the following method: constructing multiple initial models respectively using multiple machine learning algorithms, constructing a first training set and a first test set using sample data of factors affecting the breakdown pressure of the tight oil horizontal well, training the multiple initial models respectively using the first training set, and testing them using the first test set, selecting a model based on the test results, and obtaining the tight oil horizontal well breakdown pressure prediction model; Extracting factors affecting tight oil production from the tight oil horizontal well data, inputting the factors affecting tight oil production into a tight oil production prediction model to obtain tight oil production; wherein the tight oil production prediction model is trained according to the following method: constructing multiple initial models using multiple prediction algorithms, constructing a second training set and a second test set using sample data of factors affecting tight oil production, training the multiple initial models using the second training set, and testing them using the second test set, selecting a model based on the test results, and obtaining a tight oil horizontal well fracture pressure prediction model; The tight oil horizontal wells are classified according to the tight oil production to obtain a tight oil production classification result, and a tight oil horizontal well development strategy is generated according to the tight oil horizontal well fracture pressure and the tight oil production classification result.
2. The method for predicting tight oil reservoir fracture pressure and production classification according to claim 1, characterized in that: After acquiring the tight oil horizontal well data, the method further includes: Preprocessing the tight oil horizontal well data, filling in missing data, removing outliers and filling in missing data; Based on the pre-processed tight oil horizontal well data, the factors affecting the fracture pressure and production of tight oil horizontal wells were preliminarily selected. The preliminary selection method was as follows: factors were preliminarily selected through manual investigation, analysis and induction; factors not involved in the blocks and factors with abnormal data were eliminated by analyzing and collating the logging, fracturing and production data of different blocks of tight oil horizontal wells.
3. The method for predicting the fracture pressure and production classification of tight oil reservoirs according to claim 1, characterized in that: Factors affecting the fracture pressure of the tight oil horizontal well are extracted from the tight oil horizontal well data, including: For the tight oil horizontal well data, all factors affecting the fracture pressure of the tight oil horizontal well are obtained by using the random forest importance algorithm and the grey relational analysis algorithm, and the influence is comprehensively ranked; Through stepwise regression analysis, factors with unsatisfactory correlation with the breakdown pressure of tight oil horizontal wells were eliminated, and the main controlling factors affecting the breakdown pressure of tight oil horizontal wells were obtained.
4. The method for predicting tight oil reservoir fracture pressure and production classification according to claim 3, characterized in that: The main controlling factors affecting the fracture pressure of tight oil horizontal wells include at least: acoustic wave time difference, density, Poisson's ratio, brittleness index, depth, minimum principal stress, natural gamma and maximum principal stress.
5. The method for predicting tight oil reservoir fracture pressure and production classification according to claim 1, characterized in that: The training method of the tight oil horizontal well fracture pressure prediction model includes: The corresponding initial models are constructed using the XGBoost algorithm, GA-BP neural network, support vector regression algorithm, and partition fitting algorithm; After training and testing multiple initial models, the corresponding prediction accuracy was obtained, and the model with the highest prediction accuracy was selected as the fracture pressure prediction model for tight oil horizontal wells.
6. The method for predicting tight oil reservoir fracture pressure and production classification according to claim 1, characterized in that: Factors affecting tight oil production are extracted from the tight oil horizontal well data, including: For the tight oil horizontal well data, all factors affecting tight oil production are obtained by using a random forest importance algorithm and a grey relational analysis algorithm, and are comprehensively ranked by importance and relevance; Through stepwise regression analysis, factors with unsatisfactory correlation with tight oil production were eliminated to obtain the main controlling factors affecting tight oil production.
7. The method for predicting tight oil reservoir fracture pressure and production classification according to claim 6, characterized in that: The main controlling factors affecting tight oil production include at least: cluster spacing, total sand volume, permeability, number of fractures, average operation displacement, total liquid volume, average pump-off pressure, pre-fluid ratio, porosity and oil saturation.
8. The method for predicting tight oil reservoir fracture pressure and production classification according to claim 1, characterized in that: The training method of the tight oil production prediction model includes: Use the support vector machine classification algorithm, random forest classification algorithm, AdaBoos algorithm, and XGBoost algorithm to build corresponding initial models respectively; After training and testing multiple initial models respectively, the corresponding prediction accuracy, precision and recall rates are obtained, and the model with the highest average of the three is selected as the tight oil production prediction model; wherein, the accuracy rate is the proportion of samples correctly predicted by the model to the total samples; the precision rate is the proportion of samples predicted by the model to be positive that are actually positive; the recall rate is the proportion of samples that are actually positive that are correctly predicted by the model to be positive.
9. A prediction device for tight oil reservoir fracture pressure and production classification, characterized by: The device includes: Data acquisition module, used to collect tight oil horizontal well data; A tight oil horizontal well burst pressure prediction module is used to extract factors affecting the burst pressure of the tight oil horizontal well from the tight oil horizontal well data, input the factors affecting the burst pressure of the tight oil horizontal well into a tight oil horizontal well burst pressure prediction model, and obtain the tight oil horizontal well burst pressure; wherein the tight oil horizontal well burst pressure prediction model is trained according to the following method: multiple initial models are constructed respectively using multiple machine learning algorithms, a first training set and a first test set are constructed using sample data of factors affecting the burst pressure of the tight oil horizontal well, the multiple initial models are trained respectively using the first training set, and the models are tested using the first test set, and a model is selected according to the test results to obtain the tight oil horizontal well burst pressure prediction model; a tight oil production prediction module, configured to extract factors influencing tight oil production from the tight oil horizontal well data, input the factors influencing tight oil production into a tight oil production prediction model, and obtain tight oil production; wherein the tight oil production prediction model is trained according to the following method: constructing multiple initial models using multiple prediction algorithms, constructing a second training set and a second test set using sample data of factors influencing tight oil production, training the multiple initial models using the second training set, and testing them using the second test set, selecting a model based on the test results, and obtaining a tight oil horizontal well fracture pressure prediction model; The development strategy generation module is used to classify the tight oil horizontal wells according to the tight oil production, obtain the tight oil production classification results, and generate the tight oil horizontal well development strategy according to the tight oil horizontal well fracture pressure and the tight oil production classification results.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
12. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Method for predicting yield of conventional well of tight oil reservoir based on support vector machine
CN114320266A
Initiation and propagation control of vertical hydraulic fractures in unconsolidated and weakly cemented sediments
US20070199713A1