A method for predicting hydraulic fracture cross-layer in tight sandstone based on machine learning
Through machine learning-based methods, using well logging data analysis and deep learning models, an integrated classifier F(x) predicts hydraulic fracture penetration is solved, and the accuracy and complex geological conditions simulation problems of hydraulic fracture prediction in the existing technology are achieved, achieving more efficient fracturing design and oil and gas well output improvement.
Patent Information
- Application Number
- CN202411365922.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-29
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2044-09-29
AI Technical Summary
The prior art lacks accuracy and general applicability in the prediction of hydraulic fracture penetration, making it difficult to effectively simulate crack expansion under complex geological conditions, resulting in poor fracturing effect.
Using machine learning-based methods, through well logging data analysis and deep learning models, combined with base classifiers such as decision trees, support vector machines and logistic regression, an integrated classifier F(x) is built to predict the layer penetration of hydraulic fractures, and the Z-fraction method is used to remove outliers, design a recursive neural network to generate synthetic data, solve the problem of data loss, and use a weighted average method to improve prediction accuracy.
It improves the accuracy and reliability of hydraulic fracture penetration prediction, helps to formulate reasonable fracturing parameters, avoid fracturing failure caused by crack penetration, reduces oil field testing costs, and increases oil and gas well production.
Smart Images

Figure CN119538763B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting hydraulic fracture cross-layer in tight sandstone based on machine learning, belonging to the technical field of fracture cross-layer prediction. Background Art
[0002] In the process of oil and gas exploitation, hydraulic fracturing technology is widely used to improve the production of oil and gas wells, and the key lies in the accurate prediction and control of hydraulic fractures. However, the cross-layer behavior of hydraulic fractures is affected by various geological factors, and accurately predicting its cross-layer situation is crucial for improving oil well production and optimizing fracturing design. For example, improper fracturing parameter settings may cause fractures to longitudinally penetrate the interlayer, thus affecting the fracturing effect. Reasonably controlling the longitudinal extension of fractures can maximize the extension of fractures within the reservoir, increase the reservoir stimulation volume, and at the same time prevent the intrusion of water layer fluids due to penetrating the interlayer and avoid fracturing failure.
[0003] Current technologies mostly rely on empirical formulas or numerical simulation software, but they usually lack accuracy and generality. Early studies mainly relied on analysis to predict the fracture extension, such as the KGD and PKN models. These models are based on two-dimensional or three-dimensional elastic theory, but usually assume that the material is homogeneous and elastic, ignoring the complexity of actual geological conditions, such as the heterogeneity and anisotropy of rocks.
[0004] In addition, linear elastic fracture mechanics (LEFM) takes into account the stress intensity factor at the crack tip. However, due to its assumption of linear elasticity, it may not accurately describe the crack propagation in low-permeability coalbed methane reservoirs, which usually exhibit inelastic behaviors such as plasticity or toughness. To overcome the limitations of LEFM, scholars have introduced elastoplastic constitutive models, which can more accurately simulate the nonlinear damage evolution and crack propagation in coal seams. However, these models may require complex mathematical operations and consume a large amount of computing resources. In addition, there are various numerical simulation methods, including the finite element method (FEM), the finite difference method (FDM), the boundary element method (BEM), the displacement discontinuity method (DDM), the discrete element method (DEM), and the extended finite element method (XFEM). In addition to the above methods, Kresse proposed an unconventional fracture model (UFM) in 2013 to simulate the height propagation of cracks in natural fracture networks; Yamamoto described a three-dimensional fracture propagation model based on the finite element method (FEM) in 2004. Dahi-Taleghani used the extended finite element method (XFM) in 2011 to study the propagation of multi-strand hydraulic fracturing and explored the interaction between induced fractures and natural fractures; Wu proposed a three-dimensional non-planar crack height extension numerical model based on the displacement discontinuity method (DDM) in 2015; Guo used the cohesive zone model (CZM) in 2017 to study the interaction behavior between hydraulic fractures and interfaces under preset path conditions; Yan used the finite discrete element method (FDEM) in 2017 to construct a three-dimensional fully coupled model considering pore permeability to simulate the behavior of hydraulic fractures penetrating layers. These methods can simulate more complex geological conditions, but their accuracy depends on the selected material parameters and model complexity. Therefore, the present invention is proposed. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a method for predicting the penetration of hydraulic fractures in tight sandstone based on machine learning.
[0006] The technical solution of the present invention is as follows:
[0007] A method for predicting the penetration of hydraulic fractures in tight sandstone based on machine learning, the steps are as follows:
[0008] S1. Use well logging data to analyze and calculate formation parameters (parameters of reservoirs and barriers), including reservoir thickness, barrier thickness, formation tensile strength, and stress difference between reservoir and barrier;
[0009] S2. Collect relevant data of horizontal wells that have undergone hydraulic fracturing in the oilfield block, including formation parameters (reservoir thickness, barrier thickness, formation tensile strength, and stress difference between reservoir and barrier) affecting the penetration of hydraulic fractures, construction parameters (fracturing fluid displacement, fracturing fluid viscosity), and corresponding penetration results (1 represents penetration, 0 represents non-penetration);
[0010] S3. Remove outliers from the data collected in steps S1 and S2 using the Z-score method. In actual applications, due to reasons such as human error and instrument failure, logging data distortion or missing often occurs in some well sections, resulting in missing formation parameters. Therefore, a comprehensive logging data prediction model based on deep learning is designed to generate synthetic data, increase the integrity of the dataset, and generate a dataset after standardizing the entire data including the original data and the synthetic data;
[0011] S4. Select three base classifiers: decision tree, support vector machine, and logistic regression. For each base classifier, use the Bootstrap sampling method to generate multiple subsets from the training set, and use the Genetic Algorithm to automatically adjust the parameters of the base classifier;
[0012] S5. Calculate the average of the weights calculated by the error-optimal least squares method and the reciprocal of the mean square error method to obtain the average weight;
[0013] S6. Use weighted averaging for the three base classifiers to obtain the final integrated classifier F(x). Input the new formation parameters and construction parameters into the final integrated classifier F(x) to predict the situation of hydraulic fracture crossing layers.
[0014] According to the preference of the present invention, in step S1, the formation parameter calculation method is as follows:
[0015] S11. Determine the reservoir thickness and the interlayer thickness by calculating the rock porosity, permeability, and formation shale content V sh ;
[0016] S12. Use the formation tensile strength as the objective function, and perform function regression with the rock density and the longitudinal wave velocity logging data as independent variables to obtain the formation tensile strength;
[0017] S13. Calculate the stress difference between the reservoir and the interlayer using the principal stress values in the reservoir and the principal stress values in the interlayer in the logging data.
[0018] According to the preference of the present invention, in step S11, the formation shale content V sh is calculated by formula (2);
[0019]
[0020] In the formula: I GR is the natural gamma index (dimensionless); GR is the natural gamma logging value (API); GR max is the natural gamma logging value at the pure shale section (API); GR min is the natural gamma logging value at the pure sandstone section (API);
[0021]
[0022] Where: V sh is the shale content of the formation (dimensionless); I GR is the natural gamma ray index (dimensionless); GCUR is the Hilchie index (dimensionless), taking 3.7 for the Tertiary formation and 2.0 for the old formation;
[0023] A multiple regression prediction model is established using the measured rock porosity values, three rock porosity curves (acoustic, density, neutron), and the formation shale content and core porosity. After result verification, the porosity explained by the model is similar to the measured porosity, and the error is relatively small;
[0024]
[0025] Where: ρ is the rock density (g / cm 3 ); CNL is the compensated neutron (%); AC is the acoustic travel time (μs / m); is the rock porosity (%); V sh is the shale content of the formation (dimensionless);
[0026] A permeability interpretation model is established based on the correlation between rock porosity and permeability. The calculation model is:
[0027] K = 0.0036e 0.652φ (4)
[0028] Where: K is the permeability (mD); φ is the rock porosity (%);
[0029] When V sh is less than or equal to 0.2, the rock porosity is greater than or equal to 6% and the permeability is greater than or equal to 0.02 mD, then this layer is considered a reservoir layer, and the thickness of this layer is the reservoir thickness; if V sh the shale content is greater than 0.2, the porosity is less than 6% and the permeability is less than 0.02 mD, then this layer is considered an interlayer, and the thickness of this layer is the interlayer thickness.
[0030] According to the preference of the present invention, in step S12, taking the tensile strength as the research target, using the rock density and longitudinal wave velocity logging data as independent variables, through functional regression analysis, a multiple regression model is established, and the regression equation with the highest correlation is determined:
[0031]
[0032] Where: σ t is the formation tensile strength (MPa); ρ is the shale density (g / cm 3 ); V pis the longitudinal wave velocity (μs / m).
[0033] Preferably according to the present invention, in step S13, the calculation formula for the stress difference between the reservoir and the interlayer is:
[0034] Δσ = σ r - σ i (6)
[0035] In the formula: Δσ is the stress difference between the reservoir and the interlayer (MPa); σ r is the principal stress value in the reservoir (MPa); σ i is the principal stress value in the interlayer (MPa).
[0036] Preferably according to the present invention, step S3 is specifically:
[0037] S31. The Z - score is a statistical method for measuring the degree of dispersion of a data point relative to the mean of a data set, and can be used to detect outliers in the data set. The calculation formula of the Z - score is as follows:
[0038]
[0039] In the formula: X is the value of the data point; μ is the mean of the data set; σ is the standard deviation of the data set;
[0040] Data points with the absolute value of the Z - score greater than 3 are outliers;
[0041] S32. Construct a network including a one - dimensional RNN layer, a recursive module and a recursive jump module introducing Bi - LSTM, an attention mechanism, and an exogenous variable autoregressive (ARX) module. First, transform the logging data to be trained into a short - term depth sequence arranged recursively by depth and a long - term depth sequence arranged recursively by jump. Input the short - term depth sequence and the long - term depth sequence data into the RNN layer of the network respectively. After processing through time steps, input the output sequence into the recursive module and the recursive jump module. The ReLU activation function is used in the RNN operation, and the non - linear activation functions used in the recursive module and the recursive jump module are the Sigmoid and ReLU functions. The RNN layer initially extracts the local dependencies and short - term data features of the input logging data. The recursive module and the recursive jump module both add a dropout mechanism, which can effectively solve the problem of gradient disappearance that is prone to occur during training while memorizing the data features of the long - term depth sequence. This form of jump neurons saving data features can solve the problem of data information forgetting caused by too large a depth span in logging data. Then, the network inputs several time steps intercepted by ARX and the outputs of the two recursive modules into the fully - connected layer to estimate the missing logging data in a linear and non - linear fusion manner;
[0042] S33. Organize the processed data into a dataset for subsequent model training:
[0043] The formula for normalization processing is as follows:
[0044]
[0045] In the formula: x' is the value after normalization; x is the original data; min(x) and max(x) are the minimum and maximum values in the dataset respectively.
[0046] According to the preference of the present invention, step S4 is specifically as follows:
[0047] S41. Dataset division: According to the dataset processed in step S3, randomly divide it into a 70% training set and a 30% test set. 70% of the data is used to train the model, and 30% of the data is used to test the prediction ability of the model;
[0048] S42. Define parameters and their optimization ranges:
[0049] Decision tree: Define the maximum depth of the tree (the value range is from None to 20), the minimum number of samples for splitting (the value range is from 2 to 10), the minimum number of samples in a leaf node (the value range is from 1 to 5), the splitting criterion (optional gini or entropy), and the maximum number of features (the value range is integer, floating point number or None);
[0050] Support vector machine: Define the kernel function type (linear, poly, rbf, sigmoid), the regularization parameter (the value range is from 0.1 to 100), the kernel coefficient (the value range is from 0.1 to 10), and the degree of the polynomial kernel (the value range is from 2 to 3);
[0051] Logistic regression: Define the regularization type (L1, L2, elasticnet, none), the regularization parameter (the value range is from 0.01 to 100), the solver type (newton-cg, lbfgs, liblinear, sag, saga), the multi-class strategy (optional ovr or multinomial), and the maximum number of iterations (the value range is from 100 to 1000);
[0052] S43. Use the Bootstrap sampling method to generate multiple sub-training sets from the training set. For each sub-training set generated by Bootstrap, randomly generate a set of parameter values as the initial population. The chromosome encoding of each individual is a set of parameter values. Use the base classifier to evaluate the performance of each individual on each sub-training set generated by Bootstrap. The fitness value is based on the accuracy rate and F1 score of the model on the validation set;
[0053] S44. Select the top 20% of excellent individuals according to the fitness value to enter the next generation. Generate new offspring individuals through crossover operations, and perform random mutations on the newly generated offspring to increase the diversity of the population. Repeat the steps of selection, crossover, and mutation until the termination condition is met (reaching the maximum number of iterations of 500 times or the improvement of the fitness value is less than 5%). Select the individual with the highest fitness value from the final population, and its parameter combination is the optimal parameter of the base classifier;
[0054] Among them, the calculation formula of accuracy is as follows:
[0055]
[0056] In the formula: TP is the true positive example, that is, the number of samples correctly predicted as the positive class by the model; TN is the true negative example, that is, the number of samples correctly predicted as the negative class by the model; FP is the false positive example, that is, the number of samples wrongly predicted as the positive class by the model (type I error); FN is the false negative example,
[0057] that is, the number of samples wrongly predicted as the negative class by the model (type II error);
[0058] The calculation formula of the F1 score is as follows: [[ID=*17]]
[0059]
[0060] In the formula: Precision refers to the proportion of actually positive classes among the classes predicted as positive by the model, and the calculation formula is: Recall refers to the proportion of actually positive classes correctly predicted as positive by the model, and the calculation formula is:
[0061]
[0062] Preferably according to the present invention, step S5 is specifically:
[0063] S51. The calculation formula for optimizing the weights by the error-optimal least squares method is as follows:
[0064]
[0065] l 1min +l 2min +l 3min =1(12)
[0066] In the formula: n is the number of samples in the dataset; y i is the actual cross-cutting result; are the predicted values of the decision tree, support vector machine, and logistic regression base classifiers respectively; l 1min 、l 2min 、l 3minare the weights of the three base classifiers under the error-optimal least squares method;
[0067] S52. The weight optimization calculation formula of the reciprocal of the mean square error is as follows:
[0068]
[0069] l 1MSE +l 2MSE +l 3MSE = 1 (15)
[0070] In the formula, n is the number of samples in the dataset; y i is the actual cross-layer result; are the predicted values of the three base classifiers of decision tree, support vector machine, and logistic regression respectively; l 1MSE 、l 2MSE 、l 3MSE are the weights of the three base classifiers under the reciprocal of the mean square error method;
[0071] S53. Calculate the average weight:
[0072]
[0073] In the formula, l1, l2, and l3 are the average weights of the three base classifiers respectively; l 1min 、l 2min 、l 3min are the weights of the three base classifiers under the error-optimal least squares method respectively; l 1MSE 、l 2MSE 、l 3MSE are the weights of the three base classifiers under the reciprocal of the mean square error method.
[0074] According to the preference of the present invention, in step S6, the weighted average method is:
[0075] S61. For each sample x, calculate the weighted average of the prediction results of all base classifiers. The formula is as follows:
[0076]
[0077] In the formula: P(x) represents the predicted probability that the sample x belongs to the hydraulic fracture cross-layer; α i is the weight of the i-th base classifier; h i (x)
[0078] is the prediction result of the i-th base classifier for the sample x;
[0079] S61. According to the prediction result of the weighted average, select the category with the highest prediction probability as the final prediction, that is:
[0080]
[0081] Wherein: is the final prediction result.
[0082] The beneficial effects of the present invention are as follows:
[0083] By comprehensively utilizing well logging data analysis, data preprocessing, base classifier optimization, and ensemble learning techniques, the present invention effectively improves the accuracy and reliability of hydraulic fracture cross-layer prediction. Specifically, the method first determines the key parameters of reservoirs and interlayers, such as thickness, tensile strength, and stress difference, through well logging data analysis; then, by collecting actual hydraulic fracturing data of tight reservoirs, a data set containing formation parameters and construction parameters is constructed, and the Z-score method is used to remove outliers. Aiming at the problem of missing data, a well logging data prediction model based on a recurrent neural network is designed. By combining short-term and long-term depth sequences, synthetic data is generated to enhance the integrity of the data set, effectively solving the problem of missing well logging data information. On this basis, decision trees, support vector machines, and logistic regression are used as base classifiers, and weights are calculated by the error-optimal least squares method and the reciprocal of the mean square error method. Finally, the weighted average method + is used to obtain the ensemble classifier F(x) for predicting the cross-layer situation of hydraulic fractures. The implementation steps of the method are clear and easy to operate, and can quickly perform cross-layer prediction on the horizontal wells of hydraulic fracturing in a specific block, providing a scientific basis for formulating reasonable fracturing parameters, helping to improve the production of oil and gas wells, avoiding fracturing failure caused by fracture cross-layer at the same time, reducing the oilfield test cost, and having important practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 is the flow schematic diagram of the present invention;
[0085] Figure 2 is the comprehensive well logging interpretation diagram of the key cross-layer mechanical parameter analysis of the horizontal well section in the embodiment of the present invention;
[0086] Figure 3 is the data diagram of the relationship between the single-well interlayer thickness and the fracture height in the embodiment of the present invention;
[0087] Figure 4 is the data diagram of the relationship between the single-well reservoir thickness and the fracture height in the embodiment of the present invention;
[0088] Figure 5 is the statistical chart of the cross-layer stress-related data of the tight sandstone reservoir in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0089] The present invention will be further described below by way of examples in conjunction with the drawings, but not limited thereto.
[0090] Example 1:
[0091] As Figure 1 shown, this embodiment provides a method for predicting hydraulic fracture cross-layer in tight sandstone based on machine learning, and the steps are as follows:
[0092] S1. Statistically analyze the formation parameters and construction parameter data of 17 wells in the oilfield work area. As shown in the logging data graph Figure 2 shown, use logging data to analyze and calculate formation parameters (parameters of reservoir and interlayer), including reservoir thickness, interlayer thickness, formation tensile strength, and stress difference between reservoir and interlayer;
[0093] The calculation method of formation parameters is as follows:
[0094] S11. Determine the reservoir thickness and interlayer thickness by calculating the rock porosity, permeability, and formation shale content V sh ;
[0095] The formation shale content V sh is calculated by Equation (2);
[0096]
[0097] In the formula: I GR is the natural gamma index (dimensionless); GR is the natural gamma logging value (API); GR max is the natural gamma logging value at the pure shale location (API); GR min is the natural gamma logging value at the pure sandstone location (API);
[0098]
[0099] In the formula: V sh is the formation shale content (dimensionless); I GR is the natural gamma index (dimensionless); GCUR is the Hilchie index (dimensionless), taking 3.7 for the Tertiary formation and 2.0 for the old formation;
[0100] The rock porosity is obtained using logging interpretation data, and the calculation formula is as follows:
[0101] TOC = 40.6 - 15.6 * DEN + 0.013 * GR + 0.009 * CNL (19)
[0102] φ = 0.32 * AC - 0.55 * DEN + 0.84 * TOC - 0.48 * I GR + 8.67 (20)
[0103] In the formula: I GR is the natural gamma index (dimensionless); TOC is the organic carbon content (dimensionless); DEN is the rock density (g / cm 3); GR is the natural gamma-ray log value (API) of the estimated interval; CNL is the compensated neutron (%); AC is the acoustic travel time (μs / m); φ is the rock porosity (%);
[0104] Based on the correlation between rock porosity and permeability, a permeability interpretation model is established, and the calculation model is:
[0105] K = 0.0036e 0.652φ (4)
[0106] In the formula: K is the permeability (mD); φ is the rock porosity (%);
[0107] When V sh is less than or equal to 0.2, the rock porosity is greater than or equal to 6% and the permeability is greater than or equal to 0.02 mD, then this layer is considered a reservoir layer, and the thickness of this layer is the reservoir thickness; if V sh the shale content is greater than 0.2, the porosity is less than 6% and the permeability is less than 0.02 mD, then this layer is considered an interlayer, and the thickness of this layer is the interlayer thickness;
[0108] S12. The formula for calculating the formation tensile strength is:
[0109] σ t = 14.91ln(V p ) + 22.41ln(ρ) - 140.2 (5)
[0110] In the formula: σ t is the formation tensile strength (MPa); ρ is the shale density (g / cm 3 ); V p is the longitudinal wave velocity (μs / m).
[0111] S13. Calculate the stress difference between the reservoir and the interlayer using the principal stress values in the reservoir and the interlayer in the logging data;
[0112] The formula for calculating the stress difference between the reservoir and the interlayer is:
[0113] Δσ = σ r - σ i (6)
[0114] In the formula: Δσ is the stress difference between the reservoir and the interlayer (MPa); σ r is the principal stress value in the reservoir (MPa); σ i is the principal stress value in the interlayer (MPa).
[0115] S2. Collect the relevant data of horizontal wells that have undergone hydraulic fracturing in the oilfield block, including formation parameters (reservoir thickness, interlayer thickness, formation tensile strength, and stress difference between reservoir and interlayer) affecting the penetration of hydraulic fractures, construction parameters (fracturing fluid displacement, fracturing fluid viscosity), and corresponding penetration results (1 represents penetration, 0 represents non-penetration);
[0116] S3. Use the Z-score method to remove outliers from the data collected in steps S1 and S2. In actual applications, due to reasons such as human error and instrument failure, the logging data of some well sections are often distorted or missing, resulting in the lack of formation parameters. Therefore, a comprehensive logging data prediction model based on deep learning is designed to generate synthetic data to increase the integrity of the dataset. After standardizing the entire data including the original data and synthetic data, a dataset is generated. Specifically;
[0117] S31. The calculation formula of the Z-score is as follows:
[0118]
[0119] Where: X is the value of the data point; μ is the average value of the dataset; σ is the standard deviation of the dataset;
[0120] Data points with the absolute value of the Z-score greater than 3 are outliers;
[0121] S32. Construct a network that includes a one-dimensional RNN layer, a recursive module with Bi-LSTM introduced, a recursive jump module, an attention mechanism, and an exogenous variable autoregressive (ARX) module. First, transform the logging data to be trained into short-term depth sequences arranged recursively by depth and long-term depth sequences arranged recursively by jump. Input the short-term depth sequence and long-term depth sequence data into the RNN layer of the network respectively. After processing through time steps, input the output sequences into the recursive module and the recursive jump module. The ReLU activation function is used in the RNN operation, and the non-linear activation functions used in the recursive module and the recursive jump module are the Sigmoid and ReLU functions. The RNN layer initially extracts the local dependencies and short-term data features of the input logging data. The dropout mechanism is added to both the recursive module and the recursive jump module, which can effectively solve the problem of gradient disappearance that is prone to occur during training while memorizing the data features of the long-term depth sequence. This form of jump neurons preserving data features can solve the problem of data information forgetting caused by excessive depth span in logging data. Then, the network passes several time steps intercepted by the ARX and the outputs of the two recursive modules into the fully connected layer to estimate the missing logging data in a linear and non-linear fusion manner;
[0122] S33. Organize the processed data into a dataset for subsequent model training:
[0123] The formula for standardization is as follows:
[0124]
[0125] Where: x' is the value after standardization; x is the original data; min(x) and max(x) are the minimum and maximum values in the dataset respectively;
[0126] S4. Select three base classifiers: decision tree, support vector machine, and logistic regression. For each base classifier, use the Bootstrap sampling method to generate multiple subsets from the training set, and use the Genetic Algorithm to automatically adjust the parameters of the base classifier. Specifically;
[0127] S41. Dataset division: According to the dataset processed in step S3, randomly divide it into a 70% training set and a 30% test set. 70% of the data is used to train the model, and 30% of the data is used to test the prediction ability of the model;
[0128] S42. Define the parameters and their optimization ranges:
[0129] Decision tree: Define the maximum depth of the tree (value range from None to 20), the minimum number of samples for splitting (value range from 2 to 10), the minimum number of samples in a leaf node (value range from 1 to 5), the splitting criterion (optional gini or entropy), and the maximum number of features (value range can be integer, floating-point number, or None);
[0130] Support vector machine: Define the kernel function type (linear, poly, rbf, sigmoid), the regularization parameter (value range from 0.1 to 100), the kernel coefficient (value range from 0.1 to 10), and the degree of the polynomial kernel (value range from 2 to 3);
[0131] [[ID=2,4]]Logistic regression: Define the regularization type (L1, L2, elasticnet, none), the regularization parameter (value range from 0.01 to 100), the solver type (newton-cg, lbfgs, liblinear, sag, saga), the multi-class strategy (optional ovr or multinomial), and the maximum number of iterations (value range from 100 to 1000);
[0132] S43. Use the Bootstrap sampling method to generate multiple sub-training sets from the training set. For each sub-training set generated by Bootstrap, randomly generate a set of parameter values as the initial population. The chromosome encoding of each individual is a set of parameter values. Use the base classifier to evaluate the performance of each individual on each sub-training set generated by Bootstrap. The fitness value is based on the accuracy and F1 score of the model on the validation set;
[0133] S44. Select the top 20% of the excellent individuals according to the fitness value to enter the next generation. Generate new offspring individuals through crossover operations, and perform random mutations on the newly generated offspring to increase the diversity of the population. Repeat the steps of selection, crossover, and mutation until the termination condition is met (reaching the maximum number of iterations of 500 times or the improvement of the fitness value is less than 5%). Select the individual with the highest fitness value from the final population, and its parameter combination is the optimal parameter of the base classifier;
[0134] Among them, the calculation formula of accuracy is as follows:
[0135]
[0136] In the formula: TP is the true positive example, that is, the number of samples correctly predicted as the positive class by the model; TN is the true negative example, that is, the number of samples correctly predicted as the negative class by the model; FP is the false positive example, that is, the number of samples wrongly predicted as the positive class by the model (type I error); FN is the false negative example, that is, the number of samples wrongly predicted as the negative class by the model (type II error);
[0137] The calculation formula of the F1 score is as follows:
[0138]
[0139] In the formula: Precision refers to the proportion of actual positive classes among the samples predicted as positive classes by the model, and the calculation formula is: Recall refers to the proportion of actual positive classes correctly predicted as positive classes by the model, and the calculation formula is:
[0140] S5. Calculate the average of the weights calculated by the error-optimal least squares method and the reciprocal of the mean square error method to obtain the average weight. Specifically;
[0141] S51. The weight optimization calculation formula of the error-optimal least squares method is as follows:
[0142]
[0143] l 1min +l 2min +l 3min =1 (12)
[0144] In the formula: n is the number of samples in the dataset; y i is the actual cross-layer result; are the predicted values of the decision tree, support vector machine, and logistic regression, the three base classifiers respectively; l 1min 、l 2min 、l 3minis the weight of the three base classifiers under the error-optimal least squares method;
[0145] S52. The weight optimization calculation formula of the reciprocal of the mean square error is as follows:
[0146]
[0147] l 1MSE +l 2MSE +l 3MSE = 1 (15)
[0148] In the formula, n is the number of samples in the dataset; y i is the actual cross-layer result; are the predicted values of the three base classifiers of decision tree, support vector machine, and logistic regression respectively; l 1MSE 、l 2MSE 、l 3MSE are the weights of the three base classifiers under the reciprocal of the mean square error method;
[0149] S53. Calculate the average weight:
[0150]
[0151] In the formula, l1, l2, and l3 are the average weights of the three base classifiers respectively; l 1min 、l 2min 、l 3min are the weights of the three base classifiers under the error-optimal least squares method respectively; l 1MSE 、l 2MSE 、l 3MSE are the weights of the three base classifiers under the reciprocal of the mean square error method.
[0152] S6. Weightedly average the 3 base classifiers to obtain the final integrated classifier F(x);
[0153] The weighted average method is:
[0154] S61. For each sample x, calculate the weighted average of the prediction results of all base classifiers. The formula is as follows:
[0155]
[0156] In the formula: P(x) represents the predicted probability that the sample x belongs to the hydraulic fracture cross-layer; α i is the weight of the i-th base classifier; h i (x) is the prediction result of the i-th base classifier for the sample x;
[0157] S61. According to the prediction result of the weighted average, select the category with the highest prediction probability as the final prediction, that is:
[0158]
[0159] In the formula: is the final prediction result.
[0160] Based on the reservoir thickness, interlayer thickness, formation tensile strength, reservoir-interlayer stress difference, fracturing fluid displacement, and fracturing fluid viscosity data of five wells X1, X2, X3, X4, and X5, the final integrated classifier F(x) is used to predict whether the hydraulic fractures in these five wells penetrate the layers. The above are only some preferred embodiments of the present invention. Any person skilled in the art may modify the above-described technical solutions or modify them into equivalent technical solutions. Therefore, the corresponding simple modifications or equivalent transformations made according to the technical solutions of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A method for predicting hydraulic fracture cross-layer in tight sandstone based on machine learning, characterized in that, The steps are as follows: S1. Analyze and calculate formation parameters using logging data, including reservoir thickness, interlayer thickness, formation tensile strength, and reservoir-interlayer stress difference. The calculation methods of formation parameters are as follows: S11. Determine the reservoir thickness and the thickness of the interlayer by calculating the rock porosity, permeability, and the formation shale content V sh S12. Take the formation tensile strength as the objective function, and perform function regression with rock density and P-wave velocity logging data as independent variables to obtain the formation tensile strength; S13. Calculate the reservoir-interlayer stress difference using the principal stress values in the reservoir and the principal stress values in the interlayer in the logging data; S2. Collect relevant data of horizontal wells that have undergone hydraulic fracturing, including formation parameters, construction parameters, and corresponding cross-layer results that affect hydraulic fracture cross-layer; S3. Use the Z-score method to remove outliers from the data collected in steps S1 and S2, design a comprehensive logging data prediction model based on deep learning to generate synthetic data, increase the integrity of the data set, and generate a data set after normalizing the entire data including the original data and the synthetic data; S4. Select three base classifiers: decision tree, support vector machine, and logistic regression. For each base classifier, use the Bootstrap sampling method to generate multiple subsets from the training set, and use the genetic algorithm to automatically adjust the parameters of the base classifier; S5. Calculate the average of the weights calculated by the error-optimal least squares method and the reciprocal of the mean square error method to obtain the average weight; S6. Weight-average the three base classifiers to obtain the final integrated classifier F(x). Input the new formation parameters and construction parameters into the final integrated classifier F(x) to predict the situation of hydraulic fracture cross-layer.
2. The method for predicting cross-layer hydraulic fractures in tight sandstone based on machine learning according to claim 1, wherein In step S11, the shale content V of the formation sh is calculated by formula (2); Where: I GR is the natural gamma ray index; GR is the natural gamma ray log value; GR max is the natural gamma ray log value at the pure shale; GR min is the natural gamma ray log value at the pure sandstone; Where: V sh is the shale content of the formation; I GR is the natural gamma ray index; GCUR is the Hilchie index; Where: ρ is the rock density; CNL is the compensated neutron; AC is the acoustic travel time; is the rock porosity; V sh is the formation shale content; Establish a permeability interpretation model based on the correlation between rock porosity and permeability. The calculation model is: K = 0.0036e 0.652φ (4) In the formula: K is the permeability (mD); φ is the rock porosity (%); V sh Layers with a value of V less than or equal to 0.2, a rock porosity greater than or equal to 6%, and a permeability greater than or equal to 0.02 mD are reservoirs, and the thickness of this layer is the reservoir thickness; V sh Layers with a shale content greater than 0.2, a porosity less than 6%, and a permeability less than 0.02 mD are barriers, and the thickness of this layer is the barrier thickness.
3. The method for predicting the penetration of hydraulic fractures in tight sandstone based on machine learning according to claim 2, characterized in that, In step S12, taking the tensile strength as the research objective, using rock density and P-wave velocity logging data as independent variables, a multiple regression model is established through function regression analysis: Where: σ t is the formation tensile strength; ρ is the shale density; V p is the longitudinal wave velocity.
4. The method for predicting hydraulic fracture cross-layer in tight sandstone based on machine learning according to claim 3, characterized in that, In step S13, the calculation formula for the reservoir-interlayer stress difference is: Δσ = σ r -σ i (6) Where: Δσ is the stress difference between the reservoir and the barrier layer; σ r is the principal stress value in the reservoir; x i is the principal stress value in the barrier layer.
5. The method for predicting cross-layer hydraulic fractures in tight sandstone based on machine learning according to claim 4, characterized in that, Step S3 is specifically: S31. The calculation formula of the Z-score is as follows: In the formula: X is the value of the data point; μ is the mean of the data set; σ is the standard deviation of the data set; Data points with an absolute value of the Z-score greater than 3 are outliers; S32. Construct a network that includes a one-dimensional RNN layer, a recursive module with an introduced Bi-LSTM and a recursive skip module, an attention mechanism, and an exogenous variable autoregressive (ARX) module. First, transform the logging data to be trained into a short-term depth sequence arranged recursively by depth and a long-term depth sequence arranged by recursive skip. Input the short-term depth sequence and the long-term depth sequence data into the RNN layer of the network respectively. After processing through time steps, input the output sequence into the recursive module and the recursive skip module. The ReLU activation function is used in the RNN operation, and the non-linear activation functions used in the recursive module and the recursive skip module are the Sigmoid and ReLU functions. The RNN layer initially extracts the local dependencies and short-term data features of the input logging data. The dropout mechanism is added to both the recursive module and the recursive skip module, which can solve the problem of gradient disappearance that is prone to occur during training while memorizing the data features of the long-term depth sequence. This form of jump neurons preserving data features solves the problem of data information forgetting caused by too large a depth span in logging data. Then, the network passes several time steps intercepted by ARX and the outputs of the two recursive modules into the fully connected layer to estimate the missing logging data in a linear and non-linear fusion manner; S33. Organize the processed data into a dataset for subsequent model training: The formula for normalization processing is as follows: In the formula: x' is the value after normalization; x is the original data; min(x) and max(x) are the minimum and maximum values in the dataset respectively.
6. The method for predicting the hydraulic fracture cross-layer of tight sandstone based on machine learning according to claim 5, wherein Step S4 is specifically as follows: S41. Dataset division: According to the dataset processed in step S3, randomly divide it into a 70% training set and a 30% test set. 70% of the data is used to train the model, and 30% of the data is used to test the prediction ability of the model; S42. Define the parameters and their optimization ranges: Decision tree: Define the maximum depth of the tree, the minimum number of samples for splitting, the minimum number of samples for leaf nodes, the splitting criterion, and the maximum number of features; Support vector machine: Define the kernel function type, the regularization parameter, the kernel coefficient, and the degree of the polynomial kernel; Logistic regression: Define the regularization type, the regularization parameter, the solver type, the multi-class strategy, and the maximum number of iterations; S43. Use the Bootstrap sampling method to generate multiple sub-training sets from the training set. For each sub-training set generated by Bootstrap, randomly generate a set of parameter values as the initial population. The chromosome encoding of each individual is a set of parameter values. Use the base classifier to evaluate the performance of each individual on each sub-training set generated by Bootstrap. The fitness value is based on the accuracy and F1 score of the model on the validation set; S44. Select the top 20% of the excellent individuals based on the fitness value to enter the next generation. Generate new offspring individuals through crossover operations, and randomly mutate the newly generated offspring to increase the diversity of the population. Repeat the selection, crossover, and mutation steps until the termination condition is met. Select the individual with the highest fitness value from the final population, and its parameter combination is the optimal parameter of the base classifier; Among them, the formula for calculating the accuracy is as follows: Where: TP is the true positive, that is, the number of samples correctly predicted as positive by the model; TN is the true negative, that is, the number of samples correctly predicted as negative by the model; FP is the false positive, that is, the number of samples incorrectly predicted as positive by the model; FN is the false negative, that is, the number of samples incorrectly predicted as negative by the model; [[ID=~1]]The calculation formula of the F1 score is as follows: Where: Precision refers to the proportion of actual positive classes among the positive classes predicted by the model, and the calculation formula is: Recall refers to the proportion of positive classes correctly predicted by the model among the actual positive classes, and the calculation formula is:
7. The method for predicting the hydraulic fracture penetration of tight sandstone based on machine learning according to claim 6, wherein Specifically for step S5: S51. The calculation formula for optimizing the weights by the error-optimal least squares method is as follows: l 1min +l 2min +l 3min =1 (12) Where: n is the number of samples in the dataset; y i is the actual result of cross - layer; are the predicted values of the decision tree, support vector machine, and logistic regression, respectively, which are the three base classifiers; l 1min and l 2min and l 3min are the weights of the three base classifiers under the error - optimal least - squares method; S52. The calculation formula for optimizing the weights by the reciprocal of the mean square error method is as follows: where n is the number of samples in the dataset; y i is the actual result of cross - layer; are the predicted values of the decision tree, support vector machine, and logistic regression, respectively, which are three base classifiers; l 1MSE and l 2MSE and l 3MSE are the weights of the three base classifiers under the reciprocal of mean squared error method; S53. Calculate the average weight: Wherein, l1, l2, and l3 are the average weights of the three base classifiers; l 1min , l 2min , l 3min are the weights of the three base classifiers under the error-optimal least squares method; l 1MSE , l 2MSE , l 3MSE are the weights of the three base classifiers under the reciprocal mean square error method.
8. The method for predicting the penetration of hydraulic fractures in tight sandstone based on machine learning according to claim 7, wherein In step S6, the weighted average method is: S61. For each sample x, calculate the weighted average of the prediction results of all base classifiers, and the formula is as follows: where: P(x) represents the predicted probability that the sample x belongs to the hydraulic fracture penetrating the layer; α i is the weight of the i-th base classifier; h i (x) is the prediction result of the i-th base classifier for the sample x; S61. According to the prediction result of the weighted average, select the class with the highest prediction probability as the final prediction, that is: In the formula: is the final prediction result.
Citation Information
Patent Citations
Logging reservoir identification and prediction method based on machine learning
CN113642772A
Water-rich tight sandstone fluid discrimination method based on cost-sensitive learning
CN117150390A