A comprehensive analysis method for main control factors of fracturing data aiming at yield

By combining data-driven methods with multiple analytical approaches, we screened and verified the main controlling factors of fracturing data, which solved the problems of inconsistent results and reliance on human experience in traditional fracturing analysis, and achieved accurate evaluation of production control factors and optimization of fracturing schemes.

CN119849994BActive Publication Date: 2025-11-18PETROCHINA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311345539.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2025-11-18
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify the main controlling factors affecting fracturing performance. Traditional analysis methods yield inconsistent results and rely on human experience, leading to large errors in fracturing performance and production capacity prediction, as well as high verification costs.

Method used

A data-driven approach was adopted, which involved parameter screening, modeling analysis, index evaluation, and fuzzy mathematical comprehensive analysis. Combined with grey relational analysis, CatBoost ensemble learning, and SHAP evaluation, a comprehensive analysis process for the main controlling factors of fracturing data was constructed to screen and verify the main controlling factors of production.

Benefits of technology

It has achieved a relatively objective and accurate evaluation of the main factors controlling production, provided a reference for optimizing fracturing schemes, provided a basis for the development of unconventional reservoirs, and reduced the impact of human factors and verification costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119849994B_ABST
    Figure CN119849994B_ABST
Patent Text Reader

Abstract

The application provides a comprehensive analysis method for yield-oriented fracturing data main control factors, and belongs to the technical field of unconventional oil and gas reservoir exploitation. The application obtains relatively objective and accurate yield main control factor evaluation results through parameter screening, modeling analysis, index evaluation, comprehensive analysis and effect verification, and provides a reference basis for unconventional oil reservoir development. The application provides a comprehensive analysis method and process for yield-oriented fracturing data main control factors based on data driving, and tries to excavate internal relations of data as much as possible, and fully considers results of various analysis methods through mathematical means, so that comprehensive evaluation, judgment and verification of yield main control factors are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unconventional oil and gas reservoir development technology, and in particular to a comprehensive analysis method for the main controlling factors of fracturing data for production. Background Technology

[0002] Currently, unconventional oil reservoirs are characterized by well-developed microfractures, strong reservoir heterogeneity, narrow pore throats, complex seepage mechanisms, uneven fracturing stimulation, and poor production prediction. Traditional physical-driven models are insufficient to accurately represent the actual field conditions and have high requirements for physical mechanisms, making them unsuitable for guiding field fracturing operations. The development of unconventional oil and gas resources such as coalbed methane wells and sandstone / conglomerate reservoirs requires the use of horizontal well volumetric fracturing technology to increase production capacity and recovery rates.

[0003] However, fracturing effectiveness is influenced by various geological and engineering factors, such as coal seam structure, reservoir parameters, perforation parameters, and proppant dosage. Therefore, fully utilizing fracturing data, accurately identifying the main controlling factors affecting fracturing effectiveness, and optimizing fracturing schemes are key to improving fracturing efficiency and economy.

[0004] Analysis of the main controlling factors of fracturing data production refers to using fracturing data, through various mathematical methods and models, to identify the main parameters affecting production, and to evaluate and optimize them in order to improve the fracturing effect and production capacity of unconventional oil and gas wells.

[0005] Currently, various methods exist for evaluating feature importance, but the results are inconsistent. For example, methods based on grey relational analysis, linear correlation analysis, lasso regression, and random forest model weights produce different importance rankings for the same dataset and target parameters, with each method emphasizing different aspects of importance. Subsequent analysis relies heavily on manual verification and selection. Furthermore, the input parameters are not subjected to collinearity filtering and correlation testing, leading to errors in the actual controlling factors and consequently, erroneous analysis results. In indicator weight evaluation methods, fuzzy mathematics is often used, or weights are directly applied to the raw data, incorrectly linking parameter size and distribution to importance; alternatively, the indicator system is determined through qualitative analysis, heavily influenced by human experience. Finally, the analyzed controlling factors are often difficult to verify, and conducting field experiments is prohibitively expensive. Summary of the Invention

[0006] To address the problems existing in the prior art, this invention provides a comprehensive analysis method for the main controlling factors of fracturing data related to production. Through parameter selection, modeling analysis, index evaluation, comprehensive analysis, and effect verification, it obtains relatively objective and accurate evaluation results of the main controlling factors of production, providing a reference for the development of unconventional reservoirs. This invention provides a data-driven comprehensive analysis method for the main controlling factors of fracturing data related to production, maximizing the exploration of internal relationships within the data and fully considering the results of multiple analysis methods through mathematical means, thereby achieving a comprehensive evaluation, judgment, and verification of the main controlling factors of production.

[0007] This invention provides a comprehensive analysis method for the main controlling factors of fracturing data for production output, comprising the following steps:

[0008] Step S1: Preprocess the fracturing data of field production: Perform data missing and data anomaly processing on the fracturing data of field production to obtain the processed fracturing data.

[0009] Step S2 involves feature selection and transformation of the geological engineering parameters that affect fracturing data: calculating the collinearity and correlation between geological engineering parameters, screening geological engineering parameters, and extracting features of geological engineering parameters through data dimensionality reduction and data augmentation of the processed fracturing data.

[0010] Step S3: Use multiple evaluation methods to evaluate the importance of fracturing parameters for target production: Use multiple importance evaluation methods to evaluate the characteristics of target production parameters, and establish an importance evaluation matrix for geological engineering parameters by integrating the importance evaluation results of multiple importance evaluation methods. The importance evaluation methods include grey relational analysis, CatBoost ensemble learning, and SHAP evaluation.

[0011] Step S4: Perform a comprehensive evaluation of the main control factors of the evaluation matrix using fuzzy mathematics: determine the weights of the feature vectors of each importance evaluation method in the importance evaluation matrix based on the entropy weight method, calculate the final feature importance score by weighting, and rank them to obtain the main control factors of output.

[0012] Preferably, in step S1, the missing data is filled using the KNN nearest neighbor prediction method.

[0013] Preferably, in step S2, feature selection and transformation are performed on the geological and engineering parameters affecting the fracturing data, including the following steps:

[0014] Collinearity analysis was performed on geological engineering parameters using VIF analysis, and correlation analysis was performed on the parameters using principal component analysis.

[0015] The VIF analysis method is used to reduce the dimensionality of geological engineering parameters, thereby removing collinearity while preserving data interpretability.

[0016] Preferably, the variance expansion factor (VIF) is used to determine collinearity between geological engineering parameters by pairwise combination. The VIF calculation method is as follows:

[0017]

[0018] For parameter combinations with a variance inflation factor (VIF) greater than 10 exhibiting multicollinearity, principal component replacement is performed, and dimensionality reduction is achieved using principal component analysis (PCA), retaining only one principal component. The principal component calculation method is as follows:

[0019] And satisfy

[0020] in,

[0021] R j 2 It is the coefficient of determination obtained by linear regression on the other k-1 geological engineering parameters when the j-th geological engineering parameter is used as the explained variable, where k is the total number of geological engineering parameters participating in the dimensionality reduction;

[0022] Where Z1 is the first principal component, X j It is the dataset of the j-th geological engineering parameter, φ 1j It is the weight of the first principal component of the j-th geological engineering parameter;

[0023] n is the total number of sample wells;

[0024] m represents the total number of geological engineering parameters;

[0025] j is the sequence number of the geological engineering parameter;

[0026] i is the sample well number;

[0027] x ij Let be the j-th geological engineering parameter of the i-th well.

[0028] Principal Component Analysis (PCA) is a commonly used data analysis method that can reduce high-dimensional data to a lower-dimensional space while preserving the main characteristic components. The mathematical principle of PCA is to find the directions of maximum variance in the data using covariance matrix and eigenvalue analysis, and then project the data onto these axes to achieve dimensionality reduction. The advantages of PCA are that it can reduce data complexity, reduce computational overhead, remove noise, and extract important information. The disadvantages of PCA are that it is not necessarily applicable to all data, may lose useful information, and can only handle numerical data. Figure 4As shown, PCA maps two-dimensional data scatter points to one-dimensional axes while preserving the distribution information of the scatter points as much as possible.

[0029] Preferably, for parameters with weak collinearity, indicating strong independence, data is augmented through cross-multiplication, increasing data complexity. A two-dimensional transformation converts X1,X2 to X1,X2,X1 2 X2 2 There are five vectors, X1 and X2.

[0030] The dataset after dimensionality reduction and augmentation is represented as {X} 改造段长 Z 主成分1 ,…,X 用液强度}

[0031] Preferably, in step S3, the geological engineering parameter dataset after dimensionality reduction and augmentation is used to evaluate the importance of the characteristics of the target yield parameter using multiple importance evaluation methods. The importance evaluation matrix of the geological engineering parameter is established by combining the importance evaluation results of multiple importance evaluation methods. The target yield parameter is used as the dependent variable and other geological engineering parameters are used as independent variables. The characteristic importance score of each geological engineering parameter to the target yield parameter is calculated. Finally, the importance evaluation matrix E is constructed.

[0032] Preferably, in step S4, the analysis results of the importance evaluation matrix E are comprehensively evaluated using fuzzy mathematics, the final parameter importance score for the target parameter is calculated and ranked, and the main control factors of output are analyzed and judged.

[0033] Preferably, in step S4, the importance evaluation matrix E is standardized to obtain a fuzzy evaluation matrix. The entropy weight method in fuzzy mathematics is used to calculate the weights of each importance evaluation method and the weights of each geological engineering parameter to obtain a comprehensive fuzzy evaluation vector. The final importance score of each geological engineering parameter is obtained and sorted from high to low to determine the main control factors of production.

[0034] Preferably, when standardizing the importance evaluation matrix E, for any geological engineering parameter, based on the principle of whether it is closer to the true situation of the parameter, one of positive index normalization or negative index normalization is selected to normalize the importance evaluation matrix E.

[0035] Preferably, the positive index is normalized as follows:

[0036]

[0037] in,

[0038] E b Let e ​​be the data vector of evaluation index b in the evaluation matrix. ab The values ​​of evaluation parameter a in the evaluation matrix for evaluation index b, nab This is the normalized score of evaluation parameter a in evaluation index b.

[0039] Preferably, the negative index is normalized as follows:

[0040]

[0041] in,

[0042] E b Let e ​​be the data vector of evaluation index b in the evaluation matrix. ab The values ​​of evaluation parameter a in the evaluation matrix for evaluation index b, n ab This is the normalized score of evaluation parameter a in evaluation index b.

[0043] Preferably, the weights of each importance evaluation method and the weights of each geological engineering parameter are calculated using the entropy weight method, and the steps are as follows:

[0044] Calculate the value e of evaluation parameter a in evaluation index b. ab The proportion P of evaluation index b ab :

[0045]

[0046] Where h is the number of evaluation parameters;

[0047] Calculate the information entropy value r of evaluation index b. b :

[0048]

[0049] Calculate the weight ω of each evaluation index b :

[0050]

[0051] Where u represents the number of evaluation indicators;

[0052] Finally, the evaluation matrix E is multiplied by the index weight vector to obtain the fuzzy comprehensive evaluation matrix S:

[0053]

[0054] in,

[0055]

[0056] in,

[0057] [ξ], [J] and These are the feature importance score vectors obtained by the grey relational analysis method, the CatBoost ensemble learning method, and the SHAP evaluation method, respectively.

[0058] Each element in the matrix represents an element of the comprehensive evaluation matrix E, indicating the importance score of the geological engineering parameter obtained by using different importance evaluation methods.

[0059] ω is the evaluation index vector.

[0060] Weights are calculated on the principal component vectors, and the abstract, dimension-reduced, and augmented geological engineering parameters in the comprehensive fuzzy evaluation vector S are restored to their original geological engineering parameters, resulting in importance scores for the original geological engineering parameters. The relationship between the collinear combinations of geological engineering parameters and their importance scores from the dimension-reduced principal components is shown below:

[0061]

[0062] in,

[0063] S Z The comprehensive evaluation score of the principal component evaluation parameters;

[0064] j is the sequence number of the geological engineering parameter;

[0065] k represents the total number of geological engineering parameters involved in the dimensionality reduction;

[0066] S Zj The comprehensive evaluation score of the j-th geological engineering parameter after principal component reduction;

[0067] φ zj It is the weight of the j-th geological engineering parameter of the principal component.

[0068] Preferably, the method further includes step S5, establishing a regression prediction model for geological engineering parameters and target yield parameters, introducing the final feature importance score of each importance rating method as the basis for initializing the regression prediction model, and verifying the rationality of the main control factor analysis results by comparing the training results and the accuracy of the regression prediction model.

[0069] Preferably, in step S5, the rationality of the main control factor analysis results is evaluated using a prediction accuracy index that combines accuracy requirements:

[0070]

[0071] Where err is the minimum error requirement, e true e is the true value pred Here is the predicted value, n is the total number of sample wells, and acc is the predicted value. err To meet the required level of prediction accuracy, acc err A value of 1 indicates that the prediction accuracy of all samples meets the requirements.

[0072] Preferably, the method for calculating feature importance scores using the grey relational analysis method is as follows:

[0073]

[0074] Where, ξ j Let y be the grey correlation coefficient between geological engineering parameter j and production parameter y. i For the production of sample well i, x ij Here, represents the fracturing data for geological engineering parameter j of sample well i, n represents the total number of sample wells, and ρ represents the resolution coefficient, typically 0.5. A larger grey relational coefficient indicates a closer relationship with the target parameter, signifying a positive indicator.

[0075] Preferably, when using the CatBoost ensemble learning evaluation method for importance assessment, geological engineering parameters are used as input and yield parameters as output. The CatBoost ensemble learning model is trained, and the calculation method for the importance score of geological engineering parameter features is derived as follows:

[0076]

[0077]

[0078] in, This represents the importance evaluation of geological engineering parameter j in the ensemble learning model, where M is the number of trees representing geological engineering parameter j in the generative model; T represents the number of trees, T q Let q be the th tree, L be the number of leaf nodes in each tree, and L-1 be the number of non-leaf nodes in the tree. t It is a feature associated with node t, I t 2 This is the reduction in squared loss after node t is split. For example, before the split, the node has 100 data points and an MSE of 1; after the split, it becomes two nodes, one with 30 data points and an MSE of 0.9, and the other with 70 data points and an MSE of 0.8. Considering weighting, I... t 2 That equals (100 + 1 - 300.9 - 70 + 0.8) / 100 = 0.17. t 2 A larger value indicates a greater impact of the node on the accuracy of the final predicted yield. Therefore, the greater the importance of a node feature, the greater its impact on the prediction of the target parameter, making it a positive indicator.

[0079] The CatBoost model employs a fully symmetric decision tree as its base learner and enhances feature dimensionality using methods such as combining class features and target feature statistics. It also utilizes a ranking boosting algorithm to obtain unbiased gradient estimation, increasing computational complexity to improve model performance, which aligns well with the current reality of small sample data. Feature importance ranking is determined by extracting feature weights from within the CatBoost model.

[0080] Preferably, based on the trained CatBoost ensemble learning model, the SHAP value is used to analyze feature importance. The SHAP importance score calculation method is as follows:

[0081]

[0082] in, For the SHAP importance score of geological engineering parameter j, {X 改造段长 Z 主成分1 ,…,X 用液强度} is the dimensionality-reduced and augmented geological engineering dataset, X j Let be the fracturing data for geological engineering parameter j; h is the number of geological engineering parameters after dimensionality reduction and augmentation, {X 改造段长 Z 主成分1 ,…,X 用液强度}\{X j} is excluding {X j The set of all possible input geological engineering parameter features, f x (S) is the prediction of the feature subset S of geological engineering parameters. Let f be the proportion of feature combinations of a subset S of geological engineering parameter features. The sum of the proportions of feature combinations of all possible subsets S of geological engineering parameter features equals 1. x The SHAP function is the prediction function of the CatBoost model. The larger the SHAP value, the greater the influence on the target parameter, which is a positive indicator.

[0083] Shapley Additive Explanations (SHAP) is a game theory-based method that can interpret the output of any machine learning model. The core idea of ​​SHAP is to calculate the marginal contribution of each feature to the model output, i.e., the SHAP value of each feature. SHAP values ​​satisfy additivity, meaning the model output equals the sum of the SHAP values ​​of all features plus a baseline value. SHAP supports various model types, including deep learning, tree models, and linear models, and can handle high-dimensional and time-series data. This paper uses SHAP to explain the feature importance of the CatBoost model.

[0084] Fuzzy mathematics evaluation is a comprehensive evaluation method based on fuzzy mathematics, capable of handling evaluation problems involving fuzzy concepts. The basic steps of fuzzy mathematics comprehensive evaluation are as follows: 1. Determine the evaluation objective and evaluation index system, i.e., the object to be evaluated and the various factors influencing it; 2. Determine the factor set and the comment set, i.e., use ordinary sets to represent each factor and possible evaluation results; 3. Determine the weight of each factor, i.e., use fuzzy sets to represent the importance of each factor to the evaluation objective; 4. Determine the membership function of each factor, i.e., use functions to represent the membership degree of each factor to each comment, that is, the degree to which it belongs to a certain comment; 5. Construct a fuzzy comprehensive evaluation matrix, i.e., use a matrix to represent the membership degree of each factor to each comment; 6. Perform fuzzy comprehensive operations, i.e., calculate using weights and membership degrees to obtain the final evaluation result.

[0085] Preferably, the KNN nearest neighbor prediction method for filling missing data specifically includes: treating samples as a set of vectors, measuring the proximity between sample vectors using Euclidean geometric distance, and filling sample vectors with missing values ​​using the average of the known values ​​of the nearest multiple sets of sample vectors.

[0086] Preferably, the padded dataset is represented as {X} 改造段长 ,X 加砂强度 ,…,X 用液强度 The dataset {X, Y} is a column vector of production parameter data. The geological engineering parameters are from n sample wells, totaling m items. 改造段长 ,X 加砂强度 ,…,X 用液强度 The matrix formed by} is n×m, where the geological engineering parameter j of sample well i is represented by x. ij .

[0087] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0088] 1. This invention realizes a comprehensive evaluation method for main control factors that combines modeling analysis and fuzzy mathematics. It overcomes the difficulties in comprehensively considering multiple main control schemes in modeling analysis methods, as well as the problems of the influence of human factors on weights and the difficulty in obtaining index matrices in fuzzy mathematics. Furthermore, it provides a method for verifying the rationality of the combination of main control factors.

[0089] 2. This invention proposes a comprehensive analysis method for the main controlling factors of fracturing data in relation to production. It aims to uncover the internal relationships within the data, fully consider the results of multiple analysis methods, and comprehensively evaluate and judge the main controlling factors of production. By effectively utilizing fracturing data, it extracts the main characteristics related to production and integrates multiple analysis methods and models to obtain relatively objective and accurate evaluation results of the main controlling factors of production, providing a reference for the development of unconventional reservoirs. Attached Figure Description

[0090] Figure 1This is a schematic diagram of the geological engineering parameter feature transformation process according to an embodiment of the present invention;

[0091] Figure 2 This is a schematic diagram illustrating the process of constructing a geological engineering parameter importance evaluation matrix according to an embodiment of the present invention;

[0092] Figure 3 This is a schematic diagram of the fuzzy comprehensive evaluation process for the importance of geological engineering parameters according to an embodiment of the present invention;

[0093] Figure 4 A schematic diagram illustrating the reduction of two-dimensional data to one dimension using principal component analysis (significant information loss);

[0094] Figure 5 A schematic diagram illustrating the reduction of two-dimensional data to one-dimensional data using principal component analysis (with minimal information loss);

[0095] Figure 6 This is a flowchart of a comprehensive analysis method for key control factors of fracturing data related to production output, according to an embodiment of the present invention. Detailed Implementation

[0096] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0097] This invention provides a comprehensive analysis method for the main controlling factors of fracturing data for production output, comprising the following steps:

[0098] Step S1: Preprocess the fracturing data of field production: Perform data missing and data anomaly processing on the fracturing data of field production to obtain the processed fracturing data.

[0099] Step S2 involves feature selection and transformation of the geological engineering parameters that affect fracturing data: calculating the collinearity and correlation between geological engineering parameters, screening geological engineering parameters, and extracting features of geological engineering parameters through data dimensionality reduction and data augmentation of the processed fracturing data.

[0100] Step S3: Use multiple evaluation methods to evaluate the importance of fracturing parameters for target production: Use multiple importance evaluation methods to evaluate the characteristics of target production parameters, and establish an importance evaluation matrix for geological engineering parameters by integrating the importance evaluation results of multiple importance evaluation methods. The importance evaluation methods include grey relational analysis, CatBoost ensemble learning, and SHAP evaluation.

[0101] Step S4: Perform a comprehensive evaluation of the main control factors of the evaluation matrix using fuzzy mathematics: determine the weights of the feature vectors of each importance evaluation method in the importance evaluation matrix based on the entropy weight method, calculate the final feature importance score by weighting, and rank them to obtain the main control factors of output.

[0102] Furthermore, in step S1, the missing data is filled using the KNN nearest neighbor prediction method.

[0103] Furthermore, in step S2, feature selection and transformation are performed on the geological and engineering parameters affecting the fracturing data, including the following steps:

[0104] Collinearity analysis was performed on geological engineering parameters using VIF analysis, and correlation analysis was performed on the parameters using principal component analysis.

[0105] The VIF analysis method is used to reduce the dimensionality of geological engineering parameters, thereby removing collinearity while preserving data interpretability.

[0106] Furthermore, the variance expansion factor (VIF) is used to determine collinearity among geological engineering parameters by pairwise combination. The VIF calculation method is as follows:

[0107]

[0108] For parameter combinations with a variance inflation factor (VIF) greater than 10 exhibiting multicollinearity, principal component replacement is performed, and dimensionality reduction is achieved using principal component analysis (PCA), retaining only one principal component. The principal component calculation method is as follows:

[0109] And satisfy

[0110] in,

[0111] R j 2 It is the coefficient of determination obtained by linear regression on the other k-1 geological engineering parameters when the j-th geological engineering parameter is used as the explained variable, where k is the total number of geological engineering parameters participating in the dimensionality reduction;

[0112] Where Z1 is the first principal component, X j It is the dataset of the j-th geological engineering parameter, φ 1j It is the weight of the first principal component of the j-th geological engineering parameter;

[0113] n is the total number of sample wells;

[0114] m represents the total number of geological engineering parameters;

[0115] j is the sequence number of the geological engineering parameter;

[0116] i is the sample well number;

[0117] x ij Let be the j-th geological engineering parameter of the i-th well.

[0118] Furthermore, for parameters with weak collinearity, indicating strong independence, data is augmented through cross-multiplication, increasing data complexity. A two-dimensional transformation converts X1,X2 to X1,X2,X1 2 X2 2 There are five vectors, X1 and X2.

[0119] The dataset after dimensionality reduction and augmentation is represented as {X} 改造段长 Z 主成分1 ,…,X 用液强度}

[0120] Furthermore, in step S3, using the dimension-reduced and augmented geological engineering parameter dataset, the importance of the characteristics of the target yield parameter is evaluated using multiple importance evaluation methods. The importance evaluation matrix of the geological engineering parameter is established by integrating the importance evaluation results of multiple importance evaluation methods. With the target yield parameter as the dependent variable and other geological engineering parameters as independent variables, the characteristic importance score of each geological engineering parameter to the target yield parameter is calculated. Finally, the importance evaluation matrix E is constructed.

[0121] Furthermore, in step S4, the analysis results of the importance evaluation matrix E are comprehensively evaluated using fuzzy mathematics, the final parameter importance score for the target parameter is calculated and ranked, and the main control factors of output are analyzed and judged.

[0122] Furthermore, in step S4, the importance evaluation matrix E is standardized to obtain a fuzzy evaluation matrix. The entropy weight method in fuzzy mathematics is used to calculate the weights of each importance evaluation method and the weights of each geological engineering parameter to obtain a comprehensive fuzzy evaluation vector. The final importance score of each geological engineering parameter is obtained and sorted from high to low to determine the main control factors of production.

[0123] Furthermore, when standardizing the importance evaluation matrix E, for any geological engineering parameter, based on the principle of whether it is closer to the true situation of the parameter, either positive index normalization or negative index normalization is selected to normalize the importance evaluation matrix E.

[0124] Furthermore, the positive indicators are normalized as follows:

[0125]

[0126] in,

[0127] E b Let e ​​be the data vector of evaluation index b in the evaluation matrix. ab The values ​​of evaluation parameter a in the evaluation matrix for evaluation index b, n ab This is the normalized score of evaluation parameter a in evaluation index b.

[0128] Furthermore, the negative indicators are normalized as follows:

[0129]

[0130] in,

[0131] E b Let e ​​be the data vector of evaluation index b in the evaluation matrix. ab The values ​​of evaluation parameter a in the evaluation matrix for evaluation index b, n ab This is the normalized score of evaluation parameter a in evaluation index b.

[0132] Furthermore, the weights of each importance evaluation method and the weights of each geological engineering parameter are calculated using the entropy weight method, as follows:

[0133] Calculate the value e of evaluation parameter a in evaluation index b. ab The proportion P of evaluation index b ab :

[0134]

[0135] Where h is the number of evaluation parameters;

[0136] Calculate the information entropy value r of evaluation index b. b :

[0137]

[0138] Calculate the weight ω of each evaluation index b :

[0139]

[0140] Where u represents the number of evaluation indicators;

[0141] Finally, the evaluation matrix E is multiplied by the index weight vector to obtain the fuzzy comprehensive evaluation matrix S:

[0142]

[0143] in,

[0144]

[0145] in,

[0146] [ξ], [J] and These are the feature importance score vectors obtained by the grey relational analysis method, the CatBoost ensemble learning method, and the SHAP evaluation method, respectively.

[0147] Each element in the matrix represents an element of the comprehensive evaluation matrix E, indicating the importance score of the geological engineering parameter obtained by using different importance evaluation methods.

[0148] ω is the evaluation index vector.

[0149] Weights are calculated on the principal component vectors, and the abstract, dimension-reduced, and augmented geological engineering parameters in the comprehensive fuzzy evaluation vector S are restored to their original geological engineering parameters, resulting in importance scores for the original geological engineering parameters. The relationship between the collinear combinations of geological engineering parameters and their importance scores from the dimension-reduced principal components is shown below:

[0150]

[0151] in,

[0152] S Z The comprehensive evaluation score of the principal component evaluation parameters;

[0153] j is the sequence number of the geological engineering parameter;

[0154] k represents the total number of geological engineering parameters involved in the dimensionality reduction;

[0155] S Zj The comprehensive evaluation score of the j-th geological engineering parameter after principal component reduction;

[0156] φ zj It is the weight of the j-th geological engineering parameter of the principal component.

[0157] Furthermore, step S5 is included, which establishes a regression prediction model for geological engineering parameters and target yield parameters, introduces the final feature importance score of each importance rating method as the basis for initializing the regression prediction model, and verifies the rationality of the main control factor analysis results by comparing the training results and the accuracy of the regression prediction model.

[0158] Furthermore, in step S5, the rationality of the main control factor analysis results is evaluated using a prediction accuracy index that combines accuracy requirements:

[0159]

[0160] Where err is the minimum error requirement, e true e is the true value pred Here is the predicted value, n is the total number of sample wells, and acc is the predicted value. err To meet the required level of prediction accuracy, acc err A value of 1 indicates that the prediction accuracy of all samples meets the requirements.

[0161] Furthermore, the calculation method for feature importance scores using the grey relational analysis method is as follows:

[0162]

[0163] Where, ξ j Let y be the grey correlation coefficient between geological engineering parameter j and production parameter y. i For the production of sample well i, x ij Here, represents the fracturing data for geological engineering parameter j of sample well i, n represents the total number of sample wells, and ρ represents the resolution coefficient, typically 0.5. A larger grey relational coefficient indicates a closer relationship with the target parameter, signifying a positive indicator.

[0164] Furthermore, when using the CatBoost ensemble learning evaluation method for importance assessment, geological engineering parameters are used as input and yield parameters as output. The CatBoost ensemble learning model is trained, and the calculation method for the importance score of geological engineering parameter features is derived as follows:

[0165]

[0166]

[0167] in, This represents the importance evaluation of geological engineering parameter j in the ensemble learning model, where M is the number of trees representing geological engineering parameter j in the generative model; T represents the number of trees, T q Let q be the th tree, L be the number of leaf nodes in each tree, and L-1 be the number of non-leaf nodes in the tree. t It is a feature associated with node t, I t 2 This is the reduction in squared loss after node t is split. For example, before the split, the node has 100 data points and an MSE of 1; after the split, it becomes two nodes, one with 30 data points and an MSE of 0.9, and the other with 70 data points and an MSE of 0.8. Considering weighting, I... t 2 That equals (100 + 1 - 300.9 - 70 + 0.8) / 100 = 0.17. t 2 A larger value indicates a greater impact of the node on the accuracy of the final predicted yield. Therefore, the greater the importance of a node feature, the greater its impact on the prediction of the target parameter, making it a positive indicator.

[0168] Furthermore, based on the trained CatBoost ensemble learning model, the SHAP value is used to analyze feature importance. The SHAP importance score calculation method is as follows:

[0169]

[0170] in, For the SHAP importance score of geological engineering parameter j, {X 改造段长 Z 主成分1 ,…,X 用液强度} is the dimensionality-reduced and augmented geological engineering dataset, X j Let be the fracturing data for geological engineering parameter j; h is the number of geological engineering parameters after dimensionality reduction and augmentation, {X 改造段长 Z 主成分1 ,…,X 用液强度}\{X j} is excluding {X j The set of all possible input geological engineering parameter features, f x (S) is the prediction of the feature subset S of geological engineering parameters. Let f be the proportion of feature combinations of a subset S of geological engineering parameter features. The sum of the proportions of feature combinations of all possible subsets S of geological engineering parameter features equals 1. x The SHAP function is the prediction function of the CatBoost model. The larger the SHAP value, the greater the influence on the target parameter, which is a positive indicator.

[0171] Furthermore, the KNN nearest neighbor prediction method for filling missing data specifically includes: treating samples as a set of vectors, measuring the proximity between sample vectors using Euclidean geometric distance, and filling sample vectors with missing values ​​with the average of the known values ​​of the nearest multiple sets of sample vectors.

[0172] Furthermore, the padded dataset is represented as {X} 改造段长 ,X 加砂强度 ,…,X 用液强度 The dataset {X, Y} is a column vector of production parameter data. The geological engineering parameters are from n sample wells, totaling m items. 改造段长 ,X 加砂强度 ,…,X 用液强度 The matrix formed by} is n×m, where the geological engineering parameter j of sample well i is represented by x. ij .

[0173] Further, in one embodiment, the dataset includes the following geological engineering parameters: build-up point, A sounding depth, A vertical depth, B sounding depth, B vertical depth, artificial well bottom / m, horizontal section length / m, workable well section length / m, designed intervention section length / m, oil layer penetration rate / %, oil layer thickness / m, Class I oil layer / m, Class II oil layer / m, Class III oil layer / m, Φ_min / %, Φ_max / %, Φ_avg / %, K_min / mD, K_max / mD, K_avg / mD, So_min / %, So_max / %, So_avg / %, Slug-quartz sand (70 / 140 mesh), Main-quartz sand (70 / 140 mesh), Slug-quartz sand (40 / 70 mesh), Main-quartz sand (40 / 70 mesh), Slug-quartz sand (30 / 50 mesh), Main-quartz sand (30 / 50 mesh), Slug-quartz sand (20 / 40 mesh), Main-quartz sand (20 / 40 mesh), Slug-ceramsite (70 / 140 mesh), Slug-ceramsite (40 / 70 mesh), Slug-ceramsite (20 / 40 mesh), Ceramsite (20 / 40 mesh), Ceramsite (30 / 50 mesh), Total construction fluid volume / m³ 3 Total sand volume during construction / m 3 Actual modified section length, actual number of sections, actual number of clusters, number of grades with 6 or more clusters within a section, number of clusters with 6 or more clusters within a section, acid foaming grade, construction liquid-sand ratio, and construction liquid strength / m 3 / m, construction sand strength (m) 3 / m, construction slippage water volume / m 3 Construction slippery water ratio / %, pre-construction liquid ratio / %, construction uniform sand ratio / %, maximum construction sand ratio / kg / m³ 3 Average pump stop pressure / MPa, maximum construction pressure / MPa, average burst pressure / MPa, highest burst pressure / MPa, construction span (d), pure working time (d), highest number of construction sections per day, average number of construction sections per day, sand addition completion rate / %, minimum actual construction discharge / m³ / min, maximum actual construction discharge / m³ / min 3 / min, casing outer diameter / mm, casing wall thickness / mm, etc. Production indicators are selected based on early cumulative production indicators, such as 30-day cumulative production.

[0174] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.

Claims

1. A comprehensive analysis method for key controlling factors in fracturing data related to production output, characterized in that, Includes the following steps: Step S1: Preprocess the fracturing data of field production: Perform data missing and data anomaly processing on the fracturing data of field production to obtain the processed fracturing data. Step S2 involves feature selection and transformation of the geological engineering parameters that affect fracturing data: calculating the collinearity and correlation between geological engineering parameters, screening geological engineering parameters, and extracting features of geological engineering parameters through data dimensionality reduction and data augmentation of the processed fracturing data. Step S3: Use multiple evaluation methods to evaluate the importance of fracturing parameters for target production: Use multiple importance evaluation methods to evaluate the characteristics of target production parameters, and establish an importance evaluation matrix for geological engineering parameters by integrating the importance evaluation results of multiple importance evaluation methods. The importance evaluation methods include grey relational analysis, CatBoost ensemble learning, and SHAP evaluation. Step S4: Perform a comprehensive evaluation of the main control factors of the evaluation matrix using fuzzy mathematics: determine the weights of the feature vectors of each importance evaluation method in the importance evaluation matrix based on the entropy weight method, calculate the final feature importance score by weighting, and rank them to obtain the main control factors of output.

2. The comprehensive analysis method for key controlling factors of fracturing data for production output as described in claim 1, characterized in that, In step S1, the missing data is filled using the KNN nearest neighbor prediction method.

3. The comprehensive analysis method for key controlling factors of fracturing data for production output as described in claim 1, characterized in that, In step S2, feature selection and transformation are performed on the geological and engineering parameters that affect the fracturing data, including the following steps: Collinearity analysis was performed on geological engineering parameters using VIF analysis, and correlation analysis was performed on the geological engineering parameters using principal component analysis. The VIF analysis method is used to reduce the dimensionality of geological engineering parameters, thereby removing collinearity while preserving data interpretability.

4. The comprehensive analysis method for key controlling factors of fracturing data regarding production output as described in claim 3, characterized in that, The variance expansion factor (VIF) is used to determine collinearity among geological engineering parameters by pairwise combination. The VIF calculation method is as follows: For parameter combinations with a variance inflation factor (VIF) greater than 10 exhibiting multicollinearity, principal component replacement is performed, and dimensionality reduction is achieved using principal component analysis (PCA), retaining only one principal component. The principal component calculation method is as follows: And satisfy in, R j 2 It is the first j When a geological engineering parameter is used as the explained variable, other... k -1 coefficient of determination obtained by linear regression of geological engineering parameters k The total number of geological engineering parameters involved in dimensionality reduction; in, Z 1 is the first principal component. X j It is the first j A dataset of geological engineering parameters, It is the first j The first principal component weights of each geological engineering parameter; n This represents the total number of sample wells; j For geological engineering parameter serial numbers; i The sample well number; x ij For the first i Koujing's first j Geological engineering parameters.

5. The comprehensive analysis method for key controlling factors of fracturing data regarding production output as described in claim 4, characterized in that, In step S3, the geological engineering parameter dataset after dimensionality reduction and augmentation is used to evaluate the importance of the characteristics of the target yield parameter using multiple importance evaluation methods. The importance evaluation matrix of the geological engineering parameter is established by combining the importance evaluation results of multiple importance evaluation methods. With the target yield parameter as the dependent variable and other geological engineering parameters as independent variables, the characteristic importance score of each geological engineering parameter to the target yield parameter is calculated. Finally, the importance evaluation matrix E is constructed.

6. The comprehensive analysis method for key controlling factors of fracturing data for production output according to claim 1, characterized in that, In step S4, the analysis results of the importance evaluation matrix E are comprehensively evaluated using fuzzy mathematics, the final parameter importance scores for the target parameters are calculated and ranked, and the main control factors of production are analyzed and judged.

7. The comprehensive analysis method for key controlling factors of fracturing data for production output as described in claim 6, characterized in that, In step S4, the importance evaluation matrix E is standardized to obtain a fuzzy evaluation matrix. The entropy weight method in fuzzy mathematics is used to calculate the weights of each importance evaluation method and the weights of each geological engineering parameter to obtain a comprehensive fuzzy evaluation vector. The final importance score of each geological engineering parameter is obtained and sorted from high to low to determine the main control factors of production.

8. The comprehensive analysis method for key controlling factors of fracturing data for production output as described in claim 7, characterized in that, When standardizing the importance evaluation matrix E, for any geological engineering parameter, based on the principle of whether it is closer to the true situation of the parameter, either positive index normalization or negative index normalization is selected to normalize the importance evaluation matrix E.

9. The comprehensive analysis method for key controlling factors of fracturing data for production output as described in claim 8, characterized in that, The normalization process for positive indicators is as follows: in, E b Evaluation indicators in the evaluation matrix b data vectors, e ab Let the evaluation parameter 'a' in the evaluation matrix take the value of evaluation index 'b'. n ab For evaluation parameters a In evaluation indicators b The normalized score.

10. The comprehensive analysis method for key controlling factors of fracturing data for production output according to claim 8, characterized in that, The negative index is normalized as follows: in, E b Evaluation indicators in the evaluation matrix b data vectors, e ab Let the evaluation parameter 'a' in the evaluation matrix take the value of evaluation index 'b'. n ab For evaluation parameters a In evaluation indicators b The normalized score.

11. The comprehensive analysis method for key controlling factors of fracturing data for production output according to claim 7, characterized in that, The weights of each importance evaluation method and the weights of various geological engineering parameters are calculated using the entropy weight method. The steps are as follows: Calculate evaluation parameters a In evaluation indicators b The value of e ab Evaluation indicators b proportion P ab : in, h To evaluate the number of parameters; Calculate evaluation indicators b Information entropy value r b : Calculate the weight of each evaluation indicator : in, u The number of evaluation indicators; Finally, the importance evaluation matrix will be used. E The fuzzy comprehensive evaluation matrix is ​​obtained by multiplying the index weight vector. S : ; in, in, , and These are the feature importance score vectors obtained by the grey relational analysis method, the CatBoost ensemble learning method, and the SHAP evaluation method, respectively. ; Each element in the matrix represents an element of the importance evaluation matrix E, indicating the importance score of the geological engineering parameter obtained by using different importance evaluation methods. For the evaluation index vector, ; Weights are calculated for the principal component vectors, and the fuzzy comprehensive evaluation matrix is ​​then processed. S The abstract, dimension-reduced, and augmented geological engineering parameters are restored to their original counterparts, resulting in importance scores for these original parameters. The relationship between the importance scores of collinear combinations of geological engineering parameters and their dimension-reduced principal components is shown below: in, S Z The comprehensive evaluation score of the principal component evaluation parameters; j For geological engineering parameter serial numbers; k The total number of geological engineering parameters involved in dimensionality reduction; S Zj For the first phase after principal component reduction j The comprehensive evaluation score of each geological engineering parameter; It is the principal component number j The weights of each geological engineering parameter.

12. The comprehensive analysis method for key controlling factors of fracturing data for production output according to claim 1, characterized in that, It also includes step S5, which establishes a regression prediction model for geological engineering parameters and target yield parameters, introduces the final feature importance score of each importance rating method as the basis for initializing the regression prediction model, and verifies the rationality of the main control factor analysis results by comparing the training results and the accuracy of the regression prediction model.

13. The comprehensive analysis method for key controlling factors of fracturing data for production output according to claim 12, characterized in that, In step S5, the rationality of the main control factor analysis results is evaluated using the prediction accuracy index combined with the accuracy requirements: in, err To meet the minimum error requirement, e true For the true value, e pred For predicted values, n This represents the total number of sample wells. acc err To meet the required level of prediction accuracy, acc err A value of 1 indicates that the prediction accuracy of all samples meets the requirements.

Citation Information

Patent Citations

  • Method for optimizing fracturing parameters

    CN106703776A

  • Offshore oil-gas reservoir production increase scheme generation method

    CN110956388A