Shale gas well sweet spot prediction method and device based on ensemble learning

CN121413818APending Publication Date: 2026-01-27PETROCHINA CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411007632.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-25
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

同时由于大多数页岩气藏具有分布范围广、储层厚度大,且普遍含气的储层特征,使得页岩气井的稳产时间较长

Benefits of technology

[0042]本发明实施例提供的上述技术方案的有益效果至少包括:在预先训练模型的过程中,根据确定的影响甜点的主控因素,从目标区域中采样页岩气井的预测数据中提取主控因素数据;以及对所述第二数据中各参数进行特征提取,得到特征数据;基于目标区域中采样页岩气井的预测数据中的主控因素数据、第二数据中各参数的特征数据以及甜点数据组成的训练数据集训练各基模型,训练好的各个基模型的集成形成一个精度更高的模型,并且在选取训练数据的过程中,全面考虑了影响甜点的各种因素;基于训练好的甜点预测模型和预设的结合策略,通过输入待预测页岩气井的主控因素数据和特征数据得到甜点预测数据,提高了预测结果的准确性和预测效率,经实验数据表明,准确率在92%以上,提高工作效率50倍以上,进一步的,预测结果的准确性可以使得在施工过程中,优选甜点较好的压裂井层,指导压裂设计,从而降低压裂施工成本,提高压后产量。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121413818A_ABST
    Figure CN121413818A_ABST
Patent Text Reader

Abstract

The invention discloses a shale gas well dessert prediction method and device based on integrated learning. The method comprises the steps of obtaining main control factors and feature data of a shale gas well to be predicted, and inputting the main control factors and the feature data into a trained dessert prediction model to obtain dessert prediction data; the model training process comprises the following steps: determining main control factors influencing the dessert according to the correlation between each parameter in the prediction data of the sampled shale gas well in the target area and the dessert; the prediction data comprises first data and second data, the first data comprises geological data, perforation data and oil and gas production data, and the second data comprises logging data and fracturing construction data; performing feature extraction on each parameter in the second data to obtain feature data; constructing a training data set according to the main control factors and the characteristic data of the shale gas well; and training the dessert prediction model based on the training data set. The accuracy of a prediction result can be improved, and fracturing design is effectively guided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of fracturing engineering, in particular to a shale gas well sweet spot prediction method and device based on ensemble learning. BACKGROUND

[0002] With the rapid development of global economy, energy consumption and demand are growing rapidly. Unlike conventional oil and gas reservoirs, shale gas reservoirs are unconventional gas reservoirs of "self-generation and self-storage", and the reservoir gas mainly exists in the form of adsorption and free state. At the same time, due to the characteristics of most shale gas reservoirs with wide distribution, large reservoir thickness and generally gas-containing reservoirs, the stable production time of shale gas wells is longer. Therefore, for the sustainable and stable development of natural gas in China, whether the shale gas reservoir can be effectively developed and utilized plays a very important significance, but at present, a relatively perfect shale gas development block optimization system has not been established.

[0003] The shale gas "sweet spot" usually refers to the best enrichment and easy-to-develop area or horizon of shale gas exploration and development, which has high economic benefits. Therefore, it is of great significance to deeply study the shale gas "sweet spot" prediction method and accurately predict and identify the shale sweet spot. The prediction of shale gas "sweet spot" is the key to low-cost exploration and high-yield of shale gas.

[0004] Many scholars at home and abroad have studied the prediction of shale gas sweet spot, mainly through seismic data analysis, pre-stack inversion, fracture prediction, logging data prediction, core analysis, gas content analysis, reservoir modeling and other methods to predict the location of favorable exploration area of shale gas. These methods are all through the attributes obtained by correlation analysis to predict the sweet spot of shale gas. SUMMARY

[0005] The present application inventors found that most of the current methods are to realize the sweet spot prediction of shale gas well through qualitative analysis, which cannot achieve the quantitative analysis of shale gas well sweet spot prediction. In the prediction process, the geological sweet spot prediction and the engineering sweet spot prediction are separated, and the sweet spot prediction considering the combination of geology and engineering is less. Therefore, these methods are difficult to adapt to the characteristics of low porosity and low permeability, strong heterogeneity of shale gas reservoir, and cannot more accurately realize the sweet spot prediction of shale gas well, so as to more effectively guide the fracturing design to better realize the industrial application.

[0006] In view of the above problems, the present application is proposed to provide a shale gas well sweet spot prediction method and device based on ensemble learning to overcome the above problems or at least partially solve the above problems.

[0007] The present application provides a shale gas well sweet spot prediction method based on ensemble learning, comprising:

[0008] Obtaining prediction data of a target area to-be-predicted shale gas well, the prediction data comprising first data and second data, the first data comprising geological data, perforation data and oil and gas production data, and the second data comprising logging data and fracturing operation data;

[0009] According to a predetermined main control factor affecting the sweet spot, main control factor data is extracted, and feature data of each parameter in the second data is extracted;

[0010] The main control factor data and the feature data of the to-be-predicted shale gas well are input into a trained sweet spot prediction model, and sweet spot prediction data of the to-be-predicted shale gas well is output; wherein the training process of the sweet spot prediction model comprises:

[0011] According to the correlation between each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot, the main control factor affecting the sweet spot is determined, and the main control factor data is extracted from the prediction data of the sampling shale gas well in the target area; and feature data of each parameter in the second data is extracted; the sweet spot data of the sampling shale gas well is data evaluated based on each parameter in the prediction data;

[0012] According to the main control factor data and the feature data of the shale gas well, a training data set is constructed; the training data set comprises the main control factor data of the sampling shale gas well, the feature data of each parameter in the second data and the sweet spot data;

[0013] The sweet spot prediction model is trained based on the training data set; the sweet spot prediction model comprises at least three kinds of integrated learning base models and a combination strategy module connected with the output end of the base model, and the combination strategy module is used for aggregating the output results of the base model according to a preset combination strategy.

[0014] In some optional embodiments, the main control factor affecting the sweet spot is determined according to the correlation between each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot, comprising:

[0015] According to the pearson correlation analysis method, the correlation coefficient between each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot is determined, and the main control factor affecting the sweet spot is screened from each parameter based on the correlation coefficient between each parameter and the sweet spot.

[0016] In some optional embodiments, the correlation coefficient between each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot of the shale gas well is determined according to the pearson correlation analysis method, comprising:

[0017] Based on each parameter data in the prediction data of the sampling shale gas well in the target area and the sweet spot data of the shale gas well, the correlation coefficient between each parameter and the sweet spot is calculated by using the following formula:

[0018]

[0019] wherein, X represents parameter data in prediction data of a sample shale gas well in the target area, Y represents a sweet spot of the shale gas well, r xy represents a correlation coefficient of the parameter and the sweet spot, N represents a total number of parameters, and n represents an nth parameter to be processed, and n is 1 to N.

[0020] In some optional embodiments, the feature extraction on each parameter in the second data obtains feature data, and the feature data includes at least one of a maximum value, a minimum value, an average value, a three-quantile value, a median value, a seven-quantile value, a variance, a standard deviation, a skewness, a kurtosis, a change rate, an extreme value, an entropy, and a curvature.

[0021] In some optional embodiments, the feature extraction on each parameter in the second data obtains feature data, and the feature data includes at least one of a maximum value, a minimum value, an average value, a three-quantile value, a median value, a seven-quantile value, a variance, a standard deviation, a skewness, a kurtosis, a change rate, an extreme value, an entropy, and a curvature.

[0022] In some optional embodiments, the method further includes: preprocessing the obtained prediction data of the shale gas well to be predicted in the target area.

[0023] In some optional embodiments, the method further includes: supplementing the missing parameters in the logging data with data calculated by a formula corresponding to the parameters.

[0024] In some optional embodiments, the method further includes: data cleaning of the prediction data, including data deduplication, missing value processing, abnormal value processing, data normalization and standardization, and data conversion.

[0025] In some optional embodiments, the training of the sweet spot prediction model based on the training data set includes:

[0026] In some optional embodiments, the training of the sweet spot prediction model based on the training data set includes:

[0027] In some optional embodiments, the training of the sweet spot prediction model based on the training data set includes:

[0028] In some optional embodiments, the training of the sweet spot prediction model based on the training data set includes:

[0029] In some optional embodiments, the training of the sweet spot prediction model based on the training data set includes:

[0030] In some optional embodiments, the output results of the base models are aggregated according to a preset combination strategy, including:

[0031] The sweet spot prediction data of each base model in the ensemble learning is collected based on the voting regressor, and the collected sweet spot prediction data of each base model is aggregated according to a preset combination strategy;

[0032] The combination strategy includes a simple average method and a weighted average method.

[0033] The embodiment of the present application also provides a shale gas well sweet spot prediction device based on ensemble learning, including:

[0034] A first data acquisition unit is configured to acquire prediction data of a target area shale gas well to be predicted, the prediction data including first data and second data, the first data including geological data, perforation data and oil and gas production data, and the second data including logging data and fracturing operation data; main control factor data is extracted according to a pre-determined main control factor affecting a sweet spot; and feature data is obtained by performing feature extraction on each parameter in the second data;

[0035] A sweet spot prediction unit is configured to input the main control factor data and the feature data of the shale gas well to be predicted into a trained sweet spot prediction model, and output sweet spot prediction data of the shale gas well to be predicted;

[0036] A second data acquisition unit is configured to determine a main control factor affecting a sweet spot according to the correlation between each parameter in the prediction data of a target area shale gas well to be predicted and the sweet spot, extract main control factor data from the prediction data of the target area shale gas well to be predicted, and perform feature extraction on each parameter in the second data to obtain feature data;

[0037] A training data set is constructed based on the main control factor data and the feature data of the shale gas well; the training data set includes main control factor data of a sample shale gas well, feature data of each parameter in the second data and sweet spot data; the sweet spot data of the sample shale gas well is data evaluated based on each parameter in the prediction data;

[0038] A model training unit is configured to train a sweet spot prediction model based on the training data set; the sweet spot prediction model includes at least three ensemble learning base models and a combination strategy module connected to an output end of the base models, and the combination strategy module is configured to aggregate output results of the base models according to a preset combination strategy.

[0039] The embodiment of the present application also provides a computer storage medium, the computer storage medium storing computer executable instructions, the computer executable instructions being executed by a processor to implement a shale gas well sweet spot prediction method based on ensemble learning.

[0040] The embodiment of the present application also provides a prediction device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the shale gas well sweet spot prediction method based on ensemble learning when executing the program.

[0041] The embodiment of the present application also provides a computer program product, comprising a computer program, and the computer program implements the shale gas well sweet spot prediction method based on ensemble learning when executed by a processor.

[0042] The above technical solution provided by the embodiment of the present application has at least the following beneficial effects: in the process of pre-training the model, the master control factor data is extracted from the prediction data of the shale gas well sampled in the target area according to the determined master control factor affecting the sweet spot; and the feature data is obtained by performing feature extraction on each parameter in the second data; each base model is trained based on the training data set composed of the master control factor data in the prediction data of the shale gas well sampled in the target area, the feature data of each parameter in the second data and the sweet spot data, and the integration of the trained each base model forms a model with higher precision, and in the process of selecting the training data, various factors affecting the sweet spot are comprehensively considered; based on the trained sweet spot prediction model and the preset combination strategy, the sweet spot prediction data is obtained by inputting the master control factor data and the feature data of the shale gas well to be predicted, the accuracy and the prediction efficiency of the prediction result are improved, the accuracy rate is above 92%, the work efficiency is improved by more than 50 times, further, the accuracy of the prediction result can make the sweet spot better in the construction process, guide the fracturing design, thereby reducing the fracturing construction cost and improving the post-fracturing production.

[0043] Other features and advantages of the present application will be set forth in the following description, and in part will become apparent to those skilled in the art from the description, or can be learned by practice of the present application. The objects and other advantages of the present application can be realized and attained by the structure particularly pointed out in the written description and claims hereof as well as the appended drawings.

[0044] The technical solutions of the present application will be further described in detail below with the help of the accompanying drawings and embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0045] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, illustrate the present application and are used to explain the present application together with the embodiments of the present application, and do not constitute a limitation on the present application. In the drawings:

[0046] Figure 1 The flowchart of the shale gas well sweet spot prediction method based on ensemble learning in the embodiment of the present application;

[0047] Figure 2 Corresponding relationship diagram of correlation coefficient and correlation in the embodiment one of the present application;

[0048] Figure 3 Flowchart of the dessert prediction model training method in the embodiment one of the present application;

[0049] Figure 4 Principle diagram of the ensemble learning in the embodiment one of the present application;

[0050] Figure 5 Principle diagram of the ensemble learning based on the selected base model in the embodiment one of the present application;

[0051] Figure 6 Flowchart of the shale gas well dessert prediction method based on the ensemble learning in the embodiment two of the present application;

[0052] Figure 7 Principle diagram of the shale gas well dessert prediction method based on the ensemble learning in the embodiment two of the present application;

[0053] Figure 8 Principle diagram of the feature extraction in the embodiment two of the present application;

[0054] Figure 9 Principle diagram of the neural network in the embodiment two of the present application;

[0055] Figure 10 Structure schematic diagram of the shale gas well dessert prediction device based on the ensemble learning in the embodiment two of the present application. DETAILED DESCRIPTION

[0056] Exemplary embodiments of the present disclosure will be described in greater detail below with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood, and the scope of the present disclosure can be accurately conveyed to those skilled in the art.

[0057] In order to solve the problem that the dessert prediction of the shale gas well cannot be more accurately realized and the dessert prediction is not comprehensively considered in the aspects of geology and engineering in the prior art, the embodiment of the present application provides a shale gas well dessert prediction method based on ensemble learning, which can more accurately realize the dessert prediction of the shale gas well, thereby more effectively guiding the fracturing design to better realize the industrial application.

[0058] Embodiment one

[0059] The embodiment one of the present application provides a shale gas well dessert prediction method based on ensemble learning, which is based on a trained dessert prediction model and predicts the dessert of a shale gas well to be predicted, and the prediction process is as followsFigure 1 As shown, comprising:

[0060] Step S101: obtaining prediction data of a shale gas well to be predicted in a target area, the prediction data comprising first data and second data, the first data comprising geological data, perforation data and oil and gas production data, and the second data comprising logging data and fracturing operation data.

[0061] For the shale gas well to be predicted, in the process of obtaining data, part of the missing prediction data can be supplemented by corresponding data of adjacent wells, similar wells or simulation wells.

[0062] Step S102: extracting main control factor data according to a pre-determined main control factor affecting the sweet spot; and performing feature extraction on each parameter in the second data to obtain feature data.

[0063] In this step, the process of pre-determining the main control factor affecting the sweet spot comprises:

[0064] According to the pearson correlation analysis method, the correlation coefficients of each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot are determined, and based on the correlation coefficients of each parameter and the sweet spot, the main control factor affecting the sweet spot is selected from each parameter.

[0065] In this step, based on each parameter data in the prediction data of the sampling shale gas well in the target area and the sweet spot of the shale gas well, the correlation coefficient of each parameter and the sweet spot is calculated using the following formula, and the sweet spot data of the sampling shale gas well is data evaluated based on each parameter in the prediction data:

[0066]

[0067] Wherein, X represents each parameter data in the prediction data of the sampling shale gas well in the target area, Y represents the sweet spot data of the shale gas well, r xy represents the correlation coefficient of the parameter and the sweet spot, N represents the total number of parameters, and n represents the nth parameter to be processed, which takes a value of 1-N.

[0068] According to the determined correlation coefficients of each parameter in the prediction data and the sweet spot, the parameters with correlation coefficients meeting the requirements are selected from the prediction data of the sampling shale gas well in the target area, which can be selected according to certain rules, for example, selecting parameters with habit greater than a pre-set threshold value, considering that these parameters have a greater impact on the sweet spot, which can be used as the main control factor affecting the sweet spot. Referring to Figure 2For example, in the example of the correlation coefficient, different ranges of the correlation coefficient correspond to different correlations, for example, a parameter with a correlation coefficient greater than 0.8 and less than or equal to 1 has a strong correlation, a parameter with a correlation coefficient greater than 0.6 and less than or equal to 0.8 has a strong correlation, a parameter with a correlation coefficient greater than 0.4 and less than or equal to 0.6 has a moderate correlation, a parameter with a correlation coefficient greater than 0.2 and less than or equal to 0.4 has a weak correlation, and a parameter with a correlation coefficient greater than 0 and less than or equal to 0.2 has a very weak correlation or no correlation. The correlation between the range of the correlation coefficient and the correlation can be set as needed, and when selecting the main control factor, parameters with different degrees of correlation can be selected as the main control factor as needed, for example, parameters with a correlation coefficient of 0.4 or more, i.e., a correlation of moderate correlation or more, can be selected from the sampled shale gas well prediction data in the target area as the main control factor affecting the sweet spot.

[0069] In this step, the process of extracting feature data from the second data includes:

[0070] Each parameter data in the second data is converted into multi-dimensional feature data to extract more representative and distinguishable parameter data; the multi-dimensional feature data includes at least one of the maximum value, the minimum value, the average value, the three quantiles, the median, the seven quantiles, the variance, the standard deviation, the skewness, the kurtosis, the change rate, the extreme value, the entropy, and the curvature.

[0071] Each parameter data in the second data is converted into mathematical feature data of different dimensions in Table 1 below, and different dimensions of mathematical feature data can be fused based on business analysis to better represent the parameter data.

[0072] Table 1

[0073] Maximum Minimum Mean Terquartile Median Septile Variance Standard deviation Skewness Kurtosis Rate of change Extreme value Entropy Curvature ……

[0074] The above two steps of determining the main control factor affecting the sweet spot and extracting the feature data from the second data are not limited in sequence, and one of the steps can be executed first and then the other step is executed.

[0075] Step S103: Input the main control factor data and the feature data of the shale gas well to be predicted into the trained sweet spot prediction model, and output the sweet spot prediction data of the shale gas well to be predicted.

[0076] The obtained main control factor data of the shale gas well to be predicted and the feature data of each parameter in the second data of the shale gas well to be predicted are input into the ensemble learning base model included in the sweet spot prediction model; the main control factor data can be adjusted according to the actual formation characteristics to improve the prediction accuracy of the sweet spot prediction model;

[0077] The output results of the integrated learning base models are aggregated using a preset combination strategy to obtain sweet spot prediction data output by the sweet spot prediction model.

[0078] Before predicting the sweet spot of the target shale gas well, the sweet spot prediction model can be trained in advance using known data of the target area, so that the prediction model can more accurately predict the sweet spot of the target shale gas well to be predicted in the target area. The training process of the sweet spot prediction model is as shown in Figure 3

[0079] Step S201: determining the main control factors affecting the sweet spot according to the correlation between each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot, extracting the main control factor data of the sampling shale gas well in the target area from the prediction data, and extracting the features of each parameter in the second data to obtain the feature data; the prediction data includes the first data and the second data, the first data includes the geological data, the perforation data and the oil and gas production data, and the second data includes the logging data and the fracturing operation data.

[0080] In this step, the implementation process of determining the main control factors affecting the sweet spot and performing feature extraction can refer to step S102. After determining the main control factors affecting the sweet spot, the main control factor data of the shale gas well can be obtained from the prediction data.

[0081] Step S202: constructing a training data set according to the main control factor data and the feature data of the shale gas well to be predicted.

[0082] The main control factor data of each well in the sampling shale gas well in the target area and the feature data of each parameter in the second data of each well, and the sweet spot data of each well jointly constitute the training data set. The training data set constructed includes the main control factor data of the sampling shale gas well, the feature data of each parameter in the second data, and the sweet spot data. The training data set can include multiple sample data. The main control factor data, the feature data of each parameter in the second data and the corresponding sweet spot data of one shale gas well can be taken as one sample data, and the sample data of multiple shale gas wells jointly constitute the training data set.

[0083] Step S203: training the sweet spot prediction model based on the training data set; the sweet spot prediction model includes at least three integrated learning base models and a combination strategy module connected to the output end of the base model, and the combination strategy module is used to aggregate the output results of the base model according to a preset combination strategy.

[0084] ​The ensemble learning base model selected by the sweet spot prediction model can be selected according to requirements. The base model can be any model suitable for ensemble learning, or a neural network model. When the difference between the selected ensemble learning base models is not large and the homogeneity is small, the prediction effect will be better. Each ensemble learning base model is actually a group of individual learners, such as individual learner 1, individual learner 2,..., and individual learner T. The principle diagram for realizing ensemble learning based on each individual learner is shown in FIG. 1. Figure 4

[0085] In this step, the main control factor data of each well in the training data set, and the characteristic data of each parameter in the second data of each well are input into the ensemble learning base model included in the sweet spot prediction model; the main control factor data can be adjusted according to the actual formation characteristics to improve the prediction accuracy of the sweet spot prediction model;

[0086] The output results of the ensemble learning base model are aggregated using a preset combination strategy to obtain sweet spot prediction data output by the sweet spot prediction model;

[0087] Based on the sweet spot prediction data output by the sweet spot prediction model and the sweet spot data in the training data set, the loss of each base model is determined and the base model parameters are adjusted. After multiple iterations of training, a trained base model with a required model loss is obtained. The trained base models are integrated to obtain a sweet spot prediction model.

[0088] In some optional embodiments, the ensemble learning base model in the sweet spot prediction model can select multiple different base models, such as but not limited to gradient BP neural network, XGBoost algorithm model, and random forest algorithm model. The implementation principle of ensemble learning based on BP neural network, XGBoost algorithm model, and random forest algorithm model can be seen in FIG. 2. Figure 5

[0089] In some optional embodiments, the sweet spot prediction data of each base model of the ensemble learning is collected based on a voting regressor, and the sweet spot prediction data of each base model collected is aggregated according to a preset combination strategy. The combination strategy includes simple average method and weighted average method.

[0090] The weighted average method considers the weights of each base learner. These weights are obtained during the training process. When calculating the average value, different weights are assigned to the output of each base learner, and the output result of the model is confirmed by combining the weights of each base learner. The simple average method is to sum the outputs of each base learner and divide the number of base learners to obtain the data as the output result of the model.

[0091] ​​Steps S201-S203 introduce the training process of the dessert prediction model. The model is formed by integrating multiple base models to form a more accurate model. The training data is used to train each base model in turn, and the outputs of each base model are aggregated by a predetermined aggregation strategy to obtain the prediction result of the dessert, improving the accuracy and prediction efficiency of the prediction result.

[0092] In this embodiment, the principle diagram of the dessert prediction model integration learning in the dessert prediction model training and dessert prediction process can be referred to as shown in Figure 5 The training data is input into the base models including BP neural network, XGBoost algorithm model and random forest algorithm model, and the prediction results output by each base model are combined by the combination strategy module to obtain the final prediction result.

[0093] Before predicting the dessert and training the dessert prediction model, in order to improve the data quality, the prediction data of the sampling shale gas well in the target area obtained can be preprocessed in advance, the prediction data including first data and second data, the first data including geological data, perforation data and oil and gas production data, and the second data including logging data and fracturing operation data; the preprocessing process includes:

[0094] The missing parameters in the logging data are supplemented by the data calculated by the formula corresponding to the parameters;

[0095] The complete prediction data is cleaned, including:

[0096] 1) Data deduplication: remove duplicate records in the prediction data.

[0097] This is achieved by comparing unique identifiers or key fields in the records. For example, use the unique identifiers such as well number, interval number and depth to match the records in the prediction data, and delete redundant data records.

[0098] 2) Missing value processing: if the quality of the prediction data is high, fill in the missing values in the prediction data, the missing values being filled in using the difference, average, median, mode, etc. of other data records for the parameter or based on similar wells and adjacent wells data; if the quality of the prediction data is poor, the missing values are removed.

[0099] 3) Abnormal value processing: detect and process abnormal values in the prediction data, and delete or replace the abnormal values with acceptable values.

[0100] 4) Data normalization and standardization: standardize the data format in the prediction data to a consistent format.

[0101] 5) Data conversion: converting the data format in the prediction data to facilitate subsequent processing and analysis.

[0102] In the above method of the embodiment, based on the obtained main control factor data and feature data of the shale gas well to be predicted affecting the sweet spot, the prediction data of the sweet spot is determined by using the trained sweet spot prediction model, the model is constructed based on at least three kinds of integrated learning base models and a combination strategy module connected with the output end of the base model, meanwhile, various influencing factors of the sweet spot are comprehensively considered, the prediction efficiency and the accuracy of the prediction result are improved, the accurate prediction of the shale gas well sweet spot can be realized, and the fracturing design is effectively guided.

[0103] Embodiment two

[0104] The embodiment two of the present application provides a specific implementation process of a shale gas well sweet spot prediction method based on integrated learning, a flowchart is shown in Figure 6 , a module principle diagram is shown in Figure 7 , and includes the following steps:

[0105] In the embodiment, taking the integrated learning base model in the sweet spot prediction model as a BP neural network, an XGBoost algorithm model and a random forest algorithm model, and taking the weighted average as an example of the combination strategy module, the specific steps of predicting the sweet spot of the shale gas well to be predicted are introduced.

[0106] Step S301: obtaining the prediction data of the sampling shale gas well in the target area, the prediction data including first data and second data, the first data including geological data, perforation data and oil and gas production data, and the second data including logging data and fracturing operation data.

[0107] The first data can be obtained from multiple channels, the geological data including horizontal shale gas well length, horizon, reservoir effective thickness, formation temperature, lithology, etc., the perforation data including perforation cluster number, cluster spacing, hole diameter, hole density, hole number, etc., and the oil and gas production data including gas production, liquid production, water cut, etc.

[0108] In the second data, the logging data includes acoustic time difference, bulk density, natural gamma, shale content, permeability, etc., and the fracturing operation data includes pump pressure, casing pressure, displacement, sand volume, sanding intensity, liquid intensity, total amount of proppant, and amount of liquid into the ground.

[0109] Step S302: preprocessing the obtained prediction data.

[0110] The preprocessing operation has been introduced in the embodiment one, and will not be repeated here.

[0111] Step S303: According to the correlation between each parameter in the prediction data of the sampling shale gas well in the target area and the sweet spot, the main control factor affecting the sweet spot is determined, and the main control factor data is extracted.

[0112] Referring to step S102, the main control factor analysis is performed on each parameter in the prediction data and the sweet spot to determine the main control factor affecting the sweet spot, and the interference of artificial subjective selection factor is excluded: the correlation coefficient between each parameter in the prediction data and the sweet spot is determined by the pearson correlation analysis method, and the parameter with a correlation coefficient of 0.4 or more is selected as the main control factor affecting the sweet spot; the main control factor can be adaptively adjusted according to the actual formation characteristics.

[0113] Step S304: Feature extraction is performed on each parameter in the second data to obtain feature data.

[0114] Referring to step S102, each parameter in the second data is converted into mathematical feature data of different dimensions in Table 1, and the principle of feature extraction on the logging data and the fracturing operation data in the second data is as shown in Figure 8 In the following model training process, different dimensions of features can be fused, and different dimensions of features can be comprehensively utilized to improve the expression ability of the model and better predict the sweet spot.

[0115] Step S305: According to the main control factor data and the feature data of the shale gas well, a training data set is constructed.

[0116] The main control factor data of the sampling shale gas well in the prediction data, the feature data of each parameter in the second data of each well, and the corresponding sweet spot data constitute a training data set, which can be stored in a sample library, and the constructed sample library can be used in the subsequent model training process.

[0117] Step S306: The sweet spot prediction model is trained based on the training data set.

[0118] In this embodiment, the BP neural network is a forward feedback neural network model in the integrated learning base model included in the sweet spot prediction model, which is trained and optimized by a back propagation algorithm. The BP neural network is composed of an input layer, a hidden layer and an output layer, wherein the hidden layer can contain multiple levels. Each level is composed of multiple neuron nodes, and the connection between the neurons has a weight. The input layer receives the original data, the hidden layer calculates and transmits the information, and the output layer finally gives the prediction result of the network, and the principle is as shown in Figure 9 .

[0119] The XGBoost algorithm model has significant improvements in accuracy, training speed and prevention of overfitting by introducing a regularization term, optimizing a loss function, column sampling, missing value processing, parallel processing, shrinkage mechanism, approximate histogram algorithm and multi-threading technology.

[0120] The random forest algorithm model can prevent some feature data from having too much influence on the entire model by randomly selecting master factor data and randomly selecting feature data, thereby improving the diversity and robustness of the model.

[0121] The specific process of training based on the base model includes:

[0122] 1) Input the master factor data of each well in the training data set and the feature data of each parameter in the second data of each well into the BP neural network, and the training data set is forward propagated from the input layer of the BP neural network, passes through the hidden layer, and reaches the output layer: the neurons of each layer are weighted and summed, and processed by the activation function to obtain the output of the layer, and the output is passed to the next layer until the output of the BP neural network output layer is obtained;

[0123] Determine the error between the final output of the BP neural network and the actual shale gas well sweet spot value;

[0124] Based on the error, the error is propagated from the output layer to the hidden layer according to the gradient descent method until the input layer: in the back propagation process, the weights of each layer of neurons are updated according to the error, and the error of the next forward propagation is reduced;

[0125] Repeat the above steps until the output error of the BP neural network output layer reaches the preset acceptance level.

[0126] 2) Input the master factor data of each well in the training data set and the feature data of each parameter in the second data of each well into the XGBoost algorithm model, and the XGBoost algorithm model continuously adds new trees to fit the residual error of the last prediction by feature splitting, and finally trains k trees to minimize the residual error.

[0127] 3) Randomly divide the master factor data of each well in the training data set and the feature data of each parameter in the second data of each well into different sub-training sets, and the sub-training set is composed of master factor data and feature data. A decision tree is constructed on each sub-training set using a decision tree algorithm, and each decision tree is trained; the prediction results of multiple decision trees are averaged or weighted to obtain the final regression result.

[0128] The trained base models are integrated into a trained sweet spot prediction model.

[0129] The trained dessert prediction model can be evaluated by a sample data set in a sample library, or evaluated by a business expert. The qualified model can be released for user use.

[0130] Step S307: obtaining prediction data of a to-be-predicted shale gas well in a target area, determining a main control factor affecting the dessert according to the correlation between each parameter of the prediction data and the dessert, and performing feature extraction on each parameter in the second data to obtain feature data.

[0131] In this step, the obtained main control factor data and feature data of the to-be-predicted shale gas well are used as input data of the trained dessert prediction model.

[0132] Step S308: inputting the main control factor data and the feature data of the to-be-predicted shale gas well into the trained dessert prediction model to output dessert prediction data of the to-be-predicted shale gas well.

[0133] The main control factor data and the feature data of the to-be-predicted shale gas well are input into a BP neural network, an XGBoost algorithm model and a random forest algorithm model respectively, and correspondingly, first dessert prediction data, second dessert prediction data and third dessert prediction data are obtained.

[0134] The first dessert prediction data, the second dessert prediction data and the third dessert prediction data are collected based on a voting regressor, and the three dessert prediction data are aggregated according to a weighted average strategy to obtain final dessert prediction data.

[0135] In this embodiment, the shale gas well dessert model integrates the BP neural network, the XGBoost algorithm model and the random forest algorithm model, and comprehensively considers various factors affecting the dessert, which greatly improves the prediction efficiency of the shale gas well and the accuracy of the prediction result. The obtained dessert prediction result can effectively guide the fracturing design.

[0136] Based on the same inventive concept, the embodiments of the present application also provide a shale gas well dessert prediction device based on ensemble learning. The device can be arranged in a computer device with computing processing function. The structure of the device is shown in Figure 10 The device comprises:

[0137] A first data acquisition unit 10 is configured to acquire prediction data of a to-be-predicted shale gas well in a target area, wherein the prediction data comprises first data and second data, the first data comprises geological data, perforation data and oil and gas production data, and the second data comprises logging data and fracturing operation data; main control factor data is extracted according to a pre-determined main control factor affecting the dessert; and feature data is obtained by performing feature extraction on each parameter in the second data.

[0138] The dessert prediction unit 20 is configured to input the main control factor data and the feature data of the shale gas well to be predicted into the trained dessert prediction model, and output dessert prediction data of the shale gas well to be predicted.

[0139] The second data acquisition unit 30 is configured to determine the main control factor affecting the dessert according to the correlation between each parameter in the prediction data of the sampling shale gas well in the target area and the dessert, extract the main control factor data from the prediction data of the sampling shale gas well in the target area, and perform feature extraction on each parameter in the second data to obtain the feature data.

[0140] The training data set is constructed according to the main control factor data and the feature data of the shale gas well; the training data set includes the main control factor data of each sampling shale gas well, the feature data of each parameter in the second data, and the dessert data; the dessert data of the sampling shale gas well is data evaluated based on each parameter in the prediction data.

[0141] The model training unit 40 is configured to train the dessert prediction model based on the training data set; the dessert prediction model includes at least three ensemble learning base models and a combination strategy module connected to the output end of the base model, and the combination strategy module is configured to aggregate the output results of the base model according to a preset combination strategy.

[0142] As to the shale gas well dessert prediction device based on ensemble learning in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be described in detail here.

[0143] The embodiment of the present application further provides a computer storage medium, which stores computer executable instructions, and the computer executable instructions are executed by a processor to implement the shale gas well dessert prediction method based on ensemble learning.

[0144] The embodiment of the present application further provides a prediction device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the shale gas well dessert prediction method based on ensemble learning when executing the program.

[0145] The embodiment of the present application further provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the shale gas well dessert prediction method based on ensemble learning.

[0146] Unless specifically stated otherwise, terms such as processing, computing, calculating, determining, displaying, and the like, can refer to an action or process of one or more processing or computing systems, or similar devices, that manipulate or transform data represented as physical (e.g., electronic) quantities within the systems' registers or memories into other data similarly represented as physical quantities within the systems' memories, registers or other such information storage, transmission or display devices. The terms "information," "data," "instructions," “command,” “signal,” “bit,” “symbol,” and the like refer to physical quantities presumed to represent a pertinent physical reality.

[0147] It should be understood that the particular order in which the steps in the disclosed processes have been presented is exemplary. Based on design preferences, it is understood that the particular order of steps in the processes can be rearranged without departing from the scope of the disclosure. The accompanying method claims present elements of the various steps in exemplary order and are not meant to be limited to the specific order or hierarchy presented.

[0148] In the above detailed description, various features are grouped together in single embodiments for the purpose of streamlining the disclosure. This disclosed approach is not to be interpreted as reflecting an intention that the embodiments of the claimed subject matter require more features than are expressly recited in each claim. Rather, as the claims below reflect, inventive subject matter lies in fewer than all features of the disclosed single embodiments. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate preferred embodiment.

[0149] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0150] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.

[0151] For a software implementation, the techniques described herein can be implemented with modules (e.g., procedures, functions, and so on) that perform the functions described herein. The software codes can be stored in memory units and executed by processors. The memory unit can be implemented within the processor or external to the processor, in which case it can be communicatively coupled to the processor via various means as is known in the art.

[0152] The above description includes one or more examples of the embodiments. Of course, not all possible combinations of components or methods described above can be claimed as embodiments. One of ordinary skill in the art can recognize that modifications and variations of the described embodiments can be made without departing from the scope of the present disclosure. It is therefore intended that the embodiments described herein be considered in all respects as illustrative and not restrictive, particularly as numerous modifications and further embodiments can become apparent to those skilled in the art. Accordingly, the scope of the present disclosure is intended to be defined by the following claims rather than the description. Moreover, the use of the terms "first", "second", etc. do not denote any order or importance, but rather the terms are used to distinguish one element from another. Furthermore, the use of the terms "including", "containing", etc. are meant to encompass the terms "consisting of" and / or "consisting essentially of". Moreover, the use of the term "or" is meant to encompass "and / or", unless otherwise indicated.

Claims

1. A method for predicting sweet spots in shale gas wells based on ensemble learning, characterized in that, include: Acquire prediction data for shale gas wells to be predicted in the target area. The prediction data includes first data and second data. The first data includes geological data, perforation data, and oil and gas production data. The second data includes well logging data and fracturing operation data. Based on the predetermined main controlling factors affecting desserts, extract the main controlling factor data; and perform feature extraction on each parameter in the second data to obtain feature data; The main controlling factor data and feature data of the shale gas well to be predicted are input into the trained sweet spot prediction model, and the sweet spot prediction data of the shale gas well to be predicted are output; wherein: the training process of the sweet spot prediction model includes: Based on the correlation between the parameters and sweet spots in the predicted data of shale gas wells sampled in the target area, the main controlling factors affecting the sweet spots are determined, and the main controlling factor data is extracted from the predicted data of shale gas wells sampled in the target area; and feature extraction is performed on each parameter in the second data to obtain feature data; the sweet spot data of the sampled shale gas wells is data obtained based on the evaluation of each parameter in the predicted data; A training dataset is constructed based on the main control factor data and feature data of the sampled shale gas wells; the training dataset includes the main control factor data of the sampled shale gas wells, the feature data of each parameter in the second data, and the sweet spot data. The dessert prediction model is trained based on the training dataset; the dessert prediction model includes at least three ensemble learning base models and a combination strategy module connected to the output of the base models. The combination strategy module is used to aggregate the output results of the base models according to a preset combination strategy.

2. The method as described in claim 1, characterized in that, The process involves determining the main controlling factors influencing the sweet spot based on the correlation between various parameters of the predicted data from shale gas wells sampled in the target area and the sweet spot, including: The correlation coefficients between each parameter and the sweet spot in the predicted data of sampled shale gas wells in the target area are determined by Pearson correlation analysis. Based on the correlation coefficients between each parameter and the sweet spot, the main controlling factors affecting the sweet spot are screened out from each parameter.

3. The method as described in claim 2, characterized in that, The correlation coefficients between each parameter in the predicted data of sampled shale gas wells in the target area and the sweet spot of shale gas wells are determined according to the Pearson correlation analysis method, including: Based on the predicted data of shale gas wells sampled in the target area, as well as the sweet spot data of the shale gas wells, the correlation coefficient between each parameter and the sweet spot is calculated using the following formula: Where X represents the parameter data in the predicted data of sampled shale gas wells in the target area, Y represents the sweet spot of the shale gas well, and r xy The correlation coefficient between the parameter and the dessert is represented by N, where N represents the total number of parameters and n represents the nth parameter to be processed, with a value ranging from 1 to N.

4. The method as described in claim 1, characterized in that, The step of extracting features from each parameter in the second data to obtain feature data includes: The parameter data in the second data is converted into multidimensional feature data; the multidimensional feature data includes at least one of the following: maximum value, minimum value, average value, tertiles, median, quartiles, variance, standard deviation, skewness, kurtosis, rate of change, extreme values, entropy, and curvature.

5. The method as described in claim 1, characterized in that, Also includes: Preprocessing is performed on the prediction data of the shale gas wells to be predicted in the target area, and / or on the prediction data of the sampled shale gas wells in the target area: The missing parameters in the well logging data are supplemented using data calculated from the formulas corresponding to those parameters; Data cleaning is performed on the forecast data, including data deduplication, missing value handling, outlier handling, data normalization and standardization, and data transformation.

6. The method as described in claim 1, characterized in that, The training of the dessert prediction model based on the training dataset includes: The main control factor data of each shale gas well sampled from the training dataset and the feature data of each parameter in the second data of each well are input into the ensemble learning base model included in the sweet spot prediction model; The output of the ensemble learning base model is aggregated using a preset combination strategy to obtain dessert prediction data output by the dessert prediction model. Based on the dessert prediction data output by the dessert prediction model and the dessert data in the training dataset, the loss of each base model is determined and the base model parameters are adjusted. After multiple iterations of training, a well-trained dessert prediction model with a model loss that meets the requirements is obtained.

7. The method as described in claim 1, characterized in that, The ensemble learning base model includes at least three of the following: BP neural network, XGBoost algorithm model, random forest algorithm model, and gradient boosting algorithm model.

8. The method as described in claim 1, characterized in that, The outputs of the base model are aggregated according to a preset fusion strategy, including: The dessert prediction data of each base model of the ensemble learning is collected based on the voting regressor, and the collected dessert prediction data of each base model is aggregated according to the preset combination strategy. The combined strategies include simple averaging and weighted averaging.

9. A shale gas well sweet spot prediction device based on ensemble learning, characterized in that, include: The first data acquisition unit is used to acquire prediction data for shale gas wells to be predicted in the target area. The prediction data includes first data and second data. The first data includes geological data, perforation data, and oil and gas production data. The second data includes well logging data and fracturing operation data; based on the pre-determined main controlling factors affecting the sweet spot, main controlling factor data is extracted; and feature extraction is performed on each parameter in the second data to obtain feature data; The sweet spot prediction unit is used to input the main control factor data and feature data of the shale gas well to be predicted into the trained sweet spot prediction model and output the sweet spot prediction data of the shale gas well to be predicted. The second data acquisition unit is used to determine the main controlling factors affecting the sweet spot based on the correlation between each parameter in the predicted data of shale gas wells sampled in the target area and the sweet spot, extract the main controlling factor data from the predicted data of shale gas wells sampled in the target area; and perform feature extraction on each parameter in the second data to obtain feature data. A training dataset is constructed based on the main control factor data and feature data of shale gas wells. The training dataset includes the main control factor data of sampled shale gas wells, the feature data of each parameter in the second data, and the sweet spot data. The sweet spot data of sampled shale gas wells is data obtained by evaluating each parameter in the prediction data. The model training unit is used to train the dessert prediction model based on the training dataset. The dessert prediction model includes at least three ensemble learning base models and a combination strategy module connected to the output of the base models. The combination strategy module is used to aggregate the output results of the base models according to a preset combination strategy.

10. A computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which, when executed by a processor, implement the sweet spot prediction method for shale gas wells based on ensemble learning as described in any one of claims 1-8.

11. A prediction device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the shale gas well sweet spot prediction method based on ensemble learning as described in any one of claims 1-8.

12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the shale gas well sweet spot prediction method based on ensemble learning as described in any one of claims 1-8.

Citation Information

Cited By

  • Method and system for predicting gold mine target area in shallow coverage area

    CN121996991A

  • Geological double-dessert prediction method, device and equipment based on knowledge graph representation learning, medium and product

    CN122066100A

  • Geological double sweet spot prediction method and device based on knowledge graph representation learning, equipment, medium and product

    CN122066100B