A hydrocarbon source rock TOC machine learning prediction method and system

By using a seismic attribute-driven machine learning method to screen sensitive attributes from seismic data for TOC prediction of source rocks, the prediction problem in low-exploration areas is solved, and reliable prediction of 3D source rock TOC is achieved, reducing the dependence on well data.

CN116776149BActive Publication Date: 2026-02-10CHINA NATIONAL OFFSHORE OIL (CHINA) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310725124.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-19
Publication Date
2026-02-10
Estimated Expiration
2043-06-19

AI Technical Summary

Technical Problem

In areas with limited exploration and a lack of well data, traditional methods based on well logging data and geophysical inversion are insufficient for effective three-dimensional prediction of the total hydrocarbon source rock (TOC).

Method used

By using a seismic attribute-driven machine learning method, a large number of seismic attributes are calculated from seismic data, sensitive attributes are screened out, and predictions are made through machine learning models to achieve the prediction of the three-dimensional source rock TOC.

Benefits of technology

In a very limited number of wells, reliable three-dimensional source rock TOC prediction was achieved, reducing reliance on well logging data and providing a foundation for exploration and development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116776149B_ABST
    Figure CN116776149B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of hydrocarbon source rock TOC machine learning prediction method and system, comprising the following steps: based on the training sample set data obtained in advance, various prediction models previously built are trained, and according to the prediction effect of each prediction model, the optimal prediction model is screened out;Based on the optimal prediction model screened out, the current seismic data is predicted, and three-dimensional hydrocarbon source rock TOC prediction result is obtained.The present application effectively overcomes the problem that the TOC of the hydrocarbon source rock in the low exploration area is difficult to predict, reduces the dependence on well logging data, and more fully utilizes seismic data, and lays a foundation for further exploration and development work.Therefore, the present application can be widely applied in the field of oil and gas exploration and development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oil and gas exploration and development, specifically to a seismic attribute-driven machine learning method and system for predicting the Total Organic Carbon Content (TOC) of source rocks. Background Technology

[0002] The key to evaluating oil and gas basins lies in the presence of source rocks. As oil and gas exploration continues to move towards lower exploration areas (deep lower layers, deep water, or new exploration areas), the prediction of the spatial distribution of source rocks has become a crucial indicator for evaluating oil and gas basins. A common process for predicting the spatial distribution of source rocks involves using geophysical logging and seismic data to depict the spatial distribution and variability of source rock parameters. Common characteristic parameters of source rocks include TOC (Total Thermal Content), vitrinite reflectance, organic matter abundance, and peak pyrolysis temperature. Borehole TOC data is generally derived from geochemical analysis of core data.

[0003] Geophysical methods for predicting the total organic carbon (TOC) of source rocks include well logging data fitting and comprehensive elastic parameter prediction. Commonly used well logging data fitting methods include the ΔlogR method and the four-parameter method. Three-dimensional TOC prediction of source rocks relies on the numerical relationship between elastic parameters such as elastic impedance and TOC. These methods all depend on well data, and the prediction accuracy of source rock TOC is positively correlated with the number of wells in the study area. However, low-exploration areas are generally low-well areas, lacking evaluation parameters and well logging data for source rocks. Therefore, traditional TOC prediction methods based on well logging data and geophysical inversion methods are difficult to implement. Summary of the Invention

[0004] To address the aforementioned problems, the purpose of this invention is to provide a seismic attribute-driven TOC machine learning prediction method and system for source rocks. This method fully utilizes the properties of seismic data and achieves three-dimensional TOC prediction of source rocks based on only a very small number of wells. It effectively overcomes the difficulty of predicting TOC of source rocks in low-exploration areas and lays the foundation for further exploration and development.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] In a first aspect, the present invention provides a machine learning method for predicting the TOC (Total Organic Carbon) of source rocks, comprising the following steps:

[0007] Based on the pre-acquired training sample set data, various pre-built prediction models are trained, and the optimal prediction model is selected according to the prediction effect of each prediction model.

[0008] Based on the selected optimal prediction model, predictions are made on the currently acquired seismic data to obtain the three-dimensional source rock TOC prediction results.

[0009] Furthermore, the step of training various pre-built prediction models based on pre-acquired training sample set data, and selecting the optimal prediction model based on the prediction performance of each model, includes:

[0010] Obtain the training sample set;

[0011] The training sample set is divided into a test set and a training set according to a preset ratio;

[0012] Several pre-built prediction models are trained using the training set, and the prediction performance of each model is verified using the test set. The prediction model with the best prediction performance is selected as the optimal prediction model.

[0013] Furthermore, obtaining the training sample set includes:

[0014] Calculations are performed on earthquake attribute data based on earthquake data;

[0015] Based on the calculated seismic attribute data, the sensitive attributes of source rock TOC were selected and used as the training sample set.

[0016] Furthermore, the seismic attribute data includes signal analysis type and geophysical inversion type attributes. The signal analysis type seismic attributes include at least one of instantaneous amplitude, instantaneous frequency, coherence, half-time energy, and arc length. The geophysical inversion type seismic attributes include at least one of elastic impedance, density, Lamé parameter, and Young's modulus.

[0017] Furthermore, based on the calculated seismic attribute data, the sensitive attributes of the source rock TOC are selected and used as a training sample set, including:

[0018] Draw cross-plots of source rock TOC with different seismic attributes to qualitatively screen seismic attributes that are correlated with source rock TOC;

[0019] The Pearson correlation coefficients between the source rock TOC and each selected seismic attribute were calculated. Based on the magnitude of the absolute value of the correlation coefficients, the attributes most correlated with the source rock TOC were selected as the training sample set.

[0020] Furthermore, the pre-built prediction models include at least one of the following ensemble methods: gradient boosting tree, extreme gradient boosting tree, random forest, histogram-based gradient boosting tree, and local cascade ensemble method.

[0021] Secondly, the present invention provides a source rock TOC machine learning prediction system, comprising:

[0022] The model selection module is used to train various pre-built prediction models based on pre-acquired training sample set data, and select the optimal prediction model based on the prediction performance of each prediction model.

[0023] The prediction module is used to predict the current seismic data based on the selected optimal prediction model, and obtain the three-dimensional source rock TOC prediction results.

[0024] Furthermore, the model selection module includes:

[0025] The data acquisition module is used to acquire the training sample set;

[0026] The data partitioning module is used to divide the training sample set into a test set and a training set according to a preset ratio;

[0027] The model training module is used to train several pre-built prediction models using the training set, and to verify the prediction performance of each prediction model using the test set, and to select the prediction model with the best prediction performance as the optimal prediction model.

[0028] Thirdly, the present invention provides a computer-readable storage medium for storing one or more programs, said one or more programs including instructions that, when executed by a computing device, cause the computing device to perform any of the methods.

[0029] Fourthly, the present invention provides a computing device comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods.

[0030] Due to the adoption of the above technical solutions, this invention has the following advantages: This invention makes full use of the properties of seismic data, and achieves three-dimensional TOC prediction of source rocks based on only a very small number of wells. It effectively overcomes the problem of high difficulty in predicting TOC of source rocks in low exploration areas, reduces the dependence on well logging data, and makes fuller use of seismic data, laying the foundation for further exploration and development work.

[0031] Therefore, this invention can be widely applied in the field of oil and gas exploration and development. Attached Figure Description

[0032] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. In the drawings:

[0033] Figure 1 This is a flowchart of the TOC machine learning prediction method for source rocks provided in this embodiment of the invention;

[0034] Figure 2 These are the TOC curve of well A and some well-side seismic attribute data in this embodiment of the invention;

[0035] Figure 3 This is a correlation analysis between TOC and seismic attribute data in an embodiment of the present invention;

[0036] Figures 4a-4c This is a cross-sectional analysis of TOC and seismic attributes in an embodiment of the present invention, wherein, Figure 4a It is a cross-sectional analysis of TOC, arc length, EHT, instantaneous frequency, and instantaneous amplitude; Figure 4b It is a cross-section analysis of TOC, EI, Lambda, Mu, and density; Figure 4c It is a cross-analysis of TOC, kurtosis, variance, principal amplitude, and integrated energy spectrum;

[0037] Figures 5a-5d These are the four characteristic parameters (3D) of TOC in this embodiment of the invention, wherein, Figure 5a For EI; Figure 5b For Lambda; Figure 5c Density; Figure 5d For EHT;

[0038] Figure 6 This is the TOC prediction result of source rocks based on three-dimensional seismic attributes in the embodiments of the present invention;

[0039] Figure 7 This is a comparison between the predicted TOC value and the actual value at well C in this embodiment of the invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0041] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0042] Considering the lack of well data in low-exploration areas, some embodiments of this invention provide a machine learning method for predicting the TOC of source rocks. This method calculates a large number of seismic attributes through signal analysis and geophysical inversion, and combines the TOC curves of a few wells (at least one well) to screen for sensitive TOC attributes. Leveraging the excellent data fitting properties of machine learning, a reliable three-dimensional source rock TOC prediction data volume is ultimately obtained.

[0043] Correspondingly, in other embodiments of the present invention, a source rock TOC machine learning prediction system, device, and storage medium are provided.

[0044] Example 1

[0045] like Figure 1 As shown in this embodiment, a machine learning method for predicting the TOC of source rocks is provided. Starting from seismic data, a large number of seismic attributes are calculated through signal analysis and geophysical inversion to obtain the sensitive attributes of source rock TOC. Then, machine learning methods are used to predict the TOC of source rocks, thereby realizing the prediction of the spatial distribution of source rocks. Specifically, it includes the following steps:

[0046] (1) Use the training sample set to train various pre-built prediction models, and select the optimal prediction model based on the prediction effect of each prediction model.

[0047] Specifically, it includes the following steps:

[0048] (1.1) Obtain the training sample set.

[0049] (1.1.1) Calculate earthquake attribute data based on earthquake data.

[0050] Seismic data contains rich information about the subsurface medium. Different types of seismic attributes possess different properties and can reflect the characteristics of the subsurface medium from various perspectives. Therefore, in this embodiment, the seismic attributes to be calculated are divided into two types: signal analysis attributes and geophysical inversion attributes. Signal analysis attributes mainly include instantaneous amplitude, instantaneous frequency, coherence, half-time energy, and arc length; geophysical inversion attributes mainly include elastic impedance, density, Lamé parameters (λ and μ), and Young's modulus.

[0051] (1.1.2) Based on the calculated seismic attribute data, the sensitive attributes of source rock TOC are selected as the training sample set.

[0052] Specifically, it includes the following steps:

[0053] (1.1.2.1) Draw the cross-analysis diagram of source rock TOC and different seismic attributes, and qualitatively screen the seismic attributes that are correlated with source rock TOC (as shown in Figure 4).

[0054] (1.1.2.2) Calculate the Pearson correlation coefficients between the source rock TOC and the different seismic attributes selected in step (1.2.1) (e.g., Figure 3 As shown in the figure, the attributes most relevant to the source rock TOC are selected as the training sample set for machine learning by the magnitude of the absolute value of the correlation coefficient.

[0055] (1.2) Divide the training sample set into a test set and a training set according to the preset ratio.

[0056] Based on the well locations within the study area, seismic attribute data from the wellbore access points were extracted and divided into test and training sets.

[0057] (1.3) Train the various pre-built prediction models using the training set, and verify the prediction effect of each prediction model using the test set. The prediction model with the best prediction effect is taken as the optimal prediction model.

[0058] In this embodiment, the constructed prediction model includes Gradient Boosting Tree (GBDT), Extreme Gradient Boosting Tree (XGBoost), Random Forest (RF), Histogram-based Gradient Boosting Tree (HGBDT), and Local Cascaded Ensemble (LCE) methods from the ensemble approach. XGBoost and HGBDT are improvements on GBDT, while LCE is a fusion of XGBoost and RF. These methods will be evaluated using the same set of criteria, and the method with the highest prediction accuracy will be selected for application to a wider range of research areas.

[0059] (2) Based on the selected optimal prediction model, the current collected seismic data is predicted to obtain the three-dimensional source rock TOC prediction results.

[0060] Example 2

[0061] This embodiment uses a three-dimensional work area as an example to achieve three-dimensional TOC prediction of source rocks using the method provided in Embodiment 1. The prediction is then compared with that of a verification well within the work area to verify the reliability of the machine learning model.

[0062] like Figure 2 The image shows the TOC curve of well A and some seismic attribute data from the well bypass. Figure 3 This involves a correlation analysis between TOC and seismic attributes. Figure 4 shows the cross-plot analysis between TOC and seismic attributes. As can be seen from the figure, the sensitive attributes of TOC in the study area of ​​this embodiment are elastic impedance, λ, density, and half-time energy.

[0063] Figure 5 shows the three-dimensional TOC sensitive properties of the source rock, namely elastic impedance, λ, density, and half-time energy.

[0064] Figure 6 This is a TOC prediction result for source rocks based on three-dimensional seismic attributes. Figure 7 This is a comparison between the predicted and actual TOC values ​​at the location of verification well C. As can be seen from the figure, the source rock prediction method of this invention is reliable.

[0065] Example 3

[0066] The above-described embodiment 1 provides a source rock TOC machine learning prediction method. Correspondingly, this embodiment provides a source rock TOC machine learning prediction system. The system provided in this embodiment can implement the source rock TOC machine learning prediction method of embodiment 1. The system can be implemented through software, hardware, or a combination of both. For example, the system may include integrated or separate functional modules or units to execute the corresponding steps in the methods of embodiment 1. Since the system in this embodiment is basically similar to the method embodiment, the description process in this embodiment is relatively simple. For relevant details, please refer to the description of embodiment 1. The system embodiment provided in this embodiment is merely illustrative.

[0067] The source rock TOC machine learning prediction system provided in this embodiment includes:

[0068] The model selection module is used to train various pre-built prediction models based on pre-acquired training sample set data, and select the optimal prediction model based on the prediction performance of each prediction model.

[0069] The prediction module is used to predict the current seismic data based on the selected optimal prediction model, and obtain the three-dimensional source rock TOC prediction results.

[0070] Preferably, the model screening module includes:

[0071] The data acquisition module is used to acquire the training sample set;

[0072] The data partitioning module is used to divide the training sample set into a test set and a training set according to a preset ratio;

[0073] The model training module is used to train several pre-built prediction models using the training set, and to verify the prediction performance of each prediction model using the test set. The prediction model with the best prediction performance is selected as the optimal prediction model.

[0074] Example 4

[0075] This embodiment provides a processing device corresponding to the source rock TOC machine learning prediction method provided in Embodiment 1. The processing device can be a client-side processing device, such as a mobile phone, laptop, tablet computer, desktop computer, etc., to execute the method of Embodiment 1.

[0076] The processing device includes a processor, a memory, a communication interface, and a bus. The processor, memory, and communication interface are connected via the bus to enable communication between them. The memory stores a computer program that can run on the processor. When the processor runs the computer program, it executes the source rock TOC machine learning prediction method provided in Embodiment 1.

[0077] In some embodiments, the memory may be high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device.

[0078] In other embodiments, the processor can be a general-purpose processor of various types, such as a central processing unit (CPU) or a digital signal processor (DSP), and is not limited thereto.

[0079] Example 5

[0080] The source rock TOC machine learning prediction method of this embodiment 1 can be specifically implemented as a computer program product. The computer program product may include a computer-readable storage medium on which computer-readable program instructions for executing the source rock TOC machine learning prediction method of this embodiment 1 are loaded.

[0081] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A machine learning method for predicting the total organic carbon (TOC) of source rocks, characterized in that, Includes the following steps: Based on the pre-acquired training sample set data, various pre-built prediction models are trained, and the optimal prediction model is selected according to the prediction effect of each prediction model. Based on the selected optimal prediction model, the currently acquired seismic data is used to predict the TOC of the three-dimensional source rock. The process of training various pre-built prediction models based on pre-acquired training sample set data and selecting the optimal prediction model based on the prediction performance of each prediction model includes: acquiring a training sample set; dividing the training sample set into a test set and a training set according to a preset ratio; training several pre-built prediction models using the training set and verifying the prediction performance of each prediction model using the test set; and selecting the prediction model with the best prediction performance as the optimal prediction model. The process of obtaining the training sample set includes: calculating seismic attribute data based on seismic data; and selecting sensitive attributes of source rock TOC based on the calculated seismic attribute data as the training sample set. The process of selecting sensitive attributes of source rock TOC based on the calculated seismic attribute data and using them as a training sample set includes: combining the TOC curves of at least one well to draw a cross-sectional analysis diagram of source rock TOC and different seismic attributes, qualitatively screening seismic attributes that are correlated with source rock TOC; calculating the Pearson correlation coefficient between source rock TOC and each screened seismic attribute, and selecting the attributes most correlated with source rock TOC as a training sample set based on the magnitude of the absolute value of the correlation coefficient. The pre-built prediction models include at least one of the following ensemble methods: gradient boosting tree, extreme gradient boosting tree, random forest, histogram-based gradient boosting tree, and local cascade ensemble method.

2. The method for predicting TOC (Total Organic Carbon) of source rocks using machine learning as described in claim 1, characterized in that, The seismic attribute data includes signal analysis type and geophysical inversion type attributes. The signal analysis type seismic attributes include at least one of instantaneous amplitude, instantaneous frequency, coherence, half-time energy, and arc length. The geophysical inversion type seismic attributes include at least one of elastic impedance, density, Lamé parameter, and Young's modulus.

3. A machine learning prediction system for source rock TOC, characterized in that, include: The model selection module is used to train various pre-built prediction models based on pre-acquired training sample set data, and select the optimal prediction model based on the prediction performance of each prediction model. The prediction module is used to predict the current seismic data based on the selected optimal prediction model, and obtain the three-dimensional source rock TOC prediction results. The model selection module includes: The data acquisition module is used to acquire the training sample set; The data partitioning module is used to divide the training sample set into a test set and a training set according to a preset ratio; The model training module is used to train several pre-built prediction models using the training set, and to verify the prediction effect of each prediction model using the test set, and to select the prediction model with the best prediction effect as the optimal prediction model. The process of obtaining the training sample set includes: calculating seismic attribute data based on seismic data; and selecting sensitive attributes of source rock TOC based on the calculated seismic attribute data as the training sample set. The process of selecting sensitive attributes of source rock TOC based on calculated seismic attribute data and using them as a training sample set includes: drawing cross-sectional analysis diagrams of source rock TOC and different seismic attributes; qualitatively selecting seismic attributes that are correlated with source rock TOC; calculating the Pearson correlation coefficient between source rock TOC and each selected seismic attribute; and selecting the attributes most correlated with source rock TOC as a training sample set based on the magnitude of the absolute value of the correlation coefficient. The pre-built prediction models include at least one of the following ensemble methods: gradient boosting tree, extreme gradient boosting tree, random forest, histogram-based gradient boosting tree, and local cascade ensemble method.

4. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 2.

5. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 2.

Citation Information

Patent Citations

  • Method and system for predicting TOC (total organic carbon) of source rock in sparse well area, electronic equipment and medium

    CN115629414A