Random forest method-based surface porosity prediction method, apparatus and device, and storage medium

By fitting shale void ratio parameters using the random forest method, the problems of high acquisition cost and cumbersome steps in existing technologies are solved, achieving fast and accurate shale void ratio prediction and improving the efficiency of sweet spot identification.

CN121998134APending Publication Date: 2026-05-08CHINA NAT PETROLEUM CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA NAT PETROLEUM CORP
Filing Date
2024-11-04
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies for obtaining shale surface surface rate parameters are costly and involve cumbersome steps, making it difficult to efficiently identify shale deposits.

Method used

A machine learning approach based on random forests was adopted to fit the porosity parameters of shale using well logging data and core data, including organic porosity, inorganic porosity, and fracture porosity. The process of obtaining porosity parameters was simplified through model training and validation.

Benefits of technology

It enables rapid and accurate prediction of shale hole rate, reduces acquisition costs, facilitates shale oil exploration and development, and improves the efficiency of sweet spot identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998134A_ABST
    Figure CN121998134A_ABST
Patent Text Reader

Abstract

The invention discloses a surface porosity prediction method and device of a random forest method, equipment and a storage medium, and relates to the technical field of unconventional oil and gas exploration and development. The method is based on a machine learning method of a random forest, logging information and core data are used for fitting surface porosity parameters (including organic surface porosity, inorganic surface porosity and fracture surface porosity) of shale, a surface porosity RF model is trained, surface porosity data of other wells in a research block are predicted through the trained surface porosity RF model, and the surface porosity parameters of other wells in the research block are predicted. A convenient means is provided for shale dessert identification, the acquisition cost of the surface porosity parameter is reduced, and the acquisition process of the surface porosity parameter is simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unconventional oil and gas exploration and development technology, and more specifically to a face rate prediction method, apparatus, equipment and storage medium based on the random forest method. Background Technology

[0002] "Sweet spot" assessment is an important aspect of unconventional oil and gas exploration and development, and it is of great significance for the large-scale and efficient development of unconventional oil and gas. The connotation of the "sweet spot" concept is constantly expanding, and the selection of assessment parameters and the standards for their values ​​are becoming more diverse and regionally specific.

[0003] According to the dessert classification standards of a certain region, as shown in Table 1 below, the classification is required to be based on three surface area ratio parameters: inorganic surface area ratio, organic surface area ratio, and cracked surface area ratio.

[0004] Table 1 shows the criteria for classifying desserts. The three face rate parameters in the dessert classification criteria are generally obtained through digital core analysis, which requires a series of steps such as extracting the core and constructing the digital core, which is costly and cumbersome. Summary of the Invention

[0005] To overcome the defects and shortcomings of the existing technologies, this invention provides a facet ratio prediction method, apparatus, device, and storage medium based on the random forest method. The purpose of this invention is to overcome the problems of high cost and cumbersome procedures in obtaining facet ratios for three types of shale. This invention uses a machine learning method based on random forests to fit facet ratio parameters (including organic facet ratio, inorganic facet ratio, and fracture facet ratio) of shale using well logging data and core data, providing a convenient means for identifying sweet spots in shale, reducing the cost of obtaining facet ratio parameters, and simplifying the process of obtaining facet ratio parameters.

[0006] To address the problems existing in the prior art, the present invention is achieved through the following technical solution.

[0007] The first aspect of this invention provides a face ratio prediction method based on the random forest method, which includes the following steps: S1. Collect logging curves and core electron microscopy data from each well in the study block; S2. Align the core electron microscopy data of each well with the corresponding well logging data at depth; use part of the aligned data as the training sample set for the face rate RF model and the other part as the validation sample set for the trained face rate RF model. S3. Perform correlation analysis on the data in the training sample set in step S2, and select the curves with high correlation to the face rate data respectively. S4. Divide the selected curves into the training set and test set for training the face rate RF model; S5. Train the face ratio RF model using the training set, and test the trained face ratio RF model using the test set to verify the reliability of the face ratio RF model. If the prediction results deviate significantly from the actual results, readjust the parameters until the trained face ratio RF model passes the reliability test. S6. Use the validation sample set from step S2 to validate the face rate RF model trained in step S5. If the validation passes, it means that the face rate RF model can be applied in the study area; otherwise, repeat steps S3-S5. S7. Using the face rate RF model validated in step S6, predict the face rate data of other wells in the study block.

[0008] More preferably, the surface area ratio includes organic surface area ratio, inorganic surface area ratio, and fracture surface area ratio. Through RF model training and testing, organic surface area ratio RF model, inorganic surface area ratio RF model, and fracture surface area ratio RF model are obtained, and applied to other wells in the study block to predict the organic surface area ratio, inorganic surface area ratio, and fracture surface area ratio of other wells in the study block, respectively.

[0009] In a further preferred embodiment, in step S1, the logging data includes logging curves of the target layer and calculated mineral model data.

[0010] In a further preferred embodiment, in step S2, the depth of the core electron microscope data is used as the standard to align the logging data with the core electron microscope data.

[0011] In an even more preferred embodiment, in step S2, the data after aligning the core electron microscopy data with the well logging data depth is exported into a single file.

[0012] More preferably, in step S4, the ratio of the training set to the test set is 0.6~0.9:0.4~0.1; the sum of the training set ratio and the test set ratio does not exceed 1.

[0013] In a further preferred embodiment, in step S4, the ratio of the training set to the test set is 0.7:0.3.

[0014] In a further preferred embodiment, in step S5, the face rate RF model is trained by adjusting parameters, including the number of feature trees, the number of random seeds, and the number of leaves in the feature trees.

[0015] In a further preferred embodiment, in step S5, the reliability of the face rate RF model is verified by judging the regression plot or error plot of the test set. If the predicted result deviates significantly from the actual result, the parameters are readjusted.

[0016] In a further preferred embodiment, in step S6, before validating the face face ratio RF model with the data in the validation sample set, the number of curves and the order of the curves in the data input into the face face ratio RF model from the validation sample set are now aligned with the training set used when building the model.

[0017] In a further preferred embodiment, during model validation in step S6, the face rate predicted using well logging data is compared with the face rate obtained from core electron microscopy data. If the difference is small, the model is accurate and can be applied to the study area; otherwise, if the difference is large, the model is inaccurate and the face rate RF model needs to be retrained according to steps S3-S5.

[0018] In a further preferred embodiment, in step S7, when using the face surface rate RF model validated in step S6 to predict the face surface rate data of other wells in the study block, the number of curves and the order of curves in the logging data of the other wells are corresponding to the training set used when establishing the model.

[0019] A second aspect of the present invention provides a face rate prediction device based on a random forest method, the prediction device comprising: The model data collection module is used to collect logging curves and core electron microscopy data of each well in the study block, and to align the core electron microscopy data of each well with the depth of its corresponding logging data. The model data preprocessing module is used to perform correlation analysis on the depth-aligned data and select curves with high correlation to the face rate data. The model building module is used to train a face face ratio RF model based on curves with high correlation to face face ratio data, and to perform reliability and accuracy verification on the trained face face ratio RF model. If the reliability verification fails, the face face ratio RF model is retrained by adjusting the parameters; if the reliability verification passes but the accuracy verification fails, the correlation curves are re-selected through the model data preprocessing module and the model is trained again; until both the reliability verification and accuracy verification pass, a well-trained face face ratio RF model is obtained. The data prediction module uses the face rate RF model trained by the model building module, and substitutes the well logging data of other wells in the study block to predict the face rate of other wells.

[0020] A third aspect of the present invention provides a computer device including a processor, an input device, an output device, and a memory, wherein the processor, the input device, the output device, and the memory are interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to perform some or all of the steps as described in the first aspect of the present invention.

[0021] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of the present invention.

[0022] Compared with the prior art, the beneficial technical effects of the present invention are as follows: This invention utilizes a machine learning method based on random forests to fit porosity parameters (organic porosity, inorganic porosity, and fracture porosity) of shale using well logging data and core data. This provides a convenient means for identifying sweet spots in shale. All parameters in the method can be obtained from well logging data and laboratory data. The random forest fitting method using well logging data can quickly and accurately predict shale porosity, and based on this prediction, sweet spot layers can be identified, which has significant practical value in shale oil exploration and development. Attached Figure Description

[0023] Figure 1 This is a flowchart of the face rate prediction method based on the random forest method of the present invention; Figure 2 This is a heatmap of correlation analysis in Embodiment 2 of the present invention; Figure 3 This is a regression plot of the RF model test set in Embodiment 2 of the present invention; Figure 4 This is the accuracy verification result of the RF model in Embodiment 2 of the present invention. Detailed Implementation

[0024] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0025] Example 1 As a preferred embodiment of the present invention, please refer to the appendix to the specification. Figure 1 As shown, this embodiment discloses a face rate prediction method based on the random forest method, which includes the following steps: S1. Collect logging curves and core electron microscopy data from each well in the study block; S2. Align the core electron microscopy data of each well with the corresponding well logging data at depth; use part of the aligned data as the training sample set for the face rate RF model and the other part as the validation sample set for the trained face rate RF model. S3. Perform correlation analysis on the data in the training sample set in step S2, and select the curves with high correlation to the face rate data respectively. S4. Divide the selected curves into the training set and test set for training the face rate RF model; S5. Train the face ratio RF model using the training set, and test the trained face ratio RF model using the test set to verify the reliability of the face ratio RF model. If the prediction results deviate significantly from the actual results, readjust the parameters until the trained face ratio RF model passes the reliability test. S6. Use the validation sample set from step S2 to validate the face rate RF model trained in step S5. If the validation passes, it means that the face rate RF model can be applied in the study area; otherwise, repeat steps S3-S5. S7. Using the face rate RF model validated in step S6, predict the face rate data of other wells in the study block.

[0026] As an example of this embodiment, the surface rate includes organic surface rate, inorganic surface rate and fracture surface rate. Through RF model training and testing, organic surface rate RF model, inorganic surface rate RF model and fracture surface rate RF model are obtained, and applied to other wells in the study block to predict the organic surface rate, inorganic surface rate and fracture surface rate of other wells in the study block, respectively.

[0027] As another example of this embodiment, the logging data includes logging curves of the target layer and calculated mineral model data.

[0028] In one implementation of this embodiment, in step S2, the depth of the core electron microscope data is used as a standard to align the well logging data with the core electron microscope data. The data after depth alignment between the core electron microscope data and the well logging data is then exported into a single file.

[0029] In another implementation of this embodiment, the ratio of the training set to the test set is 0.6~0.9:0.4~0.1; the sum of the training set ratio and the test set ratio does not exceed 1. Preferably, the ratio of the training set to the test set is 0.7:0.3.

[0030] In another implementation of this embodiment, in step S5, the reliability of the face rate RF model is checked. The determination is made based on the regression plot or error plot of the test set. If the predicted result deviates significantly from the actual result, the parameters are readjusted.

[0031] As another implementation of this embodiment, in step S6, before validating the face face ratio RF model with the data in the validation sample set, the number of curves and the order of curves of the data input into the face face ratio RF model in the validation sample set are now aligned with the training set when the model is built.

[0032] Specifically, during model validation, the face rate predicted using well logging data is compared with the face rate obtained from core electron microscopy data. If the difference is small, the model is accurate and can be applied to the study area; otherwise, if the difference is large, the model is inaccurate and the face rate RF model needs to be retrained according to steps S3-S5.

[0033] In step S7, the face rate RF model validated in step S6 is used to predict the face rate data of other wells in the study block. The number of curves and the order of curves in the logging data of other wells are input and correspond to the training set when the model was built.

[0034] Example 2 As another preferred embodiment of the present invention, this embodiment further supplements and elaborates on the technical solution of the present invention based on the above-described embodiment 1. This embodiment uses specific examples to illustrate the technical solution of the present invention in detail.

[0035] Well logging curves of the target layer from wells 1, 2, and 3 in study block A were collected, including mineral models, with data point intervals of 0.125m. Core electron microscopy data (digital core data) of the target layer from wells 1 and 2 in study block A were also collected, including organic porosity, inorganic porosity, and fracture porosity.

[0036] The data was imported into the CIFLOG software. Using the "Data Fitting to Machine Learning" module in the preprocessing section of the CIFLOG software, the logging curves (including mineral models) of wells A1 and 2 in the study block were aligned to the core data depth, and then the data after alignment was exported.

[0037] Taking organic surface area ratio as an example, the training and prediction steps of the RF model for inorganic surface area ratio and cracked surface area ratio are the same, and will not be repeated here.

[0038] Use MATLAB software to write correlation analysis code and generate heatmaps (as shown in the instruction manual). Figure 2 As shown), combined with the geological background, several curves with good correlation to the organic surface rate were obtained: CNL, DEN, GR, KTH, POR, TOC, SW, VPYRT, VQUAZ, and VSH. Other curves with poor correlation (correlation < 0.2) were deleted, and the file was saved as a new file.

[0039] Write the code for a random forest machine learning method using MATLAB software. Read the previously saved file, where the organic face rate (OMPOR) is the predicted variable, and the other curves are variables. Set the training set ratio to 0.8 and the test set ratio to 0.2. Adjust the number of random seeds, leaves, number of nodes, and number of features. Output the regression plot and observe it (see the instruction manual appendix). Figure 3 As shown in the figure, we continue to try to adjust the parameters to make the model prediction results more accurate.

[0040] Save the adjusted model, substitute the data from well 3 into the model, predict the organic porosity curve of the target interval in well 1, save the data, and then import it into the CIFLOG software to plot the OMPOR curve. Add the core data from well 3 for comparison (refer to the instruction manual). Figure 4 As shown in the figure, the effect is good, therefore the model is qualified.

[0041] Substitute the data into the qualified organic surface area ratio RF model of another well 4 in study block A to predict the organic surface area ratio curve.

[0042] Subsequently, this embodiment also collected core electron microscopy data from well A4 in the study block, and substituted the predicted organic porosity curve into CIFLOG software for plotting. The actual organic porosity core data from well A4 were added for comparison, and the prediction results were found to be quite accurate. This embodiment demonstrates that the method of this application can effectively predict porosity.

[0043] Example 3 As another preferred embodiment of the present invention, this embodiment provides a face rate prediction device based on the random forest method, the prediction device comprising: The model data collection module is used to collect logging curves and core electron microscopy data of each well in the study block, and to align the core electron microscopy data of each well with the depth of its corresponding logging data. The model data preprocessing module is used to perform correlation analysis on the depth-aligned data and select curves with high correlation to the face rate data. The model building module is used to train a face face ratio RF model based on curves with high correlation to face face ratio data, and to perform reliability and accuracy verification on the trained face face ratio RF model. If the reliability verification fails, the face face ratio RF model is retrained by adjusting the parameters; if the reliability verification passes but the accuracy verification fails, the correlation curves are re-selected through the model data preprocessing module and the model is trained again; until both the reliability verification and accuracy verification pass, a well-trained face face ratio RF model is obtained. The data prediction module uses the face rate RF model trained by the model building module, and substitutes the well logging data of other wells in the study block to predict the face rate of other wells.

[0044] Example 4 In another preferred embodiment of the present invention, in order to achieve the above objectives, according to another aspect of the present application, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the face ratio prediction method based on the random forest method described above.

[0045] In this embodiment, the processor can be a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0046] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and units, such as the program module units corresponding to the above method embodiments of the present invention. The processor executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory, thereby implementing the methods in the above method embodiments.

[0047] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor, etc. Furthermore, the memory may include high-speed random access memory and non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0048] The one or more units are stored in the memory, and when executed by the processor, the steps of Embodiment 1 above are performed.

[0049] Example 5 As another preferred embodiment of the present invention, this embodiment discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in Embodiment 1 above.

Claims

1. A face rate prediction method based on random forest, characterized in that: The prediction method includes the following steps: S1. Collect logging curves and core electron microscopy data from each well in the study block; S2. Align the core electron microscopy data of each well with the corresponding well logging data at depth; use part of the aligned data as the training sample set for the face rate RF model and the other part as the validation sample set for the trained face rate RF model. S3. Perform correlation analysis on the data in the training sample set in step S2, and select the curves with high correlation to the face rate data respectively. S4. Divide the selected curves into the training set and test set for training the face rate RF model; S5. Train the face ratio RF model using the training set, and test the trained face ratio RF model using the test set to verify the reliability of the face ratio RF model. If the prediction results deviate significantly from the actual results, readjust the parameters until the trained face ratio RF model passes the reliability test. S6. Use the validation sample set from step S2 to validate the face rate RF model trained in step S5. If the validation passes, it means that the face rate RF model can be applied in the study area; otherwise, repeat steps S3-S5. S7. Using the face rate RF model validated in step S6, predict the face rate data of other wells in the study block.

2. The face rate prediction method based on random forest as described in claim 1, characterized in that: The surface area ratio includes organic surface area ratio, inorganic surface area ratio, and fracture surface area ratio. Through RF model training and testing, organic surface area ratio RF model, inorganic surface area ratio RF model, and fracture surface area ratio RF model are obtained. These are then applied to other wells in the study block to predict the organic surface area ratio, inorganic surface area ratio, and fracture surface area ratio of other wells in the study block, respectively.

3. A face rate prediction method based on random forest as described in claim 1 or 2, characterized in that: In step S1, the logging data includes logging curves of the target layer and calculated mineral model data.

4. A face rate prediction method based on random forest as described in claim 1 or 2, characterized in that: In step S2, the depth of the core electron microscope data is used as the standard to align the logging data with the core electron microscope data.

5. The face rate prediction method based on random forest as described in claim 4, characterized in that: In step S2, the data obtained by aligning the core electron microscopy data with the well logging data depth is exported into a single file.

6. A face rate prediction method based on random forest as described in claim 1 or 2, characterized in that: In step S4, the ratio of the training set to the test set is 0.6~0.9:0.4~0.1; the sum of the training set ratio and the test set ratio does not exceed 1.

7. The face rate prediction method based on random forest as described in claim 6, characterized in that: In step S4, the ratio of the training set to the test set is 0.7:0.

3.

8. A face rate prediction method based on random forest as described in claim 1 or 2, characterized in that: In step S5, the face rate RF model is trained by adjusting parameters, including the number of feature trees, the number of random seeds, and the number of leaves in the feature trees.

9. The face rate prediction method based on random forest as described in claim 8, characterized in that: In step S5, the reliability of the face rate RF model is tested. The result is determined by the regression plot or error plot of the test set. If the predicted result deviates significantly from the actual result, the parameters are readjusted.

10. A face rate prediction method based on random forest as described in claim 1 or 2, characterized in that: In step S6, before validating the face face ratio RF model with the data in the validation sample set, the number of curves and the order of the curves in the data input into the face face ratio RF model in the validation sample set are now correlated with the training set when building the model.

11. The face rate prediction method based on random forest as described in claim 10, characterized in that: In step S6, during model validation, the face rate predicted using well logging data is compared with the face rate obtained from core electron microscopy data. If the difference is small, the model is accurate and can be applied to the study block; otherwise, if the difference is large, the model is inaccurate and the face rate RF model needs to be retrained according to steps S3-S5.

12. A face rate prediction method based on random forest as described in claim 1 or 2, characterized in that: In step S7, the face rate RF model validated in step S6 is used to predict the face rate data of other wells in the study block. The number of curves and the order of curves in the logging data of other wells are input and correspond to the training set when the model was built.

13. A face rate prediction device based on the random forest method, characterized in that: The prediction device includes, The model data collection module is used to collect logging curves and core electron microscopy data of each well in the study block, and to align the core electron microscopy data of each well with the depth of its corresponding logging data. The model data preprocessing module is used to perform correlation analysis on the depth-aligned data and select curves with high correlation to the face rate data. The model building module is used to train a face face ratio RF model based on curves with high correlation to face face ratio data, and to perform reliability and accuracy verification on the trained face face ratio RF model. If the reliability verification fails, the face face ratio RF model is retrained by adjusting the parameters; if the reliability verification passes but the accuracy verification fails, the correlation curves are re-selected through the model data preprocessing module and the model is trained again; until both the reliability verification and accuracy verification pass, a well-trained face face ratio RF model is obtained. The data prediction module uses the face rate RF model trained by the model building module, and substitutes the logging data of other wells in the study block to predict the face rate of other wells.

14. A computer device, characterized in that: The system includes a processor, an input device, an output device, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute some or all of the steps of the face ratio prediction method based on the random forest method as described in any one of claims 1-12.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps of a face ratio prediction method based on a random forest method as described in any one of claims 1-12.