Method, device and equipment for determining oiliness of shale and storage medium

By building a machine learning-based oil saturation prediction model and utilizing full-diameter core data and logging data, the problem of inaccurate shale oil content evaluation in existing technologies has been solved, achieving more accurate prediction and quantitative analysis, and supporting reservoir evaluation and transformation decisions.

CN120296416BActive Publication Date: 2025-10-17CHINA UNIV OF PETROLEUM (BEIJING)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510351257.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-10-17
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

When determining shale oil saturation, existing technologies use core test methods that are time-consuming, expensive, and inaccurate, and well logging prediction methods that are inaccurate in mixed shales. There is a lack of methods that combine the shale's own characteristics with precise quantitative characterization, resulting in unclear oil content evaluation.

Method used

By obtaining oil saturation data and logging data of full-diameter cores on site, an oil saturation prediction model is constructed. The model is trained using machine learning methods, and the model parameters are optimized using training and validation sets to determine the optimal hyperparameter combination and improve prediction accuracy.

Benefits of technology

The accuracy of shale oil saturation prediction has been improved, and the key controlling factors of oil content in mixed shale can be quantified and analyzed more clearly, supporting geological evaluation and reservoir reconstruction decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296416B_ABST
    Figure CN120296416B_ABST
Patent Text Reader

Abstract

The application provides a shale oil-bearing property determination method, device, equipment and storage medium. The method comprises the following steps: acquiring logging data corresponding to multiple depths in a well to be measured; inputting the logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to the multiple depths output by the oil saturation prediction model; and determining the oil-bearing property of shale based on the oil saturation prediction values corresponding to the multiple depths. When the oil saturation prediction model is constructed, the model is trained by using oil saturation data of full-diameter cores in the field, which can improve the accuracy of the oil saturation prediction model, the prediction result output by the model is more accurate, the prediction accuracy of the oil saturation of shale is improved, and the method can quantitatively analyze and interpret the main control factors of the oil-bearing property.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of oil and gas exploration, and particularly relates to a shale oil-bearing property determination method, device, equipment and storage medium. BACKGROUND

[0002] The mixed-accumulation shale has the typical characteristics of extremely strong vertical and horizontal heterogeneity, development of multi-source fine particle mixed-accumulation laminae, complex micro-nano pore structure and complex fluid occurrence state, which brings great challenges to the systematic characterization and quantitative evaluation of key elements such as shale reservoir property, oil-bearing property, mobility and compressibility. The shale oil-bearing saturation is an important reservoir evaluation index, which is related to reservoir evaluation, development layer selection and geological reserve estimation.

[0003] The existing methods for determining the shale oil-bearing saturation mainly include core experiment determination method and logging prediction method. The core experiment determination method mainly carries out two-dimensional nuclear magnetic resonance experiment or other determination methods based on sealed coring data, so as to obtain shale oil-bearing saturation data. The logging prediction method mainly uses conventional logging or nuclear magnetic resonance logging data to calculate the oil-bearing saturation based on core calibration logging.

[0004] The core experiment determination method consumes a large amount of time and cost in the process of systematic sealed coring and experiment. The logging prediction method is not accurate in calculating the oil-bearing saturation because the traditional Archie model is not applicable due to the complex rock-electricity relationship in the mixed-accumulation shale, and the radial detection depth of the nuclear magnetic resonance logging is limited, so that the reflected flushed zone oil-bearing saturation value is affected by the drilling mud and cannot reflect the true formation condition, resulting in a low final calculation value.

[0005] The above methods do not consider the influence of the characteristics of the mixed-accumulation shale on the shale oil-bearing property, resulting in unclear understanding of the main control factors of the shale oil-bearing property, and lack of more accurate quantitative characterization method, and the calculation accuracy cannot meet the needs of reservoir oil-bearing property evaluation. SUMMARY

[0006] The embodiments of the present application provide a shale oil-bearing property determination method, device, equipment and storage medium, to solve the problem of inaccurate prediction result output by the existing oil-bearing saturation prediction model.

[0007] In a first aspect, the present application provides a shale oil-bearing property determination method, device, equipment and storage medium, to solve the problem of inaccurate prediction result output by the existing oil-bearing saturation prediction model.

[0008] Obtain the oil-bearing saturation data of the full-diameter core on site and the logging data corresponding to the depth;

[0009] The oil-bearing saturation data is used as a label data set, and the logging data is used as a feature data set;

[0010] determining a training set and a validation set according to the label data set and the feature data set;

[0011] performing iterative training processing on the prediction model by using the training set to obtain a trained prediction model;

[0012] performing validation processing on the trained prediction model by using the validation set;

[0013] in a case where the validation result indicates that the validation is successful, determining the prediction model as an oil saturation prediction model.

[0014] Optionally, the determining a training set and a validation set according to the label data set and the feature data set comprises:

[0015] extracting data as a first data set from the label data set and the feature data set according to a first preset proportion;

[0016] extracting data as a second data set from the label data set and the feature data set according to a second preset proportion;

[0017] processing data of the first data set and the second data set to obtain the training set and the validation set.

[0018] Optionally, the performing training processing on the prediction model by using the training set to obtain a trained prediction model comprises:

[0019] inputting feature data in the training set into the prediction model to obtain a prediction result output by the prediction model;

[0020] determining a loss function of the prediction model based on the prediction result and corresponding label data in the training set;

[0021] in a case where the loss function reaches a preset threshold value or the number of iterations reaches a preset number of times, determining that the training of the prediction model is completed;

[0022] in a case where the loss function does not reach the preset threshold value, adjusting model parameters of the prediction model, and performing iterative training processing on the adjusted prediction model based on the training set again until the loss function reaches the preset threshold value or the number of iterations reaches the preset number of times.

[0023] Optionally, the method further comprises:

[0024] determining an optimal hyperparameter combination of the prediction model from a plurality of candidate hyperparameter combinations of the prediction model based on a preset objective function;

[0025] refitting the prediction model according to the optimal hyperparameter combination to obtain a best prediction model.

[0026] In a second aspect, the present application provides a method for determining oil-bearing property of shale, comprising:

[0027] obtaining logging data corresponding to a plurality of depths in a well to be measured;

[0028] inputting the logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to the plurality of depths output by the oil saturation prediction model, wherein the oil saturation prediction model is obtained by training the oil saturation prediction model using the training method of the first aspect, and the oil saturation prediction values corresponding to the plurality of depths are used to indicate the corresponding oil-bearing property.

[0029] Optionally, the method further comprises:

[0030] inputting the oil saturation prediction model into an analysis model, and analyzing and processing the prediction results output by the oil saturation model based on the analysis model to output analysis results;

[0031] determining, according to the analysis results, the influence degree of different features and the interaction between the features in the feature data set corresponding to the oil saturation prediction model on the prediction results output by the oil saturation prediction model.

[0032] In a third aspect, the present application provides a device for constructing an oil saturation prediction model, comprising:

[0033] an obtaining module configured to obtain oil saturation data of a full-diameter core in the field and logging data corresponding to the depth;

[0034] a processing module configured to use the oil saturation data as a label data set and use the logging data as a feature data set;

[0035] a determining module configured to determine a training set and a validation set according to the label data set and the feature data set;

[0036] the processing module is further configured to use the training set to iteratively train and process a prediction model to obtain a trained prediction model;

[0037] the processing module is further configured to use the validation set to verify the trained prediction model;

[0038] the determining module is further configured to determine the prediction model as an oil saturation prediction model if the verification result indicates that the verification is successful.

[0039] Optionally, the processing module is further configured to:

[0040] extract data as a first data set from the label data set and the feature data set according to a first preset proportion;

[0041] extract data as a second data set from the label data set and the feature data set according to a second preset proportion;

[0042] process data of the first data set and the second data set to obtain a training set and a verification set.

[0043] Optionally, the processing module is further configured to input feature data in the training set into the prediction model to obtain a prediction result output by the prediction model.

[0044] The determination module is further configured to determine a loss function of the prediction model based on the prediction result and corresponding label data in the training set.

[0045] The determination module is further configured to determine that the prediction model is trained completely when the loss function reaches a preset threshold or the number of iterations reaches a preset number of times.

[0046] The processing module is further configured to adjust model parameters of the prediction model when the loss function does not reach the preset threshold, and iteratively train the prediction model based on the training set again until the loss function reaches the preset threshold or the number of iterations reaches the preset number of times.

[0047] Optionally, the determination module is further configured to determine an optimal hyperparameter combination of the prediction model from a plurality of candidate hyperparameter combinations of the prediction model based on a preset target function.

[0048] The processing module is further configured to refit the prediction model according to the optimal hyperparameter combination to obtain an optimal prediction model.

[0049] In a fourth aspect, the present application provides a device for determining oil content of shale, the device comprising:

[0050] A obtaining module is configured to obtain logging data corresponding to a plurality of depths in a well to be measured.

[0051] A processing module is configured to input the logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to the plurality of depths output by the oil saturation prediction model, wherein the oil saturation prediction model is trained by using the training method of the oil saturation prediction model of the first aspect and various possible implementation manners of the first aspect, and the oil saturation prediction values corresponding to the plurality of depths are used to indicate corresponding oil content.

[0052] In a fifth aspect, the present application provides a device for constructing an oil saturation prediction model, comprising:

[0053] a memory;

[0054] a processor;

[0055] The memory stores computer-executable instructions.

[0056] The processor executes the computer-executable instructions stored in the memory to implement the method according to the first aspect and various possible implementation manners of the first aspect.

[0057] In a sixth aspect, the present application provides a device for determining oil-bearing property of shale, comprising:

[0058] a memory;

[0059] a processor;

[0060] The memory stores computer-executable instructions.

[0061] The processor executes the computer-executable instructions stored in the memory to implement the method for determining oil-bearing property of shale according to the second aspect and various possible implementation manners of the second aspect.

[0062] In a seventh aspect, the present application provides a computer storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method according to the first aspect and various possible implementation manners of the first aspect or the second aspect and various possible implementation manners of the second aspect.

[0063] The present application provides a method, device, equipment and storage medium for determining oil-bearing property of shale. The method comprises the following steps: acquiring logging data corresponding to multiple depths in a well to be measured; inputting the logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to the multiple depths output by the oil saturation prediction model; and determining oil-bearing property of shale based on the oil saturation prediction values corresponding to the multiple depths. In the process of constructing the oil saturation prediction model, the oil saturation data of full-diameter cores in the field are used to train the model, which can improve the accuracy of the oil saturation prediction model and make the prediction results output by the model more accurate. Meanwhile, the method comprehensively considers the characteristics of the mixed accumulation shale, such as multiple-source fine-grained minerals and laminated mixed accumulation, strong heterogeneity, and micro-migration of shale oil due to source-reservoir separation, so that the key control factors of oil saturation of the mixed accumulation shale can be quantified and analyzed more clearly while improving the prediction accuracy, which has visual discussion value for the reasons of oil-bearing difference and is helpful for subsequent geological evaluation and reservoir reconstruction decision. BRIEF DESCRIPTION OF DRAWINGS

[0064] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0065] Figure 1 A method for constructing an oil saturation prediction model provided in the embodiment of this application Figure 1 ;

[0066] Figure 2 A method for constructing an oil saturation prediction model provided in the embodiment of this application Figure 2 ;

[0067] Figure 3 A flow chart of a method for determining shale oil content provided in an embodiment of the present application;

[0068] Figure 4 A ranking diagram of the importance of oil-bearing characteristics of mixed-accumulation shale based on SHAP analysis provided in an embodiment of the present application;

[0069] Figure 5 A ranking diagram of the main controlling factors of oil content of mixed shale based on SHAP analysis provided in an embodiment of the present application;

[0070] Figure 6 A scatter plot of characteristic interaction effects of factors affecting oil content in mixed shale based on SHAP analysis provided in an embodiment of the present application;

[0071] Figure 7 A schematic diagram of the structure of a device for constructing an oil saturation prediction model provided in an embodiment of the present application;

[0072] Figure 8 A schematic diagram of the structure of an oil saturation prediction model construction device provided in an embodiment of the present application.

[0073] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0074] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions in this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0075] The terms "first", "second", "third", "fourth" and the like in the description and in the claims of the present application, and above-described drawings, if any, are used to distinguish between similar objects and are not necessarily used to describe a particular sequential or chronological order. It is to be understood that the use of the terms so

[0076] In the embodiments of the present application, the words "exemplary" and "for example" are used to mean serving as an example, instance, or illustration, at 5 2 least. Any implementation or design scenario described as "exemplary" or "for example" in the present application should not be construed as being preferred or advantageous over other implementations or design scenarios. In fact, any implementation or design scenario described as "exemplary" or "for example" in the present application is intended to present concepts or features in a concrete manner, and the one skilled in the art can implement or design the present application in other concrete forms without changing the essential nature thereof.

[0077] The oil saturation of shale is a key indicator for evaluating reservoir quality, and is of great significance for reservoir evaluation, selection of development intervals, and estimation of geological reserves.

[0078] Currently, the methods for determining the oil saturation of shale mainly fall into two categories: one is experimental determination based on core samples, and the other is logging prediction. The experimental determination mainly relies on samples obtained by sealed coring, and the oil saturation data of shale are obtained by performing two-dimensional nuclear magnetic resonance experiments or other related testing means. The logging prediction is to calculate the oil saturation by using conventional logging data or nuclear magnetic resonance logging data with the help of core calibration logging technology.

[0079] However, the core experimental determination method needs to invest a large amount of time and cost in the process of sealed coring and subsequent experiments. In the environment of complex rock-electricity relationship of mixed deposition shale, the traditional Archie model is not applicable, and therefore the result of calculating the oil saturation by using the conventional logging resistivity method is often not accurate enough. In addition, although the nuclear magnetic resonance logging can provide information about fluid saturation, its radial detection depth is limited, and the reflected oil saturation of the flushed zone is easily affected by the drilling mud, which cannot truly reflect the formation conditions, thereby leading to the calculated oil saturation value being too low.

[0080] That is, the above methods do not consider the influence of the characteristics of mixed deposition shale on the oil saturation of shale, which leads to unclear understanding of the main controlling factors of the oil saturation of shale, and lack of more accurate quantitative characterization methods, and the calculation accuracy cannot meet the needs of the evaluation of reservoir oiliness.

[0081] To solve the above problems in the prior art, the present application provides a shale oil-bearing property determination method. The method comprises the following steps: acquiring logging data corresponding to multiple depths in a well to be measured; inputting the logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to the multiple depths output by the oil saturation prediction model; and determining the oil-bearing property of shale based on the oil saturation prediction values corresponding to the multiple depths. In the process of constructing the oil saturation prediction model, the model is trained by using oil saturation data of full-diameter cores in the field, so that the accuracy of the oil saturation prediction model is improved, the prediction result output by the model is more accurate, and the method comprehensively considers the characteristics of the mixed accumulation shale, such as the mixed accumulation of multi-source fine-grained minerals and laminations, strong heterogeneity, and shale oil micro-migration due to source-reservoir separation, so that the key control factors of the oil saturation of the mixed accumulation shale can be quantified and analyzed more clearly while the prediction accuracy is improved, the oil-bearing difference reasons have visual discussion value, and the method is helpful for subsequent geological evaluation and reservoir reconstruction decision-making.

[0082] The technical solutions of the present application and how the technical solutions solve the above technical problems will be described in detail below with specific examples. The following specific examples can be combined with each other, and the same or similar concepts or processes can not be described again in some examples. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0083] Figure 1 is a construction method of an oil saturation prediction model provided by an embodiment of the present application Figure 1 . As shown in Figure 1 , the construction method of the oil saturation prediction model provided by the embodiment comprises the following steps:

[0084] S101: acquiring oil saturation data of full-diameter cores in the field and logging data corresponding to the depths.

[0085] The oil saturation data refers to the proportion of oil phase fluid in the rock in a certain logging section. The logging data may, for example, be porosity, total organic carbon, free hydrocarbon, mineral content, and conventional logging curve data.

[0086] It can be understood that, in the embodiment, two-dimensional nuclear magnetic resonance measurement is performed on the full-diameter cores of the target well to obtain two-dimensional nuclear magnetic resonance spectra, and the oil saturation data is extracted from the two-dimensional nuclear magnetic resonance T1-T2 spectra.

[0087] Further, according to the logging depths corresponding to the oil saturation, the porosity, total organic carbon, free hydrocarbon, mineral content, and conventional logging curve data (spontaneous potential, deep resistivity, neutron, density, acoustic time difference, natural gamma, etc.) corresponding to the depths are acquired.

[0088] By obtaining the oil saturation data of the target well logging and the well logging data corresponding to the depth, the physical properties, oil content, fluid properties and the like of the reservoir can be comprehensively analyzed, thereby providing a basis for reservoir classification, reserve calculation and oil and gas production capacity prediction.

[0089] S102: The oil saturation data is taken as a label data set, and the well logging data is taken as a feature data set.

[0090] The label data in the label data set is an object to be predicted by the machine learning model. The feature data in the feature data set is used to train the machine learning model.

[0091] It can be understood that in the embodiment, the oil saturation data is taken as a label, and a plurality of well logging data such as porosity, total organic carbon, free hydrocarbon, mineral content, natural potential, deep resistivity and the like are taken as features. The obtained oil saturation data and well logging data are subjected to missing value or abnormal value processing to ensure the integrity and accuracy of the data, and then the label data set and the feature data set are obtained.

[0092] S103: Determine a training set and a validation set according to the label data set and the feature data set.

[0093] The training set is used to guide the model to learn the rules and patterns in the data, and contains all the features and labels that the model needs to learn. The validation set is used to monitor the performance of the model during training and help adjust the parameters of the model.

[0094] It can be understood that the label data set and the feature data set obtained in the above steps are divided into two parts according to a certain proportion. Specifically, one part is used to train the machine learning model, and the other part is used to evaluate the performance of the model.

[0095] It can also be understood that when the data set is divided, a random division method is used, that is, the probability of each data point being allocated to the training set or the validation set is equal, which helps to avoid bias in the data division process, thereby improving the generalization ability of the model.

[0096] The training set is used to train the model, and the validation set is used to evaluate the performance of the model. During training, the performance indicators on the validation set are monitored to determine whether the model has overfitting or underfitting. If the performance of the model on the validation set is poor, the parameters or structure of the model can be adjusted to improve its performance.

[0097] By determining the training set and the validation set, and training and adjusting the machine learning model through the training set and the validation set, the performance and generalization ability of the model can be improved.

[0098] S104: The training set is used to iteratively train the prediction model, and a trained prediction model is obtained.

[0099] It can be understood that, by using the acquired training data set, the internal parameters and structure of the prediction model are continuously adjusted and optimized through multiple repeated training processes, until the model reaches the predetermined performance standard or convergence condition on the training set, thereby obtaining a trained model that can be used for actual prediction.

[0100] It can also be understood that, by observing the trend of the loss function during the training process, it can be determined whether the model has converged to the optimal solution or is close to the optimal solution. If the loss function value no longer significantly decreases after multiple iterations, it is considered that the model has converged and training can be stopped.

[0101] During the iterative training process, the parameters of the model, such as learning rate, batch size, number of iterations, etc., can also be adjusted according to the performance of the model on the training set and the validation set, in order to further optimize the performance of the model.

[0102] In each iteration, a batch of samples is randomly selected from the training set, the prediction error of the model is calculated, and the parameters of the model are updated according to the gradient of the error. This process will be repeated until the predetermined number of iterations or convergence condition is reached.

[0103] The purpose of this step is to iteratively train the prediction model on the training set, thereby obtaining a trained prediction model with good performance.

[0104] S105: Verifying the trained prediction model using the validation set.

[0105] Specifically, the samples in the validation set are input into the model, and the prediction results of the model are collected. According to the prediction results of the model and the true labels of the validation set, a series of performance indicators are calculated, such as accuracy, recall, F1 score, mean square error, etc.

[0106] By using the validation set to verify the trained prediction model, the performance of the model can be comprehensively evaluated, and guidance can be provided for further model optimization.

[0107] S106: In the case where the verification result indicates that the verification is successful, the prediction model is determined as the oil saturation prediction model.

[0108] It can be understood that, by analyzing the performance indicators such as accuracy, recall, F1 score, etc. on the validation set, and comparing the prediction results with the true labels, the analysis results can be used as the basis for determining whether the model is successful.

[0109] Furthermore, based on the verification results, the model is evaluated to see whether it meets the predetermined performance standards. If the performance of the model on the verification set meets the predetermined standards and the deviation between the predicted results and the true labels is within the preset error range, the verification is considered successful and the model is determined to be the oil saturation prediction model.

[0110] When the verification result indicates success, the prediction model is determined as an oil saturation prediction model, thereby meeting the application needs in the field of oil and gas exploration and development.

[0111] This embodiment provides a method for constructing an oil saturation prediction model. The method obtains oil saturation data of a target well and well logging data at a corresponding depth; uses the oil saturation data as a label data set and the well logging data as a feature data set; determines a training set and a validation set based on the label data set and the feature data set; iteratively trains the prediction model using the training set to obtain a trained prediction model; verifies the trained prediction model using the validation set; and if the verification result indicates successful verification, determines the prediction model as an oil saturation prediction model, thereby improving the accuracy of the oil saturation prediction model.

[0112] Figure 2 A method for constructing an oil saturation prediction model provided in the embodiment of this application Figure 2 This embodiment is based on Figure 1 Based on the embodiment, a possible implementation method of the method for constructing the oil saturation prediction model is described in detail. Figure 2 As shown, the method includes:

[0113] S201: Obtain oil saturation data of the full-diameter core on site and well logging data at the corresponding depth.

[0114] Among them, step S201 is similar to the above-mentioned step S101 and will not be repeated here.

[0115] S202: The oil saturation data is used as a label data set, and the well logging data is used as a feature data set.

[0116] Among them, step S202 is similar to the above-mentioned step S102 and will not be repeated here.

[0117] S203: extracting data from the label dataset and the feature dataset according to a first preset ratio as a first dataset.

[0118] The first preset ratio is set according to the total amount of data and the requirements of the experiment.

[0119] It can be understood that, according to the preset proportion, a corresponding proportion of data is randomly selected from the label data set and the feature data set to form a new data set, which can be used for subsequent machine learning model training.

[0120] S204: Extracting data as a second data set from the label data set and the feature data set according to a second preset proportion.

[0121] It can be understood that, in the above step S203, a part of data is divided from the feature data set and the label data set according to the preset proportion as the first data set, and further, in this embodiment, the remaining part of the data in the feature data set and the label data set is taken as the second data set and used for subsequent machine learning model verification.

[0122] S205: Processing the data of the first data set and the second data set to obtain a training set and a validation set.

[0123] It can be understood that, the first data set and the second data set are preprocessed, including missing value processing, normalization, etc. Among them, the purpose of preprocessing is to improve the data quality, so that the model can learn more effectively.

[0124] Specifically, the missing value processing refers to checking whether there are missing or blank values in the data set, and taking corresponding measures to fill or delete, so as to avoid the negative impact of these missing values on model training. Common missing value processing methods include using mean, median, mode, etc. Statistics for filling, or using more complex interpolation algorithms to estimate missing values.

[0125] Normalization processing is to scale the feature values in the data set to a specified range, usually between 0 and 1. For example, using the MinMaxScaler method, the data is scaled to 0 to 1 according to the minimum and maximum value of each feature, so as to eliminate the deviation between different features due to different dimensions or value ranges, so that the model can treat each feature more equally during training.

[0126] This step processes the data of the first data set and the second data set to obtain a training set and a validation set, providing a basis for the training and evaluation of the machine learning model.

[0127] S206: Inputting the feature data in the training set into the prediction model to obtain the prediction result output by the prediction model.

[0128] In this embodiment, the prediction model is constructed based on the XGBoost algorithm and is used to predict oil saturation. The feature data in the training set is input into the prediction model, and the prediction model will calculate the predicted output value according to the input feature data using the internal learned rules and patterns.

[0129] S207: Determine the loss function of the prediction model based on the prediction result and the corresponding label data in the training set.

[0130] The loss function provides a quantitative standard for measuring the accuracy of the model prediction. During training, the loss function is used to guide the parameter update of the model. By minimizing the loss function, the optimal model parameters are found, and the prediction performance of the model is improved.

[0131] It can be understood that the calculated loss value can be used to evaluate the performance of the model. If the loss value is high, it means that the prediction performance of the model is poor, and the parameters or structure of the model need to be adjusted. Through iterative training and adjustment, the loss value can be gradually reduced, and the prediction accuracy of the model can be improved.

[0132] S208: Determine that the prediction model training is completed when the loss function reaches a preset threshold or the number of iterations reaches a preset number.

[0133] It can be understood that by setting the threshold of the loss function, it can be ensured that the model will not overfit the training data during training, thereby maintaining the generalization ability to new data. Setting an upper limit on the number of iterations can ensure that the training process will not go on indefinitely, thereby saving time and computing resources.

[0134] After each iteration, check whether the loss function value has reached the preset threshold or the number of iterations has reached the preset upper limit. If any of the conditions are met, it can be considered that the model training has been completed.

[0135] Setting the loss function threshold and the number of iterations to determine the conditions for completing the training of the prediction model helps to avoid overfitting, control the training time, and ensure that the model can converge to the optimal solution.

[0136] S209: Adjust the model parameters of the prediction model when the loss function does not reach the preset threshold, and re-iterate the training of the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

[0137] It can be understood that after each iteration training is completed, the loss function value of the current model is calculated and compared with the preset threshold. If the loss function value does not reach the preset threshold, the parameters of the model are adjusted based on the gradient information of the loss function, and the prediction value and the loss function value are calculated according to the new parameters.

[0138] After each re-training, check again whether the loss function value has reached the preset threshold or the number of iterations has reached the preset maximum value. If any of the conditions are met, stop training.

[0139] By iteratively adjusting the model parameters and retraining until the loss function reaches a preset threshold or the number of iterations reaches a preset number, the performance of the model can be optimized.

[0140] S210: The trained prediction model is verified using the validation set.

[0141] It can be understood that after training, the model is evaluated using the validation set. This includes calculating various performance indicators of the model on the validation set, such as accuracy, recall, F1 score, etc. Based on the performance indicators on the validation set, the performance of the model is analyzed.

[0142] Verifying the trained prediction model using the validation set helps to evaluate the performance of the model, prevent overfitting, adjust model parameters, and select the best model.

[0143] S211: In the case where the verification result indicates that the verification is successful, the prediction model is determined as the oil saturation prediction model.

[0144] It can be understood that before starting the verification, the performance indicators that the model needs to achieve on the validation set and the corresponding thresholds need to be determined.

[0145] Further, the calculated performance indicators are compared with the preset success criteria. If the performance indicators meet or exceed the success criteria, the verification is considered successful.

[0146] In the case where the verification is successful, the current trained model is determined as the prediction model for oil saturation prediction.

[0147] In an optional embodiment, the method for constructing the oil saturation prediction model further comprises: determining the optimal hyperparameter combination of the prediction model from multiple candidate hyperparameter combinations of the prediction model based on a preset objective function, and refitting the prediction model according to the optimal hyperparameter combination to obtain the best prediction model.

[0148] For example, the prediction model is constructed based on the XGBoost algorithm, and Bayesian optimization or other methods are used to search for the optimal hyperparameters of XGBoost. Bayesian optimization uses a probability model to approximate the objective function and guides the search process through a sampling function, thereby efficiently finding the optimal solution in the parameter space.

[0149] As you can understand, the objective function evaluates the performance of the XGBoost model given its hyperparameters. The objective function accepts a set of hyperparameters as input and returns a corresponding model performance score. Outputs of the objective function can include, for example, the coefficient of determination (R²) or the root mean square error (RMSE). The Bayesian optimization algorithm is used to search for the optimal hyperparameters for XGBoost. At each iteration, the algorithm selects the next set of hyperparameters to evaluate based on the current estimate of the objective function. Through continuous iteration, the algorithm gradually converges to the optimal solution.

[0150] Based on the trained oil saturation prediction model, the optimal parameters can be used to predict the oil saturation at multiple target logging depths, which helps to improve the efficiency and accuracy of oil exploration and production.

[0151] The present application provides a method for constructing an oil saturation prediction model. The method first obtains oil saturation data of a target well and well logging data of a corresponding depth, and uses them as a label dataset and a feature dataset, respectively. Corresponding data is extracted from the label dataset and the feature dataset according to a preset ratio as a training set and a validation set. The feature data in the training set is input into the prediction model to obtain a prediction result output by the prediction model. Based on the prediction result and the corresponding label data in the training set, a loss function of the prediction model is determined. If the loss function reaches a preset threshold or the number of iterations reaches a preset number, the prediction model training is determined to be complete. If the loss function does not reach the preset threshold, the model parameters of the prediction model are adjusted, and the adjusted prediction model is iteratively trained again based on the training set until the loss function reaches a preset threshold or the number of iterations reaches a preset number. Finally, the trained prediction model is verified using a validation set. If the verification result indicates successful verification, the prediction model is determined to be an oil saturation prediction model. This method iteratively trains the prediction model multiple times until the preset conditions are met, thereby improving the accuracy and reliability of the oil saturation prediction model.

[0152] Figure 3 This is a flow chart of a method for determining shale oil saturation provided in an embodiment of the present application. Figure 3 As shown, the method for determining shale oil saturation provided in this embodiment includes:

[0153] S301: Acquire well logging data corresponding to multiple depths in the well to be logged.

[0154] Among them, logging data include porosity, total organic carbon, free hydrocarbons, mineral content and conventional logging curves.

[0155] S302: Input the well logging data into the oil saturation prediction model to obtain the oil saturation prediction values ​​corresponding to multiple depths output by the oil saturation prediction model.

[0156] The oil saturation prediction model is trained according to the training method of the oil saturation prediction model in the above embodiment. The oil saturation prediction values corresponding to the plurality of depths are used to indicate the corresponding oil saturation.

[0157] It can be understood that the obtained logging data of the well to be measured is input into the trained oil saturation prediction model, and the oil saturation prediction model analyzes and processes the logging data at different depths according to its own structure and parameters, and then outputs the oil saturation prediction value at the corresponding depth. The porosity, total organic carbon, free hydrocarbon, mineral content and conventional logging curve data are used as characteristic inputs, which are parameters selected in consideration of the characteristics of the mixed accumulation shale.

[0158] In an optional embodiment, the shale oil saturation determination method further comprises: inputting the oil saturation prediction model into an analysis model, and analyzing and processing the prediction results output by the oil saturation model based on the analysis model to output analysis results; determining the influence degree of different features and the interaction between the features in the feature data set corresponding to the oil saturation prediction model on the prediction results output by the oil saturation prediction model according to the analysis results.

[0159] For example, based on the oil saturation prediction model that has been trained and optimized, the SHAP (SHapley Additive exPlanations) analysis method is further used to visualize and numerically evaluate the contribution of each feature to the prediction results of each sample.

[0160] The main process is explained as follows:

[0161] 1. Calculate the SHAP value: use the trained oil saturation prediction model to calculate the SHAP value of each sample and each feature for the verification set or test set data; SHAP introduces the idea of Shapley value in game theory, which measures the expected increment of the contribution of a feature from "joining" to "not joining" to the model output.

[0162] In this process, the SHAP value of feature i can be determined by the following formula, for example:

[0163]

[0164] where F is the set of all features, S is a subset that does not contain feature i, f(S) is the prediction value of the model on the feature subset S, |S|! is the factorial of the subset size, and f(S∪{i})-f(S) is the prediction change after feature i is added to the subset S.

[0165] In this formula, the combination weight is It can be ensured that all possible feature subsets are evenly weighted.

[0166] 2. Accumulation and visualization: SHAP values of each sample are summarized and a feature importance ranking chart is drawn; for example, the average contribution of each feature and the positive / negative impact can be intuitively displayed by using the visualization tool provided by SHAP.

[0167] In this process, the SHAP value corresponding to each record can be calculated for all data of the test set and the training set, and the contribution of each feature to the oil saturation and the interaction effect visualization result are obtained; then, the SHAP feature importance bar chart and scatter chart are drawn to identify the most critical control factors (such as porosity, TOC, free hydrocarbon, mineral content, and a curve in the conventional logging curve).

[0168] Figure 4 A feature importance ranking chart of mixed accumulation type shale oil-bearing property based on SHAP analysis is provided for the embodiments of the present application. As shown in Figure 4 The importance ranking of the shale oil-bearing property influencing factors is: porosity > clay mineral > total organic carbon > free hydrocarbon > spontaneous potential > analcime > deep resistivity > dolomite > neutron > quartz > density > acoustic time difference > natural gamma.

[0169] Figure 5 A main control factor ranking chart of mixed accumulation type shale oil-bearing property based on SHAP analysis is provided for the embodiments of the present application. The pink points are used to indicate that the feature value has a positive impact on the model prediction, and the blue points are used to indicate that the feature value has a negative impact on the model prediction.

[0170] The horizontal axis (SHAP value) is used to represent the influence of the feature on the prediction result, and the farther the point is from the center line (zero point), the greater the influence of the feature on the model output. A positive SHAP value indicates a positive impact, and a negative SHAP value indicates a negative impact. The vertically arranged features are sorted from large to small in terms of influence, and the features above have a greater total influence on the model output, while the features below have a smaller influence.

[0171] Therefore, in combination with Figure 5 It can be known that porosity, clay mineral, total organic carbon and free hydrocarbon content are the main control factors of shale oil-bearing property. The high value of porosity has a greater positive contribution to the model, and the low value of clay mineral has a greater positive relationship with the model.

[0172] 3. Identifying main control factors and interaction: according to the size of the feature SHAP value, the key parameters with the greatest contribution to the prediction result are identified; through the interaction SHAP value between features, the mutual coupling effect between different features can be determined, so as to show a positive or negative adjustment effect in prediction.

[0173] Among them, analyzing the scatter plot of interaction effects between features can reveal the inherent mechanism of oil content differentiation in mixed accumulation shales from a data-driven perspective.

[0174] Figure 6 A scatter plot of characteristic interaction effects of factors affecting oil content in mixed shale based on SHAP analysis is provided in an embodiment of the present application.

[0175] like Figure 6 As shown, as porosity increases, SHAP rapidly increases within the 2-4% range, reflecting a rapidly increasing contribution to the model. When porosity SHAP exhibits a positive influence, the corresponding clay content is low, reflecting a negative impact of clay minerals. However, the interaction plot between free hydrocarbons and total organic carbon exhibits an overall linear positive correlation, reflecting the positive impact of high free hydrocarbon and total organic carbon content on the model. This interaction analysis reveals the importance of the relationship between oil content differences and source-reservoir configuration in mixed-accumulation shales, demonstrating that shale oil content is not controlled by a single factor but rather by the combined effects of multiple factors.

[0176] In this process, by determining the degree of influence of different features in the feature data set corresponding to the oil saturation prediction model and the interaction between the features on the prediction results output by the oil saturation prediction model, not only the accuracy of oil saturation prediction is improved, but also the understanding of the geological characteristics of mixed accumulation shale reservoirs is deepened, providing a theoretical basis and technical support for subsequent exploration and development.

[0177] An embodiment of the present application provides a method for determining shale oil saturation. The method obtains logging data corresponding to multiple depths in a well to be logged, inputs the logging data into a trained and optimized oil saturation prediction model, and obtains oil saturation prediction values ​​corresponding to multiple depths output by the oil saturation prediction model, thereby improving the prediction accuracy of shale oil saturation.

[0178] Figure 7 This is a schematic diagram of the structure of a device for constructing an oil saturation prediction model provided in this application. Figure 7 As shown, the apparatus 400 for constructing an oil saturation prediction model provided in this embodiment includes:

[0179] Acquisition module 401, used to obtain oil saturation data of full-diameter cores on site and well logging data at corresponding depths;

[0180] A processing module 402 is configured to use the oil saturation data as a label data set and the well logging data as a feature data set;

[0181] Determining module 403, for determining a training set and a validation set based on the label data set and the feature data set;

[0182] The processing module 402 is further configured to perform iterative training on the prediction model by using the training set, to obtain a trained prediction model.

[0183] The processing module 402 is further configured to perform verification on the trained prediction model by using the verification set.

[0184] The determination module 403 is further configured to determine the prediction model as the oil saturation prediction model if the verification result indicates that the verification is successful.

[0185] Optionally, the processing module 402 is further configured to:

[0186] extract data from the label data set and the feature data set as a first data set according to a first preset proportion;

[0187] extract data from the label data set and the feature data set as a second data set according to a second preset proportion;

[0188] process data of the first data set and the second data set to obtain a training set and a verification set.

[0189] Optionally, the processing module 402 is further configured to input feature data in the training set into the prediction model to obtain a prediction result output by the prediction model.

[0190] The determination module 403 is further configured to determine a loss function of the prediction model based on the prediction result and corresponding label data in the training set.

[0191] The determination module 403 is further configured to determine that the training of the prediction model is completed if the loss function reaches a preset threshold or the number of iterations reaches a preset number.

[0192] The processing module 402 is further configured to adjust model parameters of the prediction model if the loss function does not reach the preset threshold, and perform iterative training on the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

[0193] Optionally, the determination module 403 is further configured to determine an optimal hyperparameter combination of the prediction model from a plurality of candidate hyperparameter combinations of the prediction model based on a preset target function.

[0194] The processing module 402 is further configured to re-fit the prediction model according to the optimal hyperparameter combination to obtain an optimal prediction model.

[0195] The embodiment provides a device for constructing an oil saturation prediction model, which can execute the method provided in the method embodiment, and has similar implementation principles and technical effects.

[0196] The embodiment further provides a device for determining oil content of shale.

[0197] The device comprises an acquisition module.

[0198] The device further comprises a processing module.

[0199] The device for determining oil content of shale can execute the method provided in the method embodiment, and has similar implementation principles and technical effects.

[0200] Figure 8 A structural schematic diagram of a device for constructing an oil saturation prediction model is provided. Figure 8 As shown in the structural schematic diagram of the device for constructing an oil saturation prediction model, the device for constructing an oil saturation prediction model comprises a receiver, a transmitter, a processor and a memory.

[0201] The receiver is configured to receive instructions and data.

[0202] The transmitter is configured to send instructions and data.

[0203] The memory is configured to store computer execution instructions.

[0204] The processor is configured to execute the computer execution instructions stored in the memory, so as to implement each step executed by the method for constructing an oil saturation prediction model in the above embodiment.

[0205] Optionally, the memory can be independent or integrated with the processor.

[0206] When the memory is independently arranged, the electronic device further comprises a bus for connecting the memory and the processor.

[0207] The application further provides a computer storage medium, and the computer storage medium stores computer execution instructions, and when a processor executes the computer execution instructions, a method for constructing an oil saturation prediction model is realized.

[0208] Those of ordinary skill in the art understand that all or some steps in the method disclosed above and the functional modules / units in the system and device can be implemented as software, firmware, hardware and appropriate combinations thereof. In the hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be executed by several physical components in cooperation. Certain physical components or all physical components can be implemented as software executed by a processor, such as a central processor, a digital signal processor or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer readable medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As known to those of ordinary skill in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. In addition, as known to those of ordinary skill in the art, communication media typically includes computer readable instructions, data structures, program modules or other data in modulated data signals such as carrier waves or other transport mechanisms, and can include any information delivery medium.

[0209] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the application be limited only by the scope of the claims, including any appropriate amendments thereof, and that there be no intention to limit the application to the equivalents thereof, unless the equivalents are expressly claimed. The specification and examples given herein are to be considered exemplary only, and the true scope and spirit of the application indicated by the following claims.

[0210] It should be understood that the application is not limited to the precise construction that has been described above and shown in the accompanying drawings, and that changes and modifications can be effected therein by those skilled in the art without departing from the scope of the application. The scope of the application should be limited only by the claims that follow.

Claims

1. A method for constructing an oil saturation prediction model, which is applied to the scenario of determining the oil saturation of mixed shale, characterized in that: include: Obtaining oil saturation data of full-diameter cores on site and well logging data of corresponding depths, wherein the well logging data includes porosity, total organic carbon, free hydrocarbons, mineral content, and conventional well logging curve data; the oil saturation data is extracted from a two-dimensional nuclear magnetic resonance spectrum obtained by performing two-dimensional nuclear magnetic resonance measurements on full-diameter cores of target wells; Using the oil saturation data as a label data set and the well logging data as a feature data set; Determine a training set and a validation set based on the label data set and the feature data set; The prediction model constructed based on the XGBoost algorithm is iteratively trained using the training set to obtain a trained prediction model; Using the validation set to validate the trained prediction model; If the verification result indicates that the verification is successful, the prediction model is determined as an oil saturation prediction model, and the oil saturation prediction model is used to predict and output oil saturation prediction values ​​corresponding to multiple depths in the well to be tested, and the oil saturation prediction values ​​corresponding to the multiple depths are used to indicate corresponding oil content; The method further includes: inputting the oil saturation prediction model into an analysis model, and analyzing and processing the prediction result output by the oil saturation model based on the analysis model, and outputting the analysis result; According to the analysis results, the degree of influence of different features in the feature data set corresponding to the oil saturation prediction model and the interaction between the features on the prediction results output by the oil saturation prediction model is determined. The prediction results are a scatter plot of the feature interaction effects of the factors affecting the oil content of the mixed accumulation type shale and a ranking diagram of the main controlling factors of the oil content of the mixed accumulation type shale.

2. The method according to claim 1, characterized in that The determining of a training set and a validation set based on the label data set and the feature data set includes: Extracting data from the label dataset and the feature dataset according to a first preset ratio as a first dataset; Extracting data from the label dataset and the feature dataset according to a second preset ratio as a second dataset; The data of the first data set and the second data set are processed to obtain a training set and a validation set.

3. The method according to claim 2, characterized in that The prediction model is trained using the training set to obtain a trained prediction model, including: Inputting the feature data in the training set into the prediction model to obtain the prediction result output by the prediction model; Determining a loss function of the prediction model based on the prediction result and corresponding label data in the training set; When the loss function reaches a preset threshold or the number of iterations reaches a preset number, determining that the prediction model training is completed; When the loss function does not reach the preset threshold, the model parameters of the prediction model are adjusted, and the adjusted prediction model is iteratively trained again based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

4. The method according to claim 3, characterized in that The method further comprises: Determining an optimal hyperparameter combination of the prediction model from a plurality of candidate hyperparameter combinations of the prediction model based on a preset objective function; The prediction model is refitted according to the optimal hyperparameter combination to obtain the best prediction model.

5. A method for determining the oil content of shale, characterized in that: The method comprises: Obtain logging data corresponding to multiple depths in the well to be logged; The logging data is input into an oil saturation prediction model to obtain oil saturation prediction values ​​corresponding to multiple depths output by the oil saturation prediction model, wherein the oil saturation prediction model is trained using the oil saturation prediction model training method described in any one of claims 1 to 4, and the oil saturation prediction values ​​corresponding to the multiple depths are used to indicate the corresponding oil content.

6. The method according to claim 5, characterized in that The method further comprises: Inputting the oil saturation prediction model into the analysis model, and analyzing and processing the prediction result output by the oil saturation model based on the analysis model, and outputting the analysis result; The degree of influence of different features in the feature data set corresponding to the oil saturation prediction model and the interaction between the features on the prediction result output by the oil saturation prediction model is determined based on the analysis result.

7. A device for constructing an oil saturation prediction model, the device being applied to the scenario of determining the oil saturation of mixed shale, characterized in that: The device comprises: An acquisition module is used to obtain oil saturation data of full-diameter cores on site and well logging data at corresponding depths, wherein the well logging data includes porosity, total organic carbon, free hydrocarbons, mineral content, and conventional well logging curve data; the oil saturation data is extracted from a two-dimensional nuclear magnetic resonance spectrum obtained by performing two-dimensional nuclear magnetic resonance measurements on full-diameter cores of target wells; a processing module, configured to use the oil saturation data as a label data set and the well logging data as a feature data set; A determination module, configured to determine a training set and a validation set based on the label data set and the feature data set; The processing module is further configured to perform iterative training on the prediction model constructed based on the XGBoost algorithm using the training set to obtain a trained prediction model; The processing module is further configured to use the verification set to perform verification processing on the trained prediction model; The determination module is further configured to, if the verification result indicates successful verification, determine the prediction model as an oil saturation prediction model, wherein the oil saturation prediction model is configured to predict and output oil saturation prediction values ​​corresponding to multiple depths in the well to be tested, wherein the oil saturation prediction values ​​corresponding to the multiple depths are configured to indicate corresponding oil content; The device further includes: inputting the oil saturation prediction model into an analysis model, and analyzing and processing the prediction result output by the oil saturation model based on the analysis model, and outputting the analysis result; According to the analysis results, the degree of influence of different features in the feature data set corresponding to the oil saturation prediction model and the interaction between the features on the prediction results output by the oil saturation prediction model is determined. The prediction results are a scatter plot of the feature interaction effects of the factors affecting the oil content of the mixed accumulation type shale and a ranking diagram of the main controlling factors of the oil content of the mixed accumulation type shale.

8. The device according to claim 7, characterized in that The processing module is further configured to: Extracting data from the label dataset and the feature dataset according to a first preset ratio as a first dataset; Extracting data from the label dataset and the feature dataset according to a second preset ratio as a second dataset; The data of the first data set and the second data set are processed to obtain a training set and a validation set.

9. A device for constructing an oil saturation prediction model, characterized in that: The device comprises: Memory; processor; wherein the memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method for constructing an oil saturation prediction model according to any one of claims 1 to 4.

10. A computer storage medium, characterized in that The computer storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method for constructing an oil saturation prediction model according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-scale rock physics fused reservoir saturation logging intelligent evaluation method and system

    CN118859354A

  • Machine learning-assisted rational design of separation membranes

    WO2023200979A1