Shale oiliness determination method and device, equipment and storage medium

By constructing a machine learning-based shale oil saturation prediction model, using full-diameter core data and logging data, the problem of inaccurate prediction of shale oil content in the existing technology is solved, and a higher-precision shale oil content evaluation is achieved, and reservoir transformation decisions are supported.

CN120296416AActive Publication Date: 2025-07-11CHINA UNIV OF PETROLEUM (BEIJING)

Patent Information

Application Number
CN202510351257.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

When determining the oil saturation of shale in the prior art, the core experimental determination method is time-consuming and costly. The logging prediction method is inaccurate in the mixed shale and cannot meet the needs of reservoir oil content evaluation.

Method used

By obtaining the oil saturation data and logging data of the full-diameter core on site, an oil saturation prediction model is constructed, the model is trained using machine learning methods, and the model parameters are optimized using the training set and verification set to improve the prediction accuracy, and quantitative analysis is carried out in combination with the characteristics of mixed shale.

Benefits of technology

The accuracy of shale oil saturation prediction is improved, key control factors are clearly quantified, and geological evaluation and reservoir transformation decisions are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296416A_ABST
    Figure CN120296416A_ABST
Patent Text Reader

Abstract

The invention provides a shale oiliness determination method and device, equipment and a storage medium. According to the method, logging data corresponding to multiple depths in a to-be-logged well are obtained, the logging data are input into an oil saturation prediction model, oil saturation prediction values corresponding to the multiple depths output by the oil saturation prediction model are obtained, and then the oil saturation prediction values corresponding to the multiple depths are obtained based on the oil saturation prediction values corresponding to the multiple depths. When the oil saturation prediction model is constructed, due to the fact that the model is trained through the oil saturation data of the on-site full-diameter rock core, the precision of the oil saturation prediction model can be improved, and the prediction result output through the model is more accurate; therefore, the prediction precision of the shale oil saturation is improved, and quantitative analysis and interpretability analysis can be carried out on the oiliness main control factors by the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of oil and gas exploration, and particularly to a method, device, equipment and storage medium for determining shale oiliness. Background Art

[0002] Mixed sedimentary shale has typical characteristics such as extremely strong vertical and horizontal heterogeneity, well-developed multi-source fine-grained mixed sedimentary laminations, complex micro-nano pore-fracture structures, and complex fluid occurrence states, which pose great challenges to the systematic characterization and quantitative evaluation of key elements such as shale reservoir properties, oiliness, mobility, and compressibility. Shale oil saturation is an important reservoir evaluation index, which is related to reservoir evaluation, selection of development horizons, and geological reserve assessment.

[0003] Existing methods for determining shale oil saturation are mainly divided into two categories: core experiment measurement method and logging prediction method. The core experiment measurement method mainly conducts two-dimensional nuclear magnetic resonance experiments or other measurement methods based on sealed core data to obtain shale oil saturation data. The logging prediction method mainly uses the means of core calibration logging and calculates the oil saturation using data such as conventional logging or nuclear magnetic resonance logging.

[0004] The core experiment measurement method will consume a large amount of time and cost during the process of systematic sealed coring and experiments. Due to the complex rock-electric relationship in mixed sedimentary shale, the traditional Archie model is not applicable in the logging prediction method, and the result of calculating the oil saturation using the conventional logging resistivity method is relatively inaccurate; while the radial detection depth of nuclear magnetic resonance logging is limited, and the oil saturation value in the flushed zone reflected by it will be affected by drilling mud and cannot reflect the real formation conditions, resulting in a low final calculated value.

[0005] The above methods do not consider the influence of the characteristics of mixed sedimentary shale on shale oiliness, resulting in unclear understanding of the main controlling factors of shale oiliness and lack of more accurate quantitative characterization methods, and their calculation accuracy cannot meet the requirements of reservoir oiliness evaluation. Summary of the Invention

[0006] The embodiments of the present application provide a method, device, equipment and storage medium for determining shale oiliness, so as to solve the problem that the prediction results output by the existing oil saturation prediction model are not accurate enough.

[0007] In a first aspect, the present application provides a method for constructing an oil saturation prediction model, including:

[0008] Obtaining the oil saturation data of a full-diameter core on site and the logging data corresponding to the depth;

[0009] Using the oil saturation data as a label data set and the logging data as a feature data set;

[0010] Determine a training set and a validation set according to the label data set and the feature data set;

[0011] Iteratively train the prediction model using the training set to obtain a trained prediction model;

[0012] Validate the trained prediction model using the validation set;

[0013] When the validation result indicates successful validation, determine the prediction model as an oil saturation prediction model.

[0014] Optionally, the determining a training set and a validation set according to the label data set and the feature data set includes:

[0015] Extract data from the label data set and the feature data set according to a first preset ratio as a first data set;

[0016] Extract data from the label data set and the feature data set according to a second preset ratio as a second data set;

[0017] Process the data of the first data set and the second data set to obtain a training set and a validation set.

[0018] Optionally, the iteratively training the prediction model using the training set to obtain a trained prediction model includes:

[0019] Input the feature data in the training set into the prediction model to obtain a prediction result output by the prediction model;

[0020] Determine a loss function of the prediction model based on the prediction result and the corresponding label data in the training set;

[0021] When the loss function reaches a preset threshold or the number of iterations reaches a preset number, determine that the prediction model is trained;

[0022] When the loss function does not reach the preset threshold, adjust the model parameters of the prediction model and re-iteratively train the adjusted prediction model using the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

[0023] Optionally, the method further includes:

[0024] Determine an optimal hyperparameter combination of the prediction model from multiple candidate hyperparameter combinations of the prediction model based on a preset objective function;

[0025] Re-fit the prediction model according to the optimal hyperparameter combination to obtain an optimal prediction model.

[0026] In a second aspect, the present application provides a method for determining shale oiliness, including:

[0027] Obtaining well logging data corresponding to multiple depths in a well to be measured;

[0028] Inputting the well logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model, wherein the oil saturation prediction model is trained by using the training method of the oil saturation prediction model described in the first aspect, and the oil saturation prediction values corresponding to multiple depths are used to indicate the corresponding oiliness.

[0029] Optionally, the method further includes:

[0030] Inputting the oil saturation prediction model into an analysis model, and analyzing and processing the prediction result output by the oil saturation model based on the analysis model to output an analysis result;

[0031] Determining the influence degree of different features and the interaction between features in the feature dataset corresponding to the oil saturation prediction model on the prediction result output by the oil saturation prediction model according to the analysis result.

[0032] In a third aspect, the present application provides a device for constructing an oil saturation prediction model, including:

[0033] An acquisition module, configured to acquire oil saturation data of a full-diameter core on site and well logging data corresponding to the depth;

[0034] A processing module, configured to use the oil saturation data as a label dataset and the well logging data as a feature dataset;

[0035] A determination module, configured to determine a training set and a validation set according to the label dataset and the feature dataset;

[0036] The processing module is further configured to perform iterative training processing on a prediction model by using the training set to obtain a trained prediction model;

[0037] The processing module is further configured to perform validation processing on the trained prediction model by using the validation set;

[0038] The determination module is further configured to determine the prediction model as an oil saturation prediction model when the verification result indicates successful verification.

[0039] Optionally, the processing module is further configured to:

[0040] Extract data from the label dataset and the feature dataset according to a first preset ratio as a first dataset;

[0041] Extract data from the label dataset and the feature dataset according to a second preset ratio as a second dataset;

[0042] Process the data in the first dataset and the second dataset to obtain a training set and a validation set.

[0043] Optionally, the processing module is further configured to input the feature data in the training set into the prediction model to obtain a prediction result output by the prediction model;

[0044] The determining module is further configured to determine a loss function of the prediction model based on the prediction result and the corresponding label data in the training set;

[0045] The determining module is further configured to determine that the training of the prediction model is completed when the loss function reaches a preset threshold or the number of iterations reaches a preset number;

[0046] The processing module is further configured to, when the loss function does not reach the preset threshold, adjust the model parameters of the prediction model and re-iteratively train the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

[0047] Optionally, the determining module is further configured to determine an optimal hyperparameter combination of the prediction model from multiple candidate hyperparameter combinations of the prediction model based on a preset objective function;

[0048] The processing module is further configured to refit the prediction model according to the optimal hyperparameter combination to obtain an optimal prediction model.

[0049] In a fourth aspect, the present application provides a device for determining shale oiliness, the device includes:

[0050] An acquisition module, configured to acquire well logging data corresponding to multiple depths in a well to be measured;

[0051] A processing module, configured to input the well logging data into an oil saturation prediction model to obtain oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model, where the oil saturation prediction model is trained by using the training method of the oil saturation prediction model described in the first aspect and all possible implementation manners of the first aspect, and the oil saturation prediction values corresponding to multiple depths are used to indicate the corresponding oiliness.

[0052] Fifth aspect, the present application provides a device for constructing an oil saturation prediction model, including:

[0053] A memory;

[0054] A processor;

[0055] Wherein, the memory stores computer-executable instructions;

[0056] The processor executes the computer-executable instructions stored in the memory to implement the method as described in the first aspect and various possible implementation manners of the first aspect above.

[0057] Sixth aspect, the present application provides a device for determining shale oiliness, including:

[0058] A memory;

[0059] A processor;

[0060] Wherein, the memory stores computer-executable instructions;

[0061] The processor executes the computer-executable instructions stored in the memory to implement the method for determining shale oiliness as described in the second aspect and various possible implementation manners of the second aspect above.

[0062] Seventh aspect, the present application provides a computer storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the method as described in the first aspect and various possible implementation manners of the first aspect or the second aspect and various possible implementation manners of the second aspect above.

[0063] The present application provides a method, device, equipment and storage medium for determining shale oiliness. The method obtains well logging data corresponding to multiple depths in a well to be measured, inputs the well logging data into an oil saturation prediction model, and obtains oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model. Then, based on the oil saturation prediction values corresponding to multiple depths, the oiliness of the shale is determined. Among them, when constructing the oil saturation prediction model, since the oil saturation data of the full-diameter core on site is used to train the model, the accuracy of the oil saturation prediction model can be improved, making the prediction results output by the model more accurate. At the same time, the method comprehensively considers the characteristics of the mixed sedimentary shale, such as the mixing of multi-source fine-grained minerals and laminations, extremely strong heterogeneity, and micro-migration of shale oil in the source-reservoir separation. Therefore, while improving the prediction accuracy, it can more clearly quantify and analyze the key control factors of the oil saturation of the mixed sedimentary shale, has a visual exploration value for the reasons of oil content differences, and is helpful for subsequent geological evaluation and reservoir stimulation decision-making. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0065] Figure 1 Flow chart of a method for constructing an oil saturation prediction model provided for an embodiment of the present application Figure 1 ;

[0066] Figure 2 Flow chart of a method for constructing an oil saturation prediction model provided for an embodiment of the present application Figure 2 ;

[0067] Figure 3 Flow chart of a method for determining shale oiliness provided for an embodiment of the present application;

[0068] Figure 4 Ranking diagram of the importance degree of the oiliness characteristics of mixed sedimentary shale based on SHAP analysis provided for an embodiment of the present application;

[0069] Figure 5 Ranking diagram of the main controlling factors of the oiliness of mixed sedimentary shale based on SHAP analysis provided for an embodiment of the present application;

[0070] Figure 6 Scatter plot of the characteristic interaction effect of the influencing factors of the oiliness of mixed sedimentary shale based on SHAP analysis provided for an embodiment of the present application;

[0071] Figure 7 Structure diagram of a device for constructing an oil saturation prediction model provided for an embodiment of the present application;

[0072] Figure 8 Structure diagram of a device for constructing an oil saturation prediction model provided for an embodiment of the present application.

[0073] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These accompanying drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to explain the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0074] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below in conjunction with the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0075] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.

[0076] In the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0077] The oil saturation of shale is a key indicator for evaluating reservoir quality and is of great significance for reservoir evaluation, selection of development layers and estimation of geological reserves.

[0078] At present, there are two main methods for determining shale oil saturation: one is the experimental determination method based on core samples, and the other is the logging prediction method. The experimental determination method mainly relies on samples obtained by closed coring, and obtains shale oil saturation data by conducting two-dimensional nuclear magnetic resonance experiments or other related testing methods. The logging prediction method uses core calibration logging technology and conventional logging data or nuclear magnetic resonance logging data to infer oil saturation.

[0079] However, the core test method requires a lot of time and cost in the process of closed coring and subsequent experiments. In the environment of mixed shale with complex rock-electricity relationships, the traditional Archie model is not applicable to the logging prediction method. Therefore, the results of calculating oil saturation using conventional logging resistivity method are often not accurate enough. In addition, although nuclear magnetic resonance logging can provide information about fluid saturation, its radial detection depth is limited. The oil saturation of the flushing zone reflected is easily affected by drilling mud and cannot truly reflect the formation conditions, resulting in a low calculated oil saturation value.

[0080] That is, none of the above methods consider the influence of the characteristics of mixed shale on the oil content of shale in combination with its own characteristics, resulting in unclear understanding of the main controlling factors of shale oil content and lack of more accurate quantitative characterization methods. The calculation accuracy cannot meet the needs of reservoir oil content evaluation.

[0081] In view of the above problems existing in the prior art, the present application provides a method for determining shale oiliness. The method obtains well logging data corresponding to multiple depths in a well to be measured, inputs the well logging data into an oil saturation prediction model, and obtains oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model. Then, based on the oil saturation prediction values corresponding to multiple depths, the oiliness of the shale is determined. When constructing the oil saturation prediction model, since the oil saturation data of on-site full-diameter cores is used to train the model, the accuracy of the oil saturation prediction model can be improved, making the prediction results output by the model more accurate. At the same time, the method comprehensively considers the characteristics of multi-source fine-grained minerals and laminated mixing, extremely strong heterogeneity, and micro-migration of shale oil in source-reservoir separation in mixed accumulation shales. Therefore, while improving the prediction accuracy, it can more clearly quantify and analyze the key control factors of oil saturation in mixed accumulation shales, and has visual exploration value for the reasons of oil content differences, which is helpful for subsequent geological evaluation and reservoir reconstruction decision-making.

[0082] The following uses specific embodiments to detail the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application with reference to the accompanying drawings.

[0083] Figure 1 It is a flow chart of a method for constructing an oil saturation prediction model provided by an embodiment of the present application Figure 1 As Figure 1 shown, the method for constructing an oil saturation prediction model provided in this embodiment includes:

[0084] S101: Obtain the oil saturation data of on-site full-diameter cores and the well logging data corresponding to the depths.

[0085] Among them, the oil saturation data refers to the proportion of the oil-phase fluid in the rock within a certain well logging section. The well logging data can be, for example, porosity, total organic carbon, free hydrocarbons, mineral content, and conventional well logging curve data.

[0086] It can be understood that in this embodiment, by performing two-dimensional nuclear magnetic resonance measurement on the full-diameter core of the target well logging, a two-dimensional nuclear magnetic resonance spectrum is obtained, and the oil saturation data is extracted from the two-dimensional nuclear magnetic resonance T1-T2 spectrum.

[0087] Further, according to the well logging depth corresponding to the oil saturation, the porosity, total organic carbon, free hydrocarbons, mineral content, and conventional well logging curve data (natural potential, deep resistivity, neutron, density, acoustic travel time, natural gamma, etc.) corresponding to the depth are obtained.

[0088] By obtaining the oil saturation data of the target well logging and the well logging data at the corresponding depth, the physical properties, oil-bearing properties, fluid properties, etc. of the reservoir can be comprehensively analyzed, providing a basis for reservoir classification, reserve calculation, and oil and gas production capacity prediction.

[0089] S102: Use the oil saturation data as the label data set and the well logging data as the feature data set.

[0090] Among them, the label data in the label data set is the object that the machine learning model needs to predict. The feature data in the feature data set is used to train the machine learning model.

[0091] It can be understood that in this embodiment, the oil saturation data is used as the label, and multiple well logging data, such as porosity, total organic carbon, free hydrocarbons, mineral content, spontaneous potential, deep resistivity, etc., are used as features. The obtained oil saturation data and well logging data are processed for missing values or outliers to ensure the integrity and accuracy of the data, and then the label data set and the feature data set are obtained.

[0092] S103: Determine the training set and the validation set according to the label data set and the feature data set.

[0093] Among them, the training set is used to guide the model to learn the rules and patterns in the data, and contains all the features and labels that the model needs to learn. The validation set is used to monitor the performance of the model during training and help adjust the parameters of the model.

[0094] It can be understood that the label data set and the feature data set obtained in the above steps are divided into two parts according to a certain proportion. Specifically, one part is used to train the machine learning model, and the other part is used to evaluate the model performance.

[0095] It can also be understood that when dividing the data set, a random division method is adopted, that is, the probability of each data point being assigned to the training set or the validation set is equal, which helps to avoid bias in the data division process and thus improve the generalization ability of the model.

[0096] Use the training set to train the model and use the validation set to evaluate the performance of the model. During training, monitor the performance metrics on the validation set to determine whether the model has overfitting or underfitting. If the performance of the model on the validation set is not good, the parameters or structure of the model can be adjusted to improve its performance.

[0097] By determining the training set and the validation set, and training and adjusting the machine learning model through the training set and the validation set, it is beneficial to improve the performance and generalization ability of the model.

[0098] S104: Perform iterative training processing on the prediction model using the training set to obtain a trained prediction model.

[0099] It is understandable that, by using the obtained training data set, through multiple repeated training processes, the internal parameters and structure of the prediction model are continuously adjusted and optimized until the model reaches the predetermined performance standard or convergence condition on the training set, thereby obtaining a trained model that can be used for actual prediction.

[0100] It is also understandable that, by observing the change trend of the loss function during the training process, it can be judged whether the model has converged to the optimal solution or is close to the optimal solution. If the loss function value no longer decreases significantly after multiple iterations, it is considered that the model has converged and the training can be stopped.

[0101] During the iterative training process, the parameters of the model, such as the learning rate, batch size, number of iterations, etc., can also be adjusted according to the performance of the model on the training set and the validation set to further optimize the performance of the model.

[0102] In each iteration, a batch of samples is randomly selected from the training set, the prediction error of the model is calculated, and the parameters of the model are updated according to the gradient of the error. This process will be repeated until the predetermined number of iterations or convergence condition is reached.

[0103] The purpose of this step is to perform iterative training processing on the prediction model through the training set, and then obtain a trained prediction model with good performance.

[0104] S105: Use the validation set to perform validation processing on the trained prediction model.

[0105] Specifically, the samples in the validation set are input into the model, and the prediction results of the model are collected. According to the prediction results of the model and the true labels of the validation set, a series of performance metrics are calculated, such as accuracy, recall, F1 score, mean squared error, etc.

[0106] By using the validation set to perform validation processing on the trained prediction model, the performance of the model can be comprehensively evaluated, and guidance for further model optimization can be provided.

[0107] S106: In the case where the validation result indicates successful validation, determine the prediction model as the oil saturation prediction model.

[0108] It is understandable that by analyzing performance metrics such as accuracy, recall, F1 score, etc. on the validation set, as well as the comparison between the prediction results and the true labels, and using the analysis results as the basis for judging whether the model is successful.

[0109] Further, according to the verification result, evaluate whether the model meets the predetermined performance criteria. If the performance of the model on the validation set reaches the predetermined standard and the deviation between the prediction result and the true label is within the preset error range, it is considered that the verification is successful, and the model is determined as the oil saturation prediction model.

[0110] By determining the prediction model as the oil saturation prediction model when the verification result indicates success, the application requirements in the field of oil and gas exploration and development are met.

[0111] A method for constructing an oil saturation prediction model provided in this embodiment includes obtaining the oil saturation data of the target well logging and the well logging data at the corresponding depth; using the oil saturation data as the label data set and the well logging data as the feature data set; determining the training set and the validation set according to the label data set and the feature data set; performing iterative training on the prediction model using the training set to obtain the trained prediction model; performing verification on the trained prediction model using the validation set; and determining the prediction model as the oil saturation prediction model when the verification result indicates successful verification, thereby improving the accuracy of the oil saturation prediction model.

[0112] Figure 2 The flow chart of a method for constructing an oil saturation prediction model provided in an embodiment of the present application Figure 2 This embodiment is based on Figure 1 an embodiment and details a possible implementation manner of the method for constructing an oil saturation prediction model. As Figure 2 shown, the method includes:

[0113] S201: Obtain the oil saturation data of the full-diameter core in the field and the well logging data at the corresponding depth.

[0114] Among them, step S201 is similar to the above step S101 and will not be elaborated here.

[0115] S202: Use the oil saturation data as the label data set and the well logging data as the feature data set.

[0116] Among them, step S202 is similar to the above step S102 and will not be elaborated here.

[0117] S203: Extract data from the label data set and the feature data set according to the first preset ratio as the first data set.

[0118] Among them, the first preset ratio is set according to the total amount of data and the requirements of the experiment.

[0119] Understandably, according to a preset ratio, corresponding proportions of data are randomly selected from the label dataset and the feature dataset to form a new dataset, which can be used for subsequent machine learning model training.

[0120] S204: Extract data from the label dataset and the feature dataset according to a second preset ratio as the second dataset.

[0121] Understandably, in the above step S203, a part of the data is divided from the feature dataset and the label dataset according to a preset ratio as the first dataset. Further, in this embodiment, the remaining data in the feature dataset and the label dataset are used as the second dataset and are used for subsequent machine learning model verification.

[0122] S205: Process the data in the first dataset and the second dataset to obtain a training set and a validation set.

[0123] Understandably, preprocess the first dataset and the second dataset, including missing value processing, normalization, etc. Among them, the purpose of preprocessing is to improve the data quality so that the model can learn more effectively.

[0124] Specifically, missing value processing refers to checking whether there are missing or blank values in the dataset and taking corresponding measures to fill or delete them to avoid the negative impact of these missing values on model training. Common missing value processing methods include filling with statistical quantities such as mean, median, and mode, or using more complex interpolation algorithms to estimate missing values.

[0125] Normalization processing is to scale the feature values in the dataset to a specified range, usually between 0 and 1. For example, the MinMaxScaler method is used to scale the data between 0 and 1 according to the minimum and maximum values of each feature to eliminate the bias caused by different dimensions or value ranges between different features, so that the model can treat each feature more equally during training.

[0126] This step processes the data in the first dataset and the second dataset to obtain a training set and a validation set, providing a basis for the training and evaluation of the machine learning model.

[0127] S206: Input the feature data in the training set into the prediction model to obtain the prediction result output by the prediction model.

[0128] In this embodiment, the prediction model is constructed based on the XGBoost algorithm and is used to predict the oil saturation. Input the feature data in the training set into the prediction model, and the prediction model will calculate the predicted output value according to the input feature data using the rules and patterns learned internally.

[0129] S207: Determine the loss function of the prediction model based on the prediction results and the corresponding label data in the training set.

[0130] Among them, the loss function provides a quantitative criterion for measuring the accuracy of the model prediction. During the training process, the loss function is used to guide the update of the model parameters. By minimizing the loss function, the optimal model parameters can be found, thereby improving the prediction performance of the model.

[0131] It can be understood that the calculated loss value can be used to evaluate the performance of the model. If the loss value is relatively high, it indicates that the prediction performance of the model is poor, and the parameters or structure of the model need to be adjusted. Through iterative training and adjustment, the loss value can be gradually reduced, and the prediction accuracy of the model can be improved.

[0132] S208: Determine that the training of the prediction model is completed when the loss function reaches a preset threshold or the number of iterations reaches a preset number.

[0133] It can be understood that by setting the threshold of the loss function, it can be ensured that the model will not overfit the training data during the training process, thereby maintaining the generalization ability for new data. Setting an upper limit on the number of iterations can ensure that the training process will not continue indefinitely, thereby saving time and computing resources.

[0134] After each iteration, check whether the value of the loss function has reached the preset threshold or whether the number of iterations has reached the preset upper limit. If either of these conditions is met, it can be considered that the model training is completed.

[0135] Determining the conditions for completing the training of the prediction model by setting the loss function threshold and the upper limit of the number of iterations helps to avoid overfitting, control the training time, and ensure that the model can converge to near the optimal solution.

[0136] S209: In the case where the loss function does not reach the preset threshold, adjust the model parameters of the prediction model and re-iterate the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

[0137] It can be understood that after each iteration training is completed, the value of the loss function of the current model is calculated and compared with the preset threshold. If the value of the loss function does not reach the preset threshold, the parameters of the model are adjusted based on the gradient information of the parameters of the loss function, and the predicted value and the loss function value are calculated according to the new parameters.

[0138] After each re-training, check again whether the value of the loss function has reached the preset threshold or whether the number of iterations has reached the preset maximum value. If either of these conditions is met, stop the training.

[0139] By iteratively adjusting the model parameters and retraining until the loss function reaches a preset threshold or the number of iterations reaches a preset number, it helps to optimize the performance of the model.

[0140] S210: Use the validation set to perform a validation process on the trained prediction model.

[0141] It can be understood that after training is completed, the validation set is used to evaluate the model. This includes calculating various performance metrics of the model on the validation set, such as accuracy, recall rate, F1 score, etc. According to the performance metrics on the validation set, the performance of the model is analyzed.

[0142] Using the validation set to perform a validation process on the trained prediction model helps to evaluate the performance of the model, prevent overfitting, adjust the model parameters, and select the best model.

[0143] S211: In the case where the validation result indicates successful validation, determine the prediction model as the oil saturation prediction model.

[0144] It can be understood that before starting the validation, it is necessary to clarify the performance metrics that the model needs to achieve on the validation set and the corresponding thresholds.

[0145] Furthermore, compare the calculated performance metrics with the preset success criteria. If the performance metrics reach or exceed the success criteria, it is considered that the validation is successful.

[0146] In the case of successful validation, determine the currently trained model as the prediction model for oil saturation prediction.

[0147] In an optional embodiment, the method for constructing the oil saturation prediction model further includes: determining the optimal hyperparameter combination of the prediction model from multiple candidate hyperparameter combinations of the prediction model based on a preset objective function, and refitting the prediction model according to the optimal hyperparameter combination to obtain the best prediction model.

[0148] Exemplarily, construct a prediction model based on the XGBoost algorithm, and use methods such as Bayesian optimization to search for the optimal hyperparameters of XGBoost. Among them, Bayesian optimization uses a probability model to approximate the objective function and a acquisition function to guide the search process, so as to efficiently find the optimal solution in the parameter space.

[0149] It is understandable that the objective function is used to evaluate the performance of the XGBoost model under given hyperparameters. The objective function will accept a set of hyperparameters as input and return the corresponding model performance score. The output of the objective function can be, for example, the coefficient of determination R2, the root mean square error RMSE. The Bayesian optimization algorithm is used to search for the optimal hyperparameters of XGBoost. In each iteration, the algorithm will select the next set of hyperparameters to be evaluated based on the current estimate of the objective function. Through continuous iteration, the algorithm will gradually converge to the optimal solution.

[0150] Based on the trained oil saturation prediction model, the optimal parameters can be used to predict the oil saturation at multiple target logging depths, which helps to improve the efficiency and accuracy of oil exploration and production.

[0151] A method for constructing an oil saturation prediction model provided by an embodiment of the present application. First, obtain the oil saturation data of the target well logging and the corresponding logging data of the depth, and use them as the label data set and the feature data set respectively. Extract the corresponding data from the label data set and the feature data set according to a preset ratio as the training set and the validation set. Input the feature data in the training set into the prediction model to obtain the prediction result output by the prediction model. Then, based on the prediction result and the corresponding label data in the training set, determine the loss function of the prediction model. When the loss function reaches the preset threshold or the number of iterations reaches the preset number, it is determined that the training of the prediction model is completed. When the loss function does not reach the preset threshold, adjust the model parameters of the prediction model and re-iterate the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number. Finally, use the validation set to verify the trained prediction model. When the verification result indicates successful verification, determine the prediction model as the oil saturation prediction model. This method performs multiple iterative trainings on the prediction model until the preset conditions are met, improving the accuracy and reliability of the oil saturation prediction model.

[0152] Figure 3 It is a flow chart of a method for determining the shale oil saturation provided by an embodiment of the present application. As Figure 3 shown, the method for determining the shale oil saturation provided by this embodiment includes:

[0153] S301: Obtain the logging data corresponding to multiple depths in the well to be measured.

[0154] Among them, the logging data includes porosity, total organic carbon, free hydrocarbons, mineral content, and conventional logging curves, etc.

[0155] S302: Input the logging data into the oil saturation prediction model to obtain the oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model.

[0156] Among them, the oil saturation prediction model is trained according to the training method of the oil saturation prediction model in the above embodiment. The oil saturation prediction values corresponding to multiple depths are used to indicate the corresponding oiliness.

[0157] It can be understood that the logging data of the well to be measured obtained is input into the trained oil saturation prediction model. The oil saturation prediction model analyzes and processes the logging data at different depths according to its own structure and parameters, and then outputs the oil saturation prediction value corresponding to the depth. Using porosity, total organic carbon, free hydrocarbons, mineral content, and conventional logging curve data, etc. as feature inputs is to select parameters by comprehensively considering the characteristics of mixed shale.

[0158] In an optional embodiment, the method for determining shale oil saturation further includes: inputting the oil saturation prediction model into an analysis model, and based on the analysis model, analyzing and processing the prediction result output by the oil saturation model, and outputting an analysis result; according to the analysis result, determining the influence degree of different features and the interaction between features in the characteristic dataset corresponding to the oil saturation prediction model on the prediction result output by the oil saturation prediction model.

[0159] Exemplarily, based on the oil saturation prediction model that has been trained and optimized, the SHAP (SHapley Additive exPlanations) analysis method is further used to visually and numerically evaluate the contribution degree of each feature to the prediction result of each sample.

[0160] The main process is explained as follows:

[0161] 1. Calculate SHAP values: Using the trained oil saturation prediction model, calculate the SHAP values of each sample and each feature for the validation set or test set data respectively; SHAP introduces the idea of Shapley values in game theory to measure the expected incremental contribution of a certain feature to the model output from the process of "adding" to "not adding".

[0162] In this process, for example, the following formula can be used to determine the SHAP value of feature i:

[0163]

[0164] Among them, F is the set of all features, S is the subset that does not include feature i, f(S) is the prediction value of the model on the feature subset S, |S|! is the factorial of the subset size, and f(S∪{i}) - f(S) is the prediction change after feature i is added to subset S.

[0165] In this formula, the combination weight It can ensure that all possible feature subsets are equally weighted.

[0166] 2. Cumulation and visualization: Aggregate the SHAP values of each sample and draw a graph ranking the feature importance; for example, with the visualization tools provided by SHAP, visually display the average contribution degree and positive / negative impact of each feature.

[0167] In this process, the SHAP values corresponding to each record can be calculated for all the data in the test set and the training set, obtaining the visualization results of the contribution degree of each feature to the oil saturation and the interaction effect; then draw a bar graph and a scatter plot of the feature importance of SHAP to identify the most critical control factors (such as porosity, TOC, free hydrocarbons, mineral content, a certain curve in the conventional logging curves, etc.).

[0168] Figure 4 This is a graph ranking the importance degree of the oil-bearing characteristics of the mixed shale based on SHAP analysis provided by an embodiment of the present application. As Figure 4 shown, the ranking of the importance degree of the influencing factors of shale oil-bearing property is: porosity > clay minerals > total organic carbon > free hydrocarbons > spontaneous potential > analcime > deep resistivity > dolomite > neutron > quartz > density > acoustic time difference > natural gamma ray.

[0169] Figure 5 This is a graph ranking the main control factors of the oil-bearing property of the mixed shale based on SHAP analysis provided by an embodiment of the present application. Among them, the pink dots are used to indicate that the feature values have a positive impact on the model prediction, and the blue dots are used to indicate that the feature values have a negative impact on the model prediction.

[0170] The horizontal axis (SHAP value) is used to characterize the magnitude of the influence of the feature on the prediction result. The farther the point is from the center line (zero point), the greater the influence of the feature on the model output. A positive SHAP value indicates a positive impact, and a negative SHAP value indicates a negative impact; the vertically arranged features are ranked from the largest to the smallest in terms of influence. The features above have a greater total influence on the model output, while the features below have a smaller influence.

[0171] Therefore, combined with Figure 5 it can be known that porosity, clay minerals, total organic carbon, and free hydrocarbon content are the main control factors of shale oil-bearing property. Among them, the high value of porosity has a greater positive contribution to the model, while the low value of clay minerals has a greater positive relationship with the model.

[0172] 3. Identify the main control factors and interaction effects: According to the magnitude of the feature SHAP values, identify the key parameters that contribute the most to the prediction result; through the interaction SHAP values between features, the mutual coupling effects existing between different features can be determined, thus showing positive or negative adjustment effects in the prediction.

[0173] Among them, by analyzing the scatter plot of the interaction effects between features, the internal mechanism of the difference in oil-bearing properties of mixed-source shale can be revealed from a data-driven perspective.

[0174] Figure 6 This is a scatter plot of the interaction effects of the influencing factors of the oil-bearing properties of mixed-source shale based on SHAP analysis provided by the embodiments of the present application.

[0175] As Figure 6 shown, as the porosity gradually increases, the SHAP increases rapidly in the range of 2-4%, indicating a rapid increase in the contribution to the model; when the SHAP of porosity shows a positive effect, the corresponding clay content is low, reflecting the negative effect of clay minerals; differently, the interaction diagram between free hydrocarbons and total organic carbon as a whole shows a positive linear correlation, which reflects that high free hydrocarbon and total organic carbon contents both have a positive effect on the model performance. The analysis of the interaction effects between such features can reveal the importance of the relationship between the difference in oil-bearing properties and source-reservoir configuration in mixed-source shale, and also reflects that the oil-bearing properties of shale are not controlled by a single factor, but the result of the combined action of multiple factors.

[0176] In this process, by determining the influence degree of different features and the interaction between features in the characteristic data set corresponding to the oil saturation prediction model on the prediction result output by the oil saturation prediction model, not only the accuracy of oil saturation prediction is improved, but also the understanding of the geological characteristics of mixed-source shale reservoirs is deepened, providing a theoretical basis and technical support for subsequent exploration and development.

[0177] A method for determining the oil saturation of shale provided by the embodiments of the present application improves the prediction accuracy of the oil saturation of shale by obtaining well logging data corresponding to multiple depths in a well to be measured and inputting the well logging data into a trained and optimized oil saturation prediction model to obtain the oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model.

[0178] Figure 7 This is a schematic structural diagram of a device for constructing an oil saturation prediction model provided by the present application. As Figure 7 shown, the device 400 for constructing the oil saturation prediction model provided in this embodiment includes:

[0179] An acquisition module 401, configured to acquire the oil saturation data of the full-diameter core on site and the well logging data corresponding to the depth;

[0180] A processing module 402, configured to use the oil saturation data as a label data set and the well logging data as a characteristic data set;

[0181] A determination module 403, configured to determine a training set and a validation set according to the label data set and the characteristic data set;

[0182] The processing module 402 is further configured to perform iterative training processing on the prediction model by using the training set to obtain a trained prediction model;

[0183] The processing module 402 is further configured to perform verification processing on the trained prediction model by using the verification set;

[0184] The determination module 403 is further configured to determine the prediction model as an oil saturation prediction model when the verification result indicates successful verification.

[0185] Optionally, the processing module 402 is further configured to:

[0186] Extract data from the label data set and the feature data set according to a first preset ratio as a first data set;

[0187] Extract data from the label data set and the feature data set according to a second preset ratio as a second data set;

[0188] Process the data in the first data set and the second data set to obtain a training set and a verification set.

[0189] Optionally, the processing module 402 is further configured to input the feature data in the training set into the prediction model to obtain a prediction result output by the prediction model;

[0190] The determination module 403 is further configured to determine a loss function of the prediction model based on the prediction result and the corresponding label data in the training set;

[0191] The determination module 403 is further configured to determine that the prediction model is trained when the loss function reaches a preset threshold or the number of iterations reaches a preset number;

[0192] The processing module 402 is further configured to adjust the model parameters of the prediction model when the loss function does not reach the preset threshold, and re-perform iterative training processing on the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

[0193] Optionally, the determination module 403 is further configured to determine an optimal hyperparameter combination of the prediction model from multiple candidate hyperparameter combinations of the prediction model based on a preset objective function;

[0194] The processing module 402 is further configured to refit the prediction model according to the optimal hyperparameter combination to obtain an optimal prediction model.

[0195] The construction device of an oil saturation prediction model provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0196] This embodiment also provides a device for determining shale oiliness, which includes:

[0197] An acquisition module, configured to acquire well logging data corresponding to multiple depths in a well to be measured;

[0198] A processing module, configured to input the well logging data into the oil saturation prediction model to obtain oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model, where the oil saturation prediction model is trained by using the training method of the oil saturation prediction model described in the first aspect and various possible implementation manners of the first aspect, and the oil saturation prediction values corresponding to multiple depths are used to indicate the corresponding oiliness.

[0199] The device for determining shale oiliness provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0200] Figure 8 It is a schematic structural diagram of a device for constructing an oil saturation prediction model provided by this application. As Figure 8 shown, the device for constructing an oil saturation prediction model provided by this application, the device 500 for constructing an oil saturation prediction model includes: a receiver 501, a transmitter 502, a processor 503, and a memory 504.

[0201] The receiver 501 is configured to receive instructions and data;

[0202] The transmitter 502 is configured to send instructions and data;

[0203] The memory 504 is configured to store computer execution instructions;

[0204] The processor 503 is configured to execute the computer execution instructions stored in the memory 504 to implement each step executed by the method for constructing an oil saturation prediction model in the above embodiment. Specifically, reference can be made to the relevant descriptions in the above method embodiment for constructing an oil saturation prediction model.

[0205] Optionally, the above memory 504 can be either independent or integrated with the processor 503.

[0206] When the memory 504 is independently arranged, the electronic device further includes a bus for connecting the memory 504 and the processor 503.

[0207] The present application further provides a computer storage medium storing computer-executable instructions, which, when executed by a processor, implement the method for constructing an oil saturation prediction model as executed by the device for constructing an oil saturation prediction model as described above.

[0208] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof. In the hardware implementation, the division of the functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, one physical component may have multiple functions, or one function or step may be executed by several physical components in cooperation. Some or all of the physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, magnetic disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and may include any information delivery medium.

[0209] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0210] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A method for constructing an oil saturation prediction model, characterized in that, Including: Obtaining the oil saturation data of the full-diameter core on site and the logging data at the corresponding depth; Taking the oil saturation data as a label data set and taking the logging data as a feature data set; Determining a training set and a validation set according to the label data set and the feature data set; Performing iterative training processing on the prediction model using the training set to obtain a trained prediction model; Performing validation processing on the trained prediction model using the validation set; When the validation result indicates successful validation, determining the prediction model as an oil saturation prediction model.

2. The method according to claim 1, wherein The determining the training set and the validation set according to the label data set and the feature data set includes: Extracting data from the label data set and the feature data set according to a first preset ratio as a first data set; Extracting data from the label data set and the feature data set according to a second preset ratio as a second data set; Processing the data of the first data set and the second data set to obtain a training set and a validation set.

3. The method according to claim 2, wherein Performing training processing on the prediction model using the training set to obtain a trained prediction model, including: Inputting the feature data in the training set into the prediction model to obtain a prediction result output by the prediction model; Determining a loss function of the prediction model based on the prediction result and the corresponding label data in the training set; When the loss function reaches a preset threshold or the number of iterations reaches a preset number, determining that the prediction model is trained; When the loss function does not reach the preset threshold, adjusting the model parameters of the prediction model and re-performing iterative training processing on the adjusted prediction model based on the training set until the loss function reaches the preset threshold or the number of iterations reaches the preset number.

4. The method according to claim 3, wherein The method further includes: Determining an optimal hyperparameter combination of the prediction model from multiple candidate hyperparameter combinations of the prediction model based on a preset objective function; Re-fitting the prediction model according to the optimal hyperparameter combination to obtain an optimal prediction model.

5. A method for determining shale oil content, characterized in that, The method includes: Obtaining logging data corresponding to multiple depths in a well to be measured; Inputting the logging data into the oil saturation prediction model to obtain oil saturation prediction values corresponding to multiple depths output by the oil saturation prediction model, where the oil saturation prediction model is trained by the training method of the oil saturation prediction model according to any one of claims 1-4, and the oil saturation prediction values corresponding to multiple depths are used to indicate the corresponding oiliness.

6. The method according to claim 5, wherein The method further includes: Inputting the oil saturation prediction model into an analysis model and performing analysis processing on the prediction result output by the oil saturation model based on the analysis model to output an analysis result; Determining the influence degree of different features and the interaction between features in the feature data set corresponding to the oil saturation prediction model on the prediction result output by the oil saturation prediction model according to the analysis result.

7. An apparatus for constructing an oil saturation prediction model, characterized in that The device includes: An acquisition module for acquiring the oil saturation data of the full-diameter core on site and the logging data at the corresponding depth; A processing module, configured to use the oil saturation data as a label data set and the logging data as a feature data set; A determination module, configured to determine a training set and a validation set according to the label data set and the feature data set; The processing module is further configured to perform iterative training processing on the prediction model by using the training set to obtain a trained prediction model; The processing module is further configured to perform validation processing on the trained prediction model by using the validation set; The determination module is further configured to, when the validation result indicates successful validation, determine the prediction model as an oil saturation prediction model.

8. The device according to claim 7, characterized in that, The processing module is further configured to: Extract data from the label data set and the feature data set according to a first preset ratio as a first data set; Extract data from the label data set and the feature data set according to a second preset ratio as a second data set; Process the data in the first data set and the second data set to obtain a training set and a validation set.

9. An apparatus for constructing an oil saturation prediction model, characterized in that The device includes: A memory; A processor; Wherein, the memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory to implement the method for constructing an oil saturation prediction model according to any one of claims 1-4.

10. A computer storage medium, characterized in that, Computer execution instructions are stored in the computer storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method for constructing an oil saturation prediction model according to any one of claims 1-4.

Citation Information

Patent Citations

  • Rock core saturation prediction model construction method and rock core saturation prediction method

    CN112859192A

  • Method for continuously depicting oil content distribution of mixed shale oil

    CN117784220A

  • Saturation prediction method and device, electron and storage medium

    CN117803372A

  • Multi-scale rock physics fused reservoir saturation logging intelligent evaluation method and system

    CN118859354A

  • Oil reservoir parameter prediction method, device and equipment based on petrophysics

    CN118939955A

Cited By

  • Shale oil content prediction method, device, equipment and medium

    CN121235219A