Machine learning-based lithofacies prediction method and device, electronic equipment and medium
By using a random forest algorithm to establish a lithophase recognition model based on geomorphic and logging data in shale staple facies recognition, the problem of relying on experimental samples and high-cost logging in the existing technology is solved, and the accuracy and stability of lithophase prediction are improved.
Patent Information
- Application Number
- CN202311586259.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art relies on a large number of experimental samples and high-cost logging technology in shale staphylococcus recognition, and the accuracy of a single machine learning algorithm is not high enough and is greatly affected by noise and data imbalance.
The random forest algorithm is used to build a lithophase recognition model based on geomorphology and logging data, and the logging attributes that are sensitive to lithologies are preferred. As a combination classifier, the algorithm effect of random forests is stable, has strong generalization ability, and is not prone to overfitting.
It improves the accuracy and stability of lithophago prediction, reduces experimental costs, enhances the recognition ability of longitudinal lithophago distribution of shale, and provides better reservoir dessert prediction support.
Smart Images

Figure CN120044633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of shale petrophysics, and more particularly, to a lithofacies prediction method, apparatus, electronic device, and medium based on machine learning. Background Art
[0002] Lithology and lithofacies identification and division are important tasks in formation evaluation and fine reservoir description. Continental shale formations exhibit strong heterogeneity, and the vertical lithofacies distribution is complex. Currently, the most reliable and intuitive method for lithology identification is core drilling. However, the number of core wells is generally small, and the cored intervals are limited. Traditional methods for determining shale lithofacies rely on X-ray diffraction analysis (XRD) experimental data. With the development of logging technology, it is also possible to continuously identify rock mineral components on a single well through pulsed neutron spectroscopy (PNS) logging, thereby obtaining the lithofacies types and distribution characteristics in the vertical direction of a single well. However, XRD experiments require a large number of experimental samples, and the cost of PNS logging technology is relatively high, and most old work areas are not equipped with it.
[0003] Machine learning algorithms can quickly increase the data set through the training of a certain amount of data, thus making up for the inconvenience of obtaining a large number of samples due to the strong heterogeneity of shale, reducing the interpretation cost and improving the analysis effect. After Wolf et al. first published a method for automatically determining formation lithology based on logging data in 1982, using computers to identify lithology has become an important direction in the development of logging technology. Al-Mudhafar et al. (2019) successfully identified five carbonate lithofacies of the Mishrif formation in the West Qurna oilfield in southern Iraq based on 11 logging data such as natural gamma ray, spontaneous potential, and neutron porosity, using the K-means clustering algorithm. Currently, most algorithms are single machine learning algorithms, and the accuracy is generally not high, and they are greatly affected by the unbalanced distribution of lithofacies category data. For example, boosting algorithms such as XGBoost are easily affected by noise.
[0004] Currently, there is a need to develop a lithofacies prediction method based on machine learning.
[0005] The information disclosed in the background art section of the present invention is only intended to deepen the understanding of the general background art of the present invention, and should not be regarded as an admission or any form of suggestion that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0006] The present invention provides a lithofacies prediction method, device, electronic device and medium based on machine learning, which can use the random forest algorithm, based on geochemical and logging data, optimize the logging attributes sensitive to lithology, establish a lithofacies identification model. As a combined classifier, compared with a single classifier, the random forest algorithm has stable algorithm effect, stronger generalization ability, is not prone to overfitting, and has a better overall prediction effect, providing technical support for subsequent reservoir sweet spot prediction.
[0007] In a first aspect, an embodiment of the present disclosure provides a lithofacies prediction method based on machine learning, including:
[0008] Classify the core types in the target area to obtain a logging data sample set;
[0009] Determine the sample attribute values, and normalize all the attribute values of the logging data sample set;
[0010] Perform logging attribute analysis using the Pearson coefficient to obtain the correlations between logging attributes and between logging attributes and the target lithofacies;
[0011] Establish an initial classification model, and train the initial classification model according to the logging data sample set;
[0012] Evaluate the trained classification model through precision, recall and F1-score, adjust the model parameters, and output the final model.
[0013] As a specific implementation manner of the embodiment of the present disclosure, the core types in the target area include high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, and low-carbon mudstone facies.
[0014] As a specific implementation manner of the embodiment of the present disclosure, the sample attribute values include natural gamma, potassium, caliper logging, acoustic travel time, compensated neutron, and density.
[0015] As a specific implementation manner of the embodiment of the present disclosure, the correlation is:
[0016]
[0017] where PCC is the Pearson coefficient, that is, the correlation, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B.
[0018] As a specific implementation manner of the embodiment of the present disclosure, the precision is:
[0019]
[0020] Among them, a true positive example refers to a positive sample predicted as positive, and a false positive example refers to a negative sample predicted as positive.
[0021] As a specific implementation manner of the embodiment of the present disclosure, the recall rate is:
[0022]
[0023] Among them, a true positive example refers to a positive sample predicted as positive, and a false negative example refers to a positive sample predicted as negative.
[0024] As a specific implementation manner of the embodiment of the present disclosure, the F1-score is:
[0025]
[0026] In a second aspect, an embodiment of the present disclosure further provides a lithofacies prediction device based on machine learning, including:
[0027] A classification module that classifies the core types in the target area to obtain a logging data sample set;
[0028] A normalization module that determines sample attribute values and normalizes all attribute values of the logging data sample set;
[0029] A correlation calculation module that performs logging attribute analysis using the Pearson coefficient to obtain the correlations between logging attributes and between logging attributes and the target lithofacies;
[0030] A training module that establishes an initial classification model and trains the initial classification model according to the logging data sample set;
[0031] An evaluation module that evaluates the trained classification model through precision, recall rate, and F1-score, adjusts the model parameters, and outputs the final model.
[0032] As a specific implementation manner of the embodiment of the present disclosure, the core types in the target area include high-carbon silt-laminated clay shale facies, medium-carbon silt-laminated clay shale facies, carbonaceous mudstone facies, silt-laminated interbedded mixed shale facies, and carbon-poor mudstone facies.
[0033] As a specific implementation manner of the embodiment of the present disclosure, the sample attribute values include natural gamma, potassium, well diameter logging, acoustic time difference, compensated neutron, and density.
[0034] As a specific implementation manner of the embodiment of the present disclosure, the correlation is:
[0035]
[0036] Among them, PCC is the Pearson coefficient, which is the correlation. Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B.
[0037] As a specific implementation manner of the embodiment of the present disclosure, the precision rate is:
[0038]
[0039] Among them, a true positive example refers to a positive sample predicted as positive, and a false positive example refers to a negative sample predicted as positive.
[0040] As a specific implementation manner of the embodiment of the present disclosure, the recall rate is:
[0041]
[0042] Among them, a true positive example refers to a positive sample predicted as positive, and a false negative example refers to a positive sample predicted as negative.
[0043] As a specific implementation manner of the embodiment of the present disclosure, the F1-score is:
[0044]
[0045] In a third aspect, an embodiment of the present disclosure further provides an electronic device, which includes:
[0046] A memory storing executable instructions;
[0047] A processor that runs the executable instructions in the memory to implement the above-mentioned machine learning-based lithofacies prediction method.
[0048] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned machine learning-based lithofacies prediction method is implemented.
[0049] The method and device of the present invention have other characteristics and advantages, which will be obvious from the accompanying drawings incorporated herein and the subsequent specific embodiments, or will be detailed in the accompanying drawings incorporated herein and the subsequent specific embodiments. These drawings and specific embodiments are used together to explain the specific principles of the present invention. Description of the Drawings
[0050] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above-mentioned and other objects, features, and advantages of the present invention will become more obvious. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0051] Figure 1 The flowchart showing the steps of the machine learning-based lithofacies prediction method according to an embodiment of the present invention is presented.
[0052] Figure 2 The schematic diagram showing the correlation between different logging attributes according to an embodiment of the present invention is presented.
[0053] Figure 3 The schematic diagram showing the correlation analysis between different logging attributes and lithofacies according to an embodiment of the present invention is presented.
[0054] Figure 4 The schematic diagram showing the relationship between the model error and the number of machine learning iterations according to an embodiment of the present invention is presented.
[0055] Figure 5 The schematic diagram showing the confusion matrix of different lithofacies identification models according to an embodiment of the present invention is presented.
[0056] Figure 6a and Figure 6b The schematic diagrams respectively showing the comparison between the model prediction result and the true lithofacies result according to an embodiment of the present invention are presented.
[0057] Figure 7 The block diagram showing a machine learning-based lithofacies prediction device according to an embodiment of the present invention is presented.
[0058] Explanation of reference numerals:
[0059] 201, classification module; 202, normalization module; 203, correlation calculation module; 204, training module; 205, evaluation module. Detailed implementation manners
[0060] The preferred embodiments of the present invention will be described in more detail below. Although the preferred embodiments of the present invention are described below, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0061] To facilitate understanding of the solutions and effects of the embodiments of the present invention, six specific application examples are given below. Those skilled in the art should understand that this example is only for facilitating the understanding of the present invention, and any specific details are not intended to limit the present invention in any way.
[0062] Example 1
[0063] Figure 1 The flowchart showing the steps of the machine learning-based lithofacies prediction method according to an embodiment of the present invention is presented.
[0064] As Figure 1As shown in the figure, the machine learning-based lithofacies prediction method includes: Step 101, classifying the core types in the target area to obtain a logging data sample set; Step 102, determining the sample attribute values and normalizing all the attribute values of the logging data sample set; Step 103, performing logging attribute analysis using the Pearson coefficient to obtain the correlations between logging attributes and between logging attributes and the target lithofacies; Step 104, establishing an initial classification model and training the initial classification model according to the logging data sample set; Step 105, evaluating the trained classification model through precision, recall, and F1-score, adjusting the model parameters, and outputting the final model.
[0065] In one example, the core types in the target area include high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, and low-carbon mudstone facies.
[0066] In one example, the sample attribute values include natural gamma, potassium, caliper logging, acoustic travel time, compensated neutron, and density.
[0067] In one example, the correlation is:
[0068]
[0069] where PCC is the Pearson coefficient, which is the correlation, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B.
[0070] In one example, the precision is:
[0071]
[0072] where True Positive refers to the positive sample predicted as positive, and False Positive refers to the negative sample predicted as positive.
[0073] In one example, the recall is:
[0074]
[0075] where True Positive refers to the positive sample predicted as positive, and False Negative refers to the positive sample predicted as negative.
[0076] In one example, the F1-score is:
[0077]
[0078] Specifically, based on shale geochemical data and logging data, the core types in the target area are classified into 6 major categories, including high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, carbon-poor mudstone facies, and other categories. They are marked with numbers 1-5, 10 corresponding to different lithofacies types. Then, lithology digital tags are added after the logging data at each depth of the target layer, and the marked logging data sample set is obtained.
[0079] Select the following 6 conventional logging parameters as sample attribute values, namely natural gamma (GR), potassium (K), caliper logging (CAL), acoustic time difference (AC), compensated neutron (CNL), and density (DEN).
[0080] Different logging data have different dimensions and attribute value orders of magnitude. If the logging data are directly used as input to train the lithology identification model, their influence degrees on the results are different. To eliminate this systematic error, it is necessary to standardize the data, and all attribute values of the logging data sample set are scaled to the range (-1, 1). The calculation process is as follows:
[0081]
[0082] Among them, χ is the sample data, μ is the sample data mean, and σ is the sample data standard deviation.
[0083] The magnitude of the correlation between logging attributes and lithofacies reflects the strength of the ability of the logging attribute to reflect the information of reservoir lithofacies changes. The greater the correlation, the more accurately the logging attribute can reflect the distribution of lithofacies. At the same time, the magnitude of the correlation between different logging attributes also plays an important role in lithofacies identification. When the correlation between two logging attributes is large, it means that they provide a lot of similar information in lithofacies identification and do not provide more help for distinguishing different lithofacies; when the correlation between two logging attributes is small, it is difficult for the model to learn more parameter features from the logging data. Therefore, when using multiple logging attributes for lithofacies identification, it is necessary to comprehensively consider the correlation between different logging attributes and between logging attributes and lithofacies. The Pearson (PCC) coefficient is used for logging attribute analysis. By calculating the PCC coefficients between various logging attributes and the target lithofacies and between different logging attributes, the correlation degrees between logging attributes and between logging attributes and the target lithofacies are obtained. The calculation formula is as follows:
[0084]
[0085] Among them, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B. The value range of the PCC coefficient is from -1 to 1, where 0 indicates no correlation between the two, a positive number indicates positive correlation, and a negative number indicates negative correlation. The greater the absolute value of the PCC coefficient, the stronger the correlation.
[0086] An initial classification model is established. The CART decision tree is used as the base classifier. It builds a decision tree by randomly sampling the sample set with replacement and randomly sampling the feature set without replacement. The final generated multiple decision trees form a random forest model, and the final prediction result of the model is determined by the voting of the base classifiers. The CART algorithm is one of the commonly used decision tree generation algorithms. The CART decision tree is generated from top to bottom. That is, first, a certain feature is taken from the feature set of the root node, and the data set is divided into 2 sub-nodes according to this feature. Each sub-node contains a certain number of samples. Then, the same processing is performed on each sub-node respectively until the condition for the decision tree to stop growing is reached to complete the construction of the decision tree.
[0087] The 5-fold cross-validation method is used to select appropriate parameters for the model, including the number of iterations, the minimum number of samples in the leaf node, etc. The training set is randomly divided into 5 subsets of the same size. 4 of these subsets are used to train the model, and the remaining 1 subset is used to validate the model. This process is repeated 5 times. The cross-validation score of the model is obtained by averaging the 5 validation scores.
[0088] The precision rate, recall rate, and F1-score are used to evaluate the quality of the classification model. The formulas are as follows:
[0089]
[0090]
[0091]
[0092] Among them, true positive refers to the positive sample predicted as positive, false positive refers to the negative sample predicted as positive, true negative refers to the negative sample predicted as negative, and false negative refers to the positive sample predicted as negative.
[0093] After completing the parameter tuning of the model, the test set is used to verify the prediction effect of the established lithofacies recognition model. After the verification is completed, the lithofacies is predicted according to the verified lithofacies recognition model.
[0094] Example 2
[0095] The present invention also provides a lithofacies prediction device based on machine learning, including:
[0096] A classification module that classifies the core types in the target area to obtain a well logging data sample set;
[0097] Normalization module, which determines the sample attribute values and normalizes all the attribute values of the well logging data sample set;
[0098] Correlation calculation module, which performs well logging attribute analysis using the Pearson coefficient to obtain the correlations between well logging attributes and between well logging attributes and the target lithofacies;
[0099] Training module, which establishes an initial classification model and trains the initial classification model according to the well logging data sample set;
[0100] Evaluation module, which evaluates the trained classification model through precision, recall and F1-score, adjusts the model parameters and outputs the final model.
[0101] In one example, the core types in the target area include high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, and low-carbon mudstone facies.
[0102] In one example, the sample attribute values include natural gamma, potassium, well diameter logging, acoustic time difference, compensated neutron and density.
[0103] In one example, the correlation is:
[0104]
[0105] where PCC is the Pearson coefficient, which is the correlation, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B.
[0106] In one example, the precision is:
[0107]
[0108] where True Positive refers to the positive sample predicted as positive, and False Positive refers to the negative sample predicted as positive.
[0109] In one example, the recall is:
[0110]
[0111] where True Positive refers to the positive sample predicted as positive, and False Negative refers to the positive sample predicted as negative.
[0112] In one example, the F1-score is:
[0113]
[0114] Specifically, based on shale geochemical data and logging data, the core types in the target area are classified into 6 main categories, including high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, carbon-poor mudstone facies, and other categories. They are labeled with numbers 1-5, 10 corresponding to different lithofacies types. Then, lithology digital tags are added after the logging data at each depth of the target layer, and the labeled logging data sample set is obtained.
[0115] Select the following 6 conventional logging parameters as sample attribute values, namely natural gamma (GR), potassium (K), caliper logging (CAL), acoustic time difference (AC), compensated neutron (CNL), and density (DEN).
[0116] Different logging data have different dimensions and attribute value orders of magnitude. If the logging data are directly used as input to train the lithology identification model, their influence degrees on the results are different. To eliminate this systematic error, it is necessary to standardize the data, and all attribute values of the logging data sample set are scaled to the range (-1, 1). The calculation process is as follows:
[0117]
[0118] Among them, χ is the sample data, μ is the sample data mean, and σ is the sample data standard deviation.
[0119] The magnitude of the correlation between logging attributes and lithofacies reflects the strength of the ability of the logging attribute to reflect the information of reservoir lithofacies changes. The greater the correlation, the more accurately the logging attribute can reflect the distribution of lithofacies. At the same time, the magnitude of the correlation between different logging attributes also plays an important role in lithofacies identification. When the correlation between two logging attributes is large, it means that they provide a large amount of similar information in lithofacies identification and do not provide more help for distinguishing different lithofacies; when the correlation between two logging attributes is small, it is difficult for the model to learn more parameter features from the logging data. Therefore, when using multiple logging attributes for lithofacies identification, it is necessary to comprehensively consider the correlation between different logging attributes and between logging attributes and lithofacies. The Pearson (PCC) coefficient is used for logging attribute analysis. By calculating the PCC coefficients between various logging attributes and the target lithofacies and between different logging attributes, the correlation degrees between logging attributes and between logging attributes and the target lithofacies are obtained. The calculation formula is as follows:
[0120]
[0121] Among them, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B. The value range of the PCC coefficient is from -1 to 1, where 0 indicates no correlation between the two, a positive number indicates positive correlation, and a negative number indicates negative correlation. The greater the absolute value of the PCC coefficient, the stronger the correlation.
[0122] Build an initial classification model. Use the CART decision tree as the base classifier. It constructs the decision tree by randomly sampling the sample set with replacement and randomly sampling the feature set without replacement. The finally generated multiple decision trees form a random forest model, and the final prediction result of the model is determined by the voting of the base classifiers. The CART algorithm is one of the commonly used decision tree generation algorithms. The CART decision tree is generated from top to bottom. That is, first, a certain feature is taken from the feature set of the root node, and the data set is divided into 2 sub-nodes according to this feature. Each sub-node contains a certain number of samples. Then, the same process is performed on each sub-node respectively until the condition for the decision tree to stop growing is reached to complete the construction of the decision tree.
[0123] Use the 5-fold cross-validation method to select appropriate parameters for the model, including the number of iterations, the minimum number of samples in the leaf node, etc. Randomly divide the training set into 5 subsets of the same size. Use 4 of the subsets to train the model, and the remaining 1 subset is used to verify the model. Repeat this process 5 times. The cross-validation score of the model is obtained by averaging the 5 validation scores.
[0124] Use precision, recall, and F1-score to evaluate the quality of the classification model. The formulas are as follows:
[0125]
[0126]
[0127]
[0128] Among them, true positive refers to the positive sample predicted as positive, false positive refers to the negative sample predicted as positive, true negative refers to the negative sample predicted as negative, and false negative refers to the positive sample predicted as negative.
[0129] After completing the parameter tuning of the model, use the test set to verify the prediction effect of the established lithofacies identification model. After the verification is completed, predict the lithofacies according to the verified lithofacies identification model.
[0130] Example 3
[0131] For the logging data of continental shale in a certain area, use the logging attributes of the known well A and the random forest algorithm to identify the vertical lithofacies distribution of the shale in the blind well B. The prediction results are in good agreement with the measured lithofacies data, and the accuracy rate reaches more than 92%.
[0132] Based on shale geochemical data and logging data, the core types in the target area are classified into 6 main categories, including high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, carbon-poor mudstone facies, and other categories. They are marked with numbers 1 - 5, 10 corresponding to different lithofacies types. Then, the lithology digital tags are added after the logging data at each depth in the target layer, and the marked logging data sample set is obtained, as shown in Table 1.
[0133] Table 1
[0134] Lithofacies label Lithofacies 1 High-carbon silt-laminated clay shale facies 2 Medium-carbon silt-laminated clay shale facies 3 Carbonaceous mudstone facies 4 Silt-laminated interbedded mixed shale facies 5 Carbon-poor mudstone facies 10 Others
[0135] Select the following 6 conventional logging parameters as sample attribute values, namely natural gamma (GR), potassium (K), caliper logging (CAL), acoustic time difference (AC), compensated neutron (CNL), and density (DEN).
[0136] Different logging data have different dimensions and orders of magnitude of attribute values. If the logging data are directly used as input to train the lithology identification model, their influence on the results will be different. To eliminate this systematic error, it is necessary to standardize the data, and all attribute values of the logging data sample set are scaled to the range (-1, 1). The calculation process is as follows:
[0137]
[0138] Among them, χ is the sample data, μ is the sample data mean, and σ is the sample data standard deviation.
[0139] When using multiple logging attributes for lithofacies identification, it is necessary to comprehensively consider the correlation between different logging attributes and between logging attributes and lithofacies. The Pearson (PCC) coefficient is used for logging attribute analysis. By calculating the PCC coefficients between various logging attributes and the target lithofacies and between different logging attributes, the correlation degree between logging attributes and between logging attributes and the target lithofacies is obtained. The calculation formula is as follows:
[0140]
[0141] Among them, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B. The value range of the PCC coefficient is from -1 to 1, where 0 represents no correlation between the two, a positive number represents positive correlation, and a negative number represents negative correlation. The greater the absolute value of the PCC coefficient, the stronger the correlation.
[0142] Use the 5-fold cross-validation method to select appropriate parameters for the model, including the number of iterations, the minimum number of samples in leaf nodes, etc. Randomly divide the training set into 5 subsets of the same size. Use 4 of these subsets to train the model, and use the remaining 1 subset to validate the model. Repeat this process 5 times. The cross-validation score of the model is obtained by averaging the 5 validation scores.
[0143] Adopt precision, recall, and F1-score to evaluate the quality of the classification model. The formulas are as follows:
[0144]
[0145]
[0146]
[0147] Among them, true positives refer to positive samples predicted as positive, false positives refer to negative samples predicted as positive, true negatives refer to negative samples predicted as negative, and false negatives refer to positive samples predicted as negative.
[0148] After completing the parameter tuning of the model, use the test set to verify the prediction effect of the established lithofacies identification model. After the verification is completed, predict the lithofacies according to the verified lithofacies identification model.
[0149] Figure 2 Shows a schematic diagram of the correlation between different logging attributes according to an embodiment of the present invention. In the figure, the correlation between CNL and AC is the highest, and the absolute value of PCC is 0.8729, indicating that CNL and AC provide a large amount of similar information in lithofacies identification.
[0150] Figure 3 Shows a schematic diagram of the correlation analysis between different logging attributes and lithofacies according to an embodiment of the present invention. It can be seen that the correlation between GR and lithofacies is the strongest, and the absolute value of PCC is 0.63; the correlations between AC, CNL, K, and DEN and lithofacies are secondary, and the absolute value of PCC is higher than 0.2; the correlation between CAL and lithofacies is the worst, and the absolute value of PCC is 0.05.
[0151] Figure 4 Shows a schematic diagram of the relationship between the model error and the number of machine learning iterations according to an embodiment of the present invention. When the number of decision trees exceeds 150, the model error no longer decreases significantly. Therefore, the number of decision trees is determined to be 150.
[0152] Figure 5The figure shows a schematic diagram of the confusion matrix of different lithofacies recognition models according to an embodiment of the present invention. Based on the confusion matrix, the precision rate, recall rate, and F1-score can be calculated to evaluate the classification model. The calculation results are 98.81%, 98.81%, and 97.65% respectively, indicating that the lithofacies recognition effect of the random forest algorithm is strong.
[0153] Figure 6a and Figure 6b respectively show a schematic diagram of the comparison between the model prediction result and the true lithofacies result according to an embodiment of the present invention. The coincidence degree between the two is very high, and the accuracy rate reaches more than 92%.
[0154] Example 4
[0155] Figure 7 The figure shows a block diagram of a lithofacies prediction device based on machine learning according to an embodiment of the present invention.
[0156] As Figure 7 shown, the lithofacies prediction device based on machine learning includes:
[0157] A classification module 201 that classifies the core types in the target area to obtain a logging data sample set;
[0158] A normalization module 202 that determines the sample attribute values and normalizes all the attribute values of the logging data sample set;
[0159] A correlation calculation module 203 that analyzes the logging attributes using the Pearson coefficient to obtain the correlations between logging attributes and between logging attributes and the target lithofacies;
[0160] A training module 204 that establishes an initial classification model and trains the initial classification model according to the logging data sample set;
[0161] An evaluation module 205 that evaluates the trained classification model through the precision rate, recall rate, and F1-score, adjusts the model parameters, and outputs the final model.
[0162] As an optional solution, the core types in the target area include high-carbon silt laminated clay shale facies, medium-carbon silt laminated clay shale facies, carbonaceous mudstone facies, silt laminated interbedded mixed shale facies, and low-carbon mudstone facies.
[0163] As an optional solution, the sample attribute values include natural gamma, potassium, well diameter logging, acoustic time difference, compensated neutron, and density.
[0164] As an optional solution, the correlation is:
[0165]
[0166] Among them, PCC is the Pearson coefficient, that is, the correlation. Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B.
[0167] As an alternative, the precision rate is:
[0168]
[0169] Among them, true positives refer to positive samples predicted as positive, and false positives refer to negative samples predicted as positive.
[0170] As an alternative, the recall rate is:
[0171]
[0172] Among them, true positives refer to positive samples predicted as positive, and false negatives refer to positive samples predicted as negative.
[0173] As an alternative, the F1-score is:
[0174]
[0175] Example 5
[0176] This embodiment provides an electronic device, which includes: a memory storing executable instructions; and a processor that runs the executable instructions in the memory to implement the above-mentioned lithofacies prediction method based on machine learning.
[0177] The electronic device according to the embodiment of the present disclosure includes a memory and a processor.
[0178] The memory is used to store non-temporary computer-readable instructions. Specifically, the memory may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0179] The processor may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions. In an embodiment of the present disclosure, the processor is used to run the computer-readable instructions stored in the memory.
[0180] Those skilled in the art should understand that, in order to solve the technical problem of how to obtain good user experience effects, the present embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included in the protection scope of the present disclosure.
[0181] For the detailed description of this embodiment, reference may be made to the corresponding descriptions in the foregoing embodiments, and details will not be repeated here.
[0182] Example 6
[0183] The present embodiment provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the described lithofacies prediction method based on machine learning is implemented.
[0184] According to the computer-readable storage medium of the embodiments of the present disclosure, non-temporary computer-readable instructions are stored thereon. When the non-temporary computer-readable instructions are run by a processor, all or part of the steps of the methods of the foregoing embodiments of the present disclosure are executed.
[0185] The above computer-readable storage medium includes but is not limited to: optical storage media (such as CD-ROMs and DVDs), magneto-optical storage media (such as MOs), magnetic storage media (such as magnetic tapes or external hard drives), media with built-in rewritable non-volatile memories (such as memory cards), and media with built-in ROMs (such as ROM cartridges).
[0186] Those skilled in the art should understand that the purpose of the above description of the embodiments of the present invention is only to exemplarily illustrate the beneficial effects of the embodiments of the present invention, and is not intended to limit the embodiments of the present invention to any of the examples given.
[0187] The foregoing embodiments of the present invention have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
Claims
1. A machine learning-based lithofacies prediction method, characterized in that, it includes: Classify the core types in the target area to obtain a logging data sample set; Determine the sample attribute values and normalize all the attribute values of the logging data sample set; Conduct logging attribute analysis using the Pearson coefficient to obtain the correlations between logging attributes and between logging attributes and the target lithofacies; Establish an initial classification model and train the initial classification model according to the logging data sample set; Evaluate the trained classification model through precision, recall, and F1-score, adjust the model parameters, and output the final model.
2. The machine learning-based lithofacies prediction method according to claim 1, wherein, the core types in the target area include high-carbon silt-laminated clay shale facies, medium-carbon silt-laminated clay shale facies, carbonaceous mudstone facies, silt-laminated interbedded mixed shale facies, and low-carbon mudstone facies.
3. The machine learning-based lithofacies prediction method according to claim 1, wherein, the sample attribute values include natural gamma, potassium, well diameter logging, acoustic travel time, compensated neutron, and density.
4. The machine learning-based lithofacies prediction method according to claim 1, wherein, the correlation is: where PCC is the Pearson coefficient, i.e., the correlation, Cov(A,B) is the covariance of A and B, D(A) is the standard deviation of A, and D(B) is the standard deviation of B.
5. The machine learning-based lithofacies prediction method according to claim 1, wherein, the precision is: where true positive refers to a positive sample predicted as positive, and false positive refers to a negative sample predicted as positive.
6. The machine learning-based lithofacies prediction method according to claim 1, wherein, the recall is: where true positive refers to a positive sample predicted as positive, and false negative refers to a positive sample predicted as negative.
7. The machine learning-based lithofacies prediction method according to claim 1, wherein, the F1-score is:
8. A machine learning-based lithofacies prediction device, characterized in that, it includes: A classification module that classifies the core types in the target area to obtain a logging data sample set; A normalization module that determines the sample attribute values and normalizes all the attribute values of the logging data sample set; A correlation calculation module that conducts logging attribute analysis using the Pearson coefficient to obtain the correlations between logging attributes and between logging attributes and the target lithofacies; A training module that establishes an initial classification model and trains the initial classification model according to the logging data sample set; An evaluation module that evaluates the trained classification model through precision, recall, and F1-score, adjusts the model parameters, and outputs the final model.
9. An electronic device, characterized in that, the electronic device includes: A memory that stores executable instructions; A processor that runs the executable instructions in the memory to implement the machine learning-based lithofacies prediction method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program which, when executed by a processor, implements the machine learning-based lithofacies prediction method according to any one of claims 1-7.
Citation Information
Cited By
Construction method and application of high-precision prediction model for full-well-section marine shale lithofacies
CN120744865A
Lithofacies combination identification method, apparatus and device, and storage medium
CN121051611A