Automatic annotation method and system for diagenetic facies samples integrating affinity propagation clustering and graph convolutional neural network
By integrating the method of affinity propagation clustering and graph convolutional neural network, the problem of small sample size and low accuracy in diagenetic phase sample annotation is solved, and the automatic diagenetic phase annotation effect with high accuracy and stability is achieved.
Patent Information
- Application Number
- CN202411864247.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The problem of small sample size and low accuracy in the labeling of diagenetic samples.
The diagenetic annotation model is optimized through data preprocessing, APM graph structure and multi-layer graph convolution network operations.
The accuracy and stability of diagenetic samples were improved, with Precision being above 86%, Recall being above 90%, and F1 score being above 89%, which performed better than traditional methods.
Smart Images

Figure CN119339163B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent identification of diagenetic facies. Specifically, it relates to an automatic annotation method and system for diagenetic facies samples that integrates affinity propagation clustering and graph convolutional neural network. Background Technique
[0002] The annotation of diagenetic facies samples is a key link in the intelligent identification of diagenetic facies. Traditional diagenetic facies annotation methods mostly rely on the experience of geological experts [1-3] , and it is necessary to analyze well logging curves and make manual judgments, which has problems of large data volume, time-consuming and laborious; in addition, due to the data quality problems of well logging curves themselves and manual judgment errors, it will also affect the accuracy of the final diagenetic facies identification. Therefore, carrying out the automatic annotation work of diagenetic facies samples has important practical significance for improving the speed and efficiency of reservoir evaluation [4,5] .
[0003] Traditional automatic annotation techniques can be classified into three types: semi-supervised learning, self-supervised learning, and deep learning. Semi-supervised learning trains a model through some labeled data, and then preliminarily annotates the unlabeled data, and improves the model accuracy through alternating training and correction [6] . Although such methods reduce the dependence on fully labeled data, the effect is limited by the model and data quality. Due to the relatively weak data quality or certain deviations of well logging curve data, it is difficult to play its advantages in the automatic annotation of diagenetic facies samples. Self-supervised learning is to learn through the internal structure in the data, which can greatly reduce the dependence on manual annotation [7] , but the training process is complex and requires a large amount of data sets. Due to the complex data characteristics and insufficient data volume of diagenetic facies samples, it is difficult to be effectively applied in the automatic annotation of diagenetic facies samples.
[0004] With the rise of deep learning technology, automatic annotation technology has also been significantly improved. Usually, models such as convolutional neural network and recurrent neural network are used to identify and classify time series data or pictures, so as to automatically generate labels [8,9] . Such methods can effectively identify complex features and have a relatively fast recognition speed, but they still face some challenges in the field of automatic identification of diagenetic facies samples: one is that the training process of deep learning usually requires a large number of samples, and due to cost factors, it is difficult to support the number of diagenetic facies samples; the other is that due to the complexity and diversity of the geological environment, it is difficult for the model to comprehensively capture all features, resulting in its accuracy being less than satisfactory.
[0005] In recent years, graph convolutional neural network
[10] has been used to solve the problem of sample annotation of graph-structured data [11-13], especially in the few-label learning scenario, it can also maintain high learning efficiency and annotation accuracy. The formation process of diagenetic facies is affected by sedimentation and various factors, and there are certain distribution characteristics in the spatial stratigraphic structure. Therefore, a graph convolutional neural network can be used to construct a graph structure to capture the correlation between different depths, thereby improving the accuracy of annotation.
[0006] Meanwhile, to further overcome the few-sample problem and optimize the graph structure, the Affinity Propagation (AP) algorithm is introduced. The AP algorithm works by means of "message passing", dynamically adjusting the connections according to the similarity between nodes to determine the most suitable clustering centers, rather than relying on fixed rules. [14-15] . Ordinary graph structures are usually constructed based on spatial positions or depth sequences, and the connections between nodes are based on fixed rules. However, the AP algorithm can more effectively cluster nodes with similar logging response characteristics, and the generated graph structure better reflects the relationship between different diagenetic facies and the sedimentary similarity characteristics of the strata.
[16] , thus effectively dealing with the problem of a small sample size. Summary of the Invention
[0007] The technical problem to be solved by the present invention is:
[0008] To solve the problems of a small sample size and low accuracy in the annotation work of diagenetic facies.
[0009] The present invention provides an automatic annotation method for diagenetic facies samples integrating affinity propagation clustering and graph convolutional neural network, including the following steps:
[0010] S100. Establish a data set, including dividing diagenetic facies types and preprocessing logging curve data. The logging curve data includes logging curve selection, processing missing values and outliers of the selected logging curves, and sample annotation, and then annotating the logging curve data according to diagenetic facies types.
[0011] S200. Use the affinity propagation clustering algorithm to construct an APM graph structure, connect nodes with similar features but not adjacent, for realizing diagenetic facies label sharing.
[0012] S300. Adopt a multi-layer graph convolutional network, combine the affinity propagation clustering results obtained in step S200, perform convolution operations on the logging curve data processed in step S100, and after multiple iterations, obtain an optimized model structure based on the graph convolutional diagenetic facies annotation method.
[0013] Further, in step S100, it specifically includes:
[0014] S110. Induce diagenetic facies types, classify the diagenetic facies types based on compaction, lithology, cementation and dissolution conditions, then divide the reservoirs according to reservoir properties, and correspondingly divide the diagenetic facies types according to reservoir types;
[0015] S120. Preprocess the logging curve data, including logging curve selection, logging curve preprocessing and sample annotation:
[0016] S121. Logging curve selection: Analyze the logging response characteristics of various diagenetic facies, score the importance of each logging curve using the decision tree scoring method, and select the logging curve types that have a great impact on diagenetic facies analysis;
[0017] S122. Preprocess the logging curve data obtained in step S121, including missing value and outlier processing;
[0018] Use the regression imputation method to predict the missing values of the logging curve data. For missing value processing, use formula (1) to construct a linear regression model, use formula (2) to calculate the regression coefficients during imputation, then obtain the intercept through formula (3), and use the linear regression model to predict the missing values;
[0019] Use the linear regression model to detect outliers. For outlier detection, analyze the residuals through formula (4) to identify outliers. If the residual values of data points are too large, judge these points as outliers;
[0020] The formulas are as follows:
[0021] (1)
[0022] (2)
[0023] (3)
[0024] (4)
[0025] (5)
[0026] Among them, is the response variable, is the input variable, is the logging characteristic value of the i-th logging depth node, is the response variable value of the i-th logging depth node, is the number of logging depth nodes, is the error term, is the regression coefficient, is the intercept, is the residual, is the number of the logging curve depth node, The output obtained after prediction;
[0027] S123. Sample annotation: Normalize the well logging curve data preprocessed in step S122 to make different well logging curve data within the same measurement range; analyze and annotate according to the thin section images and scanning electron microscope images at the corresponding depths to form a dataset for automatic annotation of diagenetic facies samples.
[0028] Furthermore, classify the diagenetic facies types based on compaction, lithology, cementation, and dissolution conditions, including weakly compacted weakly cemented dissolved siltstone facies, weakly compacted weakly cemented dissolved fine sandstone facies, moderately strongly compacted weakly cemented dissolved fine sandstone facies, moderately strongly compacted weakly cemented dissolved siltstone facies, moderately strongly compacted weakly dissolved cemented fine sandstone facies, and mudstone facies; classify the diagenetic facies types according to reservoir properties, including Class I, Class II, and Class III reservoirs. Class I reservoirs correspond to weakly compacted weakly cemented dissolved siltstone facies, weakly compacted weakly cemented dissolved fine sandstone facies, and moderately strongly compacted weakly cemented dissolved fine sandstone facies; Class II reservoirs correspond to moderately strongly compacted weakly cemented dissolved siltstone facies and moderately strongly compacted weakly dissolved cemented fine sandstone facies; Class III reservoirs correspond to mudstone facies.
[0029] Furthermore, select well logging curves that have a great influence on diagenetic facies analysis, including natural gamma, acoustic time difference, deep lateral resistivity, shallow lateral resistivity, spontaneous potential, and well diameter.
[0030] Furthermore, in step S200, it specifically includes:
[0031] S210. Read the well logging curve data, select the well logging curve features according to step S121, form the feature vector of each node, and generate a feature matrix as shown in the following formula (6):
[0032] (6)
[0033] Where, , is the well logging curve depth node, is the number of features;
[0034] S220. Sort the well logging curve depth nodes according to depth, connect adjacent depth nodes with edges, construct an empty undirected graph structure based on the depth order, and extract the corresponding feature vectors;
[0035] S230. Use the affinity propagation clustering algorithm to cluster non-adjacent depth nodes, select the well logging curve features according to step S121, classify each well logging curve depth node into the corresponding cluster, establish connections between nodes with similar well logging curve features, and then calculate the Euclidean distance between nodes and perform normalization processing to make each feature value within the same range:
[0036] (7)
[0037] Similarity matrix The negative distance matrix is:
[0038] (8)
[0039] (9)
[0040] S240. Generate an adjacency matrix according to the diagenetic facies labels of clustering. If two nodes belong to the same cluster, the corresponding position in the adjacency matrix is 1, otherwise it is 0:
[0041] (10)
[0042] Among them, represents the diagenetic facies label of the log curve depth node at the th;
[0043] Generate the adjacency matrix according to the above steps Feature matrix .
[0044] Furthermore, in step S300, the first layer of convolution uses the GCNConv operation, input dimension: num_features, output dimension: 64; the second layer of convolution uses the GCNConv operation, input dimension: 64, output dimension: 32; the third layer of convolution uses the GCNConv operation, input dimension: 32, output dimension: 6;
[0045] S310. Input the feature matrix and the adjacency matrix . The specific convolution operation is as follows:
[0046] (11)
[0047] (12)
[0048] (13)
[0049] Among them, is the identity matrix, is the degree of node i, is the feature matrix of the th layer, initially the input feature matrix , is the weight matrix of the th layer, is the non-linear activation function;
[0050] S320. The output layer uses the softmax function to classify the diagenetic facies:
[0051] (14)
[0052] (15)
[0053] Among them, Z outputs the category probability distribution;
[0054] S330. Update the model parameters using the Adam optimizer:
[0055] (16)
[0056] Among them, are the model parameters, is the learning rate, is the parameter gradient operation of, is the loss function.
[0057] Furthermore, set the number of iterations to 2000 times.
[0058] A diagenetic facies sample automatic annotation system integrating affinity propagation clustering and graph convolutional neural network according to the present invention, the system has program modules corresponding to the above steps, and executes the steps in the above diagenetic facies sample automatic annotation method integrating affinity propagation clustering and graph convolutional neural network when running.
[0059] A computer-readable storage medium according to the present invention, the computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the diagenetic facies sample automatic annotation method integrating affinity propagation clustering and graph convolutional neural network when called by a processor.
[0060] Compared with the prior art, the beneficial effects of the present invention are:
[0061] (1) Using the graph neural network for diagenetic facies sample annotation and combining the graph structure constructed by affinity propagation clustering can capture the complex associations and spatial distribution characteristics between diagenetic facies, and combine the graph structure with well logging curve characteristics, thereby improving the annotation accuracy.
[0062] (2) The graph structure constructed based on the depth order and affinity propagation clustering method performs best in diagenetic facies sample annotation. Compared with the graph without constructed edges and the graph structure only based on the depth order, affinity propagation clustering can more accurately capture the similarity between diagenetic facies, can better express the relationship between nodes, thereby generating a more reasonable graph structure to improve the accuracy and stability of annotation.
[0063] (3) The Precision of the APM-GCN method of the present invention is above 86%, the Recall is above 90%, and the F1 score is above 89%. The diagenetic facies annotation effect is better than that of traditional GCN, GraphSAGE, SVM, and RF methods. APM-GCN first preliminarily classifies the data through clustering technology and then effectively transmits and learns features using graph convolution, improving the annotation accuracy. However, the annotation accuracy is easily affected by the similarity of logging curve features.
[0064] (4) After adding different Gaussian noises, the standard deviations of mAA and mAP of APM-GCN are low and the fluctuations are small, indicating a high stability of the model. Description of the Drawings
[0065] Figure 1 is a flowchart of an automatic diagenetic facies sample annotation method integrating affinity propagation clustering and graph convolutional neural network in an embodiment of the present invention;
[0066] Figure 2 is a decision tree scoring diagram of logging curves in an embodiment of the present invention;
[0067] Figure 3 is a partial thin section casting diagram and scanning electron microscope diagram of the Fuyu oil layer in the Sanzhao Sag in an embodiment of the present invention. Among them, (a), (b), and (c) are all partial thin section casting diagrams of the Fuyu oil layer in the Sanzhao Sag, and (d), (e), and (f) are all partial scanning electron microscope diagrams of the Fuyu oil layer in the Sanzhao Sag;
[0068] Figure 4 is a structural diagram of the APM network model in an embodiment of the present invention;
[0069] Figure 5 is a structural diagram of the graph convolutional network model in an embodiment of the present invention;
[0070] Figure 6 is a loss iteration result diagram during the model training process in an embodiment of the present invention;
[0071] Figure 7 is a comparison diagram of Matthews coefficient, blind well accuracy, and test accuracy of the graph construction method in an embodiment of the present invention;
[0072] Figure 8 is a line graph of the accuracy of k-means clustering method, Gaussian mixture clustering method, agglomerative clustering, and affinity propagation clustering method in an embodiment of the present invention;
[0073] Figure 9This is the confusion matrix result graph in the embodiments of the present invention. Among them, (a) is the confusion matrix result graph of the APM-GCN method of the present invention, (b) is the confusion matrix result graph of the GCN method, (c) is the confusion matrix result graph of the GraphSAGE method, (d) is the confusion matrix result graph of the RF method, and (e) is the confusion matrix result graph of the SVM method;
[0074] Figure 10 This is the calculation result graph of Precision, Recall, and F1 scores of the APM-GCN method, GCN method, GraphSAGE method, RF method, and SVM method in the embodiments of the present invention. Among them, (a) is the line graph of the Precision calculation results of the above five methods, (b) is the line graph of the Recall calculation results of the above five methods, and (c) is the line graph of the F1 score calculation results of the above five methods;
[0075] Figure 11 This is the relationship graph between different data volumes and mAA and mAP in the embodiments of the present invention;
[0076] Figure 12 This is the comparison graph of the calculation results of mAA and mAP for the annotation effects of various diagenetic facies with different data volumes in the embodiments of the present invention. Among them, (a) is the calculation result graph of the Accuracy value and mAA value in the annotation effects of various diagenetic facies, and (b) is the calculation result graph of the AP' value and mAP value in the annotation effects of various diagenetic facies;
[0077] Figure 13 This is the calculation result graph of mAA and mAP after adding Gaussian noise in the embodiments of the present invention. Detailed implementation manners
[0078] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.
[0079] Combined with Figure 1As shown in the figure, the present invention proposes an automatic labeling method for petrofacies samples that combines affinity propagation clustering and graph convolutional neural network (An Automatic Labeling Method for Petrofacies Samples Using Affinity Propagation Clustering and Graph Convolutional Neural Network, APM-GCN). First, the petrofacies types are classified and key logging curves are selected for data preprocessing, and the logging curve data is labeled with different petrofacies types; secondly, the AP algorithm is used to construct a graph structure to connect nodes with similar features but non-adjacent, realizing the sharing of petrofacies labels; then, a multi-layer graph convolutional network is adopted, combined with the AP clustering results, to perform convolution operations on the processed logging curve data, improving the classification accuracy and efficiency of automatic petrofacies labeling; finally, through comparative experiments, the effectiveness of the method proposed by the present invention is verified.
[0080] Specific implementation plan 1: Combine Figures 1 to 6 As shown in the figure, the present invention provides an automatic labeling method for petrofacies samples that combines affinity propagation clustering and graph convolutional neural network, including the following steps:
[0081] S100. Establish a data set, and the establishment of the data set includes two steps: inducing petrofacies types and preprocessing logging curve data.
[0082] S110. Induce petrofacies types, taking the Fuyu oil layer in the Sanzhao sag of the Songliao Basin as the target area, and this reservoir has the characteristics of fine grain size and high shale content.
[17] . Divide the petrofacies types according to compaction, lithology, cementation and dissolution conditions. [18,19], including: Weakly compacted weakly cemented dissolved siltstone phase (Wip), Weakly compacted weakly cemented dissolved fine sandstone phase (Wap), Medium to strong compaction of weakly cemented dissolved fine sandstone phase (Map), Medium to strong compaction of weakly cemented dissolved siltstone phase (Mip), Medium to strong compaction of weakly dissolved colluvial fine sandstone phase (Msap), and Mudstone phase (Mp). According to the reservoir properties, the above diagenetic facies types are correspondingly divided into Class I, Class II, and Class III reservoirs. Class I reservoirs correspond to Wip, Wap, and Map; Class II reservoirs correspond to Mip and Msap; Class III reservoirs correspond to Mp.
[0083] S120. Pretreatment of logging curve data, including three steps: logging curve selection, logging curve pretreatment, and sample annotation.
[0084] S121. Logging curve selection: Analyze the logging response characteristics of various diagenetic facies
[20] , and use the decision tree scoring method to score the importance of each logging curve. As Figure 2 shown, the logging curves with relatively high scores for the influence on diagenetic facies analysis characteristics (Feature Importances) are selected, including natural gamma ray (GR), acoustic time difference (AC), deep lateral resistivity (LLD), shallow lateral resistivity (LLS), spontaneous potential (SP), and caliper (CAL). [21,22] ;
[0085] S122. Pretreat the logging curve data obtained in step S121, including the processing of missing values and outliers.
[0086] The regression imputation method is used to predict the missing values of well logging curve data, and then the linear regression model is used to detect outliers, and the missing values and outliers are processed. For the processing of missing values, a linear regression model is constructed using formula (1). When performing imputation, formula (2) is used to calculate the regression coefficients, and then the intercept is obtained through formula (3). The linear regression model is used to predict the missing values. For outlier detection, formula (4) is used to analyze the residuals to identify outliers. If the residual values of some data points are too large, that is, the residual values are greater than 3.5, these points are judged as outliers. The formulas are as follows:
[0087] (1)
[0088] (2)
[0089] (3)
[0090] (4)
[0091] (5)
[0092] Among them, is the response variable, is the input variable, is the well logging eigenvalue of the i-th well logging depth node, is the response variable value of the i-th well logging depth node, is the number of well logging depth nodes, is the error term, is the regression coefficient, is the intercept, is the residual, is the number of the well logging curve depth node, the output obtained after prediction;
[0093] S123, sample annotation. The well logging curve data preprocessed in step S122 is normalized to ensure that different well logging curve data are within the same measurement range, and the node depth interval is set to 0.05 meters. Finally, detailed analysis and annotation are carried out according to the thin section images and scanning electron microscope images of the corresponding depth (as Figure 3 shown) to form a dataset for automatic annotation of diagenetic facies samples;
[0094] S200, constructing an APM graph structure using the AP algorithm to connect nodes with similar features but non-adjacent, realizing the sharing of diagenetic facies labels; The design of the APM graph structure can not only form connections between well logging curve depth nodes with similar well logging curve features but non-adjacent depths, but also ensure that these well logging curve depth nodes share the same diagenetic facies label. The APM network model is as Figure 4As shown in the figure, the construction process of the APM diagram structure specifically includes the following four steps:
[0095] S210. Read the logging curve data. According to the six features of 'GR', 'AC', 'LLD', 'LLS','SP', and 'CAL' obtained in step S121, form the feature vector of each node and generate a feature matrix, as shown in the following formula 6:
[0096] (6)
[0097] Among them, , is the logging curve depth node, is the number of features;
[0098] S220. Sort the logging curve depth nodes according to the depth (DEPT). Due to the continuity of the sedimentation process, each depth node at a certain depth in the formation has diagenetic facies similarity with other depth nodes adjacent in space. Here, a certain depth refers to the depth range of a specific formation thickness or interval that shows the sedimentation process. Therefore, adjacent depth nodes are connected by edges to construct an empty undirected graph structure based on the depth order and extract the corresponding feature vectors;
[0099] S230. Use the AP clustering algorithm to cluster the nodes with non-adjacent depths. According to the six features of 'GR', 'AC', 'LLD', 'LLS','SP', and 'CAL', classify each logging curve depth node into the corresponding cluster and establish the connection between the nodes with similar logging curve features. Calculate the Euclidean distance between the nodes and perform normalization to ensure that each feature value is within the same range.
[0100] (7)
[0101] The AP algorithm clusters through the similarity matrix and will output the cluster label of each logging curve depth node. Therefore, the similarity matrix takes the negative of the distance matrix:
[0102] (8)
[0103] (9)
[0104] S240. Generate an adjacency matrix according to the diagenetic facies labels of the clusters. If two nodes belong to the same cluster, the corresponding position in the adjacency matrix is 1, otherwise it is 0:
[0105] (10)
[0106] Among them, represents the Diagenetic facies labels for logging curve depth nodes;
[0107] Generate the adjacency matrix according to the above steps Feature matrix , laying the foundation for the subsequent automatic annotation task of diagenetic facies samples.
[0108] S300. Adopt a multi-layer graph convolutional network, combine the AP clustering results, and perform convolutional operations on the processed logging curve data; The model structure of the graph convolutional diagenetic facies annotation method is as Figure 5 shown. The first layer of convolution uses the GCNConv operation, with an input dimension of num_features and an output dimension of 64; The second layer of convolution uses the GCNConv operation, with an input dimension of 64 and an output dimension of 32; The third layer of convolution uses the GCNConv operation, with an input dimension of 32 and an output dimension of 16. Specifically, it includes:
[0109] S310. Input the feature matrix and the adjacency matrix . The specific convolutional operation is as follows:
[0110] (11)
[0111] (12)
[0112] (13)
[0113] Among them, is the identity matrix, is the degree of node i, is the feature matrix of the th layer, initially the input feature matrix , is the weight matrix of the th layer, is the non-linear activation function (ReLU).
[0114] S320. The output layer uses the softmax function to classify the diagenetic facies, and Z outputs the class probability distribution:
[0115] (14)
[0116] (15)
[0117] S330. Use the Adam optimizer to update the model parameters:
[0118] (16)
[0119] Among them, is the model parameter, is the learning rate, is the parameter gradient operation, is the loss function.
[0120] The loss iteration during the model training process is as Figure 6 shown. The number of iterations is set to 2000 times. It can be observed that after about 500 iterations, the error of the model gradually approaches zero, reaching the preset training accuracy, indicating that the model can well fit the training data after sufficient training.
[0121] Specific implementation plan two: An automatic lithofacies sample annotation system integrating affinity propagation clustering and graph convolutional neural network according to the present invention. This system has program modules corresponding to the above steps and executes the steps in the above method for automatically annotating lithofacies samples integrating affinity propagation clustering and graph convolutional neural network when running.
[0122] Other combinations and connection relationships in this implementation plan are the same as those in the first specific implementation plan.
[0123] Specific implementation plan three: A computer-readable storage medium according to the present invention. The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the method for automatically annotating lithofacies samples integrating affinity propagation clustering and graph convolutional neural network when called by a processor.
[0124] Other combinations and connection relationships in this implementation plan are the same as those in the first specific implementation plan.
[0125] Simulation experiment
[0126] To verify the effectiveness of the proposed method, the present invention conducts graph structure comparison evaluation experiments, clustering method comparison evaluation experiments, model performance evaluation experiments, and model stability evaluation experiments. Specific parameters of the experimental equipment: CPU is Intel Xeon Silver 4210R, memory is 64G, GPU is RTX 6000; operating system is Ubuntu 20.04.3; experimental framework is PyTorch 1.8.0.
[0127] 1. Graph structure comparison evaluation experiment
[0128] To compare the performance of graph structures in the lithofacies sample annotation task, the present invention uses the Matthews correlation coefficient (MCC) and the precision of lithofacies annotation results as the main evaluation indicators.
[0129] MCC: A commonly used indicator for evaluating the performance of classification models. The calculation formula is:
[0130] (17)
[0131] Precision: The proportion of correctly labeled diagenetic facies types, and the calculation formula is:
[0132] (18)
[0133] Among them, TP represents the number of samples correctly labeled as the positive class, TN represents the number of samples correctly labeled as the negative class, FP represents the number of samples mislabeled as the positive class, and FN represents the number of samples mislabeled as the negative class.
[0134] During the experiment, 8000 logging curve depth nodes were used, and three graph structure construction methods were designed to analyze their effects on the diagenetic facies labeling results: the first is no edge construction (No edge construction, NE method), directly using the original logging curves as input data; the second is a graph structure based on the depth sequence (A graph structure based on deep depth sequence, DS method); the third is a graph structure based on the depth sequence and clustering method (A graph structure based on depth sequence and clustering, SC method). The same training data and test data were used during the experiment, and the MCC and Precision of different graph construction methods were calculated. To further test the generalization ability of the three graph structures, 2101 logging curve depth nodes of Well F361 were used for blind well testing. The logging curve data of F361 were used for graph construction but did not participate in the model training. Finally, the precision of the diagenetic facies type labeling results was calculated.
[0135] 2. Comparative evaluation experiment of clustering methods
[0136] The comparative experiment of clustering methods used the calculated precision (Precision) of diagenetic facies labeling as the evaluation index. Among them, the meaning and calculation method of Precision are the same as those in Experiment 1. Based on Experiment 1, the graph construction method with the best labeling effect was selected during the experiment. The same training data and test data were used in this experiment, and the selection of 4 different clustering methods was compared, namely k-means clustering, Gaussian mixture clustering, agglomerative clustering, and affinity propagation clustering. The precision of diagenetic facies labeling was compared to evaluate the influence of clustering methods on the labeling effect.
[0137] 3. Model performance evaluation test
[0138] The present invention established a confusion matrix and calculated the precision (Precision), recall (Recall), and F of the labeling results of each diagenetic facies type 1Fraction
[23] As an evaluation index, a model performance evaluation experiment is carried out.
[0139] Recall represents the proportion of the correct quantity in the actual components, which can be expressed by formula (19):
[0140] (19)
[0141] (20)
[0142] The experimental process is based on the construction method of Experiment 2 diagram to further verify the effects of the APM-GCN of the present invention and other deep learning algorithms in diagenetic facies annotation. First, the standardized (i.e., preprocessed) logging curve data is randomly divided according to the ratio of 6:4. Then, a graph structure is constructed using the depth sequence and AP clustering method, and the standardized logging curve data is used as the feature matrix. Finally, the trained graph convolutional neural network GCN, GraphSAGE (graph sampling and aggregation), SVM (support vector machine), and RF (random forest) are used to annotate diagenetic facies samples to evaluate the effect of the APM-GCN algorithm.
[0143] 4. Model Stability Evaluation Test
[0144] This experiment applies the accuracy (mean Average-Accuracy, mAA)
[24] and mAP (mean Average-Precision, mAP)
[25] As evaluation indexes.
[0145] mAA represents the average value of the recognition accuracies (Accuracy) of multiple diagenetic facies types. Accuracy represents the proportion of the correctly recognized quantity in the total quantity. The formula is as follows, where TP represents the quantity of diagenetic facies correctly recognized, and N represents the total quantity of diagenetic facies samples.
[0146] (21)
[0147] mAP represents the average value of the average precisions (Average-Precision, AP’) of multiple diagenetic facies types. AP’ represents the area under the curve plotted by Precision and Recall. Among them, the definitions of Precision and Recall are the same as those in Experiment 3 above.
[0148] The experimental process is divided into two parts:
[0149] (1) Influence of different data volumes on the annotation effect. Different numbers of datasets, namely 2000, 4000, 6000, 8000, and 10000 depth nodes, were selected from different wells in the same area. Calculate the Accuracy and AP’ of the APM-GCN algorithm in the diagenetic facies annotation, and calculate the mA value and mAP value.
[0150] (2) To further verify the stability of the model, Gaussian noise was added to the dataset with the best annotation effect. The well logging curve dataset with added noise was used for model training and evaluation. The same mAA and mAP metrics were used for performance evaluation. The changes in model performance before and after adding noise were compared, and the stability of the model was analyzed.
[0151] Experimental results
[0152] 1. Results of the graph structure comparison experiment
[0153] The Matthews correlation coefficient and accuracy of the graph construction method are as Figure 7 shown. When only the original data is input into the APM-GCN model of the present invention, the test data accuracy can reach 0.8. Although the correct labeling rate of the diagenetic facies samples is 10% lower than that of the input constructed graph structure, it is still higher than the labeling accuracies of SVM and RF, proving that using graph structure data can significantly improve the accuracy of diagenetic facies annotation. The DS method is a graph structure based on depth order, considering geological factors in the diagenetic facies formation process, and the test accuracy is 0.91. The SC method combines depth order and clustering methods, and the test accuracy is 0.9, slightly lower than DS. The main reason for this result is that the training and test data are randomly selected, so there is a high similarity between the training and test data.
[0154] From Figure 7 it can be seen that in the F361 well test, the labeling accuracies of the three graph structures are 0.62, 0.7, and 0.75 respectively. The results show that the accuracies of SC and DS are similar in the diagenetic facies label annotation task, but the generalization ability of the SC training model is significantly better than that of DS. Therefore, good accuracy and generalization ability are shown in the diagenetic facies label annotation based on the SC method.
[0155] 2. Results of the clustering method comparison experiment
[0156] The accuracies of the four clustering selection methods are as Figure 8As shown, affinity propagation clustering performs best in the accuracy of diagenetic facies annotation, reaching 0.90. In contrast, the annotation accuracies of Gaussian mixture clustering and agglomerative clustering are 0.87 and 0.86 respectively, and the accuracy of k-means clustering is the lowest, only 0.85. This indicates that although Gaussian mixture clustering and agglomerative clustering can better capture the potential structure in the data, they are still slightly inferior to the affinity propagation clustering method in processing well logging curve data. Although the k-means clustering method has great advantages in calculation speed and implementation, its effect is not as good as other methods. The calculation results show that in the test data, the model based on the affinity propagation clustering algorithm has a higher accuracy than other clustering algorithms.
[0157] 3. Experimental results of model performance evaluation
[0158] The confusion matrix results are as Figure 9 shown. The calculation results of Precision, Recall and F1 score are as Figure 10 shown. It can be seen from Figure 9 that the APM-GCN method has the best effect in diagenetic facies label annotation, followed by the GCN and GraphSAGE methods, and the SVM and RF methods perform the worst. This is because APM-GCN, GCN and GraphSAGE use graph structure data containing the relationships between samples, while SVM and RF only use tabular data of independent samples, which proves that using graph structure data can significantly improve the accuracy of diagenetic facies label annotation. Compared with the GCN method and the GraphSAGE method, the APM-GCN model performs better in the accuracy of diagenetic facies sample annotation and can better realize the diagenetic facies sample annotation work of tight reservoirs.
[0159] In addition, the experimental results show that the number of correctly classified Wip is the highest, but there are a small number of misclassifications into the Wap category. There are many misclassifications in the Wap category, misclassified as Wip and Mip. The Mip category is misclassified as Wap and Map. The Map category has good recognition performance, but there are some misclassifications into the Msap category. The Msap category has good recognition performance, but there are misclassifications into the Map category. The Mp category has poor recognition performance and a large number of misclassifications, especially misclassified as the Msap category. The reason for the misjudgment may be that the response characteristics of well logging curves are affected not only by diagenetic facies, but also by geology and drilling operations such as clay minerals, pore fluids, and mud pressure.
[0160] It can be seen from Figure 10It can be seen that the Precision, Recall, and F1-score values of the APM-GCN method are all higher than those of the GraphSAGE method, SVM method, and RF method, while the GCN method and GraphSAGE method are higher than the SVM method and RF method. The four methods used in the Mp category all have good effects, indicating that the logging response characteristics of this diagenetic facies are significantly different from those of other diagenetic facies. From Figure 9 (a), it can be seen that there are many misclassifications in the Wap category and the annotation accuracy rate is low. The main reason is that Wap and Msap have similar logging response characteristics, so there is a phenomenon that Msap is misjudged as Wap. The same phenomenon also exists in Wip and Mip. The annotation accuracy rate of the APM-GCN method for various diagenetic facies is above 86%, indicating that it has a good annotation effect. In Figure 9 (b), the Recall of the APM-GCN method for various diagenetic facies is above 90%, indicating that it has a good Recall. In Figure 9 (c), the F1-score of the APM-GCN method is the highest and is above 89% for all, indicating that this method has a good annotation effect for various diagenetic facies and basically meets the accuracy requirements for diagenetic facies label annotation.
[0161] 4. Experimental Results of Model Stability Evaluation
[0162] (1) The calculation results of mAA and mAP for the annotation effects with different data volumes are as shown in Figure 11 The relationship diagrams between different data volumes and mAA and mAP are as shown in Figure 12 .
[0163] From Figure 11 , it can be seen that as the data volume increases from 2000 to 10000, the annotation accuracy rate of diagenetic facies gradually increases, but the growth rate slows down at the data volume of 6000, indicating that increasing the data volume can effectively improve the annotation accuracy rate of diagenetic facies, but after reaching a certain data volume, the effect will gradually decrease. From Figure 12 , it can be seen that at the depth node of 8000, the annotation accuracy rate of various diagenetic facies is the best, about 90%. As the data volume increases, the Accuracy value and mAA value in the annotation effects of various diagenetic facies show an overall upward trend, and the annotation accuracy rate of most diagenetic facies has improved. The mAP value increases from 0.72 (data volume of 2000) to 0.81 (data volume of 10000), but the increase amplitude tends to decrease at large data volumes. However, the increase amplitude gradually decreases at large data volumes. Given that the depth node of 8000 has the best annotation result, this data volume is used for the accuracy experiment.
[0164] (2) The calculation results of mAA and mAP after adding Gaussian noise are as shown in Figure 13As shown. After adding Gaussian noise, the mAA value and mAP value show small fluctuations, with average values of 89.99 and 81.935 respectively, and standard deviations of 0.76 and 0.837 respectively, indicating that the model performance is stable. The maximum and minimum values of the mAA value are 90.85 and 88.48 respectively, with a fluctuation range of 2.37; the maximum and minimum values of the mAP value are 83.15 and 80.84, with a fluctuation range of 2.31. Generally speaking, although adding Gaussian noise has a certain impact on the mAA and mAP values of the model, the overall fluctuation is small, proving that the model shows strong stability and robustness.
[0165] Although the present invention is disclosed as above, the protection scope of the present invention is not limited thereto. Those skilled in the art of the present invention can make various changes and modifications without departing from the spirit and scope of the present invention, and these changes and modifications will all fall within the protection scope of the present invention.
[0166] [1] De Ros, Luiz Fernando, and Karin Goldberg. "Reservoirpetrofacies: a tool for quality characterization and prediction." AAPG, annual convention and exhibition, Long Beach, abstracts volume . 2007.
[0167] [2]Watney, W. L., et al. "Memoir 71, Chapter 6: PetrofaciesAnalysis--A Petrophysical Tool for Geologic / Engineering ReservoirCharacterization." (1999): 73-90.
[0168] [3] [1] Liu Jingjing. Prediction of high-quality tight sandstone reservoirs based on artificial intelligence [D]. Chang'an University, 2023. DOI: 10.26976 / d.cnki.gchau.2022.000152.
[0169] [4] Zou Caineng, Tao Shizhen, Zhou Hui, et al. Formation, classification and quantitative evaluation methods of diagenetic facies [J]. Petroleum Exploration and Development, 2008(05): 526-540.
[0170] [5] Cui, Yufeng, et al. "Prediction of diagenetic facies using welllogs–A case study from the upper Triassic Yanchang Formation, Ordos Basin,China." Marine and Petroleum Geology 81 (2017): 50-65.
[0171] [6] Chapelle, Olivier, Bernhard Scholkopf, and Alexander Zien. "Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]." IEEETransactions on Neural Networks 20.3 (2009): 542-542.
[0172] [7]Yann, L., and Misra Ishan. "Self-supervised learning: The darkmatter of intelligence." Meta AI Blog (2021).
[0173] [8] Krizhevsky, Alex,Ilya Sutskever, and Geoffrey E. Hinton. "Imagenet classification with deep convolutional neural networks." Advances in neural information processing systems 25 (2012).
[0174] [9] Pattnaik, Sonali, et al. "Automatic carbonate rock faciesidentification with deep learning." SPE Annual Technical Conference andExhibition?. SPE, 2020.
[0175]
[10] Wu, Zonghan, et al. "A comprehensive survey on graph neuralnetworks." IEEE transactions on neural networks and learning systems 32.1(2020): 4-24.
[0176]
[11] Kipf, Thomas N., and Max Welling. "Semi-supervised classificationwith graph convolutional networks." arxiv preprint arxiv:1609.02907 (2016).
[0177]
[12] Lewinsohn, Daniel P., et al. "Consensus label propagation withgraph convolutional networks for single-cell RNA sequencing cell typeannotation." Bioinformatics 39.6 (2023): btad360.
[0178]
[13] Lu, Guoqing, et al. "Lithology identification using graph neuralnetwork in continental shale oil reservoirs: A case study in Mahu Sag,Junggar Basin, Western China." Marine and Petroleum Geology 150 (2023):106168.
[0179]
[14] Frey, Brendan J., and Delbert Dueck. "Clustering by passingmessages between data points." science 315.5814 (2007): 972-976.
[0180]
[15] Dueck, Delbert. Affinity propagation: clustering data by passing messages . Toronto, ON, Canada: University of Toronto, 2009.
[0181]
[16] Givoni, Inmar E., and Brendan J. Frey. "A binary variable modelfor affinity propagation." Neural computation 21.6 (2009): 1589-1600.
[0182]
[17] Liu, Y. X., et al. "Diagenesis and pore evolution of reservoirof the Member4 of Lower Cretaceous Quantou Formation in Sanzhao Sag, northernSongliao Basin." J Palaeogeogr 12.4 (2010): 480-7.
[0183]
[18] Ding Sheng, Zhong Siying, Gao Guoqiang, et al. Quantitative evaluation of diagenetic facies of low-permeability reservoirs by combining logging and geology[J]. Journal of Southwest Petroleum University (Science & Technology Edition), 2012, 34(04): 83-87.
[0184]
[19] Tao Liu, Zongbao Liu, Kejia Zhang, Chunsheng Li, Yan Zhang,Zihao Mu, Fang Liu, Xiaowen Liu, Mengning Mu. Research on the generation andannotation method of thin section images of tight oil reservoir based on deeplearning. Scientific Reports, 2024.
[0185]
[20] Yang Wei, et al. "Research methods and applications of carbonate diagenetic facies - taking the oolitic beach reservoir of the Feixianguan Formation on the northern margin of the Yangtze Block as an example." Acta Petrologica Sinica 27.3 (2011): 749-756.
[0186]
[21] Zhao, M., Cao, G., Huang, X., Yang, L. (2022). HybridTransformer-CNN for Real Image Denoising. IEEE Signal Processing Letters, 29,1252-1256.
[0187]
[22] Zhang, **yu, William Ambrose, and Wei **e. "Applyingconvolutional neural networks to identify lithofacies of large-n cores fromthe Permian Basin and Gulf of Mexico: The importance of the quantity andquality of training data." Marine and Petroleum Geology 133 (2021): 105307.
[0188]
[23] Maria Navin, J. R., Pankaja, R. (2016). Performance analysis oftext classification algorithms using confusion matrix. Int. J. Eng. Tech.Res. IJETR, 6, 75-78.
[0189]
[24] Antariksa, G., Muammar, R., Lee, J. (2022). Performanceevaluation of machine learning-based classification with rock-physicsanalysis of geological lithofacies in Tarakan Basin, Indonesia. Journal ofPetroleum Science and Engineering, 208, 109250.
[0190]
[25] N. Pereira. (2022). PereiraASLNet: ASL letter recognition with YOLOX taking Mean Average Precision and Inference Time considerations. In 2022 2nd International Conference on Artificial Intelligence and Signal Processing (AISP), Vijayawada, India, 2022, pp. 1-6, doi: 10.1109 / AISP53593.2022.9760665.
Claims
1. A method for automatic annotation of diagenetic facies samples by integrating affinity propagation clustering and graph convolutional neural network, characterized in that: The following steps are involved: S100, data set establishment, including classification of diagenetic facies types and preprocessing of well logging curve data, wherein the well logging curve data includes well logging curve selection, missing value and outlier processing of the selected well logging curves, and sample labeling, and then labeling the well logging curve data according to the diagenetic facies type; Select the well logging curves that have a great influence on diagenetic phase analysis, including natural gamma, acoustic wave time difference, deep lateral resistivity, shallow lateral resistivity, natural potential and well diameter; S200, using affinity propagation clustering algorithm to build APM graph structure, connecting nodes with similar characteristics but non-adjacent, for sharing diagenetic facies labels, specifically including: S210, reading the logging curve data, selecting the logging curve features according to step S100, forming the feature vector of each node, and generating a feature matrix, as shown in the following formula (6): Where X∈R n×d , n is the number of depths of the logging curve, d is the number of features; S220, the depth nodes of the logging curve are sorted according to the depth, adjacent depth nodes are connected by edges, an empty undirected graph structure based on the depth order is constructed, and the corresponding feature vector is extracted; S230, clustering nodes with non-adjacent depths using affinity propagation clustering algorithm, selecting logging curve features, classifying each logging curve depth node into a corresponding cluster, establishing connections between nodes with similar logging curve features, and then calculating the Euclidean distance D between nodes ij And normalization is performed to make each eigenvalue in the same range: The similarity matrix S takes the negative distance D ij The matrix is: S ij =-D ij (8) L=APCluster(S) (9) S240. Generate an adjacency matrix based on the diagenetic phase labels of the clusters. If two nodes belong to the same cluster, the corresponding position in the adjacency matrix is 1, otherwise it is 0: Among them, L i Indicates the diagenetic facies label of the i-th well log depth node; The generated matrices are the adjacency matrix A and the feature matrix X; S300, using a multi-layer graph convolutional network, combined with the affinity propagation clustering result obtained in step S200, a convolution operation is performed on the well logging curve data processed in step S100, and after multiple iterations, an optimized model structure based on the graph convolution diagenetic facies annotation method is obtained, which specifically includes: The first convolution layer uses GCNConv operation, input dimension: num_features, output dimension: 64; the second convolution layer uses GCNConv operation, input dimension: 64, output dimension: 32; the third convolution layer uses GCNConv operation, input dimension: 32, output dimension: 6; S310: Input the feature matrix X and the adjacency matrix A. The specific convolution operation is as follows: Where I is the identity matrix, is the degree of node i, H (l) is the feature matrix of the lth layer, initially the input feature matrix X, W (l) is the weight matrix of the lth layer, σ is the nonlinear activation function; S320, the output layer uses the softmax function to classify the diagenetic phases: Z=softmax(H (3) ) (14) Among them, Z outputs the category probability distribution; S330. Update model parameters using the Adam optimizer: Among them, θ is the model parameter, η is the learning rate, is the gradient operation of the parameters θ and E is the loss function.
2. The automatic annotation method for diagenetic facies samples integrating affinity propagation clustering and graph convolutional neural network according to claim 1 is characterized in that: In step S100, it specifically includes: S110. Summarize the diagenetic facies types, classify the diagenetic facies types according to compaction, lithology, cementation and dissolution, classify the reservoirs according to the reservoir properties, and classify the diagenetic facies types according to the reservoir types; S120, preprocessing well logging curve data, including well logging curve selection, well logging curve preprocessing and sample labeling; S121, well logging curve selection, analyze the well logging response characteristics of various diagenetic phases, use the decision tree scoring method to score the importance of each well logging curve, and select the well logging curve type with the greatest impact on diagenetic phase analysis; S122, preprocessing the logging curve data obtained in step S121, including processing missing values and abnormal values; The missing values of the logging curve data are predicted using the regression interpolation method. The missing value processing uses formula (1) to build a linear regression model. When interpolating, the regression coefficient is calculated using formula (2). Then, the intercept is obtained using formula (3). The missing values are predicted using the linear regression model. The linear regression model is used to detect outliers. Outlier detection uses formula (4) to analyze the residuals to identify outliers. If the residual value of a data point is too large, these points are judged to be outliers. S123, sample labeling, normalizing the well logging curve data pre-processed in step S122 to make different well logging curve data within the same measurement range; analyzing and labeling the casting thin section images and scanning electron microscope images at corresponding depths to form a data set of automatically labeled diagenetic phase samples.
3. The automatic annotation method of diagenetic facies samples integrating affinity propagation clustering and graph convolutional neural network according to claim 2 is characterized by: The diagenetic facies types are divided according to compaction, lithology, cementation and dissolution, including weak compaction and weak cementation dissolved siltstone facies, weak compaction and weak cementation dissolved fine sandstone facies, medium-strong compaction and weak cementation dissolved fine sandstone facies, medium-strong compaction and weak cementation dissolved siltstone facies, medium-strong compaction and weak dissolution cementation fine sandstone facies and mudstone facies; According to reservoir properties, diagenetic facies types are divided into Type I, Type II and Type III reservoirs. Type I reservoirs correspond to weakly compacted and weakly cemented dissolved siltstone facies, weakly compacted and weakly cemented dissolved fine sandstone facies and medium-strong compacted and weakly cemented dissolved fine sandstone facies; Type II reservoirs correspond to medium-strong compacted and weakly cemented dissolved siltstone facies and medium-strong compacted and weakly dissolved and cemented fine sandstone facies; Type III reservoirs correspond to mudstone facies.
4. The automatic annotation method for diagenetic facies samples integrating affinity propagation clustering and graph convolutional neural network according to claim 3 is characterized by: Set the number of iterations to 2000.
5. An automatic annotation system for diagenetic facies samples integrating affinity propagation clustering and graph convolutional neural network, characterized by: When the system is running, the steps in the method for automatically labeling diagenetic phase samples by integrating affinity propagation clustering and graph convolutional neural network according to any one of claims 1 to 4 are executed.
6. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is configured to implement the steps of the automatic annotation method of diagenetic phase samples integrating affinity propagation clustering and graph convolutional neural network described in any one of claims 1 to 4 when called by a processor.
Citation Information
Patent Citations
Method and system for identifying lithology of rock reservoir
CN117633658A