Offshore buried hill reservoir lithology identification method based on domain adaptive network
Through the domain adaptive network method, the samples imbalance, data distribution differences and log response values during identification of igneous rocks in offshore submerged mountain reservoirs were solved, and high-precision lithologic recognition and good generalization capabilities were achieved.
Patent Information
- Application Number
- CN202510178181.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems such as uneven sample, significant data distribution differences, and overlapping log response values during the identification of igneous rocks in offshore submerged reservoirs, resulting in insufficient generalization capabilities of the model and low recognition accuracy.
A lithologic recognition model is generated by constructing well logging data sets, performing data augmentation processing, building a domain adaptive network and training. The method includes a feature extractor, a tag predictor and a domain discriminator, and dynamically adjusts feature weights through adversarial network and weighting mechanisms to achieve alignment of data distributions of source domain and target domain.
It effectively alleviates the impact of sample imbalance on model performance, improves the classification performance of the model when dealing with overlapping log response characteristics, significantly improves the identification accuracy of complex lithologies in the offshore subsidence reservoirs, and maintains good generalization capabilities.
Smart Images

Figure CN120067809A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of lithology identification, and particularly relates to a method for identifying the lithology of offshore buried hill reservoirs based on a domain adaptive network (DAEDAN). Background Art
[0002] In recent years, buried hill oil and gas reservoirs have gradually become the focus of CNOOC's offshore exploration. Significant oil and gas discoveries have been made successively in the Bohai Bay Basin, the Pearl River Estuary Basin, and the Qiongdongnan Basin, confirming the huge oil and gas exploration potential of buried hills. As an unconventional reservoir, the lithology of offshore buried hill reservoirs is extremely complex. Its bedrock is diverse in structure, tectonics, chemical composition, and mineral composition, with a rich variety of rock types, mainly igneous rocks, specifically including tectonic schist, diorite (including alteration), diabase (including alteration), andesite (including alteration), and granite (including alteration, weathering, etc.).
[0003] Currently, common methods for identifying igneous rocks include crossplot methods, formation element logging, and imaging logging, etc. However, crossplot methods are only applicable to dealing with simple linear or basic non-linear relationships and are difficult to accurately handle the interactions between complex minerals; although formation element logging can provide element content information, it cannot directly and accurately reflect lithology characteristics, especially in rocks with complex compositions or uneven structures, where its performance is limited and it is easily interfered by downhole environmental conditions; although imaging logging can clearly display the structural characteristics of the wellbore, it lacks the ability to analyze mineral compositions. Generally speaking, these traditional methods have low identification accuracy, high uncertainty, and cannot effectively handle complex non-linear mineral relationships when dealing with complex lithologies and multi-mineral combinations.
[0004] With the rapid development of intelligent oil and gas technologies, the application of artificial intelligence in lithology identification has become a research hotspot. Current intelligent logging lithology identification methods usually rely on the quantitative analysis of logging data or expert experience. Researchers select logging curves sensitive to lithology, construct the response values of each curve into a feature vector, and combine the cuttings interpretation results as labels to input into the model for training, thereby generating a classifier that can effectively distinguish lithologies. Compared with traditional logging methods, intelligent lithology identification has the characteristics of simplicity, high efficiency, and the ability to make full use of various information. Intelligent lithology identification methods are mainly divided into two categories: traditional machine learning and deep learning. Traditional machine learning algorithms (such as support vector machines, k-means clustering, and decision trees, etc.) perform well in the identification of conventional lithology reservoirs. In complex lithology reservoirs, due to the heterogeneity of the formation spatial scale and distribution, a single algorithm often has difficulty learning such complex non-linear relationships, and the identification results are also unstable. Therefore, researchers have gradually adopted methods such as ensemble learning and hybrid models to improve the performance of the model. For example, Wang et al. successfully solved the deficiencies of traditional algorithms in aspects such as the selection of initial centers and the deviation of clustering centers caused by outliers by introducing the weighted cosine theory into the KNN model. Yi et al. applied the XGBoost algorithm to volcanic rock identification, constructed three lithology identification schemes based on XGBoost, and achieved good results in practical applications. Given the complex non-linear relationship between logging responses and geological factors in unconventional oil and gas reservoirs, deep learning methods have gradually received attention. Min et al. applied the deep belief network to volcanic rock identification and verified its advantages in dealing with the non-linear relationship between logging responses and complex formations. With the in-depth research, intelligent lithology identification has gradually focused on specific neural network architectures. For example, Lin et al. proposed an automatic lithofacies identification system for borehole data based on the LSTM network and achieved good results in tight sandstone reservoirs. Zhu et al. extracted the curve feature relationships through a convolutional neural network and significantly improved the accuracy of lithology identification. In addition, researchers have also combined unconventional data to further improve the identification accuracy. Valentín M et al. combined conventional logging and micro-resistivity images, adopted a deep residual neural network, and developed a method for automatic lithofacies identification of borehole imaging logging, with remarkable results. Some researchers have also improved the recognition accuracy by adding some additional information to the network. X Ren et al. significantly improved the recognition accuracy by integrating logging data and sedimentary patterns and combining probability statistics.
[0005] Although the above methods have achieved certain results in the prediction of complex lithologies, however, in the identification of buried hill reservoirs, especially igneous rocks, these methods still have some deficiencies, and the main reasons are as follows:
[0006] 1) Problem of sample imbalance: The distribution range of different lithologic reservoirs in buried hills is limited, the vertical thickness is relatively thin, and the available conventional logging information is limited, resulting in insufficient training samples. In addition, there are various types of igneous rocks, which are prone to data imbalance and insufficient data volume for some lithologies. This may cause the model to tend to predict the majority class during training and have poor prediction effects on the minority class.
[0007] 2) Significantly different data distributions: The origin of igneous rocks is extremely complex and is mainly formed by mantle-derived magma and crust-derived magma. Mantle-derived magma is generated by partial melting deep in the mantle and undergoes evolutionary processes and fractional crystallization during the ascent, forming various types of igneous rocks; crust-derived magma is formed by partial melting within the crust, and the differences in crust types (continental crust and oceanic crust) affect the types of igneous rocks formed. This series of complex evolutionary processes results in a wide variety of igneous rocks. The diagenetic process is affected by multiple factors such as temperature, pressure, magma composition, cooling rate, and assimilation, resulting in diverse structures and textures of igneous rocks, and their output forms are mostly irregular bodies such as dikes, sills, and veins, with a lack of regularity in distribution. This complex distribution characteristic is likely to cause the data distribution spaces of the source domain and the target domain to be inconsistent during model training, affecting the generalization ability of the model. Although some scholars have tried to use methods based on instance weighting to reduce the differences between domains, this method depends on target domain labels, is easily affected by human factors, and for the case of large differences in the distribution of igneous rocks, adjusting the weights may not be able to effectively adapt to these complex changes.
[0008] 3) Overlap of logging response value intervals: During the formation of igneous rocks, affected by ore-bearing hydrothermal fluids and then undergoing weathering and leaching after diagenesis, some rocks undergo alteration and secondary minerals are produced. The alteration process will cause significant changes in the physical properties of the bedrock and the composition of some dark minerals, increasing the difficulty of lithology discrimination. Conventional logging response values usually provide a basis for lithology identification from aspects such as the radioactivity, physical properties, pore structure, and facies belt differences of rocks. However, the logging response values of altered rocks will also change accordingly. Coupled with the complexity of the structure and texture of igneous rocks themselves and the influence of various geological factors (such as pressure, temperature, humidity, etc.), the logging response value intervals of different lithologies are prone to overlap. Therefore, simply relying on logging response values as the judgment criterion for lithology identification is difficult to achieve accurate classification of lithologies. Summary of the Invention
[0009] Aiming at the above deficiencies in the prior art, the method for identifying buried hill lithology in the sea based on a domain adaptation network provided by the present invention solves the problems that although existing related methods can improve the identification accuracy by learning complex non-linear features, there are still problems of insufficient generalization ability and model deviation when dealing with data imbalance, feature cross-overlap, and significant offset in well-to-well data distribution.
[0010] To achieve the above invention object, the technical solution adopted by the present invention is: A method for identifying the lithology of buried hill reservoirs at sea based on a domain adaptation network, comprising the following steps:
[0011] S1. Construct a logging dataset, including a target domain dataset and a source domain dataset;
[0012] The data in the source domain dataset includes the logging response curves in the source wells and their corresponding lithology category labels, and the data in the target domain dataset includes the logging response curves in the target wells;
[0013] S2. Perform optimal data augmentation processing on the source domain dataset to obtain a source domain augmented dataset;
[0014] S3. Construct a domain adaptation network, and use the source domain augmented dataset and the target domain dataset to train it to obtain a lithology identification model;
[0015] The domain adaptation network includes a feature extractor, a label predictor, and a domain discriminator. The feature extractor is coupled with the label predictor to form a feedforward neural network; the domain discriminator is connected to the feature extractor through a gradient reversal layer to form an adversarial network;
[0016] S4. Use the label predictor in the lithology identification model to perform lithology classification and identification on the data in the target domain to be identified.
[0017] Further, the step S2 includes the following sub-steps:
[0018] S21. Use a multi-mineral optimal interpretation model to process the logging response curves in the source domain dataset, and calculate and obtain mineral component curves;
[0019] S22. Based on the mineral component curves and lithology sensitive curves in the source domain, and their corresponding lithology category labels, perform oversampling dataset augmentation to obtain a source domain augmented dataset;
[0020] S23. Combine the source domain augmented dataset and the target domain dataset to form an input dataset.
[0021] Further, in the step S21, the expression of the multi-mineral optimization interpretation model is:
[0022]
[0023] In the formula, represents the objective function, represents the response value of the jth component to the ith logging instrument, represents the response value of the formation to the ith logging instrument, represents the relative content of the jth component, represents the maximum relative volume of the j-th component. represents the constant 1, m represents the number of logging instruments, and n represents the number of components in the formation.
[0024] Furthermore, in the step S22, the method for oversampling the dataset for enhancement is specifically as follows:
[0025] Calculate the k-nearest neighbors for each minority class sample in the source domain, randomly select one of the neighbors, and generate new samples through the data enhancement formula to achieve oversampling data enhancement;
[0026] Among them, the data enhancement formula is:
[0027]
[0028] In the formula, represents the generated new sample, represents the original minority class sample, represents the neighbor sample of the original minority class sample, represents a random number within the range of (0, 1).
[0029] Furthermore, in the step S3, in the domain adaptation network:
[0030] The feature extractor is used to map the data in the input source domain enhanced dataset and target domain dataset to a shared feature space;
[0031] The label predictor is used to identify the lithology class labels of the data in the source domain enhanced dataset;
[0032] The domain discriminator is used to classify the data in the feature space and identify its source domain.
[0033] Furthermore, the feature extractor includes a first feature extractor, a second feature extractor, and a feature fusion device;
[0034] The first feature extractor is used to extract mineral features from the mineral data of the source domain enhanced dataset and target domain dataset. The second feature extractor is used to extract logging features from the logging response values of the source domain enhanced dataset and target domain dataset. The feature fusion device is used to dynamically adjust the weights of the extracted features through a weighting mechanism and fuse them to form a joint feature representation ;
[0035] Among them, the joint feature representation is:
[0036]
[0037]
[0038]
[0039]
[0040]
[0041] In the formula, and respectively represent the mineral features and logging features extracted by the first feature extractor and the second feature extractor The learning weight parameters corresponding to the mineral features and respectively represent the mineral features and logging features The corresponding weights, and respectively represent the mineral features and logging features Of the learning weight parameters, and respectively represent the mineral features and logging features The corresponding bias terms.
[0042] Furthermore, in step S3, during the training process of the domain adaptation network, the domain adaptation network continuously minimizes the classification loss of the label predictor for the data with lithology category labels in the source domain, and at the same time minimizes the domain adversarial loss of the domain discriminator on all data in the source domain and the target domain, and finally converges to a stable state to achieve effective transfer learning between the source domain and the target domain.
[0043] Furthermore, during the training process of the domain adaptation network, based on the joint feature representation extracted by the feature extractor, through domain adversarial training, the feature distributions of its source domain and target domain are ensured to have similarity;
[0044] During the domain adversarial training process, the loss function of the domain discriminator Is:
[0045]
[0046] In the formula, Represents the domain label of the i-th sample, the source domain is 1, and the target domain is 0, Represents the predicted probability of the domain predictor for the i-th sample, Represents the joint representation value of the i-th sample, Represents the domain label distribution value, and N represents the total number of samples.
[0047] Further, during the training process of the domain adaptive network, the label predictor classifies the lithology of the data in the target domain based on the joint feature representation, and its classification loss function is:
[0048]
[0049] In the formula, represents the true category of the i-th sample, represents the predicted probability that the i-th sample belongs to the lithology category c, represents the joint representation value of the i-th sample, represents the true category distribution value.
[0050] The beneficial effects of the present invention are as follows:
[0051] (1) The data augmentation strategy in the method of the present invention effectively alleviates the influence of sample data imbalance on the model performance; the experimental results show that in some ensemble learning methods, through the voting mechanism and adaptive sample weight adjustment, more attention and learning weights can be given to the minority class samples, enhancing the anti-imbalance ability of the model.
[0052] (2) In the method of the present invention, the mineral content data calculated based on the optimization interpretation method is fused with the logging response data as the input, providing additional discrimination information for the model. Compared with the data set without using the optimization method, in BP, SVM, and the method of the present invention, the macro-average recall rates are increased by 11%, 11%, and 15% respectively, indicating that this strategy significantly improves the classification performance of the model in dealing with the overlapping of logging response features.
[0053] (3) The domain adversarial training based on the automatic weighting mechanism in the method of the present invention effectively aligns the data distributions of the source domain and the target domain, reducing the influence of data distribution shift; this method is especially suitable for identifying complex lithologies in offshore buried hill reservoirs and can still maintain a high recognition accuracy even in the case of severe data drift.
[0054] (4) Experiments based on logging data show that the lithology identification method proposed in this paper is superior to existing methods in accurately identifying complex lithologies such as igneous rocks, and the model exhibits good generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flowchart of the offshore lithology identification method based on the domain adaptive network provided by the present invention.
[0056] Figure 2 is a structural diagram of the domain adaptive network framework provided by the present invention.
[0057] Figure 3 is a lithology distribution characteristic diagram of the buried hill reservoir in Area YL8 provided by the present invention.
[0058] Figure 4 It is the t-SNE visualization distribution feature map of the test data set provided by the present invention.
[0059] Figure 5 It is the identification result map of the lithology model of Well X provided by the present invention. Specific embodiments
[0060] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0061] An embodiment of the present invention provides a method for identifying the lithology of a buried hill reservoir in the sea based on a domain adaptation network, as Figure 1 shown, including the following steps:
[0062] S1. Construct a logging data set, including a target domain data set and a source domain data set;
[0063] S2. Perform optimal data augmentation processing on the source domain data set to obtain a source domain augmented data set;
[0064] S3. Construct a domain adaptation network and train it using the source domain augmented data set and the target domain data set to obtain a lithology identification model;
[0065] S4. Use the label predictor in the lithology identification model to perform lithology classification and identification on the data in the target domain to be identified.
[0066] In step S1 of the embodiment of the present invention, in the constructed logging data set, the data in the source domain data set includes the logging response curves in the source well and their corresponding lithology category labels, and the data in the target domain data set includes the logging response curves in the target well; wherein, the target well and the source well are wells belonging to the same area or wells that are close to each other in geographical space;
[0067] Specifically, the source domain data set is expressed as , is the training sample, corresponding to the label , C is the number of lithology categories; the target domain data set is expressed as , is the number of target unlabeled samples. In data neighborhood adaptation, and are sampled from different marginal probability distributions.
[0068] Step S2 of the embodiment of the present invention includes the following sub-steps:
[0069] S21. Process the logging response curves in the source domain dataset using the multi-mineral optimal interpretation model, and calculate the mineral component curves through solution.
[0070] S22. Based on the mineral component curves and lithology-sensitive curves in the source domain, and their corresponding lithology class labels, perform oversampling dataset enhancement to obtain the source domain enhanced dataset.
[0071] S23. Combine the source domain enhanced dataset and the target domain dataset to form the input dataset.
[0072] In step S21, during the process of calculating the mineral component curves, the buried hill formation component model can be regarded as composed of various components with different properties, including oil, water, natural gas, shale, and various rock skeleton minerals. The optimal interpretation method is to establish an uncorrelated function between each measured value of the detection instrument and the formation components, and then use the optimization method to find the minimum solution of the uncorrelated function and obtain the optimal approximate solution of each component.
[0073] Suppose the relative contents of various rock skeleton minerals and other components (oil, water, natural gas, shale) in the formation components are respectively: , the response equations of various logging instruments can be written as:
[0074]
[0075] To make full use of logging information and improve the reliability of interpretation, the case of m > n is usually adopted to construct an overdetermined linear equation system to make it have an optimal solution in the sense of least squares. Since the optimal solution may occur < 0 or > 1, to make the solution result conform to geological significance, the expression of the multi-mineral optimization interpretation model is:
[0076]
[0077] In the formula, represents the objective function, represents the response value of the j-th component to the i-th logging instrument, represents the response value of the formation to the i-th logging instrument, represents the relative content of the j-th component, represents the maximum relative volume of the j-th component, represents the constant 1, m represents the number of logging instruments, and n represents the number of components in the formation.
[0078] In step S22 of the embodiment of the present invention, when solving the problem of imbalanced data sets, the SMOTE (Synthetic Minority Over-sampling Technique) algorithm expands the training data set by synthesizing new minority class samples, thereby effectively improving the performance of the classifier. Its core idea is to randomly sample minority class samples and generate new samples by interpolation based on the feature relationships between minority class samples to increase the quantity of data; based on this, the method for oversampling data set enhancement in this embodiment is specifically as follows:
[0079] Calculate the k-nearest neighbors for each minority class sample in the source domain, randomly select one of its neighbors, and generate new samples through the data augmentation formula to achieve oversampling data augmentation;
[0080] Among them, the data augmentation formula is:
[0081]
[0082] In the formula, represents the generated new sample, represents the original minority class sample, represents the neighbor sample of the original minority class sample, represents a random number within the range of (0, 1).
[0083] The above oversampling data augmentation method enriches the distribution of minority class samples and improves the classification performance of the model.
[0084] In step S3 of the embodiment of the present invention, the domain adaptation network (DANN) is a representation learning technology for domain adaptation, aiming to learn feature representations that can be shared between the source domain and the target domain, so as to achieve effective model migration
[40] . DANN adopts an adversarial training method by simultaneously considering the domain invariance and discriminability of features, so that the features generated by the feature extractor are distributed uniformly between different domains, while maintaining good classification performance on the source domain.
[0085] As Figure 2 shown, the domain adaptation network (DANN) architecture includes three core components: a feature extractor, a label predictor, and a domain discriminator; among them, the feature extractor is coupled with the label predictor to form a feedforward neural network, which is mainly responsible for accurately classifying source domain data; the domain discriminator is connected to the feature extractor through a gradient reversal layer to form an adversarial network;
[0086] Among them, the feature extractor is used to map the data in the input source domain enhanced dataset and target domain dataset to a shared feature space, enabling the label predictor to effectively distinguish the data categories in the source domain while making it difficult for the domain discriminator to distinguish the source domain of the data; the label predictor is used to identify the lithology category labels of the data in the source domain enhanced dataset; the domain discriminator is used to classify the data in the feature space and identify its source domain.
[0087] In this embodiment, the feature extractor includes a first feature extractor, a second feature extractor, and a feature fusion device;
[0088] The first feature extractor is used to extract mineral features from the mineral data in the source domain enhanced dataset and the target domain dataset, the second feature extractor is used to extract logging features from the logging response values in the source domain enhanced dataset and the target domain dataset, and the feature fusion device is used to dynamically adjust the weights of the extracted features through a weighting mechanism and fuse them to form a joint feature representation ;
[0089] Among them, the joint feature representation is:
[0090]
[0091]
[0092]
[0093]
[0094]
[0095] In the formula, and respectively represent the mineral features and logging features extracted by the first feature extractor and the second feature extractor , and respectively represent the weights corresponding to the mineral features and the logging features , and respectively represent the learning weight parameters of the mineral features and the logging features , and respectively represent the bias terms corresponding to the mineral features and the logging features .
[0096] Through the above-mentioned feature extraction and weight-based multi-source fusion process by the feature extractor, the model can automatically adjust the influence of the two feature sources according to the input data.
[0097] In step S3 of the embodiment of the present invention, during the training process of the domain adaptation network, the domain adaptation network continuously minimizes the classification loss of the label predictor for the data with lithology category labels in the source domain, and at the same time minimizes the domain adversarial loss of the domain discriminator on all the data in the source domain and the target domain, and finally converges to a stable state to achieve effective transfer learning between the source domain and the target domain.
[0098] In this embodiment, during the training process of the domain adaptation network, based on the joint feature representation extracted by the feature extractor, through domain adversarial training, the feature distributions of the source domain and the target domain are ensured to make the extracted features have similarity between the two source domains and the target domain;
[0099] During the domain adversarial training process, the loss function of the domain discriminator is:
[0100]
[0101] In the formula, represents the domain label of the i-th sample, the source domain is 1, and the target domain is 0, represents the predicted probability of the domain predictor for the i-th sample, represents the joint representation value of the i-th sample, represents the domain label distribution value, and N represents the total number of samples.
[0102] In this embodiment, the core of the above-mentioned domain adversarial training is to optimize the feature extractor through the Gradient Reversal Layer (GRL), making it difficult for the domain discriminator to distinguish the source of the samples; specifically, while maximizing the loss of the domain discriminator, the feature extractor generates feature representations that are as domain-invariant as possible.
[0103] In this embodiment, during the training process of the domain adaptation network, the label predictor classifies the lithology of the data in the target domain based on the joint feature representation, and its classification loss function is:
[0104]
[0105] In the formula, represents the true category of the i-th sample, represents the predicted probability that the i-th sample belongs to the lithology category c, represents the joint representation value of the i-th sample, represents the true category distribution value.
[0106] Based on the domain adversarial loss of the above domain discriminator and the classification loss of the label predictor, the total loss function of the domain adaptation network is obtained as follows:
[0107]
[0108] wherein, represents the weight hyperparameter that controls the weight between the classification loss and the domain adversarial loss; by alternately optimizing the domain predictor and the feature extractor, the model can reduce the distribution difference between the source domain and the target domain while ensuring the classification accuracy.
[0109] In the embodiment of the present invention, an effect verification experimental example of the above lithology identification method is provided.
[0110] In this embodiment, the data is selected from 7 exploration wells in the YL8 area. The curve types and sample numbers of each well dataset are shown in Table 1 in detail;
[0111] Table 1: Curve types and sample numbers of well datasets
[0112]
[0113] All experiments were carried out on a PC equipped with an Intel Core i5-12500H processor. The model was implemented using Python 3.8, and the libraries used included PyTorch, Scikit-learn, and Imblearn. For the hyperparameter setting of the DANN network, its feature extractor, domain classifier, and label classifier have the same structure, and are all composed of two fully connected layers containing 64 nodes. The ReLU activation function is used between the hidden layers. The Adam algorithm is used as the optimizer. Before each experiment, the data was subjected to maximum-minimum normalization and standardization processing. The hyperparameters of the intelligent lithology identification model used in the experiment were all tuned to the optimal by the grid search method.
[0114] In order to optimize the hyperparameters of the model, in this embodiment, a grid search algorithm was used to traverse and search each model within the specified parameter range. Specifically:
[0115] Learning rate: Search between 0.001 and 0.01 with a step size of 0.001.
[0116] Number of training iterations (epochs): Vary between 50 and 200 times with a step size of 25.
[0117] Regularization parameter α (alpha): Search between 0.001 and 0.1 with a step size of 0.005.
[0118] Domain adaptation parameter λ: Vary between 1 and 10 with a step size of 1.
[0119] Through a comprehensive search of the above hyperparameter space, the impacts of different parameter combinations on the model performance were evaluated, and the optimal parameter configuration was finally determined as follows: the learning rate was set to 0.005, the number of training iterations was 150, the regularization parameter α was 0.01, and the domain adaptation parameter λ was 4.
[0120] To evaluate the accuracy and applicability of the domain adaptation network for data augmentation (DAEDAN) method for identifying complex lithologies in buried hill reservoirs, the DAEDAN model was compared with four models: support vector machine (SVM), random forest (RF), XGBoost, and BP genetic neural network. Figure 3 80% of the lithology dataset was used for training and 20% for testing. To ensure the robustness and credibility of the experimental results, each group of experiments was independently run 10 times. Each experiment used different random number seeds to partition the dataset, and the model performance was compared and analyzed by calculating the macro recall rate.
[0121] Figure 3 Figure 10 shows the lithology distribution characteristics map of the buried hill reservoir in Area YL8. It can be seen from the figure that there is an obvious problem of dataset imbalance in the lithology samples in Area YL8. The number of diorite samples is 8 times that of the diabase samples, which has a significant impact on the model performance. To verify the effectiveness of the SMOTE data augmentation method, we conducted comparative experiments on six lithology identification methods with SMOTE augmentation. The results are shown in Tables 2 and 3. In the tables, Macro-R represents the macro-average recall rate. It can be seen that data imbalance has a significant impact on various intelligent lithology identification methods. In particular, the classification accuracy of SVM, BP, and DAEDAN networks in diabase identification almost drops to zero. The model misclassifies almost all diabase samples as diorite, mainly for two reasons: one is that the logging response values of diabase and diorite overlap greatly, and the other is that due to the large number of diorite samples, the model has an excessive bias towards it, resulting in a decline in the diabase identification performance.
[0122] Table 2: Experimental results of different prediction models without SMOTE augmentation
[0123]
[0124] Table 3: Experimental results of different prediction models with SMOTE augmentation
[0125]
[0126] In this embodiment, the influence of the mineral content curve calculated by the optimized interpretation method on the performance of each classification model was also evaluated. First, the mineral content curve of the dataset was generated using the optimized interpretation method. Then, through feature splicing or multi-source feature fusion, the mineral content curve was integrated with the well logging response data to generate a new multi-source feature vector. Next, the Synthetic Minority Over-sampling Technique (SMOTE) was applied to enhance the training dataset to mitigate the adverse effects of class imbalance on model training, and it was respectively input into each classification model for training and performance evaluation. As shown in Table 4, compared with the dataset without the mineral curve (see Table 3), the classification accuracy of each classification model has been significantly improved. Especially in the BP, SVM, and the proposed DAEDAN model in this paper, the macro-average recall rates have increased by 11%, 11%, and 15% respectively. Specifically, especially for the two types of lithologies that are difficult to distinguish, diabase and diorite, the recognition accuracy of each model has been significantly improved after introducing the mineral curve. Figure 4 It is the t-SNE visualization distribution feature map of the test dataset in Area YL8. The t-SNE (t-distributed Stochastic Neighbor Embedding) visualization method was used to compare the distribution of feature scatter points before and after introducing the mineral curve (in the figure, gray represents diorite, light blue represents diabase, dark blue represents dacite, brown represents andesite, and green represents granite). Figure 4 In (a) is the feature distribution map of the test dataset without fusing the optimized interpretation method. Figure 4 In (b) is the feature distribution map of the test dataset fusing the optimized interpretation method. The results show that after adding the mineral curve, the sample distribution of each lithology category is more compact, and the boundary between categories is clearer. This indicates that the mineral curve calculated by the optimized interpretation method provides additional discrimination information for the model, significantly improving the classification performance of the model in dealing with the overlap of well logging response features.
[0127] Table 4: Comparison table of model recognition results integrating SMOTE enhancement and the optimization method
[0128]
[0129] In practical applications, data from different regions are usually fused to achieve accurate identification of the lithology in this area. And the depth of the lithology well section to be predicted is continuously distributed, usually covering all lithology types in this area. Therefore, 100 depth-continuous samples of each lithology were extracted respectively, and the comprehensive data X well set was constructed by splicing for model verification. As Figure 5As shown (in the figure, purple represents rhyolite, orange represents andesite, gray represents granite, blue represents diorite, and green represents diabase), the recognition results of XGBoost, RF, SVM, and DAEDAN are shown from left to right in the figure. Table 5 is the evaluation table of the recognition results of the four models. It can be seen that the model proposed in this paper shows excellent performance in this mixed data, and its macro-precision, macro-recall, and macro F1-score reach 92.0%, 90.8%, and 90.8% respectively. In addition, since the test set is the same lithology for a continuous section in a well, the data distribution differences between different wells are extremely large; and the number of samples in diabase itself is small, and the discriminator has limited learnable patterns; coupled with the large overlap of eigenvalue between diabase and diorite, the prediction effect of each model for this lithology is very poor, and a large number of diabases are identified as diorites. When DAEDAN identifies the lithologies of diorite and diabase with greater recognition difficulty, the recall rates of recognition still reach 87% and 73%.
[0130] Table 5: Evaluation Table of Recognition Results of Four Classifiers
[0131]
[0132] In the present invention, specific embodiments are used to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
[0133] Those of ordinary skill in the art will realize that the embodiments described herein are for helping readers understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.
Claims
1. A method for identifying lithology of offshore buried hill reservoirs based on domain adaptive network, characterized in that: The following steps are involved: S1. Construct a well logging dataset, including a target domain dataset and a source domain dataset; The data in the source domain data set include the well logging response curve in the source well and its corresponding lithology category label, and the data in the target domain data set include the well logging response curve in the target well; S2, performing optimal data enhancement processing on the source domain dataset to obtain a source domain enhanced dataset; S3, construct a domain adaptive network, and train it using the source domain enhanced data set and the target domain data set to obtain a lithology recognition model; The domain adaptation network includes a feature extractor, a label predictor and a domain discriminator, wherein the feature extractor is coupled with the label predictor to form a feedforward neural network; The domain discriminator is connected to the feature extractor through a gradient reversal layer to form an adversarial network; S4. Use the label predictor in the lithology recognition model to classify and recognize the lithology of the data in the target domain to be identified.
2. The offshore buried hill lithology identification method based on domain adaptive network according to claim 1 is characterized in that: The step S2 comprises the following sub-steps: S21, using a multi-mineral optimal interpretation model to process the logging response curve in the source domain data set, and obtain a mineral component curve; S22, based on the mineral component curve and lithology sensitivity curve of the source domain and their corresponding lithology category labels, an oversampled data set is enhanced to obtain a source domain enhanced data set; S23: The source domain enhanced dataset and the target domain dataset are combined to form an input dataset.
3. The offshore buried hill lithology identification method based on domain adaptive network according to claim 2 is characterized in that: In step S21, the expression of the multi-mineral optimization interpretation model is: In the formula, represents the objective function, represents the response value of the jth component to the i-th logging tool, represents the response value of the formation to the i-th logging tool, represents the relative content of the jth component, represents the maximum relative volume of the jth component, represents the constant 1, m represents the number of logging instruments, and n represents the number of components in the formation.
4. The offshore buried hill lithology identification method based on domain adaptive network according to claim 2 is characterized in that: In step S22, the method for performing oversampled data set enhancement is specifically as follows: Calculate the k-nearest neighbors of each minority class sample in the source domain, randomly select one of the neighbors, generate a new sample through the data enhancement formula, and achieve oversampling data enhancement; Among them, the data enhancement formula is: In the formula, represents the generated new sample, represents the original minority class samples, represents the neighbor samples of the original minority class samples, Represents a random number in the range (0,1).
5. The offshore buried hill lithology identification method based on domain adaptive network according to claim 1 is characterized in that: In step S3, in the domain adaptive network: The feature extractor is used to map data in the input source domain enhanced dataset and the target domain dataset to a shared feature space; The label predictor is used to identify the lithology category labels of the data in the source domain enhanced dataset; The domain discriminator is used to classify the data in the feature space and identify its source domain.
6. The offshore buried hill lithology identification method based on domain adaptive network according to claim 5 is characterized in that: The feature extractor includes a first feature extractor, a second feature extractor and a feature fusion device; The first feature extractor is used to extract mineral features from the mineral data of the source domain enhanced data set and the target domain data set, the second feature extractor is used to extract well logging features from the well logging response values of the source domain enhanced data set and the target domain data set, and the feature fusion device is used to dynamically adjust the weights of the extracted features through a weighting mechanism and fuse them to form a joint feature representation ; Among them, the joint feature representation for: In the formula, and Represent the first feature extractor and the second feature extractor Extracted mineral features and well logging features, and Represents mineral characteristics and logging characteristics The corresponding weight, and Represents mineral characteristics and logging characteristics The learning weight parameters are and Represents mineral characteristics and logging characteristics The corresponding bias term.
7. The offshore buried hill lithology identification method based on domain adaptive network according to claim 5 is characterized in that: In step S3, during the training process of the domain adaptation network, the domain adaptation network continuously minimizes the classification loss of the label predictor for the data with lithology category labels in the source domain, and at the same time minimizes the domain adversarial loss of the domain discriminator on all data in the source domain and the target domain, and finally converges to a stable state, thereby realizing effective transfer learning between the source domain and the target domain.
8. The offshore buried hill lithology identification method based on domain adaptive network according to claim 7 is characterized in that: In the training process of the domain adaptation network, based on the joint feature representation extracted by the feature extractor, the feature distribution of the source domain and the target domain is adjusted through domain adversarial training to ensure that the extracted features have similarities between the two source domains and the target domain; During domain adversarial training, the loss function of the domain discriminator is for: In the formula, represents the domain label of the i-th sample, the source domain is 1 and the target domain is 0. represents the predicted probability of the domain predictor for the i-th sample, represents the joint representation value of the i-th sample, represents the domain label distribution value, and N represents the total number of samples.
9. The offshore buried hill lithology identification method based on domain adaptive network according to claim 8, characterized in that: During the training process of the domain adaptation network, the label predictor classifies the lithology of the data in the target domain based on the joint feature representation, and its classification loss function is: In the formula, represents the true category of the i-th sample, represents the predicted probability that the i-th sample belongs to lithology category c, represents the joint representation value of the i-th sample, Represents the true category distribution value.
Citation Information
Cited By
Slope rock mass structural surface generalization identification method and system based on multi-level domain adaptation
CN121904740A