Three-dimensional intelligent interpolation method based on deep learning

By combining a deep learning-based 3D intelligent interpolation method with class-based non-equilibrium modeling and deep learning, the inapplicability problem in complex reservoir lithology identification is solved, achieving high-precision reservoir lithology prediction and 3D model construction, thus improving the success rate of reservoir exploration.

CN115880455BActive Publication Date: 2026-05-01CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA PETROLEUM & CHEMICAL CORP
Filing Date
2021-09-26
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies are not applicable to the identification of complex reservoir lithology, making it difficult to achieve high-precision reservoir lithology prediction. Furthermore, traditional logging lithology interpretation methods are inefficient and greatly affected by human factors, and cannot effectively handle unbalanced logging data.

Method used

A deep learning-based 3D intelligent interpolation method is adopted. Through quantitative evaluation of the feature matching relationship of multi-type geological heterogeneous data, combined with classification and non-equilibrium, intelligent identification of logging lithology is carried out. Delaunay triangulation and CNN convolutional neural network are used for 3D intelligent interpolation to establish high-quality input data of reservoir parameters.

Benefits of technology

It achieves more accurate intelligent identification of well logging lithology, effectively reduces the ambiguity of seismic prediction, improves the accuracy of reservoir prediction and the success rate of exploration of complex oil reservoirs, and provides high-quality data support for the construction of 3D models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115880455B_ABST
    Figure CN115880455B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional intelligent interpolation method based on deep learning, which comprises the following steps: 1, performing quantitative evaluation on the feature and matching relationship of multi-type geological heterogeneous data; 2, performing intelligent identification of well logging lithology based on the combination of class removal non-uniformity and deep learning; and 3, performing three-dimensional intelligent interpolation based on deep learning. The three-dimensional intelligent interpolation method based on deep learning effectively fuses a large amount of well-seismic data and geological research results based on artificial intelligence technology, and the result can be provided to geophysical personnel for three-dimensional model construction and reservoir prediction research, thereby laying a solid foundation for the next step research of determining favorable reservoirs, assisting well location design and calculating reserves.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of petroleum exploration technology, and in particular to a three-dimensional intelligent interpolation method based on deep learning. Background Technology

[0002] Big data is impacting our lives and has significant implications for our research into new problems and our review of old ones. Key scientific and technological issues in the development of geological science in the era of big data include the integrated storage, management, and processing of structured and semi-structured / unstructured data, big and small data, mixed and precise data, models and data, static exploration models and dynamic monitoring models, the combination of data mining and data analysis, the unification of correlation and causation, and the in-depth mining and visualization of geological big data. The key to analyzing geoscientific big data lies in conducting comprehensive, multi-faceted, and multi-dimensional analysis of diverse and heterogeneous geoscientific data. Zuo Renguang analyzed the identification patterns of weak geochemical anomalies and applied the local RX method for multivariate dimensionality reduction and extraction of weak and slow geochemical anomalies, achieving good results. This is due to the semantic complexity of the spatiotemporal changes of geological bodies and phenomena, the nonlinearity and uncertainty of processes, and the multi-dimensional, multi-scale, and real-time nature of change information. Wu Chonglong proposed the integration and utilization of geological spatiotemporal big data, involving a series of theoretical, methodological, and technical issues, mainly including an integrated spatial reference system for geology and geography; the coupling of static structural exploration models and dynamic change monitoring models; and data mining, data fusion, and intelligent processing technologies for geological big data. Developing intelligent data processing methods can keep pace with the extraordinary growth of big data; therefore, artificial intelligence geology should be an important direction for development. In the field of artificial intelligence research, through long-term practice, many scientific methods and application technologies have been accumulated, such as information extraction, cluster analysis, quantitative interpolation, and automatic dimensionality expansion.

[0003] Accurate reservoir lithology prediction and identification are prerequisites for successful oil and gas exploration and development. Well logging, as an important oil and gas exploration tool, can provide high-resolution and large-volume reservoir rock physical response information for effective reservoir lithology identification. Traditional well logging lithology interpretation is mostly based on the professional knowledge and regional experience of the interpreters. Due to noise interference during the acquisition of various types of well logging data, and the overlapping phenomenon of well logging responses of different lithologies, traditional well logging lithology interpretation methods are often inefficient and greatly affected by human factors. In recent years, machine learning, as an intelligent and efficient data mining method, has been increasingly applied in well logging interpretation. Dubois et al. (2007) and Adrielle et al. proposed using neural networks (NN) to achieve reservoir lithofacies classification based on well logging data. Anazi et al. (2010), Sebtosheikh et al. (2015), and Zhang Xiangjun et al. (2018) successively used the support vector machine algorithm (SVM) for well logging reservoir prediction. (2015) Li et al. (2010) and Shi et al. (2015) used the decision tree (DT) algorithm for well logging identification of complex lithology. Real-world well logging data often exhibits imbalance, with some classes having significantly more samples than others. Single-model algorithms have limitations in interpreting well logging data. Ensemble learning (EL) methods can effectively address this issue. Li Yijing et al. (2016) and Chai Mingrui et al. (2017) applied boosting algorithms from ensemble learning to classification in imbalanced data and machine learning methods for well logging prediction of sandstone and conglomerate cuttings composition. Cracknell et al. (2013) and Zhou Xueqing et al. (2017) respectively applied random forest (RF) classification and regression methods to reservoir prediction and lithofacies classification. Extreme gradient boosting (XGBoost) is an optimization of the boosting algorithm (Chen et al., 2016; Torlay et al., 2017), possessing higher model accuracy and generalization ability. It incorporates regularization terms to prevent overfitting and supports parallel computation, finding widespread application in image recognition. Yan et al. (2019) applied the XGBoost algorithm to well logging reservoir classification with some success. However, with the deepening of exploration and the increase in oil and gas development costs, the inapplicability of shallow machine learning models in the identification of complex reservoir lithology stands in stark contrast to the higher requirements for the accuracy and reliability of reservoir lithology predictions in exploration and development. Faced with the more complex and subtle nonlinear mapping relationship between well logging response and geological lithology information in well logging interpretation, there is an urgent need to propose an intelligent well logging lithology interpretation method with stronger data mining capabilities to achieve high-precision identification of complex reservoir lithology.

[0004] To visualize geoscientific big data, geological simulation techniques can be used to construct the data. Geological simulation requires transforming discrete, finite-space sample point data into continuous, visible geological profiles or geological bodies. In some cases, data gaps or low collection rates during historical processes prevent the data from achieving the goals of big data visualization. Geoscientific sampling point data processing falls into two categories: one involves processing the data into uniform or regular grid structured data, which facilitates stereoscopic visualization, offers a wide variety of algorithms, and has high efficiency in real-time image rendering; the other involves directly processing the estimated values ​​obtained from interpolation points.

[0005] Because geoscientific data are often unevenly distributed and limited in quantity, but certain connections and regularities still exist among them, "we can combine geological principles and relevant geological conditions to select interpolation methods" to obtain continuous geological profiles or geological bodies through geological data interpolation. One paper proposes a 5D interpolation algorithm based on least-squares offset-driven interpolation, which interpolates data to missing offsets and azimuths by de-rasterizing the current subsurface image. The first approach is Generalized Cross-Validation (GCV), a guided method for surface prediction errors that does not require prior knowledge of noise levels. By extending existing fast interpolation frameworks, it can construct realistic and smooth fits for the very large datasets typically collected in geophysics. To balance geological accuracy requirements, surface continuity, and data storage, Li Mingchao et al. proposed and implemented a complex geological surface interpolation approximation fitting construction method based on NURBS technology. This method employs NURBS skinning interpolation for concentrated and uniformly distributed raw data in key engineering areas, ensuring the surface strictly passes through these data points. For discretely distributed data in surrounding areas, NURBS approximation fitting is used to ensure the surface fully approximates the original data with a given accuracy. Finally, the geological structure rationality, geometry, and accuracy of the overall surface are checked, analyzed, and adjusted. For three-dimensional geological bodies, 3D interpolation differs from 2D interpolation in many details, involving fitting calculations in multiple directions. Spatial interpolation converts discretely distributed borehole exploration data into continuous geological surface data. Using geostatistical models or mathematical functions, based on known data and their interrelationships, data values ​​at unknown points and within regions are interpolated or extrapolated to form a continuous surface. Commonly used geometric methods include the Shepard method, the Delaunay tetrahedral partitioning method, and the Kriging method, a representative spatial statistical method. However, these methods are not highly accurate. Using borehole data to interpolate and construct geological surfaces is currently the most effective and reliable method for 3D geological modeling. Based on the distribution characteristics of borehole data and the actual conditions of the geological body, selecting a suitable and effective interpolation method is fundamental to ensuring the accuracy of the interpolation results.

[0006] The existing technologies described above are significantly different from the present invention and have failed to solve the technical problem we want to address. Therefore, we have invented a new three-dimensional intelligent interpolation method based on deep learning. Summary of the Invention

[0007] The purpose of this invention is to provide a deep learning-based three-dimensional intelligent interpolation method that is efficient and accurate in reflecting the spatial variation laws and characteristics of complex lithology.

[0008] The objective of this invention can be achieved through the following technical measures: a deep learning-based 3D intelligent interpolation method, which includes:

[0009] Step 1: Conduct a quantitative evaluation of the characteristics and matching relationships of various types of geological heterogeneous data;

[0010] Step 2: Perform intelligent identification of logging lithology based on a combination of class-based non-equilibrium classification and deep learning;

[0011] Step 3: Perform 3D intelligent interpolation based on deep learning.

[0012] The objective of this invention can also be achieved through the following technical measures:

[0013] Step 1 includes:

[0014] Step 1.1: Prepare the input data;

[0015] Step 1.2: Construct a multidisciplinary data analysis and knowledge base;

[0016] Step 1.3: Optimize and construct reservoir lithofacies logging data;

[0017] Step 1.4: Establish the rock physical relationship between elastic parameters and reservoir physical property parameters;

[0018] Step 1.5: Establish quantitative evaluation formulas for geological elements and reservoir lithology and physical properties.

[0019] In step 1.1, the collected multi-type geological heterogeneous data are imported into the database, and preprocessing such as data cleaning and data transformation is performed on the various information. These geological heterogeneous data are then categorized according to three-dimensional coordinates for the next matching step.

[0020] In step 1.2, the characteristics of multidisciplinary data such as well logging, drilling, geophysical exploration, and geology are comprehensively analyzed to study the connotation of heterogeneous data representation; a knowledge base of lithology, reservoir, reservoir framework, and reservoir physical properties is constructed based on the reservoir target geology model.

[0021] In step 1.3, well logging information analysis of geological reservoirs and interlayers is performed to find the correlation between reservoirs and multiple disciplines; principal component analysis and curve optimization and correlation analysis based on long short-term memory neural network are carried out on a large amount of well logging information to realize the optimization and construction of lithofacies well logging data.

[0022] In step 1.4, based on the effective medium theory, adaptive theory, contact theory, and anisotropic model, the rock physics relationship between elastic parameters and physical property parameters is established.

[0023] In step 1.5, based on the quantitative evaluation formulas for various parameters established through extensive statistical analysis, and with the detailed analysis of the development area, the matching relationship between the macro-geological elements of the study area and various reservoir parameters is established through the organic combination of the two, thus forming a quantitative evaluation formula.

[0024] Step 2 includes:

[0025] Step 2.1: Create learning samples;

[0026] Step 2.2: Perform data de-imbalance processing based on the MAHAKIL (Kruskal) method;

[0027] Step 2.3: Group the learning samples. The training set is used for model training, the validation set is used for model hyperparameter tuning, and the validation set is used for model optimization.

[0028] Step 2.4: Build a deep learning network of appropriate size based on the number of learning samples and train the network;

[0029] Step 2.5: Use the optimized deep learning model to perform well logging lithology or intelligent lithology identification for unknown wells.

[0030] In step 2.1, learning samples are established based on the logging response curves and well string interpretation results. Each sample includes logging data such as sonic transit time, bulk density, compensated neutrons, natural gamma, deep induction resistivity, depth, and interpretation conclusions.

[0031] In step 2.2, the MAHAKIL oversampling method is used to generate new samples for the minority class by simulating the reproductive process in genetics. This method removes imbalances while maximizing the diversity of the new samples. The process consists of three steps:

[0032] ① Separate the minority class samples from the dataset that needs to be processed, denoted as Cmin. For each minority class sample in this class, calculate its Mahalanobis distance.

[0033] ② Sort Cmin according to Mahalanobis distance, and divide Cmin in half from the median of the sequence number, denoted as Cmin1 and Cmin2 respectively;

[0034] ③ Select data samples from Cmin1 and Cmin2 in pairs, calculate the mean of the pair, and use it as the new sample to generate;

[0035] If the number of new samples generated in one round is insufficient to meet the requirement of the highest probability, the newly generated samples and samples in Cmin1 and Cmin2 will continue to reproduce. If it is still insufficient, the two newly generated samples will then reproduce with samples in their respective classes, and so on.

[0036] Step 3 includes:

[0037] Step 3.1: Prepare the input data by using different reservoir parameters, including: velocity, density, porosity, clay content, fluid saturation, or the logging lithology intelligently identified in Step 2, as input data.

[0038] Step 3.2: Regularization of input data. The random forest regression algorithm is used to regularize the input data.

[0039] Step 3.3: Input the structural interpretation horizons of different vertical target segments to determine the three-dimensional volume data interpolation framework model;

[0040] Step 3.4: Perform piecewise interpolation based on Delaunay triangulation;

[0041] Step 3.5: Perform CNN-based 3D intelligent interpolation.

[0042] Step 3.4 includes:

[0043] ① Use different reservoir parameters or intelligently identified logging lithology data as input to construct a large triangle containing all scattered points and put them into a triangle linked list;

[0044] ② Insert the scattered points in the point set one by one, find the circumcircle of the triangle in the triangle linked list, the triangle containing the insertion point, delete the common edges that affect the triangle, and connect all the vertices of the triangle that affect the insertion point to complete the insertion of a point in the Delaunay triangle linked list.

[0045] ③ Optimize the newly formed triangles locally according to the optimization criteria, and put the newly formed triangles into the Delaunay triangle linked list.

[0046] Step 3.5 includes:

[0047] ① Extraction of training samples: The piecewise interpolated Delaunay triangle obtained in step 3.4 is used as the effective label, and the earthquake amplitude value within the triangle range is used as the sample feature value;

[0048] ②Establish network architecture: Establish a 100-layer Unet network. The first 50 layers are downsampled. During the sampling process, in order to prevent the loss of edge information of some seismic data, a padding operation is added to the input data before convolution. In the last 50 layers, in order to obtain more low-frequency information, the downsampled low-frequency information and the upsampled high-frequency information are channel-merged.

[0049] ③ Training the model: Divide the sample data in an 8:2 ratio, with 80% used for training and 20% for testing;

[0050] ④ Model Output: When the loss function of the trained model meets the requirements, the model is output;

[0051] ⑤ Model application: Extend the model to three-dimensional space to obtain three-dimensional lithological bodies.

[0052] This invention presents a deep learning-based 3D intelligent interpolation method for reservoir prediction in oil exploration, encompassing 3D geological body modeling, seismic inversion, lithology identification, and fluid identification. The method includes: a well logging lithology intelligent identification method combining class-based de-equilibriumization and deep learning. While de-equilibriumization, it maximizes the diversity of new samples. The equalized new samples are then used for deep learning to obtain the complex response relationship between well logging response values ​​and well logging lithology, thereby achieving more accurate well logging lithology intelligent identification. A multidisciplinary heterogeneous big data structure is established, and a deep learning-based reservoir parameter interpolation method is built on this geoscience big data structure. This effectively reduces the ambiguity of seismic prediction, improves the accuracy of reservoir prediction, and increases the success rate of complex reservoir exploration. Based on artificial intelligence technology, this method effectively integrates a large amount of well and seismic data with geological research results. The results can be provided to geophysicists for 3D model construction and reservoir prediction research, laying a solid foundation for geologists to identify favorable reservoirs, assist in well location design, and calculate reserves.

[0053] This invention addresses the limitations of machine learning in well logging lithology identification or classification, which are constrained by a large number of classes, insufficient training samples, and uncontrollable class imbalance. It proposes an intelligent well logging lithology identification method based on a combination of class imbalance removal and deep learning. The imbalance removal process employs the Mahakil oversampling method, which generates new samples for minority classes by simulating the reproductive process in genetics, maximizing the diversity of new samples while removing imbalance. The balanced new samples are then used for deep learning to obtain the complex response relationship between well logging response values ​​and well logging lithology, thereby achieving more accurate intelligent well logging lithology identification.

[0054] This study investigates deep learning network structures with spatiotemporal characteristics, constructs an intelligent relationship model between seismic data and reservoir parameters based on a deep learning platform, and optimizes network coefficients using the intelligent optimizer of the deep learning platform. This results in the construction of an optimized deep learning network for three-dimensional intelligent interpolation of reservoir parameters such as velocity, density, porosity, clay content, and fluid saturation.

[0055] This paper systematically analyzes the advantages and disadvantages of geostatistical interpolation algorithms, combines multidisciplinary heterogeneous data analysis to explore the intrinsic relationship between reservoir parameters and heterogeneous data, and uses random forest regression algorithm based on different disciplinary geoscience models to regularize heterogeneous data, providing high-quality input data for three-dimensional intelligent interpolation.

[0056] This paper discusses the construction of deep learning network models considering spatiotemporal relationships, the programming of deep learning network algorithms, the design of CNN deep learning methods, and the completion of three-dimensional intelligent interpolation. Attached Figure Description

[0057] Figure 1 This is a flowchart of a specific embodiment of the deep learning-based 3D intelligent interpolation method of the present invention;

[0058] Figure 2 This is a schematic diagram of the MAHAKIL method flow in a specific embodiment of the present invention;

[0059] Figure 3 This is a diagram of the intelligent identification network structure for well logging lithology in a specific embodiment of the present invention. Detailed Implementation

[0060] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0061] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments of the present invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, and / or combinations thereof.

[0062] like Figure 1 As shown, Figure 1 This is a flowchart of the deep learning-based 3D intelligent interpolation method of the present invention.

[0063] 1. Quantitative evaluation of the characteristics and matching relationships of multi-type geological heterogeneous data

[0064] The system analyzes the characteristics of multidisciplinary data such as well logging, drilling, geophysical exploration, and geology, studies the data representation connotation of geological knowledge in heterogeneous data, investigates the multi-scale matching relationship of multidisciplinary data at different resolutions, constructs a heterogeneous data matching method under a multi-scale framework, and builds databases (knowledge bases) of lithology, reservoir, reservoir framework, and reservoir physical properties based on reservoir target geology as models.

[0065] Taking reservoir targets as the research object, and combining geological knowledge, we conduct data feature analysis of well logging, petrography, and seismic attributes. Targeting reservoir rock characteristics, we analyze reservoir fabrication, rock composition, rock structure, and mineral composition, and conduct in-depth mining of well logging information for complex geological reservoirs and interlayers. We analyze the correlation and variation patterns of electrical properties in the reservoir, and perform principal component analysis (PCA) and curve optimization and correlation analysis based on long short-term memory neural networks on a large amount of well logging data to achieve the optimization and construction of lithofacies well logging data.

[0066] For different reservoir types, the main structure of complex lithologies is determined through rock physical data. Based on effective medium theory, adaptive theory, contact theory, anisotropy model, etc., a series of steps are taken to establish the rock physical relationship between elastic parameters and physical property parameters. Based on geological understanding, quantitative evaluation formulas for various parameters are established through extensive statistical analysis. Combined with detailed analysis of the development area, the matching relationship between macroscopic geological elements and various reservoir parameters in the study area is established through the organic integration of the two, forming a quantitative evaluation formula.

[0067] 2. Intelligent identification of logging lithology based on the combination of class-based disequilibrium classification and deep learning

[0068] With the continuous development of the exploration field, the increase in measurement data has significantly increased the workload of traditional manual well logging lithology identification. Conventional machine learning methods have, to some extent, achieved intelligent well logging interpretation. However, well logging lithology identification or classification problems are characterized by a large number of classes, insufficient training samples, and uncontrollable class imbalance. Conventional machine learning methods have low prediction accuracy or even show inapplicability in this problem. Therefore, it is necessary to study a well logging lithology intelligent identification method based on the combination of class imbalance removal and deep learning. The training sample imbalance removal process adopts the MAHAKIL oversampling method. This method generates new samples for the minority class by simulating the reproduction process in genetics, ensuring the diversity of new samples to the greatest extent while removing imbalance. The new samples after imbalance are used for deep learning to obtain the complex response relationship between well logging response values ​​and well logging lithology, thereby achieving more accurate intelligent well logging lithology identification. The establishment of a deep learning-based intelligent identification model for well logging lithology is based on learning samples composed of existing well logging lithology interpretation results. These samples consist of data pairs composed of various well logging response values ​​and their corresponding lithology labels. The quantity and quality of these samples largely determine the scale of the deep learning model that can be constructed and its final prediction accuracy. During training, the learning samples are divided into a training set and a validation set. The training set is used for model training, while the validation set is used for hyperparameter tuning and model optimization. In actual prediction, a deep learning network of appropriate size is established based on the number of learning samples, and the network is trained. Finally, the optimized deep learning model is used for intelligent identification of well logging lithology.

[0069] 3. Deep Learning-Based 3D Intelligent Interpolation

[0070] For different reservoir parameters, including velocity, density, porosity, clay content, fluid saturation, or intelligently identified well logging lithology data, a deep learning network structure with spatiotemporal characteristics is studied. The optimized data is regularized using a random forest regression algorithm to regularize heterogeneous data, providing high-quality input data for 3D intelligent interpolation. Based on the construction of triangulation interpolation grids for reservoir parameters such as velocity, density, porosity, clay content, and fluid saturation, or well logging lithology, a 3D intelligent interpolation method based on a CNN convolutional neural network is used to complete the 3D intelligent interpolation.

[0071] Example 1:

[0072] In a specific embodiment 1 of the present invention, the deep learning-based three-dimensional intelligent interpolation method of the present invention includes the following steps:

[0073] Step 1: Quantitative evaluation of geological heterogeneity data characteristics and matching relationships of the lower section of Sha-3 turbidite.

[0074] Step 1.1: Input Data Preparation

[0075] The collected geological heterogeneity data of the lower section of the Shahejie Formation turbidite was imported into the database, and preprocessing was performed on the various information, including data cleaning and transformation. This geological heterogeneity data was then organized according to three-dimensional coordinates for the next matching step.

[0076] Step 1.2: Multidisciplinary Data Analysis and Knowledge Base Construction

[0077] This study comprehensively analyzes the characteristics of data from multiple disciplines, including well logging, drilling, geophysical exploration, and geology, to investigate the representational meaning of heterogeneous data. A knowledge base is constructed based on reservoir target geology, categorized into lithology, reservoir, reservoir framework, and reservoir physical properties.

[0078] Step 1.3: Optimization and Construction of Lithofacies Logging Data for Turbidite Reservoirs

[0079] Well logging data analysis of turbidite reservoirs, mudstone layers, and calcareous mudstone layers is used to identify correlations between reservoirs and multiple disciplines. Principal component analysis (PCA) and curve optimization and correlation analysis based on long short-term memory neural networks are conducted on a large amount of well logging data to achieve the optimization and construction of lithofacies well logging data.

[0080] Step 1.4: Establish the rock physical relationship between elastic parameters and reservoir physical property parameters.

[0081] Based on the effective medium theory, adaptive theory, contact theory, anisotropic model, etc., the rock physical relationship between the elastic parameters and physical property parameters of turbidites is established.

[0082] Step 1.5: Establish quantitative evaluation formulas for geological elements and lithological and physical parameters of turbidite reservoirs.

[0083] Based on extensive statistical analysis to establish quantitative evaluation formulas for various parameters, and with detailed analysis of the development area, the two are organically combined to establish a matching relationship between macroscopic geological elements and turbidite reservoir parameters in the study area, thus forming a quantitative evaluation formula.

[0084] Step 2: Intelligent identification of logging lithology based on the combination of class-based disequilibrium reduction and deep learning

[0085] Step 2.1: Establishing Learning Samples

[0086] Learning samples are established based on well logging response curves and well string interpretation results. Each sample includes well logging data such as sonic transit time (AC), bulk density (DEN), compensated neutron (CNL), natural gamma (GR), deep inductive resistivity (RILD), depth, and interpretation conclusions (turbidite type).

[0087] Step 2.2: Data de-imbalance processing based on the MAHAKIL method

[0088] There are three main methods to address class imbalance: undersampling, oversampling, and threshold shifting. Undersampling involves sampling fewer samples from the majority class, addressing the core issue of preventing information loss due to neglecting some samples. Oversampling and threshold shifting both add samples to the minority class set to achieve dynamic balance with the majority class set, maximizing sample accuracy. Threshold shifting, however, doesn't mechanically equalize the number of minority and majority class samples; instead, it makes classification easier and more accurate. The Mahakil oversampling method, by simulating the reproductive process in genetics, generates new samples for the minority class, eliminating imbalance while maximizing the diversity of new samples. One of Mahakil's goals is to reduce the percentage of fault-prone individuals (Pfp). Its process consists of three steps:

[0089] ① Separate the minority class samples from the dataset that needs to be processed, denoted as Cmin. For each minority class sample in this class, calculate its Mahalanobis distance.

[0090] ② Sort Cmin according to Mahalanobis distance, and divide Cmin in half from the median of the sequence number, denoted as Cmin1 and Cmin2 respectively.

[0091] ③ Select data samples from Cmin1 and Cmin2 in pairs, calculate the mean of the pair, and use it as the new sample to generate.

[0092] If the number of new samples generated in one round is insufficient to meet the requirements of Pfp, the newly generated samples are then used to breed with samples from Cmin1 and Cmin2. If this is still insufficient, the two newly generated samples are then used to breed with samples from their respective classes, and so on. A flowchart of the algorithm is shown below. Figure 2 As shown in the figure, the gray dots represent newly generated samples in each round.

[0093] Step 2.3: Learning Sample Grouping

[0094] The training set is used for model training, the validation set is used for model hyperparameter tuning, and the validation set is used for model optimization.

[0095] Step 2.4: Build a deep learning network of appropriate size based on the number of learning samples and train the network.

[0096] Step 2.5: Use the optimized deep learning model to perform well logging lithology or intelligent lithology identification for unknown wells.

[0097] A well logging lithology intelligent identification network structure based on the combination of class-based disequilibrium reduction and deep learning is as follows: Figure 3 As shown.

[0098] Step 3: Deep Learning-Based 3D Intelligent Interpolation

[0099] Step 3.1: Input Data Preparation

[0100] Use different reservoir parameters of turbidites, including velocity, density, porosity, clay content, fluid saturation, or the logging lithology intelligently identified in step 2, as input data.

[0101] Step 3.2: Regularization of input data

[0102] The random forest regression algorithm is mainly used to regularize the input data.

[0103] Step 3.3: Input the structural interpretation horizons of different vertical sand layers to determine the three-dimensional volume data interpolation framework model.

[0104] Step 3.4: Piecewise interpolation based on Delaunay triangulation, which consists of three steps:

[0105] ① Take different reservoir parameters or intelligently identified logging lithology data as input, construct a large triangle containing all scattered points, and put them into a triangle linked list.

[0106] ② Insert the scattered points in the point set one by one, find the circumcircle of the triangle in the triangle linked list, and the triangle containing the insertion point (called the influence triangle of the point). Delete the common edges of the influence triangles, and connect all the vertices of the triangles that the insertion point influences to complete the insertion of a point in the Delaunay triangle linked list.

[0107] ③ Optimize the newly formed triangles locally according to the optimization criteria, and put the newly formed triangles into the Delaunay triangle linked list.

[0108] Step 3.5: CNN-based 3D intelligent interpolation

[0109] ① Extraction of training samples: The Delaunay triangle obtained from piecewise interpolation in step 3.4 is used as the effective label, and the earthquake amplitude values ​​within the triangle range are used as the sample feature values.

[0110] ② Network Architecture: A 100-layer Unet network was established. The first 50 layers were downsampled. To prevent the loss of edge information from the seismic data during sampling, padding was added to the input data before convolution. This mitigated the drawback of information loss, or more accurately, the relatively small role of corner or image edge information. In the last 50 layers, upsampling was performed to obtain more low-frequency information by merging the downsampled low-frequency information with the upsampled high-frequency information.

[0111] ③ Training the model: Divide the sample data into an 8:2 ratio, with 80% used for training and 20% for testing.

[0112] ④ Model output: When the loss function of the trained model meets the requirements, the model is output.

[0113] ⑤ Model application: Extend the model to three-dimensional space to obtain three-dimensional turbidite lithology.

[0114] Example 2:

[0115] In a specific embodiment 2 of the present invention, the deep learning-based three-dimensional intelligent interpolation method of the present invention includes the following steps:

[0116] Step 1: Quantitative evaluation of geological heterogeneity data characteristics and matching relationships of the upper section of the Sha-4 sandstone and conglomerate mass.

[0117] Step 1.1: Input Data Preparation

[0118] The collected geological heterogeneity data of the upper section of the Sha-4 sandstone and conglomerate mass was imported into the database, and preprocessing was performed on the various information, including data cleaning and transformation. This geological heterogeneity data was then organized according to three-dimensional coordinates for the next matching step.

[0119] Step 1.2: Multidisciplinary Data Analysis and Knowledge Base Construction

[0120] This study comprehensively analyzes the characteristics of data from multiple disciplines, including well logging, drilling, geophysical exploration, and geology, to investigate the connotation of heterogeneous data representation. A knowledge base is constructed based on the target geology of sandstone and conglomerate reservoirs in steep slope zones, categorized into lithology, reservoir, reservoir framework, and reservoir physical properties.

[0121] Step 1.3: Optimization and Construction of Lithofacies Logging Data for Sandstone and Conglomerate Reservoirs

[0122] Well logging data analysis of sandstone and conglomerate reservoirs and mudstone interlayers is used to identify correlations between reservoirs and multiple disciplines. Principal component analysis (PCA) and curve optimization and correlation analysis based on long short-term memory neural networks are performed on a large amount of well logging data to achieve the optimization and construction of lithofacies well logging data.

[0123] Step 1.4: Establish the rock physical relationship between elastic parameters and reservoir physical property parameters.

[0124] Based on the effective medium theory, adaptive theory, contact theory, anisotropic model, etc., the rock physical relationship between the elastic parameters and physical property parameters of sandstone and conglomerate is established.

[0125] Step 1.5: Establish quantitative evaluation formulas for geological elements and lithological and physical property parameters of sandstone and conglomerate reservoirs.

[0126] Based on extensive statistical analysis to establish quantitative evaluation formulas for various parameters, and with detailed analysis of the development area, the two are organically combined to establish the matching relationship between macroscopic geological elements and sandstone and conglomerate reservoir parameters in the study area, thus forming a quantitative evaluation formula.

[0127] Step 2: Intelligent identification of logging lithology based on the combination of class-based disequilibrium reduction and deep learning

[0128] Step 2.1: Establishing Learning Samples

[0129] Learning samples are established based on well logging response curves and well string interpretation results. Each sample includes well logging data such as sonic transit time (AC), bulk density (DEN), compensated neutron (CNL), natural gamma (GR), deep inductive resistivity (RILD), depth, and interpretation conclusions (sandstone and conglomerate body type).

[0130] Step 2.2: Data de-imbalance processing based on the MAHAKIL method

[0131] There are three main methods to address class imbalance: undersampling, oversampling, and threshold shifting. Undersampling involves sampling fewer samples from the majority class, addressing the core issue of preventing information loss due to neglecting some samples. Oversampling and threshold shifting both add samples to the minority class set to achieve dynamic balance with the majority class set, maximizing sample accuracy. Threshold shifting, however, doesn't mechanically equalize the number of minority and majority class samples; instead, it makes classification easier and more accurate. The Mahakil oversampling method, by simulating the reproductive process in genetics, generates new samples for the minority class, eliminating imbalance while maximizing the diversity of new samples. One of Mahakil's goals is to reduce the percentage of fault-prone individuals (Pfp). Its process consists of three steps:

[0132] ① Separate the minority class samples from the dataset that needs to be processed, denoted as Cmin. For each minority class sample in this class, calculate its Mahalanobis distance.

[0133] ② Sort Cmin according to Mahalanobis distance, and divide Cmin in half from the median of the sequence number, denoted as Cmin1 and Cmin2 respectively.

[0134] ③ Select data samples from Cmin1 and Cmin2 in pairs, calculate the mean of the pair, and use it as the new sample to generate.

[0135] If the number of new samples generated in one round is insufficient to meet the requirements of Pfp, the newly generated samples are then used to breed with samples from Cmin1 and Cmin2. If this is still insufficient, the two newly generated samples are then used to breed with samples from their respective classes, and so on. A flowchart of the algorithm is shown below. Figure 2 As shown in the figure, the gray dots represent newly generated samples in each round.

[0136] Step 2.3: Learning Sample Grouping

[0137] The training set is used for model training, the validation set is used for model hyperparameter tuning, and the validation set is used for model optimization.

[0138] Step 2.4: Build a deep learning network of appropriate size based on the number of learning samples and train the network.

[0139] Step 2.5: Use the optimized deep learning model to perform well logging lithology or intelligent lithology identification on the unknown well conglomerate body.

[0140] A well logging lithology intelligent identification network structure based on the combination of class-based disequilibrium reduction and deep learning is as follows: Figure 3 As shown.

[0141] Step 3: Deep Learning-Based 3D Intelligent Interpolation

[0142] Step 3.1: Input Data Preparation

[0143] Use different reservoir parameters of the sandstone and conglomerate body, including velocity, density, porosity, clay content, fluid saturation, or the logging lithology intelligently identified in step 2, as input data.

[0144] Step 3.2: Regularization of input data

[0145] The random forest regression algorithm is mainly used to regularize the input data.

[0146] Step 3.3: Input the structural interpretation stratigraphic positions of the multi-stage sandstone and conglomerate blocks in the upper vertical direction of Sha-4, and determine the three-dimensional volume data interpolation framework model.

[0147] Step 3.4: Piecewise interpolation based on Delaunay triangulation, which consists of three steps:

[0148] ① Take different reservoir parameters or intelligently identified logging lithology data as input, construct a large triangle containing all scattered points, and put them into a triangle linked list.

[0149] ② Insert the scattered points in the point set one by one, find the circumcircle of the triangle in the triangle linked list, and the triangle containing the insertion point (called the influence triangle of the point). Delete the common edges of the influence triangles, and connect all the vertices of the triangles that the insertion point influences to complete the insertion of a point in the Delaunay triangle linked list.

[0150] ③ Optimize the newly formed triangles locally according to the optimization criteria, and put the newly formed triangles into the Delaunay triangle linked list.

[0151] Step 3.5: CNN-based 3D intelligent interpolation

[0152] ① Extraction of training samples: The Delaunay triangle obtained from piecewise interpolation in step 3.4 is used as the effective label, and the earthquake amplitude values ​​within the triangle range are used as the sample feature values.

[0153] ② Network Architecture: A 200-layer Unet network was established. The first 100 layers were downsampled. To prevent the loss of edge information in the seismic data during sampling, padding was added to the input data before convolution. This mitigated the drawback of information loss, or more accurately, the relatively small role of corner or image edge information. In the last 100 layers, upsampling was performed to obtain more low-frequency information by merging the downsampled low-frequency information with the upsampled high-frequency information.

[0154] ③ Training the model: Divide the sample data into an 8:2 ratio, with 80% used for training and 20% for testing.

[0155] ④ Model output: When the loss function of the trained model meets the requirements, the model is output.

[0156] ⑤ Model application: Extend the model to three-dimensional space to obtain the lithological body of three-dimensional sandstone and conglomerate.

[0157] Example 3:

[0158] In a specific embodiment 3 of the present invention, the deep learning-based 3D intelligent interpolation method of the present invention includes the following steps:

[0159] Step 1: Quantitative Evaluation of Geological Heterogeneity Data Characteristics and Matching Relationships of the Sha-4 Lower-Kongdian Formation Red Beds Step 1.1: Input Data Preparation

[0160] The collected geological heterogeneity data of the Sha-Si-Xia-Kongdian Formation red beds were imported into the database, and preprocessing was performed on various information, including data cleaning and transformation. This geological heterogeneity data was then organized according to three-dimensional coordinates for the next matching step.

[0161] Step 1.2: Multidisciplinary Data Analysis and Knowledge Base Construction

[0162] This study comprehensively analyzes the characteristics of data from multiple disciplines, including well logging, drilling, geophysical exploration, and geology, to investigate the representational connotations of heterogeneous data. A knowledge base is constructed based on the target geology of thin interbedded reservoirs in red beds, categorized into lithology, reservoir, reservoir framework, and reservoir physical properties.

[0163] Step 1.3: Optimization and Construction of Lithofacies Logging Data for Sandstone and Conglomerate Reservoirs

[0164] This study analyzes well logging information from red-bed thin-interbedded reservoirs and complex interlayers such as mudstone and igneous rocks to identify correlations between reservoirs and multiple disciplines. Principal component analysis (PCA) and curve optimization and correlation analysis based on long short-term memory neural networks are performed on a large amount of well logging data to achieve optimal and constructed lithofacies well logging data.

[0165] Step 1.4: Establish the rock physical relationship between elastic parameters and reservoir physical property parameters.

[0166] Based on the effective medium theory, adaptive theory, contact theory, anisotropic model, etc., the rock physical relationship between elastic parameters and physical property parameters of thin interbedded red beds is established.

[0167] Step 1.5: Establish quantitative evaluation formulas for geological elements and lithological and physical property parameters of red-bed thin interbedded reservoirs.

[0168] Based on extensive statistical analysis to establish quantitative evaluation formulas for various parameters, and with detailed analysis of the development area, the two are organically combined to establish the matching relationship between macroscopic geological elements and parameters of red-bed thin interbedded reservoirs in the study area, thus forming a quantitative evaluation formula.

[0169] Step 2: Intelligent identification of logging lithology based on the combination of class-based disequilibrium reduction and deep learning

[0170] Step 2.1: Establishing Learning Samples

[0171] Learning samples are established based on well logging response curves and well string interpretation results. Each sample includes well logging data such as acoustic transit time (AC), bulk density (DEN), compensated neutron (CNL), natural gamma (GR), deep inductive resistivity (RILD), depth, and interpretation conclusions (red layer type).

[0172] Step 2.2: Data de-imbalance processing based on the MAHAKIL method

[0173] There are three main methods to address class imbalance: undersampling, oversampling, and threshold shifting. Undersampling involves sampling fewer samples from the majority class, addressing the core issue of preventing information loss due to neglecting some samples. Oversampling and threshold shifting both add samples to the minority class set to achieve dynamic balance with the majority class set, maximizing sample accuracy. Threshold shifting, however, doesn't mechanically equalize the number of minority and majority class samples; instead, it makes classification easier and more accurate. The Mahakil oversampling method, by simulating the reproductive process in genetics, generates new samples for the minority class, eliminating imbalance while maximizing the diversity of new samples. One of Mahakil's goals is to reduce the percentage of fault-prone individuals (Pfp). Its process consists of three steps:

[0174] ① Separate the minority class samples from the dataset that needs to be processed, denoted as Cmin. For each minority class sample in this class, calculate its Mahalanobis distance.

[0175] ② Sort Cmin according to Mahalanobis distance, and divide Cmin in half from the median of the sequence number, denoted as Cmin1 and Cmin2 respectively.

[0176] ③ Select data samples from Cmin1 and Cmin2 in pairs, calculate the mean of the pair, and use it as the new sample to generate.

[0177] If the number of new samples generated in one round is insufficient to meet the requirements of Pfp, the newly generated samples are then used to breed with samples from Cmin1 and Cmin2. If this is still insufficient, the two newly generated samples are then used to breed with samples from their respective classes, and so on. A flowchart of the algorithm is shown below. Figure 2 As shown in the figure, the gray dots represent newly generated samples in each round.

[0178] Step 2.3: Learning Sample Grouping

[0179] The training set is used for model training, the validation set is used for model hyperparameter tuning, and the validation set is used for model optimization.

[0180] Step 2.4: Build a deep learning network of appropriate size based on the number of learning samples and train the network.

[0181] Step 2.5: Use the optimized deep learning model to perform well logging lithology or intelligent lithology identification on thin interbedded red layers in unknown wells.

[0182] A well logging lithology intelligent identification network structure based on the combination of class-based disequilibrium reduction and deep learning is as follows: Figure 3 As shown.

[0183] Step 3: Deep Learning-Based 3D Intelligent Interpolation

[0184] Step 3.1: Input Data Preparation

[0185] Use different reservoir parameters of the red layer, including: velocity, density, porosity, clay content, fluid saturation, or the logging lithology intelligently identified in step 2, as input data.

[0186] Step 3.2: Regularization of input data

[0187] The random forest regression algorithm is mainly used to regularize the input data.

[0188] Step 3.3: Input the structural interpretation horizon of the vertical multi-sand layer group of Sha-4-Kongdian, and determine the three-dimensional volume data interpolation framework model.

[0189] Step 3.4: Piecewise interpolation based on Delaunay triangulation, which consists of three steps:

[0190] ① Take different reservoir parameters or intelligently identified logging lithology data as input, construct a large triangle containing all scattered points, and put them into a triangle linked list.

[0191] ② Insert the scattered points in the point set one by one, find the circumcircle of the triangle in the triangle linked list, and the triangle containing the insertion point (called the influence triangle of the point). Delete the common edges of the influence triangles, and connect all the vertices of the triangles that the insertion point influences to complete the insertion of a point in the Delaunay triangle linked list.

[0192] ③ Optimize the newly formed triangles locally according to the optimization criteria, and put the newly formed triangles into the Delaunay triangle linked list.

[0193] Step 3.5: CNN-based 3D intelligent interpolation

[0194] ① Extraction of training samples: The Delaunay triangle obtained from piecewise interpolation in step 3.4 is used as the effective label, and the earthquake amplitude values ​​within the triangle range are used as the sample feature values.

[0195] ② Network Architecture: A 240-layer Unet network was established. The first 120 layers were downsampled. To prevent the loss of edge information in the seismic data during sampling, padding was added to the input data before convolution. This mitigated the drawback of information loss, or more accurately, the relatively small role of corner or image edge information. During the last 120 layers of upsampling, to obtain more low-frequency information, the downsampled low-frequency information was combined with the upsampled high-frequency information through channel merging.

[0196] ③ Training the model: Divide the sample data into a 7:3 ratio, with 70% used for training and 30% for testing.

[0197] ④ Model output: When the loss function of the trained model meets the requirements, the model is output.

[0198] ⑤ Model Application: Extend the model to three-dimensional space to obtain three-dimensional red-bed thin interbedded lithological bodies.

[0199] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0200] Except for the technical features described in the specification, all other technologies are known to those skilled in the art.

Claims

1. A three-dimensional intelligent interpolation method based on deep learning, characterized in that, This deep learning-based 3D intelligent interpolation method includes: Step 1: Conduct a quantitative evaluation of the characteristics and matching relationships of various types of geological heterogeneous data; Step 2: Perform intelligent identification of logging lithology based on a combination of class-based non-equilibrium classification and deep learning; Step 3, performing deep learning-based 3D intelligent interpolation includes: Step 3.1: Prepare the input data by using different reservoir parameters, including: velocity, density, porosity, clay content, fluid saturation, or the logging lithology intelligently identified in Step 2, as input data. Step 3.2: Regularization of input data. The random forest regression algorithm is used to regularize the input data. Step 3.3: Input the structural interpretation horizons of different vertical target segments to determine the three-dimensional volume data interpolation framework model; Step 3.4: Performing piecewise interpolation based on Delaunay triangulation includes: ① Use different reservoir parameters or intelligently identified logging lithology data as input to construct a large triangle containing all scattered points and put them into a triangle linked list; ② Insert the scattered points in the point set one by one, find the circumcircle of the triangle in the triangle linked list, the triangle containing the insertion point, delete the common edges that affect the triangle, and connect all the vertices of the triangle that affect the insertion point to complete the insertion of a point in the Delaunay triangle linked list. ③ Optimize the newly formed triangles locally according to the optimization criteria, and put the newly formed triangles into the Delaunay triangle linked list; Step 3.5: Performing CNN-based 3D intelligent interpolation includes: ① Extraction of training samples: The piecewise interpolated Delaunay triangle obtained in step 3.4 is used as the effective label, and the earthquake amplitude value within the triangle range is used as the sample feature value; ②Establish network architecture: Establish a 100-layer Unet network. The first 50 layers are downsampled. During the sampling process, in order to prevent the loss of edge information of some seismic data, a padding operation is added to the input data before convolution. In the last 50 layers, in order to obtain more low-frequency information, the downsampled low-frequency information and the upsampled high-frequency information are channel-merged. ③ Training the model: Divide the sample data in an 8:2 ratio, with 80% used for training and 20% for testing; ④ Model Output: When the loss function of the trained model meets the requirements, the model is output; ⑤ Model application: Extend the model to three-dimensional space to obtain three-dimensional lithological bodies.

2. The deep learning-based three-dimensional intelligent interpolation method according to claim 1, characterized in that, Step 1 includes: Step 1.1: Prepare the input data; Step 1.2: Construct a multidisciplinary data analysis and knowledge base; Step 1.3: Optimize and construct reservoir lithofacies logging data; Step 1.4: Establish the rock physical relationship between elastic parameters and reservoir physical property parameters; Step 1.5: Establish quantitative evaluation formulas for geological elements and reservoir lithology and physical properties.

3. The deep learning-based three-dimensional intelligent interpolation method according to claim 2, characterized in that, In step 1.1, the collected multi-type geological heterogeneous data are imported into the database, and the various information is cleaned and preprocessed by data transformation. The geological heterogeneous data are then categorized according to three-dimensional coordinates for the next matching step.

4. The deep learning-based three-dimensional intelligent interpolation method according to claim 2, characterized in that, In step 1.2, the characteristics of multidisciplinary data from logging, drilling, geophysical exploration, and geology are comprehensively analyzed to study the connotation of heterogeneous data representation; a knowledge base is constructed based on the reservoir target geology as a model, including lithology, reservoir, reservoir framework, and reservoir physical properties.

5. The deep learning-based three-dimensional intelligent interpolation method according to claim 2, characterized in that, In step 1.3, well logging information analysis of geological reservoirs and interlayers is performed to find the correlation between reservoirs and multiple disciplines; principal component analysis and curve optimization and correlation analysis based on long short-term memory neural network are carried out on a large amount of well logging information to realize the optimization and construction of lithofacies well logging data.

6. The deep learning-based three-dimensional intelligent interpolation method according to claim 2, characterized in that, In step 1.4, based on the effective medium theory, adaptive theory, contact theory, and anisotropic model, the rock physics relationship between elastic parameters and physical property parameters is established.

7. The deep learning-based three-dimensional intelligent interpolation method according to claim 2, characterized in that, In step 1.5, based on the quantitative evaluation formulas for various parameters established through extensive statistical analysis, and with the detailed analysis of the development area, the matching relationship between the macro-geological elements of the study area and various reservoir parameters is established through the organic combination of the two, thus forming a quantitative evaluation formula.

8. The deep learning-based three-dimensional intelligent interpolation method according to claim 1, characterized in that, Step 2 includes: Step 2.1: Create learning samples; Step 2.2: Perform data de-imbalance processing based on the MAHAKIL method; Step 2.3: Group the learning samples. The training set is used for model training, the validation set is used for model hyperparameter tuning, and the validation set is used for model optimization. Step 2.4: Build a deep learning network of appropriate size based on the number of learning samples and train the network; Step 2.5: Use the optimized deep learning model to perform well logging lithology or intelligent lithology identification for unknown wells.

9. The deep learning-based three-dimensional intelligent interpolation method according to claim 8, characterized in that, In step 2.1, a learning sample is established based on the logging response curve and the well string interpretation results. The logging data of each sample includes sonic transit time, bulk density, compensated neutron, natural gamma, deep induction resistivity, depth, and interpretation conclusion.

10. The deep learning-based three-dimensional intelligent interpolation method according to claim 8, characterized in that, In step 2.2, the MAHAKIL oversampling method is used to generate new samples for the minority class by simulating the reproductive process in genetics. This method removes imbalances while maximizing the diversity of the new samples. The process consists of three steps: ① Separate the minority class samples from the dataset that needs to be processed, denoted as Cmin. For each minority class sample in this class, calculate its Mahalanobis distance. ② Sort Cmin according to Mahalanobis distance, and divide Cmin in half from the median of the sequence number, denoted as Cmin1 and Cmin2 respectively; ③ Select data samples from Cmin1 and Cmin2 in pairs, calculate the mean of the pair, and use it as the new sample to generate; If the number of new samples generated in one round is insufficient to meet the requirement of the highest probability, the newly generated samples and samples in Cmin1 and Cmin2 will continue to reproduce. If it is still insufficient, the two newly generated samples will then reproduce with samples in their respective classes, and so on.

Citation Information

Patent Citations

  • Multi-scale well logging curve automatic identification method based on deep learning

    CN110619353A

  • Three-dimensional complex geologic model label manufacturing method suitable for machine learning algorithm

    CN111815773A