Geology table data enhancement method, device and system based on deep learning and storage medium

Through an improved conditional generation adversarial network (ICG-GAN) processing geotechnical tabular data, using continuous features as conditional input and combined with a supervised classifier, the problems of scarcity and multi-scale features of geotechnical tabular data are solved. The generated data significantly improves prediction performance under multiple classifiers.

CN120277557AActive Publication Date: 2025-07-08INSTITUTE OF GEOLOGY AND GEOPHYSICS CHINESE ACADEMY OF SCIENCES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510766148.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Geological tabular data is scarce and has multi-scale characteristics and strong correlation. Existing data amplification methods are difficult to effectively maintain the authenticity and applicability of the data, especially when dealing with mixed tabular data of discrete and continuous variables.

Method used

An improved conditional generation adversarial network (ICG-GAN) is used to generate simulated continuous feature data by taking the continuous features of geologic data as conditional input, combining the supervisory classifier and discriminator, and completing discrete labels through anti-normalization, and building a multi-dimensional evaluation system to evaluate the effectiveness of generated data.

Benefits of technology

The generated data distribution is closer to the original data, significantly improving the prediction performance under multiple classifiers, improving the diversity and authenticity of the data, and is suitable for tasks such as geological maps and mineral resource prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277557A_ABST
    Figure CN120277557A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of resource exploration, geological mapping and environment monitoring, and discloses a geoscience table data enhancement method, device and system based on deep learning and a storage medium. According to the method, continuous feature vectors in a data set serve as condition input, and physical attributes and multi-scale relevance of the continuous feature vectors are reserved; and a classification voter based on the random forest, the SVM and the XGBoost is constructed, prediction and complementation of discrete features are realized, and the problem of insufficient discrete labels of small samples is effectively solved. And a multi-dimensional evaluation system is also constructed for evaluating the model performance of the system. Experimental results taking multiple groups of rock core analysis data as examples show that compared with a current optimal CTGAN model, data distribution generated by ICG-GAN is closer to original data, and performance is remarkably improved on six application indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of resource exploration, geological mapping and environmental monitoring, and in particular to a method, device, system and storage medium for enhancing geological table data based on deep learning. Background Art

[0002] The widespread application of artificial intelligence and big data technologies in the fields of mineral resource evaluation, oil and gas exploration, seismic data interpretation and imaging, and environmental monitoring has provided a new analytical paradigm and technical path for geological research. However, due to constraints such as exploration costs and geological conditions, obtaining sample data with clear geological attributes or feature identifiers (i.e., labeled data) often requires manual annotation, which is inefficient and leads to a shortage of labeled data, which seriously restricts the generalization ability and prediction accuracy of the model. Therefore, how to effectively expand the geological labeled data set and improve the performance and stability of data-driven models has become a key scientific and engineering challenge that needs to be solved urgently.

[0003] Geological data mainly include two types: image data and tabular data. The present invention focuses on the amplification of tabular data. Tabular data is usually stored in a matrix form, with rows representing different samples and columns corresponding to various types of geological features, including both discrete features (such as lithology types, geological classification labels) and continuous features (such as physical parameters, geochemical element content). The notable feature of geological tabular data is that the sample size is usually small, and it has multi-scale characteristics and strong correlation (including spatial correlation and feature correlation).

[0004] In response to the scarcity of geological table data, existing technologies have proposed a variety of data augmentation strategies. For example, new samples are created near the center of a known mineral deposit through location perturbation; some technologies have proposed a window-based augmentation method (WBDA) to generate more geologically constrained label samples by gradually subdividing the initial window. In the field of seismic data, a translation-based augmentation strategy is adopted to generate additional seismic characteristics and porosity label data by moving a fixed-size sampling window along the well trajectory. Some scholars have proposed a random reduction and increase (RDI) algorithm, combined with a physical constraint closed-loop (PC-CL) strategy, to improve the accuracy of seismic impedance inversion in a semi-supervised manner. These methods maintain the spatial distribution characteristics of geological data to a certain extent, but are still limited in improving sample diversity and have failed to effectively solve the problem of data imbalance.

[0005] To address the issue of data imbalance, existing technologies have introduced augmentation methods applicable to imbalanced datasets. For example, the ADASYN method is used to generate new synthetic samples near minority-class samples, effectively balancing the lithology distribution in the drilling dataset. To address the overfitting problem of the coal seam wettability prediction model, researchers have adopted the SMOTE method to augment small-sample imbalanced datasets. However, such traditional data augmentation methods are difficult to maintain the strong correlation of geoscience data, and their interpolation strategies are prone to introducing noise, reducing the authenticity and applicability of the generated data.

[0006] With the continuous development of machine learning and deep learning technologies, various new data generation strategies have emerged. For example, an improved Variational Autoencoder (VAE) model is used to model the random noise in geophysical data, thereby generating simulated data consistent with the characteristics of real noise. Based on the Improved Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (ICEEMDAN) method and combined with the Pearson correlation coefficient screening strategy, more representative training samples are constructed. Existing technologies have proposed the Geo-TabGAN (Geological Tabular data Generative Adversarial Network) model for geoscience tabular data augmentation, and through the local augmentation strategy, effectively alleviated the problem of uneven distribution of generated label samples. These deep learning-based generation models have shown significant advantages in enhancing the flexibility and high fidelity of data augmentation, but still face certain challenges when dealing with multi-scale features, especially mixed tabular data containing discrete and continuous variables.

[0007] With the development of machine learning and deep learning technologies, various new data generation strategies have emerged continuously. For example, an improved variational autoencoder (VAE) model is used to model the random noise in geophysical data to generate simulated data consistent with the characteristics of real noise. Based on the improved complete ensemble empirical mode decomposition (ICEEMDAN) method and combined with the Pearson correlation coefficient screening strategy, more representative training samples are constructed. A controllable geophysical model generation method is constructed using StyleGAN2 and StyleGAN2-ADA, and the generation results cover complex geological structures (such as faults, sedimentary bedding, etc.), providing an effective data augmentation means for geophysical inversion tasks. Some technologies have proposed a semi-supervised seismic facies classification framework based on conditional GAN. By synthesizing labeled seismic facies images through LoGANv2, the classification accuracy under small sample conditions is significantly improved, verifying the practicality of generative adversarial networks in geoscience image data augmentation. Some scholars have proposed the Geo-TabGAN model for geoscience tabular data augmentation, and effectively alleviated the problem of uneven distribution of generated label samples through local augmentation strategies. These methods based on deep generative models show significant advantages in improving the flexibility and high fidelity of data augmentation, providing diverse solutions to solve the problem of data scarcity in the geoscience field. However, challenges still exist in dealing with multi-scale features, especially mixed tabular data containing discrete and continuous variables.

[0008] In recent years, generative adversarial networks (GANs) have shown advantages in dealing with the relationship between discrete and continuous features. For example, by introducing a classifier in discriminator training, the problem of generating mixed tabular data is effectively solved; CTGAN applies conditional generative adversarial networks to tabular data augmentation, improving the authenticity of generated data. However, there are still limitations when these methods are directly applied to geoscience data.

[0009] There is no large-scale standardized dataset for geoscience tabular data. Geoscience tabular data usually has characteristics such as small sample size, significant multi-scale features, strong spatial correlation and inter-feature correlation. Therefore, there is an urgent need to develop data generation strategies specifically for geoscience tabular data. Summary of the Invention

[0010] The present invention is provided to solve the above problems existing in the prior art. Therefore, there is a need for a method, device, system, and storage medium for enhancing geoscience tabular data based on deep learning. By taking the continuous feature vectors in the dataset as conditional inputs, their physical properties and multi-scale correlations are retained; and a classification voting mechanism based on random forest, SVM, and XGBoost is constructed to realize the prediction and completion of discrete features, effectively addressing the problem of insufficient discrete labels in small samples. To systematically evaluate the model performance, the present invention constructs a multi-dimensional evaluation system. Experimental results using multiple groups of core analysis data show that compared with the current optimal CTGAN model, the data distribution generated by ICG-GAN is closer to the original data, and significant performance improvements are achieved in six application metrics.

[0011] According to the first aspect of the present invention, there is provided a method for enhancing geoscience tabular data based on deep learning, the method comprising: Establishing a data enhancement model; wherein, the data enhancement model includes a generator, a supervised classifier, and a discriminator; the generator responds to an input random noise vector and conditional information, and outputs simulated continuous feature data, the conditional information being a vector obtained by preprocessing the continuous variables of the geoscience tabular data, the continuous variables of the geoscience tabular data including magnetic susceptibility, apparent resistivity, and element content, the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels, forms a sample pair with the simulated continuous feature data and its corresponding class labels, and through inverse normalization to the original scale, realizes the process of complementing discrete labels to obtain generated samples; the discriminator is used to output a probability estimate of the data source according to real data samples or generated samples and conditional information; Training the data enhancement model to obtain a trained data enhancement model; Constructing a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects: data similarity and the performance of downstream prediction tasks; wherein, the augmented data is obtained based on the trained data enhancement model.

[0012] Furthermore, the supervised classifier includes multiple base classifiers, and the manner in which the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels includes: Each base classifier outputs a probability distribution for all classes for the input simulated continuous feature data; Based on the output of each base classifier, the class label is determined.

[0013] Furthermore, based on the output of each base classifier, the class label is determined through the following formula: ; In the formula, Represents the class label, is the number of classifiers participating in the voting, represents the th classifier's predicted probability for class X represents the input sample, k represents the number of classes.

[0014] Furthermore, when training the data augmentation model, the binary cross-entropy loss function is used to model the adversarial objective between the generator and the discriminator. The learning rate is decayed to 90% of the initial value every set number of training epochs. In each mini-batch, the discriminator is updated first, and then the generator is updated to maintain the dynamic balance of the training process. The weight decay parameter is set and L2 regularization is introduced to suppress the overfitting of the discriminator. The loss value is recorded every set number of training epochs to monitor the training process.

[0015] Furthermore, the multi-dimensional evaluation system includes two dimensions: qualitative evaluation and quantitative evaluation.

[0016] Furthermore, the qualitative evaluation includes: using a single-feature distribution histogram and a two-feature group marginal plot for visual analysis; wherein, the single-feature distribution histogram intuitively presents the similarities and differences in the feature distributions of different data sets; the two-feature group marginal plot includes a central main plot, an upper box plot, and a right-side normal distribution curve. The central main plot shows the scatter distribution of two key features and their linear regression fitting curve, which is used to characterize the feature correlation and distribution law of different data sets; the upper box plot shows the distribution characteristics, dispersion degree, and potential outliers of the abscissa feature to quickly identify the median, interquartile range, and distribution symmetry of the data; the right-side normal distribution curve describes the central tendency and dispersion degree of the ordinate feature.

[0017] Furthermore, the quantitative evaluation includes: using multiple evaluation metrics to compare the performance of the original data set and the augmented data set on multiple classifiers; the multiple evaluation metrics include a confusion matrix, accuracy, precision, recall, F1-score, Matthews correlation coefficient, and the area under the ROC curve.

[0018] According to the second technical solution of the present invention, a geoscience tabular data augmentation device based on deep learning is provided. The device includes a processor, and the processor is configured to: The model building module is configured to build a data augmentation model; wherein, the data augmentation model includes a generator, a supervised classifier, and a discriminator; the generator outputs simulated continuous feature data in response to an input random noise vector and conditional information, and the conditional information is a vector obtained by preprocessing the continuous variables of the geoscience tabular data, and the continuous variables of the geoscience tabular data include magnetic susceptibility, apparent resistivity, and element content; the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels, forms a sample pair from the simulated continuous feature data and its corresponding class labels, and realizes the process of complementing discrete labels by inverse normalization to the original scale to obtain generated samples; the discriminator is used to output a probability estimate of the data source according to real data samples or generated samples and conditional information. The model training module is configured to train the data augmentation model to obtain a trained data augmentation model. The model evaluation module is configured to construct a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects: data similarity and downstream prediction task performance; wherein, the augmented data is obtained based on the trained data augmentation model.

[0019] According to the third technical solution of the present invention, a geoscience tabular data augmentation system based on deep learning is provided, and the system includes: a memory for storing a computer program; a processor for executing the computer program to implement the method as described above.

[0020] According to the fourth technical solution of the present invention, a non-transitory computer-readable storage medium storing instructions is provided, and when the instructions are executed by a processor, the method as described above is executed.

[0021] According to the deep learning-based geoscience tabular data augmentation method, device, system, and storage medium of each solution of the present invention, it has at least the following technical effects: 1) The present invention proposes a tabular data augmentation scheme based on an improved conditional generative adversarial network (ICG-GAN) for small-sample and feature-complex-coupled geoscience data. Compared with existing tabular data generation models (such as CTGAN), ICG-GAN has been optimized in network structure and conditional setting, and is more adaptable to the multi-source nature, feature correlation, and physical property constraints of geoscience data. Through verification experiments on core analysis data, the results show that this method is closer to real data in statistical distribution and significantly improves in downstream prediction tasks of various classifiers (such as balanced random forest, random forest, support vector machine (SVM), and XGBoost). This improvement is not only reflected in single indicators such as accuracy, precision, recall, F1 score, and Matthews correlation coefficient (MCC), but also in a multi-dimensional, qualitative and quantitative combined comprehensive evaluation system.

[0022] 2) The advantage of ICG-GAN lies in making full use of the physical constraints and feature correlations in geoscience data, and it can effectively generate synthetic data with real statistical characteristics under the condition of small samples. This characteristic is crucial for geoscience tasks such as geological mapping, mineral resource prediction, and environmental monitoring, where the data demand is large and the distribution is complex. By improving the data volume and quality, ICG-GAN creates conditions for constructing a robust and well-generalized data-driven model, enabling researchers to obtain reliable and practical prediction results even in the case of data scarcity. Description of the Drawings

[0023] Figure 1 Shows the network structure diagram of the data enhancement model ICG-GAN according to an embodiment of the present invention; Figure 2 Shows the histogram comparison schematic diagram of each feature according to an embodiment of the present invention, including the original data set, the data set generated by CTGAN, and the data set generated by ICG-GAN; Figure 3 Shows the group marginal plots of four groups of feature pairs for three data sets according to an embodiment of the present invention: (a) and (b) show the geophysical parameter feature pairs; (c) and (d) show the geochemical parameter feature pairs. In each group of group marginal plots, the main plot is a scatter distribution plot with two key features as the coordinate axes, and a linear regression fitting curve is superimposed. Above the main plot is the box plot of the horizontal axis feature, and on the right is the probability distribution curve (normal distribution fitting result) of the vertical axis feature; Figure 4 Shows the performance comparison results of the generated data and the original data under three mainstream classification models according to an embodiment of the present invention: (a1–a3) are for Balanced Random Forest; (b1–b3) are for Random Forest; (c1–c3) are for SVM; (d1–d3) are for XGBoost. In each group, the first group of data is the original data, the second group of data is the data generated by CTGAN, and the third group of data is the data generated by the method proposed in the present invention; Figure 5 Shows the radar chart according to an embodiment of the present invention, demonstrating the performance of three data sets under different classification models: (a) Balanced Random Forest; (b) Random Forest; (c) SVM; (d) XGBoost; Figure 6 Shows the ROC-AUC curve according to an embodiment of the present invention, comparing the prediction performance of three data sets under four classification models. Detailed Embodiment

[0024] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0025] In recent years, the application of artificial intelligence in the field of earth science has become increasingly widespread, but the scarcity of labeled data severely restricts its application effect. Although the existing data augmentation method based on conditional generative adversarial network (CGAN) has achieved successful applications in fields such as finance and medicine, it is difficult to be directly applied to the earth science field because it fails to fully consider the multi-scale characteristics and strong correlation of earth science data.

[0026] Based on this, the embodiment of the present invention provides a method for enhancing earth science tabular data based on deep learning. This method designs an improved conditional generative adversarial network model, namely the data augmentation model (Improved Conditional Generative Adversarial Network, ICG-GAN), which is specifically used for the augmentation of earth science tabular data. This method includes the following three key designs: (1) Taking continuous features as conditional inputs to retain the physical properties and internal correlations of earth science data; (2) Introducing a classification voting mechanism based on the "ensemble idea" to generate more consistent discrete labels for each group of continuous features, improving the label reliability of the generated data; (3) Constructing a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects: data similarity and the performance of downstream prediction tasks.

[0027] Specifically, the method includes the following steps: S100. Establish a data augmentation model; wherein, the data augmentation model includes a generator, a supervised classifier, and a discriminator; the generator responds to the input random noise vector and conditional information, and outputs simulated continuous feature data. The conditional information is a vector obtained by preprocessing the continuous variables of the earth science tabular data. The continuous variables of the earth science tabular data include magnetic susceptibility, apparent resistivity, and element content. The supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels, forms a sample pair with the simulated continuous feature data and its corresponding class labels, and through anti-normalization to the original scale, realizes the process of complementing discrete labels to obtain generated samples; the discriminator is used to output the probability estimate of the data source according to the real data sample or the generated sample and the conditional information.

[0028] Earth science tabular data usually consists of two types of variables: continuous variables and discrete variables . Among them, in this embodiment, a data set containing samples, and each sample has features is defined as a continuous variable; while information representing a category or label is regarded as a discrete variable 。

[0029] Conditional Generative Adversarial Network (cGAN for short) is a type of model that introduces conditional information (such as labels, specific attributes, or context information) on the basis of the traditional GAN (Generative Adversarial Network) framework to guide the generation process. Conditional Generative Adversarial Networks (CGANs) consist of two parts: a generator and a discriminator Generator: The generator receives a noise vector ( is the noise dimension) and conditional information (i.e., the feature vector of the training data), and outputs generated data 。The mathematical expression is: ;

[0030] where are the generator parameters.

[0031] Discriminator: The discriminator receives a continuous variable or the data generated by the generator as well as conditional information and outputs a probability estimate of the data source 。Its mathematical expression is: ;

[0032] where are the discriminator parameters. When the input is the generated data , the discriminator will try to judge it as a fake sample; when the input is real data , the discriminator tries to identify it as a real sample.

[0033] The training objective of CGAN is to make the samples generated by the generator indistinguishable from the real data in the eyes of the discriminator, while the discriminator continuously improves its ability to distinguish between true and false. The optimization objective function of adversarial training: ; where is a min-max game form, G represents the generator, and the goal is to minimize the value function V; D represents the discriminator, and the goal is to maximize the value function V. The two are continuously optimized through adversarial training; E represents the mathematical expectation, which is used to average the values of random variables according to their probability distributions; Pdata (X, C) represents a continuous variable and the conditional information of the joint probability distribution, which describes the probability law of their occurrence in the real scenario; represents the probability output (taking values between 0 and 1, the closer to 1 indicates that the discriminator is more convinced that it is real data) of the discriminator D for the continuous variable combined with the conditional information judged as "real data", is to take the logarithm of this judgment probability, which is used to construct the objective function to make the discriminator as much as possible to judge real data as true; represents the probability distribution of the noise vector Z, which is used for the initialization generation process of the generator; G ( Z , C ) represents the sample generated by the generator G after receiving the noise vector Z and the conditional information ; D ( G ( Z , C ), C ) represents the probability output of the discriminator for the generated sample G ( Z , C ) combined with the conditional information judged as real data, indicating that the discriminator should try its best to judge the generated sample as fake (maximizing this expectation, corresponding to the discriminator distinguishing between true and false; while the generator should minimize the corresponding part of this expectation, making it difficult for the discriminator to distinguish). Through this confrontation, the generator is promoted to generate more and more realistic samples.

[0034] Although conditional generative adversarial networks (CGANs) were initially mainly applied to image data generation, from its theoretical framework, this generation mechanism can be extended to tabular data augmentation. The present invention enables CGANs to effectively handle the generation task of geoscience structured tabular data through three key innovations: optimizing the model structures of the generator and the discriminator, improving the representation method of conditional information, and proposing a discrete feature completion strategy.

[0035] The structure of the improved conditional generative adversarial network (i.e., the data augmentation model ICG - GAN) is as Figure 1 shown, mainly consisting of a generator (Generator) and a discriminator It consists of a generator and a discriminator. The generator contains four fully connected layers. The input is the concatenation of a 100 - dimensional random noise and conditional information, which passes through hidden layers with dimensions of 128, 256, and 512 in sequence. The activation function is ReLU, and finally, it outputs synthetic samples with the same dimension as the conditional information. The discriminator has a similar structure, also consisting of four fully connected layers. The input is the concatenation of sample data and conditional information. The dimensions of the hidden layers are 512, 256, and 128 respectively. The LeakyReLU activation function (with a negative slope of 0.2) is used, and finally, it outputs the probability judgment of whether the sample is real or generated through the Sigmoid function.

[0036] Traditional conditional generative adversarial networks (cGANs) usually use discrete features as conditional information. However, due to the complex feature correlations in geoscience data, using discrete features as conditional inputs is difficult to fully express the internal relationships between samples in multi - scale and highly complex geological environments. Therefore, in this embodiment, the continuous variables of geoscience data are used as conditional information and input into the generative model. It should be particularly noted that the continuous variables are not the original continuous variables (such as magnetic susceptibility, apparent resistivity, element content, etc.), but the "conditional information" obtained after pre - processing operations such as outlier removal, normalization, or standardization. Using continuous features as conditional inputs aims to fully explore the physical and chemical properties and the high correlations between their features in geoscience data, thereby improving the approximation ability of generated samples to the real distribution in the multi - variable space. The specific implementation process is as follows: ;

[0037] This conditional setting enables the generator to refer to the feature distribution and correlations of the original data when generating new samples, thus better retaining the physical feature constraints of geoscience data. Of course, in order to more effectively maintain the physical feature constraints of geoscience data, the present invention introduces regularization strategies such as Batch Normalization and Weight Decay in the generator network to prevent abnormal deviations from occurring in the generated data during training.

[0038] Tabular data usually contains columns of mixed types, i.e., both discrete variables and continuous variables coexist. This mixed characteristic makes it difficult for generative adversarial networks to capture their respective distribution laws and the correlations between variables simultaneously, thus affecting the quality and consistency of the generated data. In addition, in the scenario of small-sample labeled data, if the number of samples of a certain category in the discrete column is too small, the traditional cGAN model with discrete variables as conditional inputs is difficult to effectively learn the statistical features of this category. The present invention uses a classification voting mechanism to complete the discrete label complementation for the continuous feature data generated by cGAN. This process essentially uses a supervised classifier constructed from real samples to infer the "latent category" of the generated samples. Specifically, the generator receives a random noise vector and a conditional information ( being the continuous features of real samples), and outputs simulated continuous feature data , and the calculation process is as follows: ;

[0039] Since the generated data only contains but lacks the corresponding class labels , in this embodiment, a supervised classifier is used to predict its labels. In the classifier design, the soft voting strategy in ensemble learning is adopted, and the prediction results of multiple base classifiers are weighted and averaged according to the class probabilities to finally generate class labels. Suppose each base classifier outputs a probability distribution for all classes for the input sample, i.e., the continuous variable : ;

[0040] Then the final prediction result of the voting classifier for the sample can be expressed as: ;

[0041] wherein, is the number of classifiers participating in the voting, represents the prediction probability of the -th classifier for the class . In this embodiment, three commonly used supervised learning models are selected as base classifiers: Random Forest (RF), Support Vector Machine (SVM) and XGBoost, and they are regarded as equally important, i.e., the voting weights are the same.

[0042] After the voting classifier is trained, the present invention applies it to the continuous feature samples generated by the generator, i.e., the generated data for predicting its corresponding class label thus constructing a generated sample pair with a complete structure (continuous features + discrete labels) , and the calculation process is as shown in Equation (8). This method not only maintains the diversity of generated samples in the feature space but also combines the discriminative ability of the original label information, significantly improving the usability of the generated data.

[0043] ; In the formula,[[]]END]] h represents the trained classifier.

[0044] Finally, the continuous feature part of the generated sample will be inverse-normalized to the original scale according to Equation (9), and the process of complementing the discrete label will be completed.

[0045] ; In the formula, InverseTransformer represents the inverse transformation function / operation,[[]]END]] represents the data normalized to the original scale.

[0046] S200. Train the data augmentation model to obtain a trained data augmentation model.

[0047] During the training process, ICG-GAN uses the Binary Cross Entropy (BCE) loss function to model the adversarial objective between the generator and the discriminator. BCE is widely used in binary classification tasks and can effectively measure the difference in discriminative probabilities between generated samples and real samples, which helps to accelerate the convergence speed of the discriminator and improve the stability of generator training. In terms of the optimizer, the Adam optimizer with an adaptive learning rate mechanism is adopted. This method shows good stability and convergence performance when dealing with non-convex optimization problems and is thus widely used in generative adversarial networks. However, in actual training, traditional GANs are still vulnerable to problems such as unstable generated distributions and mode collapse.

[0048] To improve the training stability and efficiency in the small-sample scenario, the present invention introduces the StepLR learning rate scheduling strategy, which decays the learning rate to 90% of the initial value every 500 epochs to balance the rapid convergence in the initial stage of training and the fine-tuning in the later stage. In each mini-batch, the discriminator is updated first, and then the generator is updated to maintain the dynamic balance of the training process. In addition, the weight decay parameter is set to 1e-4, and L2 regularization is introduced to suppress the overfitting of the discriminator. The mini-batch size of the network is set to 64, and the loss value is recorded every 500 epochs to monitor the training process. Through the collaborative optimization of the network structure and training strategy, ICG-GAN demonstrates good training stability and high-quality sample diversity in the task of generating geoscience tabular data, alleviating the common problems of training instability and mode collapse in traditional conditional generative adversarial networks for small-sample geoscience data.

[0049] S300. Construct a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects: data similarity and the performance of downstream prediction tasks; wherein, the augmented data is obtained based on the trained data augmentation model.

[0050] In this embodiment, a complete evaluation system is constructed to comprehensively evaluate the quality of the generated data from both qualitative and quantitative dimensions (see Table 1).

[0051] The quantitative evaluation verifies the effectiveness of the data augmentation method from the application effect level, mainly by comparing the performance of the original dataset and the augmented dataset on multiple classifiers. In this study, four machine learning classifiers with a wide application basis in the geoscience field are selected: Balanced Random Forest (BRF), Random Forest (RF), Support Vector Machine (SVM), and XGBoost. These classifiers have their own advantages: BRF effectively addresses the class imbalance problem through balanced sampling; RF has strong feature selection and anti-noise capabilities; SVM is suitable for high-dimensional feature spaces and has excellent generalization performance; XGBoost combines the advantages of decision trees and gradient boosting and performs outstandingly in many machine learning competitions. To comprehensively evaluate the model performance, six evaluation indicators listed in Table 1 are adopted. The value ranges of these indicators are all [0,1], and the closer the value is to 1, the better the prediction performance. Through comprehensive comparison of multiple indicators, the improvement effect of data augmentation on the classification performance can be comprehensively evaluated. In addition, k-fold cross-validation is used for model training and testing, effectively reducing the randomness of the evaluation results.

[0052] Table 1 Evaluation indicators of the generated data

[0053] Next, the embodiments of the present invention will fully illustrate the feasibility and progressiveness of the present invention in combination with a specific implementation case.

[0054] Two sets of drilling data are used in this embodiment, both collected from different distribution areas of a certain mining area. The first set of data comes from two scientific research deep boreholes KY15 - 03 - 01 and KY14 - 02 - 01, located between the main mining area and the eastern mining area respectively, with drilling depths of 1762.75 m and 2022.18 m respectively, and a total of 9 characteristic variables are included. The second set of data includes 4 boreholes distributed in different areas, namely TK13 - 4, TK8 - 2, TK13 - 3, and TK7 - 1, with a total of 6 characteristic variables. The physical property parameters and chemical element contents in the data used in this embodiment are the comprehensive processing results after multiple measurements. The label information of the data is jointly completed by experts from fields such as geology, logging, and core analysis. To ensure the traceability of the annotation work, the entire annotation process is detailedly recorded, including time, work area, instrument model, determination personnel, verification personnel, specimen number, borehole depth, core name, reference basis for lithology description, and annotation version, etc., for subsequent verification and revision.

[0055] The first set of core analysis data comes from two scientific research deep boreholes (KY15 - 03 - 01 and KY14 - 02 - 01), and a total of 690 core samples are collected. Each sample contains 9 characteristic variables, including 3 geophysical parameters (magnetic susceptibility, polarizability, and resistivity) and 6 element contents (Mg, Al, Si, Fe, Ni, Nb), and is accompanied by lithology labels. The sample proportion distribution of lithology types is as follows: slate (41.2%), (magnetite) mineralized dolomite (22.6%), magnetite (14.9%), dolomite (2.3%), quartz sandstone (4.2%), and dolomite - slate interbed (14.8%). Some statistical information is shown in Table 2. This data set is selected because it covers typical geoscience information characteristics and has attributes such as multi - source, multi - dimensional, and multi - scale, integrating geophysical and geochemical characteristics, and has strong representativeness and application value.

[0056] Table 2. Partial Core Analysis Table Data Set

[0057] Magnetic susceptibility Polarizability Resistivity Mg Al Si Fe Ni Nb <![CDATA[Class Don't > 3.07 21.8285 122303.345 31824 83474 227998 67984 53 102 1 9.18 62.935 149.77 15868 82472 351379 25930 26 84 1 10.83 155.457 0.415 8801 74532 286398 37598 23 87 1 28.81 53.1315 1005.49 0 19586 456890 18853 15 7 1 10.61 72.337 77.085 84013 8743 35113 118374 18 14 2 80.89 52.6265 3405.36 67864 4129 21154 96099 35 175 2 67.05 57.1405 1820.89 80245 7155 25069 86097 32 284 2 38.46 44.7215 706.45 72617 5335 15095 102921 0 46 2 24009 124.3465 43.23 50336 13581 42638 121809 25 29 3 32801 59.988 0.85 74834 16785 52195 176772 0 152 3 57632 64.749 62.255 54971 13203 40801 182376 0 128 3 36368 121.825 16.815 26742 15355 50794 180548 75 185 3 7.32 20.2655 486.075 36416 6388 35097 188115 32 99 4 2.61 6.669 1268.965 81273 9649 26526 60407 23 26 4 0.87 13.204 5346.04 7862 7911 448785 2409 0 0 5 1.98 61.599 82.515 0 18107 486542 4222 10 7 5 1.06 30.6135 92.665 0 36625 418936 30182 83 172 5 837 32.647 66.15 51320 23906 81470 48334 34 394 6 259 18.0165 104.22 86948 1195 38500 55125 22 197 6 190 12.306 88.115 50791 17649 79167 46792 23 214 6 243 16.8575 1187.22 84261 24043 175799 67281 29 491 6

[0058] The experimental environment configuration of this embodiment is as follows: The hardware platform is a Lenovo P920 graphics workstation, equipped with an Intel Xeon Gold 6226R CPU (2.90 GHz), 64 GB of memory, and an NVIDIA RTX A2000 GPU; the software environment is based on Python 3.8 and the PyTorch deep learning framework. The model uses the Adam optimizer, and the initial learning rates of the generator and discriminator are both set to 0.0002, the batch size is 128, the number of training epochs is 10,000, and the latent space dimension is 100. To improve the model performance, the StepLR learning rate scheduling strategy is introduced, and binary cross-entropy is selected as the loss function. In terms of data preprocessing, the Min-Max normalization method is used to standardize the data, and the distribution consistency between the generated data and the original data is ensured through feature accuracy matching. In the dataset division, the ratio of the training set to the test set in the main experiment is 7:3.

[0059] Figure 2 Figure 4 shows the distribution comparison of nine key features (including magnetic susceptibility, polarizability, resistivity, and element contents (Mg, Al, Si, Fe, Ni, Nb)) among real data, CTGAN-generated data, and ICG-GAN-generated data proposed in this study. Generally speaking, the data generated by ICG-GAN is closer to the real distribution in terms of morphology and range. For most features (such as magnetic susceptibility, polarizability, Mg, Al, Si, Fe, Ni, Nb), ICG-GAN can better reproduce the skewness and kurtosis characteristics of the real data. For example, in terms of the tail behavior and overall morphology of element features such as Al, Fe, Ni, and Nb, ICG-GAN has a higher consistency with the real data. In contrast, the data generated by CTGAN tends to be concentrated and deviate from the real distribution. Especially in the distribution of polarizability and Mg, Al elements, the CTGAN data is overly concentrated in a smaller numerical range. In addition, for some features (such as resistivity), abnormal spikes and unreasonable distributions appear in the distribution of the data generated by CTGAN, while ICG-GAN effectively avoids this situation. Generally speaking, these comparison results show that ICG-GAN shows more excellent performance in terms of the approximation of data distribution and the retention of real sample characteristics.

[0060] Figure 3 The four groups of marginal plots in Figure 5 show the comparison of the original data (red), CTGAN-generated data (blue), and ICG-GAN-generated data (green) in two-dimensional statistics. The following analyzes the four subplots (a), (b), (c), and (d) in Figure 5: Figure 3 The four subplots (a), (b), (c), and (d) in Figure 5 are analyzed as follows: Figure 3(a) shows the marginal plot of magnetic susceptibility and resistivity. It can be seen that the horizontal distribution ranges of the original data and the ICG-GAN generated data are narrower than those of the CTGAN generated data. There are several high-resistivity outliers in the low magnetic susceptibility region of the original data, which is consistent with the ICG-GAN generated data, while the outliers of the CTGAN generated data are more dispersed. By linearly regressing and fitting the resistivity and magnetic susceptibility of different datasets, it is found that the correlation between the ICG-GAN generated data and the original data is closer. The box plot above the main figure shows that the distribution of magnetic susceptibility of the CTGAN generated data is more discrete, with larger differences and more outliers; while the ICG-GAN generated data and the original data are more similar in terms of dispersion and concentration. The normal distribution plot on the right shows that the means of the three datasets are not very different, but the variances of the two generated datasets are both smaller than that of the original data.

[0061] Figure 3 (b) shows the correlation between magnetic susceptibility and polarizability. The original data shows an obvious positive correlation trend in this feature combination, and some data points have a higher degree of dispersion on the magnetic susceptibility axis, showing a larger linear expansion. In contrast, the data generated by CTGAN is significantly different from the original data. Especially in the region of high magnetic susceptibility and low polarizability, there are points in the generated data that do not exist in the original data, resulting in the regression curve of magnetic susceptibility and polarizability deviating from the original relationship. Relatively speaking, the scatter points of the data generated by ICG-GAN are more concentrated in the low magnetic susceptibility and low polarizability interval, and are consistent with the mean value and change trend of the real data. A certain number of high polarizability points can also be reproduced in the high magnetic susceptibility interval, thus maintaining the diversity and authenticity of the overall distribution. In terms of marginal distribution, the box plot of magnetic susceptibility of the ICG-GAN generated data shows a data concentration range and discrete points similar to the original data; the generated polarizability distribution curve is consistent with the overall shape of the original data. However, the polarizability distribution curve generated by CTGAN is significantly shifted, and its box plot of magnetic susceptibility shows a large number of outliers that do not match the real data. The length of the box plot also indicates a higher degree of dispersion of the generated data and larger differences between the data.

[0062] Figure 3(c) shows the relationship between magnesium and silicon. The original data showed a clear negative correlation trend in this feature combination. However, the data generated by CTGAN differed significantly from the original data, especially in the high-silicon region, where discrete points that did not exist in the original data appeared, causing the regression curve of magnesium and silicon to deviate seriously from the original relationship. In addition, the magnetic susceptibility box plot generated by CTGAN shows that the numerical range is quite different from the real data. In contrast, the concentrated distribution area of ​​the data scatter points generated by ICG-GAN is consistent with the real data, especially for the discrete points in the high-magnesium and low-silicon region, and the regression curve trends of the two are consistent. In terms of marginal distribution, the concentrated range and discrete points of the magnetic susceptibility box plot of the data generated by ICG-GAN are similar to the original data.

[0063] Figure 3 (d) shows the distribution relationship between iron and nickel content. The original data show that the iron content covers a wide range of values, while the nickel content shows a certain degree of discreteness in the medium and high value areas. Overall, there are a small number of low iron (Fe) and high nickel (Ni) value points in the data, and the distribution is relatively scattered. Compared with the original data, the data generated by CTGAN shifts toward high nickel in the Fe-Ni plane, resulting in a regression relationship that is significantly higher than the original data. The box plot of Fe content shows that its value range is significantly higher than the original data, and there are a large number of discrete points, indicating that CTGAN introduces more noise data. However, the mean and variance of the normal distribution of nickel content are similar to those of the original data. In contrast, the data generated by ICG-GAN show a more balanced distribution on this feature combination, and its Fe value range is similar to the original data, and the dispersion of nickel is consistent with the original data in different Fe intervals. The box plot of Fe shows a distribution feature that is closer to the real data, and the trend of the regression fitting line is also consistent with the original data. However, in the distribution diagram of nickel, the variance of the data set is different from that of the original data, mainly due to the influence of the data volume.

[0064] Overall, Figure 3 The difference between CTGAN and ICG-GAN in the ability to fit multidimensional spatial data distribution is clearly demonstrated. The results show that ICG-GAN exhibits higher fidelity in the reproduction of multidimensional feature joint distribution, marginal feature distribution and regression relationship, which is close to the statistical characteristics and intrinsic structure of the original data. Therefore, the method proposed in this paper is more suitable for the amplification of geoscience data.

[0065] In order to verify the effect of the generated data set on improving the model performance, this example compares the application effects of the original data set, the data set amplified based on CTGAN, and the data set amplified based on the method of the present invention under different classifiers. This example randomly samples 70% of the original data set as a training set and 30% as a test set for prediction. The amplified training data set is composed of 70% of the original data set combined with the generated data set, and the test set is still the original 30%. Figure 4The confusion matrices under four classifiers (Balanced Random Forest, Random Forest, Support Vector Machine, and XGBoost) are shown. Figure 4 Among them, (a)-(d) correspond to one classifier respectively, where a is Balanced Random Forest, b is Random Forest, c is Support Vector Machine, and d is XGBoost. Subscripts 1, 2, and 3 represent the confusion matrices of the original data, the data amplified based on CTGAN, and the data amplified based on the method of this study respectively. Each column of the confusion matrix represents the predicted class, and the total number of columns is the number of data predicted as this class; each row represents the true class, and the total number of rows is the number of data instances of this class; the values on the diagonal represent the number of true data predicted as this class. It can be seen from the confusion matrices that the models trained on the dataset amplified by ICG-GAN show better performance in all four classifiers. Subsequently, this embodiment further analyzes other classification metrics to more clearly and intuitively show the impact of the generated data on the prediction model (see Figure 5 )

[0066] Figure 5 The performance evaluation results of four classifiers - Balanced Random Forest (Balanced RF), Random Forest (RF), Support Vector Machine (SVM), and XGBoost - on different generated datasets (CTGAN, ICG-GAN, original dataset) are shown. The experiments show that the classifiers trained on the ICG-GAN dataset have significantly improved performance in all evaluation metrics.

[0067] In the quantitative analysis, the classification accuracy of Balanced RF on the ICG-GAN dataset reaches 89.37%, which is 2.41% and 4.35% higher than that of CTGAN (86.96%) and the original dataset (85.02%) respectively. As shown in Figure 5 (a)_ in the figure, the order of model performance improvement is ICG-GAN > CTGAN > original dataset. Compared with CTGAN, the precision, recall, F1-score, and Matthews correlation coefficient (MCC) are improved by 2.35%, 2.41%, 2.35%, and 5.30% respectively. The RF classifier shows significant improvement on the ICG-GAN dataset, and the classification accuracy reaches 96.14%, which is 8.70% and 7.25% higher than that of CTGAN (87.44%) and the original dataset (88.89%) respectively. As shown in Figure 5As shown in Fig. (b), the data generated by CTGAN failed to improve the model performance but instead led to a performance decline. Analysis suggests that the main reason for this phenomenon is that CTGAN uses discrete features as conditional inputs, and the discrete features are unevenly distributed in the original data. During the training process, it is easy to introduce generation biases, resulting in a deviation between the distribution of the generated samples and the real data. In addition, discrete features have obvious limitations in expressing the complex feature correlations and physical properties in geoscience data and are difficult to fully capture the physical and chemical coupling relationships between variables. Therefore, these factors jointly affect the actual performance of the samples generated by CTGAN in the lithology classification task. The accuracy of XGBoost on the ICG-GAN dataset is 94.69%, which is 7.73% and 3.39% higher than that of CTGAN (86.96%) and the original dataset (91.30%) respectively. Other evaluation metrics also show significant advantages. In contrast, although the indicators of SVM have improved on the ICG-GAN dataset, its overall performance is still lower than that of RF and XGBoost.

[0068] From the analysis of algorithm characteristics, Random Forest (RF) and XGBoost perform excellently on the ICG-GAN dataset, mainly due to their advantages in handling complex feature interactions and nonlinear relationships. Balanced Random Forest effectively alleviates the class imbalance problem through sampling strategies. Although Support Vector Machine (SVM) has theoretical advantages in high-dimensional feature spaces, its performance is limited by kernel function selection and parameter optimization. Regardless of the classifier used, the ICG-GAN dataset significantly improves the classifier performance, mainly because it better preserves the statistical characteristics and inter-class relationships of the original data during the data generation process. In summary, the experimental results fully demonstrate the superiority of the dataset generated by ICG-GAN in improving classifier performance.

[0069] Figure 6 Shows the training and test results of the original data, the data generated by CTGAN, and the data generated by ICG-GAN proposed in this study under four classifiers (Balanced Random Forest, Random Forest, Support Vector Machine, and XGBoost), including the comparison of ROC curves and AUC values.

[0070] The result analysis shows that both the original data (Original) and the ICG-GAN generated data proposed by the present invention (ICG-GAN) exhibit excellent classification performance in most cases. The AUC values are generally high, and the curves are closely fitted to the upper left corner, indicating a good balance between the true positive rate (TPR) and the false positive rate (FPR). Specifically, under the Random Forest classifier, the AUC value of the original data is 0.9934, while that of the ICG-GAN generated data is improved to 0.9978, showing an improvement. This indicates that ICG-GAN not only approaches the real data in terms of data distribution but also may provide a more stable decision boundary for the classifier. The Support Vector Machine (SVM) and Extreme Gradient Boosting (XGBoost) classifiers also show this trend: under the SVM classifier, the AUC of the ICG-GAN generated data is 0.9879, exceeding 0.9822 of the original data; under the XGBoost classifier, the AUC of the ICG-GAN generated data is 0.9980, further improved compared to 0.9933 of the original data. Although the Balanced Random Forest classifier has obtained a relatively high AUC (0.9842) on the original data, the ICG-GAN generated data still improves it to 0.9890, further demonstrating the general performance gain of the ICG-GAN generated data.

[0071] Compared with the original data and the ICG-GAN generated data, the AUC of the CTGAN generated data performs poorly under the three classifiers. For example, in the Support Vector Machine (SVM) classifier, the AUC of the CTGAN generated data is only 0.9687, significantly lower than the original data and the ICG-GAN data. In other classifiers such as Random Forest and XGBoost, although the AUC of the CTGAN generated data is still relatively high (0.9927 for Random Forest and 0.9902 for XGBoost), it is always lower than the ICG-GAN and the original data. Only in the Balanced Random Forest model, the AUC of the CTGAN generated data is slightly higher than that of the ICG-GAN, with an increase of only 0.14%.

[0072] Overall, ICG-GAN performs excellently in data synthesis and augmentation, not only generating more realistic distribution features but also providing training samples with high resolution and stability for mainstream classification algorithms. Under various model evaluation metrics, the classification performance of the ICG-GAN generated data reaches or is slightly higher than that of the original data, and is significantly better than the commonly used generation model CTGAN. This further verifies the application value of ICG-GAN in data enhancement and improving classification performance.

[0073] From the above implementation cases, it can be seen that compared with the existing tabular data generation models (such as CTGAN), the ICG-GAN in the present invention optimizes the network structure and condition setting, and is more adaptable to the multi-source nature, feature correlation and physical property constraints of geoscience data. Through the verification experiment of core analysis data, the results show that this method is closer to the real data in statistical distribution and significantly improves in the downstream prediction tasks of various classifiers (such as balanced random forest, random forest, support vector machine (SVM) and XGBoost). This improvement is not only reflected in single indicators such as accuracy, precision, recall, F1 score and Matthews correlation coefficient (MCC), but also reflected in the comprehensive evaluation system that combines multi-dimensions, qualitative and quantitative aspects.

[0074] The embodiment of the present invention also provides a geoscience tabular data augmentation device based on deep learning, and the device includes: A model establishment module, configured to establish a data augmentation model; wherein, the data augmentation model includes a generator, a supervised classifier and a discriminator; the generator responds to an input random noise vector and condition information, and outputs simulated continuous feature data, and the condition information is a vector obtained by preprocessing the continuous variables of the geoscience tabular data, and the continuous variables of the geoscience tabular data include magnetic susceptibility, apparent resistivity and element content, the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels, forms a sample pair with the simulated continuous feature data and its corresponding class labels, and realizes the process of complementing discrete labels by inverse normalization to the original scale to obtain generated samples; the discriminator is used to output a probability estimate of the data source according to real data samples or generated samples and condition information; A model training module, configured to train the data augmentation model to obtain a trained data augmentation model; A model evaluation module, configured to construct a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects of data similarity and downstream prediction task performance; wherein, the augmented data is obtained based on the trained data augmentation model.

[0075] The embodiment of the present invention also provides a geoscience tabular data augmentation system based on deep learning, and the system includes: a memory for storing a computer program; a processor for executing the computer program to implement the method as described in the above embodiment.

[0076] The embodiment of the present invention also provides a non-transitory computer-readable storage medium storing instructions, and when the instructions are executed by a processor, the method as described in the above embodiment is executed.

[0077] The above embodiments are only used to illustrate the present invention and are not intended to limit the present invention. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the present invention. The patent protection scope of the present invention shall be defined by the claims.

Claims

1. A method for enhancing geoscience tabular data based on deep learning, characterized in that, The method includes: Building a data augmentation model; wherein, the data augmentation model includes a generator, a supervised classifier, and a discriminator; the generator responds to an input random noise vector and conditional information, and outputs simulated continuous feature data, the conditional information is a vector obtained by preprocessing continuous variables of geoscience tabular data, the continuous variables of the geoscience tabular data include magnetic susceptibility, apparent resistivity, and element content, the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels, forms a sample pair with the simulated continuous feature data and its corresponding class labels, and through inverse normalization to the original scale, realizes the process of complementing discrete labels to obtain generated samples; the discriminator is used to output a probability estimate of the data source according to real data samples or generated samples and conditional information; Training the data augmentation model to obtain a trained data augmentation model; Constructing a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects of data similarity and downstream prediction task performance; wherein, the augmented data is obtained based on the trained data augmentation model.

2. The method according to claim 1, wherein The supervised classifier includes multiple base classifiers, and the way that the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels includes: Each base classifier outputs a probability distribution for all classes for the input simulated continuous feature data; Determining class labels based on the outputs of each base classifier.

3. The method according to claim 2, wherein Based on the outputs of each base classifier, determine class labels through the following formula: ; Wherein, represents the class label, is the number of classifiers participating in the voting, represents the th classifier's predicted probability for class , X represents the input sample, k represents the number of classes.

4. The method according to claim 1, wherein When training the data augmentation model, a binary cross-entropy loss function is used to model the adversarial objective between the generator and the discriminator. Every set number of training rounds, the learning rate is decayed to 90% of the initial value. In each mini-batch, first update the discriminator, and then update the generator to maintain the dynamic balance of the training process; Set a weight decay parameter and introduce L2 regularization to suppress overfitting of the discriminator, and record the loss value once every set number of training rounds to monitor the training process.

5. The method according to claim 1, characterized in that, The multi-dimensional evaluation system includes two dimensions: qualitative evaluation and quantitative evaluation.

6. The method according to claim 5, wherein The qualitative evaluation includes: using a single-feature distribution histogram and a two-feature group marginal map for visual analysis; wherein, the single-feature distribution histogram intuitively presents the similarities and differences in feature distributions of different data sets; the two-feature group marginal map includes a central main graph, an upper box plot, and a right normal distribution curve. The central main graph shows the scatter distribution of two key features and its linear regression fitting curve, which is used to characterize the feature correlation and distribution law of different data sets; the upper box plot shows the distribution characteristics, dispersion degree, and potential outliers of the abscissa feature to quickly identify the median, interquartile range, and distribution symmetry of the data; the right normal distribution curve describes the central tendency and dispersion degree of the ordinate feature.

7. The method according to claim 5, wherein The quantitative evaluation includes: comparing the performance of the original dataset and the augmented dataset on multiple classifiers using multiple evaluation metrics; the multiple evaluation metrics include confusion matrix, accuracy, precision, recall, F1-score, Matthews correlation coefficient, and area under the ROC curve.

8. A device for enhancing geoscience tabular data based on deep learning, characterized in that, The device includes: A model establishment module, configured to establish a data augmentation model; wherein, the data augmentation model includes a generator, a supervised classifier, and a discriminator; the generator responds to an input random noise vector and conditional information, and outputs simulated continuous feature data, the conditional information is a vector obtained by preprocessing the continuous variables of the geoscience tabular data, the continuous variables of the geoscience tabular data include magnetic susceptibility, apparent resistivity, and element content, the supervised classifier classifies the simulated continuous feature data to obtain corresponding class labels, forms a sample pair of the simulated continuous feature data and its corresponding class labels, and through inverse normalization to the original scale, realizes the process of supplementing discrete labels to obtain generated samples; the discriminator is used to output a probability estimate of the data source according to real data samples or generated samples and conditional information; A model training module, configured to train the data augmentation model to obtain a trained data augmentation model; A model evaluation module, configured to construct a multi-dimensional evaluation system to comprehensively evaluate the actual utility of the augmented data from two aspects of data similarity and downstream prediction task performance; wherein, the augmented data is obtained based on the trained data augmentation model.

9. A system for enhancing geoscience tabular data based on deep learning, characterized in that, The system includes: A memory, used to store computer programs; A processor, used to execute the computer program to implement the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing instructions, characterized in that, When the instruction is executed by the processor, the method according to any one of claims 1 to 7 is executed.

Citation Information

Patent Citations

  • Table data enhancement method and device, equipment and medium

    CN115983210A

  • Form data enhancement method and device based on reinforcement learning

    CN118446189A

  • Data enhancement model-based transformer fault diagnosis method and related equipment

    CN119598201A

  • Diversity-aware weighted majority vote classifier for imbalanced datasets

    US20220222931A1