A Method for Ore-Forming Data Augmentation Based on Generative Adversarial Networks
By generating high-quality and diverse mineralized samples by generating adversarial networks, the problem of data imbalance in traditional exploration methods is solved and the performance of mineral resource prediction models is improved.
Patent Information
- Application Number
- CN202411339942.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-09-25
AI Technical Summary
Traditional exploration methods are time-consuming and costly, with far fewer ore-forming samples than unmineralized samples, resulting in imbalance in the data set and affecting the training and prediction accuracy of machine learning models.
Using a method based on generative adversarial network, samples are classified through hierarchical clustering algorithms, category frequency distribution and weights are calculated, standardized weight parameters are constructed, synthetic data that meets weight requirements are generated, and high-quality and diverse mineralized samples are generated using weighted sampling strategies.
It significantly improves the performance of the mineral resource prediction model, solves the problem of data imbalance, and the generated synthetic data is of high quality, which can effectively improve the prediction accuracy.
Smart Images

Figure CN119272049B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of metallogenic prediction, and in particular to a method for augmenting metallogenic data based on a generative adversarial network. Background Art
[0002] Metallogenic prediction is an important part of the field of geological exploration. Traditional exploration methods rely on a large amount of fieldwork, which is often time-consuming, laborious and costly. Obtaining a large number of high-quality ore sample data is often restricted by factors such as geographical location, economic conditions and time. Therefore, there is often a problem of dataset imbalance in mineral resource exploration data. Specifically, the number of metallogenic samples is far less than that of non-metallogenic samples, and this imbalance has a significant negative impact on the training and prediction accuracy of machine learning models.
[0003] In order to address the data imbalance problem, especially for the augmentation of metallogenic data, many researchers have proposed a series of methods. Such as undersampling method, oversampling method, synthetic minority over-sampling technique SMOTE, and adaptive synthetic sampling ADASYN and other ensemble learning methods. Although the various methods proposed by predecessors have made certain progress, there are still certain challenges in the case of the high complexity unique to mineral exploration and the scarcity of metallogenic data. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for augmenting metallogenic data based on a generative adversarial network, which can generate high-quality and diverse metallogenic samples, significantly improve the performance of the mineral resource prediction model, and has great significance in the field of metallogenic prediction and broad application prospects.
[0005] To achieve the above object, the present invention provides a method for augmenting metallogenic data based on a generative adversarial network, including the following steps:
[0006] S1. Collect the geophysical, geochemical, remote sensing data of known ore points and non-ore points, preprocess the collected geophysical, geochemical, remote sensing data, and construct a metallogenic prediction dataset;
[0007] S2. Classify the samples in the metallogenic prediction dataset by using a hierarchical clustering algorithm, and classify the samples;
[0008] S3. Determine the corresponding class weights by calculating the frequency distribution of each class, and perform normalization processing on the class weights to obtain normalized weights;
[0009] S4. Construct an optional parameter including the normalized weights in the generative adversarial network model for transmitting the weight information of each class;
[0010] S5. According to the optional parameters of the standardized weights, adopt a weighted sampling strategy, and combine the preset specified conditions, such as the number of generated data, specific category values, etc., to generate synthetic data;
[0011] S6. Conduct quality assessment on the synthetic data.
[0012] Preferably, in step S1, the geological, geophysical, geochemical and remote sensing data include regional geological data, remote sensing data, soil geochemical data and geophysical data.
[0013] Preferably, in step S1, the preprocessing includes missing value processing and reclassification.
[0014] Preferably, in step S3, the frequency distribution of each category is as follows:
[0015]
[0016] Among them, f i is the frequency distribution of category i; N i is the sample number of each category i; N total is the total sample number;
[0017] The category weight is set to the frequency distribution of each category:
[0018] w i = f i (2)
[0019] Among them, w i is the category weight;
[0020] The formula for normalizing the category weight is:
[0021]
[0022] Among them, w' i is the normalized category weight; is the sum of all category weights; k is the total number of categories.
[0023] Preferably, in step S4, in the generative adversarial network model, construct an optional parameter containing the standardized weights to transmit the weight information of each category. The specific operation is as follows:
[0024] First, add an optional parameter weights in the constructor of the generative adversarial network model. The optional parameter weights adjusts the generation strategy of the generator and the discrimination strategy of the discriminator according to the standardized weights of each category during the training process;
[0025] Then, adjust the generator and the discriminator;
[0026] The generator generates samples based on the input random noise and conditional information. During training, the generator uses the optional parameter weights for weighted sampling to generate samples that meet the weight requirements.
[0027] The loss function of the generator is as follows:
[0028]
[0029] Among them, G(z, c i ) is the sample generated by the generator, D is the output of the discriminator, and c i is the class condition; L G is the loss function of the generator; is the expected value operator, representing the average of all possible values of the random noise vector z; z is the random noise vector; p z (z) is the probability distribution of the random noise vector z;
[0030] The discriminator distinguishes real samples and generated samples.
[0031] The loss function of the discriminator is as follows:
[0032]
[0033] Among them, p data (x) is the real sample distribution; D(x, c i ) is the predicted probability of the discriminator for the real sample x; G(z, c i ) is the sample generated by the generator; L D is the loss function of the discriminator; is the expected value operator for the real sample x, representing the average of all possible values of the real sample; x is the real sample; p data (x) is the probability distribution of the real data.
[0034] Preferably, in step S6, the quality of the synthetic data is evaluated, and the evaluation metrics used include mean square error, AUC, and F1 score.
[0035] Therefore, the present invention adopts the above-mentioned ore-forming data augmentation method based on a generative adversarial network, which can generate high-quality and diverse ore-forming samples, significantly improving the performance of the mineral resource prediction model. The present invention provides an effective new method for solving the data imbalance problem in mineral resource exploration, which has great significance in the field of ore-forming prediction and has a wide application prospect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of an ore-forming data augmentation method based on a generative adversarial network according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0037] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0038] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention pertains.
[0039] Embodiment 1
[0040] In this embodiment, a certain manganese mining area is taken as the research area, and the weights of ore-forming prediction and ore-controlling factors are determined for the multi-source data such as geological, geophysical, geochemical, and remote sensing data collected.
[0041] As Figure 1 shown, it is a flowchart of a method for expanding ore-forming data based on a generative adversarial network according to the present invention, specifically including the following steps:
[0042] Step S1, collect the geological, geophysical, geochemical, and remote sensing data of known ore points and non-ore points, and construct a model data set. According to the ore prospecting marks, preprocessing work such as missing value processing and reclassification is carried out on the regional geological data, remote sensing data, soil geochemical data, and geophysical data. Finally, the ore-forming element information of the four types of data is extracted in numerical form as features and incorporated into the data set, and an ore-forming prediction data set is established according to the ore occurrence situation of the borehole data. The data is stored in tabular form, with each row representing a set of geological data and each column representing a feature variable. Among them, a single piece of data has 27-dimensional features, and the features include physical features (XY coordinates, upward continuation of 2 km, etc.), chemical components (Mn, Cu, etc. and oxides Al2O3, Fe2O3, etc.), geological features (strata, faults, etc.), and remote sensing features. Part of the data is shown in Table 1:
[0043] Table 1 Partial data display
[0044] id X coordinate Y coordinate 2 km Mn Cu <![CDATA[Fe2O3]]> <![CDATA[Al2O3]]> … stratum fault 1 887122 3120015 6 4 4 4 7 … 3 10 2 887452 3121005 5 5 4 4 6 … 11 7 3 887572 3120015 6 4 5 5 7 … 3 1 … … … … … … … … … … … 647 866032 3130005 4 8 3 3 6 … 8 9
[0045] Step S2, use the hierarchical clustering algorithm to classify the existing ore-forming samples, and divide the samples into multiple categories. The agglomerative hierarchical clustering is adopted in this embodiment. Each ore-forming data point is used as an independent clustering object, the distance between clusters is calculated, and the most similar clusters are gradually merged until a single cluster containing all data points is finally formed. The distance metric used in this embodiment is the Euclidean Distance. For two points p = (p1, p2,..., p n ) and q = (q1, q2,..., q n ) in the n-dimensional space, the Euclidean distance calculation formula is as follows:
[0046]
[0047] where qi is the coordinate of point q in the i-th dimension; p i is the coordinate of point p in the i-th dimension; d Euclidean represents the Euclidean distance between two points; i is an index indicating the dimension currently being calculated.
[0048] Next, select the two clusters with the smallest distance for merging to form a new cluster. The merging strategy adopted in this embodiment includes but is not limited to the Ward linkage method. Specifically, the core formula of the Ward linkage method is as follows:
[0049]
[0050] where x i and x j are data points in cluster C; |x i - x j | represents the Euclidean distance between these two points. After merging, update the distance matrix to reflect the distance of the new cluster. This process is repeated continuously until all ore-forming data objects reach the predetermined number of clusters.
[0051] Step S3, by calculating the frequency distribution of each category, determine the corresponding category weights and standardize the weights. The following details this process:
[0052] First, it is necessary to count the number of samples in each category in the dataset. Suppose there are k categories, and the number of samples in each category i is N i , and the total number of samples is N total . The formula for the frequency distribution of each category i is:
[0053]
[0054] where f i is the frequency distribution of category i; N i is the number of samples in each category i; N total is the total number of samples.
[0055] Secondly, determine the category weights. In this embodiment, the number of samples of each category in the generated dataset is consistent with the category distribution in the actual dataset. The category weight w i can be set to the frequency distribution of each category:
[0056] w i = f i (4)
[0057] where w i is the category weight;
[0058] In this way, the weight of each category is proportional to its frequency in the original dataset.
[0059] Finally, standardize the weights. The standardized weights ensure that the sum of the weights for all classes is 1, enabling the generative model to sample according to the specified ratio when generating synthetic samples. The formula for standardizing the weights is:
[0060]
[0061] where w' i is the standardized class weight; is the sum of all class weights; k is the total number of classes.
[0062] Step S4, construct an optional parameter in the generative adversarial network model that contains the standardized weights to transmit the weight information for each class.
[0063] First, construct the optional parameter weights. To introduce the standardized weights, an optional parameter weights needs to be added to the constructor of the generative adversarial network model to transmit the weight information for each class. This parameter allows adjusting the generation strategy of the generator and the discrimination strategy of the discriminator according to the weights of each class during training.
[0064] Then, adjust the generator and the discriminator. The generator generates samples based on the input random noise and conditional information. During training, the generator will use the standardized weights for weighted sampling to generate samples that meet the class weight requirements. The loss function of the generator may consider the class weights to increase the emphasis on generating samples of minority classes. The loss function of the generator can be expressed as:
[0065]
[0066] where G(z, c i ) is the sample generated by the generator; D is the output of the discriminator; c i is the class condition; L G is the loss function of the generator; is the expectation operator, representing the average over all possible values of the random noise vector z; z is the random noise vector; p z (z) is the probability distribution of the random noise vector z.
[0067] The goal of the discriminator is to distinguish real samples from generated samples. Considering the class weights, the loss function of the discriminator will also include factors of class weights to balance the discrimination ability for samples of different classes. The calculation formula for the loss function of the discriminator is:
[0068]
[0069] Among them, p data (x) is the true sample distribution; D(x, c i ) is the predicted probability of the discriminator for the true sample x; G(z, c i ) is the sample generated by the generator; L D is the loss function of the discriminator; is the expected value operator for the true sample x, representing the average of all possible values of the true sample; x is the true sample data; p data (x) is the probability distribution of the true data.
[0070] Step S5: According to the standardized class weights, adopt a weighted sampling strategy to generate synthetic data with higher balance. The weighted sampling strategy means that when generating synthetic samples, samples are sampled according to the pre-set class weights. Classes with larger weights will be sampled more, while classes with smaller weights will be sampled less. This strategy ensures that the number of samples of each class in the generated synthetic dataset is more in line with the expected ratio, avoiding the situation where some classes in the generated dataset are too many or too few.
[0071] Suppose there are k classes in the dataset, and the standardized weight of each class is w i (i from 1 to k), and the sum of the standardized weights is W: The sampling probability p i of class i is calculated as follows:
[0072]
[0073] Here, p i represents the probability that class i is selected when generating samples.
[0074] When generating synthetic samples, use the calculated sampling probability p i to determine the class of the generated samples. By the method of random sampling, select the class according to the sampling probability p i . Generate a random number r in the interval [0, 1], and then determine the class according to the standardized weights: if 0 ≤ r < p1, then select class 1; if p1 ≤ r < p1 + p2, then select class 2; if p1 + p2 + … + p n-1 ≤ r < p1 + p2 + … + p n-1 + p n , then select class n; and so on until a class is selected.
[0075] Step S6: Evaluate the quality of the synthetic data.
[0076] After the generative adversarial network model generates synthetic data, in order to ensure that the generated data can be effectively used as the input of the prediction model, its quality must be comprehensively evaluated. The essential purpose of the data augmentation method is to improve the performance of the prediction model. Therefore, the generated data is used to train or test the prediction model, and the performance metrics of the model are evaluated. In this implementation, MSE, AUC, and F1 score are selected as evaluation metrics to judge the impact of the generated data on the model prediction effect. The calculation formula of the Mean Squared Error (MSE) is:
[0077]
[0078] where y i is the true value; is the model prediction value; m is the number of ore-forming samples. The smaller the MSE, the more helpful the data generated by this method is for the model to accurately predict. AUC is a measure of the area under the ROC curve, indicating the overall performance of the model in binary classification problems. The closer the AUC value is to 1, the better the classification effect of the model. The F1 score is the harmonic mean of the model's precision and recall, used to balance the accuracy and coverage of the model. The higher the F1 score, the better the model's ability to identify positive classes. The F1 calculation formula is as follows:
[0079]
[0080] where the precision and recall are respectively: TP is the true positive; FP is the false positive; FN is the false negative.
[0081] The generated dataset and the real dataset are compared through the above metrics respectively. If the corresponding values of the generated data are close to or higher than the real dataset, it indicates that the generated data has high quality in model prediction. If it is significantly lower than the real dataset, it may be necessary to improve the generative adversarial network (GAN) model or the data generation process.
[0082] Therefore, the present invention adopts the above-mentioned ore-forming data augmentation method based on a generative adversarial network, which can generate high-quality and diverse ore-forming samples, and significantly improve the performance of the mineral resource prediction model. This research provides an effective new method for solving the data imbalance problem in mineral resource exploration, has great significance in the field of ore-forming prediction, and has broad application prospects.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for expanding ore-forming data based on a generative adversarial network, characterized in that It includes the following steps: S1. Collect the geophysical, geochemical, remote sensing data of known ore points and non-ore points, preprocess the collected geophysical, geochemical, remote sensing data, and construct a metallogenic prediction data set; S2. Use the hierarchical clustering algorithm to classify the samples in the metallogenic prediction data set and perform category division; S3. Determine the corresponding category weights by calculating the frequency distribution of each category, and perform normalization processing on the category weights to obtain the normalized weights; S4. Construct an optional parameter including the normalized weights in the generative adversarial network model for transmitting the weight information of each category. The specific operation is as follows: First, add an optional parameter weights in the constructor of the generative adversarial network model. The optional parameter weights adjusts the generation strategy of the generator and the discrimination strategy of the discriminator according to the normalized weights of each category during the training process; Then, adjust the generator and the discriminator; The generator generates samples according to the input random noise and conditional information. During the training process, the generator uses the optional parameter weights for weighted sampling to generate samples that meet the weight requirements; The loss function of the generator is as follows: Among them, G(z, c i ) is the sample generated by the generator; D is the output of the discriminator; c i is the class condition; L G is the loss function of the generator; is the expectation operator, representing the average over all possible values of the random noise vector z; z is the random noise vector; p z (z) is the probability distribution of the random noise vector z; w i is the class weight; The discriminator distinguishes real samples and generated samples; The loss function of the discriminator is as follows: Among them, p data (x) is the true sample distribution; D(x, c i ) is the predicted probability of the discriminator for the true sample x; G(z, c i ) is the sample generated by the generator; L D is the loss function of the discriminator; is the expectation operator for the true sample x, representing the average of all possible values of the true sample; x is the true sample; p data (x) is the probability distribution of the true data; S5. According to the optional parameter of the normalized weights, adopt a weighted sampling strategy and combine the preset specified conditions to generate synthetic data; S6. Evaluate the quality of the synthetic data.
2. The ore-forming data augmentation method based on a generative adversarial network according to claim 1, characterized in that In step S1, the geophysical, geochemical, remote sensing data includes regional geological data, remote sensing data, soil geochemical data, and geophysical data.
3. The ore-forming data augmentation method based on a generative adversarial network according to claim 2, wherein In step S1, the preprocessing includes missing value processing and reclassification.
4. A method for expanding ore-forming data based on a generative adversarial network according to claim 3, characterized in that, In step S3, the frequency distribution of each category is as follows: where f i is the frequency distribution of class i; N i is the number of samples for each class i; N total is the total number of samples; The category weights are set as the frequency distribution of each category: w i = f i where, w i is the category weight; The formula for normalizing the category weights is: where, w' i is the normalized class weight; is the sum of all class weights; k is the total number of classes.
5. A method for expanding ore-forming data based on a generative adversarial network according to claim 1, characterized in that, In step S6, when evaluating the quality of the synthetic data, the evaluation indicators used include mean square error, AUC, and F1 score.
Citation Information
Patent Citations
Mineral prediction method and system based on multi-scale sample unevenness
CN115511214A
Robot autonomous exploration method and system based on generative adversarial network
CN117369455A