Method for constructing three-dimensional metallogenic prediction balanced data set based on KMEDS algorithm

The three-dimensional mineralization prediction and balance data set is constructed through the KMEDS algorithm, which solves the problem of data category imbalance and missing values, and improves the prediction accuracy and model accuracy.

CN119338112BActive Publication Date: 2025-06-17CHINA UNIV OF GEOSCIENCES (WUHAN) +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411371053.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-06-17
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

There are data class imbalance and missing values ​​in three-dimensional mineralization prediction, resulting in a degradation of classifier performance, and it is difficult for the existing technology to solve these problems in essence.

Method used

The method based on KMEDS algorithm is adopted to group the three-dimensional attribute data sets through the K-Means clustering algorithm, select a few groups with a high proportion of samples, and use the improved SMOTE oversampling algorithm to generate new samples, balance the data set, and finally input the LightGBM model for three-dimensional mineralization prediction.

Benefits of technology

The accuracy of three-dimensional mineralization prediction is improved, high-quality new samples are generated through clustering and improved SMOTE algorithm, which reduces noise generation and improves the training efficiency and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119338112B_ABST
    Figure CN119338112B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing a three-dimensional metallogenic prediction balanced data set based on the KMEDS algorithm, belonging to the technical field of three-dimensional metallogenic prediction, and comprising the following steps: Step S1, collecting geo-physical-chemical-remote sensing data and performing preprocessing to construct a three-dimensional attribute data set; Step S2, using the K-Means clustering algorithm to cluster all samples in the three-dimensional attribute data set, dividing all samples into different groups, and selecting the groups in which the proportion of minority class samples meets the requirements; Step S3, applying an improved SMOTE oversampling algorithm to the selected groups to generate new samples, and adding the new samples to the three-dimensional attribute data set in Step S1 to form a new data set; Step S4, inputting the new data set into the LightGBM model for three-dimensional metallogenic prediction. The present invention can solve the problems of missing metallogenic prediction sample data and unbalanced sample data in the prior art, thereby improving the accuracy of three-dimensional metallogenic prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional metallogenic prediction, and in particular to a method for constructing a balanced dataset for three-dimensional metallogenic prediction based on the KMEDS algorithm. Background Technique

[0002] With the development of computer technology, remarkable achievements have been made in the application research of machine learning in three-dimensional metallogenic prediction. For example, in 2015, Carranza and Laborte carried out metallogenic prediction on gold mines in the Baguio area of the Philippines using the random forest algorithm. In 2021, Fu Guangming et al. obtained 5 groups of features based on three-dimensional geological modeling and gravity and magnetic inversion, and then constructed a metallogenic prediction model for the study area based on the random forest algorithm. The prediction results have important guiding significance for the exploration of characteristic minerals in the Zhuxi tungsten mine in northeastern Jiangxi.

[0003] However, in actual engineering, the dataset for three-dimensional metallogenic prediction still faces the dilemma of unbalanced data categories. In addition, due to exploration difficulties, it is inevitable that there are a large number of missing values in the geological dataset. Using such a dataset to train a classifier will lead to a decline in the performance of the classifier. At present, oversampling or undersampling algorithms are usually used to solve such problems in other fields. In the field of three-dimensional metallogenic prediction, some studies generate negative samples through random undersampling to balance the proportion of positive and negative samples, and some studies assign higher weights to minority-class samples to make the classifier pay more attention to them. However, these techniques either only consider the balance of the number of samples and do not pay attention to the quality of the generated new samples, or only use the original sample data for prediction and do not fundamentally solve the problem of data imbalance. With the gradual in-depth application of machine learning technology in the field of three-dimensional metallogenic prediction, the uneven distribution of geological dataset categories has also become an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for constructing a balanced dataset for three-dimensional metallogenic prediction based on the KMEDS algorithm, which can solve the problems of missing metallogenic prediction sample data and unbalanced sample data in the prior art, thereby improving the accuracy of three-dimensional metallogenic prediction.

[0005] To achieve the above object, the present invention provides a method for constructing a balanced dataset for three-dimensional metallogenic prediction based on the KMEDS algorithm, including the following steps:

[0006] Step S1: Collect geo-physical-chemical-remote sensing data and perform preprocessing to construct a three-dimensional attribute dataset;

[0007] Step S2: Use the K-Means clustering algorithm to cluster all samples in the three-dimensional attribute dataset, divide all samples into different groups, and select the groups where the ratio of the number of minority-class samples to the total number of samples in the group is higher than n, and the value range of n is 0.3 to 1;

[0008] Step S3: Apply the improved SMOTE oversampling algorithm to all the groups selected in Step S2 to generate new samples, and add the new samples to the three-dimensional attribute dataset in Step S1 to form a new dataset;

[0009] The improved SMOTE oversampling algorithm is as follows:

[0010] Any one of the groups selected in Step S2 consists of N minority class samples and W features. Use the minority class samples in this group to generate an Euclidean matrix M, and the dimension of matrix M is represented by [N, W];

[0011] Assume that the mean of each dimension of the sample set X is m and the standard deviation is s. Then the standardized variable of the sample set X is:

[0012]

[0013] where, X * is the sample set after standardization;

[0014] Calculate the distances between any two sample points in the standardized sample set in sequence to obtain a matrix with the dimension of [N, N], denoted as matrix A;

[0015] The distance calculation formula between any two sample points is as follows:

[0016]

[0017] where, x i , y i represent two sample points in the sample set; i represents the i-th dimension of the sample point; ED represents the Euclidean distance between two sample points;

[0018] Calculate the probability weight P of each sample:

[0019]

[0020] where, ES is the sum of the elements in each row of matrix A, and the calculation formula is as follows:

[0021] ES = [ΣED 1i , ΣED 2i , ΣED 3i , ……, ΣED Ni 1×N

[0022] where, ΣED Ni is the sum of the Euclidean distances of the elements in N rows of matrix A;

[0023] ​Assume that the total number of minority class samples to be generated is S, then the oversampling times required for each sample are as follows:

[0024]

[0025] Among them, T is the oversampling times required for each sample; S is the total number of minority class samples; U is the number of groups of minority class samples that need to be oversampled, that is, the number of selected groups after cluster analysis;

[0026] Sort the samples in descending order of T values, and oversample each sample within the group to generate new samples;

[0027] Step S4: Input the new data set into the LightGBM model for three-dimensional mineralization prediction.

[0028] Preferably, in step S1, the geological, geophysical, geochemical and remote sensing data include geological basic data, geophysical data, geochemical data and remote sensing data.

[0029] Preferably, in step S1, the preprocessing includes grid meshing, normalization processing and inverse distance spatial interpolation.

[0030] Based on the above method for constructing a three-dimensional mineralization prediction balanced data set based on the KMEDS algorithm, the present invention also provides a system for balancing a three-dimensional mineralization prediction data set based on the KMEDS algorithm, including a data set construction module, a clustering module, an oversampling module and a prediction module;

[0031] Among them, the data set construction module is used to collect geological, geophysical, geochemical and remote sensing data and perform preprocessing to construct a three-dimensional attribute data set;

[0032] The clustering module is used to use the K-means algorithm to screen out groups that meet the requirements for generating new samples;

[0033] The oversampling module is used to apply the improved SMOTE oversampling algorithm to all selected groups to generate new samples, and add the new samples to the three-dimensional attribute data set in step S1 to form a new data set;

[0034] The prediction module is used to input the new data set into the LightGBM model for three-dimensional mineralization prediction.

[0035] Therefore, the present invention adopts the above method for constructing a three-dimensional mineralization prediction balanced data set based on the KMEDS algorithm, and the beneficial technical effects are as follows:

[0036] (1) Group the data set through the clustering algorithm, and only select the groups with a high proportion of minority class samples, so as to realize oversampling only in the safe area and avoid the generation of noise;

[0037] (2) By calculating the Euclidean distance of each sample point in the group from other sample points, the sampling value of the sample points is weighted, reducing the possibility of overlap between the newly generated samples and other samples, while ensuring that the total amount of newly generated samples by oversampling remains unchanged;

[0038] (3) Select the LightGBM model for ore-forming prediction, which realizes the efficient processing of the three-dimensional attribute dataset, reduces memory consumption and improves the training speed.

[0039] In summary, the present invention not only has low complexity and is easy to implement, but also generates samples based on the clustering distribution, well preserves the statistical characteristics of most geoscience dimensions in the original data, focuses on solving the imbalance between categories and within categories, improves the quality of the geological three-dimensional attribute dataset, overcomes the uncertainty of the model prediction results, and improves the training efficiency and model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 is a flowchart of a method for constructing a three-dimensional ore-forming prediction balanced dataset based on the KMEDS algorithm according to the present invention;

[0041] Figure 2 is a structural block diagram of a system for balancing a three-dimensional ore-forming prediction dataset based on the KMEDS algorithm. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0043] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the field to which the present invention belongs.

[0044] Embodiment 1

[0045] As Figure 1 shown, the embodiment of the present invention discloses a method for constructing a three-dimensional ore-forming prediction balanced dataset based on the KMEDS algorithm, which specifically includes the following operations:

[0046] Step S1, collect geo-physical-chemical-remote sensing data and perform preprocessing to construct a three-dimensional attribute dataset;

[0047] Collect multi-source attribute data such as geological basic data, geophysics, geochemistry, and remote sensing, and construct a three-dimensional attribute dataset through operations such as grid meshing, normalization, and inverse distance spatial interpolation. In the field of three-dimensional ore-forming prediction, positive samples mean there is ore, and negative samples mean there is no ore.

[0048] Step S2: Use the K-Means clustering algorithm to cluster all samples in the three-dimensional attribute dataset, divide all samples into different groups, and select the groups where the ratio of the number of minority-class samples to the total number of samples within the group is higher than n. The value range of n is 0.3 to 1;

[0049] The value of n can be adjusted according to the sample ratio. Here, the minority-class samples refer to the category with a smaller number of samples among the positive and negative samples. For example, if the ratio of the number of positive and negative samples is 100:1, then the negative samples are the minority-class samples; otherwise, the positive samples are the minority-class samples;

[0050] Step S3: Apply the improved SMOTE oversampling algorithm to all the groups selected in Step S2 to generate new samples, and add the new samples to the three-dimensional attribute dataset in Step S1 to form a new dataset;

[0051] The improved SMOTE oversampling algorithm is as follows:

[0052] Any one of the groups selected in Step S2 consists of N minority-class samples and W features. Use the minority-class samples in this group to generate an Euclidean matrix M, and the dimension of matrix M is represented by [N, W];

[0053] Assume that the mean of each dimension of the sample set X is m and the standard deviation is s. Then the standardized variable of the sample set X is:

[0054]

[0055] where, X * is the sample set after standardization;

[0056] Calculate the distance between any two sample points in the standardized sample set in sequence to obtain a matrix with the dimension of [N, N], denoted as matrix A;

[0057] The distance calculation formula between any two sample points is as follows:

[0058]

[0059] where, x i , y i represent two sample points in the sample set; i represents the i-th dimension of the sample point; ED represents the Euclidean distance between two sample points;

[0060] Calculate the probability weight P of each sample. The range of P is within (0, 1), which is used to determine the importance of each sample:

[0061]

[0062] where, ES is the sum of the elements in each row of matrix A, and the calculation formula is as follows:

[0063] ES = [∑ED 1i , ∑ED 2i , ∑ED 3i , ………, ΣED Ni 1×N

[0064] Among them, ΣED Ni is the sum of the Euclidean distances of the N - row elements in matrix A;

[0065] Assume that the total number of minority - class samples to be generated is S, then the oversampling times required for each sample are:

[0066]

[0067] Among them, T is the oversampling times required for each sample; S is the total number of minority - class samples; U is the number of groups of minority - class samples that need to be oversampled, that is, the number of selected groups after cluster analysis;

[0068] Sort the samples in descending order according to the T value, and oversample each sample within the group to generate new samples;

[0069] Step S4: Input the new data set into the LightGBM model for three - dimensional mineralization prediction.

[0070] The present invention will be further described below through specific examples.

[0071] Taking a certain area as the research area, based on the collected geological basic data, geophysics, geochemistry, remote sensing and other multi - source attribute data, negative samples are generated, and then three - dimensional mineralization prediction is carried out.

[0072] Pre - process the collected original data, including inverse - distance spatial interpolation, dividing quantitative data and categorical data, normalizing the quantitative data, and performing one - hot encoding on the categorical data.

[0073] Divide the quantitative data and categorical data, as shown in Table 1.

[0074] Table 1 Data Division

[0075] Quantitative type data Categorical type data F Layer H.P. Name I Fault P Density Magnetism Resistivity

[0076] In Table 1, F, H.P., I, P are geochemical data, which are fluorine, acid - insoluble matter, iodine and phosphorus respectively, Density, Magnetism, Resistivity are geophysical data, which are density, magnetic susceptibility, resistivity respectively, Layer is stratigraphic data, Fault is fault data, and Name is lithologic data.

[0077] Normalize the quantitative data, as shown in Table 2.​

[0078] Table 2 Quantitative data

[0079] F H.P. I P Density Magnetism Resistivity 2.36 0.0 0.003 28.487499 2.78 24 2018 2.254 0.0 0.00308 15.1237 2.68 34 3214 2.17571 0.0 0.0028 24.1087 2.78 24 2018 2.0375 0.0 0.00265 23.79001 2.68 34 3214 1.935 0.0 0.0021 8.40875 2.02 56 643 1.47333 0.0 2.00785 19.205 2.78 24 2018

[0080] One-hot encoding is performed on the categorical data, as shown in Table 3.

[0081] Table 3 Categorical data

[0082] Layer_K Layer_Nh1c Layer_Nh2f ··· Fault_F81 Fault_F82 Fault_F9 0 0 0 ··· 0 0 0 0 0 0 ··· 0 0 0 0 0 1 ··· 0 0 0 0 0 1 ··· 0 0 0 0 0 1 ··· 0 0 0

[0083] A dataset is constructed based on the known labeled positive and negative samples. The dataset has a total of 30,063 positive samples and 5,450 negative samples.

[0084] First, apply the K-means clustering algorithm to the samples, divide the samples into K groups, and select all groups with a minority class sample ratio greater than or equal to n. The setting of the K value affects the effect of the K-means clustering algorithm. Secondly, the sample ratio threshold n determines the sampled groups. In this embodiment, the sample threshold is set to twice the positive and negative sample ratio, and the K value is set to an integer between [2, 6] to explore the sampling effect of the KMED-SMOTE algorithm on the ore-forming unbalanced dataset. Apply the improved SMOTE oversampling algorithm to these selected groups, add the generated new samples to the original dataset to form a new dataset. The positive and negative sample ratio of the new dataset is 1:1. Randomly divide the dataset into a training set and a test set in a ratio of 8:2, construct a LightGBM model (parameters objective=binary, max_depth=-1, boosting_type=goss, learning_rate=0.03), train the model and make predictions, and compare the ore-forming prediction accuracies with and without constructing the dataset using the embodiment of the present invention, as shown in Table 4.

[0085] Table 4 Comparison of experimental results

[0086] Oversampling algorithm Prediction accuracy Not adopted 0.7845 Conventional Smote 0.8237 KMED - SMOTE 0.8563

[0087] Embodiment 2

[0088] As Figure 2 shown, the present invention also provides a system for balancing a three-dimensional ore-forming prediction dataset based on the KMEDS algorithm, including a dataset construction module, a clustering module, an oversampling module, and a prediction module;

[0089] Among them, the dataset construction module is used to collect geophysical, geochemical, and remote sensing data and perform preprocessing to construct a three-dimensional attribute dataset;

[0090] The clustering module is used to use the K-means algorithm to screen out groups that meet the requirements for generating new samples;

[0091] An oversampling module, which is used to apply an improved SMOTE oversampling algorithm to all selected groups to generate new samples, and add the new samples to the three-dimensional attribute dataset in step S1 to form a new dataset;

[0092] A prediction module, which is used to input the new dataset into the LightGBM model for three-dimensional mineralization prediction.

[0093] The system described in this embodiment can implement the three-dimensional mineralization prediction method based on the LightGBM model described above, which will not be elaborated here.

[0094] It should be noted that the content not elaborated in detail in the present invention is prior art and well-known to those skilled in the art.

[0095] Therefore, by adopting the above method for constructing a three-dimensional mineralization prediction balanced dataset based on the KMEDS algorithm, the present invention can solve the problems of missing mineralization prediction sample data and unbalanced sample data in the prior art, thereby improving the accuracy of three-dimensional mineralization prediction.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for constructing a three-dimensional mineralization prediction equilibrium data set based on the KMEDS algorithm, characterized in that: The following steps are involved: Step S1, collecting geophysical data and performing preprocessing to construct a three-dimensional attribute data set; geophysical data and remote sensing data include basic geological data, geophysical data, geochemical data, and remote sensing data; preprocessing includes grid generation, normalization processing, and inverse distance space interpolation; Step S2: Use K-Means clustering algorithm to cluster all samples in the three-dimensional attribute data set, divide all samples into different groups, and select the samples whose ratio of the number of minority class samples to the total number of samples in the group is higher than Group, The value range of is 0.3~1; Step S3, applying the improved SMOTE oversampling algorithm to all the groups selected in step S2 to generate new samples, and adding the new samples to the three-dimensional attribute data set in step S1 to form a new data set; The improved SMOTE oversampling algorithm is as follows: Any group selected in step S2 has minority class samples and The minority class samples in this group are used to generate a Euclidean matrix M ,matrix M The dimension is [ N , W ]express; Assume sample set X The mean of each dimension is , the standard deviation is , then the sample set X The standardized variables are: ; in, is the sample set after standardization; Calculate the distance between any two sample points in the standardized sample set in turn, and get a [ N , N ] dimension matrix, denoted as matrix A ; The distance calculation formula between any two sample points is as follows: ; in, , Represents two sample points in the sample set; Represents the sample point Dimension; Represents the Euclidean distance between two sample points; Calculate the probability weight of each sample : ; in, For the matrix A The sum of the elements in each row of is calculated as follows: ; in, For the matrix A middle N The sum of the Euclidean distances of row elements; Assume that the total number of minority class samples to be generated is S , then the number of oversampling times required for each sample is: ; in, The number of oversampling times required for each sample; is the total number of minority class samples; is the number of groups of minority class samples that need to be oversampled; according to The samples are prioritized in descending order of value, and each sample in the group is oversampled to generate new samples; Step S4: Input the new data set into the LightGBM model to perform three-dimensional mineralization prediction.

Citation Information

Patent Citations

  • Network traffic data enhancement method oriented to sample imbalance

    CN114781492A

  • Novel oversampling method for software defect prediction

    CN115543776A