Deep learning prospecting model fair comparison evaluation method and system based on multiple constraints
By employing a fair comparative evaluation method for deep learning mineral exploration models with multiple constraints, the fairness issue in the evaluation of deep learning mineral exploration models is resolved, enabling more accurate model selection and decision support, and improving the reliability and guidance of mineral exploration prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AERIAL PHOTOGRAMMETRY & REMOTE SENSING CO LTD
- Filing Date
- 2026-04-21
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to establish fair benchmarks in evaluating deep learning-based mineral exploration models, leading to inaccurate performance comparisons between models and impacting the selection of the optimal model and technical decisions.
A fair comparative evaluation method for deep learning mineral exploration models with multiple constraints is proposed, including data preprocessing, sample library construction, model training, performance evaluation, and decision logic consistency analysis. A fair comparative analysis report is generated by adopting a multi-dimensional performance evaluation system and decision logic consistency analysis.
It significantly improves the depth and reliability of comparative evaluation of deep learning mineral exploration models, can comprehensively reveal the actual performance characteristics of models in complex mineral exploration tasks, provides an assessment of the reliability of model decision-making, and enhances the guiding value and application credibility of evaluation conclusions.
Smart Images

Figure CN122064922A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of interdisciplinary technology of artificial intelligence and geological exploration, and more specifically, to a fair comparative evaluation method and system for deep learning mineral exploration models based on multiple constraints. Background Technology
[0002] In the field of mineral resource prediction, various deep learning models, such as convolutional neural networks, graph neural networks, and Transformers, can be integrated to predict prospective mineralized areas. To select the optimal model suitable for a specific geological scenario, existing technologies typically conduct comparative experiments on these models. In this process, existing technologies strive to evaluate model performance through a unified data preprocessing workflow, the construction of standardized training and testing sample libraries, and the use of multiple accuracy metrics, aiming to compare the merits of different models under consistent conditions.
[0003] However, due to the inherent differences in the internal structure, algorithm principles, and input data requirements of different deep learning models, simply unifying data input and evaluation metrics is still insufficient to completely eliminate the systematic bias caused by model heterogeneity in the evaluation process. This bias means that the performance comparison between models is not based on a completely fair benchmark, which may make the evaluation conclusions unable to truly and reliably reflect the actual capabilities and applicability of the models in mineral exploration prediction tasks, thereby affecting the accurate selection of the optimal model and technical decisions. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art, this invention provides a fair comparative evaluation method and system for deep learning prospecting models based on multiple constraints to address the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] A fair comparative evaluation method for deep learning-based mineral exploration models based on multiple constraints includes the following steps:
[0007] S1: Preprocess the multi-source mineral exploration data to obtain preprocessed data;
[0008] S2: Based on the preprocessed data, the difficulty level is divided according to the intensity of mineralization information, and a sample library containing training set, validation set, test set and extrapolation area samples is constructed accordingly.
[0009] S3: Train multiple deep learning mineral exploration prediction models using training and validation sets;
[0010] S4: The performance of the trained deep learning mineral exploration prediction model is evaluated using the test set and extrapolated regional samples to obtain multi-dimensional performance evaluation results.
[0011] S5: Perform a decision logic consistency analysis on the trained deep learning mineral exploration prediction model and obtain the decision logic consistency comparison results.
[0012] S6: Generate a fair comparison analysis report by integrating the multi-dimensional performance evaluation results with the consistency of decision-making logic.
[0013] Furthermore, the multi-source mineral exploration data is preprocessed to obtain preprocessed data, including:
[0014] The hyperspectral remote sensing data in the multi-source mineral exploration data is spatially aligned with other auxiliary data, and the aligned data is then subjected to noise removal processing.
[0015] The data after noise removal is normalized to obtain preprocessed data.
[0016] Furthermore, multi-source mineral exploration data includes at least two of the following: hyperspectral remote sensing data, multispectral remote sensing data, geochemical data, and geophysical data.
[0017] Furthermore, based on the preprocessed data, the difficulty levels are divided according to the intensity of mineralization information, and a sample library is constructed accordingly, including a training set, a validation set, a test set, and extrapolated regional samples, including:
[0018] Extract feature indicators reflecting the intensity of mineralization information from the preprocessed data;
[0019] The samples are divided into multiple preset difficulty levels based on feature indicators;
[0020] Samples are drawn from each difficulty level according to a preset ratio and allocated to the training set, validation set, and test set respectively.
[0021] And data from the preprocessed data that are spatially independent of the regions corresponding to the training set, validation set, and test set are used as extrapolation region samples.
[0022] Furthermore, multiple deep learning mineral exploration prediction models were trained using training and validation sets, including:
[0023] Set the same initial training conditions and hyperparameter search strategy for multiple deep learning mineral exploration prediction models;
[0024] The training set was used to iteratively optimize the parameters of multiple deep learning mineral exploration prediction models, and the validation set was used to monitor the model performance during the iterative optimization process.
[0025] When the model performance meets the preset convergence or early stopping conditions, the training is completed, and the trained deep learning mineral exploration prediction model is obtained.
[0026] Furthermore, the performance of the trained deep learning mineral exploration prediction model was evaluated using the test set and extrapolated region samples, yielding multi-dimensional performance evaluation results, including:
[0027] The test set and the extrapolated region samples are respectively input into each trained deep learning mineral exploration prediction model to obtain the corresponding prediction results.
[0028] Based on the prediction results and real labels, indicators reflecting the basic prediction accuracy of the model, the ability of the model to spatially characterize mineralized areas, and the ability of the model to generalize in unseen areas are calculated to obtain multi-dimensional performance evaluation results.
[0029] Furthermore, a decision logic consistency analysis was performed on the trained deep learning mineral exploration prediction model to obtain the decision logic consistency comparison results, including:
[0030] Consensus samples with consistent prediction results and disputed samples with differing prediction results are selected from the test set to form a sample subset for analysis.
[0031] By using model interpretability techniques, the visual or spectral features on which each trained deep learning mineral exploration prediction model makes decisions for each sample in the sample subset are extracted, thus obtaining the decision-making basis features corresponding to each model.
[0032] The consistency measure of decision-making basis features among different trained deep learning mineral exploration prediction models is calculated and compared to obtain the results of the consistency comparison of decision logic.
[0033] Furthermore, model interpretability techniques include gradient-weighted class activation mapping or attention weight visualization techniques.
[0034] Furthermore, by integrating the multi-dimensional performance evaluation results with the consistency comparison results of the decision-making logic, a fair comparison analysis report is generated, including:
[0035] Integrate the results of multi-dimensional performance evaluation with the results of consistency comparison of decision-making logic;
[0036] Based on the integrated results, the performance of different trained deep learning mineral exploration prediction models is ranked and compared across various performance dimensions. Combined with the differences in the degree of consistency of decision logic, a fair comparative analysis report is generated, which includes a ranking of the overall model performance and an evaluation of decision reliability.
[0037] On the other hand, this invention provides a fair comparative evaluation system for deep learning mineral exploration models based on multiple constraints, comprising the following modules:
[0038] The data processing module is used to preprocess multi-source mineral exploration data to obtain preprocessed data;
[0039] The sample library construction module is used to classify the difficulty level according to the intensity of mineralization information based on the preprocessed data, and construct a sample library containing training set, validation set, test set and extrapolation region samples accordingly.
[0040] The model training module is used to train multiple deep learning mineral exploration prediction models using training and validation sets.
[0041] The performance evaluation module is used to evaluate the performance of the trained deep learning mineral exploration prediction model using the test set and extrapolated regional samples, and obtain multi-dimensional performance evaluation results.
[0042] The logic analysis module is used to perform decision logic consistency analysis on the trained deep learning mineral exploration prediction model and obtain the decision logic consistency comparison results.
[0043] The report generation module is used to generate a fair comparative analysis report by integrating the results of multi-dimensional performance evaluation with the results of the consistency of decision-making logic.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. By implementing a chain of multiple constraints from data construction to model analysis, the depth and reliability of comparative evaluation of deep learning mineral exploration models are significantly improved. Not only is a standardized sample library containing extrapolation regions constructed, but the samples are also innovatively stratified by difficulty based on the intensity of mineralization information. This ensures that the models undergo balanced testing across a continuous spectrum of cases, from typical to fuzzy, thus breaking through the limitations of traditional evaluations that rely solely on random data segmentation. A fair testing benchmark with greater geological representativeness and discriminative power is established from the source. At the same time, the multi-dimensional performance evaluation system introduced comprehensively considers basic accuracy, spatial characterization ability, and generalization ability in unknown areas, making the comparison of the merits of different models no longer singular and able to more comprehensively reveal the actual performance characteristics of different models in complex mineral exploration tasks.
[0046] 2. By extracting and comparing the decision-making basis characteristics of different models on consensus and disputed samples, this method can effectively identify models with similar prediction results but different decision-making logics, thereby assessing the reliability and robustness of their decisions. This makes the final fair comparative analysis report not only include performance rankings but also provide comments on the reliability of model decisions. It provides a key basis for selecting mineral exploration prediction models that are not only highly accurate but also have reasonable decision-making logic and strong interpretability. Ultimately, it achieves a leap from superficial performance comparison to fair comparison of internal mechanisms, enhancing the guiding value and application credibility of the evaluation conclusions. Attached Figure Description
[0047] Figure 1 The flowchart shows the fair comparison and evaluation method for deep learning prospecting models based on multiple constraints according to the present invention.
[0048] Figure 2 This is a schematic diagram of the structure of the fair comparison and evaluation system for deep learning prospecting models based on multiple constraints according to the present invention. Detailed Implementation
[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0050] Example 1: Figure 1 This invention presents a fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints, which includes the following steps:
[0051] S1: Preprocess the multi-source mineral exploration data to obtain preprocessed data;
[0052] S2: Based on the preprocessed data, the difficulty level is divided according to the intensity of mineralization information, and a sample library containing training set, validation set, test set and extrapolation area samples is constructed accordingly.
[0053] S3: Train multiple deep learning mineral exploration prediction models using training and validation sets;
[0054] S4: The performance of the trained deep learning mineral exploration prediction model is evaluated using the test set and extrapolated regional samples to obtain multi-dimensional performance evaluation results.
[0055] S5: Perform a decision logic consistency analysis on the trained deep learning mineral exploration prediction model and obtain the decision logic consistency comparison results.
[0056] S6: Generate a fair comparison analysis report by integrating the multi-dimensional performance evaluation results with the consistency of decision-making logic.
[0057] S1: Preprocess the multi-source mineral exploration data to obtain preprocessed data. The specific implementation is as follows:
[0058] Multi-source mineral exploration data specifically includes hyperspectral remote sensing data, as well as at least one other auxiliary data selected from multispectral remote sensing data, geochemical data, and geophysical data. For example, hyperspectral remote sensing data can be acquired from a spaceborne hyperspectral imager. An example of other auxiliary data is multispectral remote sensing data. An example of geochemical data is the content of multiple elements obtained from laboratory analysis of samples collected from the Earth's surface or boreholes. An example of geophysical data is space field data obtained through gravity measurements. These data, acquired with their own independent spatial reference coordinate systems and spatial resolutions, require unified preprocessing.
[0059] This process spatially aligns hyperspectral remote sensing data from multi-source mineral exploration data with other auxiliary data. A target spatial reference coordinate system, such as a unified geodetic coordinate system, is selected. The original coordinates of all other auxiliary data are transformed to this target spatial reference coordinate system. A target spatial resolution is selected, which can be based on the spatial resolution of the hyperspectral remote sensing data. For other auxiliary data with spatial resolutions different from the target spatial resolution, a resampling method is used to adjust their pixel size to the target spatial resolution. The resampling method can be bilinear interpolation. In calculating new pixel values, bilinear interpolation requires finding the floating-point coordinates of the target pixel in the original image and then obtaining the values of the four nearest original pixels around that location. The new pixel value is calculated by performing two linear interpolations on these four original pixel values based on their relative distances to the target floating-point coordinates. First, two linear interpolations are performed in the horizontal direction to obtain two intermediate values; then, a third linear interpolation is performed on these two intermediate values in the vertical direction to obtain the final new pixel value. Through coordinate transformation and resampling operations, the hyperspectral remote sensing data is spatially aligned with other auxiliary data.
[0060] Noise removal is performed on the spatially aligned data. The specific content of noise removal varies depending on the data type. For hyperspectral and multispectral remote sensing data, noise removal includes bad band removal. Bad bands are spectral bands whose information is invalidated due to atmospheric absorption. Bad bands are identified by calculating the signal-to-noise ratio (SNR) of each band. The SNR can be calculated by dividing the average signal intensity of the band by the standard deviation of the band noise. Bands with an SNR below a preset SNR threshold are marked as bad bands and removed from the data. The preset SNR threshold can be set based on prior knowledge of the inherent noise level of the imaging sensor. For example, if the noise equivalent radiative value of the sensor within a specific wavelength range is known, a minimum acceptable SNR can be derived as the preset SNR threshold. Alternatively, the preset SNR threshold can be set by analyzing the statistical distribution of the SNR across all bands of the entire image; for example, the median SNR of all bands can be multiplied by a coefficient, such as 0.5, to determine the preset SNR threshold. For geochemical data, noise removal primarily targets outliers. Outlier identification can employ statistical methods, such as calculating the mean and standard deviation of each element's content. The standard deviation, a statistic measuring data dispersion, is calculated by first determining the difference between each sample's content and the mean, squaring these differences, summing them, dividing by the sample size minus one, and finally taking the square root. Samples with content values exceeding the mean plus three times the standard deviation or falling below the mean minus three times the standard deviation are considered potential outliers. Outliers can be addressed using a truncation method, replacing values outside these ranges with boundary values within that range. For geophysical data, noise removal includes trend surface analysis. Trend surface analysis calculates the regional background field using polynomial fitting. Polynomial fitting involves establishing a polynomial mathematical model to approximate the spatial field's changing trend, and solving for the polynomial coefficients using the least squares method. The polynomial order is chosen to effectively characterize the regional background morphology without fitting local details; for example, a 2nd or 3rd order polynomial is selected. Subtracting this background field from the original data yields the remaining outliers.
[0061] The data after noise removal is normalized. Normalization is performed independently for each data feature. For each spectral band of hyperspectral and multispectral remote sensing data, normalization uses a standardization method. The standardization method requires calculating the mean and standard deviation of all pixel values in a given band. The mean is the sum of all pixel values divided by the total number of pixels. The calculation of the standard deviation has been described previously. Each pixel value is subtracted from the mean of the band and then divided by the standard deviation of the band. For the content values of each element in geochemical data, normalization can use a standardization method, i.e., calculating the mean and standard deviation of the element's content in all samples, and then subtracting the mean and dividing by the standard deviation for each sample's content value. For geophysical data, such as magnetic data, normalization also uses a standardization method, i.e., calculating the mean and standard deviation of the physical field data for the entire study area, and then subtracting the mean and dividing by the standard deviation for each pixel value. After normalization, multi-source mineral exploration data from different sources are transformed to a unified numerical scale, collectively forming the preprocessed data. In the preprocessed data, each pixel or sampling point corresponds to a set of aligned, denoised, and normalized feature values.
[0062] S2: Based on the preprocessed data, the difficulty levels are divided according to the intensity of mineralization information, and a sample library containing training set, validation set, test set, and extrapolated regional samples is constructed accordingly. The specific implementation is as follows:
[0063] The preprocessed data includes multi-source mineral exploration data features that have undergone alignment, denoising, and normalization. A sample refers to an independent spatial unit in the preprocessed data; a spatial unit is a pixel, and each sample corresponds to a set of feature values and a known mineralized or non-mineralized label.
[0064] Feature indicators reflecting the intensity of mineralization information are extracted from preprocessed data. These indicators quantify the strength or typicality of the mineralization information contained in a sample. The extraction of feature indicators is based on hyperspectral remote sensing data and geochemical data included in the preprocessed data. One method for extracting feature indicators is calculating the mineralization alteration mineral index. For hyperspectral remote sensing data, the depth absorption characteristics of specific alteration minerals related to mineralization in the shortwave infrared band are calculated. For a sample, the reflectance values of the central band and two shoulder bands corresponding to the characteristic absorption valley of a specific alteration mineral are selected. An alteration mineral index value is obtained by calculating the ratio of the reflectance of the absorption valley band to the average reflectance of the two shoulder bands. The magnitude of this index value reflects the spectral development intensity of the alteration mineral in the sample. Another method for extracting feature indicators is calculating the intensity of geochemical element assemblage anomalies. For multiple elements closely related to the target mineral, their normalized content values are calculated, and the content information of these elements is combined into a geochemical composite anomaly value through a weighted summation method. In the weighted summation, the weight of each element is set according to its statistical correlation with mineralization, with elements having higher statistical correlations assigned higher weights. The alteration mineral index extracted from hyperspectral remote sensing data and the geochemical anomaly value extracted from geochemical data are averaged to form a final comprehensive characteristic index value. During averaging, the alteration mineral index and the geochemical anomaly value are given the same weight. This comprehensive characteristic index value is the characteristic index reflecting the intensity of mineralization information.
[0065] The samples are divided into multiple preset difficulty levels based on feature indicators. The difficulty level categorizes the samples according to the ease or difficulty of identifying mineralization information. The preset number of difficulty levels can be set to three, representing high-difficulty, medium-difficulty, and low-difficulty samples, respectively. The division is based on the comprehensive feature indicator value and the sample's true mineralization label. The division process involves setting a threshold. For samples known to be mineralized, the difficulty level is determined by their comprehensive feature indicator value. A high mineralization intensity threshold is set. Mineralized samples with comprehensive feature indicator values higher than the high mineralization intensity threshold are classified as low-difficulty samples. A low mineralization intensity threshold is set. Mineralized samples with comprehensive feature indicator values lower than the high mineralization intensity threshold but higher than the low mineralization intensity threshold are classified as medium-difficulty samples. Mineralized samples with comprehensive feature indicator values lower than the low mineralization intensity threshold are classified as high-difficulty samples. For samples known to be non-mineralized, the difficulty level division logic is the opposite of that for mineralized samples. A high background interference threshold is set. Non-mineralized samples with comprehensive characteristic index values higher than the high background interference threshold are classified as high-difficulty samples. A low background interference threshold is set. Non-mineralized samples with comprehensive characteristic index values lower than the high background interference threshold but higher than the low background interference threshold are classified as medium-difficulty samples. Non-mineralized samples with comprehensive characteristic index values lower than the low background interference threshold are classified as low-difficulty samples. The specific values of the high mineralization intensity threshold, low mineralization intensity threshold, high background interference threshold, and low background interference threshold are determined by statistical quantile analysis of the comprehensive characteristic index values of all known samples in the study area. The comprehensive characteristic index values of all mineralized samples are sorted in ascending order, and the value at the 75th percentile is taken as the high mineralization intensity threshold, and the value at the 25th percentile is taken as the low mineralization intensity threshold. For non-mineralized samples, their comprehensive characteristic index values are sorted in ascending order, and the value at the 75th percentile is taken as the high background interference threshold, and the value at the 25th percentile is taken as the low background interference threshold. Based on the judgment of characteristic index values and thresholds, all samples are classified into multiple preset difficulty levels.
[0066] Samples are drawn from each difficulty level according to a preset ratio and allocated to the training set, validation set, and test set, respectively. The preset ratio is a pre-defined allocation of sample sizes for the training, validation, and test sets; for example, 60% for the training set, 20% for the validation set, and 20% for the test set. The sampling process is conducted independently within each difficulty level. For a specific difficulty level, the total number of samples within that level is counted. The theoretical number of samples that level should be allocated to the training, validation, and test sets is calculated based on the preset ratio. For example, if a difficulty level has 100 samples, 60 samples should be allocated to the training set, 20 to the validation set, and 20 to the test set. Using random sampling, the required number of samples for the training set are drawn without replacement from the total samples for that level, and these drawn samples are marked as belonging to the training set. The required number of samples for the validation set are drawn without replacement from the remaining samples for that level, and these samples are marked as belonging to the validation set. The remaining samples for that level are then all assigned to the test set. For each preset difficulty level, the random sampling and allocation process is repeated. Samples assigned to the training set for all difficulty levels are merged to form a complete training set; samples assigned to the validation set for all difficulty levels are merged to form a complete validation set; and samples assigned to the test set for all difficulty levels are merged to form a complete test set.
[0067] Extrapolated region samples are extracted from the preprocessed data, separating them spatially from the regions corresponding to the training, validation, and test sets. Spatially independent means that the geographic area represented by the extrapolated region samples does not overlap with the regions represented by any samples in the training, validation, and test sets. The method for dividing the extrapolated region samples involves selecting one or more independent sub-regions within the entire study area covered by the preprocessed data. These sub-regions are spatially discontinuous with the main areas where known mineral deposits or mineral occurrences are distributed within the study area. All samples located within these selected sub-regions are extracted as a whole to form the extrapolated region samples. During extraction, the original feature data and label information of these samples are ensured to be complete. The extrapolated region samples are used as an independent whole for testing the generalization ability of the final model performance. The training set, validation set, test set, and extrapolated region samples together constitute a complete sample library for subsequent model training and fair comparative evaluation.
[0068] S3: Train multiple deep learning mineral exploration prediction models using training and validation sets. The specific implementation is as follows:
[0069] The training set and validation set are two independent subsets of the sample library constructed according to the aforementioned steps, used respectively for learning model parameters and monitoring the training process. A deep learning mineral exploration prediction model refers to a mathematical model built on a deep neural network architecture to predict the probability of mineralized prospective areas from multi-source mineral exploration data. Multiple deep learning mineral exploration prediction models mean using various neural network models with different structures simultaneously in comparative experiments, such as convolutional neural networks, recurrent neural networks, and graph neural networks. Each model will be trained independently.
[0070] To establish identical initial training conditions and hyperparameter search strategies for multiple deep learning-based mineral exploration prediction models, the initial training conditions refer to a series of fixed parameters that need to be pre-set before model training begins. These parameters must be consistent across all models participating in the comparison to ensure a fair starting point for training. Initial training conditions include the selection of the optimizer, the initial value of the learning rate, the batch size, and the upper limit on the total number of training epochs. For example, the optimizer can be an adaptive moment estimation optimizer. The initial learning rate can be set to 0.001. The batch size can be set according to the computer's memory capacity, such as 32 or 64. The upper limit on the total number of training epochs can be set to 100 epochs. The hyperparameter search strategy refers to the method of systematically adjusting certain key parameters to find the optimal performance of the model. These hyperparameters to be searched may include the learning rate, the dropout rate in the network layers, the number of convolutional kernels, etc. The hyperparameter search strategy can employ grid search or random search. For example, using a random search strategy means defining a reasonable range of values for each hyperparameter to be searched, such as a logarithmically uniform sampling of the learning rate between 0.0001 and 0.01, and a uniform sampling of the dropout rate between 0.1 and 0.5. In each independent training round, a set of specific values is randomly selected from the range of each hyperparameter as the configuration for this training. Each deep learning mineral exploration prediction model independently performs the same hyperparameter random search process and records the performance results under each set of hyperparameter configurations.
[0071] The training set is used to iteratively optimize the parameters of multiple deep learning mineral exploration prediction models, and the validation set is used to monitor the model performance during the iterative optimization process. Parameter iterative optimization refers to the process of gradually adjusting the internal weight parameters of the model by repeatedly iterating through the training data. One complete iteration is called a training epoch. In a training epoch, the samples in the training set are first divided into multiple small batches according to a set batch size. For each small batch, forward propagation is performed. Forward propagation means that the feature data of the small batch samples are input into the current model structure, processed by linear transformations and nonlinear activation functions of each layer of the network, and finally outputting the predicted probability value for each sample. Then, based on the model's predicted probability value and the true label of the sample, the loss function value is calculated. The loss function is used to measure the degree of error in the model's prediction; for example, the binary cross-entropy loss function can be used. Next, backpropagation is performed. Backpropagation calculates the gradient of the loss function with respect to the weight parameters of each layer of the model, layer by layer, from the output layer to the input layer, based on the loss function value and using the chain rule. The gradient represents the direction and rate of change of the loss function with respect to the parameters. Finally, based on the calculated gradient, the optimizer updates the model's weight parameters. The optimizer updates parameters by subtracting the product of the learning rate and the gradient from the current parameter values, thereby adjusting the loss function value in a decreasing direction. After each training epoch, the performance of the optimized model in that epoch is monitored using a validation set. The monitoring process involves inputting all samples from the validation set into the current model, performing forward propagation to obtain prediction results, and then calculating the model's performance metrics on the validation set, such as accuracy or F1 score. The current training epoch, the corresponding hyperparameter configuration, and the validation set performance metrics are recorded. This training epoch process is repeated until the preset maximum number of training epochs is reached. Throughout the iterative optimization process, the validation set does not participate in the calculation and updating of parameter gradients; its role is solely to provide performance feedback on unseen data, guiding the selection of hyperparameters and determining whether training should be stopped.
[0072] Training is complete when the model performance meets the preset convergence or early stopping conditions, resulting in the trained deep learning mineral exploration prediction model. The convergence condition refers to the model's loss function value on the training set decreasing to a stable and sufficiently low level, or the validation set performance metric increasing to a stable and sufficiently high level, and then remaining largely unchanged in subsequent training epochs. For example, a convergence threshold for the loss function value can be preset; when the training set loss function value falls below this threshold, the model is considered to have converged. The convergence threshold can be empirically set based on the task difficulty and the theoretical minimum value of the loss function. Early stopping is a training strategy to prevent overfitting. Specifically, it involves continuously monitoring the model's performance metric on the validation set. When the validation set performance metric stops improving or even begins to decline in several consecutive training epochs, the training process is terminated early, even before the maximum number of training epochs has been reached. For example, a preset early stopping patience value, such as 10 training epochs, can be used. During training, a record of the best performance on the validation set and a count of epochs since that record were maintained. Each time the validation set performance metric exceeds the historical best record after a new training epoch, the best record is updated and the count is reset to zero. If the validation set performance metric does not exceed the best record, the count is incremented. When this count reaches a preset early stopping threshold, the early stopping condition is triggered, and the training process terminates immediately. Upon training termination, the model weight parameters corresponding to the epoch with the best validation set performance are used as the final training result. For each deep learning mineral exploration prediction model, and for each independent hyperparameter search experiment, the above convergence condition or early stopping condition is applied independently to determine the training endpoint. After training is complete, the final network structure and weight parameters of each model are saved; these saved model instances constitute the trained deep learning mineral exploration prediction model. The trained deep learning mineral exploration prediction model possesses the ability to predict mineralization prospects from new multi-source mineral exploration data and will be used for subsequent performance evaluation and decision logic analysis.
[0073] S4: The performance of the trained deep learning mineral exploration prediction model is evaluated using the test set and extrapolated region samples to obtain multi-dimensional performance evaluation results. The specific implementation is as follows:
[0074] The test set and the extrapolated region samples are two independent subsets of the sample library constructed according to the aforementioned steps. The trained deep learning mineral exploration prediction model refers to multiple neural network models whose parameters have been optimized and whose weights have been fixed.
[0075] The test set and extrapolated region samples are input into each trained deep learning mineral exploration prediction model to obtain the corresponding prediction results. The input process involves feeding the feature data of each sample in the test set or extrapolated region samples into the model's input layer. The model performs forward propagation calculations on the input feature data. The output layer generates one or more numerical values. For binary classification mineral exploration prediction tasks, the output layer uses the sigmoid activation function, which maps the raw output value of the neuron to a real number between 0 and 1. This real number is interpreted as the predicted probability that the sample belongs to the mineralization category. This predicted probability value obtained after inputting a sample into the model is the preliminary prediction result for that sample. For each sample in the test set or extrapolated region samples, the above input and forward propagation process is repeated to obtain the set of predicted probability values for all samples. To convert the predicted probability values into the final category label, a classification decision threshold needs to be set. The classification decision threshold is a value between 0 and 1, used to define the lower limit of the predicted probability belonging to the mineralization category. The threshold can be adjusted based on the model's performance on the validation set to balance precision and recall. However, to simplify the evaluation process, an empirical value of 0.5 is often used directly as the classification decision threshold because 0.5 is the midpoint of the sigmoid function's output range, representing a completely uncertain boundary. When a sample's predicted probability is greater than or equal to the classification decision threshold, it is classified as a predicted mineralized category; when its predicted probability is less than the threshold, it is classified as a predicted non-mineralized category. Each trained deep learning mineral exploration prediction model produces two sets of results for the test set and extrapolated region samples: one set is a continuous set of predicted probability values, and the other set is a binary set of predicted category labels.
[0076] Based on the prediction results and ground truth labels, indicators reflecting the model's basic prediction accuracy, its ability to spatially characterize mineralized areas, and its generalization ability in unseen areas are calculated to obtain multi-dimensional performance evaluation results. Ground truth labels are pre-known category markers in the sample library. The calculation process is performed separately for the test set and the extrapolated area samples, and the corresponding indicators are calculated independently for each trained deep learning mineral exploration prediction model.
[0077] Calculate the metrics reflecting the model's basic prediction accuracy. The basic prediction accuracy metric measures the model's overall ability to correctly classify samples, primarily calculated based on the confusion matrix between the predicted class labels and the true labels. The confusion matrix is a 2x2 table, where rows represent the true class and columns represent the predicted class. The four cells record the number of true positives, false positives, true negatives, and false negatives, respectively. True positives are samples that are actually mineralized but predicted as mineralized by the model. False positives are samples that are actually non-mineralized but predicted as mineralized by the model. True negatives are samples that are actually non-mineralized but predicted as non-mineralized by the model. False negatives are samples that are actually mineralized but predicted as non-mineralized by the model. Based on the confusion matrix, calculate the accuracy metric. Accuracy is the proportion of correctly classified samples to the total number of samples. It is calculated by adding the number of true positives and true negatives, then dividing by the total number of samples in the test set or extrapolation region. Calculate the precision metric. Precision refers to the proportion of samples that are actually mineralized out of all samples predicted as mineralized by the model. It is calculated by dividing the number of true positives by the sum of the number of true positives and false positives. Recall is the proportion of samples that are correctly predicted as mineralized out of all samples that are actually mineralized. It is calculated by dividing the number of true positives by the sum of the number of true negatives and false negatives. The F1 score is the harmonic mean of precision and recall, used to balance these two metrics. It is calculated by multiplying 2 by precision and recall, then dividing by the sum of precision and recall. These basic prediction accuracy metrics quantify the model's classification performance from different perspectives.
[0078] The model's ability to spatially characterize mineralized regions is calculated. This metric primarily assesses the model's accuracy in predicting the spatial extent of mineralized regions, with the core indicator being the Intersection over Union (IoU). The IoU is calculated based on the model's predictions for the test set and the spatially binarized map of the ground truth labels. First, the model's predicted probability values for all samples in the entire test set are reconstructed into a prediction probability map based on their spatial coordinates. A spatial binarization probability threshold is set, which is used to convert the continuous prediction probability map into a binary map of mineralized / non-mineralized regions. The spatial binarization probability threshold can be the same as the aforementioned classification decision threshold, such as 0.5, or it can be set independently based on statistical analysis of the prediction probability distribution, for example, using the median of the predicted probabilities for all samples as the spatial binarization probability threshold. Pixels in the prediction probability map with probability values greater than or equal to the spatial binarization probability threshold are classified as predicted mineralized regions, and pixels with probability values less than the spatial binarization probability threshold are classified as predicted non-mineralized regions, thus obtaining a binarized map of predicted mineralized regions. The ground truth labels are also constructed into a binarized map of the actual mineralized regions based on their spatial coordinates. To calculate the intersection-union ratio (IUGR), first, identify the spatially overlapping areas between the predicted and actual mineralization regions—regions where corresponding pixel values are all 1 in both maps. The total number of pixels in this region is the intersection count. Then, find the union of the spatial extents occupied by all pixels marked as mineralized in both maps—regions where at least one corresponding pixel value is 1 in both maps. The total number of pixels in this region is the union count. The IUGR is the ratio of the intersection count to the union count.
[0079] A metric reflecting the model's generalization ability in unseen regions is calculated. This metric is specifically used to evaluate the model's performance on extrapolated region samples, with the core metric being the extrapolated region intersection-union ratio (IU). The calculation principle of the extrapolated region IU is exactly the same as that of the IU, but the data objects relied upon for the calculation are the extrapolated region samples and their corresponding model prediction results. The predicted probability values of the extrapolated region samples from the trained deep learning mineral exploration prediction model are reconstructed into a predicted probability map of the extrapolated region, and then binarized into a predicted mineralization region map using the same spatial binarization probability threshold as when calculating the IU. The true labels of the extrapolated region samples are used to construct the true mineralization region map of the extrapolated region. The number of intersection pixels and the number of union pixels between this predicted mineralization region map and the true mineralization region map are calculated, and the difference between the two yields the extrapolated region IU. In addition, the generalization ability can also be evaluated by comparing the relative decrease in the model's baseline prediction accuracy on the test set with that on the extrapolated region samples. For example, the difference between the F1 score of the model on the test set and the F1 score on the extrapolated region samples can be used as an additional measure of generalization stability. By systematically organizing all the calculated metrics according to the model and dataset, a multi-dimensional performance evaluation result is formed to comprehensively evaluate each trained deep learning mineral exploration prediction model.
[0080] S5: Perform a decision logic consistency analysis on the trained deep learning mineral exploration prediction model to obtain the decision logic consistency comparison results. The specific implementation is as follows:
[0081] The test set is used to select consensus samples with consistent predictions and disputed samples with differing predictions, forming a subset for analysis. The test set is an independent subset of the sample library constructed according to the aforementioned steps, containing multiple samples, each with a known true label. Each trained deep learning mineral exploration prediction model generates a predicted class label for each sample in the test set. Consensus samples with consistent predictions are those where all trained deep learning mineral exploration prediction models involved in the comparison predict the same class label for a given sample. Disputed samples with differing predictions are those where different trained deep learning mineral exploration prediction models do not predict the same class label for a given sample. The specific selection process is as follows: For each sample in the test set, check the predicted class labels for that sample from all trained deep learning mineral exploration prediction models. If all trained deep learning mineral exploration prediction models predict the same class label, the sample is marked as a consensus sample. If at least two trained deep learning mineral exploration prediction models predict different class labels, the sample is marked as a disputed sample. To ensure the representativeness and appropriate size of the sample subset, a certain number of samples are randomly selected from both the consensus sample and the disputed sample. For example, 20 samples are randomly selected from the consensus sample and 20 samples are randomly selected from the disputed sample to form the sample subset used for analysis. The number of randomly selected samples can be adjusted according to actual needs, but it should be ensured that both types of samples account for a certain proportion in the sample subset.
[0082] Model interpretability techniques are used to extract the visual or spectral features upon which each trained deep learning mineral exploration prediction model bases its decisions for each sample in a subset of samples, thus obtaining the decision-making basis features for each model. Model interpretability techniques refer to techniques that can reveal the input features upon which a deep learning model relies when making a specific prediction. Model interpretability techniques include gradient-weighted class activation mapping (GFRP) or attention weight visualization techniques. For trained deep learning mineral exploration prediction models using convolutional neural networks, gradient-weighted class activation mapping is used. The implementation process of gradient-weighted class activation mapping is as follows: Given a sample and a trained deep learning mineral exploration prediction model, the sample is input into the trained deep learning mineral exploration prediction model for forward propagation until the last convolutional layer of the trained deep learning mineral exploration prediction model. Simultaneously, the gradient of the trained deep learning mineral exploration prediction model's output for the predicted class of the sample is calculated. These gradients are then propagated back to the last convolutional layer to obtain the gradient weights of each feature map. The feature maps of the last convolutional layer are then weighted and summed; the weights are the mean gradients of the corresponding feature maps. The weighted feature map is linearly interpolated and upsampled to match the original input image size, and then normalized to obtain a heatmap. This heatmap is the gradient-weighted class activation map. Brighter areas in the heatmap represent regions that positively contribute to the prediction of mineralization categories by the trained deep learning mineral exploration prediction model, and are the visual features upon which the model's decisions are based. For trained deep learning mineral exploration prediction models with attention mechanisms, such as the Transformer, attention weight visualization techniques are employed. The implementation process of attention weight visualization is as follows: Given a sample and a trained deep learning mineral exploration prediction model, the sample is input into the model, and the attention weight matrix in the key attention layer is extracted. The attention weight matrix reflects the correlation strength between different input features when the model processes the sample. For visual tasks, attention weights can be mapped back to the spatial location of the input image, generating an attention heatmap. For spectral sequence tasks, attention weights can be mapped to different spectral bands, generating band importance curves. Both gradient-weighted activation maps and attention heatmaps visually demonstrate which parts of the input data the trained deep learning mineral exploration prediction model focuses on when making decisions. These visualizations are then used as the decision-making features of the trained deep learning mineral exploration prediction model for individual samples. For each sample in the subset, and for each trained deep learning mineral exploration prediction model, the aforementioned interpretability techniques are applied to obtain the corresponding decision-making features. These decision-making features can be represented as a heatmap in image form or a weight vector in numerical form.
[0083] The consistency measure of decision-making criteria features between deep learning mineral exploration prediction models trained at different times is calculated and compared to obtain the results of the consistency comparison of decision logic. The consistency measure quantifies the similarity of two trained deep learning mineral exploration prediction models in terms of decision-making criteria features. For image-based decision-making criteria features, such as heatmaps, the structural similarity index is used as the consistency measure. The structural similarity index is calculated based on the brightness, contrast, and structural information of the two heatmaps. Specifically, the two heatmaps are divided into multiple local windows, and the mean brightness, variance, and covariance of brightness are calculated within each window. Then, these statistics are combined to calculate the structural similarity of that window. Finally, the average of the structural similarities of all windows is taken to obtain the overall structural similarity index. The structural similarity index ranges from 0 to 1; a larger value indicates greater similarity between the two heatmaps. For numerical decision-making criteria features, such as band importance vectors, cosine similarity is used as the consistency measure. Cosine similarity calculates the difference in direction between two vectors, without considering their absolute magnitude. Specifically, it is calculated by dividing the dot product of the two vectors by the product of their respective norms. The cosine similarity value ranges from -1 to 1, with values closer to 1 indicating more consistent directions between the two vectors. For each sample in the sample subset, a set of similarity values is obtained by calculating the consistency of decision-making features between all trained deep learning mineral exploration prediction models. The mean and standard deviation of these similarity values are calculated for both consensus and disputed samples. For consensus samples, a higher consistency of decision-making features among different trained deep learning mineral exploration prediction models is expected, indicating that the models focus on similar features in easily identifiable samples. For disputed samples, a lower consistency of decision-making features may indicate that the models focus on different features in difficult-to-identify samples. By comparing the consistency statistics on consensus and disputed samples, the degree of consistency in decision-making logic among different trained deep learning mineral exploration prediction models is assessed. Finally, these statistical results and analytical conclusions are compiled into a comparison of decision-making logic consistency, which may include tables, charts, and textual descriptions to illustrate the degree of consistency in the decision-making logic of different trained deep learning mineral exploration prediction models and under what circumstances disagreements may occur.
[0084] S6: Generate a fair comparison analysis report by integrating the multi-dimensional performance evaluation results with the consistency comparison results of the decision-making logic. The specific implementation is as follows:
[0085] The multi-dimensional performance evaluation results are integrated with the decision logic consistency comparison results. The multi-dimensional performance evaluation results are a series of quantitative indicators calculated for each trained deep learning mineral exploration prediction model. These indicators include, but are not limited to, accuracy, precision, recall, F1 score, and intersection-over-union ratio (IoU) on the test set, as well as corresponding indicators and IoU of the extrapolated region samples. The decision logic consistency comparison results are the conclusions of the decision similarity analysis between different trained deep learning mineral exploration prediction models, including statistical values of the consistency measures of decision-making basis features calculated on consensus and disputed samples, such as the average structural similarity index or average cosine similarity. The integration process organizes these heterogeneous data and information into a unified structured framework for systematic analysis. Specifically, a comprehensive evaluation data table is created. In this comprehensive evaluation data table, each row represents a trained deep learning mineral exploration prediction model participating in the comparison. Each column represents a specific evaluation dimension, including key indicator columns selected from multi-dimensional performance evaluation results, such as test set F1 score, test set intersection-union ratio, and extrapolation region intersection-union ratio, as well as key indicator columns selected from decision logic consistency comparison results, such as average consistency metric of consensus samples and average consistency metric of disputed samples. The numerical values of these indicators for each trained deep learning mineral exploration prediction model are filled into the corresponding cells of the comprehensive evaluation data table. In addition, qualitative analysis descriptions of decision differences between models in the decision logic consistency comparison results, such as textual summaries of differences in the features that different models focus on in disputed samples, are stored as supplementary information and associated with the comprehensive evaluation data table. Through this tabular and information-associative approach, the integration of multi-dimensional performance evaluation results and decision logic consistency comparison results is achieved.
[0086] Based on the integrated results, the performance of different trained deep learning mineral exploration prediction models across various performance dimensions is ranked and compared. This, combined with differences in the degree of consistency in decision-making logic, generates a fair comparative analysis report that includes a comprehensive ranking of model performance and an assessment of decision reliability. The ranking and comparative analysis process is a systematic comparison. First, each performance indicator column in the comprehensive evaluation data table is independently ranked. For example, for the F1 score column of the test set, all trained deep learning mineral exploration prediction models are ranked from highest to lowest according to their F1 scores, with the top-ranked model performing best in that indicator. Similarly, each performance indicator, such as the intersection-union ratio of the test set and the intersection-union ratio of the extrapolated regions, is independently ranked in a similar manner, generating a model ranking sequence for each indicator. To obtain a comprehensive evaluation, a comprehensive score for each trained deep learning mineral exploration prediction model can be calculated. The comprehensive score is typically calculated using a weighted average method. Specifically, a weight is assigned to each performance indicator participating in the comprehensive score. The weight assignment can be based on the importance of the indicator. For example, the F1 score on the test set and the intersection-union ratio (IU) of the extrapolated regions can be considered core indicators and assigned high weights, such as 0.4 each; the IU can be considered an important indicator and assigned a weight of 0.2. The sum of all weights is 1. Then, the raw values in each performance indicator column are normalized to eliminate differences in the units and ranges of different indicators. Normalization can be performed using range normalization, mapping the maximum value to 1, the minimum value to 0, and other values to a linear range between 0 and 1. Next, the normalized values for each indicator are multiplied by their corresponding weights, and all weighted values are summed to obtain the model's overall score. Finally, all trained deep learning mineral exploration prediction models are ranked from highest to lowest based on their overall scores to obtain a ranking of the models' overall performance.
[0087] Based on ranking and comparative analysis, and considering the differences in the degree of consistency of decision logic, a decision reliability assessment is generated. The decision reliability assessment is a qualitative or semi-quantitative evaluation of the reliability of the model's prediction results. Key data in the comparison results of decision logic consistency are analyzed, particularly the average consistency metric for consensus samples and the average consistency metric for disputed samples. The average consistency metric for consensus samples reflects the degree of convergence of the model's decision logic on simple and clear samples. A high consistency threshold is set. If the average consistency metric of a trained deep learning mineral exploration prediction model on consensus samples is higher than the high consistency threshold compared to most other models, it indicates that the model's decision logic is consistent with mainstream judgments in simple cases, and its prediction results have high reliability under normal circumstances. Conversely, if its average consistency metric is low, it indicates that its decision logic may be unconventional, and it is necessary to examine whether it relies on unconventional features. The average consistency metric for disputed samples reflects the degree of difference in the model's decision logic on complex and ambiguous samples. This metric is usually low. By comparing the relative values of different models on this metric, the stability of their decisions when dealing with difficult samples can be evaluated. If a model's average consistency metric with other models on controversial samples is significantly lower than other models, it may mean that its decision-making logic is more unstable or more prone to making unique judgments when faced with ambiguous information. This could indicate stronger exploratory capabilities in some cases, but also higher randomness. When generating a decision reliability assessment, the above consistency analysis should be combined with the model's overall performance ranking. For example, for a model ranking high overall and with a high average consistency metric on consensus samples, a assessment could be given as "This model has excellent overall predictive performance, and its decision-making logic is highly consistent with mainstream models on clear cases, resulting in high reliability of the prediction results." For a model ranking high overall but with a low average consistency metric on consensus samples, a assessment could be given as "This model has excellent overall predictive performance, but its decision-making logic exhibits some uniqueness; further analysis is recommended to determine whether the features it relies on have a reasonable geological explanation." For a model ranking low overall and with a very low average consistency metric on controversial samples, a assessment could be given as "This model has poor overall predictive performance, and its decision-making logic fluctuates greatly on complex cases, resulting in low reliability of the prediction results."
[0088] The high consistency threshold is set based on statistical analysis of the consistency measures of decision-making criteria among all models calculated on the consensus samples. For example, the mean and standard deviation of these consistency measures can be calculated, and the high consistency threshold can be set as the mean plus one standard deviation. Alternatively, by analyzing the cumulative distribution of the consistency measures, a higher quantile can be selected as the threshold, such as the 75th percentile. The specific value of this threshold is not fixed, but dynamically determined according to the actual distribution of the consistency measures in each experiment.
[0089] This report systematically organizes the overall performance ranking of models, the independent ranking results for each performance dimension, and the decision reliability assessment for each model, forming a comprehensive and fair comparative analysis report. The report can be presented in a combination of text, tables, and charts. For example, the report could begin with a summary table of overall performance rankings, followed by sections on the ranking and analysis of each sub-indicator, along with detailed data on the consistency of decision logic and corresponding reliability assessments. The core conclusion of the report is to identify which one or more trained deep learning mineral exploration prediction models perform best in terms of performance and reliability under specific data conditions and evaluation frameworks, as well as the respective advantages, disadvantages, and applicable scenarios of different models. This provides a decision-making basis based on multiple constraints and in-depth analysis for model selection in mineral exploration prediction practice.
[0090] Example 2: Figure 2 A schematic diagram of the fair comparative evaluation system for deep learning mineral exploration models based on multiple constraints is provided. The system includes the following modules:
[0091] The data processing module is used to preprocess multi-source mineral exploration data to obtain preprocessed data;
[0092] The sample library construction module is used to classify the difficulty level according to the intensity of mineralization information based on the preprocessed data, and construct a sample library containing training set, validation set, test set and extrapolation region samples accordingly.
[0093] The model training module is used to train multiple deep learning mineral exploration prediction models using training and validation sets.
[0094] The performance evaluation module is used to evaluate the performance of the trained deep learning mineral exploration prediction model using the test set and extrapolated regional samples, and obtain multi-dimensional performance evaluation results.
[0095] The logic analysis module is used to perform decision logic consistency analysis on the trained deep learning mineral exploration prediction model and obtain the decision logic consistency comparison results.
[0096] The report generation module is used to generate a fair comparative analysis report by integrating the results of multi-dimensional performance evaluation with the results of the consistency of decision-making logic.
[0097] All calculations involved in the embodiments are dimensionless numerical calculations, and the preset parameters and thresholds in the calculations are set by those skilled in the art according to the actual situation.
[0098] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.
[0099] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and inventive constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0100] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0101] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0102] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0103] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A fair comparative evaluation method for deep learning-based mineral exploration models based on multiple constraints, characterized in that, Includes the following steps: S1: Preprocess the multi-source mineral exploration data to obtain preprocessed data; S2: Based on the preprocessed data, the difficulty level is divided according to the intensity of mineralization information, and a sample library containing training set, validation set, test set and extrapolation area samples is constructed accordingly. S3: Train multiple deep learning mineral exploration prediction models using training and validation sets; S4: The performance of the trained deep learning mineral exploration prediction model is evaluated using the test set and extrapolated regional samples to obtain multi-dimensional performance evaluation results. S5: Perform a decision logic consistency analysis on the trained deep learning mineral exploration prediction model and obtain the decision logic consistency comparison results. S6: Generate a fair comparison analysis report by integrating the multi-dimensional performance evaluation results with the consistency of decision-making logic.
2. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints as described in claim 1, characterized in that, Multi-source mineral exploration data is preprocessed to obtain preprocessed data, including: The hyperspectral remote sensing data in the multi-source mineral exploration data is spatially aligned with other auxiliary data, and the aligned data is then subjected to noise removal processing. The data after noise removal is normalized to obtain preprocessed data.
3. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints according to claim 2, characterized in that, Multi-source mineral exploration data includes at least two of the following: hyperspectral remote sensing data, multispectral remote sensing data, geochemical data, and geophysical data.
4. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints as described in claim 1, characterized in that, Based on the preprocessed data, the difficulty levels are divided according to the intensity of mineralization information, and a sample library is constructed accordingly, including a training set, a validation set, a test set, and extrapolated regional samples, including: Extract feature indicators reflecting the intensity of mineralization information from the preprocessed data; The samples are divided into multiple preset difficulty levels based on feature indicators; Samples are drawn from each difficulty level according to a preset ratio and allocated to the training set, validation set, and test set respectively. And data from the preprocessed data that are spatially independent of the regions corresponding to the training set, validation set, and test set are used as extrapolation region samples.
5. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints according to claim 1, characterized in that, Multiple deep learning mineral exploration prediction models were trained using training and validation sets, including: Set the same initial training conditions and hyperparameter search strategy for multiple deep learning mineral exploration prediction models; The training set was used to iteratively optimize the parameters of multiple deep learning mineral exploration prediction models, and the validation set was used to monitor the model performance during the iterative optimization process. When the model performance meets the preset convergence or early stopping conditions, the training is completed, and the trained deep learning mineral exploration prediction model is obtained.
6. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints according to claim 1, characterized in that, The performance of the trained deep learning mineral exploration prediction model was evaluated using the test set and extrapolated region samples, yielding multi-dimensional performance evaluation results, including: The test set and the extrapolated region samples are respectively input into each trained deep learning mineral exploration prediction model to obtain the corresponding prediction results. Based on the prediction results and real labels, indicators reflecting the basic prediction accuracy of the model, the ability of the model to spatially characterize mineralized areas, and the ability of the model to generalize in unseen areas are calculated to obtain multi-dimensional performance evaluation results.
7. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints according to claim 1, characterized in that, A decision logic consistency analysis was performed on the trained deep learning mineral exploration prediction model, and the results of the decision logic consistency comparison were obtained, including: Consensus samples with consistent prediction results and disputed samples with differing prediction results are selected from the test set to form a sample subset for analysis. By using model interpretability techniques, the visual or spectral features on which each trained deep learning mineral exploration prediction model makes decisions for each sample in the sample subset are extracted, thus obtaining the decision-making basis features corresponding to each model. The consistency measure of decision-making basis features among different trained deep learning mineral exploration prediction models is calculated and compared to obtain the results of the consistency comparison of decision logic.
8. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints according to claim 7, characterized in that, Model interpretability techniques include gradient-weighted class activation mapping or attention weight visualization techniques.
9. The fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints according to claim 1, characterized in that, Based on a comprehensive comparison of multi-dimensional performance evaluation results and decision-making logic consistency, a fair comparative analysis report is generated, including: Integrate the results of multi-dimensional performance evaluation with the results of consistency comparison of decision-making logic; Based on the integrated results, the performance of different trained deep learning mineral exploration prediction models is ranked and compared across various performance dimensions. Combined with the differences in the degree of consistency of decision logic, a fair comparative analysis report is generated, which includes a ranking of the overall model performance and an evaluation of decision reliability.
10. A fair comparative evaluation system for deep learning mineral exploration models based on multiple constraints, used to implement the fair comparative evaluation method for deep learning mineral exploration models based on multiple constraints as described in any one of claims 1-9, characterized in that, Includes the following modules: The data processing module is used to preprocess multi-source mineral exploration data to obtain preprocessed data; The sample library construction module is used to classify the difficulty level according to the intensity of mineralization information based on the preprocessed data, and construct a sample library containing training set, validation set, test set and extrapolation region samples accordingly. The model training module is used to train multiple deep learning mineral exploration prediction models using training and validation sets. The performance evaluation module is used to evaluate the performance of the trained deep learning mineral exploration prediction model using the test set and extrapolated regional samples, and obtain multi-dimensional performance evaluation results. The logic analysis module is used to perform decision logic consistency analysis on the trained deep learning mineral exploration prediction model and obtain the decision logic consistency comparison results. The report generation module is used to generate a fair comparative analysis report by integrating the results of multi-dimensional performance evaluation with the results of the consistency of decision-making logic.