Method and system for identifying long-tail distributed mineral image
By adjusting the mineral image recognition model logit fraction and inverse probability sampling, the problem of mineral image data imbalance is solved, the recognition accuracy of rare minerals is improved, and the efficient recognition of mineral images is achieved.
Patent Information
- Application Number
- CN202510483397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, there is an imbalance problem in mineral image data used to train artificial intelligence models. There are many common mineral images and few rare mineral images, resulting in low accuracy of rare mineral recognition, and it is difficult to improve the overall mineral recognition accuracy without losing the accuracy of common mineral recognition.
The cross entropy loss function adjusted by logit score was used to train the CNN or Transformer model for 60 rounds, the invariant feature learning loss function was added for 40 rounds, and the classifier was retrained for 10 rounds after freezing the feature extractor weights, and the training samples of rare minerals were increased through inverse probability sampling, and various categories of minerals were trained using the same probability sampling.
While maintaining the accuracy of common mineral recognition, it significantly improves the accuracy of rare minerals and improves the overall recognition rate of mineral images.
Smart Images

Figure CN120339712A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mineral identification, and particularly to a method and system for identifying mineral images with long-tail distribution. Background Art
[0002] Mineral image recognition based on artificial intelligence is an important method for mineral identification. However, there is an imbalance problem in the number of mineral images used to train the artificial intelligence model, that is, there are many images of common minerals and few images of rare minerals. Therefore, the recognition accuracy of common minerals is high and the recognition accuracy of rare minerals is low. However, people usually have a higher demand for the recognition of rare minerals. Therefore, there is an urgent need for a method to improve the recognition accuracy of rare minerals without sacrificing the recognition accuracy of common minerals, so as to improve the overall mineral recognition accuracy. Summary of the Invention
[0003] Therefore, the present invention provides a method and system for identifying mineral images with long-tail distribution to solve the problem of low recognition accuracy of rare minerals in the existing recognition methods.
[0004] To achieve the above object, the embodiments of the present invention provide the following technical solutions:
[0005] The present invention provides a method and system for identifying mineral images with long-tail distribution, which is characterized by including:
[0006] S1: Obtain various mineral images and divide them into a training set, a validation set, and a test set;
[0007] S2: Use the cross-entropy loss function adjusted by the logit score shown in formula (1) to train the CNN or Transformer image classification model adopted for 60 epochs. In formula (1), f y (x) is the score of the model predicting the input x as the true class y, f y' (x) is the score of the model predicting the input x as other classes y', π y is the proportion of class y in the training data, [L] is the training set, and during training, each training sample is randomly sampled for each batch;
[0008]
[0009] S3: After the training in step S2 is completed, use the loss function shown in formula (2) to train the CNN or Transformer image classification model adopted for 40 epochs. In formula (2), is the classification loss function shown in formula (1), is the invariant feature learning loss function shown in formula (3), and α is the balance and The hyperparameters in formula (3) is the feature vector of the model for the sample , θ is the learnable parameter of the model, is the average value of the feature vectors of each sample in category y at the end of S2 training, and it is updated every 20 rounds. During training, in addition to randomly sampling each sample to form the training sample set e1 for each batch, an inverse probability sampling weight s as shown in formula (4) is also added to each mineral image i . In formula (4), p i is the recognition probability of mineral image i in the previous round of training, β is the scaling factor, γ is the proportionality factor, μ i and μ tail are respectively the average recognition probabilities of the 80% samples with the smallest p head and the 20% samples with the largest p i . Therefore, the numerator of the scaling factor β is set to 4 (i.e., 8:2), that is, following the 80 / 20 rule, the sampling times of the 80% mineral pictures with low recognition accuracy are increased to obtain an additional training sample set e2. e1 and e2 are combined to form a training batch e = {e1, e2}, providing comprehensive learning samples with typical and rare attributes;
[0010]
[0011] S4: Freeze the weights of the feature extractor in the used CNN or Transformer model at the end of training in step S3, and use the logit score-adjusted cross-entropy loss function to train the fully connected layer of the classifier of the used CNN or Transformer image classification model for 10 epochs. During training, for each batch, each type of mineral is sampled with a probability of 1 / N (N is the number of mineral categories), and the samples within each type of mineral are randomly sampled;
[0012] S5: Use the trained model for actual mineral image recognition.
[0013] The embodiments of the present invention have the following advantages:
[0014] The embodiments of the present invention disclose a method and system for identifying mineral images with long-tailed distribution. By adjusting the logit score of the cross-entropy loss function of the model, adding an invariant feature learning loss after 60 rounds of training and training the mineral samples with low recognition accuracy for 40 rounds, then freezing the weights of the model feature extractor, using the logit score-adjusted cross-entropy loss function to retrain the model classifier with the same probability sampling method for each type of mineral for 10 rounds, and using the trained model for actual mineral image recognition after training, the recognition accuracy of rare minerals can be improved while maintaining the recognition accuracy of common minerals, thereby improving the overall recognition rate of mineral images. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0016] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0017] Figure 1 It is an explanatory diagram of a method for identifying mineral images with a long-tailed distribution of the present invention;
[0018] Figure 2 It is a comparison chart of the accuracy rates of a method and a system for identifying mineral images with a long-tailed distribution of the present invention. Specific Embodiments
[0019] The following specific embodiments illustrate the embodiments of the present invention. Those who are familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.
[0020] Embodiment 1
[0021] This embodiment discloses a method and a system for identifying mineral images with a long-tailed distribution. The method and system for identifying mineral images with a long-tailed distribution are as Figure 1 shown and include:
[0022] S1: Crawl 36 minerals from the mineral resources database platform (Mindat.org). After cleaning, there are a total of 113,569 mineral images, which are randomly divided into a training set, a validation set, and a test set according to a ratio of 8:1:1;
[0023] S2: Use the logit score-adjusted cross-entropy loss function shown in formula (1) to train the ResNet50 model for 60 epochs. In formula (1), f y (x) is the score of the model predicting the input x as the true class y, fy' (x) is the score that the model predicts the input x as other class y', and π y is the proportion of class y in the training data, [L] is the training set, and during training, each training sample is randomly sampled for each batch;
[0024]
[0025] S3: After the training in step S2 ends, the ResNet50 model is trained for 40 rounds using the loss function shown in formula (2). In formula (2), is the classification loss function shown in formula (1), is the invariant feature learning loss function shown in formula (3), and α is the hyperparameter that balances and , which is set to 0.0003. In formula (3), is the feature vector of the model for the sample , θ is the learnable parameter of the model, is the average value of the feature vectors of each sample in class y at the end of the S2 training. During training, in addition to randomly sampling each sample to form the training sample set e1 for each batch, an inverse probability sampling weight s shown in formula (4) is also added to each mineral image. i In formula (4), p i is the recognition probability of mineral image i in the previous round of training, β is the scaling factor, γ is the proportionality factor, μ i and μ tail are the average values of the recognition probabilities of the 80% samples with the smallest p head and the 20% samples with the largest p i respectively. Therefore, the numerator of the scaling factor β is also set to 4 (i.e., 8:2), that is, following the 80 / 20 rule, the sampling times of the 80% mineral pictures with low recognition accuracy are increased to obtain an additional training sample set e2. e1 and e2 are combined to form a training batch e = {e1, e2}, providing comprehensive learning samples of typical and rare attributes;
[0026]
[0027]
[0028] S4: When the training in step S3 ends, the weights of the feature extractor in the ResNet50 model are frozen, and the fully connected layer of the classifier of the ResNet50 model is trained for 10 epochs using the logit score adjusted cross-entropy loss function shown in formula (1). During training, for each batch, each type of mineral is sampled with a probability of 1 / N (N is the number of mineral classes, which is 36), and the samples within each type of mineral are randomly sampled;
[0029] S5: Use the trained ResNet50 model for actual mineral image recognition.
[0030] Although the present invention has been described in detail with general descriptions and specific embodiments above, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of the present invention claimed.
Claims
1. A method and system for identifying images of minerals with long-tailed distributions, characterized in that Including: S1: Obtain various mineral images and divide them into a training set, a validation set, and a test set; S2: Train the CNN or Transformer image classification model adopted for 60 epochs using the logit score-adjusted cross-entropy loss function shown in formula (1). In formula (1), f y (x) is the score predicted by the model for the input x as the true class y, and f y' (x) is the score predicted by the model for the input x as other class y'. π y is the proportion of class y in the training data, [L] is the training set, and during training, each training sample is randomly sampled for each batch; S3: After the training in step S2 ends, use the loss function shown in formula (2) to train the adopted CNN or Transformer image classification model for 40 rounds. In formula (2), is the classification loss function shown in formula (1), is the invariant feature learning loss function shown in formula (3), and α is the hyperparameter that balances and . In formula (3), is the feature vector of the model for sample , θ is the learnable parameter of the model, is the average value of the feature vectors of each sample in class y i at the end of the S2 training. During training, in addition to randomly sampling each sample to form the training sample set e1 for each batch, an inverse probability sampling weight s shown in formula (4) is also added to each mineral image. i . In formula (4), p i is the recognition probability of mineral image i in the previous round of training, β is the scaling factor, γ is the proportionality factor, μ tail and μ head are respectively the average recognition probabilities of the 80% samples with the smallest p i and the 20% samples with the largest p i . Therefore, the numerator of the scaling factor β is also set to 4 (i.e., 8:2), that is, following the 80 / 20 rule, increasing the sampling times of the 80% mineral pictures with low recognition accuracy to obtain the additional training sample set e2. e1 and e2 are combined to form a training batch e = {e1, e2}, providing comprehensive learning samples with typical and rare attributes; S4: At the end of the training in step S3, freeze the weights of the feature extractor in the used CNN or Transformer model, and use the logit score shown in formula (1) to adjust the cross-entropy loss function to train the fully connected layer of the classifier of the CNN or Transformer image classification model for 10 epochs. During training, for each batch, sample each type of mineral with a probability of 1 / N (N is the number of mineral categories), and randomly sample the samples within each type of mineral; S5: Use the trained model for actual mineral image recognition.