Mineral Identification Method, Device, and Storage Medium
By using the YOLOv8 model of global and multi-scale attention mechanism in mineral recognition, combined with edge detection and area masking treatment, the problem of time-consuming and labor-intensive and poor identification accuracy of traditional mineral recognition methods is solved, and high accuracy recognition of minerals of different sizes is achieved.
Patent Information
- Application Number
- CN202510131184.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-06
AI Technical Summary
Traditional mineral identification methods are time-consuming and labor-intensive, require high expertise, and when facing complex samples, the identification results are poorly accurate. The existing deep learning models have problems such as slow detection speed, difficulty in identifying minerals of different sizes, insufficient detection capabilities of small-particle minerals, limited feature extraction and fusion capabilities, and lack of model stability and generalization capabilities.
A mineral recognition method is adopted to input the image to be identified into the pre-trained target large-particle mineral recognition model and the target small-particle mineral recognition model, combining the global attention mechanism and multi-scale attention mechanism, large and small-particle mineral recognition mechanism are used to identify large particles and small-particle minerals, and the recognition accuracy is improved through edge detection and area masking.
Effective identification of mineral types for minerals of different sizes is achieved, the accuracy of identification is improved, the problem of time-consuming and labor-intensive and poor identification accuracy of traditional methods is overcome, and the stability and generalization ability of the model are improved.
Smart Images

Figure CN119580016B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular, to a method, device, and storage medium for mineral recognition. Background Art
[0002] Mineral recognition is a key step in mining exploration and resource development. Accurately identifying the types and distributions of minerals is of great significance for improving the utilization rate of mineral resources and exploration efficiency. Traditional mineral recognition methods mainly rely on manual observation and chemical analysis. These methods are not only time-consuming and laborious, but also require operators to have a high level of professional knowledge. Moreover, when facing complex samples, there are also significant differences in the accuracy of the recognition results.
[0003] With the development of deep learning technology, the mainly adopted deep learning technologies include Convolutional Neural Networks (CNN), Region-CNN (R-CNN), and the YOLO (You Only Look Once) series of models. However, the above models still have problems such as slow detection speed, difficulty in identifying minerals of different sizes, insufficient detection ability for small-particle minerals, limited feature extraction and fusion ability, and lack of model stability and generalization ability.
[0004] In view of this, the present invention is specifically proposed. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a method, device, and storage medium for mineral recognition, achieving the effect of effectively identifying the types of minerals for minerals of different sizes and improving the accuracy of recognition.
[0006] An embodiment of the present invention provides a method for mineral recognition, which includes:
[0007] Inputting the image to be recognized into a pre-trained target large-particle mineral recognition model to obtain the first mineral type and the target detection boxes corresponding to each first mineral type;
[0008] Performing edge detection on the image to be recognized according to each target detection box to obtain each edge detection box, and performing region masking processing on the image to be recognized according to each edge detection box to obtain an intermediate image;
[0009] Inputting the intermediate image into a pre-trained target small-particle mineral recognition model to obtain the second mineral type;
[0010] Determining the target mineral type corresponding to the image to be recognized according to the first mineral type and the second mineral type;
[0011] Among them, the target large particle mineral recognition model is a YOLOv8 model introduced with a global attention mechanism, and the target small particle mineral recognition model is a YOLOv8 model introduced with a multi-scale attention mechanism and a small target detection layer.
[0012] An embodiment of the present invention provides an electronic device, and the electronic device includes:
[0013] a processor and a memory;
[0014] The processor is used to execute the steps of the mineral recognition method described in any embodiment by calling the program or instruction stored in the memory.
[0015] An embodiment of the present invention provides a computer-readable storage medium, and the computer-readable storage medium stores a program or instruction, and the program or instruction enables a computer to execute the steps of the mineral recognition method described in any embodiment.
[0016] The embodiment of the present invention has the following technical effects:
[0017] By inputting the image to be recognized into the pre-trained target large particle mineral recognition model, the first mineral type and the target detection box corresponding to each first mineral type are obtained. To recognize large particle minerals through the target large particle mineral recognition model, that is, the YOLOv8 model introduced with a global attention mechanism. Furthermore, edge detection is performed on the image to be recognized according to each target detection box to obtain each edge detection box, and region masking processing is performed on the image to be recognized according to each edge detection box to obtain an intermediate image, so as to mask large particle minerals to prevent interference with the recognition of small particle minerals. Further, the intermediate image is input into the pre-trained target small particle mineral recognition model to obtain the second mineral type. To recognize small particle minerals through the target small particle mineral recognition model, that is, the YOLOv8 model introduced with a multi-scale attention mechanism and a small target detection layer. Finally, according to the first mineral type and the second mineral type, the target mineral type corresponding to the image to be recognized is determined, achieving the effect of effectively recognizing the mineral type for minerals of different sizes and improving the recognition accuracy. Description of the Drawings
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1It is a flowchart of a mineral identification method provided by an embodiment of the present invention;
[0020] Figure 2 It is a network structure diagram of a target large particle mineral identification model provided by an embodiment of the present invention;
[0021] Figure 3 It is a schematic structural diagram of a global attention mechanism layer provided by an embodiment of the present invention;
[0022] Figure 4 It is a schematic structural diagram of a channel attention sub-module provided by an embodiment of the present invention;
[0023] Figure 5 It is a schematic structural diagram of a spatial attention sub-module provided by an embodiment of the present invention;
[0024] Figure 6 It is a network structure diagram of a target small particle mineral identification model provided by an embodiment of the present invention;
[0025] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope protected by the present invention.
[0027] Embodiment 1
[0028] The mineral identification method provided by the embodiment of the present invention is mainly applicable to the situation of identifying mineral types by size in a mineral image. The mineral identification method provided by the embodiment of the present invention can be executed by an electronic device.
[0029] Figure 1 It is a flowchart of a mineral identification method provided by an embodiment of the present invention. Refer to Figure 1 , the mineral identification method specifically includes:
[0030] S110. Input the image to be identified into a pre-trained target large particle mineral identification model to obtain a first mineral type and a target detection frame corresponding to each first mineral type.
[0031] Among them, the image to be recognized is an image for which the mineral type needs to be recognized. The target large particle mineral recognition model is a YOLOv8 model introduced with a global attention mechanism, which is used to recognize large particle minerals in the image to be recognized. The first mineral type is various mineral types of large particle minerals in the image to be recognized. The target detection box is the detection box corresponding to each recognized large particle mineral in the image to be recognized.
[0032] Specifically, the image to be recognized is input into the pre-trained target large particle mineral recognition model, and the target large particle mineral recognition model analyzes and recognizes the mineral types of large particle minerals in the image to be recognized to obtain the first mineral type, and the target detection boxes corresponding to each first mineral type can be obtained, that is, the detection boxes corresponding to each large particle mineral are recognized.
[0033] Based on the above example, the target large particle mineral recognition model includes: a global attention layer is introduced after the basic fast spatial pyramid pooling layer in the basic backbone network of the YOLOv8 model, and the global attention layer replaces the basic fast spatial pyramid pooling layer and is connected to the corresponding basic upsampling layer and basic connection layer.
[0034] Among them, the YOLOv8 model includes a basic backbone network, a basic neck network, and a basic head network.
[0035] Based on the above example, the target large particle mineral recognition model includes: a first backbone network, a first neck network, and a first head network.
[0036] Among them, the first backbone network includes a first convolutional layer, four large particle backbone convolutional feature fusion modules, a first fast spatial pyramid pooling layer, and a global attention mechanism layer connected in sequence;
[0037] The first neck network includes a first upsampling layer connected to the global attention mechanism layer, a first connection layer connected to the first upsampling layer and the third large particle backbone convolutional feature fusion module, a first feature fusion concatenated convolutional layer connected to the first connection layer, a second upsampling layer connected to the first feature fusion concatenated convolutional layer, a second connection layer connected to the second upsampling layer and the second large particle backbone convolutional feature fusion module, a second feature fusion concatenated convolutional layer connected to the second connection layer, a second convolutional layer connected to the second feature fusion concatenated convolutional layer, a third connection layer connected to the second convolutional layer, a third feature fusion concatenated convolutional layer connected to the third connection layer, a third convolutional layer connected to the third feature fusion concatenated convolutional layer, a fourth connection layer connected to the third convolutional layer and the global attention mechanism layer, and a fourth feature fusion concatenated convolutional layer connected to the fourth connection layer;
[0038] The first head network includes a first detection layer connected to the fourth feature fusion tandem convolutional layer, a second detection layer connected to the third feature fusion tandem convolutional layer, and a third detection layer connected to the second feature fusion tandem convolutional layer.
[0039] Among them, the large-particle backbone convolutional feature fusion module includes a convolutional layer and a feature fusion tandem convolutional layer connected in sequence.
[0040] For the target large-particle mineral recognition model for large-particle minerals, a global attention mechanism is added to the backbone network (Backbone) of the baseline model YOLOv8. The global attention mechanism is used to calculate the correlation between each element in the input sequence and the target element, and then assigns weights according to these correlations, enabling the model to pay more attention to elements with higher correlations to the target. The network structure diagram of the target large-particle mineral recognition model is as Figure 2 shown. Among them, the first convolutional layer, the second convolutional layer, and the third convolutional layer are convolutional layers (Convolution, Conv). The large-particle backbone convolutional feature fusion module includes a convolutional layer and a feature fusion tandem convolutional layer connected in sequence. The first fast spatial pyramid pooling layer is the fast spatial pyramid pooling layer (Spatial Pyramid Pooling-Fast, SPPF). The global attention mechanism layer is GAM (Global Attention Mechanism), which effectively reduces the dispersion of information in the global dimension and enhances the interaction between global features. The first upsampling layer and the second upsampling layer are Upsample, which is a technique for increasing the image resolution and is mainly used to enlarge the image. The first connection layer, the second connection layer, the third connection layer, and the fourth connection layer are Concat, which are used to connect two or more arrays. The first feature fusion tandem convolutional layer, the second feature fusion tandem convolutional layer, the third feature fusion tandem convolutional layer, and the fourth feature fusion tandem convolutional layer are the concatenated convolutional layers for feature fusion (Concatenate ConvLayers with Feature Fusion, C2f), which are used for feature fusion. The first detection layer, the second detection layer, and the third detection layer are detection heads (Detect), and their functions include predicting bounding boxes and class probabilities.
[0041] The structural schematic diagram of the global attention mechanism layer is as Figure 3As shown, it is a mechanism for extracting important information in image features through global observation and modeling, and a method for enhancing the interaction between different channels and spaces. Among them, given the input feature map F1, the intermediate state feature map F2, and the output feature map F3, specifically, F1 will perform a per-channel multiplication operation with the channel attention map Mc to highlight the features of important channels; then, the obtained features will perform an element-wise multiplication operation with the spatial attention map Ms. The structural schematic diagram of the channel attention sub-module is as shown in Figure 4 As shown, the channel attention sub-module uses a 3D permutation method to store information in three dimensions. After that, a two-layer Multilayer Perceptron (MLP) is used to enhance the connection between channels and spaces. The structural schematic diagram of the spatial attention sub-module is as shown in Figure 5 As shown, the spatial attention sub-module uses two 7×7 convolutional layers (Conv) to fuse the spatial information of the input image. To ensure consistency with the channel attention sub-module, the same reduction ratio is used for both modules. Among them, C is the Channel, that is, the channel dimension, H is the Height, that is, the image height, and W is the Width, that is, the image width.
[0042] S120. Perform edge detection on the image to be recognized according to each target detection box to obtain each edge detection box, and perform region masking processing on the image to be recognized according to each edge detection box to obtain an intermediate image.
[0043] Among them, the edge detection box is the edge box corresponding to the large particle minerals in each target detection box in the image to be recognized, and it fits the edge of the large particle minerals more closely than the target detection box. The intermediate image is the image obtained by masking the large particle minerals in the image to be recognized.
[0044] Specifically, perform edge detection on the part of each target detection box corresponding to the image to be recognized to obtain the edge of the large particle minerals, and form an edge detection box. Furthermore, perform region masking processing on the part of each edge detection box corresponding to the image to be recognized, and an intermediate image can be obtained after the masking processing.
[0045] Based on the above example, the following method can be used to perform edge detection on the image to be recognized according to each target detection box to obtain each edge detection box:
[0046] For each target detection box, perform Gaussian filtering on the detection image corresponding to the target detection box in the image to be recognized to obtain a filtered image;
[0047] Determine the strong edges and weak edges according to the preset operator and the filtered image, and determine the edge detection box corresponding to the target detection box according to the strong edges and weak edges.
[0048] Among them, the detected image is the partial image corresponding to the target detection box in the image to be recognized. The filtered image is the image obtained by performing Gaussian filtering on the detected image. The preset operator is an operator set in advance, which can be a Sobel operator, etc. The strong edge is the pixel point higher than the preset highest threshold, and the weak edge is the pixel point located between the preset highest threshold and the preset lowest threshold.
[0049] Specifically, for each target detection box, the part of the image to be recognized corresponding to the target detection box is used as the detected image, and the detected image is subjected to Gaussian filtering to obtain the filtered image. Furthermore, each pixel point in the filtered image is processed according to the preset operator to obtain the gradient value of each pixel point, and the gradient value of each pixel point is compared with the preset highest threshold and the preset lowest threshold respectively. Furthermore, the strong edge and the weak edge corresponding to the detected image can be obtained, and by connecting the strong edge and the weak edge connected thereto, the edge detection box corresponding to the target detection box is obtained.
[0050] Exemplarily, first, a Gaussian filter is applied to smooth the detected image to remove noise. The role of the Gaussian filter is to reduce the random noise in the detected image through convolution operations. Furthermore, the intensity gradient and direction of the detected image at each pixel point are calculated. Among them, gradient calculation usually uses a preset operator (such as the Sobel operator) to estimate the derivatives of the detected image in the horizontal and vertical directions. Non-maximum suppression is performed on the gradient magnitude to retain the local maximum and suppress non-edge points to refine the edge and eliminate false edges in the gradient map. Further, two thresholds (the preset highest threshold and the preset lowest threshold) are used to distinguish strong edges and weak edges. Pixel points higher than the high threshold (the preset highest threshold) are considered strong edges; pixel points lower than the low threshold (the preset lowest threshold) are considered non-edges; pixel points between the high and low thresholds (the preset highest threshold and the preset lowest threshold) are considered weak edges. Finally, by connecting the strong edge and the weak edge connected thereto, a complete edge detection box is formed.
[0051] Based on the above example, the image to be recognized can be subjected to region masking processing according to each edge detection box in the following manner to obtain an intermediate image:
[0052] Delete the regional images corresponding to each edge detection box in the image to be recognized to obtain the intermediate image; or,
[0053] Fill the regional images corresponding to each edge detection box in the image to be recognized with a preset color to obtain the intermediate image.
[0054] Among them, the regional image is the partial image in the image to be recognized corresponding to the edge detection box. The preset color is a pure color set in advance, which can be white, etc.
[0055] Specifically, the partial images corresponding to the edge detection boxes in the image to be recognized are used as regional images. Pixel deletion processing can be performed on the regional images in the image to be recognized, and the remaining part is used as the intermediate image; alternatively, the regional images in the image to be recognized can be filled with a preset color, and the filled image is the intermediate image.
[0056] S130: Input the intermediate image into the pre-trained target small particle mineral recognition model to obtain the second mineral type.
[0057] Among them, the target small particle mineral recognition model is a YOLOv8 model introducing a multi-scale attention mechanism and a small target detection layer, which is used to recognize small particle minerals in the intermediate image. The second mineral type is various mineral types of small particle minerals in the image to be recognized.
[0058] Specifically, input the intermediate image into the pre-trained target small particle mineral recognition model, and analyze and recognize the mineral types of small particle minerals in the intermediate image through the target small particle mineral recognition model to obtain the second mineral type.
[0059] Based on the above example, the target small particle mineral recognition model includes: introducing a multi-scale attention mechanism layer after the basic fast spatial pyramid pooling layer in the basic backbone network of the YOLOv8 model, and the multi-scale attention mechanism layer replaces the basic fast spatial pyramid pooling layer and is connected to the corresponding basic upsampling layer and basic connection layer; introducing a small target detection layer in the basic neck network and basic head network of the YOLOv8 model.
[0060] Based on the above example, the target small particle mineral recognition model includes: a second backbone network, a second neck network, and a second head network.
[0061] Among them, the second backbone network includes a fourth convolutional layer, four small particle backbone convolutional feature fusion modules, a second fast spatial pyramid pooling layer, and a multi-scale attention mechanism layer connected in sequence;
[0062] The second neck network includes a third upsampling layer connected to the multi-scale attention mechanism layer, a fifth connection layer connected to the third upsampling layer and the third small particle backbone convolution feature fusion module, a fifth feature fusion concatenated convolution layer connected to the fifth connection layer, a fourth upsampling layer connected to the fifth feature fusion concatenated convolution layer, a sixth connection layer connected to the fourth upsampling layer and the second small particle backbone convolution feature fusion module, a sixth feature fusion concatenated convolution layer connected to the sixth connection layer, a fifth upsampling layer connected to the sixth feature fusion concatenated convolution layer, a seventh connection layer connected to the fifth upsampling layer and the first small particle backbone convolution feature fusion module, a seventh feature fusion concatenated convolution layer connected to the seventh connection layer, a fifth convolution layer connected to the seventh feature fusion concatenated convolution layer, an eighth connection layer connected to the fifth convolution layer, an eighth feature fusion concatenated convolution layer connected to the eighth connection layer, a sixth convolution layer connected to the eighth feature fusion concatenated convolution layer, a ninth connection layer connected to the sixth convolution layer, a ninth feature fusion concatenated convolution layer connected to the ninth connection layer, a seventh convolution layer connected to the ninth feature fusion concatenated convolution layer, a tenth connection layer connected to the seventh convolution layer and the multi-scale attention mechanism layer, and a tenth feature fusion concatenated convolution layer connected to the tenth connection layer;
[0063] The second head network includes a fourth detection layer connected to the seventh feature fusion concatenated convolution layer, a fifth detection layer connected to the eighth feature fusion concatenated convolution layer, a sixth detection layer connected to the ninth feature fusion concatenated convolution layer, and a seventh detection layer connected to the tenth feature fusion concatenated convolution layer.
[0064] Among them, the small particle backbone convolution feature fusion module includes a convolution layer and a feature fusion concatenated convolution layer connected in sequence.
[0065] For the target small particle mineral recognition model, a multi-scale attention mechanism is added to the backbone network (Backbone) of the baseline model YOLOv8. Aiming to retain the information on each channel and reduce the computational overhead, part of the channels are reshaped into the batch dimension, and the channel dimension is grouped into multiple sub-features, so that the spatial semantic features are evenly distributed in each feature group. Moreover, small object detection layers are added to the neck network (Neck) of the baseline model YOLOv8 to enhance the detection ability of the model for small particle minerals. The network structure diagram of the target small particle mineral recognition model is as Figure 6As shown. Among them, the fourth convolutional layer, the fifth convolutional layer, the sixth convolutional layer, and the seventh convolutional layer are convolutional layers (Convolution, Conv). The small particle backbone feature fusion tandem convolutional layer includes a convolutional layer and a feature fusion tandem convolutional layer connected in sequence. The second fast spatial pyramid pooling layer is a fast spatial pyramid pooling layer (Spatial Pyramid Pooling-Fast, SPPF). The multi-scale attention mechanism layer (Efficient Multi-Scale Attention, EMA) is used to achieve better performance while maintaining the model size and computational efficiency. The third upsampling layer, the fourth upsampling layer, and the fifth upsampling layer are Upsample. The fifth connection layer, the sixth connection layer, the seventh connection layer, the eighth connection layer, the ninth connection layer, and the tenth connection layer are Concat, which is used to connect two or more arrays. The fifth feature fusion tandem convolutional layer, the sixth feature fusion tandem convolutional layer, the seventh feature fusion tandem convolutional layer, the eighth feature fusion tandem convolutional layer, the ninth feature fusion tandem convolutional layer, and the tenth feature fusion tandem convolutional layer are concatenate convolutional layers with feature fusion (C2f), which are used for feature fusion. The fourth detection layer, the fifth detection layer, the sixth detection layer, and the seventh detection layer are detection heads (Detect), and their functions include predicting bounding boxes and class probabilities.
[0066] In the backbone network (BackBone) of the YOLOv8 model, the EMA module, that is, the multi-scale attention mechanism, is added. The EMA mechanism improves the CA (Coordinate Attention) mechanism, adds 1 parallel subnet to capture multi-scale feature relationships, and uses a cross-space learning method to fuse the feature information between the two subnets.
[0067] In the YOLOv8 model, a small object detection layer is added. Specifically, the output of the second layer of the Backbone (backbone network) is increased, and the feature map of this output is concatenated with the deepest layer feature of the FPN (Feature Pyramid Network). Subsequently, the concatenated feature is added to the path of the PAN (Path Aggregation Network) to output feature maps of four scales to the Head (head network) structure. The added fourth-scale feature map can be 160*160, which has a smaller receptive field and is suitable for detecting small objects, and can more accurately identify small particle minerals.
[0068] Based on the above examples, the target large particle mineral recognition model and the target small particle mineral recognition model are trained in the following way:
[0069] For each sample mineral image in the sample image set, obtain a sample first type, a sample second type, and a sample mineral type through an initial large particle mineral recognition model and an initial small particle mineral recognition model;
[0070] Determine each model metric parameter according to the first sample type, the second sample type, the sample mineral type, and the calibrated mineral type corresponding to each sample mineral image in the sample image set;
[0071] In response to each model metric parameter meeting the model performance requirements, use the initial large particle mineral recognition model as the pre-trained target large particle mineral recognition model, and use the initial small particle mineral recognition model as the pre-trained target small particle mineral recognition model;
[0072] In response to at least one of each model metric parameter not meeting the model performance requirements, adjust the model parameters based on each model metric parameter, the initial large particle mineral recognition model, and / or the initial small particle mineral recognition model, and return to execute the step of obtaining a sample first type, a sample second type, and a sample mineral type for each sample mineral image in the sample image set through the initial large particle mineral recognition model and the initial small particle mineral recognition model until each model metric parameter meets the model performance requirements.
[0073] Based on the above example, the following method can be used to obtain a sample first type, a sample second type, and a sample mineral type for each sample mineral image in the sample image set through an initial large particle mineral recognition model and an initial small particle mineral recognition model:
[0074] For each sample mineral image in the sample image set, input the sample mineral image into the initial large particle mineral recognition model to obtain a sample first type and a sample detection frame corresponding to each sample first type;
[0075] Perform edge detection according to each sample detection frame to obtain each mineral edge frame, and perform regional masking processing on the sample mineral image according to each mineral edge frame to obtain a sample process image;
[0076] Input the sample process image into the initial small particle mineral recognition model to obtain a sample second type;
[0077] Determine the sample mineral type according to the sample first type and the sample second type.
[0078] Among them, the sample image set is a data set used for training the target large particle mineral recognition model and the target small particle mineral recognition model. The sample image set includes each sample mineral image and the calibrated mineral type corresponding to each sample mineral image. The sample mineral image is an image collected or an image containing minerals in an existing database. The calibrated mineral type is the result of calibrating the types of each mineral in the sample mineral image. The initial large particle mineral recognition model and the initial small particle mineral recognition model are the models at the start of training and during the training process. The sample first type is various mineral types of large particle minerals in the sample mineral image output by the model. The sample detection frame is the detection frame corresponding to each identified large particle mineral in the sample mineral image. The mineral edge frame is the edge frame corresponding to the large particle mineral in each sample detection frame in the sample mineral image, which is more fitting to the edge of the large particle mineral compared to the sample detection frame. The sample process image is the image obtained by masking the large particle minerals in the sample mineral image. The sample second type is various mineral types of small particle minerals in the sample mineral image output by the model. The sample mineral type is the set of mineral types of various sized minerals in the sample mineral image identified by the model. The model metric parameter is a parameter for measuring whether the model passes, such as average precision, recall rate, mean average precision, etc. The model performance requirement is the data threshold for judging whether each model metric parameter is trained successfully.
[0079] Specifically, for each sample mineral image in the sample image set, the sample mineral image is input into the initial large particle mineral recognition model. The initial large particle mineral recognition model analyzes and identifies the mineral type of the large particle minerals in the sample mineral image to obtain the sample first type, and the sample detection frame corresponding to each sample first type can be obtained. Edge detection is performed on the part of the sample mineral image corresponding to each sample detection frame to obtain the edge of the large particle mineral, forming a mineral edge frame. Furthermore, the part of each mineral edge frame corresponding to the sample mineral image is subjected to region masking processing, and the sample process image can be obtained after the masking processing. The sample process image is input into the initial small particle mineral recognition model. The initial small particle mineral recognition model analyzes and identifies the mineral type of the small particle minerals in the sample process image to obtain the sample second type. The sample first type and the sample second type are combined and duplicates are removed, that is, only one of the same mineral types in the sample first type and the sample second type is retained. Accordingly, the sample mineral type corresponding to the sample mineral image can be obtained. Based on the first sample type, the second sample type, the sample mineral type, and the calibrated mineral type corresponding to each sample mineral image in the sample image set, each model metric parameter is calculated according to different calculation methods of the model metric parameters. If all the model metric parameters meet the model performance requirements, it means that each trained model can already meet the usage requirements. Therefore, the initial large particle mineral recognition model is used as the pre-trained target large particle mineral recognition model, and the initial small particle mineral recognition model is used as the pre-trained target small particle mineral recognition model. If at least one of the model metric parameters does not meet the model performance requirements, it means that at least one model is not yet ready for use. Therefore, training adjustment needs to be continued. Therefore, based on each model metric parameter, the model parameters of the initial large particle mineral recognition model and / or the initial small particle mineral recognition model are adjusted, and training is performed again, that is, return to execute the step of inputting the sample mineral image into the initial large particle mineral recognition model to obtain the sample first type and the sample detection frame corresponding to each sample first type until all the model metric parameters meet the model performance requirements, and the initial large particle mineral recognition model can be used as the pre-trained target large particle mineral recognition model, and the initial small particle mineral recognition model is used as the pre-trained target small particle mineral recognition model.
[0080] Exemplarily, an image set of samples is input into the model for training. Each sample mineral image is input into the model to obtain a feature map, and the feature map is input into the detection decoupling branch, which is divided into a classification branch and a regression branch. The regression branch is used to perform regression prediction on the sample detection boxes, and the classification branch is used to classify the mineral categories to obtain the probability of the types of minerals in the sample detection boxes. Then, non-maximum suppression is used to sort the classification scores of the sample detection boxes. The loss value between the sample mineral types output by the model and the sample calibration types is calculated according to the loss function, and then backpropagation is performed through the optimization algorithm to reduce the loss value until the number of iterations is completed, at which point the model training ends.
[0081] The model metric parameters include the following parameters:
[0082]
[0083]
[0084]
[0085]
[0086] Among them, is the precision rate, is the recall rate, is the average precision, is the mean average precision. Correctly classifying a positive example as a positive example is denoted as TP (True Positive), misclassifying a positive example as a negative example is denoted as FN (False Negative), correctly classifying a negative example as a negative example is denoted as TN (True Negative), and misclassifying a negative example as a positive example is denoted as FP (False Positive). is the Precision value corresponding to the R value in the precision-recall curve, and n is the number of mineral categories.
[0087] Based on the above example, the model metric parameters include the first metric parameters corresponding to the initial large particle mineral recognition model and the second metric parameters corresponding to the initial small particle mineral recognition model. The model parameters can be adjusted based on each model metric parameter for the initial large particle mineral recognition model and / or the initial small particle mineral recognition model in the following manner:
[0088] In response to at least one of each first metric parameter not meeting the model performance requirements, the model parameters of the initial large particle mineral recognition model are adjusted based on each first metric parameter, and the initial large particle mineral recognition model is updated;
[0089] In response to at least one of the second metric parameters not meeting the model performance requirements, the initial small particle mineral recognition model is adjusted based on the second metric parameters, and the initial small particle mineral recognition model is updated.
[0090] Among them, the first metric parameter is a partial parameter in the model metric parameters corresponding to the initial large particle mineral recognition model, and the second metric parameter is a partial parameter in the model metric parameters corresponding to the initial small particle mineral recognition model. It can be understood that in addition to the first metric parameter and the second metric parameter, the model metric parameters can also include a third metric parameter for measuring the overall recognition effect, which will not be elaborated here.
[0091] Specifically, if at least one of the first metric parameters does not meet the model performance requirements, it indicates that the initial large particle mineral recognition model cannot be used in practice. Therefore, the initial large particle mineral recognition model needs to be adjusted based on the first metric parameters, and the initial large particle mineral recognition model is updated. If at least one of the second metric parameters does not meet the model performance requirements, it indicates that the initial small particle mineral recognition model cannot be used in practice. Therefore, the initial small particle mineral recognition model needs to be adjusted based on the second metric parameters, and the initial small particle mineral recognition model is updated.
[0092] It can be understood that if there are parameters in both the first metric parameters and the second metric parameters that do not meet the model performance requirements, or at least one of the third metric parameters does not meet the model performance requirements, it indicates that both the initial large particle mineral recognition model and the initial small particle mineral recognition model need to be adjusted, that is, the initial large particle mineral recognition model and the initial small particle mineral recognition model are adjusted based on the model metric parameters, and the initial large particle mineral recognition model and the initial small particle mineral recognition model are updated.
[0093] Based on the above example, the sample image set can be constructed in the following way. Specifically, it can be:
[0094] Obtain each initial mineral image and determine the corresponding initial mineral type of each initial mineral image;
[0095] According to the preset combination quantity, determine the mineral images to be combined from each initial mineral image, perform preset operations on each mineral image to be combined, and combine the mineral images to be combined after the preset operations to obtain a sample mineral image, and determine the calibrated mineral type corresponding to the sample mineral image according to the initial mineral type corresponding to each mineral image to be combined.
[0096] Determine the number of sample mineral images. In response to the number of samples being less than the preset number of images, return to perform the operation of determining the mineral images to be combined from each initial mineral image according to the preset combination number until the number of samples is greater than or equal to the preset number of images;
[0097] Construct a sample image set based on each sample mineral image and the calibrated mineral types corresponding to each sample mineral image.
[0098] Among them, the initial mineral images can be photographed or publicly available mineral images. The initial mineral types are the various mineral types calibrated for the initial mineral images. The preset combination number is the number preset for combining into one sample mineral image, for example, it can be 2 - 4 images. The mineral images to be combined are selected from the initial mineral images according to the preset combination number and are used for image combination. The preset operations can include operations such as random cropping, flipping, and shrinking. The number of samples is the number of sample mineral images. The preset number of images is the minimum number of sample mineral images in the sample image set.
[0099] Specifically, obtain each initial mineral image and determine the corresponding initial mineral types of each initial mineral image. Select the mineral images to be combined with the preset combination number from each initial mineral image, and perform preset operations on each mineral image to be combined to adjust the mineral images to be combined and improve the generalization degree. Furthermore, combine the mineral images to be combined after the preset operations, that is, merge them into one image to obtain the sample mineral image. Integrate the initial mineral types corresponding to the mineral images to be combined that form the sample mineral image to obtain the calibrated mineral type corresponding to the sample mineral image. Count the number of samples of the sample mineral image. If the number of samples is less than the preset number of images, it means that the number of samples is insufficient and additional sample mineral images need to be supplemented. Therefore, return to perform the operation of determining the mineral images to be combined from each initial mineral image according to the preset combination number until the number of samples is greater than or equal to the preset number of images. Then, construct a sample image set based on each sample mineral image and the calibrated mineral types corresponding to each sample mineral image.
[0100] Exemplarily, an initial mineral image is obtained, and the obtaining method is a publicly available dataset or by capturing and collecting mineral images. The initial mineral image is labeled with the initial mineral type using an image annotation tool. After the labeling is completed, an Extensible Markup Language (XML) file is generated, and the corresponding xml file is converted into a txt (text document) file required by the YOLOv8 network. All the initial mineral images are randomly grouped into groups of two to four (preset combination quantity), and random cropping, flipping, and shrinking operations (preset operations) are performed on the mineral images to be combined in each group. The mineral images to be combined after the operation transformation in each group are merged into one image to obtain a sample mineral image, and a sample image set is constructed.
[0101] S140. Determine the target mineral type corresponding to the image to be recognized according to the first mineral type and the second mineral type.
[0102] Among them, the target mineral type is a set of the mineral types of minerals of various sizes in the image to be recognized.
[0103] Specifically, the first mineral type and the second mineral type are combined, and duplicate items are removed, that is, only one of the same mineral types in the first mineral type and the second mineral type is retained. Accordingly, the target mineral type corresponding to the image to be recognized can be obtained.
[0104] The present invention has the following technical effects: By inputting the image to be recognized into a pre-trained target large particle mineral recognition model, the first mineral type and the corresponding target detection frames of each first mineral type are obtained. Through the target large particle mineral recognition model, that is, the YOLOv8 model introducing the global attention mechanism, large particle mineral recognition is performed. Furthermore, edge detection is performed on the image to be recognized according to each target detection frame to obtain each edge detection frame, and region masking processing is performed on the image to be recognized according to each edge detection frame to obtain an intermediate image to mask the large particle minerals and prevent interference with the recognition of small particle minerals. Further, the intermediate image is input into a pre-trained target small particle mineral recognition model to obtain the second mineral type. Through the target small particle mineral recognition model, that is, the YOLOv8 model introducing the multi-scale attention mechanism and the small target detection layer, small particle mineral recognition is performed. Finally, according to the first mineral type and the second mineral type, the target mineral type corresponding to the image to be recognized is determined, achieving the effect of effectively recognizing the mineral types of minerals of different sizes and improving the recognition accuracy.
[0105] Embodiment 2
[0106] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7As shown, the electronic device 200 includes one or more processors 201 and a memory 202.
[0107] The processor 201 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 200 to perform desired functions.
[0108] The memory 202 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 201 can run the program instructions to implement the mineral identification method of any embodiment of the present invention described above and / or other desired functions. Various contents such as initial external parameters, thresholds, etc. can also be stored in the computer-readable storage media.
[0109] In one example, the electronic device 200 can further include: an input device 203 and an output device 204, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). The input device 203 can include, for example, a keyboard, a mouse, etc. The output device 204 can output various information to the outside, including warning prompt information, braking force, etc. The output device 204 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0110] Of course, for simplicity, Figure 7 only some of the components related to the present invention in the electronic device 200 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 200 can further include any other appropriate components.
[0111] In addition to the above methods and devices, an embodiment of the present invention can also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps of the mineral identification method provided by any embodiment of the present invention.
[0112] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present invention. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0113] In addition, an embodiment of the present invention may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps of the mineral identification method provided by any embodiment of the present invention.
[0114] The computer-readable storage medium may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0115] It should be noted that the terms used in the present invention are only for describing specific embodiments and do not limit the scope of the present application. As shown in the specification of the present invention, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. The term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element.
[0116] It should also be noted that the orientation or positional relationship indicated by terms such as "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. Unless otherwise clearly specified and defined, terms such as "installed", "connected", "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A mineral identification method, characterized in that: include: Inputting the image to be identified into a pre-trained target large-grain mineral identification model to obtain the first mineral type and the target detection frame corresponding to each first mineral type; Performing edge detection on the image to be identified according to each target detection frame to obtain each edge detection frame, and performing region masking processing on the image to be identified according to each edge detection frame to obtain an intermediate image; Inputting the intermediate image into a pre-trained target small-particle mineral recognition model to obtain a second mineral type; Determining a target mineral type corresponding to the image to be identified according to the first mineral type and the second mineral type; The target large-grain mineral recognition model is a YOLOv8 model that introduces a global attention mechanism, and the target small-grain mineral recognition model is a YOLOv8 model that introduces a multi-scale attention mechanism and a small target detection layer; The performing a regional masking process on the image to be identified according to each edge detection frame to obtain an intermediate image comprises: Deleting the area image corresponding to each edge detection frame in the image to be identified to obtain an intermediate image; or, The area image corresponding to each edge detection frame in the image to be identified is filled with a preset color to obtain an intermediate image.
2. The method according to claim 1, characterized in that The target large-particle mineral recognition model and the target small-particle mineral recognition model are trained based on the following method: For each sample mineral image in the sample image set, the sample first type, the sample second type and the sample mineral type are obtained through an initial large-grain mineral recognition model and an initial small-grain mineral recognition model; Determine each model metric parameter according to the first sample type, the second sample type, the sample mineral type and the calibration mineral type corresponding to each sample mineral image in the sample image set; In response to each model metric parameter satisfying the model performance requirement, the initial large-grain mineral recognition model is used as a pre-trained target large-grain mineral recognition model, and the initial small-grain mineral recognition model is used as a pre-trained target small-grain mineral recognition model; In response to at least one of the model metric parameters not meeting the model performance requirements, the initial large-grain mineral identification model and / or the initial small-grain mineral identification model are adjusted based on the model metric parameters, and the step of obtaining the first sample type, the second sample type, and the sample mineral type for each sample mineral image in the sample image set through the initial large-grain mineral identification model and the initial small-grain mineral identification model is returned to be executed until all model metric parameters meet the model performance requirements.
3. The method according to claim 2, characterized in that The model metric parameters include a first metric parameter corresponding to the initial large-grain mineral identification model and a second metric parameter corresponding to the initial small-grain mineral identification model; The adjusting of model parameters of the initial large-grain mineral identification model and / or the initial small-grain mineral identification model based on each model metric parameter includes: In response to at least one of the first metric parameters not meeting the model performance requirement, adjusting the model parameters of the initial large-grained mineral identification model based on the first metric parameters to update the initial large-grained mineral identification model; In response to at least one of the second metric parameters not meeting the model performance requirement, the initial small-grained mineral identification model is adjusted based on the second metric parameters to update the initial small-grained mineral identification model.
4. The method according to claim 2, characterized in that: Also includes: Acquire each initial mineral image, and determine the initial mineral type corresponding to each initial mineral image; According to a preset number of combinations, a mineral image to be combined is determined from each initial mineral image, a preset operation is performed on each mineral image to be combined, and each mineral image to be combined after the preset operation is performed is combined to obtain a sample mineral image, and according to the initial mineral type corresponding to each mineral image to be combined, a calibration mineral type corresponding to the sample mineral image is determined; Determine the sample quantity of the sample mineral images, and in response to the sample quantity being less than the preset image quantity, return to perform the operation of determining the mineral images to be combined from the initial mineral images according to the preset combination quantity, until the sample quantity is greater than or equal to the preset image quantity; A sample image set is constructed according to each sample mineral image and the calibrated mineral type corresponding to each sample mineral image.
5. The method according to claim 1, characterized in that The target large-grain mineral recognition model includes: introducing a global attention layer after the basic fast spatial pyramid pooling layer in the basic backbone network of the YOLOv8 model, and the global attention layer replaces the basic fast spatial pyramid pooling layer and is connected to the corresponding basic sampling layer and basic connection layer.
6. The method according to claim 1, characterized in that The target small-particle mineral recognition model includes: introducing a multi-scale attention mechanism layer after the basic fast spatial pyramid pooling layer in the basic backbone network of the YOLOv8 model, and the multi-scale attention mechanism layer replaces the basic fast spatial pyramid pooling layer and is connected to the corresponding basic sampling layer and basic connection layer; introducing a small target detection layer in the basic neck network and the basic head network of the YOLOv8 model.
7. The method according to claim 1, characterized in that The performing edge detection on the image to be identified according to each target detection frame to obtain each edge detection frame includes: For each target detection frame, Gaussian filtering is performed on the detection image corresponding to the target detection frame in the image to be identified to obtain a filtered image; A strong edge and a weak edge are determined according to a preset operator and the filtered image, and an edge detection frame corresponding to the target detection frame is determined according to the strong edge and the weak edge.
8. An electronic device, characterized in that: The electronic device comprises: Processor and memory; The processor is used to execute the steps of the mineral identification method according to any one of claims 1 to 7 by calling the program or instruction stored in the memory.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program or instruction, which enables a computer to execute the steps of the mineral identification method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Remote sensing image target detection method and system based on improved YOLOV8 algorithm
CN117197669A
Sit-up action quality automatic evaluation method based on posture key points
CN118968629A
Power transmission line unmanned aerial vehicle inspection picture defect identification method and system
CN119131621A