Deep learning-based glow-flash mineral identification method and system
Through deep learning-based methods, rock images are collected and fused, combined with the modification of YOLO-V8 network structure, the MF-YOLO-V8 model is proposed, which solves the problem of time-consuming and prone to deviation in traditional rock classification methods, and realizes efficient identification and naming of gabbro-diorite.
Patent Information
- Application Number
- CN202510027465.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional rock classification methods rely on manual recognition, which is time-consuming and labor-intensive and prone to artificial deviations. Most mineral recognition methods based on artificial intelligence stay on the images of small-field mineral particles, lacking support for large-field whole-field image recognition under polarized microscopes.
Using a deep learning-based method, the image fusion is performed by collecting orthogonal and single polarizer images, generating fusion images and labeling data sets, and modifying them based on the YOLO-V8 network structure, the MF-YOLO-V8 model is proposed for mineral prediction and segmentation of gabbro-diorite.
It realizes efficient identification and naming of gabbro-diorite, improves the accuracy and efficiency of mineral identification, reduces artificial deviations, and supports large-scale vision whole map recognition.
Smart Images

Figure CN119942539A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mineral category identification, and in particular to a method and system for identifying gabbro-diorite minerals based on deep learning. Background Art
[0002] Gabbro is a basic plutonic intrusive rock, mainly composed of basically equal content of monoclinic pyroxene (diopside, isopyrite, common pyroxene, etc.) and basic plagioclase (labradorite and pyroxene), and minor minerals include amphibole, olivine, biotite, orthopyroxene, and a small amount of quartz and alkaline feldspar. Gabbro is gray-black, with a medium-grained to coarse-grained structure, and associated minerals include iron, titanium, copper, nickel, phosphorus, etc. Diorite is a representative rock of fully crystalline neutral plutonic rocks and one of the main rock types in granite. It is mainly composed of plagioclase (medium-longer plagioclase) and one or several dark minerals, of which the total amount of dark minerals is generally around 20-35%. It contains no or only a small amount of potassium feldspar, generally not more than 10% of the total amount of feldspar. It contains no or very little quartz, and its amount does not exceed 5% of the total amount of light-colored minerals. The dark minerals are mainly amphibole, and sometimes pyroxene and biotite.
[0003] In current scientific research and industrial applications, deep learning technology has become one of the key technologies in the field of image classification and recognition. Especially in geological research, identifying rock types by observing rock slices under a microscope is of great significance for understanding the composition and evolutionary history of the earth. However, traditional rock classification methods usually rely on manual identification, especially the quantitative statistics of minerals, which is not only time-consuming and labor-intensive, but also may be biased due to human factors. At present, most of the artificial intelligence-based mineral recognition stays on simple small-field mineral particle images, lacking the recognition of the entire image with a large field of view under a polarizing microscope. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for identifying gabbro-diorite minerals based on deep learning.
[0005] To achieve the purpose of the present invention, the technical solution provided by the present invention is as follows:
[0006] First aspect
[0007] The present invention provides a method for identifying gabbro-diorite minerals based on deep learning, comprising the following steps:
[0008] Step 1: Collect image data and generate data sets;
[0009] Step 2: Train and optimize the recognition model;
[0010] Step 3: Use the recognition model to identify and name gabbro-diorite.
[0011] Furthermore, the step 1 specifically includes the following:
[0012] Step 1.1: Collect images of gabbro-diorite under orthogonal and single polarizing microscopes;
[0013] Step 1.2: Annotate the images under orthogonal and single polarizing microscopes;
[0014] Step 1.3: fuse the images under the same field of view and the same angle under the orthogonal single polarizer at a transparency of 50%;
[0015] Step 1.4: Generate fused image annotation information based on the annotation information under the orthogonal and single polarizers, and generate a mask to obtain a fused image and annotation data set.
[0016] Furthermore, in step 1.2, artificial intelligence and manual work are used to annotate the images under the orthogonal and single polarizing microscopes.
[0017] Furthermore, the step 2 specifically includes the following:
[0018] Step 2.1: Modify the YOLO-V8 network structure according to the mean field theory. In the backbone network structure, use the mean field module to replace the original BN module, add a layer of connection to the feature fusion part, and add a pixel prediction branch to the classification prediction part to obtain the MF-YOLO-V8 network model.
[0019] Step 2.2: Perform network training and perform accuracy evaluation;
[0020] Step 2.3: Compare the accuracy with the YOLO-V8 network structure before modification.
[0021] Furthermore, the step 3 specifically includes the following:
[0022] Step 3.1: Use the MF-YOLO-V8 network model to predict and segment minerals on a gabbro and a diorite thin section;
[0023] Step 3.2: Calculate the average percentage of each mineral pixel level based on the prediction results;
[0024] Step 3.3: Project the calculation results onto the QAPF diagram to complete the naming of the rock thin sections.
[0025] Second aspect
[0026] The present invention provides a gabbro-diorite mineral identification system based on deep learning, comprising a data acquisition unit, a model building unit and an identification unit;
[0027] The data acquisition unit is used to acquire image data and generate a data set;
[0028] The model building unit is used to train and optimize the recognition model;
[0029] The identification unit is used to identify and name gabbro-diorite using the identification model.
[0030] Furthermore, the data acquisition unit is specifically used to perform the following:
[0031] Step 1.1: Collect images of gabbro-diorite under orthogonal and single polarizing microscopes;
[0032] Step 1.2: Annotate the images under orthogonal and single polarizing microscopes;
[0033] Step 1.3: fuse the images under the same field of view and the same angle under the orthogonal single polarizer at a transparency of 50%;
[0034] Step 1.4: Generate fused image annotation information based on the annotation information under the orthogonal and single polarizers, and generate a mask to obtain a fused image and annotation data set.
[0035] Furthermore, in step 1.2, artificial intelligence and manual work are used to annotate the images under the orthogonal and single polarizing microscopes.
[0036] Furthermore, the model building unit is specifically used to perform the following:
[0037] Step 2.1: Modify the YOLO-V8 network structure according to the mean field theory. In the backbone network structure, use the mean field module to replace the original BN module, add a layer of connection to the feature fusion part, and add a pixel prediction branch to the classification prediction part to obtain the MF-YOLO-V8 network model.
[0038] Step 2.2: Perform network training and perform accuracy evaluation;
[0039] Step 2.3: Compare the accuracy with the YOLO-V8 network structure before modification.
[0040] Furthermore, the identification unit is specifically configured to perform the following:
[0041] Step 3.1: Use the MF-YOLO-V8 network model to predict and segment minerals on a gabbro and a diorite thin section;
[0042] Step 3.2: Calculate the average percentage of each mineral pixel level based on the prediction results;
[0043] Step 3.3: Project the calculation results onto the QAPF diagram to complete the naming of the rock thin sections.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) The data acquisition scheme was improved. In order to meet the needs of whole-image recognition of gabbro-diorite images under a 2x microscope, two images under a crossed polarizer and a single polarizer were collected to enrich the collected data information. Through image fusion, the mineral features under the crossed polarizer and the single polarizer were fused to form new features of the mineral image, making some minerals that were originally easy to confuse easier to be recognized by the model. In order to reduce the impact of extinction on mineral recognition, images at different angles in the same field of view were collected. To a certain extent, the difficulty of whole-image recognition was reduced.
[0046] (2) In the data labeling process, in order to make the labeling more accurate, a combination of manual labeling and artificial intelligence labeling was adopted to make the labeling information more accurate and improve efficiency.
[0047] (3) According to the idea of mean field theory, based on the YOLO-v8 network structure, the backbone network structure (Backbone), feature fusion structure (Neck) and classification prediction structure (Head) were modified, and the MF-YOLO-v8 model was proposed. In the backbone network structure (Backbone), the ConvModule module was modified, the idea of mean field theory was added, and the original BN part was replaced with MFConv. According to the characteristics of gabbro-diorite images, the feature fusion structure (Neck) was optimized, and a connection layer was added to make the model more adaptable to images with complex information. The classification prediction structure (Head) divides the original detection head into three parts, so that the classification and segmentation tasks can be performed simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 A schematic diagram of the method steps provided by an embodiment of the present invention;
[0049] Figure 2 Schematic diagram of the fusion of a single polarization image and an orthogonal polarization image in an embodiment of the present invention;
[0050] Figure 3 This is a structural diagram of the MF-Yolo-v8 network model in an embodiment of the present invention;
[0051] Figure 4 This is a structural diagram of the MFConv module in an embodiment of the present invention;
[0052] Figure 5 This is a structural diagram of a C2F module in an embodiment of the present invention;
[0053] Figure 6It is a structural diagram of the SPPF module in an embodiment of the present invention;
[0054] Figure 7 Schematic diagram of the Neck structure of MF-YOLO-v8 in an embodiment of the present invention;
[0055] Figure 8 Schematic diagram of the decoupling head structure of MF-YOLO-v8 in an embodiment of the present invention;
[0056] Fig. 9 It is a confusion matrix of various mineral pixel recognition results of MF-YOLO-v8 training in an embodiment of the present invention;
[0057] Fig.10 A schematic diagram of the average accuracy of the regression boxes of various minerals in the confusion matrix of the pixel recognition of various minerals in the MF-YOLO-v8 training results in an embodiment of the present invention;
[0058] Fig.11 It is the confusion matrix of the identification of various mineral pixels in the MF-YOLO-v8 training results in the embodiment of the present invention, and the average accuracy of various mineral pixels. DETAILED DESCRIPTION
[0059] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] like Figure 1 As shown, the present invention provides a method for identifying gabbro-diorite minerals based on deep learning, comprising the following steps:
[0061] Step 1: Collect image data and generate data sets;
[0062] The step 1 specifically includes the following:
[0063] Step 1.1: Collect images of gabbro-diorite under orthogonal and single polarizing microscopes;
[0064] Step 1.2: Use artificial intelligence and manual work to annotate the images under orthogonal and single polarizing microscopes;
[0065] Step 1.3: fuse the images under the same field of view and the same angle under the orthogonal single polarizer at a transparency of 50%;
[0066] Step 1.4: Generate fused image annotation information based on the annotation information under the orthogonal and single polarizers, and generate a mask to obtain a fused image and annotation data set.
[0067] It should be noted that the present invention collects 32 rock slices including gabbro, diorite, monzogranite, and peridotite. For each slice, 1 to 9 fields with high mineral content are randomly selected for shooting. 462 images are collected under cross polarizers, and 462 images are collected under single polarization, for a total of 924 RGB images. In order to retain more mineral particle information, the acquisition data selects the objective lens magnification of 2x, the camera acquisition mode is pixel offset, and the image resolution is 5760*3600.
[0068] Taking the gabbro image as an example, according to the feature that the image overlaps every 90° under the orthogonal polarizer, the upper and lower polarizers are rotated synchronously from 0° during image acquisition, and rotated clockwise to 90°. The rock slices are photographed every 15°, and two images under the orthogonal polarizer and the single polarizer are collected as a group. After that, the overlapping images taken at 90° are deleted, and the images collected by the polarizer at 0°, 15°, 30°, 45°, 60°, and 75° are retained. These 12 images are renamed according to the slice number, field of view number, shooting order, and picture type. These 12 images can be called an image sequence under the same field of view. The present invention collects a total of 77 groups of rock image sequences.
[0069] Use labelme to annotate the mineral information in the captured image. Load the SAM model in labelme, use the SAM model to annotate the dark minerals with clear boundaries in the image, and outline the contours of the dark minerals. When annotating light-colored minerals, use a curve connected by points to outline their contours. After circling the outline of the mineral, use masks of different colors to cover it according to the type of mineral annotated. This completes the annotation of the category, outline and location information of individual mineral particles. This application divides the objects of mineral annotation into 10 categories, namely monoclinic pyroxene (Cpx), orthopyroxene (Opx), basic plagioclase (Pl), andesine (Pl1), olivine (Ol), hornblende (Hbl), biotite (Bi), quartz (Qtz), opaque minerals (Np) and background (background).
[0070] It should be noted that the present invention adopts a new idea, which is to fuse the images collected under the orthogonal polarizer and the single polarizer at the same angle according to the transparency of each image of 50% to generate a new image. Figure 2 shown.
[0071] The generated new image combines all the features of the orthogonal polarizer and single polarizer images, generating new mineral image features. Many minerals that are not easy to identify become obvious after image fusion. Some features have changed, and the mineral particles cannot be judged according to the original experience of mineral identification. However, since the position information of the mineral particles in the fused image has not changed, the previous annotation information is also applicable to the new image after fusion, and the annotation information of the single polarizer image and the orthogonal polarizer image are applicable to the new image. The annotation information of the original single polarizer image at the same angle in the same field of view is superimposed with the annotation information of the orthogonal polarizer image to generate a new set of annotation information. This set of annotation information is applied to the fused image, and all annotations of the fused image are completed.
[0072] Step 2: Train and optimize the recognition model;
[0073] The step 2 specifically includes the following:
[0074] Step 2.1: Modify the YOLO-V8 network structure according to the mean field theory. In the backbone network structure, use the mean field module to replace the original BN module, add a layer of connection to the feature fusion part, and add a pixel prediction branch to the classification prediction part to obtain the MF-YOLO-V8 network model.
[0075] Step 2.2: Perform network training and perform accuracy evaluation;
[0076] Step 2.3: Compare the accuracy with the YOLO-V8 network structure before modification.
[0077] It should be noted that applying mean field theory to deep learning can help us understand and predict the behavior of the network from a macro perspective, such as how the initialization of parameters affects the training effect of the network, or how the depth and width of the network determine the learning ability and generalization ability. In addition, mean field theory can also be used to study the dynamic properties of the network optimization process, such as the gradient disappearance or explosion problem, and how to optimize these properties by adjusting the network structure or learning rate.
[0078] Mean field theory is a method of collectively processing the effects of the environment on individuals, replacing the sum of individual effects with the average effect. This method can simplify the study of complex problems and transform a high-order, multi-dimensional, difficult-to-solve problem into a low-dimensional problem, which is equivalent to integrating the effects of the environment on the research object and then interacting with the research object. It is observed that the effects between neurons in the same layer are symmetrical. Once this symmetry is achieved through proper initialization, it is expected to remain unchanged at all subsequent times. This especially shows the following two characteristics:
[0079] (1) Marginal uniformity: If the control of neuron i in layer l depends on the other neurons in this layer only through the global statistics of layer l, then this law applies to all neurons in layer l.
[0080] (2) Self-averaging: The sum of a sufficient number of terms can be replaced by an appropriate integral that reflects symmetry in their roles, with each term corresponding to a neuron in the same layer. More specifically, if we compare neuron i among n neurons in the same layer with g(x i ,A i ) association, where A i is a measure μ independent of neuron i i The random amount sampled, then as n approaches infinity, for an appropriate probability measure ρ, it acts as a replacement for the set of neurons in that layer.
[0081]
[0082] Where n: number of neurons; i: the i-th neuron; g(x i ,A i ): the output of the i-th convolution layer; g(x i ,a i ): the output of one of the neurons in the i-th layer; μ i : is a function of the convolution kernel w; ρ; probability measure.
[0083] According to the mean field theory, when n→∞, a mean field multi-layer fully connected neural network can be constructed.
[0084] The forward propagation is defined as follows:
[0085]
[0086] The inductive definition is as follows:
[0087] H 1 (w; x) =<w,x>
[0088] H 2 (f; x, ρ 1 )=∫f(θ,w)σ 1 (θ,H 1 (w; x))ρ 1 (dθ,dw)
[0089]
[0090] in, Specifically, ρ is the probability of occurrence in the system, x∈R d is the input item, is the output item.
[0091] The mean field multilayer neural network also calculates the gradient of each variable in the back propagation based on the gradient descent algorithm, and calculates the loss function based on the learning rounds. Determine the learning rate α, α>0.
[0092] The difference between the mean field multilayer fully connected neural network and the ordinary multilayer fully connected neural network is that when the number of neurons in a single layer n→∞, the mean field multilayer fully connected neural network uses the average value of all weights of this layer instead of the value of its weight for calculation.
[0093] YOLO-v8 is a real-time target detection system. As an efficient deep learning model, YOLOv8 has achieved significant performance improvements and technological innovations on the basis of inheriting the advantages of its predecessors. YOLOv8 adopts a deeper neural network architecture, combined with an attention mechanism and an improved residual module, which greatly enhances the model's detection accuracy and adaptability to complex environments. In addition, by introducing new optimization techniques and algorithms, such as automatic label smoothing and recursive neural network training, YOLOv8 not only improves training efficiency, but also optimizes running performance on different hardware platforms, including GPUs, TPUs, and mobile devices. YOLO-v8 can be mainly divided into three parts: backbone network structure (Backbone), feature fusion structure (Neck), and classification prediction structure (Head). This application will use the mean field theory to modify the backbone network structure (Backbone), such as Figure 3 As shown in the figure, the feature fusion structure (Neck) and the classification prediction structure (Head) are optimized according to the particularity of the image, thereby improving the learning efficiency of the model.
[0094] (1) Backbone network structure
[0095] Combining the idea of mean field theory, this application modifies the backbone network structure of YOLO-v8. The modified MF-YOLO-v8 backbone network structure mainly includes three basic components: MFConvModule module (such as Figure 4 As shown), C2F module (as Figure 5 As shown) and SPPF module (as Figure 6 shown).
[0096] (1.1) MFConvModule
[0097] The main function of the MFConvModule module is to extract feature information from the image. The MFConvModule module consists of three parts: convolution layer, scaling layer, and SILU activation function layer. Combining the idea of the mean field theory, compared with the ConvModule module of the original YOLO-v8 backbone network structure (Backbone), an MFConv layer is added after the convolution layer as a scaling layer, and the BN (BatchNormalization) layer is removed. According to the mean field formula, starting from the second layer of convolution, the output result of the previous layer is scaled. Through this change, it is hoped that the accuracy of feature extraction will be enhanced, without losing model efficiency, and as many image features as possible will be retained. In addition, the MFConvModule module retains the SILU function in the ConvModule module as the activation layer. Its expression is:
[0098] y=x*sigmoid(x)
[0099] Activation functions play a very important role in neural networks. Here are some of the main functions of activation functions:
[0100] Introducing nonlinearity: The main goal of the activation function is to introduce nonlinearity in the model. This is because, without an activation function, no matter how many layers the neural network has, it can only represent linear functions. By introducing nonlinearity, we can make the neural network better adapt to complex data and simulate more complex functions.
[0101] Determine whether a neuron should be activated: The activation function defines the form of the neuron's output given an input (including bias). In other words, the activation function determines whether a neuron should be activated. This is determined based on whether the input information is important and needs to be further propagated.
[0102] Helps with optimization: Activation functions and their derivatives (gradients) play a key role in the back-propagation process. During the back-propagation process, the gradients are used to update the weights and biases of the network. Choosing the right activation function can help the network converge faster and reduce problems that occur during training, such as gradient disappearance or explosion.
[0103] SiLU is an improved version of Sigmoid and ReLU. SiLU has the characteristics of no upper bound, lower bound, smoothness, and non-monotonicity. SiLU works better than ReLU on deep models. It can be regarded as a smoothed ReLU activation function. It has some advantages of the ReLU activation function: it can alleviate the gradient vanishing problem; it can also solve some shortcomings of the ReLU function: the ReLU function is not centered on zero, and the gradient in the negative part is zero. In addition, the SiLU function is a smooth function, which means that it has derivatives in the entire domain, which is conducive to optimization.
[0104] (1.2) C2F module
[0105] Compared with previous generations of YOLO, especially YOLO-v5, the C2F module in the YOLO-v8 backbone network is a relatively big change. The C2F module has fewer parameters and better feature extraction capabilities. In the C2F module, it can be seen that the output of the previous round is first processed by a convolution module, then a split process, and after being processed by n DarknetBottleneck modules, the results of the residual module and the backbone module are concat spliced, and then processed by a convolution module for output.
[0106] The residual connection mentioned in it is a technique for building deep networks. Its core idea is to pass residuals or errors by skipping layer connections. In traditional neural networks, information flows through layers of the network, and each layer is transformed and features are extracted through nonlinear activation functions. However, as the neural network deepens, the problem of gradient disappearance or "gradient explosion" may occur, resulting in difficulty in network convergence or performance degradation.
[0107] Residual connections solve the gradient vanishing and exploding problems by introducing cross-layer connections to pass the original input information directly to subsequent layers. Specifically, it adds the input of the network to the output of the intermediate layer to form a "shortcut or jump connection", allowing the gradient to propagate more easily. Mathematically, suppose we have an input x, which is processed through multiple network layers to obtain a predicted value H(x). Then the expression of the residual connection is:
[0108] F(x)=H(x)+x
[0109] Among them, F(x) is the output of the residual block, H(x) is the feature representation obtained after processing through a series of network layers, and x is the jump connection from the input directly connected to the residual block. Through the residual connection, the network can learn the residual or error more easily, so that the deeper features of the network can be more accurately expressed. This is very useful for training deep neural networks and can improve the performance and convergence speed of the network.
[0110] The essence of C2F is to use a 1×1 convolution kernel to reduce the number of input channels to the original The purpose is to reduce calculation and memory, and then use multiple 3×3 convolution kernels to perform convolution operations, extract feature information, add the input directly to the output through residual connection, thus forming a cross-layer connection, and finally use 1×1 convolution kernel to restore the number of channels of the feature map.
[0111] (1.3) SPPF module
[0112] In traditional convolutional neural networks, fully connected layers are usually used to map features extracted by convolution and pooling layers to vectors of fixed length for classification or detection. However, fully connected layers require fixed-size inputs, which makes the network unable to process input images of different sizes. Therefore, in 2014, He Kaiming et al. proposed a layer called Spatial Pyramid Pooling (SPP) to solve the problem of different input image sizes and achieve accurate classification and detection on input images of different sizes. Before the emergence of SPP, region-based deep neural networks (such as R-CNN) usually used a selective search (sliding window) algorithm to generate candidate regions, and then performed separate convolution and pooling operations on each candidate region, and finally input the pooled feature vector into the classifier for classification. The method of using selective search has the following disadvantages:
[0113] Large amount of computation: The selective search algorithm needs to segment and merge each pixel in the image to generate a large number of candidate regions. Convolution and pooling operations on each candidate region require a lot of computing resources, making this method slow.
[0114] There are overlapping areas: Since the selective search algorithm is based on regions, there may be overlapping parts between different regions. This will cause different parts of the same target to be detected repeatedly, increasing the amount of calculation and the false detection rate.
[0115] Poor invariance: The selective search algorithm may generate too many or too few candidate regions, which will result in different numbers of candidate regions for the same object in different images, making the model less invariant to the target.
[0116] To solve these problems, YOLO-v8 introduced the SPPF module. In SPPF, the input image is first extracted through a convolutional neural network, and then the output feature map of the convolutional layer is divided into grids of different scales. For each grid, a pooling operation is performed inside it, and the pooling output results are spliced into a feature vector of fixed length.
[0117] (2) Feature Fusion Structure (Neck)
[0118] The feature fusion structure (Neck) in YOLO-v8 is an important component in the model. It still adopts the idea of PAN-FPN and plays an important role in feature extraction and fusion. PAN-FPN (Path Aggregation Network with Feature Pyramid Network) is a neural network architecture for object detection in computer vision. It combines the Feature Pyramid Network (FPN) with the Path Aggregation Network (PAN) to improve the accuracy and efficiency of target detection. FPN is used to extract features from images of different scales, while PAN is used to aggregate these features across different layers of the network. This allows the network to detect objects of different sizes and resolutions and handle complex scenes with multiple objects.
[0119] The Neck part of YOLO-v8 is responsible for further processing the features from Backbone to improve the accuracy and robustness of target detection. It fuses multi-scale feature maps by introducing different structures and techniques to better capture the information of targets of different scales. The feature pyramid network structure (FPN) is used to process the multi-scale feature maps from Backbone. FPN enables the model to detect targets at different scales by building feature pyramids at different levels. The Neck part also includes feature fusion operations to fuse features from different levels. This feature fusion helps improve the model's detection accuracy of targets, especially for targets of different scales. The Neck part usually uses upsampling and downsampling operations to adjust the scale and resolution of the feature map. The upsampling operation can enlarge the low-resolution feature map to the same size as the high-resolution feature map to retain more detail information. The downsampling operation can reduce the size of the high-resolution feature map to reduce the amount of computation and memory consumption.
[0120] Compared with YOLO-v8, the steps of the Neck part of MF-YOLO-v8 are as follows: the outputs P3, P4, and P5 of the feature extraction network (Backbone) are input into the PAN-FPN network structure to fuse the feature maps of multiple scales; P5 is fused with P4 after upsampling to obtain F1, F1 is fused with P3 after a C2F layer and an upsampling to obtain T1, T1 is fused with F1 after a convolution layer to obtain F2, F2 is fused with a C2F layer to obtain T2, T2 is fused with P5 after a convolution layer to obtain F3, F3 is fused with a C2f layer to obtain T3, and finally T1, T2, and T3 are the products of the entire Neck. The process is as follows Figure 7 shown.
[0121] The modified Neck part effectively extracts and fuses multi-scale features through operations such as feature pyramid network and feature fusion, thereby improving the performance and robustness of object detection. This enables the model to better adapt to objects of different scales and sizes and achieve more accurate detection results in complex scenarios.
[0122] (3) Classification prediction structure (Head)
[0123] The Backbone and Neck parts can be understood as preparations for the Head part. After the previous preparations, the outputs of the Neck part, T1, T2, and T3, represent feature maps of different levels. The Head part is a process of processing these three feature maps to generate the output results of the model. The most core change in the Neck part of MF-YOLO-v8 is reflected in the decoupled head (Decoupled-Head), which decomposes the original detection head into three parts. The three decoupled heads in the Head part correspond to the feature map outputs T1, T2, and T3 of the Neck part respectively. The decoupled head is as follows Figure 8 As shown. The workflow of the decoupling head is: the feature maps T1, T2, and T3 obtained by the network are respectively input into the decoupling head for prediction. The detection head contains 4 3x3 convolutions and 2 1x1 convolutions. At the same time, the CIOU loss function is added to the regression branch of the detection head. The regression head needs to calculate the position offset between the predicted box and the real box, and then send the offset to the regression head for loss calculation, and then output a four-dimensional vector, which represents the upper left corner coordinates x, y and the lower right corner coordinates x, y of the target box. The classification head performs a unified size (Rol Pooling) and convolution operation on each candidate box extracted without an anchor point (AnchorFree) to obtain a classifier output tensor. The value at each position represents the probability that the candidate box belongs to each category. Finally, the final detection result is screened out by maximum value suppression.
[0124] Because the Head part decomposes the original detection head into two parts, one is semantic segmentation and the other is target detection.
[0125] Object detection uses cross-entropy loss as a loss function, which is a commonly used loss function in classification problems. It measures the difference between the probability distribution predicted by the model and the actual label, and is used to measure the accuracy of the model prediction. The expression is:
[0126]
[0127] M represents the number of categories; y ic represents a sign function (0 or 1), if the true category of sample i is equal to c, it takes 1, otherwise it takes 0; pic It represents the predicted probability that the observed sample i belongs to category c.
[0128] The loss functions for semantic segmentation are VFL loss function and CIOU loss function. The VFL loss function is as follows:
[0129]
[0130] q is the IoU (intersection over union) of bbox (prediction box) and gt (real box). IoU is the intersection of the prediction box and the real box divided by the union of the two boxes. Then p is the score, that is, the probability. The improvement of this formula is to change the division by N in the basic formula to multiplication by q. Then the two boxes intersect, that is, q>0, which is a positive sample. If the two boxes do not intersect, let q=0, which is a negative sample.
[0131] The CIOU loss function is as follows:
[0132]
[0133] In this formula, IoU is the intersection over union ratio, b and b gt represents the center point of the two rectangular boxes, ρ represents the Euclidean distance between the two rectangular boxes, c represents the diagonal distance of the closed area of the two rectangular boxes, v is used to measure the consistency of the relative proportions of the two rectangular boxes, and α is the weight coefficient.
[0134] Experimental results and analysis
[0135] (1) Training environment and hyperparameter settings
[0136] The collected data set is divided into training set and test set in a ratio of 9:1.
[0137] The initial learning rate of the optimized MF-Yolo-v8 network training is 0.01, the warmup learning strategy is adopted, the batch size is 16, and the total number of epoch training rounds is 500.
[0138] Training environment: CPU: 16 cores, Xeon(R) Gold 6430, GPU: RTX 4090 / 24GB.
[0139] (2) Evaluation indicators
[0140] This application uses three indicators as the basis for evaluating model training, namely precision P (Precision), recall R (Recall), and category average precision MAP (MeanAverage Precision).
[0141] The samples in the experiment can be divided into four situations according to the true category and the predicted category. TP: the number of positive categories predicted by actual positive categories, FN: the number of negative categories predicted by actual positive categories, FP: the number of positive categories predicted by actual negative categories, TN: the number of negative categories predicted by actual negative categories. The classification result confusion matrix established by sample type can intuitively reflect the training classification of the model.
[0142] (3) Experimental results and analysis
[0143] The various indicators of MF-YOLO-v8 after training are as follows:
[0144] This is the confusion matrix of various mineral pixel recognition results of MF-YOLO-v8 training results, such as Fig. 9 As shown. By normalizing the confusion matrix, the probability of each mineral being correctly identified in the confusion matrix is shown. The horizontal axis of the confusion matrix represents the actual category of the mineral, and the vertical axis represents the predicted category of the mineral. The main diagonal of the confusion matrix represents the situation where the mineral category is accurately identified, and the numbers in the matrix represent the probability of this situation. The remaining unmarked situations are situations where the probability is less than 1%.
[0145] After 500 generations of training, the average precision (B-AP) of various mineral regression boxes on the validation set is as follows: Fig.10 As shown, MAP = 91.9% (Box); clinopyroxene (Cpx): 0.953, orthopyroxene (Opx): 0.970, basic plagioclase (Pl): 0.855, andesine (Pl1): 0.840, biotite (Bi): 0.902, amphibole (Hbl): 0.977, olivine (Ol): 0.922, quartz (Qtz): 0.891, opaque mineral (Np): 0.960. Average precision mean MAP: 0.919.
[0146] After 500 generations of training, the pixel average precision (M-AP) of various minerals on the validation set is as follows: Fig.11 As shown, MAP = 91.4% (Mask). Clinopyroxene (Cpx): 0.951, orthopyroxene (Opx): 0.969, basic plagioclase (Pl): 0.843, andesine (Pl1): 0.818, biotite (Bi): 0.893, amphibole (Hbl): 0.973, olivine (Ol): 0.922, quartz (Qtz): 0.895, opaque mineral (Np): 0.961. Average intersection ratio MIou: 91.9%.
[0147] Step 3: Use the recognition model to identify and name gabbro-diorite.
[0148] Furthermore, the step 3 specifically includes the following:
[0149] Step 3.1: Use the MF-YOLO-V8 network model to predict and segment minerals on a gabbro and a diorite thin section;
[0150] Step 3.2: Calculate the average percentage of each mineral pixel level based on the prediction results;
[0151] Step 3.3: Project the calculation results onto the QAPF diagram to complete the naming of the rock thin sections.
[0152] It should be noted that the QAPF diagram is QAPF [quartz (Q), alkali feldspar (A), plagioclase (P), para-feldspar (F)] diagram. When dark minerals are <90%, the QAPF diagram recommended by the International Union of Geological Sciences (IUGS) is used as the classification standard for intrusive rocks, while intrusive rocks with dark minerals >90% should use the Ol-Opx-Cpx or Ol-Px-Hbl triangle diagram. Gabbro determined by QAPF should be further subdivided using Pl-Px-Ol or Pl-Px-Hbl.
[0153] The present invention provides a gabbro-diorite mineral identification system based on deep learning, comprising a data acquisition unit, a model building unit and an identification unit;
[0154] The data acquisition unit is used to acquire image data and generate a data set;
[0155] The model building unit is used to train and optimize the recognition model;
[0156] The identification unit is used to identify and name gabbro-diorite using the identification model.
[0157] Furthermore, the data acquisition unit is specifically used to perform the following:
[0158] Step 1.1: Collect images of gabbro-diorite under orthogonal and single polarizing microscopes;
[0159] Step 1.2: Annotate the images under orthogonal and single polarizing microscopes;
[0160] Step 1.3: fuse the images under the same field of view and the same angle under the orthogonal single polarizer at a transparency of 50%;
[0161] Step 1.4: Generate fused image annotation information based on the annotation information under the orthogonal and single polarizers, and generate a mask to obtain a fused image and annotation data set.
[0162] Furthermore, in step 1.2, artificial intelligence and manual work are used to annotate the images under the orthogonal and single polarizing microscopes.
[0163] Furthermore, the model building unit is specifically used to perform the following:
[0164] Step 2.1: Modify the YOLO-V8 network structure according to the mean field theory. In the backbone network structure, use the mean field module to replace the original BN module, add a layer of connection to the feature fusion part, and add a pixel prediction branch to the classification prediction part to obtain the MF-YOLO-V8 network model.
[0165] Step 2.2: Perform network training and perform accuracy evaluation;
[0166] Step 2.3: Compare the accuracy with the YOLO-V8 network structure before modification.
[0167] Furthermore, the identification unit is specifically configured to perform the following:
[0168] Step 3.1: Use the MF-YOLO-V8 network model to predict and segment minerals on a gabbro and a diorite thin section;
[0169] Step 3.2: Calculate the average percentage of each mineral pixel level based on the prediction results;
[0170] Step 3.3: Project the calculation results onto the QAPF diagram to complete the naming of the rock thin sections.
[0171] Finally, it should be noted that the above embodiments are only used to illustrate and explain the present invention, and are not intended to limit the present invention to the scope of the described embodiments. In addition, those skilled in the art can understand that the present invention is not limited to the above embodiments, and more variations and modifications can be made according to the teachings of the present invention, and these variations and modifications all fall within the scope of the protection claimed by the present invention.
Claims
1. A method for identifying gabbro-diorite minerals based on deep learning, characterized in that: The steps include: Step 1: Collect image data and generate data sets; Step 2: Train and optimize the recognition model; Step 3: Use the recognition model to identify and name gabbro-diorite.
2. The method for identifying gabbro-diorite minerals based on deep learning according to claim 1, characterized in that: The step 1 specifically includes the following: Step 1.1: Collect images of gabbro-diorite under orthogonal and single polarizing microscopes; Step 1.2: Annotate the images under orthogonal and single polarizing microscopes; Step 1.3: fuse the images under the same field of view and the same angle under the orthogonal single polarizer at a transparency of 50%; Step 1.4: Generate fused image annotation information based on the annotation information under the orthogonal and single polarizers, and generate a mask to obtain a fused image and annotation data set.
3. The method for identifying gabbro-diorite minerals based on deep learning according to claim 2, characterized in that: In step 1.2, artificial intelligence and manual work are used to annotate the images under the orthogonal and single polarizing microscopes.
4. The method for identifying gabbro-diorite minerals based on deep learning according to claim 3, characterized in that: The step 2 specifically includes the following: Step 2.1: Modify the YOLO-V8 network structure according to the mean field theory. In the backbone network structure, use the mean field module to replace the original BN module, add a layer of connection to the feature fusion part, and add a pixel prediction branch to the classification prediction part to obtain the MF-YOLO-V8 network model. Step 2.2: Perform network training and perform accuracy evaluation; Step 2.3: Compare the accuracy with the YOLO-V8 network structure before modification.
5. The method for identifying gabbro-diorite minerals based on deep learning according to claim 4, characterized in that: The step 3 specifically includes the following: Step 3.1: Use the MF-YOLO-V8 network model to predict and segment minerals on a gabbro and a diorite thin section; Step 3.2: Calculate the average percentage of each mineral pixel level based on the prediction results; Step 3.3: Project the calculation results onto the QAPF diagram to complete the naming of the rock thin sections.
6. A gabbro-diorite mineral identification system based on deep learning, characterized in that: It includes a data collection unit, a model building unit and an identification unit; The data acquisition unit is used to acquire image data and generate a data set; The model building unit is used to train and optimize the recognition model; The identification unit is used to identify and name gabbro-diorite using the identification model.
7. The gabbro-diorite mineral identification system based on deep learning according to claim 6, characterized in that: The data acquisition unit is specifically used to perform the following: Step 1.1: Collect images of gabbro-diorite under orthogonal and single polarizing microscopes; Step 1.2: Annotate the images under orthogonal and single polarizing microscopes; Step 1.3: fuse the images under the same field of view and the same angle under the orthogonal single polarizer at a transparency of 50%; Step 1.4: Generate fused image annotation information based on the annotation information under the orthogonal and single polarizers, and generate a mask to obtain a fused image and annotation data set.
8. The deep learning-based gabbro-diorite mineral identification system according to claim 7, characterized in that: In step 1.2, artificial intelligence and manual work are used to annotate the images under the orthogonal and single polarizing microscopes.
9. The deep learning-based gabbro-diorite mineral identification system according to claim 8, characterized in that: The model building unit is specifically used to perform the following: Step 2.1: Modify the YOLO-V8 network structure according to the mean field theory. In the backbone network structure, use the mean field module to replace the original BN module, add a layer of connection to the feature fusion part, and add a pixel prediction branch to the classification prediction part to obtain the MF-YOLO-V8 network model. Step 2.2: Perform network training and perform accuracy evaluation; Step 2.3: Compare the accuracy with the YOLO-V8 network structure before modification.
10. The deep learning-based gabbro-diorite mineral identification system according to claim 9, characterized in that: The identification unit is specifically used to perform the following: Step 3.1: Use the MF-YOLO-V8 network model to predict and segment minerals on a gabbro and a diorite thin section; Step 3.2: Calculate the average percentage of each mineral pixel level based on the prediction results; Step 3.3: Project the calculation results onto the QAPF diagram to complete the naming of the rock thin sections.