A rock thin section mineral identification method, medium and device
By using the U-Net architecture in rock sheet mineral recognition combined with Swin-Transformer and convolutional attention mechanism module, the problem of manual identification in traditional methods is solved, and efficient and accurate mineral recognition effect is achieved.
Patent Information
- Application Number
- CN202510221839.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Traditional rock sheet mineral recognition methods rely on manual identification, making it difficult to quickly and accurately identify the coexistence of multiple minerals in complex rock samples, and are limited by the recognition ability of the human eye and human error.
A rock sheet mineral recognition method is proposed. Using the U-Net architecture combined with Swin-Transformer and convolutional attention mechanism module, the channel and spatial attention weighting are obtained by acquiring image features at multiple polarization angles, fusing features of different polarization angles, acquiring global information of the image, and generating fine segmentation results through multi-layer decoder.
This method can effectively identify mineral images under polarization at different angles, improve the efficiency and accuracy of mineral recognition, especially in complex mineral areas, showing higher accuracy and robustness.
Smart Images

Figure CN119723582B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of mineral identification, and in particular to a rock slice mineral identification method, medium and equipment. Background Art
[0002] Shale, sandstone and carbonate rock are important types of sedimentary rocks, and their mineral composition is of great significance for geological research and resource development. Traditional mineral identification methods rely on manual observation and chemical analysis, such as rock hand specimen identification, rock thin section identification, X-ray fluorescence analysis technology (XRF), X-ray diffraction technology (XRD), etc. Among them, rock hand specimen identification is the simplest and most direct method of lithology identification, but it is impossible to identify the mineral composition inside the rock. Rock thin section identification is a more traditional lithology identification technology. By using a polarizing microscope to observe the mineral crystal characteristics and optical properties in the rock thin section, the type of rock and its genetic characteristics can be analyzed, thereby realizing the identification of the lithology. Thin section identification is an important technical means in geological prospecting. It can show the microstructure and characteristics of the rock in detail, and has advantages in lithology identification that cannot be replaced by other methods.
[0003] With the rapid development of artificial intelligence technology, especially breakthroughs in image processing and deep learning, automation and intelligent technology have gradually replaced traditional manual identification. However, despite the significant progress of artificial intelligence technology, mineral identification in rock thin sections still faces many challenges, especially in the analysis of complex rock samples such as clastic rocks and carbonate rocks. Traditional rock thin section analysis relies on manual identification under a microscope. Although this method has high accuracy, it relies on the experience and skills of professionals and is limited by the recognition ability of the human eye and human errors. Especially in the case of coexistence of multiple minerals in the sample and complex crystal morphology, manual identification is often difficult to complete quickly and accurately. Therefore, how to solve this problem and improve the efficiency and accuracy of mineral identification has become a key technical problem in mineralogy, geology and related fields.
[0004] Although deep learning and image segmentation technologies have been successful in many fields, the automated application of rock thin section mineral identification still faces limitations. In particular, due to the complexity of minerals in rock thin sections and their changes under different lighting conditions, traditional convolutional neural network (CNN) models are difficult to accurately identify all mineral species.
[0005] Patent application CN118840749A used Swin-Transformer and graph convolution to segment pathological WSI images and achieved good results. Lin et al. used a new deep medical image segmentation framework of dual Swin Transformer U-Net (DS-TransUNet) to achieve automatic medical image segmentation. However, these are mainly used in the medical field and do not involve the field of mineral identification.
[0006] The EfficientNetV2 feature extractor is used to more effectively extract mineral edge information at different scales, but the EfficientNetV2 encoding is not as flexible as the Swin-Transformer architecture in large-scale image classification problems, especially tasks that require multi-scale, global context awareness. In addition, the author only used extinction feature images at 7 angles, which has certain limitations.
[0007] Patent application CN118506068A optimizes the MobileViT model. By introducing super-resolution reconstruction, GELU activation function, coordinated attention mechanism, transfer learning and cosine annealing algorithm, it not only improves the recognition accuracy, but also accelerates the training process. Through a series of optimization experiments, it verifies the significant impact of different technologies on recognition accuracy and training time, and provides an efficient and accurate method for automatic lithology recognition of rock thin section images. However, it does not consider the influence of rock extinction, and has weak recognition ability for mineral images under orthogonal polarization at different angles. Summary of the invention
[0008] The purpose of the present invention is to solve the problem of complex mineral identification of mineral images under different polarization angles, and to propose a rock thin section mineral identification method, comprising the following steps:
[0009] S1. Obtain single polarization images and orthogonal polarization images of rock slices at multiple polarization angles, and annotate the images to obtain a set of labeled images;
[0010] S2. A rock thin section mineral recognition model is constructed based on the U-Net architecture. The left part of the rock thin section mineral recognition model of the U-Net architecture is connected to five feature extraction blocks through four downsampling blocks. Each feature extraction block is composed of a convolution and a Swin-Transformer Stage connected together, and the Swin-Transformer Stage outputs features.
[0011] S3. Use the label image set to train and test the rock thin section mineral recognition model, input the rock thin section image to be detected into the trained rock thin section mineral recognition model, and obtain the mineral segmentation result of the rock thin section to be detected.
[0012] Further,
[0013] There is also a convolutional attention mechanism module between the connected convolution and Swin-Transformer Stage in the rock thin section mineral recognition model.
[0014] Furthermore, the decoder part of the rock thin section mineral recognition model based on the U-Net architecture is expressed as follows:
[0015]
[0016] in, represents the decoding result of the l-th layer decoder, is the convolution kernel weight of the l-th layer decoder, is the feature map of the jump connection to the l-th layer decoder, represents the encoding result of the l+1th layer decoder, Represents a deconvolution operation.
[0017] Furthermore, the boundary loss is used to optimize the prediction results of the mineral segmentation map for the contour:
[0018]
[0019] in, represents the boundary loss, represents the region boundary, represents the boundary gradient, It represents the probability mapping of the segmented area predicted by the mineral segmentation map, and T(x) is the label mapping of the true segmented area.
[0020] Furthermore, Dice loss is used to balance the similarity between the mineral segmentation map and the true segmentation image:
[0021]
[0022] in, represents the Dice loss, represents the value of the i-th pixel in the mineral segmentation map, yes The weight of Represents the value of the i-th pixel in the true annotation image.
[0023] Furthermore, the results of the mineral segmentation map are optimized using boundary loss and Dice loss:
[0024]
[0025] Among them, L represents the total loss function, and They represent weight coefficients respectively.
[0026] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned rock thin section mineral identification method is implemented.
[0027] The present invention also proposes an electronic device, comprising a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the above-mentioned rock thin section mineral identification method.
[0028] The beneficial effects brought by the technical solution provided by the present invention are:
[0029] The present invention first extracts the features of rock slice images with different polarizations through feature extraction, and performs channel attention and spatial attention weighting, and fuses the weighted features, fuses the features of images with different polarization angles, and inputs the fused images through a multi-layer Swin-Transformer Stage to obtain the global information of the image, and then restores the image to the original resolution through a multi-layer decoder of U-Net, and uses jump connections to fuse high-level and low-level features to generate a more refined segmentation result. The solution of the present invention has a strong ability to recognize mineral images under polarization of different angles, and can handle complex mineral image segmentation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of a rock thin section mineral identification method according to an embodiment of the present invention;
[0031] Figure 2 It is a schematic diagram of a rock thin section mineral identification model constructed based on the U-Net architecture according to an embodiment of the present invention;
[0032] Figure 3 It is the structural diagram of two consecutive Swin Transformer blocks;
[0033] Figure 4 It is the structure diagram of the convolutional attention mechanism module (CBAM);
[0034] Figure 5 is a schematic diagram of a rock slice image to be identified in an embodiment of the present invention;
[0035] Figure 6 It is a schematic diagram of feature extraction results of a rock thin section mineral identification model constructed based on a U-Net architecture according to an embodiment of the present invention;
[0036] Figure 7is a schematic diagram of the image segmentation results of the rock thin section mineral recognition model constructed based on the U-Net architecture in an embodiment of the present invention; wherein, Figure 7 (a) is a rock slice image after different polarization images are fused. Figure 7 (b) is the quartz mineral edge map extracted in the embodiment;
[0037] Figure 8 It is a block diagram of an electronic device in an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0038] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0039] Example 1: A flow chart of a rock thin section mineral identification method according to an embodiment of the present invention is shown in FIG. Figure 1 , specifically including the following steps:
[0040] S1. Obtain single polarization images and orthogonal polarization images of rock slices at multiple polarization angles, and annotate the images to obtain a set of labeled images.
[0041] Use a polarizing microscope to collect images of rock thin section samples, and collect sample images at different polarization angles (such as 1°, 5°, 15°, etc.) to ensure that the mineral characteristics under different lighting conditions are covered. These images include rock thin section images under single polarization (PPL) and orthogonal polarization (XPL), which represent different optical properties. Manually annotate the minerals in the images, generate label images through annotation tools, and each pixel is assigned a corresponding label according to the type of mineral annotated (such as different pixel values for different minerals). Generate the corresponding segmentation mask, and perform quality assessment on the collected images, remove low-quality images, ensure the validity of the data set, and finally review the annotation results to ensure the accuracy and consistency of the annotation.
[0042] S2. A rock thin section mineral recognition model is constructed based on the U-Net architecture. The left part of the rock thin section mineral recognition model with the U-Net architecture connects five feature extraction blocks through four downsampling blocks. Each feature extraction block is composed of a convolution and a Swin-Transformer Stage. The original two consecutive convolutions are replaced by a convolution and a Swin-Transformer Stage, and the Swin-Transformer Stage outputs the features.
[0043] The rock thin section mineral identification model constructed based on the U-Net architecture in the embodiment of the present invention is referenced Figure 2 The rock thin section mineral recognition model consists of an encoder and a decoder connected by skip connections.
[0044] Encoder part: The images of rock slices with different polarizations are sequentially subjected to the first convolution and the first Swin-Transformer Stage of the model encoder to extract features and obtain the first feature map; after the first feature map is subjected to maximum pooling (downsampling), it is sequentially subjected to the second convolution and the second Swin-Transformer Stage of the encoder to extract features and obtain the second feature map; after the second feature map is subjected to maximum pooling (downsampling), it is sequentially subjected to the third convolution and the third Swin-Transformer Stage of the encoder to extract features and obtain the third feature map; after the third feature map is subjected to maximum pooling (downsampling), it is sequentially subjected to the fourth convolution and the fourth Swin-Transformer Stage of the encoder to extract features and obtain the fourth feature map; after the fourth feature map is subjected to maximum pooling (downsampling), it is sequentially subjected to the fifth convolution and the fifth Swin-TransformerStage of the encoder to extract features and obtain the fifth feature map.
[0045] Decoder part: The fifth feature map is obtained by the first upsampling (deconvolution operation) to obtain the sixth feature map, and the fourth feature map is connected to the sixth feature map through a jump connection to obtain the first fused feature map; the first fused feature map is successively passed through the first convolution and the second upsampling (deconvolution operation) of the decoder to obtain the seventh feature map, and the third feature map is connected to the seventh feature map through a jump connection to obtain the second fused feature map; the second fused feature map is successively passed through the second convolution and the third upsampling (deconvolution operation) of the decoder to obtain the eighth feature map, and the second feature map is connected to the eighth feature map through a jump connection to obtain the third fused feature map; the third fused feature map is successively passed through the third convolution and the fourth upsampling (deconvolution operation) of the decoder to obtain the ninth feature map, and the first feature map is connected to the ninth feature map through a jump connection to obtain the fourth fused feature map; the fourth fused feature map passes through the fourth convolution and a 1×1 convolution of the decoder to output the image segmentation result.
[0046] Except for the last convolution of the encoder which uses 1×1 convolution, all other convolutions use 3×3 convolution.
[0047] The first Swin-Transformer Stage includes a linear embedding layer (i.e., fully connected layer) and two consecutive Swin Transformer blocks; the second Swin-Transformer Stage includes a Patch Merging layer and two consecutive Swin Transformer blocks; the third Swin-Transformer Stage includes a Patch Merging layer and two consecutive Swin Transformer blocks; the fourth Swin-Transformer Stage includes a Patch Merging layer and four consecutive Swin Transformer blocks; the fifth Swin-Transformer Stage includes a Patch Merging layer and two consecutive Swin Transformer blocks. For the structural diagram of two consecutive Swin Transformer blocks, please refer to Figure 3 .
[0048] Swin-Transformer provides powerful global feature modeling capabilities, especially through hierarchical nested sliding windows (Window-based Attention), which can capture local and global features of images. However, the feature maps of Swin-Transformers are also downsampled layer by layer, which may lead to the dilution of detail information. Using jump connections to input images, the image generates multiple layers of features through multiple stages of Swin-Transformers. The features of each stage are directly sent to the corresponding stage of the U-Net decoder through jump connections. The decoder uses the high-resolution details provided by the jump connection to fuse with the features obtained by its own upsampling, thus avoiding the loss of details caused by multiple downsampling and ensuring that the mineral boundary and pore information can be fully transmitted. In addition, the jump connection provides an additional gradient path to avoid the gradient vanishing problem, which helps the network to train more stably. Therefore, the jump connection ensures that the detail information of the mineral boundary and pore will not be lost, and improves the accuracy and robustness of the segmentation.
[0049] In another embodiment of the present invention, a convolutional attention mechanism module (CBAM) is further included between the nth (n=1, 2, 3, 4) convolution of the encoder and the nth Swin-Transformer Stage, that is, a CBAM is further included between the convolution of each feature extraction block and the Swin-Transformer Stage. The structural diagram of the convolutional attention mechanism module is shown in Figure 1. Figure 4 .
[0050] The feature map input to the Swin-Transformer Stage is divided into several local windows (window size is w×w). The self-attention mechanism of each window can help the model capture long-distance dependency information. The attention calculation formula within each window is as follows:
[0051]
[0052] Among them, Q, K, V are query vector, key vector and value vector respectively, d is the dimension of the vector, and the softmax operation normalizes the calculated similarity to generate the attention weight of each pixel. For each window, the self-attention mechanism can enhance the global dependency of features. For the entire image, Swin-Transformer will perform this windowed self-attention operation multiple times to obtain the global information of the image.
[0053] Convolutional Attention Mechanism Module of The channel attention mechanism dynamically adjusts the importance of each channel's features. By weighting the feature map of each channel, the model can automatically focus on important channel features and ignore irrelevant channels. The spatial attention mechanism is used to adaptively adjust different areas in the image, especially mineral boundaries and important areas. By analyzing spatial context information, the model can focus on the mineral area and ignore the background or noise.
[0054] The decoder part of U-Net restores the image to its original resolution through deconvolution operations and uses skip connections to fuse high-level and low-level features to generate a more refined segmentation result. Specifically, the expression of the decoder part of U-Net is as follows:
[0055]
[0056] in, represents the decoding result of the l-th layer decoder, is the convolution kernel weight of the l-th layer decoder, is the feature map of the jump connection to the l-th layer decoder, represents the encoding result of the l+1th layer decoder, Represents a deconvolution operation.
[0057] In order to optimize the segmentation effect, the boundary loss is used to optimize the prediction results of the mineral segmentation map for the contour:
[0058]
[0059] in, represents the boundary loss, represents the region boundary, represents the boundary gradient, It represents the probability mapping of the segmented area predicted by the mineral segmentation map, and T(x) is the label mapping of the true segmented area.
[0060] Use Dice loss to balance the similarity between the mineral segmentation map and the true segmentation image:
[0061]
[0062] in, represents the Dice loss, represents the value of the i-th pixel in the mineral segmentation map, yes The weight of Represents the value of the i-th pixel in the true annotation image.
[0063] The results of the mineral segmentation maps are optimized using boundary loss and Dice loss to ensure that the model performs well on different tasks.
[0064]
[0065] Among them, L represents the total loss function, and They represent weight coefficients respectively.
[0066] Through the above steps, combined with Swin-Transformer and U-Net, the model can effectively segment minerals, especially in complex mineral areas, showing higher accuracy and robustness.
[0067] S3. Use the label image set to train and test the rock thin section mineral recognition model, input the rock thin section image to be detected into the trained rock thin section mineral recognition model, and obtain the mineral segmentation result of the rock thin section to be detected.
[0068] refer to Figure 5 , Figure 5 Schematic diagram of a rock slice image to be identified in an embodiment of the present invention, including images at multiple polarization angles. Figure 6 It is a schematic diagram of feature extraction results of a rock thin section mineral identification model constructed based on a U-Net architecture according to an embodiment of the present invention, which extracts a high-dimensional feature map representing the integration of global information and local features. Figure 7 is a schematic diagram of the image segmentation results of the rock thin section mineral recognition model constructed based on the U-Net architecture in an embodiment of the present invention, wherein: Figure 7 (a) is a rock slice image after different polarization images are fused. Figure 7 (b) is the quartz mineral edge map extracted in the embodiment.
[0069] Embodiment 2: In an exemplary embodiment, a computer-readable storage medium is included, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned rock thin section mineral identification method is implemented.
[0070] Example 3: Please refer to Figure 8 In an exemplary embodiment, an electronic device is also included, including at least one processor, at least one memory, and at least one communication bus.
[0071] Wherein, a computer program is stored in the memory, and the computer program includes computer-readable instructions. The processor calls the computer-readable instructions stored in the memory through the communication bus to execute the above-mentioned rock thin section mineral identification method.
[0072] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A rock thin section mineral identification method, characterized in that: The following steps are involved: S1. Obtain single polarization images and orthogonal polarization images of rock slices at multiple polarization angles, and annotate the images to obtain a set of labeled images; S2. A rock thin section mineral recognition model is constructed based on the U-Net architecture. The left part of the rock thin section mineral recognition model of the U-Net architecture is connected to five feature extraction blocks through four downsampling blocks. Each feature extraction block is composed of a convolution and a Swin-Transformer Stage connected together, and the Swin-Transformer Stage outputs features. The first Swin-Transformer Stage includes a linear embedding layer and two consecutive Swin Transformer blocks; the second, third, and fifth Swin-Transformer Stages each include a Patch merging layer and two consecutive Swin Transformer blocks; the fourth Swin-Transformer Stage includes a Patch merging layer and four consecutive Swin Transformer blocks; There is also a convolutional attention mechanism module between the connected convolution and Swin-Transformer Stage in the rock thin section mineral recognition model; S3. Use the label image set to train and test the rock thin section mineral recognition model, input the rock thin section image to be detected into the trained rock thin section mineral recognition model, and obtain the mineral segmentation result of the rock thin section to be detected.
2. A rock thin section mineral identification method according to claim 1, characterized in that: The decoder part of the rock thin section mineral recognition model based on the U-Net architecture is expressed as follows: in, represents the decoding result of the l-th layer decoder, is the convolution kernel weight of the l-th layer decoder, is the feature map of the jump connection to the l-th layer decoder, represents the encoding result of the l+1th layer decoder, Represents a deconvolution operation.
3. A rock thin section mineral identification method according to claim 1, characterized in that: Optimizing the mineral segmentation map using boundary loss to predict the contours: in, represents the boundary loss, represents the region boundary, represents the boundary gradient, It represents the probability mapping of the segmented area predicted by the mineral segmentation map, and T(x) is the label mapping of the true segmented area.
4. A rock thin section mineral identification method according to claim 3, characterized in that: Use Dice loss to balance the similarity between the mineral segmentation map and the true segmentation image: in, represents the Dice loss, represents the value of the i-th pixel in the mineral segmentation map, yes The weight of Represents the value of the i-th pixel in the true annotation image.
5. A rock thin section mineral identification method according to claim 4, characterized in that: Results of optimizing the mineral segmentation map using boundary loss and Dice loss: Among them, L represents the total loss function, and They represent weight coefficients respectively.
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
7. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises computer-readable instructions, and the processor is configured to call the computer-readable instructions to execute the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Deep learning-based rock slice image lithology identification method
CN118506068A
WSI image classification system based on deep learning
CN118840749A
Intelligent evaluation method for weathering degree of tunnel surrounding rock
CN118675031A