Lithology identification method based on rock slice single polarized light image and orthogonal polarized light image feature fusion
By fusing single-polarized and orthogonal polarized image features in rock sheet image recognition, and using the WTConvNeXt-Inception model and cross-attention mechanism, the problem of insufficient accuracy and information utilization of existing lithologic recognition methods is solved, and efficient and accurate automatic lithologic recognition is achieved.
Patent Information
- Application Number
- CN202510299280.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-01
AI Technical Summary
The existing rock lithologic identification methods are limited by the observer's experience and expertise, making it difficult to achieve high-precision classification, and the single-mode image information of single-polarized or orthogonal polarized is insufficiently utilized.
Using the identification method based on the feature fusion of single polarized images and orthogonal polarized images of rock sheets, features are extracted through the WTConvNeXt-Inception model, and feature fusion is performed using the cross attention mechanism, and finally classified through global pooling and full connection layer.
It improves the accuracy of lithology recognition and the feature recognition capabilities of the model, realizes fast, accurate and fully automatic identification of rock flake images, and reduces labor costs and time costs.
Smart Images

Figure CN120236127A_ABST
Abstract
Description
Technical Field:
[0001] The present invention relates to a method for identifying the lithology of rocks in the field of geological exploration, and specifically to a method for identifying lithology by fusing the characteristics of single-polarized light images and cross-polarized light images of rock thin sections. Background Art:
[0002] The identification of rock lithology is an essential part in the process of oil exploration and development. According to the unique lithological characteristics of rocks such as color, composition, and structure, reservoir rocks can be divided into different units. Based on similar lithologies, reservoir rocks in the same unit have similar geological and diagenetic conditions, which are closely related to porosity and oil content, and are also the premise for the characteristic research of oil reservoirs, the calculation of oil reserves, and geological modeling.
[0003] Thin section identification is one of the important methods for lithology identification. Making rock samples into thin sections and observing and analyzing the mineral composition and content, structure and texture, formation sequence, alteration, and secondary changes of rocks with the aid of a polarized light microscope is a basic research method and a conventional observation means for related research such as rock type identification, rock genesis analysis, geological structure interpretation, and paleo-geological environment inversion, and is a basic skill for geological researchers. However, identifying mineral combinations under the microscope requires a large amount of labor cost and time cost, and its correctness is often affected by the experience and professional limitations of observers.
[0004] With the rapid development of computer technology, artificial intelligence has been more closely integrated with oil geological exploration, and intelligent identification methods based on machine learning and deep learning have begun to be applied to thin section identification. Traditional machine learning methods require manual extraction of shallow features such as color, texture, and contour in images, and then use computer technology for data organization and processing. The performance of the algorithm depends on the accuracy of feature recognition and extraction. The lithology identification method based on deep learning can adaptively form filters sensitive to various key deep features during the learning process, and can better perform feature extraction and dimensionality reduction compared with traditional machine learning identification methods, thereby realizing lithology identification. However, most methods only use single-polarized light or cross-polarized light single-modal images for identification, resulting in insufficient information utilization, and the classification accuracy and generalization ability of the model are limited. The color and morphological features of single-polarized light images are relatively obvious, but they are insufficient in distinguishing fine mineral structures and features, while using only cross-polarized light images alone is easily interfered by similar mineral morphologies and it is difficult to achieve high-precision classification.
[0005] In view of this, it is necessary to develop a method for identifying lithology by fusing the characteristics of single-polarized light images and cross-polarized light images of rock thin sections to solve the above problems. Summary of the Invention:
[0006] The object of the present invention is to provide a lithology identification method based on the feature fusion of single-polarized light images and cross-polarized light images of rock thin sections. This lithology identification method based on the feature fusion of single-polarized light images and cross-polarized light images of rock thin sections is used to solve the problem that the existing rock lithology identification methods are affected by the experience and professionalism of observers, resulting in incorrectness or difficulty in achieving high-precision classification.
[0007] The technical solution adopted by the present invention to solve its technical problems is as follows: This lithology identification method based on the feature fusion of single-polarized light images and cross-polarized light images of rock thin sections includes the following steps;
[0008] S1: Collect single-polarized light images and cross-polarized light images of rock thin sections. After preprocessing, perform lithology annotation on them and establish a rock image data set;
[0009] S2: Construct a lithology identification model, including WTConvNeXt-Inception as the backbone network to extract features, and a cross-attention mechanism to fuse the features of single-polarized light images and cross-polarized light images, and classify the spliced features; specifically:
[0010] S21. Extract the features of single-polarized light images and cross-polarized light images of rock thin sections;
[0011] S22. Use a cross-attention mechanism to fuse the feature information of the extracted single-polarized and cross-polarized rock thin section images;
[0012] S23. After performing global pooling on the fused and spliced features, use a fully connected layer operation to obtain the rock thin section identification and classification results;
[0013] S3: Input the rock image to be identified into the trained lithology identification model and output the lithology identification result.
[0014] In the above solution, S21 is specifically as follows:
[0015] Input the single-polarized light image and cross-polarized light image of the rock thin section into the shared WTConvNeXt-Inception model;
[0016] The WTConvNeXt-Inception model replaces the DWConv in the ConvNeXt convolution block with WTConv and decomposes it into four parallel branches along the channel dimension, namely small 3×3 convolution kernels, 1×11, 11×1 rectangular convolution kernels, and an identity mapping, including an initial convolution layer and four convolution block layers.
[0017] In the above solution, S22 is specifically as follows:
[0018] The extracted single-polarized light image and orthogonal-polarized light image feature channel cross-attention module perform global average pooling and max pooling operations on each channel, and input them into a two-layer fully connected network with shared parameters to generate two channel attention weight vectors. Multiply the original feature map with the two channel attention weight vectors channel by channel to obtain a weighted feature map;
[0019] Input the weighted feature map into the spatial fusion attention module. Perform average pooling and max pooling respectively in the channel dimension. Concatenate the channel average pooling map and the channel max pooling map in the channel dimension to obtain a feature map with two channels. Convolve the concatenated feature map with a 3×3 convolution kernel to obtain a spatial attention weight. Multiply the original feature map element by element with the spatial attention weight after Softmax to obtain a new weighted feature map, and then add the new feature map to the original feature map to obtain the final feature map.
[0020] The steps of preprocessing the single-polarized light image and orthogonal-polarized light image of the rock thin section in the above scheme S1:
[0021] Perform data elimination and data augmentation processing on the single-polarized light image and orthogonal-polarized light image data of the rock thin section to improve the clarity and accuracy of the image, and obtain standardized rock thin section image data. Use the standardized rock thin section image data to establish a rock thin section image dataset;
[0022] The process of data elimination for the rock thin section image is to remove damaged rock images and blurred images; the rock thin section images include single-polarized light images and orthogonal-polarized light images of dolomite, limestone, sandstone, conglomerate and shale;
[0023] The data augmentation process is to randomly extract rock thin section images for random flipping, mirror inversion, random translation, chromaticity change, brightness change and contrast processing.
[0024] Beneficial effects:
[0025] 1. In the feature extraction part of the present invention, wavelet convolution is adopted, which can significantly increase the receptive field, realize multi-frequency response to rock thin section images, enhance the feature recognition ability of the model, and introduce the Inception structure to decompose large kernel convolution into multiple small kernel convolutions, improving the calculation efficiency.
[0026] 2. In the feature fusion part of the present invention, a cross-attention fusion module is introduced, which can effectively utilize the features of single-polarized light images and orthogonal-polarized light images, realize information complementarity, and improve the recognition accuracy.
[0027] 3. The optimized automatic lithology recognition model of rock thin section images in the present invention performs automatic lithology recognition on the rock thin section dataset, can rapidly improve the accuracy rate during short-term training, accelerates the convergence process of the model, shortens the training cycle, avoids the model falling into local optimality, effectively improves the accuracy rate of lithology recognition, has high computing power, can be deployed to mobile terminals and PC terminals, and constructs a rock thin section recognition system to achieve rapid recognition.
[0028] 4. Aiming at the problem of low accuracy rate of lithology recognition, the present invention proposes a lightweight model by fusing the features of single-polarized light images and cross-polarized light images, realizes rapid, accurate and fully automatic recognition of rock thin section images, and effectively improves the accuracy rate of lithology recognition. BRIEF DESCRIPTION OF THE DRAWINGS:
[0029] Figure 1 is the flow chart of the present invention;
[0030] Figure 2 is the network model framework diagram of the lithology recognition method based on the feature fusion of single-polarized light images and cross-polarized light images of rock thin sections in the present invention;
[0031] Figure 3 is the structure diagram of the feature extraction network WTConvNeXt-Inception in the present invention;
[0032] Figure 4 is the structure diagram of the cross-attention mechanism feature fusion in the present invention. DETAILED DESCRIPTION OF THE INVENTION:
[0033] The following further describes the present invention in conjunction with the attached drawings:
[0034] As Figure 1 shown, this lithology recognition method based on the feature fusion of single-polarized light images and cross-polarized light images of rock thin sections includes the following contents:
[0035] S1: Collect single-polarized light images and cross-polarized light images of rock thin sections, perform lithology annotation on them, and establish a rock image dataset;
[0036] S1 also includes the step of preprocessing the collected single-polarized light images and cross-polarized light images of rock thin sections, including: performing data elimination and data enhancement processing on the data of the single-polarized light images and cross-polarized light images of rock thin sections to improve the clarity and accuracy of the images, obtaining standardized rock thin section image data, and establishing a rock thin section image dataset using the standardized rock thin section image data.
[0037] In S1, the process of data elimination for the thin-section rock images is to remove damaged rock images and blurred images; the data augmentation process is to randomly select thin-section rock images and perform random flipping, mirror inversion, random translation, chromaticity change, brightness change, and contrast change, etc.
[0038] The thin-section rock images include single-polarized light images and cross-polarized light images of dolomite, limestone, sandstone, conglomerate, and shale.
[0039] S2: Construct a lithology recognition model, including WTConvNeXt-Inception as the backbone network to extract features, and a cross-attention mechanism to fuse the features of single-polarized light images and cross-polarized light images, and classify the spliced features.
[0040] S3: Input the rock image to be recognized into the trained lithology recognition model, and output the lithology recognition result.
[0041] In the second implementation mode of this application:
[0042] The said S2 includes the following steps, as Figure 2 shown:
[0043] S21. Extract the features of the single-polarized light image and cross-polarized light image of the thin-section rock;
[0044] S22. Use the cross-attention mechanism to fuse the feature information of the extracted single-polarized light and cross-polarized light thin-section rock images;
[0045] S23. After performing global pooling on the fused and spliced features, use operations such as a fully connected layer to obtain the recognition and classification result of the thin-section rock.
[0046] In the said S21, when taking the features of the single-polarized light image and cross-polarized light image of the thin-section rock, the extraction steps are as Figure 3 shown,
[0047] Input the single-polarized light image and cross-polarized light image of the thin-section rock into the shared WTConvNeXt-Inception model;
[0048] Convert the sizes of the single-polarized light image and cross-polarized light image into images of 224×224×3;
[0049] The image sequentially passes through the Conv2D and LayerNorm layers and enters the Stage1 layer;
[0050] Repeat the operation of the block, dim = 96 layers for 3 times. The block first decomposes it into four parallel branches along the channel dimension, namely small 3×3 WT convolutional kernels, 1×11 and 11×1 rectangular WT convolutional kernels, and an identity mapping. Then, it passes through the Concat, LayerNorm, and MLP layers in sequence. Finally, the result is concatenated with the original input block image for output;
[0051] Input the result after repeating the operation of the block, dim = 96 layers for 3 times, into the Downsample layer. At this time, the size of the feature map is 56×56×96;
[0052] The result enters the Stage2 layer, and the block with dim = 192 layers is repeated for 3 times. The block first decomposes it into four parallel branches along the channel dimension, namely small 3×3 WT convolutional kernels, 1×11 and 11×1 rectangular WT convolutional kernels, and an identity mapping. Then, it passes through the Concat, LayerNorm, and MLP layers in sequence. Finally, the result is concatenated with the original input block image for output;
[0053] Input the result after repeating the operation of the block, dim = 192 layers for 3 times, into the Downsample layer. At this time, the size of the feature map is 28×28×192;
[0054] The result enters the Stage3 layer, and the block with dim = 384 layers is repeated for 9 times. The block first decomposes it into four parallel branches along the channel dimension, namely small 3×3 WT convolutional kernels, 1×11 and 11×1 rectangular WT convolutional kernels, and an identity mapping. Then, it passes through the Concat, LayerNorm, and MLP layers in sequence. Finally, the result is concatenated with the original input block image for output;
[0055] Input the result after repeating the operation of the block, dim = 384 layers for 3 times, into the Downsample layer. At this time, the size of the feature map is 14×14×384;
[0056] The result enters the Stage4 layer, and the block with dim = 768 layers is repeated for 3 times. The block first decomposes it into four parallel branches along the channel dimension, namely small 3×3 WT convolutional kernels, 1×11 and 11×1 rectangular WT convolutional kernels, and an identity mapping. Then, it passes through the Concat, LayerNorm, and MLP layers in sequence. Finally, the result is concatenated with the original input block image for output;
[0057] Input the result after repeating the operation of the block, dim = 768 layers for 3 times, into the Downsample layer. At this time, the size of the feature map is 7×7×768.
[0058] In the third embodiment of the present application:
[0059] The S22 cross-attention mechanism feature fusion is as follows Figure 4 :
[0060] Put the two features F1 and F2 of size c×h×w of the single-polarized image and the orthogonal-polarized image extracted by the backbone network into the channel cross-attention module;
[0061] Perform global average pooling and max pooling operations on each channel of the input feature map to obtain two global feature descriptions F avg and F max ,
[0062]
[0063] Subsequently, the feature vectors after global max pooling and average pooling are input into a two-layer fully connected network with shared parameters to learn the attention weights of each channel, generating two channel attention weight vectors, MLP(F) = W1·ReLU(W0·F);
[0064] After adding the two channel attention vectors, in order to ensure that the attention weights are between 0 and 1, apply the Sigmoid activation function to generate the channel attention weight M c , M c = σ(MLP(F avg ) + MLP(F max ));
[0065] Obtain the channel weight values M1 c and M2 c of the input feature maps F1 and F2, multiply them to obtain the cross-weight matrix M Cross = M1 c (M2 c ), T Multiply the original feature maps F1 and F2 with the cross-weight matrix M Cross after Softmax channel by channel to obtain the weighted feature maps F'1 and F'2,
[0066] F1' = Softmax(M Cross )·F1, F2' = Softmax(M Cross );
[0067] Input the obtained F'1 and F'2 into the spatial fusion attention module;
[0068] Perform average pooling and max pooling on the input feature map in the channel dimension respectively to obtain two spatial feature maps of size 1×H×W and
[0069]
[0070] The channel average pooling map and the channel maximum pooling map are concatenated in the channel dimension to obtain a feature map with two channels The concatenated feature map is convolved through a 3×3 convolution kernel to obtain a spatial attention weight M s ,
[0071] The original feature maps F1 and F2 and the spatial attention weight M after Softmax s are multiplied element by element to obtain the weighted feature maps F”1 and F”2, F”1 = Softmax(M s )·F'1, F”2 = Softmax(M s )·F'2;
[0072] The idea of the residual network is introduced into the fusion model to obtain f1 and f2, f1 = F”1 + F1, f2 = F”2 + F2.
[0073] The above has introduced in detail the method for lithology identification based on the feature fusion of single-polarized light images and cross-polarized light images of rock thin sections. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A rock property identification method based on the fusion of single polarization image and orthogonal polarization image features of rock thin sections, characterized in that The steps include: S1: Collect single polarization images and orthogonal polarization images of rock slices, annotate them with lithology after preprocessing, and establish rock image dataset; S2: Construct a lithology recognition model, including WTConvNeXt-Inception as the backbone network to extract features, use the cross-attention mechanism to fuse the features of single polarization images and orthogonal polarization images, and classify the spliced features; specifically: S21. Extract features of single polarization image and orthogonal polarization image of rock thin section; S22. Use the cross attention mechanism to fuse the feature information of the extracted single polarization and orthogonal polarization rock thin section images; S23. After globally pooling the fused and spliced features, a fully connected layer operation is used to obtain the rock slice recognition and classification results; S3: Input the rock image to be identified into the trained lithology identification model and output the lithology identification result.
2. The lithology identification method based on the fusion of single polarization image and orthogonal polarization image features of rock thin sections according to claim 1 is characterized by: The S21 is specifically as follows: Input the single polarization image and orthogonal polarization image of the rock slice into the shared WTConvNeXt-Inception model; The WTConvNeXt-Inception model replaces the DWConv in the ConvNeXt convolution block with WTConv and decomposes it into four parallel branches along the channel dimension, namely a small 3×3 convolution kernel, 1×11, 11×1 rectangular convolution kernels and an identity mapping, including an initial convolution layer and four convolution block layers.
3. The lithology identification method based on the fusion of single polarization image and orthogonal polarization image features of rock thin sections according to claim 2 is characterized by: The S22 is specifically as follows: The extracted single polarization image and orthogonal polarization image feature channel cross attention module performs global average pooling and maximum pooling operations on each channel, and inputs them into a two-layer fully connected network with shared parameters to generate two channel attention weight vectors. The original feature map and the two channel attention weight vectors are multiplied channel by channel to obtain the weighted feature map. The weighted feature map is input into the spatial fusion attention module, and average pooling and maximum pooling are performed on the channel dimension respectively. The channel average pooling map and the channel maximum pooling map are spliced on the channel dimension to obtain a feature map with two channels. The spliced feature map is convolved with a 3×3 convolution kernel to obtain a spatial attention weight. The original feature map is multiplied element by element by the spatial attention weight after Softmax to obtain a weighted new feature map, and then the new feature map is added to the original feature map to obtain the final feature map.
4. The lithology identification method based on the fusion of single polarization image and orthogonal polarization image features of rock thin sections according to claim 3 is characterized by: The step of preprocessing the single polarization image and the orthogonal polarization image of the rock slice in S1 is as follows: Performing data elimination and data enhancement processing on the rock thin section single polarization image and orthogonal polarization image data to improve the clarity and accuracy of the image, obtaining standardized rock thin section image data, and establishing a rock thin section image data set using the standardized rock thin section image data; The data removal process of rock thin section images is to remove damaged rock images and blurred images; rock thin section images include: single polarization images and orthogonal polarization images of dolomite, limestone, sandstone, conglomerate and shale; The data enhancement process is to randomly extract rock thin section images and perform random flipping, mirror inversion, random translation, color change, brightness change and contrast processing.
Citation Information
Cited By
Remote sensing image cloud type inversion model and system based on symmetric cross attention
CN120562477A
Multi-modal rock labeling method and device based on heterogeneous adaptive network
CN121170802A
A multi-modal rock labeling method and device based on a heterogeneous adaptive network
CN121170802B