Advanced feature fusion-based bimodal rock slice image semantic segmentation method

By adopting the advantageous feature fusion method of the dual-modal channel separation module and the dual-branch spatial attention module in semantic segmentation, the shortcomings of single-modal semantic segmentation in complex scenarios are solved, and more efficient and accurate rock sheet recognition is achieved.

CN120125822APending Publication Date: 2025-06-10CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510255590.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has poor single-modal semantic segmentation effect in complex scenarios, and the semantic segmentation method based on a single type of rock sheet polarized image has problems such as insufficient segmentation accuracy and large differences in results under different optical conditions.

Method used

A dual-modal rock sheet image semantic segmentation method based on the fusion of advantageous features is adopted. The dual-modal channel separation module and the dual-branch spatial attention module extract and fuse advantageous feature information from the channel and spatial dimensions, and a dual-channel feature extraction network is constructed, and it is replaced with the backbone network of the DeepLab V3+ network.

Benefits of technology

It improves the speed and accuracy of rock sheet recognition, effectively solves the correlation and consistency problems between multimodal information, and improves the semantic segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125822A_ABST
    Figure CN120125822A_ABST
Patent Text Reader

Abstract

The invention discloses a bimodal rock slice image semantic segmentation method based on dominant feature fusion, and belongs to the technical field of image processing, and the method comprises the steps: obtaining a data set D; constructing an AFFM module comprising a dual-mode channel separation module and a dual-branch space attention module; constructing a dual-channel feature extraction network based on a MobileNetV3 network and an AFFM module; replacing a backbone network of the DeepLab V3 + network with a dual-channel feature extraction network to obtain an improved semantic segmentation network, and training to obtain a dual-mode rock slice image semantic segmentation model; the method is used for semantic segmentation of to-be-identified rock slices. According to the method, the AFFM module is designed, a new backbone network is designed based on the AFFM module and the MobileNetV3, advantage feature information can be extracted and fused from channel and space dimensions, and the speed and accuracy of rock slice recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a dual-modal rock thin section image semantic segmentation method based on dominant feature fusion. Background Art

[0002] With the development of computer vision and artificial intelligence technologies, deep learning algorithms have been widely used to assist in the identification of rock thin sections. The single-modal deep learning semantic segmentation algorithm based on a single type of polarized light image of rock thin sections is one of the most important research points in the current semantic segmentation field. This type of network usually also has a certain adaptability to datasets with a large span of application scenarios. For example, the Non-Local Network (NLNet) Non-Local significantly increases the accuracy of the semantic segmentation network at that time, but there is also the problem of huge computational complexity caused by learning global context without position dependence, which makes it difficult to be implemented in actual engineering projects. The real-time semantic segmentation model, the Bilateral Segmentation Network (BiSeNet), performs well in general semantic segmentation scenarios, but still has poor segmentation effects in complex scenarios with a large amount of details and occlusions. Thus, semantic segmentation in a single modality often faces the problem of insufficient information provided by the modality in complex scenarios. In addition, the semantic segmentation method based on a single type of polarized light image of rock thin sections also has problems such as insufficient segmentation accuracy and large differences in semantic segmentation results for polarized light images collected under different optical conditions, and still requires a large amount of correction of the segmentation results by the identification workers.

[0003] DeepLab V3+ is a semantic segmentation network belonging to the encoder-decoder architecture. It extracts image features through the encoder and then maps these features back to the original image size through the decoder to achieve pixel-level classification. The encoder includes a backbone network and ASPP. The input passes through the backbone network to obtain two outputs: one is the low-level feature, which is directly provided to the decoder, and the other is the high-level feature, which captures multi-scale context information through ASPP. The outputs of the low-level feature and ASPP are both input into the decoder for processing.

[0004] MobileNetV3 is a lightweight convolutional neural network, which is divided into two versions: MobileNetV3-Small and MobileNetV3-Large. The overall structures of the two are similar, both including convolutional layers, multiple bneck layers, pooling layers, convolutional layers, etc. Among them, the Small version includes 11 bneck layers, and the Large version includes 15 bneck layers. As the core module and the basic module of the MobileNetV3 network, the bneck layer mainly implements depthwise separable convolution + SE channel attention mechanism + residual connection. In the present invention, the multiple bneck layers are divided into 5 stages from beginning to end to output feature maps of 5 different scales.

[0005] ASPP, the full English name is Atrous Spatial Pyramid Pooling, and the Chinese name is Atrous Spatial Pyramid Pooling. It captures object information at different scales by using atrous convolutional layers with different dilation rates to process the input feature map in parallel. Summary of the Invention

[0006] The purpose of the present invention is to provide a dual-modal rock thin section image semantic segmentation method based on the fusion of dominant features, which solves the above problems, extracts and fuses dominant feature information from the channel and spatial dimensions respectively, and improves the speed and accuracy of rock thin section recognition.

[0007] To achieve the above purpose, the technical solution adopted by the present invention is as follows: A dual-modal rock thin section image semantic segmentation method based on the fusion of dominant features, including the following steps; S1, obtain the dataset D, where the dataset D includes multiple sample pairs, and one sample pair includes the plane-polarized light image and the cross-polarized light image of the same rock thin section; S2, construct the AFFM module, including a dual-modal channel separation module, a dual-branch spatial attention module, and a fusion module; The dual-modal channel separation module is used to input the first input feature 、the second input feature , after extracting the complementary mapping relationship between the two from the channel dimension, generate the corresponding first attention vector 、the first channel weight and the first output feature , generate the corresponding second attention vector 、the second channel weight and the second output feature ; The dual-branch spatial attention module is used to separately process 、 Perform spatial attention and output the corresponding spatial attention weights and ; The fusion module is used to generate the output of the AFFM module , , where is element-wise multiplication, is element-wise addition; S3. Construct a dual-channel feature extraction network, including S31 to S32; S31. Obtain two MobileNetV3 networks with the same structure, respectively labeled as and . The MobileNetV3 network includes a first convolutional layer, n bneck layers, and a second convolutional layer. The n + 2 layers are divided into 5 stages from top to bottom, and each stage outputs a feature map. The is used to input a single polarized light map and sequentially output the i-th single polarized light feature map through the i-th stage . The is used to input an orthogonally polarized light map and sequentially output the i-th orthogonally polarized light feature map through the i-th stage , 1 ≤ i ≤ 5; S32. Set the i-th AFFM module between , and use and as the first input feature and the second input feature of respectively. For the first 4 stages, the output of is added to as the input of the (i + 1)-th stage, the output of is added to as the input of the (i + 1)-th stage. For the 5th stage, the output of is used as the output of the dual-channel feature extraction network; S4. Obtain a DeepLab V3+ network, replace its backbone network with the dual-channel feature extraction network, and obtain an improved semantic segmentation network;

[0008] Preferably: The dual-modal channel separation module is used to input the first input feature and the second input feature , spliced into a splicing feature , One path generates a first aggregation feature through global max pooling and MLP , and the other path generates a second aggregation feature through average pooling and MLP , and then , Generate a bimodal attention vector through element-wise addition and the Sigmoid function , and perform channel splitting to obtain a first attention vector , a second attention vector ; , Generate a first channel weight through the Softmax function respectively , a second channel weight , Multiply element-wise with and then with Splice into a second output feature , Multiply element-wise with and then with Splice into a first output feature .

[0009] Preferably: The double-branch spatial attention module includes a first spatial attention unit and a second spatial attention unit with the same structure;

[0010] The first spatial attention unit includes a global max pooling layer, an average pooling layer, a splicing layer, a convolutional layer, and a sigmoid function layer; Perform global max pooling operation on the channel dimension through the global max pooling layer to obtain the output , Perform average pooling operation on the channel dimension through the average pooling layer to obtain the output , and the splicing layer is used to , Splice to obtain a feature , reduce the dimension through the convolutional layer, and then output through the sigmoid function layer ; The second spatial attention unit is used for input , output .

[0011] Preferably: S4 specifically includes steps S41~S43; S41, obtain a DeepLab V3+ network, including a backbone network, ASPP, and a decoder; S42, replace the backbone network with a dual-channel feature extraction network, and The output of ASPP is output ; S43, will and All are sent to the decoder to obtain the semantic segmentation results.

[0012] Preferably, the MobileNetV3 network adopts MobileNetV3-Large, including a first convolutional layer, 15 bneck layers, and a second convolutional layer, a total of 17 layers, which are divided into 5 stages from the 3rd, 5th, 8th, 14th, and 17th layers respectively.

[0013] Preferably, the MobileNetV3 network adopts MobileNetV3-Small, including a first convolutional layer, 11 bneck layers, and a second convolutional layer, a total of 13 layers, which are divided into 5 stages from the 3rd, 5th, 8th, 10th, and 13th layers respectively.

[0014] Preferably, the dataset D is the rock thin section microscopic image dataset of Nanjing University, the rock thin section microscopic image dataset of PaddlePaddle AIStudio Galaxy Community, or the rock thin section microscopic image dataset of the Xu Zhuang Formation of the Middle Cambrian in the Ordos Basin.

[0015] Compared with the prior art, the advantages of the present invention are: a dual-modal deep learning algorithm is designed to improve the accuracy of semantic segmentation and the efficiency of rock thin section identification. (1) An AFFM module was designed. The module includes a dual-modal channel separation module and a dual-branch spatial attention module. The dual-modal channel separation module extracts the complementary mapping relationship between the two modalities from the channel dimension and refines the effective features; the dual-branch spatial attention module extracts the spatial attention weights of the extracted effective features according to the respective advantages of the single polarization and orthogonal polarization images. The final AFFM module can extract the attention vectors (first attention vector, second attention vector) of the two modal inputs from the channel and spatial dimensions respectively, generate the weights of the two modalities (first channel weight, second channel weight), and fuse the advantageous feature information of different modalities according to the weights, which effectively solves the correlation and consistency problems between multimodal information.

[0016] (2) A dual-channel feature extraction network is designed. The network uses two MobileNetV3 networks with the same structure as two branches to extract features from single polarization images and orthogonal polarization images respectively, and uses the AFFM module to fuse the dominant features of the feature maps of the two branches at different stages to finally obtain high-level features. The MobileNetV3 uses a depth-wise separable convolution consisting of two convolution operations, depth-wise convolution and point-wise convolution, connected in series.

[0017] (3) An improved semantic segmentation network with a dual-branch encoder-decoder structure is constructed based on a dual-channel feature extraction network. Among them, the original semantic segmentation network selects the DeepLab V3+ network, which includes an encoder and a decoder. The encoder includes a backbone network and ASPP. The backbone network only uses a deep neural network to extract features from a feature map. In the present invention, the backbone network is replaced with a dual-channel feature extraction network, and the output of the dual-channel feature extraction network is used as the high-level feature, and the output is obtained after passing through ASPP. . The decoder first performs upsampling on and combines the result with for splicing. Finally, the spliced result is further upsampled to restore the image resolution, and the segmentation predictions of different classes are output by channel, and finally the unified semantic segmentation of the two-modal information is realized. This segmentation network can effectively improve the speed and accuracy of rock thin section recognition.

[0018] (4) Based on the above structure, through network structure experiments, it is determined that the dual-channel feature fusion segmentation network has the best performance when using MobileNetV3-Large as the backbone network and the hyperparameter setting of the fusion ratio is 0.2. Its mean intersection over union (MIoU) in the dual-modal rock thin section dataset can reach 73.43%. Two ablation experiments are carried out to verify the rationality, effectiveness and plug-and-play characteristics of the setting of the dominant feature fusion module. The comparison experiment with other networks shows that the dual-channel feature fusion segmentation network achieves the highest MIoU performance index while having fewer parameters. Description of the Drawings

[0019] Figure 1 is the structural diagram of the dual-modal channel separation module; Figure 2 is the flow chart of the dual-modal channel separation module; Figure 3 is the structural diagram of the dual-branch spatial attention module; Figure 4 is the flow chart of the dual-branch spatial attention module; Figure 5 is the structural diagram of the AFFM module; Figure 6 is the schematic diagram of the improved semantic segmentation network structure; Figure 7 is the encoder structure of the dual-channel feature extraction network; Figure 8 is the flow chart of the present invention. Detailed Embodiments

[0020] The present invention will be further described below in conjunction with the embodiments and the drawings.

[0021] Embodiment 1: Refer toFigures 1 to 8 , a dual-modal thin rock slice image semantic segmentation method based on dominant feature fusion, comprising the following steps; S1. Obtain a data set D, where the data set D includes a plurality of sample pairs, and one sample pair includes a single-polarized light map and a cross-polarized light map of the same thin rock slice; S2. Construct an AFFM module, including a dual-modal channel separation module, a dual-branch spatial attention module, and a fusion module; The dual-modal channel separation module is used to input a first input feature and a second input feature . After extracting the complementary mapping relationship between the two from the channel dimension, generate the corresponding first attention vector , the first channel weight and the first output feature . Generate the corresponding second attention vector , the second channel weight and the second output feature ; The dual-branch spatial attention module is used to perform spatial attention on and respectively, and output the corresponding spatial attention weights and ; The fusion module is used to generate the output of the AFFM module , , where is element-wise multiplication, and is element-wise addition; S3. Construct a dual-channel feature extraction network, including S31~S32; S31. Obtain two MobileNetV3 networks with the same structure, which are respectively labeled as and . The MobileNetV3 network includes a first convolutional layer, n bneck layers, and a second convolutional layer. The n+2 layers are divided into 5 stages from top to bottom, and each stage outputs a feature map. The is used to input the single-polarized light map and sequentially output the i-th single-polarized light feature map through the i-th stage. The is used to input the cross-polarized light map and sequentially output the i-th cross-polarized light feature map , 1≤i≤5; S32. Set the i-th AFFM module between , and and and As the first input feature and the second input feature respectively, for the first 4 stages, the output of is added to as the input of the (i + 1)-th stage, the output of is added to as the input of the (i + 1)-th stage. For the 5th stage, the output of is used as the output of the dual-channel feature extraction network; S4. Obtain a DeepLab V3+ network, replace its backbone network with the dual-channel feature extraction network to obtain an improved semantic segmentation network; S5. Train the improved semantic segmentation network with the dataset D until convergence to obtain a dual-modal rock thin section image semantic segmentation model;

[0022] In this embodiment: The dual-modal channel separation module is used to input the first input feature and the second input feature and splice them into a spliced feature . One path passes through global max pooling and MLP to generate a first aggregated feature , and the other path passes through average pooling and MLP to generate a second aggregated feature . Then, and are element-wise added and passed through the Sigmoid function to generate a dual-modal attention vector , and channel splitting is performed to obtain a first attention vector and a second attention vector ; and respectively pass through the Softmax function to generate a first channel weight and a second channel weight . is multiplied element-wise with and then spliced with to form a second output feature . is multiplied element-wise with and then spliced with to form a first output feature .

[0023] The dual-branch spatial attention module includes a first spatial attention unit and a second spatial attention unit with the same structure; The first spatial attention unit includes a global maximum pooling layer, an average pooling layer, a concatenation layer, a convolutional layer, and a sigmoid function layer; The global maximum pooling operation in the channel dimension is performed by the global maximum pooling layer to obtain an output , The average pooling operation in the channel dimension is performed by the average pooling layer to obtain an output , and the concatenation layer is used to , concatenate to obtain a feature , reduce the dimension through the convolutional layer, and then output through the sigmoid function layer ; The second spatial attention unit is used for input , output .

[0024] S4 specifically includes steps S41 to S43; S41, obtain a DeepLab V3+ network, including a backbone network, ASPP, and a decoder; S42, replace the backbone network with a dual-channel feature extraction network, and obtain an output from the output of through ASPP; S43, send both and into the decoder to obtain the semantic segmentation result.

[0025] The MobileNetV3 network uses MobileNetV3-Large, including a first convolutional layer, 15 bneck layers, and a second convolutional layer, with a total of 17 layers, which are divided into 5 stages from the 3rd, 5th, 8th, 14th, and 17th layers respectively.

[0026] The dataset D is the Nanjing University rock teaching thin section microscopic image dataset, the PaddlePaddle AI Studio Star River Community rock thin section microscopic image dataset, or the Ordos Basin Middle Cambrian Xuzhuang Formation rock thin section microscopic image dataset.

[0027] Regarding the Nanjing University Petrology Thin Section Micrograph Dataset, this dataset covers 28 sedimentary rocks, 40 igneous rocks, and 40 metamorphic rocks, covering more than 90% of the common rock types, as well as more than 95% of the common mineral types and rock structures required to be mastered in geology majors. It contains a total of 2,634 polarized micrographs, with 1 transmitted single polarized light photo and 7 - 8 transmitted cross - polarized light photos taken for each rock thin section. The single polarized light photo is the single polarized light image described in the present invention, and the transmitted cross - polarized light photo is the cross - polarized light image described in the present invention.

[0028] Regarding the Rock Thin Section Micrograph Dataset of the Star River Community on PaddlePaddle AI Studio: This dataset is divided into 10 folders according to mineral types, and among them, 7 types of mineral thin sections are photographed under two light source environments of single polarized light and cross - polarized light.

[0029] Regarding the Rock Thin Section Micrograph Dataset of the Middle Cambrian Xuzhuang Formation in the Ordos Basin: The dataset consists of polarized micrographs of 192 rock thin sections, and each rock thin section contains one thin section cross - polarized light micrograph and one single polarized light micrograph with the same field of view.

[0030] Example 2: Refer to Figures 1 to 8 , based on Example 1, a more specific structure of the dual - mode channel separation module is given. The dual - mode channel separation module includes a first splicing layer, a global maximum pooling layer, an average pooling layer, an MLP, a first addition layer, a Sigmoid function layer, a channel splitting layer, a Softmax function layer, a first multiplication layer, a second multiplication layer, a second splicing layer, and a third splicing layer. The first splicing layer is used to splice the first input feature and the second input feature into a spliced feature ; the global maximum pooling layer is used to perform a global maximum pooling operation on in the channel dimension to obtain the corresponding output ; the average pooling layer is used to perform an average pooling operation on in the channel dimension to obtain the corresponding output ; the MLP inputs , , and outputs the corresponding first aggregated feature , the second aggregated feature ; the first addition layer is used to add , element - by - element; the Sigmoid function layer is used to generate a dual - mode attention vector by performing a Sigmoid function operation on the output of the first addition layer; the channel splitting layer is used to split Channel splitting is performed to obtain the first attention vector and the second attention vector ; the Softmax function layer is used to perform the Softmax function operation on and respectively, and correspondingly generate the first channel weight and the second channel weight ; the first multiplication layer is used to perform element-wise multiplication on and to obtain the output ; the second multiplication layer is used to perform element-wise multiplication on and to obtain the output ; the second concatenation layer is used to concatenate and into the first output feature ; the third concatenation layer is used to concatenate and into the second output feature .

[0031] Example 3: Refer to Figures 1 to 8 , the MobileNetV3 network uses MobileNetV3-Small, including a first convolutional layer, 11 bneck layers, and a second convolutional layer, a total of 13 layers, and is divided into 5 stages from the 3rd, 5th, 8th, 10th, and 13th layers respectively. The rest is the same as in Example 1.

[0032] Example 4: Refer to Figures 1 to 8 , on the basis of Example 1, in step S4 of the dual-channel feature extraction network constructed by the present invention, it is used to replace the backbone network of the DeepLab V3+ network. Therefore, the dual-channel feature extraction network can also be regarded as the backbone network of the improved semantic segmentation network. Through step S5, training the improved semantic segmentation network can obtain a bimodal rock thin section image semantic segmentation model. In this embodiment, when constructing the dual-channel feature extraction network, MobileNetV3 uses MobileNetV3-Large and MobileNetV3-Small respectively, and different hyperparameters are used for experiments. The experimental data is as follows in Table 1: Table 1. Comparison table of experimental results of selecting different MobileNetV3 networks and hyperparameters , In Table 1, mPA is Mean Pixel Accuracy in English, which is the average pixel accuracy of the category in Chinese; MIoU is Mean Intersection over Union in English, which is the average intersection over union ratio in Chinese. As can be seen from Table 1, when constructing a dual-channel feature extraction network, the present invention has the best performance when using MobileNetV3-Large and setting the fusion ratio hyperparameter to 0.2, and its average intersection over union ratio MIoU in the dual-modal rock thin section dataset can reach 73.43%.

[0033] Example 5: Based on Example 1, we use different semantic segmentation models to conduct performance comparison experiments, and obtain Table 2; Table 2. Comparative experiments on performance of different semantic segmentation networks Network Input image size <![CDATA[FLOPs(1×10 10 )]]> Model size (MB) Inference time (ms) mPA (%) MIoU (%) DANet <![CDATA[3×480 2 > 189.43 189 11.48 83.94 71.59 Unet <![CDATA[3×512 2 > 225.41 209 12.25 83.79 70.92 PSPNet <![CDATA[3×473 2 > 193.96 178 11.62 83.86 71.46 DeepLabV3+ <![CDATA[3×512 2 > 242.00 22.4 19.85 84.69 71.82 OCNet <![CDATA[3×512 2 > 138.35 137 10.15 83.89 70.61 Bi-RGB (Small) <![CDATA[2×3×512 2 > 10.99 30.5 15.45 84.50 72.46 Bi-RGB (Large) <![CDATA[2×3×512 2 > 32.33 56.5 17.23 85.79 73.43 In Table 2, Bi-RGB (Small) is a model obtained by using MobileNetV3-Small for two MobileNetV3 networks in the dual-channel feature extraction network of the present invention, and Bi-RGB (Large) is a model obtained by using MobileNetV3-Large for two MobileNetV3 networks in the dual-channel feature extraction network of the present invention. FLOPs is Floating Point Operations, which refers to the number of floating point operations per second and is an indicator of the computing power of a computer or model.

[0034] It can be seen from Table 2 that compared with other semantic segmentation networks, the dual-channel feature fusion segmentation network of the present invention can achieve the highest MIoU performance index while having fewer parameters in the Bi-RGB (Large) state.

[0035] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A dual-modal rock thin section image semantic segmentation method based on dominant feature fusion, characterized by: The steps include: S1, obtaining a data set D, wherein the data set D includes a plurality of sample pairs, wherein one of the sample pairs includes a single polarization image and an orthogonal polarization image of the same rock thin section; S2, construct the AFFM module, including the bimodal channel separation module, the dual-branch spatial attention module and the fusion module; The dual-modal channel separation module is used to input the first input feature , the second input feature , after extracting the complementary mapping relationship between the two from the channel dimension, generate The corresponding first attention vector , first channel weight and the first output feature ,generate The corresponding second attention vector , the second channel weight and the second output feature ; The dual-branch spatial attention module is used to respectively , Perform spatial attention and output the corresponding spatial attention weights , ; The fusion module is used to generate the output of the AFFM module , , where is element-wise multiplication, Add the elements together; S3, constructing a dual-channel feature extraction network, including S31~S32; S31, obtain two MobileNetV3 networks with the same structure, marked as , The MobileNetV3 network includes the first convolutional layer, the n+2 layer, and the second convolutional layer. The n+2 layer is divided into 5 stages from top to bottom, and each stage outputs a feature map. Used to input a single polarization image and output the i-th single polarization feature image in sequence after the i-th stage , Used to input the orthogonal polarization image and output the i-th orthogonal polarization feature image in sequence after the i-th stage , 1≤i≤5; S32, in and Set the i-th AFFM module , and and As The first input feature and the second input feature of the first four stages, The output of Add as The input of the i+1th stage, The output of Add as The input of the i+1th stage, for the 5th stage, The output of is used as the output of the dual-channel feature extraction network; S4, obtain a DeepLab V3+ network, replace its backbone network with a dual-channel feature extraction network, and obtain an improved semantic segmentation network; S5, using the dataset D to train the improved semantic segmentation network until convergence, and obtain the bimodal rock thin section image semantic segmentation model; S6, obtaining a single polarization image and an orthogonal polarization image of the rock slice to be identified, inputting them into a dual-modal rock slice image semantic segmentation model, and outputting a semantic segmentation result.

2. The dual-modal rock thin section image semantic segmentation method based on dominant feature fusion according to claim 1 is characterized in that: The dual-modal channel separation module is used to input the first input feature , the second input feature , spliced ​​as splicing features , The first aggregate feature is generated through global maximum pooling and MLP , the other path generates the second aggregate feature through average pooling and MLP , and then , Generate a bimodal attention vector by element addition and Sigmoid function , perform channel splitting to obtain the first attention vector , the second attention vector ; , The first channel weight is generated by the Softmax function respectively , the second channel weight , and Multiply the elements and then add Concatenate as the second output feature , and Multiply the elements and then add Concatenate as the first output feature .

3. The dual-modal rock thin section image semantic segmentation method based on dominant feature fusion according to claim 1, characterized in that: The dual-branch spatial attention module includes a first spatial attention unit and a second spatial attention unit with the same structure; The first spatial attention unit includes a global maximum pooling layer, an average pooling layer, a splicing layer, a convolution layer and a sigmoid function layer; The global maximum pooling layer performs the global maximum pooling operation in the channel dimension to obtain the output , The average pooling layer performs the average pooling operation in the channel dimension to obtain the output , the splicing layer is used to , Features obtained by splicing , reduced in dimension by the convolutional layer, and then output by the sigmoid function layer ; The second spatial attention unit is used to input , Output .

4. The dual-modal rock thin section image semantic segmentation method based on dominant feature fusion according to claim 1, characterized in that: S4 specifically includes steps S41 to S43; S41, obtain a DeepLab V3+ network, including a backbone network, ASPP and decoder; S42, replace the backbone network with a dual-channel feature extraction network The output of ASPP is output ; S43, and All are sent to the decoder to obtain the semantic segmentation results.

5. The dual-modal rock thin section image semantic segmentation method based on dominant feature fusion according to claim 1, characterized in that: The MobileNetV3 network adopts MobileNetV3-Large, including the first convolutional layer, 15 bneck layers, and the second convolutional layer, a total of 17 layers, which are divided into 5 stages from the 3rd, 5th, 8th, 14th, and 17th layers respectively.

6. The dual-modal rock thin section image semantic segmentation method based on dominant feature fusion according to claim 1, characterized in that: The MobileNetV3 network adopts MobileNetV3-Small, including the first convolutional layer, 11 bneck layers, and the second convolutional layer, a total of 13 layers, which are divided into 5 stages from the 3rd, 5th, 8th, 10th, and 13th layers respectively.

7. The dual-modal rock thin section image semantic segmentation method based on dominant feature fusion according to claim 1, characterized in that: The dataset D is the rock thin section microscopic image dataset of Nanjing University, the rock thin section microscopic image dataset of PaddlePaddle AI Studio Galaxy Community, or the rock thin section microscopic image dataset of the Xuzhuang Formation of the Middle Cambrian in the Ordos Basin.

Citation Information

Cited By

  • Lithium ore microscopic image segmentation method and system based on improved Unet model

    CN121304707A

  • Bridge crack image segmentation system and method thereof

    CN121982045A