An RGB-T image semantic segmentation method and device

By using a dual-branch RGB-T semantic segmentation network for spatial cross-modal information fusion and multi-scale feature iterative fusion, combined with RGB image random mask data enhancement, the problem of not fully utilizing modal complementarity in RGB-T semantic segmentation is solved, achieving high-performance semantic segmentation and low-cost annotation under poor lighting conditions.

CN116091765BActive Publication Date: 2026-01-06TSINGHUA UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211715697.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-01-06
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

Existing RGB-T semantic segmentation methods fail to fully utilize the spatial complementarity between RGB features and thermal infrared features, resulting in poor semantic segmentation performance under adverse lighting conditions.

Method used

A dual-branch RGB-T semantic segmentation network with spatial cross-modal information fusion and multi-scale feature iterative fusion is adopted, combined with the RGB image random mask data augmentation method, to train the RGB-T image semantic segmentation model and deeply mine cross-modal spatial complementary information.

Benefits of technology

It improves semantic segmentation performance under poor lighting conditions, reduces annotation costs, and achieves more accurate semantic segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091765B_ABST
    Figure CN116091765B_ABST
Patent Text Reader

Abstract

This invention provides an RGB-T image semantic segmentation method and apparatus, comprising: pre-training an RGB-T image semantic segmentation model on a semi-annotated RGB-T image pair dataset using spatial cross-modal information fusion, multi-scale feature iterative fusion, and RGB image random mask data augmentation methods, thereby improving the RGB-T image semantic segmentation model's ability to mine cross-modal spatial complementary information and its semantic segmentation performance under poor lighting conditions, while reducing annotation costs. In the application phase, the RGB-T image semantic segmentation model is used to generate semantically segmented images of the target RGB-T image pairs, improving the accuracy of the semantic segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an RGB-T image semantic segmentation method and apparatus. Background Technology

[0002] Semantic segmentation aims to assign a category label to each pixel in an RGB image. As one of the key technologies for scene perception, it plays a vital role in computer vision tasks such as autonomous driving, pedestrian detection, and remote sensing image analysis.

[0003] Because RGB images may lack texture information in some areas under poor lighting conditions (too low brightness or overexposure), directly performing semantic segmentation on RGB images with missing texture information will inevitably lead to unreliable semantic segmentation results. Therefore, RGB-T semantic segmentation, which uses thermal infrared images to supplement the texture information of RGB images, has emerged. Most existing RGB-T semantic segmentation methods employ modal feature self-enhancement followed by additive fusion or modal feature alignment followed by channel-dimensional fusion to achieve the fusion of RGB features and thermal infrared features, and then use the fused features to complete the image semantic segmentation.

[0004] However, neither modal feature self-enhancement followed by additive fusion nor modal feature alignment followed by channel dimension fusion fully utilizes the spatial complementarity between modal features, resulting in poor performance of existing RGB-T semantic segmentation. Summary of the Invention

[0005] This invention provides an RGB-T image semantic segmentation method and apparatus to address the problem of poor semantic segmentation performance in existing technologies due to insufficient utilization of the spatial complementarity between RGB features and thermal infrared features. By employing spatial cross-modal information fusion, multi-scale feature iterative fusion, and RGB image random mask data augmentation methods, an RGB-T image semantic segmentation model is trained on a semi-annotated RGB-T image pair dataset. This enhances the RGB-T image semantic segmentation model's ability to mine cross-modal spatial complementary information, enabling the RGB-T image semantic segmentation model to have low annotation costs and high semantic segmentation performance under poor lighting conditions. Consequently, accurate semantic segmentation can be achieved using the RGB-T image semantic segmentation model.

[0006] In a first aspect, the present invention provides an RGB-T image semantic segmentation method, the method comprising:

[0007] The RGB-T image semantic segmentation model is invoked; wherein, the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing an RGB branch and a thermal infrared branch using the RGB-T image semantic segmentation dataset;

[0008] The target RGB-T image is input into the RGB-T image semantic segmentation model to obtain the first semantic segmentation image output by the RGB branch and the second semantic segmentation image output by the thermal infrared branch.

[0009] Based on the missing texture information of the target RGB-T image pair, one of the first semantic segmentation image and the second semantic segmentation image is selected as the semantic segmentation image of the target RGB-T image pair;

[0010] The RGB-T image semantic segmentation dataset is obtained by augmenting the first dataset with RGB image random masking; the first dataset is obtained by pixel-level semantic segmentation annotation of a portion of the RGB-T image pairs in the dataset composed of RGB-T image pairs.

[0011] The dual-branch RGB-T semantic segmentation network deeply mines the cross-modal spatial complementary texture features of the input RGB-T image pairs through spatial cross-modal information fusion and multi-scale feature iterative fusion.

[0012] In a second aspect, the present invention provides an RGB-T image semantic segmentation apparatus, the apparatus comprising:

[0013] The calling module is used to call the RGB-T image semantic segmentation model; wherein, the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing RGB branch and thermal infrared branch using the RGB-T image semantic segmentation dataset;

[0014] The generation module is used to input the target RGB-T image pair into the RGB-T image semantic segmentation model to obtain the first semantic segmentation image output by the RGB branch and the second semantic segmentation image output by the thermal infrared branch;

[0015] The selection module is used to select one of the first semantic segmentation image and the second semantic segmentation image as the semantic segmentation image of the target RGB-T image pair based on the missing texture information of the target RGB-T image pair.

[0016] The RGB-T image semantic segmentation dataset is obtained by augmenting the first dataset with RGB image random masking; the first dataset is obtained by pixel-level semantic segmentation annotation of a portion of the RGB-T image pairs in the dataset composed of RGB-T image pairs.

[0017] The dual-branch RGB-T semantic segmentation network deeply mines the cross-modal spatial complementary texture features of the input RGB-T image pairs through spatial cross-modal information fusion and multi-scale feature iterative fusion.

[0018] The present invention provides an RGB-T image semantic segmentation method and apparatus, which pre-trains a dual-branch RGB-T semantic segmentation network containing RGB and thermal infrared branches using an RGB-T image semantic segmentation dataset to obtain an RGB-T image semantic segmentation model. Since (1) the dual-branch RGB-T semantic segmentation network adaptively and complementaryly fuses the RGB modal features and thermal infrared modal features of the input RGB-T image pair through spatial cross-modal information fusion, and compensates for spatial information loss during feature extraction through multi-scale feature iterative fusion, it has the ability to deeply mine the cross-modal spatial complementary texture features of RGB-T image pairs. (2) The dual-branch RGB-T semantic segmentation network enables the RGB-T image semantic segmentation model to better cope with the loss of texture signals in single-modal data. (3) The RGB-T image semantic segmentation dataset is obtained by data augmentation of a semi-annotated RGB-T image pair dataset using an RGB modal random mask method, which introduces new intermodal spatial complementary regions, which is beneficial for fully utilizing the annotated data to train the RGB-T image semantic segmentation model. Therefore, the resulting RGB-T image semantic segmentation model can better utilize the complementary information of cross-modal data space. Compared with existing RGB-T semantic segmentation techniques, it achieves better semantic segmentation performance under adverse lighting conditions and incurs lower annotation costs, contributing to cost reduction and efficiency improvement in fine-grained perception of complex environments. By inputting the target RGB-T image pair into the RGB-T image semantic segmentation model, a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch are obtained. Based on the missing texture information of the target RGB-T image pair, one of the first and second semantic segmentation images is selected as the semantic segmentation image for the target RGB-T image pair, resulting in more accurate semantic segmentation results. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating an RGB-T image semantic segmentation method provided by the present invention;

[0021] Figure 2 This is a schematic diagram of the structure of the dual-branch RGB-T semantic segmentation network provided by the present invention;

[0022] Figure 3 This is an example diagram of fully supervised learning provided by the present invention;

[0023] Figure 4 This is an example diagram of cross-modal mutual learning provided by the present invention;

[0024] Figure 5 This is the RGB-T feature stream provided by the present invention;

[0025] Figure 6 This is a schematic diagram of the spatial cross-modal information fusion module provided by the present invention;

[0026] Figure 7 This is a schematic diagram of the structure of the multi-scale feature iterative fusion module provided by the present invention;

[0027] Figure 8 This is a schematic flowchart of an RGB-T image semantic segmentation device provided by the present invention;

[0028] Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention;

[0029] Figure label:

[0030] 910: Processor; 920: Communication interface; 930: Memory; 940: Communication bus. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0032] First, the definitions of abbreviations and key terms in this field will be explained:

[0033] mIoU: Mean intersection of union.

[0034] SCF: Spatial-wise cross-modal fusion.

[0035] RMM: Recursive multi-scale meshing, iterative fusion of features at multiple scales.

[0036] M-CutOut: Mono-modal CutOut, a method for augmenting data using a single-modal random mask.

[0037] Conv: Convolution, the convolution operation

[0038] CD: Channe-wise denoising, channel-adaptive noise reduction

[0039] ADM: Attentive Demand Map, Spatial Adaptive Demand Map Assessment

[0040] CF: Cross-modal fusion

[0041] ASPP: Atrous spatial pyramid pooling

[0042] SF: Spatial-wise fusion

[0043] CA: Channel-wise attention.

[0044] SA: Spatial-wise attention, spatial attention mechanism

[0045] BN: Batch normalization.

[0046] MLP: Multilayer Perceptron

[0047] The following is combined with Figures 1-9 This invention describes an RGB-T image semantic segmentation method and apparatus.

[0048] In a first aspect, the present invention provides an RGB-T image semantic segmentation method, such as... Figure 1 As shown, it includes:

[0049] S11, Call the RGB-T image semantic segmentation model;

[0050] Specifically, the RGB-T image semantic segmentation model is pre-built, and its construction process includes:

[0051] Construct an RGB-T image semantic segmentation dataset; the RGB-T image semantic segmentation dataset can be obtained as follows:

[0052] Step A: Collect RGB-T data pairs to form dataset T;

[0053] Optionally, the dataset T can be open source data, such as the MFNet dataset for road scenes and the PST900 dataset for underground scenes.

[0054] Step B: Perform pixel-level semantic segmentation and annotation on a portion of the RGB-T image pairs in dataset T to obtain the first dataset;

[0055] Since pixel-level semantic segmentation annotation of RGB-T image pairs is very costly, this invention can be directly applied to semi-supervised tasks through cross-modal mutual learning (i.e., some RGB-T image pairs are used as labeled image pairs, and other RGB-T image pairs are used as unlabeled image pairs) to reduce the annotation cost of RGB-T image semantic segmentation models.

[0056] Step C: Use RGB image random masking to augment the first dataset to obtain the RGB-T image semantic segmentation dataset;

[0057] To fully utilize labeled image pairs, this invention proposes a single-modal random masking data augmentation method, M-CutOut. This method performs random masking operations on the RGB images in the RGB-T image pairs in the first dataset to artificially introduce new spatial complementary information regions, enabling the RGB-T image semantic segmentation model to better utilize cross-modal spatial complementary texture information during training.

[0058] Furthermore, a random masking operation is performed on the RGB images in the RGB-T image pair, including:

[0059] First, initialize a mask M with all 1s. Then, determine a rectangular region at a random location that is proportional to the image scale and set the interior of this region of M to 0. Finally, multiply the mask pixel by pixel with the original RGB image to form a new region lacking visible light texture information.

[0060] A dual-branch RGB-T semantic segmentation network is constructed; both the RGB branch and the thermal infrared branch in the dual-branch RGB-T semantic segmentation network adopt an encoder-decoder structure, wherein:

[0061] The RGB branch includes a first input layer, a first feature encoder, a first feature decoder, a first pixel-level classifier, and a main prediction output layer;

[0062] The thermal infrared branch includes a second input layer, a second feature encoder, a second feature decoder, a second pixel-level classifier, and an auxiliary prediction output layer;

[0063] The first feature encoder includes K feature extraction layers;

[0064] The second feature encoder includes K feature extraction layers;

[0065] The first feature encoder and the second feature encoder together include K spatial cross-modal information fusion modules;

[0066] The first feature decoder includes a first multi-scale feature iterative fusion module;

[0067] The second feature decoder includes a second multi-scale feature iterative fusion module. Figure 2 A schematic diagram of a dual-branch RGB-T semantic segmentation network is provided (without loss of generality, all figures in this invention are exemplified using K=5). In the diagram, SCF (Spatial-wise Cross-modal Fusion) represents the spatial cross-modal information fusion module, RMM (Recursive Multi-scale Meshing) represents the multi-scale feature iterative fusion module, and L... i This represents the feature extraction layer with index i.

[0068] from Figure 2 It can be seen that the RGB branch and thermal infrared branch in the dual-branch RGB-T semantic segmentation network progressively extract features using K-layer feature extraction layers in the encoding structure, and gradually perform spatial cross-modal information fusion using SCF modules embedded after each feature extraction layer. In the decoding structure, the RMM module performs iterative feature fusion of spatial dimensions on multi-scale fused features to compensate for the information loss caused by spatial downsampling during feature encoding. During feature decoding, scale transformation is achieved through upsampling methods (such as bilinear upsampling). Spatial cross-modal information fusion and multi-scale feature iterative fusion together enable the dual-branch RGB-T semantic segmentation network to have the ability to deeply mine cross-modal spatial complementary texture features of RGB-T image pairs.

[0069] Construct a loss function; wherein, as mentioned above, when training the RGB-T image semantic segmentation model, this invention directly applies cross-modal mutual learning to the semi-supervised semantic segmentation task, that is, using... G represents an RGB-T image pair with pixel-level semantic segmentation annotations. Pixel-level semantic segmentation annotation, with Represents RGB-T image pairs without pixel-level semantic segmentation annotations, and for Adopting such Figure 3 The fully supervised learning method shown is for Adopting such Figure 4 The cross-modal mutual learning method is shown.

[0070] Therefore, the loss function of the RGB-T image semantic segmentation model The expression can be:

[0071]

[0072]

[0073]

[0074]

[0075]

[0076] in, The total training loss is denoted as RGB-T image pairs with pixel-level semantic segmentation annotations. h is the total training loss for RGB-T image pairs without pixel-level semantic segmentation annotations. rgb (·) represents the RGB branch of the RGB-T image semantic segmentation model, h th e(·) represents the thermal infrared branch of the RGB-T image semantic segmentation model, CE(·) represents the cross-entropy loss function, and Y rgb for The corresponding pseudo-tag, Ythe, is The corresponding pseudo-tag, M is Random mask image, Can be seen as The result after M-CutOut data augmentation, y represents the semantic segmentation prediction for a single pixel. Pseudo-label Y rgb and Y the It is generated using the semantic segmentation prediction results of conventional weak data augmentation (such as flipping).

[0077] Train an RGB-T image semantic segmentation model; wherein, the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing RGB branch and thermal infrared branch using the RGB-T image semantic segmentation dataset and loss function.

[0078] The RGB-T image semantic segmentation model trained by this invention has low annotation cost and high semantic segmentation performance under poor lighting conditions.

[0079] Three existing solutions are proposed: Solution 1 (patent CN112991350A), Solution 2 (patent CN113362349A), and Solution 3 (patent CN113781504A). All of these existing solutions use fully supervised training to build the model. In contrast, this invention directly applies a cross-modal mutual learning method to a semi-supervised task (using only half of the MFNet dataset as labeled data and the other half as unlabeled data) to train the model.

[0080] Table 1 shows the comparison of model test results on the MFNet RGB-T image dataset for road scenes;

[0081] Table 2 shows a comparison of semi-supervised semantic segmentation performance on the MFNet dataset of RGB-T images of road scenes.

[0082] Tables 1 and 2 show that the method of the present invention has better semantic segmentation performance under the same experimental settings, and can achieve the effect of the existing technical solution using all labeled data by using only half of the labeled data.

[0083] Table 1

[0084]

[0085] Table 2

[0086]

[0087] S12. Input the target RGB-T image pair into the RGB-T image semantic segmentation model to obtain the first semantic segmentation image output by the RGB branch and the second semantic segmentation image output by the thermal infrared branch;

[0088] Specifically, the first feature encoder includes K feature extraction layers, denoted as L. rgb,0 To L rgb,K-1 ;

[0089] The second feature encoder includes K feature extraction layers, denoted as L. the,0 To L the,K-1 ;

[0090] The K spatial cross-modal information fusion modules jointly included by the first feature encoder and the second feature encoder are denoted as SCF0 to SCF0, respectively. K -1;

[0091] S12 includes:

[0092] The RGB image in the target RGB-T image pair is transmitted to the first feature encoder through the first input layer, and the thermal infrared image in the target RGB-T image pair is transmitted to the second feature encoder through the second input layer;

[0093] In the combined structure of the first feature encoder and the second feature encoder, the L is utilized rg b ,i Feature extraction is performed on the input RGB information to obtain fragb,i, and then the L... the,i Feature extraction is performed on the input thermal infrared information to obtain f the,i and utilizing the SCF i For fr gb i and f the i performs spatial cross-modal information fusion to obtain f rgb, i and f the ,i; where i∈[0, (K-1)], and when i=0, L rgb,i The RGB information input is the RGB image in the target RGB-T image pair, L rgb,i The input thermal infrared information is the thermal infrared image in the target RGB-T image pair, where L is the thermal infrared image when i∈[1, (K-1)]. rgb,i The RGB information input is obtained using the SCF. i -1 yields f'r gb,i -1,L rgb The thermal infrared information input in i is obtained using the SCF. i -1 yields f the,i -1;

[0094] Using the first multi-scale feature iterative fusion module to process fr gb,K -1、f m,k -2、f m,k -3 and f m K-4 performs spatial dimension iterative fusion to obtain decoded features. Using the second multi-scale feature iterative fusion module to process f the K-1f m,k -2、f m,k -3 and f m,k -4. Spatial dimension iterative fusion is performed to obtain decoded features. Where, j∈(2~K), f m,k-j It is fr gb ,k -j with f the ,k- j Features obtained after additive fusion;

[0095] Will and After additive fusion, a first additive fusion feature is obtained. The first pixel-level classifier is used to process the first additive fusion feature to obtain the first semantic segmentation image yr. gb The first semantic segmentation image yr is output through the main prediction output layer. gb Processing using a second-pixel-level classifier The second semantic segmentation image ythe is obtained, and the second semantic segmentation image y is output through the auxiliary prediction output layer. the ;

[0096] Here, the pixel-level classifier mainly performs convolution operations, that is:

[0097] Furthermore, Figure 5The RGB-T feature flow is illustrated in the feature encoding and decoding stages. Here, CD (Channel-wise Denoising) is the channel-adaptive denoising module, ADM (Attentive Demand Map) is the spatial adaptive demand map estimator, CF (Cross-modal fusion) is the cross-modal fusion module, ASPP (Atrous Spatial Pyramid Pooling) is the spatial pyramid pooling module, and SF (Spatial-wise Fusion) is the spatial dimension fusion module. CD, ADM, and CF together form the SCF module, while ASPP and SF form the RMM module.

[0098] Specifically, Figure 6 A schematic diagram of the spatial cross-modal information fusion module is provided. In real-world scenarios, both RGB images and infrared thermal images are inevitably subject to noise interference from complex environments, such as temperature fluctuations caused by strong light irradiation and abnormal heat sources. Considering this noise, the proposed SCF module employs two attention mechanisms: the channel-adaptive noise denoiser (CD) uses channel-wise attention (CA) to determine which channel is more reliable, and the spatial-adaptive demand graph estimator (ADM) uses spatial-wise attention (SA) to determine which region has a greater need for spatially complementary information fusion. These two attention mechanisms have various engineering implementations, such as the implementation in CBAM (S.Woo, J.Park, J-YLee, and ISKweon. CBAM: Convolutional block attention module. In ECCV, pages 3-19, 2018.), which are easily understood, extended, and implemented by those skilled in the art.

[0099] Therefore, the use of the SCF i For fr gb,i and f the,i Spatial cross-modal information fusion is performed to obtain fr g b ,i and f the i, including:

[0100] In the channel adaptive noise denoiser, for fr gb,i The first maximum pooling feature, MaxPool(f), is obtained by performing max pooling and mean pooling. rg b,i) and the first mean pooling feature MeanPool(f rgb,i Based on the MaxPool(f) rgb,i ) and the MeanPool(f rgbi) Generate the first channel attention map using the channel attention mechanism. and the With the aforementioned fr gb,i The product of is used as the fr gb,i noise reduction features

[0101] At the same time, for f the,i The second maximum pooling feature, MaxPool(f), is obtained by performing max pooling and mean pooling. the,i ) and the second mean pooling feature MeanPool(f the,i Based on the MaxPool(f) the,i ) and the MeanPool(f the,i The channel attention mechanism generates a second channel attention map. and the With the f the The product of i is used as f. the,i noise reduction features

[0102] Here, the stated The calculation formula is as follows:

[0103]

[0104] Here, the stated The calculation formula is as follows:

[0105]

[0106] Where MLP stands for Multilayer Perceptron and Sigmoid(·) is the Sigmoid function.

[0107] In the spatial adaptive demand graph estimator, based on Know Generate a first spatial adaptive demand graph using spatial attention mechanism Simultaneously based on Know Generating a second-space adaptive demand graph using spatial attention mechanism

[0108] Here, the stated The calculation formula is as follows:

[0109]

[0110] The The calculation formula is as follows:

[0111]

[0112] Wherein, Conv(·) is the convolution function, using a 7*7 convolution kernel, and Sigmoid(·) is the Sigmoid function. The role of the spatial adaptive demand map is to represent the spatial complementary information fusion demand of the region.

[0113] In the cross-modal fusion machine, and dot product result and Perform additive fusion to obtain the frgb,i, and then... and The dot product result and The f is obtained by performing additive fusion. the,i .

[0114] Figure 7 A schematic diagram of the multi-scale feature iterative fusion module is provided. In the diagram, the Conv unit consists of convolution operation, batch normalization operation, and ReLU activation. Spatial downsampling in the feature encoding process inevitably causes information loss, while pixel-level semantic segmentation tasks rely on detailed texture features. The proposed RMM module compensates for information loss by iteratively fusing multi-scale features.

[0115] Therefore, the first multi-scale feature iterative fusion module is used to process fr g b, K-1, f m,K -2、f m K-3 and f m,K -4. Spatial dimension iterative fusion is performed to obtain decoded features. include:

[0116] Feature encoding result fr g b, K-1 first embeds more global features ASPP(f′rgb, K-1) through ASPP, where the ASPP uses an inflation coefficient d = 2, 4, 8, and the number of output feature channels is 256. Then, iteratively fuses the upsampled features frgb, uK-1 of ASPP(f′rgb, K-1) and the fused feature f containing rich texture information. m,K -2、f m,K -3 and f m,K -4 obtained Considering computational complexity, the aforementioned fusion features utilize a single Conv unit to achieve channel dimensionality reduction. Specifically, the... The calculation formula is as follows:

[0117]

[0118]

[0119]

[0120] Where z∈[1,3], For cascading operations, Up(·) is upsampling, MeanPool(·) is mean pooling, MaxPool(·) is max pooling, Sigmoid(-) is the Sigmoid function, and · is dot product. To An adaptive mask obtained through spatial attention mechanism evaluation is used to represent where and how much information compensation is needed. To pass f through Conv unit m,K Features obtained by channel dimensionality reduction of _z-1.

[0121] Similarly, the second multi-scale feature iterative fusion module is used to process f the K-1, f m K-2, f m,K -3 and f m,K -4. Spatial dimension iterative fusion is performed to obtain decoded features. include:

[0122] Feature encoding result f the K-1 first embeds more global features through ASPP(f′). the,K-1 Then iteratively fuse ASPP(f) the,K-1 upsampling features and the fusion feature f containing rich texture information m,K -2、f m,K -3 and f m,K -4 obtained

[0123] Specifically, the aforementioned The calculation formula is as follows:

[0124]

[0125]

[0126]

[0127] To An adaptive mask obtained by evaluating spatial attention mechanisms.

[0128] S13. Based on the missing texture information of the target RGB-T image pair, select one of the first semantic segmentation image and the second semantic segmentation image as the semantic segmentation map of the target RGB-T image pair.

[0129] Specifically, S13 includes:

[0130] When neither the RGB image nor the thermal infrared image in the target RGB-T image pair has missing texture information, the first semantic segmentation image is used as the semantic segmentation image of the target RGB-T image pair.

[0131] If either the RGB image or the thermal infrared image in the target RGB-T image pair lacks texture information, the first semantic segmentation image is used as the semantic segmentation image of the target RGB-T image pair for daytime scenes, and the second semantic segmentation image is used as the semantic segmentation image of the target RGB-T image pair for nighttime scenes.

[0132] The dual-branch structure of this invention enables the RGB-T semantic segmentation model to better handle single-modal data signal loss. Table 3 shows a comparison of the signal loss robustness experimental results on the MFNet RGB-T image dataset for road scenes. The experiments demonstrate that this invention can achieve good semantic segmentation results even with signal loss.

[0133] Table 3

[0134]

[0135] In summary, the RGB-T image semantic segmentation method provided by this invention pre-trains a dual-branch RGB-T semantic segmentation network containing RGB and thermal infrared branches using an RGB-T image semantic segmentation dataset to obtain an RGB-T image semantic segmentation model. Since (1) the dual-branch RGB-T semantic segmentation network adaptively and complementaryly fuses the RGB modal features and thermal infrared modal features of the input RGB-T image pair through spatial cross-modal information fusion, and compensates for spatial information loss during feature extraction through multi-scale feature iterative fusion, it has the ability to deeply mine the cross-modal spatial complementary texture features of RGB-T image pairs. (2) The dual-branch RGB-T semantic segmentation network enables the RGB-T image semantic segmentation model to better cope with the loss of texture signals in single-modal data. (3) The RGB-T image semantic segmentation dataset is obtained by data augmentation of a semi-annotated RGB-T image pair dataset using an RGB modal random mask method. This introduces new intermodal spatial complementary information regions, which is beneficial for fully utilizing the annotated data to train the RGB-T image semantic segmentation model. Therefore, the resulting RGB-T image semantic segmentation model can better utilize the complementary information of cross-modal data space. Compared with existing RGB-T semantic segmentation techniques, it achieves better semantic segmentation performance under adverse lighting conditions and incurs lower annotation costs, contributing to cost reduction and efficiency improvement in fine-grained perception of complex environments. By inputting the target RGB-T image pair into the RGB-T image semantic segmentation model, a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch are obtained. Based on the missing texture information of the target RGB-T image pair, one of the first and second semantic segmentation images is selected as the semantic segmentation image for the target RGB-T image pair, resulting in more accurate semantic segmentation results.

[0136] Furthermore, based on this approach, similar results can be achieved by using different feature extraction backbone networks and modifying the parameters in each module (such as the number of convolutional layers, the number of channels, and activation functions). Similarly, similar semi-supervised semantic segmentation training can also be achieved by using different combinations of strong and weak data augmentation.

[0137] Secondly, the present invention provides an RGB-T image semantic segmentation apparatus, which can be referred to in correspondence with the RGB-T image semantic segmentation method described above. Figure 8 This is a flowchart illustrating an RGB-T image semantic segmentation device provided by the present invention, as shown below. Figure 8 As shown, the device includes:

[0138] The calling module is used to call the RGB-T image semantic segmentation model; wherein, the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing RGB branch and thermal infrared branch using the RGB-T image semantic segmentation dataset;

[0139] The generation module is used to input the target RGB-T image pair into the RGB-T image semantic segmentation model to obtain the first semantic segmentation image output by the RGB branch and the second semantic segmentation image output by the thermal infrared branch;

[0140] The selection module is used to select one of the first semantic segmentation image and the second semantic segmentation image as the semantic segmentation image of the target RGB-T image pair based on the missing texture information of the target RGB-T image pair.

[0141] The RGB-T image semantic segmentation dataset is obtained by augmenting the first dataset with RGB image random masking; the first dataset is obtained by pixel-level semantic segmentation annotation of a portion of the RGB-T image pairs in the dataset composed of RGB-T image pairs.

[0142] The dual-branch RGB-T semantic segmentation network deeply mines the cross-modal spatial complementary texture features of the input RGB-T image pairs through spatial cross-modal information fusion and multi-scale feature iterative fusion.

[0143] The present invention provides an RGB-T image semantic segmentation device, which pre-trains a dual-branch RGB-T semantic segmentation network containing RGB and thermal infrared branches using an RGB-T image semantic segmentation dataset to obtain an RGB-T image semantic segmentation model. (1) The dual-branch RGB-T semantic segmentation network adaptively and complementaryly fuses the RGB modal features and thermal infrared modal features of the input RGB-T image pair through spatial cross-modal information fusion, and compensates for the spatial information loss in the feature extraction stage through multi-scale feature iterative fusion, thus having the ability to deeply mine the cross-modal spatial complementary texture features of RGB-T image pairs. (2) The dual-branch RGB-T semantic segmentation network enables the RGB-T image semantic segmentation model to better cope with the loss of texture signals in single-modal data. (3) The RGB-T image semantic segmentation dataset is obtained by data augmentation of the semi-annotated RGB-T image pair dataset through RGB modal random masking, which introduces new intermodal spatial complementary information regions, which is beneficial to make full use of the annotated data for training the RGB-T image semantic segmentation model. Therefore, the resulting RGB-T image semantic segmentation model can better utilize the complementary information of cross-modal data space. Compared with existing RGB-T semantic segmentation techniques, it achieves better semantic segmentation performance under adverse lighting conditions and incurs lower annotation costs, contributing to cost reduction and efficiency improvement in fine-grained perception of complex environments. By inputting the target RGB-T image pair into the RGB-T image semantic segmentation model, a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch are obtained. Based on the missing texture information of the target RGB-T image pair, one of the first and second semantic segmentation images is selected as the semantic segmentation image for the target RGB-T image pair, resulting in more accurate semantic segmentation results.

[0144] Thirdly, Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940, wherein the processor 910, the communication interface 920, and the memory 930 communicate with each other through the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute an RGB-T image semantic segmentation method. This method includes: calling an RGB-T image semantic segmentation model; wherein the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing an RGB branch and a thermal infrared branch using an RGB-T image semantic segmentation dataset; inputting a target RGB-T image pair into the RGB-T image semantic segmentation model to obtain a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch; and based on the texture of the target RGB-T image pair... In the case of missing information, one of the first semantic segmentation image and the second semantic segmentation image is selected as the semantic segmentation image of the target RGB-T image pair; wherein, the RGB-T image semantic segmentation dataset is obtained by data augmentation of the first dataset using a random RGB image mask method; the first dataset is obtained by pixel-level semantic segmentation annotation of a portion of the RGB-T image pairs in the dataset composed of RGB-T image pairs; the dual-branch RGB-T semantic segmentation network deeply mines the cross-modal spatial complementary texture features of the input RGB-T image pairs through spatial cross-modal information fusion and multi-scale feature iterative fusion.

[0145] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0146] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute an RGB-T image semantic segmentation method provided by the above methods. The method includes: calling an RGB-T image semantic segmentation model; wherein the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing an RGB branch and a thermal infrared branch using an RGB-T image semantic segmentation dataset; inputting a target RGB-T image pair into the RGB-T image semantic segmentation model to obtain a first semantic segmentation image output by the RGB branch and the thermal infrared image. The second semantic segmentation image output by the infrared branch; based on the missing texture information of the target RGB-T image pair, one of the first semantic segmentation image and the second semantic segmentation image is selected as the semantic segmentation image of the target RGB-T image pair; wherein, the RGB-T image semantic segmentation dataset is obtained by performing data augmentation on the first dataset using a random RGB image mask; the first dataset is obtained by performing pixel-level semantic segmentation annotation on a portion of the RGB-T image pairs in the dataset composed of RGB-T image pairs; the dual-branch RGB-T semantic segmentation network deeply mines the cross-modal spatial complementary texture features of the input RGB-T image pairs through spatial cross-modal information fusion and multi-scale feature iterative fusion.

[0147] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an RGB-T image semantic segmentation method provided by the methods described above. This method includes: invoking an RGB-T image semantic segmentation model; wherein the RGB-T image semantic segmentation model is obtained by training a two-branch RGB-T semantic segmentation network containing an RGB branch and a thermal infrared branch using an RGB-T image semantic segmentation dataset; inputting a target RGB-T image pair into the RGB-T image semantic segmentation model to obtain a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch. Based on the missing texture information of the target RGB-T image pair, one of the first semantic segmentation image and the second semantic segmentation image is selected as the semantic segmentation image of the target RGB-T image pair; wherein, the RGB-T image semantic segmentation dataset is obtained by performing data augmentation on the first dataset using a random RGB image mask method; the first dataset is obtained by performing pixel-level semantic segmentation annotation on a portion of the RGB-T image pairs in the dataset composed of RGB-T image pairs; the dual-branch RGB-T semantic segmentation network deeply mines the cross-modal spatial complementary texture features of the input RGB-T image pairs through spatial cross-modal information fusion and multi-scale feature iterative fusion.

[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0149] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for RGB-T image semantic segmentation, characterized in that, The method comprises: calling an RGB-T image semantic segmentation model; wherein the RGB-T image semantic segmentation model is obtained by training a double-branch RGB-T semantic segmentation network containing an RGB branch and a thermal infrared branch using an RGB-T image semantic segmentation dataset; inputting a target RGB-T image pair into the RGB-T image semantic segmentation model to obtain a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch; selecting one of the first semantic segmentation image and the second semantic segmentation image as a semantic segmentation image of the target RGB-T image pair according to a texture information loss condition of the target RGB-T image pair; wherein the RGB-T image semantic segmentation dataset is obtained by performing data enhancement on a first dataset using an RGB image random mask method; and the first dataset is obtained by performing pixel-level semantic segmentation annotation on part of the RGB-T image pairs in a dataset composed of RGB-T image pairs; The double-branch RGB-T semantic segmentation network deeply excavates the cross-modal spatial complementary texture features of the input RGB-T image pair through spatial cross-modal information fusion and multi-scale feature iterative fusion. The RGB branch comprises a first input layer, a first feature encoder, a first feature decoder, a first pixel-level classifier, and a main prediction output layer. The thermal infrared branch comprises a second input layer, a second feature encoder, a second feature decoder, a second pixel-level classifier, and an auxiliary prediction output layer. The first feature encoder includes K feature extraction layers, respectively denoted as to ; The second feature encoder comprises K feature extraction layers, respectively denoted as to ; The first feature encoder and the second feature encoder jointly comprise K spatial cross-modal information fusion modules, respectively denoted as to ; In the combined structure of the first feature encoder and the second feature encoder, the following is utilized: Feature extraction is performed on the input RGB information to obtain Using the Feature extraction is performed on the input thermal infrared information to obtain and using the right and Spatial cross-modal information fusion is obtained and ;in, , and when hour, The RGB information input is the RGB image in the target RGB-T image pair. The input thermal infrared information is the thermal infrared image in the target RGB-T image pair, when hour, The RGB information input is obtained by using the above. Received , The thermal infrared information input is used by the above Received ; The spatial cross-modal information fusion module comprises a channel adaptive denoiser, a spatial adaptive demand map evaluator, and a cross-modal fusioner. The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: In the channel adaptive noise denoiser, for The first max pooling feature is obtained by performing max pooling and mean pooling. and the first mean pooling feature Based on the above and stated The first channel attention map is generated using the channel attention mechanism. and the With the The product of the products is used as the product of the products. noise reduction features ; Meanwhile, maximum pooling and mean pooling are performed on the first feature to obtain a second maximum-pooled feature and a second mean-pooled feature ;​​​​​​​​​ In the spatial adaptive demand graph evaluator, based on and , a first spatial adaptive demand graph is generated using a spatial attention mechanism ; meanwhile, based on and a second spatial adaptive demand graph is generated using a spatial attention mechanism ; In the cross-modal fuser, the and point multiplication results are additively fused to obtain the , and the and point multiplication results are additively fused to obtain the .​​ 2.The RGB-T image semantic segmentation method of claim 1, characterized in that, The first feature decoder comprises a first multi-scale feature iterative fusion module, and the second feature decoder comprises a second multi-scale feature iterative fusion module. The inputting of the target RGB-T image pair into the RGB-T image semantic segmentation model to obtain the first semantic segmentation image output by the RGB branch and the second semantic segmentation image output by the thermal infrared branch comprises: transmitting the RGB image in the target RGB-T image pair to the first feature encoder through the first input layer, and transmitting the thermal infrared image in the target RGB-T image pair to the second feature encoder through the second input layer; The first multi-scale feature iterative fusion module is used for iteratively fusing the spatial dimensions of the features 、 、 and to obtain decoding features , the second multi-scale feature iterative fusion module is used for iteratively fusing the spatial dimensions of the features 、 and to obtain decoding features ; wherein , is a feature obtained by additively fusing and . and obtaining a first additive fusion feature after additive fusion, processing the first additive fusion feature by using a first pixel-level classifier to obtain the first semantic segmentation image , and outputting the first semantic segmentation image by the main prediction output layer ;​ processing with a second pixel-level classifier obtaining the second semantic segmentation image and outputting the second semantic segmentation image through the auxiliary prediction output layer . 3.The RGB-T image semantic segmentation method of claim 1, wherein, The The calculation formula is as follows: ; The The calculation formula is as follows: ; where MLP is a multi-layer perceptron, Sigmoid is a sigmoid function, and Softmax is a softmax function. 4.The RGB-T image semantic segmentation method of claim 1, wherein, The The calculation formula is as follows: ; The The calculation formula is as follows: ; wherein, is a convolution function, Sigmoid is a Sigmoid function. 5.The RGB-T image semantic segmentation method of claim 2, wherein, The first multi-scale feature iterative fusion module and the second multi-scale feature iterative fusion module each comprise a spatial pyramid pooler and a spatial dimension fusioner. The first multi-scale feature iterative fusion module is used for 、 、 and spatial dimension iterative fusion to obtain the decoding feature , comprising: The first multi-scale feature iterative fusion module is used for generating the corresponding global feature by using a spatial pyramid pooling device in the first multi-scale feature iterative fusion module corresponding global feature; In the spatial dimension fuser in the first multi-scale feature iterative fusion module, the up-sampled features , , , and are iteratively fused in spatial dimension to obtain a decoded feature . The second multi-scale feature iterative fusion module is used for 、 、 and spatial dimension iterative fusion to obtain the decoding feature , comprising: determined by a spatial pyramid pooler in the second multi-scale feature iterative fusion module corresponding global features In the spatial dimension fuser in the second multi-scale feature iterative fusion module, the up-sampled features , , , and are iteratively fused in spatial dimension to obtain the decoding feature . 6.The RGB-T image semantic segmentation method of claim 5, characterized in that, The The calculation formula is as follows: ; ; ; The The calculation formula is as follows: ; ; ; wherein, , is a concatenation operation, is an up-sampling operation, is a mean pooling operation, is a max pooling operation, Sigmoid( ) is a Sigmoid function, is a point multiplication operation, is an adaptive mask obtained by evaluating a spatial attention mechanism on , is an adaptive mask obtained by evaluating a spatial attention mechanism on , is a feature obtained by reducing the channel dimension of through a Conv unit, which includes a convolution operation, a batch normalization operation, and a Relu activation. 7.The RGB-T image semantic segmentation method of claim 1, wherein, The loss function of the RGB-T image semantic segmentation model The expression is as follows: ; ; ; ; ; wherein , is the total training loss for RGB-T image pairs with pixel-level semantic segmentation annotation, is the total training loss for RGB-T image pairs without pixel-level semantic segmentation annotation, is a RGB-T image pair with pixel-level semantic segmentation annotation pixel-level semantic segmentation annotation, is the RGB branch of the RGB-T image semantic segmentation model, is the thermal infrared branch of the RGB-T image semantic segmentation model, denotes the cross-entropy loss function, is a RGB-T image pair without pixel-level semantic segmentation annotation, is the corresponding pseudo label, is the corresponding pseudo label, is a random mask map of is the semantic segmentation prediction corresponding to a single pixel. 8.The RGB-T image semantic segmentation method of any one of claims 1-7, characterized in that, The selecting of one of the first semantic segmentation image and the second semantic segmentation image as the semantic segmentation image of the target RGB-T image pair according to the texture information loss condition of the target RGB-T image pair comprises: in the case where neither the RGB image nor the thermal infrared image in the target RGB-T image pair has texture information loss, taking the first semantic segmentation image as the semantic segmentation image of the target RGB-T image pair; In a case that texture information is missing in either of the RGB image and the thermal infrared image in the target RGB-T image pair, the daytime scene takes the first semantic segmentation image as the semantic segmentation image of the target RGB-T image pair, and the nighttime scene takes the second semantic segmentation image as the semantic segmentation image of the target RGB-T image pair.

9. An RGB-T image semantic segmentation apparatus, characterized in that, The device comprises: a calling module configured to call an RGB-T image semantic segmentation model, wherein the RGB-T image semantic segmentation model is obtained by training a double-branch RGB-T semantic segmentation network comprising an RGB branch and a thermal infrared branch using an RGB-T image semantic segmentation dataset; a generating module configured to input a target RGB-T image pair into the RGB-T image semantic segmentation model to obtain a first semantic segmentation image output by the RGB branch and a second semantic segmentation image output by the thermal infrared branch; a selecting module configured to select one of the first semantic segmentation image and the second semantic segmentation image as a semantic segmentation image of the target RGB-T image pair according to a texture information missing condition of the target RGB-T image pair; wherein the RGB-T image semantic segmentation dataset is obtained by performing data enhancement on a first dataset in an RGB image random mask manner; and the first dataset is obtained by performing pixel-level semantic segmentation annotation on part of RGB-T image pairs in a dataset composed of RGB-T image pairs; the double-branch RGB-T semantic segmentation network deeply mines cross-modal spatial complementary texture features of the input RGB-T image pair through spatial cross-modal information fusion and multi-scale feature iterative fusion; wherein the RGB branch comprises a first input layer, a first feature encoder, a first feature decoder, a first pixel-level classifier, and a main prediction output layer; the thermal infrared branch comprises a second input layer, a second feature encoder, a second feature decoder, a second pixel-level classifier, and an auxiliary prediction output layer; The first feature encoder includes K feature extraction layers, respectively denoted as to ; The second feature encoder comprises K feature extraction layers, respectively denoted as to ; The first feature encoder and the second feature encoder jointly comprise K spatial cross-modal information fusion modules, respectively denoted as to ; wherein, in the combined structure of the first feature encoder and the second feature encoder, the feature extraction is performed on the input RGB information to obtain , the feature extraction is performed on the input thermal infrared information to obtain , and the spatial cross-modal information fusion is performed on and to obtain and ; wherein, , and when , the input RGB information in the is the RGB image in the target RGB-T image pair, the input thermal infrared information in the is the thermal infrared image in the target RGB-T image pair, and when , the input RGB information in the is obtained by the , and the input thermal infrared information in the is obtained by the wherein the spatial cross-modal information fusion module comprises a channel adaptive noise reducer, a spatial adaptive demand map evaluator, and a cross-modal fusioner. The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: The method comprises the following steps: In the channel adaptive noise reducer, the maximum pooling and mean pooling processing are performed on the to obtain a first maximum pooling feature and a first mean pooling feature , based on the and the , a first channel attention map is generated by using a channel attention mechanism , and a product of the and the is taken as a noise reduction feature of the ;​ At the same time, The second maximum pooling feature is obtained by performing max pooling and mean pooling. Second mean pooling feature Based on the above and stated The channel attention mechanism generates the second channel attention map. and the With the The product of as stated noise reduction features ; In the spatial adaptive demand graph evaluator, based on and , a first spatial adaptive demand graph is generated using a spatial attention mechanism ; while based on and a second spatial adaptive demand graph is generated using a spatial attention mechanism ; In the cross-modal fuser, the and point multiplication results are additively fused to obtain the , and the and point multiplication results are additively fused to obtain the .​​

Citation Information

Patent Citations

  • RGB-T image semantic segmentation method based on modal difference reduction

    CN112991350A

  • Road scene semantic segmentation method based on boundary guidance

    CN113781504A

  • Cross guide fusion RGB-T image saliency detection system

    CN113076947A

  • Video semantic segmentation method based on active learning

    US20220215662A1