Non-aligned visible light-thermal infrared image semantic segmentation method and system
By processing non-aligned visible light and thermal infrared images through the deformation-aware alignment enhancement module and the complementary feature aggregation module, the problems of poor segmentation accuracy and time consumption in the existing technology are solved, efficient semantic segmentation is achieved, and the application effect in actual scenarios is improved.
Patent Information
- Application Number
- CN202511115486.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-09-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing semantic segmentation methods for visible-light-thermal infrared images have poor accuracy when processing non-aligned images, and pixel-by-pixel alignment is time-consuming, making it difficult to achieve real-time deployment.
The deformation-aware alignment enhancement module and the complementary feature aggregation module are adopted to estimate the deformation field of thermal infrared features and perform spatial correction. Multi-scale contextual features and channel attention are combined to achieve cross-modal feature aggregation and reduce the interference of non-aligned images on the segmentation task.
The segmentation accuracy and efficiency of non-aligned visible-thermal infrared images are improved, the limitations of image alignment preprocessing are broken through, and the applicability and efficiency in practical scenarios are enhanced.
Smart Images

Figure CN120612490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a method and system for semantic segmentation of non-aligned visible light-thermal infrared images. Background Art
[0002] Visible-thermal (Red, Green, Blue, and Thermal) image semantic segmentation technology, by fusing the rich texture details of visible light images with the radiant temperature information of thermal infrared images, offers irreplaceable application value in complex environmental perception. This technology is primarily applied in four core scenarios: autonomous driving, security monitoring, industrial inspection, and medical navigation. In autonomous driving, especially in low-visibility conditions such as at night or in rain and fog, thermal infrared images can penetrate darkness and smoke to accurately capture temperature-sensitive targets such as pedestrians and animals, while visible light images provide detailed information such as road structure and traffic signs. The complementary integration of the two significantly reduces missed detection rates. In security and rescue scenarios, thermal infrared images can quickly locate hidden personnel or high-temperature fire sources in forest fires, while visible light images assist in recognizing facial features and environmental structures. This dual-modal collaboration significantly improves search and rescue efficiency. In industrial scenarios, this technology can simultaneously identify overheated areas of equipment (relying on thermal infrared images) and details of faulty components (relying on visible light images), enabling millimeter-level anomaly detection in critical facilities such as power plant pipelines, effectively preventing explosions. In medical navigation scenarios, this technology uses thermal infrared images to mark the boundaries of tissue blood flow changes and visible light images to locate anatomical structures, providing real-time navigation support for delicate surgeries such as tumor resection.
[0003] Therefore, improving the semantic segmentation accuracy of visible-thermal infrared imagery has far-reaching implications for practical applications. First, improved segmentation accuracy can directly enhance system safety and reliability. For example, accurately segmenting overheated areas in industrial equipment can prevent tens of millions of dollars in equipment damage. Second, improved segmentation accuracy can push the limits of environmental perception, enabling the perception system to maintain stable performance in extreme conditions such as day-night transitions and rain and fog, achieving true all-weather operation.
[0004] However, existing semantic segmentation methods for visible-thermal infrared images are mainly trained on aligned visible-thermal infrared images. However, real-world image pairs are often misaligned. Therefore, models trained in this way perform poorly when processing misaligned visible-thermal infrared images in real applications. The main problems are: (1) In non-aligned visible-thermal infrared images, objects at the same spatial location may appear at different locations in the two modal images, thereby reducing the accuracy of the fusion of the two modal images and resulting in a decrease in the accuracy of semantic segmentation; (2) Since the same object may have different shapes, textures, or appearances in non-aligned visible-thermal infrared images, valuable key information may be lost during image fusion, misleading the model to make incorrect judgments and reducing the segmentation accuracy after image fusion.
[0005] Furthermore, pixel-by-pixel alignment of visible-light and thermal-infrared images consumes significant computational power and time, making real-time deployment in real-world scenarios difficult. Therefore, achieving semantic segmentation of non-aligned visible-light and thermal-infrared images is a challenge that existing technologies need to address. Summary of the Invention
[0006] Therefore, the technical problem to be solved by the present invention is to overcome the problem of poor accuracy when directly processing non-aligned visible light-thermal infrared images in the prior art.
[0007] To solve the above technical problems, the present invention provides a semantic segmentation method for non-aligned visible light-thermal infrared images, comprising: The visible light image and thermal infrared image are passed through their corresponding feature encoders to obtain five levels of visible light features. and thermal infrared characteristics , is the hierarchical index; The visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ; Align enhancement features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ; The fifth level feature decoding block uses the visible light features of the current level , thermal infrared characteristics and cross-modal aggregation features As input, get the decoding features of the fifth level ; The feature decoding block from the fourth to the first level uses the visible light feature of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ; The decoded features of the first level The predicted segmentation result is obtained through the output layer .
[0008] Preferably, the visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ,include: Visible light features and thermal infrared characteristics Enter the alignment unit together to get the calibration feature ; Visible light features and thermal infrared characteristics After the spatial attention mechanism, the visible light spatial features are obtained and thermal infrared spatial characteristics , and then respectively with the calibration characteristics Multiply to get the visible light alignment feature and thermal infrared alignment features ; Aligning visible light to features , visible light characteristics And the output features of the visible light feature after the 1×1 convolution layer Add together to get the visible light space correction feature ; Thermal infrared alignment features , thermal infrared characteristics And the output features of thermal infrared features after 1×1 convolution layer Add together to get the thermal infrared spatial correction feature ; Visible light space correction features and thermal infrared spatial correction features After splicing, it passes through a 3×3 convolution layer and a Sigmoid activation function to obtain a weight tensor ; The weight tensor Respectively with visible light characteristics and thermal infrared characteristics After multiplication, concatenation is performed and then a 3×3 convolution layer is performed to obtain the alignment enhancement feature. .
[0009] Preferably, the visible light feature and thermal infrared characteristics Enter the alignment unit together to get the calibration feature ,include: Visible light features and thermal infrared characteristics Input deformation estimator to get deformation field ; The deformation field and thermal infrared characteristics Input the deformation function to obtain the calibrated thermal infrared characteristics ; The calibrated thermal infrared signature Visible light characteristics After splicing, it passes through a 3×3 convolution layer to obtain the calibration features .
[0010] Preferably, the alignment enhancement features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ,include: Align Enhancement Features and visible light characteristics Input the ASPP layer separately to obtain the multi-scale context features of the aligned enhanced features Multi-scale contextual features and visible light features , k=1,2,3 is the layer index of ASPP layer; Multi-scale contextual features that will align enhanced features Multi-scale contextual features of visible light features Subtract the corresponding ones to obtain multi-scale difference features , after splicing, input 1×1 convolution layer to obtain the target difference feature ; Multi-scale contextual features that will align enhanced features Multi-scale contextual features of visible light features Fusion is performed separately to obtain thermal infrared fusion features and visible light fusion features ; The target difference feature Fusion features with thermal infrared After multiplication, the features are fused with visible light. Add them together and input them into the 1×1 convolution layer to obtain the visible light supplementary features ; The target difference feature Fusion features with visible light After multiplication, it is fused with thermal infrared features Add them together and input them into the 1×1 convolution layer to get the thermal infrared supplementary features ; Supplementary features of visible light and thermal infrared supplementary features Visible light weight value is obtained through gating operation ; The visible light weight value Complementary features with visible light After multiplication, fusion features with thermal infrared Add together to get the visible light polymerization characteristics ; Will The value and thermal infrared supplementary characteristics After multiplication, the features are fused with visible light Add together to get the thermal infrared polymerization characteristics ; Visible light polymerization features and thermal infrared polymerization characteristics After splicing, it passes through a 1×1 convolution layer to obtain cross-modal aggregation features .
[0011] Preferably, the feature decoding block from the fourth to the first level is based on the visible light feature of the current level. , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ,include: The visible light features of the current level and thermal infrared characteristics After splicing, the input deformation estimator outputs the decoded deformation field Thermal infrared characteristics Then pass through the deformation function together, and the output decoded and corrected thermal infrared features Visible light characteristics After splicing, it passes through a 3×3 convolution layer to obtain the first branch feature ; Aggregate features across modalities and the decoding features of the previous layer After splicing, it passes through the residual channel attention block, 3×3 convolution layer and 3×3 dilated convolution in sequence to obtain the second branch feature ; The first branch feature and the second branch features After splicing, it passes through the residual channel attention block and the 3×3 convolution layer in turn to obtain the decoding features of the current level .
[0012] Preferably, when training the segmentation model consisting of the feature encoder, deformation-aware alignment enhancement module, complementary feature aggregation module, feature decoding block and output layer, the total loss function Including semantic segmentation loss , class-independent saliency loss and class-aware semantic margin loss .
[0013] Preferably, the class-independent significance loss include: The decoded features of the fifth level and the decoding features of the fourth level After passing through the 1×1 convolution layer and the 3×3 convolution layer respectively, the corresponding output features are obtained and ; Will After upsampling Splicing, sequentially passing through the residual channel attention block, 3×3 convolution layer and upsampling layer, to obtain the predicted position segmentation mask ; Based on the real segmentation annotation Get the true position segmentation mask ; Segmentation mask at predicted position and ground-truth segmentation mask The two-way cross entropy loss between the two classes calculates the class-independent significance loss .
[0014] Preferably, the class-aware semantic edge loss include: The decoded features of the third level, the second level, and the first level 、 and After passing through the 1×1 convolution layer and the 3×3 convolution layer respectively, the corresponding output features are obtained 、 and ; Will 、 and After splicing, it passes through the residual channel attention block, 3×3 convolution layer and upsampling layer in sequence to obtain the predicted semantic edge ; Get the real semantic edge based on the real segmentation annotation ; To predict semantic edges and true semantic edge The cross entropy loss between them calculates the class-aware semantic edge loss .
[0015] Preferably, the semantic segmentation loss To predict the segmentation results and ground-truth segmentation annotations The cross entropy loss between .
[0016] The present invention also provides a non-aligned visible light-thermal infrared image semantic segmentation system, comprising: The encoding module is used to pass the visible light image and thermal infrared image through their corresponding feature encoders to obtain five levels of visible light features. and thermal infrared characteristics , is the hierarchical index; Feature alignment module, used to align visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ; Feature aggregation module, used to align enhanced features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ; Decoding module, used for the fifth level feature decoding block to decode the visible light features of the current level , thermal infrared characteristics and cross-modal aggregation features As input, get the decoding features of the fifth level ; The feature decoding block from the fourth to the first level uses the visible light feature of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ; Output module, used to decode the features of the first level The predicted segmentation result is obtained through the output layer .
[0017] The above technical solution of the present invention has the following beneficial effects compared with the prior art: The present invention discloses a method for semantic segmentation of non-aligned visible light and thermal infrared images. First, a deformation-aware alignment enhancement module is used to process non-aligned visible light features and thermal infrared features. The module can estimate the deformation field of the thermal infrared features and combine the deformation field with the visible light features to achieve spatial alignment, thereby improving the spatial consistency between the visible light features and the thermal infrared features. A complementary feature aggregation module is then used to capture the unique clues in the two modalities of alignment enhancement features and visible light features, respectively, to promote cross-modal complementary interaction and provide comprehensive cross-modal aggregation features for the decoder. The present invention achieves direct semantic segmentation of non-aligned visible light and thermal infrared images, breaking through the limitations of image alignment preprocessing. It can effectively reduce the interference of non-aligned images on the segmentation task, improve the accuracy of non-aligned visible light and thermal infrared image segmentation, and significantly enhance the applicability and efficiency of visible light and thermal infrared image semantic segmentation technology in actual scenarios.
[0018] In order to alleviate the positional difference between non-aligned thermal infrared images and RGB images, the present invention introduces a deformation-aware alignment enhancement module. First, the thermal infrared features are corrected by the alignment unit to obtain the resampled deformation-aware calibration features. Then, the visible light features and thermal infrared features are further processed symmetrically. The calibration features are used to further explore the spatial position correlation between the original visible light features and thermal infrared features, and enhance the spatial consistency of cross-modal features. Finally, the characteristics of the visible light and thermal infrared modalities are fused, and the visible light features are used for spatial correction to obtain the alignment enhancement features, so as to reduce the interference of deformation between different modalities and achieve spatial alignment.
[0019] In order to explore the cross-modal complementary cross-cues between visible light features and thermal infrared features at different scales, the present invention introduces a complementary feature aggregation module. First, the ASPP layer is used to obtain multi-scale contextual features of different modalities, and the difference between features of the same scale is used as the cross-modal difference clue. The fused target difference features are used to enhance the correlation between the difference clues and the existing features, and visible light supplementary features and thermal infrared supplementary features are obtained; then, the gating operation is used to extract multimodal global information, which is used to adjust the channel relationship in each modal feature. The obtained visible light aggregated features and thermal infrared aggregated features can effectively enhance the independence of each modal information within the semantic channel; the final cross-modal aggregated features effectively integrate multimodal information across different scales and channels, and can significantly reduce the difference between different modal information.
[0020] Furthermore, the total loss function constructed by the present invention can effectively guide the segmentation model to pay more attention to the location and boundary information of the object by introducing class-independent saliency loss and class-aware semantic edge loss, which helps to reduce the interference caused by misalignment and improve the accuracy of predicted semantic segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to make the content of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings, wherein: Figure 1 It is a flow chart of a semantic segmentation method of non-aligned visible light-thermal infrared images of the present invention; Figure 2 This is the structural diagram of the deformation-aware alignment enhancement module; Figure 3 It is the structural diagram of the complementary feature aggregation module; Figure 4 It is the structural diagram of the feature decoding block; Figure 5 It is the structural diagram of the residual channel attention block; Figure 6 Schematic diagram of the class-independent saliency prediction layer and the class-aware edge generation layer; Figure 7 is an example of a deformed thermal infrared image obtained according to different deformation intensities, where Figure 7 Column (a) in is the visible light image and the true label, Figure 7 Column (b) is the original thermal infrared image. Figure 7 Column (c) shows the deformed thermal infrared image with a deformation intensity of 5. Figure 7 Column (d) shows the deformed thermal infrared image with a deformation intensity of 10. Figure 7 Column (e) is the deformed thermal infrared image obtained when the deformation intensity is 20. Figure 7 Column (f) shows the deformed thermal infrared image obtained when the deformation intensity is 50; Figure 8 This is a comparison of the segmentation results of the DAC-Net of the present invention and eight alignment methods on the U-MFNet dataset. Figure 8 Column (a) is the visible light image. Figure 8 Column (b) is a thermal infrared image. Figure 8 The (c) column in the figure is the segmentation result of MFNet. Figure 8 The (d) column in the figure is the segmentation result of RTFNet. Figure 8 Column (e) is the segmentation result map of FEANet. Figure 8 The (f) column in the figure is the segmentation result of GMNet. Figure 8 The (g) column in the figure is the segmentation result of EGFNet. Figure 8 The (h) column is the segmentation result map of MFTNet. Figure 8 The (i) column in the figure is the segmentation result map of LASNet. Figure 8 The (j) column in the figure is the segmentation result of MDRNet. Figure 8 The (k) column in the figure is the segmentation result of the DAC-Net of the present invention. Figure 8 The (l) column in is the true label. DETAILED DESCRIPTION
[0022] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention.
[0023] Semantic segmentation (SS) has advanced with the development of deep fully convolutional networks (FCNs). The DeepLab series introduced dilated convolutions, which are used in the Atrous Spatial Pyramid Pooling (ASPP) method to extract multi-scale information. DANet uses both positional and channel attention to model spatial and channel dependencies, respectively, improving segmentation results by capturing rich contextual dependencies.
[0024] However, these methods mainly rely on RGB images with good illumination, and their segmentation performance often degrades in harsh environments such as night, rain, and fog. Thermal infrared images provide valuable complementary information to RGB images, making up for the limitations of the RGB modality.
[0025] The researchers introduced a spatial transformer network (STN), which uses an affine transformation model to deform objects in an image, aligning them to a standard size and orientation, which significantly enhances the spatial invariance of the network.
[0026] However, achieving pixel-level alignment between RGB-T images in multi-sensor imaging is often a challenging and time-consuming task, which hinders the deployment of RGB-T image semantic segmentation in practical applications.
[0027] Therefore, existing RGB-T image semantic segmentation methods typically rely on aligned RGB-T images, leveraging the strengths of both modalities to provide discriminative features for complex scenes with similar textures and low-light backgrounds. For example, RTFNet employs upsampling skip connections to fuse RGB-T image information, leveraging the strengths of thermal infrared images. GMNet introduces a hierarchical feature extraction strategy, dividing multi-level features into primary, intermediate, and high-level representations. It further employs semantic, saliency, and boundary supervision to facilitate multi-task learning. LASNet decomposes RGB-T image semantic segmentation into three distinct steps: localization, activation, and sharpening. This method generates clear segmentation masks through a step-by-step process. EGFNet incorporates the Sobel operation to embed prior edge maps into feature maps, effectively capturing detailed information in the predicted masks. MFTNet develops a Transformer-based multimodal fusion method to explore the complementary nature of RGB-T modalities. MDRNet proposes a modality difference reduction module to adaptively select discriminative features for RGB-T fusion.
[0028] In summary, existing methods rely on well-aligned RGB-T image pairs. However, RGB-T image pairs in real scenes are usually not aligned. Pixel-by-pixel alignment of RGB-T images is both challenging and time-consuming, making it difficult to achieve real-time deployment in real scenes.
[0029] To solve this problem, this paper proposes a semantic segmentation method for non-aligned visible light and thermal infrared images based on deformation perception and complementarity (DAC-Net). Figure 1 Shown, including: S1: The visible light image and thermal infrared image are respectively passed through their corresponding feature encoders to obtain five levels of visible light features and thermal infrared features; S2: Input the visible light features and thermal infrared features of the same level into the Deformation-aware Alignment Enhancement (DAE) module of the same level to obtain the alignment enhancement features of each level; S3: Input the alignment enhancement features and visible light features of the same level into the complementary feature aggregation (CFA) module of the level to obtain cross-modal aggregated features of each level; S4: The Feature Decoding Block (FDB) at each level takes the visible light features, thermal infrared features, cross-modal aggregation features and decoding features of the previous level as input to obtain the decoding features of the current level; S5: Pass the decoded features of the first level through the output layer to obtain the predicted segmentation result.
[0030] Specifically, in S1, the feature encoders of visible light images and thermal infrared images both use the ResNet-152 backbone network, and the output visible light features are represented as , the thermal infrared characteristics are expressed as ,in is the level index, the height of the visible light image and thermal infrared image ,width , number of channels .
[0031] To alleviate the position difference between thermal infrared images and RGB images, the present invention introduces a deformation-aware alignment enhancement module in S2. The DAE module uses the alignment unit (AU) to estimate the thermal infrared features. and visible light characteristics The deformation field between , the deformation field The thermal infrared features are then used to correct the distortion; the DAE module then focuses on modeling the spatial correlation between RGB-T features, ultimately generating well-aligned and robust alignment-enhanced features. The formula of the deformation-aware alignment enhancement module can be expressed as: ; in, is the alignment enhancement feature of the i-th level, is the thermal infrared feature of the i-th level, is the visible light feature of the i-th level, It is a deformation-aware alignment enhancement module.
[0032] Specifically, refer to Figure 2 As shown, in S2, the visible light features of the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ,include:
[0033] S21: Visible light features and thermal infrared characteristics Enter the alignment unit together to get the calibration feature .
[0034] The core of the DAE module is to develop an efficient method to relocate the offset area to the correct position. In order to offset the negative impact of spatial misalignment on subsequent feature aggregation and feature decoding, the present invention uses an alignment unit to first learn the RGB-T feature pairs from the input. To the spatial deformation field A differentiable resampling operation, i.e., a deformation function, is then used to achieve end-to-end deformation calibration.
[0035] Specifically, S21 includes:
[0036] S21-1: Visible light characteristics and thermal infrared characteristics Input deformation estimator to get deformation field , the formula is: ; in, is the deformation field of the i-th level, is a deformation estimator based on STN.
[0037] S21-2: Transformation Field and thermal infrared characteristics Input the deformation function to obtain the calibrated thermal infrared characteristics , the formula is: ; in, is the calibrated thermal infrared feature of the i-th level, For the deformation function, you can call the integrated deformation function of the Pytorch library.
[0038] S21-3: The calibrated thermal infrared signature Visible light characteristics After splicing, it passes through a 3×3 convolution layer to obtain the calibration features , the formula is: ; in, represents the calibration feature of the i-th level, It is a 3×3 convolutional layer.
[0039] The alignment unit can be inserted into other modules as an independent component. In the feature decoder, AU is again used to smooth the deformation deviation in the RGB-T encoder features.
[0040] S22: Visible light features and thermal infrared characteristics After the spatial attention mechanism, the visible light spatial features are obtained and thermal infrared spatial characteristics , and then respectively with the calibration characteristics Multiply to get the visible light alignment feature and thermal infrared alignment features , the formula is: ; ; ; ;
[0041] Visible light characteristics and thermal infrared characteristics After 1×1 convolution layers respectively, the visible light convolution output features are obtained and thermal infrared convolution output features , the formula is: ; ; Aligning visible light to features , visible light characteristics And the visible light convolution output features Add together to get the visible light space correction feature , the formula is: ; Thermal infrared alignment features , thermal infrared characteristics And thermal infrared convolution output features Add together to get the thermal infrared spatial correction feature , the formula is: ; in, is the visible light spatial feature of the i-th level, is the thermal infrared spatial feature of the i-th level, is the Spatial Attention (SA) mechanism, is the element-wise product operation, is the visible light alignment feature of the i-th level, is the thermal infrared alignment feature of the i-th level, is a 1×1 convolutional layer, is the visible light convolution output feature of the i-th level, is the thermal infrared convolution output feature of the i-th level, and are the visible light space correction features and thermal infrared space correction features of the i-th level respectively.
[0042] The present invention adopts a symmetric structure to further explore the spatial position correlation of original features and enhance the spatial consistency of cross-modal features.
[0043] S23: Correction of visible light space features and thermal infrared spatial correction features After splicing, it passes through a 3×3 convolution layer and a Sigmoid activation function to obtain a weight tensor , the formula is: ; in, is the weight tensor of the i-th level, is the Sigmoid activation function.
[0044] S24: weight tensor Respectively with visible light characteristics and thermal infrared characteristics After multiplication, concatenation is performed and then a 3×3 convolution layer is performed to obtain the alignment enhancement feature. , the formula is: ; in, is the alignment enhancement feature of the i-th level.
[0045] The present invention combines the characteristics of visible light and thermal infrared modes, uses visible light features to perform spatial correction, and thus generates alignment enhancement features.
[0046] In S2, the DAE module uses an alignment unit including a spatial transformation network to calculate thermal infrared features. and visible light characteristics The deformation field between , effectively thermal infrared characteristics Correction is performed, which enhances the thermal infrared signature Visible light characteristics The spatial consistency of the visible light space can be enhanced by aligning the RGB-T features in the first stage. The DAE module integrates the visible light space correction features. and thermal infrared spatial correction features The characteristics of each branch, using visible light characteristics Thermal infrared characteristics Perform spatial adjustments to generate alignment enhancement features , which can be used to estimate the deformation in RGB-T features, thereby enhancing spatial consistency between RGB-T modalities. Therefore, the DAE module can alleviate the misalignment interference between modalities and obtain deformation-aware alignment capabilities.
[0047] Existing studies have attempted to capture single-scale cross-complementary information between RGB-T images, such as modal cross-connection or modal subtraction techniques. However, how to mine cross-modal complementary clues hidden in objects of different scales remains to be studied. To address this issue and further utilize the complementarity between RGB-T, the present invention introduces an effective complementary feature aggregation module to explore the feature changes of different modalities at different scales. The CFA module determines the channel weight of each modality based on the features of different scales, and can enhance the complementarity and integrity of RGB-T features by adjusting the channel relationship in the features of each modality.
[0048] The formula of the complementary feature aggregation module can be expressed as: ; in, is the cross-modal aggregation feature of the i-th level, It is a complementary feature aggregation module.
[0049] Specifically, refer to Figure 3 As shown, in S3, the alignment enhancement features of the same level are and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ,include: S31: Align Enhancement Features and visible light characteristics Input the ASPP layer separately to obtain the multi-scale context features of the aligned enhanced features Multi-scale contextual features of visible light features .
[0050] The present invention uses ASPP with expansion rates of 1, 2, and 4 to extract multi-scale context from RGB-T features , thereby enhancing the multi-level understanding of the object, the formula is expressed as: ; ; in, is the dilated spatial pyramid pooling layer, is the index of the number of layers of the dilated spatial pyramid pooling layer, is the k-th scale context feature of the visible light feature of the i-th level, is the k-th scale context feature of the alignment enhancement feature at the i-th level.
[0051] S32: Multi-scale contextual features that align enhanced features Multi-scale contextual features of visible light features Subtract the corresponding ones to obtain multi-scale difference features , after splicing, input 1×1 convolution layer to obtain the target difference feature , the formula is: ; ; in, is the target difference feature of the i-th level, is the k-th scale difference feature of the i-th level. The convolution layer is 1 × 1 with BN and ReLU. Connecting multi-scale difference cues.
[0052] S33: Aligning the multi-scale context features of enhanced features and multi-scale context features of visible light features Fusion is performed separately to obtain thermal infrared fusion features and visible light fusion features .
[0053] S34: To enhance the correlation between the difference clues and the existing features, the target difference feature Fusion features with thermal infrared After multiplication, the features are fused with visible light. Add them together and input them into the 1×1 convolution layer to obtain the visible light supplementary features , the formula is: ; The target difference feature Fusion features with visible light After multiplication, it is fused with thermal infrared features Add them together and input them into the 1×1 convolution layer to get the thermal infrared supplementary features , the formula is: ; in, is the visible light supplementary feature of the i-th level, is the thermal infrared supplementary feature of the i-th level, is the visible light fusion feature of the i-th level, is the thermal infrared fusion feature of the i-th level.
[0054] S35: Supplementing features with visible light and thermal infrared supplementary features Visible light weight value is obtained through gating operation , the formula is: ; in, is the visible light weight value of the i-th level, It is a gating operation used to extract multimodal global information; For linear mapping, this embodiment uses the Sigmoid function to project the dimension of the input feature to half the size of the original image; is the global average pooling operation.
[0055] Visible light weight value is directly related to the importance of RGB information. Similarly, we can use Describe the relative importance of thermal infrared information.
[0056] S36: Visible light weight value Complementary features with visible light After multiplication, fusion features with thermal infrared Add together to get the visible light polymerization characteristics , the formula is: ; Will The value and thermal infrared supplementary characteristics After multiplication, the features are fused with visible light Add together to get the thermal infrared polymerization characteristics , the formula is: ; in, and are the visible light aggregation features and thermal infrared aggregation features of the i-th level, which are used to enhance the independence of RGB-T information within the semantic channel.
[0057] By multiplying the RGB-T features element by element, the weight information is further consolidated.
[0058] S37: Polymerizing Visible Light Features and thermal infrared polymerization characteristics After splicing, it passes through a 1×1 convolution layer to obtain cross-modal aggregation features , the formula is: ; in, is the cross-modal aggregation feature of the i-th level.
[0059] The CFA module utilizes the atrous spatial pyramid pooling layer to promote robust cross-modal and cross-scale interactions and enhance modal complementarity, effectively integrating multimodal information across different scales and channels. The resulting cross-modal aggregation features play a key role in significantly reducing the differences between different modal information, and can provide comprehensive complementary RGB-T features for the subsequent decoding process, thereby enabling accurate prediction of semantic segmentation masks.
[0060] The decoding process of the present invention involves step-by-step prediction of semantic segmentation masks, with feature decoding blocks at each level simultaneously processing the decoded features of the previous level. and the visible light features of the current level , thermal infrared characteristics and cross-modal aggregation features , the formula is: ; in, and are the decoding features of the i-th level and the i+1-th level respectively, is the feature decoding block.
[0061] The decoding process is from high level to low level. The feature decoding block of the fifth level is based only on the visible light features of the current level. , thermal infrared characteristics and cross-modal aggregation features As input, the formula is expressed as .
[0062] Specifically, refer to Figure 4 As shown, in S4, the fifth level feature decoding block uses the visible light feature of the current level , thermal infrared characteristics and cross-modal aggregation features As input, get the decoding features of the fifth level ; The feature decoding block from the fourth to the first level uses the visible light feature of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ,include: S41: In the deformation-aware alignment branch, the visible light features of the current level are aligned and thermal infrared characteristics After splicing, input the deformation estimator to obtain the decoded deformation field ; will decode the deformation field Thermal infrared characteristics Then pass through the deformation function together to obtain the decoded and corrected thermal infrared features ; Decode and correct thermal infrared characteristics Visible light characteristics After splicing, it passes through a 3×3 convolution layer to obtain the first branch feature , the formula is: ; ; ; in, is the decoded deformation field of the i-th level, Corrected thermal infrared signature for decoding of level i, is the first branch feature of the i-th level.
[0063] S42: In the supplementary aggregation branch, cross-modal aggregation features are and the decoding features of the previous layer After splicing, it passes through the residual channel attention block (RCA), 3×3 convolution layer and 3×3 dilated convolution to obtain the second branch feature , the formula is: ; in, is the second branch feature of the i-th level, is the residual channel attention block, is a 3×3 dilated convolution with a dilation factor of 2.
[0064] Concatenating cross-modal aggregate features and the decoding features of the previous layer Before Magnify it by 2 times Same size.
[0065] Residual channel attention block structure reference Figure 5 As shown in Figure 1, the input features of the residual channel attention block pass through the 3×3 convolution layer and the channel attention mechanism (CA) in sequence, and are added to the input features of the residual channel attention block to obtain the output features of the residual channel attention block.
[0066] For the fifth-level feature decoding block, only the cross-modal aggregation features of the current level are used in the supplementary aggregation branch. As input, the formula is .
[0067] S43: The first branch feature and the second branch features After splicing, it passes through the residual channel attention block and the 3×3 convolution layer in turn to obtain the decoding features of the current level , the formula is: ; in, is the decoding feature of the i-th level.
[0068] The decoder constructed with feature decoding blocks in the present invention combines deep and shallow features, uses residual channel attention blocks to capture the semantic correlation between channels, and then gradually predicts the segmentation results in a coarse-to-fine manner.
[0069] Preferably, refer to Figure 6 As shown in the figure, in the decoding stage, the present invention also develops two auxiliary tasks: Class-agnostic Saliency Prediction (SP) and Class-aware Edge Generation (EG), to construct the class-agnostic saliency loss and class-aware semantic margin loss , to further provide the model with clues of object location and boundaries and improve the semantic segmentation performance of non-aligned RGB-T images.
[0070] When the method of the present invention trains the segmentation model composed of the feature encoder, deformation-aware alignment enhancement module, complementary feature aggregation module, feature decoding block and output layer, the total loss function Including semantic segmentation loss , class-independent saliency loss and class-aware semantic margin loss , the formula is: ; in, is the total loss function, is the semantic segmentation loss, is the class-independent saliency loss, is the class-aware semantic margin loss, and Control and The contribution coefficient is set through ablation experiment analysis.
[0071] Specifically, the semantic segmentation loss To predict the segmentation results and ground truth segmentation annotations The cross entropy loss between is expressed as: ; in, is the cross entropy loss, To predict the segmentation results, is the true segmentation annotation, that is, the true label.
[0072] Shallow features are rich in local detail information, while high-level features tend to contain deeper global information. Therefore, high-level cues are particularly good at effectively representing object locations. In the class-independent saliency prediction layer, the present invention utilizes high-level features to Work together to jointly identify the object position and improve the intra-class density of RGB-T features to obtain the predicted position segmentation mask , and then use the predicted position segmentation mask Constructing class-independent saliency loss .
[0073] The class-independent saliency prediction task is expressed as: ; in, is the class-independent saliency prediction layer, To predict the location segmentation mask, and These are the decoding features of the 4th and 5th levels respectively.
[0074] Specifically, the class-independent saliency loss include: The decoded features of the fifth level and the decoding features of the fourth level After passing through the 1×1 convolution layer and the 3×3 convolution layer respectively, the corresponding output features are obtained and , the formula is: ; Will Upsampling, that is Afterwards, with Splicing, and then passing through the residual channel attention block, 3×3 convolution layer and upsampling layer in sequence to obtain the predicted position segmentation mask , the formula is: ; in, is n upsampling operations; Based on the real segmentation annotation Get the true position segmentation mask ; The true position segmentation mask is a mask map that does not include the object category; Segmentation mask at predicted position and ground-truth segmentation mask The two-way cross entropy loss between the two classes calculates the class-independent significance loss , the formula is: ; in, is the two-way cross entropy loss.
[0075] The class-independent saliency prediction task can provide position clues to the decoder, help improve intra-class compactness, and promote multimodal spatial alignment, thereby improving the accuracy of object localization in the semantic segmentation task of non-aligned RGB-T images.
[0076] Since pixels close to object boundaries are prone to mis-segmentation, it is crucial to leverage boundary cues to improve segmentation accuracy in uncertain regions. Objects of different categories typically exhibit different sizes and shapes. Compared to traditional edge-guided detection and segmentation research, this paper focuses on drawing edges for different semantic objects.
[0077] In the class-aware edge generation layer, the present invention first labels the real segmentation Get the real semantic edge The semantic edge includes both object category information and edge information, providing semantic dependencies and boundary clues. Because boundary clues contain local, fine-grained details, they can further enhance the inter-class separability of decoded features, providing more customized support for semantic segmentation.
[0078] The class-aware edge generation task is used to mine low-level features , to obtain the predicted semantic edge , and then use the predicted semantic edge Constructing class-aware semantic margin loss .
[0079] The class-aware edge generation task is formulated as: ; in, To predict semantic edges, Generates a layer for class-aware edges.
[0080] Specifically, the class-aware semantic edge loss include: The decoded features of the third level, the second level, and the first level 、 and After passing through the 1×1 convolution layer and the 3×3 convolution layer respectively, the corresponding output features are obtained 、 and , the formula is: ; Will 、 and After splicing, it passes through the residual channel attention block, 3×3 convolution layer and upsampling layer in sequence to obtain the predicted semantic edge , the formula is: ; Based on the real segmentation annotation Get the real semantic edge ; The true semantic edge includes the category information and edge information of the object; To predict semantic edges and true semantic edge The cross entropy loss between them calculates the class-aware semantic edge loss , the formula is: ; in, is the cross entropy loss.
[0081] The class-aware edge generation task can improve the inter-class separability of RGB-T decoding features, promote the semantic alignment of decoding features, and generate segmentation results with clear boundaries.
[0082] The total loss function constructed by the present invention By introducing class-independent saliency loss and class-aware semantic margin loss , which can effectively guide the segmentation model to pay more attention to the location of the object and the boundary information of the object, help to alleviate the interference caused by misalignment, and improve the accuracy of the predicted semantic segmentation results.
[0083] In summary, the method for semantic segmentation of non-aligned visible light and thermal infrared images described in the present invention first uses a deformation-aware alignment enhancement module to process non-aligned visible light features and thermal infrared features. The module can estimate the deformation field of the thermal infrared features and combine the deformation field with the visible light features to achieve spatial alignment, thereby improving the spatial consistency between the visible light features and the thermal infrared features. The complementary feature aggregation module is then used to capture the unique clues in the two modalities of alignment enhancement features and visible light features, promote cross-modal complementary interaction, and provide comprehensive cross-modal aggregation features for the decoder. The present invention realizes direct semantic segmentation of non-aligned visible light and thermal infrared images, can effectively reduce the interference of non-aligned images on the segmentation task, and improves the accuracy of non-aligned visible light and thermal infrared image segmentation. The present invention can break through the limitations of image alignment preprocessing and significantly improve the applicability and efficiency of visible light and thermal infrared image semantic segmentation technology in actual scenes.
[0084] This example uses Accuracy (Acc) and Intersection over Union (IoU) to quantitatively evaluate the semantic segmentation results of the segmentation model. The formula is as follows: ; ; in, and are the total accuracy and total intersection-over-union ratio, respectively. is the number of categories excluding background categories, 、 and are the true positives, false positives, and false negatives of class p, respectively.
[0085] This example uses a computer equipped with a 3.0 GHz CPU, 128 GB of RAM, and four NVIDIA GeForce RTX3090 graphics cards. The segmentation model was trained and tested on the U-MFNet (unaligned RGB-T SS) dataset. The U-MFNet dataset is obtained by applying an affine transformation to the thermal infrared images in the MFNet dataset while keeping the RGB images unchanged. Figure 7 is an example of a deformed thermal infrared image obtained according to different deformation intensities, where Figure 7 Column (a) in is the visible light image and the true label, Figure 7 Column (b) is the original thermal infrared image. Figure 7 Column (c) shows the deformed thermal infrared image with a deformation intensity of 5. Figure 7 Column (d) shows the deformed thermal infrared image with a deformation intensity of 10. Figure 7 Column (e) is the deformed thermal infrared image obtained when the deformation intensity is 20. Figure 7 Column (f) in the figure shows the deformed thermal infrared image with a deformation strength of 50. The mini-batch size is 4 and the initial learning rate of the network is 1×10 −4 DAC-Net is implemented in PyTorch using the Adam optimizer. During training, this example uses data augmentation, including random flipping and cropping. During testing, this example directly inputs raw RGB-T images without any post-processing.
[0086] On the U-MFNet dataset, the proposed DAC-Net is compared with eight recent SOTA (state-of-the-art) RGB-T SS methods, including MFNet, RTFNet, FEANet, GMNet, EGFNet, MFTNet, LASNet and MDRNet.
[0087] Tables 1 and 2 show the quantitative performance of the proposed DAC-Net and other methods on 8 semantic categories and overall, respectively. The best result in each column is shown in bold, and the second result is underlined.
[0088] Table 1. Accuracy comparison of various methods on the U-MFNet dataset
[0089] Table 2. Intersection-over-Union comparison of various methods on the U-MFNet dataset
[0090] As can be seen from Tables 1 and 2, the DAC-Net of the present invention achieves the best overall performance. Compared with the second-best LASNet, the DAC-Net of the present invention improves by 1.3% in mAc and 2.6% in mIoU. Methods such as RTFNet, FEANet, GMNet, and LASNet do not consider deformation recovery between modalities, but directly fuse RGB-T features, so their performance is poor. Specifically, the DAC-Net of the present invention ranks first in the segmentation of 5 categories and second in 1 category. In mAc, the DAC-Net of the present invention improves by 1.7%, 7.0%, 8.9%, and 4.1% in the car, person, bicycle, and curve categories respectively over the state-of-the-art method MDRNet. In terms of mIoU, the DAC-Net of the present invention improves by 1.4% over the previous state-of-the-art method MFTNet.
[0091] In addition, this embodiment also selects 12 challenging samples, which cover different lighting conditions, deformation levels, scenes and object sizes, and can provide a comprehensive evaluation. Figure 8 shown. Figure 8 This is a comparison of the segmentation results of the DAC-Net of the present invention and eight alignment methods on the U-MFNet dataset. Figure 8 Column (a) is the visible light image. Figure 8 Column (b) is a thermal infrared image. Figure 8 The (c) column in the figure is the segmentation result of MFNet. Figure 8 The (d) column in the figure is the segmentation result of RTFNet. Figure 8Column (e) is the segmentation result map of FEANet. Figure 8 The (f) column in the figure is the segmentation result of GMNet. Figure 8 The (g) column in the figure is the segmentation result of EGFNet. Figure 8 The (h) column is the segmentation result map of MFTNet. Figure 8 The (i) column in the figure is the segmentation result map of LASNet. Figure 8 The (j) column in the figure is the segmentation result of MDRNet. Figure 8 The (k) column in the figure is the segmentation result of the DAC-Net of the present invention. Figure 8 The (l) column in is the true label.
[0092] Compared with other methods, the results of the proposed DAC-Net are closely consistent with the true labels. Figure 8 The first three rows of
[15] depict a typical urban street scene, mainly featuring cars and pedestrians. The DAC-Net of our invention accurately predicts large-scale cars, small-scale pedestrians and bicycles with fewer errors and omissions. Figure 8 In the images with significant deformation in the 2nd, 4th, 6th, 11th and 12th rows of
[15] , our DAC-Net enhances modality consistency in the deformation-aware alignment module, producing accurate semantic segmentation results. Figure 8 The lower half of the figure shows various night scenes. The DAC-Net of our invention effectively aggregates modality complementarity by calculating the difference of RGB-T features, thus achieving competitive results.
[0093] Based on the above-mentioned non-aligned visible light-thermal infrared image semantic segmentation method, the present invention also provides a non-aligned visible light-thermal infrared image semantic segmentation system, comprising: The encoding module is used to pass the visible light image and thermal infrared image through their corresponding feature encoders to obtain five levels of visible light features. and thermal infrared characteristics , is the hierarchical index; Feature alignment module, used to align visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ; Feature aggregation module, used to align enhanced features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ; Decoding module, used for the fifth level feature decoding block to decode the visible light features of the current level , thermal infrared characteristics and cross-modal aggregation features As input, get the decoding features of the fifth level ; The feature decoding block from the fourth to the first level uses the visible light feature of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ; Output module, used to decode the features of the first level The predicted segmentation result is obtained through the output layer .
[0094] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.
Claims
1. A semantic segmentation method for non-aligned visible light-thermal infrared images, characterized in that: include: The visible light image and thermal infrared image are passed through their corresponding feature encoders to obtain five levels of visible light features. and thermal infrared characteristics , is the hierarchical index; The visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ; Align enhancement features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ; The fifth level feature decoding block uses the visible light features of the current level , thermal infrared characteristics and cross-modal aggregation features As input, get the decoding features of the fifth level ; The feature decoding block from the fourth to the first level uses the visible light feature of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ; The decoded features of the first level The predicted segmentation result is obtained through the output layer .
2. The semantic segmentation method of non-aligned visible light-thermal infrared image according to claim 1, characterized in that: The visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ,include: Visible light features and thermal infrared characteristics Enter the alignment unit together to get the calibration feature ; Visible light features and thermal infrared characteristics After the spatial attention mechanism, the visible light spatial features are obtained and thermal infrared spatial characteristics , and then respectively with the calibration characteristics Multiply to get the visible light alignment feature and thermal infrared alignment features ; Aligning visible light to features , visible light characteristics And the output features of the visible light feature after the 1×1 convolution layer Add together to get the visible light space correction feature ; Thermal infrared alignment features , thermal infrared characteristics And the output features of thermal infrared features after 1×1 convolution layer Add together to get the thermal infrared spatial correction feature ; Visible light space correction features and thermal infrared spatial correction features After splicing, it passes through a 3×3 convolution layer and a Sigmoid activation function to obtain a weight tensor ; The weight tensor Respectively with visible light characteristics and thermal infrared characteristics After multiplication, concatenation is performed and then a 3×3 convolution layer is performed to obtain the alignment enhancement feature. .
3. The semantic segmentation method of non-aligned visible light-thermal infrared images according to claim 2, characterized in that: Visible light features and thermal infrared characteristics Enter the alignment unit together to get the calibration feature ,include: Visible light features and thermal infrared characteristics Input deformation estimator to get deformation field ; The deformation field and thermal infrared characteristics Input the deformation function to obtain the calibrated thermal infrared characteristics ; The calibrated thermal infrared signature Visible light characteristics After splicing, it passes through a 3×3 convolution layer to obtain the calibration features .
4. The semantic segmentation method of non-aligned visible light-thermal infrared images according to claim 1, characterized in that: Align enhancement features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ,include: Align Enhancement Features and visible light characteristics Input the ASPP layer separately to obtain the multi-scale context features of the aligned enhanced features Multi-scale contextual features of visible light features , k=1,2,3 is the layer index of ASPP layer; Multi-scale contextual features that will align enhanced features Multi-scale contextual features of visible light features Subtract the corresponding ones to obtain multi-scale difference features , after splicing, input 1×1 convolution layer to obtain the target difference feature ; Multi-scale contextual features that will align enhanced features Multi-scale contextual features of visible light features Fusion is performed separately to obtain thermal infrared fusion features and visible light fusion features ; The target difference feature Fusion features with thermal infrared After multiplication, the features are fused with visible light. Add them together and input them into the 1×1 convolution layer to obtain the visible light supplementary features ; The target difference feature Fusion features with visible light Multiply and fuse features with thermal infrared Add them together and input them into the 1×1 convolution layer to get the thermal infrared supplementary features ; Supplementary features of visible light and thermal infrared supplementary features Visible light weight value is obtained through gating operation ; The visible light weight value Complementary features with visible light After multiplication, fusion features with thermal infrared Add together to get the visible light polymerization characteristics ; Will The value of thermal infrared supplementary characteristics After multiplication, the features are fused with visible light Add together to get the thermal infrared polymerization characteristics ; Visible light polymerization features and thermal infrared polymerization characteristics After splicing, it passes through a 1×1 convolution layer to obtain cross-modal aggregation features .
5. The semantic segmentation method of non-aligned visible light-thermal infrared images according to claim 1, characterized in that: The feature decoding block from the fourth to the first level uses the visible light features of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ,include: The visible light features of the current level and thermal infrared characteristics After splicing, the input deformation estimator outputs the decoded deformation field Thermal infrared characteristics Then pass through the deformation function together, and the output decoded and corrected thermal infrared features Visible light characteristics After splicing, it passes through a 3×3 convolution layer to obtain the first branch feature ; Aggregate features across modalities and the decoding features of the previous layer After splicing, it passes through the residual channel attention block, 3×3 convolution layer and 3×3 dilated convolution in sequence to obtain the second branch feature ; The first branch feature and the second branch features After splicing, it passes through the residual channel attention block and the 3×3 convolution layer in turn to obtain the decoding features of the current level .
6. The method for semantic segmentation of non-aligned visible light and thermal infrared images according to claim 1, characterized in that: When training the segmentation model consisting of the feature encoder, deformation-aware alignment enhancement module, complementary feature aggregation module, feature decoding block and output layer, the total loss function is Including semantic segmentation loss , class-independent saliency loss and class-aware semantic margin loss .
7. The method for semantic segmentation of non-aligned visible light and thermal infrared images according to claim 6, characterized in that: The class-independent significance loss include: The decoded features of the fifth level and the decoding features of the fourth level After passing through the 1×1 convolution layer and the 3×3 convolution layer respectively, the corresponding output features are obtained and ; Will After upsampling Splicing, sequentially passing through the residual channel attention block, 3×3 convolution layer and upsampling layer, to obtain the predicted position segmentation mask ; Based on the real segmentation annotation Get the true position segmentation mask ; Segmentation mask at predicted position and ground-truth segmentation mask The two-way cross entropy loss between the two classes calculates the class-independent significance loss .
8. The method for semantic segmentation of non-aligned visible light and thermal infrared images according to claim 6, characterized in that: The class-aware semantic margin loss include: The decoded features of the third level, the second level, and the first level 、 and After passing through the 1×1 convolution layer and the 3×3 convolution layer respectively, the corresponding output features are obtained 、 and ; Will 、 and After splicing, it passes through the residual channel attention block, 3×3 convolution layer and upsampling layer in sequence to obtain the predicted semantic edge ; Get the real semantic edge based on the real segmentation annotation ; To predict semantic edges and true semantic edge The cross entropy loss between them calculates the class-aware semantic edge loss .
9. The method for semantic segmentation of non-aligned visible light and thermal infrared images according to claim 6, characterized in that: The semantic segmentation loss To predict the segmentation results and ground-truth segmentation annotations The cross entropy loss between .
10. A non-aligned visible light-thermal infrared image semantic segmentation system, characterized by: include: The encoding module is used to pass the visible light image and thermal infrared image through their corresponding feature encoders to obtain five levels of visible light features. and thermal infrared characteristics , is the hierarchical index; Feature alignment module, used to align visible light features at the same level and thermal infrared characteristics Input the deformation-aware alignment enhancement module of this level to obtain the alignment enhancement features of each level ; Feature aggregation module, used to align enhanced features at the same level and visible light characteristics Input the complementary feature aggregation module of this level to obtain cross-modal aggregation features of each level ; Decoding module, used for the fifth level feature decoding block to decode the visible light features of the current level , thermal infrared characteristics and cross-modal aggregation features As input, get the decoding features of the fifth level ; The feature decoding block from the fourth to the first level uses the visible light feature of the current level , thermal infrared characteristics , cross-modal aggregation features and the decoding features of the previous layer As input, get the decoding features of the current level ; Output module, used to decode the features of the first level The predicted segmentation result is obtained through the output layer .
Citation Information
Patent Citations
Power equipment semantic segmentation method based on visible light and infrared image feature fusion
CN118196405A
Cited By
Small target detection method based on visible light and thermal infrared bidirectional supervised alignment
CN122049699A
A small target detection method based on visible light and thermal infrared bidirectional supervision alignment
CN122049699B
Cascade infrared and visible image registration method and apparatus for autonomous driving
CN122347602A