Lightweight RGB-T semantic segmentation method based on multi-modal feature fusion

By adopting the lightweight RGB-T semantic segmentation method with multimodal feature fusion in the semantic segmentation model, the multi-scale feature fusion and hierarchical feature aggregation module are used to solve the problem of degradation in semantic segmentation performance in low light and harsh environments, and the segmentation effect of high precision and low computational complexity is achieved.

CN120219752APending Publication Date: 2025-06-27HANGZHOU DIANZI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510381218.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

Smart Images

  • Figure CN120219752A_ABST
    Figure CN120219752A_ABST
Patent Text Reader

Abstract

The invention discloses a lightweight RGB-T semantic segmentation method based on multi-modal feature fusion, and the method comprises the steps: firstly selecting a pair of visible light RGB and thermal infrared TIR images from an RGB-T data set, and taking the images as an input training set; secondly, constructing a lightweight RGB-T semantic segmentation model based on multi-modal feature fusion, wherein the lightweight RGB-T semantic segmentation model comprises an encoder part, a multi-scale feature fusion part, a hierarchical feature aggregation part and a decoder part; and finally, inputting the RGB-T image pair of the training set into the model for training, and generating a semantic segmentation result. According to the method, effective fusion from fine-grained features to coarse-grained features is ensured, the features can be effectively represented in a complex segmentation scene, and the method has relatively high segmentation precision and calculation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a lightweight semantic segmentation method based on multi-modal feature fusion, which is applicable to the semantic segmentation task of RGB-T (visible light and thermal infrared) images. Background Art

[0002] Semantic segmentation is an important task in computer vision, aiming to divide an image into different regions with specific semantics and accurately assign class labels to each region at the pixel level. This technology is crucial for accurate image understanding and is widely used in fields such as obstacle detection in autonomous driving, land cover classification in remote sensing, and lesion recognition in medical imaging. Traditional semantic segmentation technologies perform well under the conditions of high-quality RGB images, sufficient lighting environments, and rich visual information. However, in low-light conditions or environments with obstructed visibility, the performance of semantic segmentation models mainly relying on RGB images often significantly degrades.

[0003] To address this challenge, researchers have begun to explore the use of multi-modal data to enhance the robustness of semantic segmentation models. By combining visible light images (RGB) and thermal infrared images (TIR), RGB-T based semantic segmentation can capture the temperature information of objects in low-light environments, thus obtaining more reliable segmentation results at night or under adverse weather conditions. This multi-modal approach has been proven to effectively improve the segmentation accuracy and robustness in fields such as wildlife monitoring and security systems.

[0004] However, integrating the information of RGB and thermal infrared images into an efficient segmentation network still faces many challenges. First, due to the information differences between different modalities, the segmentation model needs to effectively combine and utilize data from different modalities. Second, features at different levels have different characteristics: high-level features contain more semantic information, while low-level features are rich in spatial details. How to effectively integrate these multi-level features so that it can maintain the global perspective of objects and capture complex spatial details is a key issue. Finally, processing multi-modal data usually requires a large amount of computing resources, which may be a major burden for real-time applications (such as autonomous driving).

[0005] Although many existing studies prioritize improving segmentation performance, they often ignore the parameter scale and computational efficiency, thus limiting the applicability of these models in resource-constrained devices. In addition, although some models introduce lightweight architectures to balance segmentation accuracy and computational efficiency, they often encounter difficulties in capturing small or occluded objects and maintaining detail accuracy when dealing with complex scenes. These challenges highlight the necessity of further optimization in cross-modal complementarity and computational efficiency. Summary of the Invention

[0006] In view of the deficiencies in the existing technology, the present invention proposes a lightweight semantic segmentation method based on multi-modal feature fusion. This method adopts an encoder-fusion-decoder network architecture, which mainly consists of an encoder part, a multi-scale feature fusion part, a hierarchical feature aggregation part, and a decoder part. By training the neural network, optimal parameters are obtained to achieve efficient semantic segmentation of RGB-T (visible light and thermal infrared) images.

[0007] To solve the above technical problems, the technical solution of the present invention is as follows:

[0008] A lightweight RGB-T semantic segmentation method based on multi-modal feature fusion, comprising the following steps:

[0009] S1. Select paired visible light RGB and thermal infrared TIR images from the RGB-T dataset as the input training set.

[0010] S2. Construct a lightweight RGB-T semantic segmentation model based on multi-modal feature fusion. The model includes an encoder part, a multi-scale feature fusion part, a hierarchical feature aggregation part, and a decoder part. The encoder part respectively constructs an RGB encoder and a TIR encoder through an encoder branch based on MobileNetV2 to perform multi-level feature extraction on RGB and TIR images respectively to obtain rich modal information; the multi-scale feature fusion part is used to perform multi-scale feature fusion on the middle and high-level features extracted by the encoder to enhance the complementarity of cross-modal information; the hierarchical feature aggregation part is used to effectively aggregate the second-layer RGB encoder features and the fused multi-scale features to enhance the feature expression ability between modalities; the decoder part decodes based on the aggregated features and the low-level features extracted by the first-layer RGB encoder to restore fine-grained spatial information, and finally generates a high-precision semantic segmentation result.

[0011] S3. Input the RGB-T image pairs of the training set into the model for training, specifically including:

[0012] Input the RGB image and the TIR image into the encoder module respectively, and extract 5-level encoder features with different resolutions through five layers of encoders;

[0013] In order to effectively bridge the modal differences between RGB features and thermal infrared (TIR) features, the present invention deploys a multi-scale feature fusion module on the middle and high-level multi-modal features to fully exploit cross-modal complementary information. Specifically, the third-layer RGB encoder features and the third-layer TIR encoder features are fused through a multi-scale feature fusion (MSEF) module to generate third-layer fused features. The fourth-layer RGB encoder features and the fourth-layer TIR encoder features are fused through a multi-scale feature fusion (MSEF) module to generate the fourth-layer fused features. The fifth-layer RGB encoder features and the fifth-layer TIR encoder features are fused through a multi-scale feature fusion (MSEF-H) module to generate the fifth-layer fused features. In this way, the multi-scale feature fusion module can fuse RGB and TIR features layer by layer to generate multi-level fused features. The MSEF module and the MSEF-H module have the same structure, thus effectively enhancing the complementarity of multi-modal features and the segmentation accuracy.

[0014] To aggregate the multi-level fused features, the present invention designs a hierarchical feature aggregation module, which can effectively integrate features at different levels. Specifically, the fused features and are integrated into an aggregated feature by the first hierarchical feature aggregation module HFE-1. The fused features and the second-layer RGB encoder features are integrated by the second hierarchical feature aggregation module HFE-2 to generate an aggregated feature

[0015] In this way, the hierarchical feature aggregation module can effectively combine low-level spatial details and high-level semantic information to improve the segmentation effect in complex scenes.

[0015] To generate the final semantic segmentation result, the present invention deploys three convolutional blocks with the same structure in the decoder part. Each convolutional block consists of a dropout layer, two depthwise separable convolutions with dilation rates of 3 and 1 respectively, and an upsampling operation. First, the first convolutional block processes the aggregated feature output by the hierarchical feature aggregation module to generate a decoded feature Then, the second convolutional block adds element-wise to the aggregated feature and then performs convolution processing to generate a feature Finally, the third convolutional block adds element-wise to the first-layer RGB encoder feature and then performs convolution processing, and finally refines the features through a depthwise separable convolution with a dilation rate of 1 and a 1×1 convolution to generate the final prediction map S. Through this step-by-step processing method, the present invention can effectively integrate multi-level features to generate high-quality semantic segmentation results.

[0016] Specifically: First, the RGB-T image pair is input into the model for training, and multi-level encoded features (RGB features and TIR features ). Among them, the low-level features (the 1st and 2nd level features) contain rich spatial detail information, while the middle and high-level features (the 3rd to 5th level features) contain more global semantic information. Then, in the fusion stage, the present invention focuses on the middle and high-level encoder features (i.e., the 3rd to 5th layer features), and adopts a multi-scale feature fusion (MSEF) module to enhance the complementarity of RGB and TIR features. This module extracts cross-modal features through multi-scale depthwise separable convolution (DSConv) and gradually fuses information at different scales to enhance the network's adaptability to complex environments. Specifically in the fusion stage, the encoded RGB features and TIR features are respectively divided into four sub-feature groups and Each subgroup corresponds to a feature extraction path at a different scale. In the first branch, depthwise separable convolution with a dilation rate of 1 (DSConv-1) is applied to and respectively to generate preliminary features and Subsequently, in the second, third, and fourth branches, DSConv with dilation rates of 2, 3, and 5 are respectively adopted, and the outputs of the previous stage are gradually fused. For example, the second group of features is obtained from , and so on, to obtain Through this strategy, the low-dilation rate branch is responsible for extracting local details, while the high-dilation rate branch can capture a larger range of context information.

[0017] After completing multi-scale feature extraction, the MSEF module adopts a two-stage fusion strategy. The first stage: At each scale, the preliminary fusion features of RGB and TIR and are concatenated and processed through 1×1 convolution, batch normalization (BN), and ReLU activation function to generate preliminary multi-scale fusion features The second stage: The preliminary fusion features at four scales and are concatenated and globally fused through 1×1 convolution, batch normalization (BN), and ReLU activation function to generate the final fusion features Finally, the original RGB encoded features are added to the fusion features through a residual connection to further enhance the expression ability of the features. Through the two-stage fusion strategy, the MSEF module can effectively combine local details and global context information, enhance the complementarity of multi-modal features, and enhance the network's segmentation ability for complex scenes.

[0018] Secondly, in the decoding stage, considering that the quality of TIR images is usually lower than that of RGB images, only the low-level RGB features (i.e., and ) and the fused multi-modal features are decoded layer by layer, excluding the TIR features of the first and second layers (i.e., and ). During this process, this study uses two hierarchical feature aggregation (HFE) modules to integrate multi-level features to enhance the expressive power of the decoder. Specifically, the HFE-1 module is mainly used to integrate high-level fused features and First, these features are dimensionally standardized and feature-enhanced through a fully connected layer to ensure seamless fusion of features at different levels. Then, the features are adjusted to the same resolution through an upsampling operation and concatenated. Subsequently, the concatenated features are processed by a 1×1 convolution, batch normalization (BN), and ReLU activation function to generate a unified aggregated feature This process can effectively integrate high-level semantic information and enhance the network's understanding of the global context. The HFE-2 module focuses on integrating middle-level and low-level features, and the inputs include the RGB encoded features of the second layer Middle-level fusion and high-level fused features Through a fully connected layer and an upsampling operation, the HFE-2 module adjusts the dimensions and concatenates these features, and then processes them through a 1×1 convolution, batch normalization (BN), and ReLU activation function to generate an aggregated feature This module can effectively combine low-level spatial details and mid-high-level semantic information to improve the network's segmentation ability for object boundaries and small targets. Through these two hierarchical feature aggregation modules, the present invention can effectively integrate multi-level features to generate high-quality semantic segmentation results.

[0019] To generate the final semantic segmentation result, the present invention introduces three convolution blocks with the same structure in the decoder part to gradually fuse feature information at different levels and achieve high-quality segmentation results. Each convolution block consists of a dropout layer, two depthwise separable convolutions (DSConv) with dilation rates of 3 and 1 respectively, and an upsampling operation. This structure can effectively extract local and global information, reduce computational complexity, and improve model efficiency. Specifically, first, the decoding process starts from the aggregated feature output by the hierarchical feature aggregation module, and generates a preliminary decoded feature after being processed by the first convolution block. Then, the second convolution block adds element-wise to the aggregated feature and performs convolution to obtain the second-level decoded feature Next, the third convolutional block will perform feature fusion with the features of the first-layer RGB encoder and generate the third-level decoding features after convolutional processing. After obtaining the final decoding features, further feature refinement is performed through a depthwise separable convolution (DSConv-1) with a dilation rate of 1 and a 1×1 convolution to enhance the accuracy of the final prediction and the clarity of the boundaries. After this series of processes, the prediction map S is finally obtained. This step-by-step fusion strategy ensures that the decoding process can fully combine high-level semantic information and low-level spatial detail information, enabling the segmentation result to have both rich context information and precise boundary details.

[0020] The present invention has the following characteristics and beneficial effects:

[0021] This study proposes a novel lightweight RGB-T semantic segmentation network to address the challenge of multi-modal feature integration. The proposed solution not only improves the segmentation performance but also maintains the computational efficiency. The proposed method utilizes two key modules: the multi-scale feature fusion module and the hierarchical feature aggregation module. The multi-scale feature fusion module plays a key role in effectively extracting and fusing multi-level features from the RGB and TIR modalities. This module extracts features at different levels using depthwise separable convolutions and integrates them through parallel concatenation operations, thus ensuring the effective fusion of features from fine-grained to coarse-grained. The hierarchical feature aggregation module integrates features at different levels through the use of fully connected layers and concatenation operations, further enhancing the feature representation ability. This module combines low-level spatial details with high-level semantic information, enabling it to effectively represent features in complex segmentation scenarios. Extensive experiments conducted on the MFNet and PST900 datasets show that the present invention is superior to existing state-of-the-art methods in terms of both segmentation accuracy and computational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. The following drawings are only some embodiments of the present invention:

[0023] Figure 1 is the overall network block diagram of the present invention;

[0024] Figure 2 is the multi-scale feature fusion module;

[0025] Figure 3 is the hierarchical feature aggregation module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention provides a lightweight RGB-T semantic segmentation method based on multi-modal feature fusion, as Figure 1 shown, which includes the following steps:

[0027] Obtain the RGB-T image pairs for training input. The present invention mainly selects the MFNet dataset as the main dataset for model training, among which 784 pairs (410 pairs during the day and 374 pairs at night) are used as the training set for training, 392 pairs (205 pairs during the day and 187 pairs at night) are used as the validation set to detect the effect of the trained model and find the optimal model parameter weights, and the remaining 393 pairs (205 pairs during the day and 187 pairs at night) are used as the test set. Among them, the first 600 videos are used as the training set for training.

[0028] As Figure 1 shown, the present invention adopts an encoding-fusion-decoding network architecture. First, paired RGB and TIR images in the training dataset are input into the encoder part of the model. The encoder module is constructed based on MobileNetV2, and through multiple convolutional layers and downsampling operations, low-level, middle-level, and high-level features are gradually extracted. Specifically, the RGB image and the TIR image are input into the encoder, and then multi-level features (RGB features and TIR features ) are extracted through the encoding branches of RGB and TIR. Among them, H and W respectively represent the height and width of the image. i represents the feature level, where the number of channels c i is set to {16, 24, 32, 96, 320}.

[0029] As is well known, Atrous Spatial Pyramid Pooling (ASPP) is widely used in semantic segmentation tasks. ASPP captures multi-scale context information by applying convolutional operations in parallel on multiple atrous convolutional layers with different dilation rates. Inspired by this, this paper deploys a Multi-Scale Feature Fusion (MSEF) module on the middle-level and high-level encoder features. As Figure 2 shown, the MSEF module adopts a parallel cascade structure and obtains multi-scale features through dilated depthwise separable convolution (DSConv). Given the RGB encoded feature of the i-th layer with the number of channels c, it is first divided into four groups, and each group contains c / 4 channels. The overall process is as follows:

[0030]

[0031] Among them, Split represents the channel splitting operation. This channel-based partitioning method distributes the input features into different convolutional paths, thereby achieving information separation and aggregation at the feature level. This approach can significantly reduce the computational complexity while maintaining the integrity of feature extraction.

[0032] Specifically, in the first group of features a depthwise separable convolution with a dilation rate of 1 (i.e., DSConv-1) is applied to generate the output feature Subsequently, the second group of features and after being added together, a depthwise separable convolution with a dilation rate of 2 (i.e., DSConv-2) is applied to generate the output feature For the third group of features and after being added together, a depthwise separable convolution with a dilation rate of 3 (i.e., DSConv-3) is applied to generate the output feature Finally, for the fourth group of features and after being added together, a depthwise separable convolution with a dilation rate of 5 (i.e., DSConv-4) is applied to obtain the output feature Through the depthwise separable convolution operations with dilation rates of 1, 2, 3, and 5, the receptive field of each scale can be extended to 3, 7, 13, and 23 respectively. These different dilation rates enable the network to capture both fine-grained details and sufficient context information. Therefore, the MESF module can effectively fuse RGB and TIR features, thereby providing an effective representation of the target object in complex scenes. Similarly, for the TIR encoded features the same operations are performed to obtain four-scale features and The whole process is defined as:

[0033]

[0034] represents the dilated depthwise separable convolution (DSConv) for the j-th group of features, represents the element-wise multiplication operation.

[0035] Then, in order to effectively fuse the RGB features and the TIR features a two-stage fusion strategy is adopted. Specifically, the first stage occurs within each group of features (i.e., each scale). At this stage, the RGB features and the TIR features of each scale are concatenated, and then processed through a 1×1 convolutional layer, a batch normalization (BN) layer, and a ReLU activation function, as shown in Figure 2as shown by “CC” in. This way can generate preliminary multi-scale fusion features (i.e., and ), where the number of channels of each scale feature is c / 4.

[0036] In the second stage of fusion, the fusion features of four scales are combined. In this stage, the four-scale features are concatenated and further processed through a 1×1 convolutional layer, a batch normalization (BN) layer, and a ReLU activation function, as Figure 2 shown by “CC” in. Through this two-stage fusion strategy, the MESF module can effectively integrate multi-scale features from RGB and TIR modalities. This hierarchical fusion method combines continuous convolutional operations, enhancing the feature representation ability and its complementarity.

[0037] Finally, the original RGB encoded features are reintroduced through residual connections to finally obtain the fused features (with the number of channels being c). The whole process of the MESF module can be expressed as:

[0038]

[0039] where [,] represents the concatenation operation, f ReLU () represents the ReLU activation function, BN represents the batch normalization layer, and f 1×1 (·) represents the 1×1 convolutional operation.

[0040] In addition, different from the third and fourth layers, the RGB and TIR encoded features of the fifth layer are fused through a variant of the MESF module, which is marked as the MESF-H module. Specifically, in the MESF-H module, the dilation rates of dilated depthwise separable convolutions (marked as DSConv-1, DSConv-2, DSConv-3, and DSConv-4 respectively) are set to 1, 2, 3, and 4. In this way, features with receptive fields of 3, 5, 9, and 15 are obtained respectively. The modified dilation rates ensure that the receptive field of the fifth layer features does not exceed the spatial size (15×20) of the input feature map, thus preventing information loss and alleviating the edge effect caused by too large receptive fields.

[0041] After obtaining the multi-level fusion features, an attempt is made to integrate these multi-level features. For this purpose, this study proposes a hierarchical feature aggregation (HFE) module, as Figure 3 shown, and this module includes two main steps. Here, taking the HFE-1 module as an example, as Figure 2 shown, first, the input features and Processed by the fully connected layer and the upsampling operation. The fully connected layer plays a key role in normalizing the feature dimensions and capturing the complex non-linear relationships between features. Dimension normalization ensures that features from different levels can be seamlessly integrated. This is crucial for effective cross-layer integration as it addresses the issues of size and channel mismatches. During this process, features with a larger number of channels are dimensionally compressed to reduce computational costs, while features with fewer channels (such as and ) are expanded to enhance their information content. This selective adjustment achieves a balance between the complexity and richness of the features.

[0042] In the second step, the adjusted features and are concatenated together to achieve cross-layer feature integration. Subsequently, the combined features pass through a 1×1 convolutional layer, a batch normalization (BN) layer, and a ReLU activation function to coordinate the channel dimensions. Then, a dropout layer is used for regularization. Through these two steps, HFE-1 can effectively integrate features from different levels to generate unified and strongly expressive aggregated features. The whole process is summarized as follows:

[0043]

[0044] where, f fc represents the fully connected layer, and f up represents the upsampling operation.

[0045] Note that, different from HFE-1 which focuses on high-level feature integration, HFE-2 emphasizes the integration of middle-level and low-level features. Specifically, HFE-2 receives three inputs, including RGB-encoded features middle-level fusion features and high-level fusion features This combination ensures that HFE-2 can integrate spatial details and semantic cues simultaneously. To balance the computational complexity and feature representation ability, the number of output channels of HFE-2 is set to 96, which ensures that this module has sufficient ability in integrating multi-level features without causing excessive consumption of computational resources.

[0046] Finally, three convolutional blocks are deployed to integrate multi-level features and and generate the final prediction map S. The whole process can be expressed as:

[0047]

[0048] where, represents the element-wise summation operation, Denotes a depthwise separable convolution with a dilation rate of 1, f 1×1 Denotes a 1×1 convolution operation. f Convblock It consists of a dropout layer, two depthwise separable convolutions with dilation rates of 3 and 1 respectively, and an upsampling operation. Here, the depthwise separable convolution is derived from MobilenetV2, which significantly reduces the number of parameters and computational overhead compared to traditional convolutions.

[0049] To improve the accuracy of the segmentation results and accelerate the convergence of the proposed network, this study adopted a hybrid loss function for the final semantic supervision. This loss function consists of weighted cross-entropy loss and Lovasz-softmax loss. This combination strategy fully utilizes the advantages of the two loss functions, helps to solve the class imbalance problem, and directly optimizes the intersection over union (IoU) of the segmentation results. The total loss function is expressed as follows:

[0050]

[0051] Where, Denotes the weighted binary cross-entropy loss, Corresponds to the Lovasz-softmax loss, and GT is the ground truth label for semantic segmentation. K represents the total number of classes, Denotes the subgradient of IoU, and N represents the total number of training samples. In addition, G j Denotes the ground truth label (0 or 1) of the j-th sample, while Y j Denotes the predicted probability of the current j-th sample.

[0052] Table 1 Quantitative comparison results of the present invention on the MFNet dataset

[0053]

[0054] On the MFNet dataset, the present invention was compared with 7 state-of-the-art models, and the specific results are shown in Table 1. Among them, "-" indicates that the author did not provide the corresponding results, and the best results in each column are marked in red. These models can be divided into two groups. The first group is models focusing on RGB-D semantic segmentation, including FuseNet, RedNet, and TSNet. In this study, these models were modified to perform RGB-T semantic segmentation by replacing the input depth image with a thermal infrared image. These models were retrained with their default parameter settings in the training set, and their segmentation performance was evaluated on the test set. The second group contains 4 models specifically designed for RGB-T semantic segmentation, including MFNet, AFNet, MMNet, and ABMDRNet. The segmentation results of existing RGB-T semantic segmentation models are from public benchmarks or directly provided by the authors.

[0055] In Table 1, a comprehensive comparison of the performance of the present invention and seven existing semantic segmentation models on the MFNet dataset is carried out. This evaluation covers the single performance of five different categories and the overall average performance. Generally speaking, the proposed model shows significant advantages in terms of mIoU and mAcc. It is worth noting that the present invention exceeds the second-best model (i.e., ABMDRNet) by 2.3% in terms of mIoU. In the recognition of specific categories, the present invention also performs excellently. Specifically, the present invention achieves the highest IoU score in all categories and realizes competitive Acc. These results confirm the effectiveness of the present invention in dealing with complex multi-category scenarios.

[0056] Table 2 Comparison of the complexity of the present invention on the MFNet dataset

[0057]

[0058] Table 2 compares the complexity of the proposed model with that of other models, including TSNet, TCD, RedNet, FuseNet, RTFNet, FuseSeg, EGFNet and GCGLNet. The evaluation metrics are the computational complexity and segmentation accuracy on the MFNet dataset. It can be seen from Table 2 that the present invention shows the lowest computational complexity, with the number of parameters being 4.15M and the FLOPs being 6.15G, significantly lower than those of EGFNet and FuseSeg. In addition, the present invention also achieves the best mAcc and mIoU metrics, which are 72.3% and 57.1% respectively. At the same time, compared with GCGLNet, which also focuses on lightweight design, the present invention is superior in all key metrics, including higher mIoU, lower computational complexity and faster inference speed. These results clearly show that the present invention can ensure high efficiency and high precision while maintaining lightweight design, making it suitable for real-time application scenarios.

Claims

1. A lightweight RGB-T semantic segmentation method based on multimodal feature fusion, characterized in that: The steps include: S1, select pairs of visible light RGB and thermal infrared TIR images from the RGB-T dataset as the input training set; S2. Build a lightweight RGB-T semantic segmentation model based on multimodal feature fusion, including encoder part, multi-scale feature fusion part, hierarchical feature aggregation part and decoder part; S3. Input the RGB-T image pairs of the training set into the lightweight RGB-T semantic segmentation model for training to generate semantic segmentation results.

2. The lightweight RGB-T semantic segmentation method based on multimodal feature fusion according to claim 1, characterized in that: The lightweight RGB-T semantic segmentation model of multimodal feature fusion includes an encoder part, a multi-scale feature fusion part, a hierarchical feature aggregation part and a decoder part; The encoder part constructs RGB encoder and TIR encoder based on the encoder branch of MobileNetV2, and performs multi-level feature extraction on RGB and TIR images respectively; The multi-scale feature fusion part is used to perform multi-scale feature fusion on the features extracted by the encoder; The hierarchical feature aggregation part is used to aggregate the second-layer RGB encoder features with the fused multi-scale features; The decoder part decodes based on the aggregated features and the low-level features extracted by the first-layer RGB encoder, and finally generates the semantic segmentation result.

3. The lightweight RGB-T semantic segmentation method based on multimodal feature fusion according to claim 2, characterized in that: The specific implementation process of the multi-scale feature fusion part is as follows: The third-layer RGB encoder features and the third-layer TIR encoder features are fused through the multi-scale feature fusion MSEF module to generate the third-layer fusion features. The fourth-layer RGB encoder features and the fourth-layer TIR encoder features are fused through the multi-scale feature fusion MSEF module to generate the fourth-layer fusion features. The fifth-layer RGB encoder features and the fifth-layer TIR encoder features are fused through the multi-scale feature fusion MSEF-H module to generate the fifth-layer fusion features. Generate multi-level fusion features The MSEF module has the same structure as the MSEF-H module.

4. The lightweight RGB-T semantic segmentation method based on multimodal feature fusion according to claim 3, characterized in that: The specific implementation process of the hierarchical feature aggregation is as follows: Fusion Features and It is integrated into aggregated features through the first hierarchical feature aggregation module HFE-1 Fusion Features and the second layer RGB encoder features The second hierarchical feature aggregation module HFE-2 is used to integrate and generate aggregated features.

5. The lightweight RGB-T semantic segmentation method based on multimodal feature fusion according to claim 4, characterized in that: The specific implementation process of the decoder is as follows: The decoder part deploys three convolutional blocks with the same structure. Each convolutional block consists of a dropout layer, two depth-separable convolutions, and an upsampling operation. First, the first convolutional block aggregates the features output by the hierarchical feature aggregation module. Processing to generate decoding features Next, the second convolutional block will With aggregate features After adding each element, convolution is performed to generate features Finally, the third convolutional block will With the first layer RGB encoder feature F1 r After element-by-element addition, convolution is performed, and finally, features are refined through a depth-wise separable convolution with a dilation rate of 1 and a 1×1 convolution to generate the final prediction graph S.

6. The lightweight RGB-T semantic segmentation method based on multimodal feature fusion according to claim 5, characterized in that: The MSEF module adopts a two-stage fusion strategy; The first stage: At each scale, the initial fusion features of RGB and TIR are and Splicing is performed and processed through 1×1 convolution, batch normalization and ReLU activation function to generate preliminary multi-scale fusion features The second stage: the initial fusion features of the four scales and Splicing is performed and global fusion is performed through 1×1 convolution, batch normalization and ReLU activation function to generate fusion features Finally, the original RGB encoding features are transformed into With fusion features Add.

7. The lightweight RGB-T semantic segmentation method based on multimodal feature fusion according to claim 6, characterized in that: The HFE-1 module is used to integrate fusion features and First, these features are dimensionally normalized and feature enhanced through a fully connected layer. Then, the features are adjusted to the same resolution through upsampling and concatenated. The concatenated features are processed by 1×1 convolution, batch normalization, and ReLU activation function to generate unified aggregate features. The HFE-2 module input includes the second layer RGB encoded features and Through the fully connected layer and upsampling operation, the HFE-2 module resizes and concatenates these features, and then processes them through 1×1 convolution, batch normalization, and ReLU activation function to generate aggregate features.

Citation Information

Cited By

  • Dual-light image fusion method for light power equipment

    CN120807315A

  • Particle occlusion discrimination and PSD measurement method and system under decoupling neural network

    CN121384732A