Low-illumination multi-modal scene understanding method based on dynamic feature selection complementation

By building a multimodal selective attention semantic segmentation network based on Transformer, the problem of insufficient utilization of multimodal feature information in low-illumination environments is solved, efficient multimodal fusion is achieved, and the scene understanding ability of unmanned systems under complex low-illumination conditions is improved, and the calculation and parameter burden is reduced.

CN120339774AActive Publication Date: 2025-07-18CHINA UNIV OF MINING & TECH
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510499360.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The prior art cannot effectively utilize multimodal feature information complementarity in low-illumination environments, resulting in insufficient understanding of scenes in unmanned systems under complex low-illumination conditions, and there are problems of parameter redundancy and increased calculation amount.

Method used

A multimodal selective attention semantic segmentation network based on Transformer is adopted, including visible light encoding network, thermal image encoding network, short-distance selective attention module and channel selective self-attention module to build an efficient multimodal fusion network. The short-distance selective attention module reduces information interference between modes, and the channel selective self-attention module enhances the information integrity of the characteristic channel dimension, and combines a lightweight decoding network for feature reconstruction.

Benefits of technology

It significantly improves the utilization efficiency of multimodal feature information, reduces the amount of parameters and calculations, and improves the scene understanding ability of the unmanned system in low-illumination environments, especially in nighttime urban road autonomous driving and dark unmanned navigation systems in underground spaces.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339774A_ABST
    Figure CN120339774A_ABST
Patent Text Reader

Abstract

The invention discloses a low-illumination multi-modal scene understanding method based on dynamic feature selection complementation, which belongs to the artificial intelligence technology and comprises the following steps: constructing a Transform-based parallel feature extraction network according to the characteristic of multi-modal feature information complementation; a short-distance selective attention module is constructed to reduce information interference between the same modals, and complementarity information in different modals is kept; constructing a channel selectivity self-attention module to enhance the information integrity of the feature channel dimension; constructing a lightweight full-connection multi-layer perceptron decoding module which is used for receiving coding network feature information and performing feature reconstruction on feature maps of different scales; according to the method, the complementary advantage of multi-modal feature information can be fully utilized, the situation that the low-illumination scene understanding ability is insufficient due to single-modal information is avoided, and the method can be applied to unmanned system scene analysis and an unmanned navigation system in a complex low-illumination scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and in particular relates to a low-illumination multimodal scene understanding method based on dynamic feature selection complementation. Background Art

[0002] With the rapid development of artificial intelligence technology, unmanned systems have higher and higher requirements for the perception and understanding of complex scenes, especially for autonomous vehicles, intelligent robots, drones and other unmanned systems in complex low-light environments. Single visible light image vision sensors are easily affected by complex low-light environments, so thermal image sensors are combined to make up for the shortcomings of visible light image sensors to ensure that unmanned systems have the same perception capabilities during the day and night and in severe weather conditions. Visible light images can provide rich semantic information, while thermal images can provide stable scene images. Research on efficient fusion methods of visible light images and thermal images can effectively improve the stability of unmanned systems in scene understanding under complex low-light conditions.

[0003] Xie et al. proposed a novel semantic segmentation framework in "SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers". The authors first proposed a hierarchical Transformer encoder, which can obtain multi-scale features without using positional encoding and adapt to inputs of various resolutions. Then the authors proposed a lightweight fully connected multi-layer perceptron decoder, which generates segmentation masks by fusing information at different levels and combining local attention and global attention. Although this method has achieved excellent semantic segmentation performance, it ignores the complementary characteristics of multimodal feature information. The proposed fusion method cannot effectively utilize the complementary information between modalities, and will be affected by the noise interference of different information within and between modalities, which limits the performance of multimodal semantic segmentation to a certain extent. In addition, there are parameter redundancy and increased computational complexity, which is not conducive to deployment on unmanned system equipment. Summary of the invention

[0004] The purpose of the present invention is to provide a low-light multimodal scene understanding method based on dynamic feature selection complementarity, which efficiently utilizes the complementary characteristics of multimodal feature information, adopts different fusion strategies to construct an efficient multimodal joint representation network, solves the problems of insufficient utilization of multimodal characteristic information and low fusion efficiency, and avoids the phenomenon of parameter redundancy and increased calculation caused by inefficient network modules. The present invention provides a low-light multimodal scene understanding method based on dynamic feature selection complementarity, which can be applied in nighttime urban road automatic driving and dark unmanned navigation systems in underground spaces.

[0005] The technical solution for achieving the object of the present invention is: a low-light multi-modal scene understanding method based on dynamic feature selection and complementarity, comprising the following steps:

[0006] Step 1, perform normalization processing on 1569 images in the MFNet dataset, unify the pixel size to H×W, where H represents the length and W represents the width; divide the images with the unified size into a training dataset and a test dataset according to the ratio of 784 / 393, perform data augmentation on the training dataset to form a network training dataset, and transfer to Step 2;

[0007] Step 2, construct a multi-modal selective attention semantic segmentation network:

[0008] The multi-modal selective attention semantic segmentation network includes a visible light encoding network, a thermal image encoding network, a short-distance selective attention module, a channel selective self-attention module, and a lightweight decoding network; among them, both the visible light encoding network and the thermal image encoding network are composed of a Transformer network pre-trained on the ImageNet dataset and serve as encoding networks for extracting features; the short-distance selective attention module is used to retain multi-modal complementary feature information; the channel selective self-attention module is used to enhance the information integrity of the feature channel dimension; the decoding network is composed of a lightweight fully connected multi-layer perceptron and is used to receive the feature information of the encoding network and perform feature reconstruction on feature maps of different scales, and transfer to Step 3;

[0009] Step 3, input the network training dataset into the multi-modal selective attention semantic segmentation network, use the network training dataset to train the multi-modal selective attention semantic segmentation network, and obtain a trained multi-modal selective attention semantic segmentation network model:

[0010] S31, divide the visible light encoding network feature extraction into four stages, and extract four different scales of visible light features corresponding to each stage, namely (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), (H / 32)×(W / 32); correspondingly, divide the thermal image encoding network feature extraction into four stages, and extract four different scales of thermal image features corresponding to each stage, namely (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), (H / 32)×(W / 32), and transfer to S32;

[0011] S32, input the visible light features and thermal image features extracted in each stage into the short-distance selective attention module to obtain short-distance fusion features, and then input the short-distance fusion features into the channel selective self-attention module to obtain the features after interaction and fusion of visible light and thermal images at this stage, and transfer to S33;

[0012] S33. Transfer the features obtained by the interactive fusion of visible light and thermal images with a scale of (H / 4)×(W / 4) in the first stage in S32 and the features obtained by the interactive fusion of visible light and thermal images with scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) obtained in the following three stages respectively to the lightweight fully-connected multi-layer perceptron decoding module to obtain four decoded features with a pixel size of (H / 4)×(W / 4), splice the four decoded features to obtain an aggregated feature, and transfer to S34;

[0013] S34. Perform a channel dimensionality reduction operation on the aggregated feature output in S33 through a multi-layer perceptron network to obtain a dimensionality-reduced feature, calculate the cross-entropy loss between the dimensionality-reduced feature and the corresponding image label in the MFNet dataset, and update the network parameters of the multi-modal selective attention semantic segmentation network with this to obtain a trained multi-modal selective semantic segmentation model, and transfer to step 4.

[0014] Step 4. Input the test dataset into the trained multi-modal selective attention semantic segmentation model, output the prediction result corresponding to each sample in the test set, and test the accuracy of the trained multi-modal selective attention semantic segmentation model.

[0015] Compared with the prior art, the advantages of the present invention are as follows:

[0016] (1) The present invention proposes a low-light multi-modal scene understanding method based on dynamic feature selection and complementarity, which can efficiently utilize the characteristics of multi-modal feature information complementarity, constructs an efficient multi-modal fusion network by adopting different fusion strategies, solves the problems of insufficient utilization of multi-modal characteristic information and low fusion efficiency, and significantly reduces the number of parameters and the amount of calculation.

[0017] (2) The present invention proposes a short-distance selective attention module that fuses the selective attention within and between modalities, thereby reducing information interference in the same modality and ensuring the retention of complementary information and the suppression of noise information in different modalities.

[0018] (3) The present invention proposes a channel selective self-attention module to enhance the information integrity of the feature channel dimension and improve the generalization and robustness. Description of the Drawings

[0019] Figure 1 It is a model diagram of a low-light multi-modal scene understanding method based on dynamic feature selection and complementarity.

[0020] Figure 2 It is an experimental result diagram of the urban road scene of the MFNet dataset.

[0021] Figure 3 It is a result diagram of the urban road scene experiment for the FMB dataset. Specific implementation manners

[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the implementation manners of the present invention in detail:

[0023] Combined with Figure 1 , a low-light multi-modal scene understanding method based on dynamic feature selection complementarity, comprising the following steps:

[0024] Step 1: Normalize 1569 images in the MFNet dataset, unify the pixel size to H×W, where H represents the length and W represents the width; divide the images with the unified size into a training dataset and a test dataset according to the ratio of 784 / 393, perform data augmentation on the training dataset to form a network training dataset, and then proceed to Step 2.

[0025] Step 2: Construct a multi-modal selective attention semantic segmentation network:

[0026] The multi-modal selective attention semantic segmentation network includes a visible light encoding network, a thermal image encoding network, a short-distance selective attention module, a channel selective self-attention module, and a lightweight decoding network; among them, both the visible light encoding network and the thermal image encoding network are composed of a Transformer network pre-trained on the ImageNet dataset and serve as encoding networks for feature extraction; the short-distance selective attention module is used to retain multi-modal complementary feature information; the channel selective self-attention module is used to enhance the information integrity of the feature channel dimension; the decoding network is composed of a lightweight fully-connected multi-layer perceptron and is used to receive the feature information of the encoding network and perform feature reconstruction on feature maps of different scales, and then proceed to Step 3.

[0027] Step 3: Input the network training dataset into the multi-modal selective attention semantic segmentation network, use the network training dataset to train the multi-modal selective attention semantic segmentation network, and obtain a trained multi-modal selective attention semantic segmentation network model:

[0028] S31. Divide the visible light encoding network feature extraction into four stages, and extract four different scales of visible light features corresponding to each stage, namely (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32); correspondingly, divide the thermal image encoding network feature extraction into four stages, and extract four different scales of thermal image features corresponding to each stage, namely (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32), and then proceed to S32.

[0029] S32. Input the visible light features and thermal image features extracted in each stage into the short - distance selective attention module to obtain short - distance fusion features, and then input the short - distance fusion features into the channel - selective self - attention module to obtain the features after the interaction and fusion of visible light and thermal images at this stage:

[0030] Input the visible light and thermal image features extracted in the first stage of S31 into the short - distance selective attention module to obtain fusion features, and then input the fusion features into the channel - selective self - attention module to obtain the features with a scale of (H / 4)×(W / 4) after the interaction and fusion of visible light and thermal images; input the visible light and thermal image features extracted in the second stage into the short - distance selective attention module to obtain fusion features, and then input the fusion features into the channel - selective self - attention module to obtain the features with a scale of (H / 8)×(W / 8) in the second stage; input the visible light and thermal image features extracted in the third stage into the short - distance selective attention module to obtain fusion features, and then input the fusion features into the channel - selective self - attention module to obtain the features with a scale of (H / 16)×(W / 16) in the third stage; input the visible light and thermal image features extracted in the fourth stage into the short - distance selective attention module to obtain fusion features, and then input the fusion features into the channel - selective self - attention module to obtain the features with a scale of (H / 32)×(W / 32) in the fourth stage.

[0031] Among them, input the visible light features and thermal image features into the short - distance selective attention module to obtain short - distance fusion features, specifically as follows:

[0032] S32 - 1. Define the query vector mapped from the visible light features as Q r and the key vector as K r and the value vector as V r Similarly, define the query vector mapped from the thermal image features as Q t and the key vector as K t and the value vector as V t; Split the query vector, key vector, and value vector respectively, and calculate the mean and standard deviation of the split query vector, as well as the mean and standard deviation of the split key vector. Normalize the split query vector and key vector according to the mean and standard deviation respectively:

[0033]

[0034] Where the modality x ∈ {r, t}; r is the visible light feature; t is the thermal image feature; is the split query vector; is the split key vector; is the split value vector; Partition is the split function; Q' x is the normalized query vector; K' x is the normalized key vector; mean is the mean function; std is the standard deviation function.

[0035] S32-2. According to the normalized query vector and the normalized key vector, calculate the visible light self-attention A rr and the thermal image self-attention A tt respectively. According to the visible light self-attention A rr and the thermal image self-attention A tt calculate the visible light modality-internal selective attention A' rr and the thermal image modality-internal selective attention A' tt respectively:

[0036]

[0037] A' rr = Select(σ(A rr ) > α, A rr , -∞)

[0038] A' tt = Select(σ(A tt ) > α, A tt , -∞)

[0039] Where Q' r is the query vector after normalization of the visible light feature; K' r is the key vector after normalization of the visible light feature; Q' t is the query vector after normalization of the thermal image feature; K' t is the key vector after normalization of the thermal image feature; T represents transpose; τ1, τ2 are learnable parameters; Select(σ(A rr ) > α, A rr , -∞) is a selection mechanism, indicating that when σ(A rr) When >α, return A rr , otherwise return -∞; Select(σ(A tt )>α, A tt , -∞) is the same; σ is the sigmoid function; α is a hyperparameter, set to 0.2.

[0040] S32-3, according to the query vector Q' after normalization of visible light features r and the key vector K' after normalization of thermal image features t , calculate the attention A between the visible light and thermal image modalities rt , according to the query vector Q' after normalization of thermal image features t and the key vector K' after normalization of visible light features r , calculate the attention A between the thermal image and visible light modalities tr , and according to the attention A between the visible light and thermal image modalities rt and the attention A between the thermal image and visible light modalities tr respectively calculate the selective attention A' between the visible light and thermal image modalities rt and the selective attention A' between the thermal image and visible light modalities tr :

[0041]

[0042] A' rt = Select(σ(A rt )>α, A rt , -∞)

[0043] A' tr = Select(σ(A tr )>α, A tr , -∞)

[0044] In the formula, T represents transpose; τ3 and τ4 are both learnable parameters; Select(σ(A rt )>α, A rt , -∞) is a selection mechanism, indicating that when σ(A rt )>α, return A rt , otherwise return -∞; Select(σ(A tr )>α, A tr , -∞) is the same; σ is the sigmoid function; α is a hyperparameter, set to 0.2.

[0045] According to the selective attention A' within the visible light modality rt , the selective attention A' within the thermal image modality tt , and the selective attention A' between the visible light and thermal image modalities rtAnd selective attention between thermal images and visible light modalities A' tr Calculate the short - distance selective attention score S s :

[0046] S s = Softmax(A′ rr + A′ tt + A′ rt + A′ tr )

[0047] Where Softmax is the normalized exponential function.

[0048] S32 - 4, using the segmented value vector and the short - distance selective attention score S s , calculate the short - distance preliminary fusion feature Flip the short - distance preliminary fusion feature , then add it to the visible light feature F r and the thermal image feature F t , and perform layer normalization to obtain the feature F s ":

[0049]

[0050] Wherein, is the segmented value vector of the visible light feature; is the segmented value vector of the thermal image feature; LN is the layer normalization operation; Linear is the linear mapping; Reverse is the flip operation.

[0051] Pass the feature F s " through a multi - layer perceptron and perform layer normalization to obtain the short - distance fusion feature F s :

[0052] F s = LN(MLP(F s ″)+ F s ″)

[0053] Where MLP represents the multi - layer perceptron.

[0054] Among them, input the short - distance fusion feature into the channel - selective self - attention module to obtain the feature after the interaction and fusion of visible light and thermal images at this stage, specifically as follows:

[0055] S32 - A, the short - distance fusion feature F s output by the short - distance selective attention module passes through a linear mapping to obtain the query vector Q s , the key vector K s and the value vector V s, and normalize the query vector Q s and the key vector K s to obtain the normalized query vector Q' s and the key vector K' s . Then, calculate the channel self-attention A s and the channel selective self-attention A' s according to the normalized query vector Q' cc and the key vector K' cc , and normalize the channel selective self-attention to obtain the channel selective self-attention score S c :

[0056]

[0057] A' cc = Select(σ(A cc ) > β, A cc , -∞)

[0058] S c = Softmax(A' cc )

[0059] In the formula, τ5 is a learnable parameter; Select(σ(A cc ) > β, A cc , -∞) is a selection mechanism, indicating that when σ(A cc ) > β, return A cc , otherwise return -∞; σ is the sigmoid function; β is a hyperparameter set to 0.05; Softmax is the normalization function.

[0060] S32 - B, calculate the preliminary channel fusion feature F s ' by using the value vector V c and the channel selective self-attention score S c . Add the preliminary channel fusion feature F c ' to the short-distance fusion feature F s , and obtain the feature F c " through linear mapping and layer normalization operations:

[0061] F′ c = V s S c

[0062] F″ c = LN(Linear(F c ′) + F s )

[0063] In the formula, Linear represents linear mapping; LN is the layer normalization operation.

[0064] Finally, feature F c " is obtained through a multi-layer perceptron and a layer normalization operation to get the channel fusion feature F c :

[0065] F c = LN(MLP(F c ″)+F c ″)

[0066] In the formula, MLP is the multi-layer perceptron.

[0067] Proceed to S33;

[0068] S33. The interaction fusion features of visible light and thermal images with a scale of (H / 4)×(W / 4) obtained in the first stage in S32 and the interaction fusion features of visible light and thermal images with scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) obtained in the subsequent three stages are respectively transmitted to the lightweight fully connected multi-layer perceptron decoding module to obtain four decoding features with a pixel size of (H / 4)×(W / 4), and the four decoding features are concatenated to obtain an aggregated feature, specifically as follows:

[0069] The interaction fusion features of visible light and thermal images with scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) obtained in the subsequent three stages are respectively subjected to upsampling operations to obtain three decoding features with a scale of (H / 4)×(W / 4). The above three decoding features with a scale of (H / 4)×(W / 4) are concatenated with the interaction fusion feature of visible light and thermal images with a scale of (H / 4)×(W / 4) obtained in the first stage to obtain the aggregated feature X C :

[0070] X C = Concat(H 1 ,H 2 ,H 3 ,H 4 )

[0071] In the formula, Concat represents the concatenation operation; H 1 ,H 2 ,H 3 ,H 4 respectively represent the decoding features of the four stages of the encoder.

[0072] Proceed to S34.

[0073] S34. Perform a channel dimensionality reduction operation on the aggregated features output in S33 through a multi-layer perceptron network to obtain the dimensionality-reduced features. Calculate the cross-entropy loss between the dimensionality-reduced features and the corresponding image labels in the MFNet dataset, and update the network parameters of the multi-modal selective attention semantic segmentation network with this to obtain a trained multi-modal semantic segmentation model. Specifically as follows:

[0074]

[0075] In the formula, Loss is the cross-entropy loss; is the label value; y i is the sample prediction value output by the model, and i is the category index of each pixel point.

[0076] Proceed to step 4.

[0077] Step 4. Input the test dataset into the trained multi-modal selective attention semantic segmentation model, output the prediction results corresponding to each sample in the test set, and test the accuracy of the trained multi-modal selective attention semantic segmentation model.

[0078] Embodiment 1

[0079] A method for low-light night vision scene understanding based on multi-modal image fusion according to the present invention is as follows:

[0080] Step 1. Normalize 1569 images in the MFNet dataset, unify the pixel size to H×W, where H represents the length and W represents the width; divide the images with the unified size into a training dataset and a test dataset according to the ratio of 784 / 393, perform data augmentation on the training dataset to form a network training dataset, and proceed to step 2.

[0081] Step 2. Construct a multi-modal selective attention semantic segmentation network:

[0082] The multi-modal selective attention semantic segmentation network includes a visible light encoding network, a thermal image encoding network, a short-distance selective attention module, a channel selective self-attention module, and a lightweight decoding network; among them, both the visible light encoding network and the thermal image encoding network are composed of a Transformer network pre-trained on the ImageNet dataset and serve as encoding networks for feature extraction; the short-distance selective attention module is used to retain multi-modal complementary feature information; the channel selective self-attention module is used to enhance the information integrity of the feature channel dimension; the decoding network is composed of a lightweight fully-connected multi-layer perceptron and is used to receive the feature information of the encoding network and perform feature reconstruction on feature maps of different scales, and proceed to step 3.

[0083] Step 3: Use the network training dataset to train the selective attention multi-modal semantic segmentation network to obtain a trained multi-modal selective attention semantic segmentation network model:

[0084] S31: Divide the visible light encoding network feature extraction into four stages, and extract four different scales of visible light features corresponding to each stage, which are (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), (H / 32)×(W / 32); correspondingly, divide the thermal image encoding network feature extraction into four stages, and extract four different scales of thermal image features corresponding to each stage, which are (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), (H / 32)×(W / 32), and then go to S32.

[0085] S32: Input the visible light features and thermal image features extracted in each stage into the short-distance selective attention module to obtain short-distance fusion features, and then input the short-distance fusion features into the channel selective self-attention module to obtain the features after the interaction and fusion of visible light and thermal images in this stage, and then go to S33.

[0086] S33: Transmit the features after the interaction and fusion of visible light and thermal images with the scale of (H / 4)×(W / 4) obtained in the first stage of S32 and the features after the interaction and fusion of visible light and thermal images with the scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) obtained in the following three stages to the lightweight fully connected multi-layer perceptron decoding module respectively to obtain four decoding features with the pixel size of (H / 4)×(W / 4), and splice the four decoding features to obtain an aggregated feature, and then go to S34.

[0087] S34: Perform a channel dimensionality reduction operation on the aggregated feature output in S33 through a multi-layer perceptron network to obtain a dimensionality reduction feature, calculate the cross-entropy loss between the dimensionality reduction feature and the corresponding image label in the MFNet dataset, and update the network parameters of the multi-modal selective attention semantic segmentation network with this to obtain a trained multi-modal selective semantic segmentation model, and then go to Step 4.

[0088] Step 4: Input the test dataset into the trained multi-modal selective attention semantic segmentation model, output the prediction results corresponding to each sample in the test set, and test the accuracy of the trained multi-modal selective attention semantic segmentation model.

[0089] The method of the present invention conducts relevant experiments on a computer configured with an Intel Xeon CPU and a Tesla V100 GPU using a network of the Python programming language and the Pytorch deep learning framework. During the training process, the batch size is set to 8, the AdamW optimizer with a weight decay of 0.01 is used as the optimizer, and the Warm up ploy learning rate adjustment method is adopted for the learning rate adjustment method, with an initial learning rate of 0.00006. Training multiple batches on the training sample set yields the low-light multi-modal scene understanding method based on dynamic feature selection and complementarity described in the present invention. The visualized experimental results are as shown in Figure 2 , Figure 3 shown.

[0090] To demonstrate the superior performance of the present invention, the present invention selects the most advanced multi-modal semantic segmentation method as a comparison model recently. The comparison results are shown in Table 1. The accuracy, intersection over union, number of parameters, and computational complexity of the model are evaluated on the MFNet dataset, and its input is a visible light image (480×640×3) and a thermal image (480×640×3).

[0091] Table 1 Comparison experimental results of different methods on the MFNet database

[0092]

[0093] It can be seen from the experimental results that the method of the present invention achieves a segmentation accuracy of 60.7%, the computational complexity FLOPs drops to 74.86G, the number of parameters Params drops to 41.08M, and the low-light multi-modal scene understanding method based on dynamic feature selection and complementarity can be applied in autonomous driving on urban roads at night and the dark and weak unmanned navigation system in underground spaces.

Claims

1. A low - illumination multi - modal scene understanding method based on complementary dynamic feature selection, characterized in that, The steps are as follows: Step 1: Normalize 1,569 images in the MFNet dataset, and unify the pixel size to H×W, where H represents the length and W represents the width; Divide the images with unified size into a training dataset and a test dataset according to the ratio of 784 / 393. Perform data augmentation on the training dataset to form a network training dataset, and then go to Step 2; Step 2: Construct a multi-modal selective attention semantic segmentation network: The multi-modal selective attention semantic segmentation network includes a visible light encoding network, a thermal image encoding network, a short-distance selective attention module, a channel-selective self-attention module, and a lightweight decoding network. Among them, both the visible light encoding network and the thermal image encoding network are composed of a Transformer network pre-trained on the ImageNet dataset and serve as encoding networks for feature extraction. The short-distance selective attention module is used to retain multi-modal complementary feature information. The channel-selective self-attention module is used to enhance the information integrity of the feature channel dimension. The decoding network is composed of a lightweight fully-connected multi-layer perceptron and is used to receive the feature information of the encoding network and perform feature reconstruction on feature maps of different scales, and then go to Step 3; Step 3: Input the network training dataset into the multi-modal selective attention semantic segmentation network, and use the network training dataset to train the multi-modal selective attention semantic segmentation network to obtain a trained multi-modal selective attention semantic segmentation network model, and then go to Step 4; Step 4: Input the test dataset into the trained multi-modal selective attention semantic segmentation model, output the prediction results corresponding to each sample in the test set, and test the accuracy of the trained multi-modal selective attention semantic segmentation model.

2. The low-illumination multi-modal scene understanding method based on complementary dynamic feature selection according to claim 1, characterized in that In Step 3, input the network training dataset into the multi-modal selective attention semantic segmentation network, and use the network training dataset to train the multi-modal selective attention semantic segmentation network to obtain a trained multi-modal selective attention semantic segmentation network model, specifically as follows: S31: Divide the visible light feature extraction of the visible light encoding network into four stages, and extract four different scales of visible light features corresponding to each stage, namely (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), (H / 32)×(W / 32); Correspondingly, divide the thermal image feature extraction of the thermal image encoding network into four stages, and extract four different scales of thermal image features corresponding to each stage, namely (H / 4)×(W / 4), (H / 8)×(W / 8), (H / 16)×(W / 16), (H / 32)×(W / 32), and then go to S32; S32: Input the visible light features and thermal image features extracted in each stage into the short-distance selective attention module to obtain short-distance fusion features, and then input the short-distance fusion features into the channel-selective self-attention module to obtain the features after the interaction and fusion of visible light and thermal images at this stage, and then go to S33; S33. The features of the visible light and thermal images with a scale of (H / 4)×(W / 4) obtained in the first stage of S32 after interactive fusion and the features of the visible light and thermal images with scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) obtained in the following three stages after interactive fusion are respectively transmitted to the lightweight fully connected multi-layer perceptron decoding module to obtain four decoded features with a pixel size of (H / 4)×(W / 4), and the four decoded features are concatenated to obtain an aggregated feature, then go to S34; S34. The aggregated feature output in S33 is subjected to a channel dimensionality reduction operation through a multi-layer perceptron network to obtain a dimensionality-reduced feature. Calculate the cross-entropy loss between the dimensionality-reduced feature and the corresponding image label in the MFNet dataset, and use this to update the network parameters of the multi-modal selective attention semantic segmentation network to obtain a trained multi-modal selective semantic segmentation model.

3. A low-light multi-modal scene understanding method based on complementary dynamic feature selection according to claim 2, characterized in that In S32, the visible light feature and the thermal image feature are input into the short-distance selective attention module to obtain a short-distance fusion feature, specifically as follows: S32-1, define the query vector mapped from the visible light feature as Q r , define the key vector as K r , define the value vector as V r . Similarly, define the query vector mapped from the thermal image feature as Q t , define the key vector as K t , define the value vector as V t ; split the query vector, key vector, and value vector respectively, and calculate the mean and standard deviation of the split query vector, as well as the mean and standard deviation of the split key vector. Normalize the split query vector and key vector respectively according to the mean and standard deviation: where the modality x ∈ {r, t}; r is the visible light feature; t is the thermal image feature; is the segmented query vector; is the segmented key vector; is the segmented value vector; Partition is the segmentation function; Q' x is the normalized query vector; K' x is the normalized key vector; mean is the mean function; std is the standard deviation function; S32-2, calculate the visible light self-attention A and the thermal image self-attention A respectively according to the normalized query vector and the normalized key vector rr and calculate the selective attention A' within the visible light modality and the selective attention A' within the thermal image modality respectively according to the visible light self-attention A tt and the thermal image self-attention A rr : tt Calculate the selective attention A' within the visible light modality and the selective attention A' within the thermal image modality respectively rr and tt ​ A rr = Q' r K' r T * τ1 A tt = Q' t K' t T * τ2 A' rr = Select(σ(A rr ) > α, A rr , -∞) A' tt = Select(σ(A tt ) > α, A tt , -∞) where Q' r is the query vector after normalization of visible light features; K' r is the key vector after normalization of visible light features; Q' t is the query vector after normalization of thermal image features; K' t is the key vector after normalization of thermal image features; T represents transpose; τ1, τ2 are learnable parameters; Select(σ(A rr )>α,A rr ,-∞) is a selection mechanism, indicating that when σ(A rr )>α, return A rr , otherwise return -∞; Select(σ(A rr )>α,A tt ,-∞) is the same; σ is the sigmoid function; α is a hyperparameter, set to 0.2; The query vector Q' normalized according to the visible light features r and the key vector K' normalized according to the thermal image features t , calculate the attention A between the visible light and thermal image modalities rt , the query vector Q' normalized according to the thermal image features t and the key vector K' normalized according to the visible light features r , calculate the attention A between the thermal image and visible light modalities tr , and according to the attention A between the visible light and thermal image modalities rt and the attention A between the thermal image and visible light modalities tr , respectively calculate the selective attention A' between the visible light and thermal image modalities rt and the selective attention A' between the thermal image and visible light modalities tr : A rt = Q' r K' t T *τ3 A tr = Q' t K' r T * τ4 A' rt = Select(σ(A rt ) > α, A rt , -∞) A' tr = Select(σ(A tr ) > α, A tr , -∞) where T represents transpose; both τ3 and τ4 are learnable parameters; Select(σ(A tr )>α,A tr , -∞) is a selection mechanism, indicating that when σ(A rt )>α, it returns A rt , otherwise it returns -∞; Select(σ(A tr )>α,A tr , -∞) is the same; σ is the sigmoid function; α is a hyperparameter, set to 0.2; According to the selective attention A' within the visible light modality rr , the selective attention A' within the thermal image modality tt , the selective attention A' between the visible light and thermal image modalities rt and the selective attention A' between the thermal image and visible light modalities tr calculate the short - distance selective attention score S s :[[]]END]] S s = Softmax(A′ rr +(A′ tt +A rt +A tr ) where Softmax is the normalized exponential function; S32-4, using the segmented value vector and the short-distance selective attention score S s , calculate the short-distance preliminary fusion feature Flip the short-distance preliminary fusion feature , then add it to the visible light feature F r and the thermal image feature F t , and perform layer normalization after addition to obtain the feature F s ": In the formula, is the value vector after visible light feature segmentation; is the value vector after thermal image feature segmentation; LN is the layer normalization operation; Linear is the linear mapping; Reverse is the flipping operation; Feature F s “The short-distance fusion feature F is obtained through a multi-layer perceptron and layer normalization operations s : F s = LN(MLP(F s ″) + F s ″) where MLP represents the multi-layer perceptron.

4. A low-light multi-modal scene understanding method based on complementary dynamic feature selection according to claim 3, characterized in that In S32, the short-distance fusion feature is input into the channel selective self-attention module to obtain the features of the visible light and thermal images after interactive fusion in this stage, specifically as follows: S32-A, linearly map the short-distance fusion feature F output by the short-distance selective attention module to obtain the query vector Q s , the key vector K s , and the value vector V s . Then normalize the query vector Q s and the key vector K s to obtain the normalized query vector Q' s and the key vector K' s . Next, calculate the channel self-attention A s and the channel selective self-attention A' s according to the normalized query vector Q' s and the key vector K'. Then normalize the channel selective self-attention to obtain the channel selective self-attention score S cc : cc c c : A cc = Q' s K' s T * τ5 A' cc = Select(σ(A cc ) > β, A cc , -∞) S c = Softmax(A' cc ) where τ5 is a learnable parameter; Select(σ(A cc )>β,A cc ,-∞) is a selection mechanism, indicating that when σ(A cc )>β, return A cc , otherwise return -∞; σ is the sigmoid function; β is a hyperparameter set to 0.05; Softmax is the normalization function; S32-B, using the value vector V s and the channel selective self-attention score S c to calculate the preliminary channel fusion feature F c '. Add the preliminary channel fusion feature F c ' to the short-distance fusion feature F s and obtain the feature F c " through linear mapping and layer normalization operations: F′ c = V s S c F″ c = LN(Linear(F c ′)+F s ) where Linear represents the linear mapping; LN is the layer normalization operation; Finally, feature F c " passes through a multi-layer perceptron and undergoes layer normalization operation to obtain the channel fusion feature F c : F c = LN(MLP(F c ″)+F c ″) where MLP is the multi-layer perceptron.

5. A low-light multi-modal scene understanding method based on complementary dynamic feature selection according to claim 4, characterized in that In S33, the features of the visible light and thermal images with a scale of (H / 4)×(W / 4) obtained in the first stage of S32 after interactive fusion and the features of the visible light and thermal images with scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) obtained in the following three stages after interactive fusion are respectively transmitted to the lightweight fully connected multi-layer perceptron decoding module to correspondingly obtain four decoded features with a pixel size of (H / 4)×(W / 4), and the four decoded features are concatenated to obtain an aggregated feature, specifically as follows: The upsampling operations are respectively performed on the features obtained by the interaction fusion of visible light and thermal images with scales of (H / 8)×(W / 8), (H / 16)×(W / 16), and (H / 32)×(W / 32) in the latter three stages, to obtain three decoded features with a scale of (H / 4)×(W / 4). The above three decoded features with a scale of (H / 4)×(W / 4) are concatenated with the features with a scale of (H / 4)×(W / 4) obtained by the interaction fusion of visible light and thermal images in the first stage, to obtain the aggregated feature X C : X C = Concat(H 1 , H 2 , H 3 , H 4 ) Wherein, Concat represents a concatenation operation; H 1 、H 2 、H 3 、H 4 respectively represent the decoded features of the four stages of the encoder.

6. A low-light multi-modal scene understanding method based on complementary dynamic feature selection according to claim 5, characterized in that In S34, the aggregated feature output in S33 is subjected to a channel dimensionality reduction operation through a multi-layer perceptron network to obtain a dimensionality-reduced feature. Calculate the cross-entropy loss between the dimensionality-reduced feature and the corresponding image label in the MFNet dataset, and use this to update the network parameters of the multi-modal selective attention semantic segmentation network to obtain a trained multi-modal selective semantic segmentation model, specifically as follows: Where Loss is the cross-entropy loss; is the label value; y i is the sample prediction value output by the model, and i is the class index of each pixel point.

Citation Information

Patent Citations

  • Multi-modal data fusion method based on semantic information amount and application

    CN115470856A

  • RGB-D multi-modal semantic segmentation method

    CN116597135A

  • Low-light night vision scene understanding method based on multi-modal image fusion

    CN117853856A

  • Online vectorization map construction method based on lightweight prior semantic map

    CN118864646A

  • Medical image segmentation method fusing multi-modal graphic and text information

    CN119205800A