An RGB-T semantic segmentation method, system, device and storage medium
By adopting the methods of feature extraction, probability feature fusion and preference fusion in RGB-T semantic segmentation, the problem of low semantic segmentation accuracy in the existing technology under unfavorable lighting conditions is solved, high-precision RGB-T semantic segmentation is achieved, and the problem of modal preference is alleviated.
Patent Information
- Application Number
- CN202310406514.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-10
AI Technical Summary
The existing RGB semantic segmentation methods cannot provide high-accurate semantic segmentation results under adverse lighting conditions such as low light, rainy, and fog. In addition, deep learning models are easily affected by modal preference problems when fusing RGB and infrared images, resulting in insufficient information extraction and utilization.
By acquiring RGB images and infrared images, feature extraction is performed separately to obtain RGB modal features and infrared modal features. سپس inputs these features into probability feature fusion module for fusion feature preference calculation, obtaining spatial fusion factor, and performing preference fusion based on this factor to obtain fusion features. Finally, the input decoder is decoded to obtain segmentation results.
High-precision semantic segmentation under unfavorable lighting conditions is achieved, modal preference problem is alleviated, information of RGB and infrared images is fully utilized, and segmentation accuracy is improved.
Smart Images

Figure CN116664829B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an RGB-T semantic segmentation method, system, device and storage medium. Background Art
[0002] Semantic segmentation is one of the advanced tasks in computer vision. The main task of semantic segmentation is to assign a classification label to each image pixel. At present, semantic segmentation plays an important role in many applications, including but not limited to medical image analysis, autonomous driving, indoor parsing, etc. Especially in the field of autonomous driving, semantic segmentation has become a basic task. However, most semantic segmentation applications use three-channel RGB images captured by visible cameras, and the images are often taken under sufficient light conditions. In environments such as low light, rainy, and foggy conditions, due to the limitations of visible light imaging, existing RGB semantic segmentation methods cannot provide the expected performance under adverse lighting conditions. Especially in dim lighting conditions, using RGB semantic segmentation methods, high-accuracy semantic segmentation results cannot be obtained.
[0003] In recent years, the localization of infrared chips has greatly reduced the price of infrared products. Infrared products have begun to enter the civilian field, especially in the field of vehicle driving. Due to the advantages of low price and convenient deployment, infrared cameras have begun to be widely used. Infrared cameras have become one of the essential hardware for the new generation of vehicles. In order to overcome the limitations of RGB semantic segmentation, an economical and effective method is to combine RGB images and infrared thermal images (Thermal Image, referred to as T) for semantic segmentation, that is, RGB-T semantic segmentation. Compared with RGB images, infrared images have their own advantages and disadvantages. In terms of advantages, infrared images have two points: 1) In theory, objects with temperatures above absolute zero can be captured by infrared cameras, so infrared images can capture many key information that RGB images miss, such as pedestrians in the dark. 2) The wavelength of infrared images is between 0.1 and 100 microns, and the images are not affected by visible light conditions. Even under the influence of various strong light, accurate information can still be accurately received. In terms of disadvantages, the defects of infrared images are also very obvious. That is, due to the influence of thermal cross-talk between different objects, the boundary information between objects in infrared images is easily blurred and lost, but RGB images can provide details and boundary information of target objects. Obviously, in theory, RGB images and infrared images can form a complementary relationship in terms of information advantages and disadvantages. So far, academic research on RGB-T semantic segmentation based on deep learning has developed rapidly, and various outstanding algorithms have emerged in an endless stream. However, first, existing algorithms often use simple methods of addition or connection to directly fuse the features of the two modalities. Second, when exploring the fusion mechanism of various RGB images and infrared images, the current algorithms all potentially assume that the fusion results of different modal features are unique. This assumption makes the model performance susceptible to the problem of modal preference, which will lead to insufficient utilization of the model's information extraction of the modality, affecting the representation ability of the fused features, and thus affecting the segmentation accuracy.
[0004] In view of this, how to effectively fuse the information of RGB images and infrared images to achieve high-precision semantic segmentation is an urgent problem to be solved. Summary of the invention
[0005] In view of this, the embodiments of the present invention provide a RGB-T semantic segmentation method, system, device and storage medium, which can efficiently and accurately implement RGB-T semantic segmentation.
[0006] On the one hand, an embodiment of the present invention provides an RGB-T semantic segmentation method, comprising:
[0007] Obtain RGB images and infrared images of the scene to be identified;
[0008] Extract features from the RGB image and the infrared image respectively through an encoder to obtain RGB modality features and infrared modality features; among them, both the RGB modality features and the infrared modality features include multi-level modality features.
[0009] Input the RGB modality features and the infrared modality features into a probability feature fusion module to calculate the fusion feature preference, and obtain spatial fusion factors at each level.
[0010] Based on the spatial fusion factors, perform preference fusion on the RGB modality features and the infrared modality features to obtain fusion features at each level.
[0011] Input the fusion features at each level into a decoder for decoding processing to obtain a segmentation result.
[0012] Optionally, in the step of obtaining the RGB image and the infrared image of the scene to be recognized, it further includes:
[0013] Based on a preset resolution, adjust and unify the resolutions of the RGB image and the infrared image.
[0014] Optionally, the method further includes:
[0015] Construct a ResNet-50 network, remove the fully connected layer and the pooling layer in the ResNet-50 network to obtain a feature extraction network.
[0016] Optionally, the encoder includes two symmetric feature extraction networks. Extracting features from the RGB image and the infrared image respectively through the encoder to obtain RGB modality features and infrared modality features includes:
[0017] Extract low-level texture information and high-level semantic classification information from the RGB image and the infrared image respectively through the two feature extraction networks to obtain multi-level RGB modality features and infrared modality features.
[0018] Among them, the feature extraction network is a multi-level structure, and each level includes bottleneck layers of different specifications.
[0019] Optionally, inputting the RGB modality features and the infrared modality features into a probability feature fusion module to calculate the fusion feature preference and obtain spatial fusion factors at each level includes:
[0020] Perform intermediate fusion processing on the RGB modality features and the infrared modality features in a hierarchical correspondence to obtain multi-level intermediate fusion features; the intermediate fusion processing includes channel dimension connection, first convolution processing, batch normalization, and first activation processing.
[0021] Based on the Gaussian distribution, determine the target samples at each level according to the intermediate fusion features.
[0022] Preference calculation is performed according to the target samples to obtain the spatial fusion factors at each level; the preference calculation includes a second convolution process and a second activation process.
[0023] Optionally, based on the spatial fusion factors, preference fusion is performed on the RGB modality features and the infrared modality features to obtain the fusion features at each level, including:
[0024] Based on the spatial fusion factors at each level, the fusion weight values of the RGB modality features and the infrared modality features at the corresponding levels are set, and weighted fusion is performed to obtain the fusion features at each level;
[0025] Among them, the expression of the fusion feature is:
[0026]
[0027] Among them, represents the fusion feature, W represents the spatial fusion factor, represents the RGB modality feature, represents the infrared modality feature, and i, j, k are all indices.
[0028] Optionally, the decoder includes multiple decoding modules, and the fusion features at each level are input into the decoder for decoding processing to obtain the segmentation result, including:
[0029] The fusion features at each level are respectively input into each layer decoding module in the multiple decoding modules for decoding processing to obtain the segmentation result;
[0030] Among them, the input data of each decoding module in the multiple decoding modules includes the result of channel dimension connection of the output of the previous decoding module and the fusion features at the corresponding level.
[0031] On the other hand, an embodiment of the present invention provides an RGB-T semantic segmentation system, including:
[0032] The first module is used to obtain the RGB image and the infrared image of the scene to be recognized;
[0033] The second module is used to respectively perform feature extraction on the RGB image and the infrared image through the encoder to obtain the RGB modality features and the infrared modality features; among them, both the RGB modality features and the infrared modality features include multi-level modality features;
[0034] The third module is used to input the RGB modality features and the infrared modality features into the probability feature fusion module for fusion feature preference calculation to obtain the spatial fusion factors at each level;
[0035] The fourth module is used to perform preference fusion on the RGB modality features and the infrared modality features based on the spatial fusion factors to obtain the fusion features at each level;
[0036] The fifth module is used to input the fusion features at each level into a decoder for decoding processing to obtain a segmentation result.
[0037] On the other hand, an embodiment of the present invention provides an RGB-T semantic segmentation device, including a processor and a memory;
[0038] The memory is used to store programs;
[0039] The processor executes the program to implement the method as described above.
[0040] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, and the storage medium stores a program, and the program is executed by a processor to implement the method as described above.
[0041] An embodiment of the present invention also discloses a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method as described above.
[0042] An embodiment of the present invention first obtains an RGB image and an infrared image of a scene to be recognized; respectively performs feature extraction on the RGB image and the infrared image through an encoder to obtain RGB modality features and infrared modality features; wherein, both the RGB modality features and the infrared modality features include multi-level modality features; through multi-level modality feature extraction, the embodiment of the present invention fully extracts and utilizes the information of each modality to obtain more semantic information; inputs the RGB modality features and the infrared modality features into a probability feature fusion module for fusion feature preference calculation to obtain spatial fusion factors at each level; based on the spatial fusion factors, performs preference fusion on the RGB modality features and the infrared modality features to obtain fusion features at each level; by introducing spatial fusion factors, the embodiment of the present invention breaks the uniqueness assumption of fusion features, regards various different preference results of fusion features as samples of a certain probability distribution, thereby transforming the learning problem of fusion features into the learning problem of the probability distribution of fusion features; finally, inputs the fusion features at each level into a decoder for decoding processing to obtain a segmentation result. The embodiment of the present invention can effectively fuse the information of the RGB image and the infrared image to achieve high-precision semantic segmentation. Description of the Drawings
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0044] Figure 1 Schematic diagram of the effect that the existing fusion result is affected by modal preference;
[0045] Figure 2 Schematic flow diagram of an RGB-T semantic segmentation method provided by an embodiment of the present invention;
[0046] Figure 3 Schematic diagram of the flow architecture of an RGB-T semantic segmentation method provided by an embodiment of the present invention;
[0047] Figure 4 Schematic diagram of the structural parameters of the feature extraction network provided by an embodiment of the present invention;
[0048] Figure 5 Schematic diagram of the structure of the probability feature fusion module provided by an embodiment of the present invention;
[0049] Figure 6 Schematic diagram of the structure of the decoding module provided by an embodiment of the present invention;
[0050] Figure 7 Schematic diagram of the parameter settings of the decoding module provided by an embodiment of the present invention;
[0051] Figure 8 Schematic diagram of an example image of the MFNet dataset provided by an embodiment of the present invention;
[0052] Figure 9 Schematic diagram of the prediction result for the MFNet dataset provided by an embodiment of the present invention. Detailed implementation manners
[0053] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0054] For the characteristic that the existing method assumes that the fusion result of different modal features is unique, an example may be given to illustrate, as Figure 1 shown, the prominent railing object in the RGB image is very blurred in the infrared image. On the contrary, the prominent pedestrian in the infrared image is hardly visible in the RGB image. For the fused features, the railing is hardly visible. This is because the result of the fused features is unique. After being trained, the deep learning model will prefer to use the stable and easy-to-learn infrared image, resulting in insufficient extraction and utilization of the information of the RGB image by the fused features, the railing target being ignored, and the segmentation result being inaccurate.
[0055] In view of the problem that the fusion method of the existing method is too simple, resulting in the fusion result not being able to fully utilize the features of each modality, on the one hand, as Figure 2 and Figure 3 shown, an embodiment of the present invention provides an RGB-T semantic segmentation method, including:
[0056] S100. Obtain the RGB image and the infrared image of the scene to be recognized;
[0057] It should be noted that this step further includes: adjusting and unifying the resolutions of the RGB image and the infrared image based on a preset resolution.
[0058] Specifically, by unifying the resolutions of the RGB image and the infrared image, the adaptability of subsequent data processing is ensured.
[0059] S200. Respectively perform feature extraction on the RGB image and the infrared image through an encoder to obtain RGB modality features and infrared modality features;
[0060] Among them, both the RGB modality features and the infrared modality features include multi-level modality features; it should be noted that in some embodiments, it further includes: constructing a ResNet-50 network, removing the fully connected layer and the pooling layer in the ResNet-50 network to obtain a feature extraction network.
[0061] In some embodiments, the encoder includes two symmetric feature extraction networks. Respectively performing feature extraction on the RGB image and the infrared image through the encoder to obtain RGB modality features and infrared modality features includes: respectively performing low-level texture information extraction and high-level semantic classification information extraction on the RGB image and the infrared image through the two feature extraction networks to obtain multi-level RGB modality features and infrared modality features; among them, the feature extraction network is a multi-level structure, and each level includes bottleneck layers of different specifications.
[0062] Specifically, in some specific embodiments, as Figure 3 shown, input the RGB image and the infrared image into two encoders (or two feature extraction networks in one encoder), and extract the corresponding modality features through the encoder network to obtain RGB modality features (including multi-level to ) and infrared modality features (including multi-level to ). Among them, the functions of the encoder (feature extraction network) are: ① Extract the information for semantic segmentation in the image. ② Gradually compress the resolution to obtain a larger receptive field, so that the model can obtain more comprehensive semantic information.
[0063] For the encoder part, the traditional ResNet-50 can be used as the feature extraction network, but with slight modifications. The fully connected layer and pooling layer of ResNet-50 are removed to maintain the resolution of the feature map. The final structure is as shown in Figure 4 . It is divided into 5 stages (STAGE, translated as "stage" here to correspond with the above text). Each stage consists of a different number of Bottlenecks (Bottleneck, namely BTNK1 and BTNK2 in Figure 4 . The key lies in improvement and deletion, and specific parameters will not be elaborated. BN represents the batch normalization layer, CONV represents the convolutional layer, MAXPOOL represents the max pooling layer, and RELU represents the activation function. The specific quantities and parameter settings are shown in Figure 4 ). The features output by each stage have a resolution that is 1 / 2 of the features of the previous stage, but carry more semantic information compared to the features of the previous stage.
[0064] In the semantic segmentation task, ① features at different stages have different meanings. Low-level features (with large resolution, stages 0, 1, 2) contain texture information, and high-level features (with small resolution, stages 3, 4) contain semantic classification information. These features are all beneficial to improving the accuracy of semantic segmentation. ② There are modality differences between RGB images and infrared images. Due to the differences in imaging principles, using the same feature extraction network to process RGB images and infrared images will cause semantic confusion. Therefore, in the embodiments of the present invention, two independent and symmetric ResNet-50s are used as the feature extraction networks to gradually extract multi-level modality features from RGB and infrared images respectively. The modality features of the corresponding stages will be sent to the corresponding probability feature fusion module for fusion.
[0065] S300: Input the RGB modality features and infrared modality features into the probability feature fusion module to calculate the fusion feature preference and obtain the spatial fusion factors for each stage;
[0066] It should be noted that in some embodiments, the RGB modality features and infrared modality features are subjected to intermediate fusion processing corresponding to the stages to obtain multi-level intermediate fusion features; the intermediate fusion processing includes channel dimension connection, first convolutional processing, batch normalization, and first activation processing; based on the Gaussian distribution, the target samples for each stage are determined according to the intermediate fusion features; preference calculation is performed according to the target samples to obtain the spatial fusion factors for each stage; the preference calculation includes second convolutional processing and second activation processing.
[0067] S400: Based on the spatial fusion factors, perform preference fusion on the RGB modality features and infrared modality features to obtain the fusion features for each stage;
[0068] It should be noted that in some embodiments, based on the spatial fusion factors at each level, the fusion weight values of the RGB modality features and the infrared modality features at the corresponding level are set, and weighted fusion is performed to obtain the fusion features at each level; among them, the expression of the fusion feature is:
[0069]
[0070] Among them, represents the fusion feature, W represents the spatial fusion factor, represents the RGB modality feature, represents the infrared modality feature, and i, j, k are all indices.
[0071] Specifically, for steps S300 and S400, in some specific embodiments, first, the modality features are fed into the Probabilistic Feature Fusion Module (PFFM) to generate the spatial fusion factor W, and then the modality features and are fused using the spatial fusion factor W to obtain the fusion feature
[0072] including the following steps: To implement the deep probabilistic feature semantic segmentation model, the fusion feature and are regarded as random variables, and the Probabilistic Feature Fusion Module (PFFM) is designed to fuse and
[0073]
[0074] In the formula, W is the spatial fusion factor, representing the fusion weight value at each spatial position, and the value range is (0, 1). i, j, k are indices. When W = 0, it represents the infrared modality feature; when W = 1, it represents the RGB modality feature. W can affect the preference of the fusion feature. To obtain different fusion features, W ij is regarded as a random variable, and then the key problem is transformed into how to generate the spatial fusion factor W. The model PFFM for generating the spatial fusion factor W is as Figure 5 shown, and the input is the modality features at the same level with the same dimension size and with the dimension size of (c, h, w).
[0075] First, calculate the intermediate fusion feature
[0076]
[0077] This calculation step is connected with in the channel dimension by concatenation (Cat), and the concatenated features are successively passed through a convolution (Conv s to perform the first convolution process), batch normalization (BN, Batchnorm), and a LeakyReLU activation function with a value of 0.2 (σ L to perform the first activation process), and an intermediate fusion feature with a channel dimension of R / r is output, where r is set to 16.
[0078] After that, use to generate variables that follow a Gaussian distribution corresponding mean and log variance, that is:
[0079]
[0080]
[0081]
[0082] The calculation step is to use a 1×1 convolution (Conv 1 ) to generate the mean and the log variance representing a Gaussian distribution with a mean of and a variance of , where d represents the output channel dimension. Then, through sampling, a sample of is obtained (Sample, that is, the target sample) to generate the spatial fusion factor.
[0083]
[0084] Finally, successively pass through a 1×1 convolution (Conv 1 to perform the second convolution process) and a Sigmoid activation function (σ S to perform the second activation process), and after normalization, a spatial fusion factor W with a dimension is obtained. Combining with the formula of the fusion feature, the final fusion feature
[0085] Since each sampling result is different, the generated spatial fusion factor W is also different, and the obtained fusion feature samples also have different preferences (W determines the preference). During the training process of the model, it is necessary to generate the optimal semantic segmentation results for these fusion feature samples that are generated from the same input but have different preferences. Therefore, the model cannot only prefer a certain modality, but needs to make full use of the information of each modality, and the problem of modality preference is solved.
[0086] S500. Input the fusion features of each level into the decoder for decoding to obtain the segmentation result;
[0087] It should be noted that the decoder includes multiple decoding modules. In some embodiments, it includes: inputting the fusion features of each level into each decoding module of the multiple decoding modules for decoding to obtain the segmentation result; wherein, the input data of each decoding module in the multiple decoding modules includes the result after concatenating the output of the previous decoding module and the fusion features of the corresponding level in the channel dimension.
[0088] Specifically, the functions of the decoder include: ① According to the input features, gradually convert the information contained in the features into the category information corresponding to each pixel; ② Upsampling to restore the resolution of the image. In some specific embodiments, Send it into the decoder to obtain the final prediction result, including the following steps:
[0089] For the decoder, the decoding module Upception Block can be used. Upception Block is composed of Transposed Block1 and Transposed Block2 connected in sequence. The structure of Upception Block is as Figure 6 shown. The parameter settings of the two sub-modules Transposed Block1 and Transposed Block2 are as Figure 7 shown. KernelSize is the convolution kernel size, Stride is the stride (sampling interval during convolution), and Padding is the padding (adding a certain number of rows and columns to each side of the input feature map so that the size of the output and input feature maps is the same). After the input features are processed by Upception Block, the decoded features can be obtained
[0090] After the dual-modal features of the encoder are processed by each PFFM, 5 different levels of fusion features can be obtained The deepest fusion feature ( Figure 3 in ) is directly input into Upception Block for decoding to obtain the decoded feature After that, the decoded feature will be concatenated with the fusion features of the corresponding resolution in the channel dimension (Cat) and then input into the next Upception Block. Repeat this step to finally obtain the output segmentation result
[0091] In some specific embodiments, it further includes an evaluation step for the segmentation result, including:
[0092] The experimental data of the embodiments of the present invention are from the RGB-T semantic segmentation public dataset MFNet-Dataset. The dataset consists of 9 types of objects (including the background class), with a total of 1569 pairs of images. Among them, 820 pairs were taken during the day and 749 pairs were taken at night. The lighting conditions are complex and challenging. Example images are as Figure 8 shown. From left to right, they represent the RGB image, the infrared image, and the ideal output result respectively. It can be seen that the infrared image maintains stable performance in different lighting environments. Therefore, during the training process of the model, if no intervention is made, the fused features will tend to use the stable and easy-to-learn infrared image, resulting in insufficient utilization of the RGB image.
[0093] In the semantic segmentation task, the mean accuracy (mAcc) and the mean intersection over union (mIoU) are often used to comprehensively evaluate the effectiveness of a model. The higher these two metrics are, the better the performance. Suppose the model needs to calculate Acc and IoU, where the correctly predicted part is denoted as TP (True Positive); the incorrectly predicted part is denoted as FP (False Positive); the part not predicted is denoted as FN (False Negative). The calculation formulas for Acc and IoU are respectively:
[0094]
[0095]
[0096] mAcc is the average of Acc (Accuracy). A dataset contains multiple classes, and Acc is the accuracy of a certain class (such as the accuracy of pedestrians, the accuracy of cars, etc.). Calculate the Acc of each class respectively, add them up and divide by the number of classes to get mAcc. Similarly, mIoU can be obtained according to the IoU (Intersection over Union) of each class.
[0097] The model proposed in the embodiments of the present invention is verified on the MFNet dataset. The results are shown in Table 1, and the prediction results are as Figure 9 shown. Among them, (a) the original RGB image; (b) the original infrared image; (c) the label (ideal prediction result); (d) the prediction result of the model of the present invention (deep probability model, including the encoder, PFFM, and decoder mentioned above); (e) the prediction result of the original model.
[0098] Table 1
[0099] Model mAcc mIoU Original Model 69.96 55.99 Model of the Present Invention 71.08 56.78
[0100] The original model refers to removing the PFFM (feature fusion module proposed in the introduction to the principle of the deep probability feature model) and directly fusing the RGB features and infrared features in an additive manner. As can be seen from Table 1, the deep probability model proposed in the embodiments of the present invention can effectively improve mAcc and mIoU. Moreover, from Figure 9 the visualization results, since the original model prefers infrared images, warning cones (the pink objects in the first row), curbs (the dark blue objects in the second row), and bicycles (the light blue objects in the second row) that are clearly present in the RGB images cannot be accurately and completely predicted. However, the deep probability model proposed in the embodiments of the present invention can utilize both RGB images and infrared images to accurately predict the results, which shows that the method of the present invention can solve the problem of modality preference.
[0101] In summary, the present invention proposes an RGB-T semantic segmentation method through a spatially adaptive weight fusion method. According to the characteristics of the input image, this method adaptively generates the fusion weights of the two modalities at each pixel in the fusion feature, giving full play to the advantages of each modality. Aiming at the problem that the unique assumption of the fusion feature leads to the fusion result being vulnerable to modality preference interference, the present invention proposes an RGB-T semantic segmentation method. This method breaks the unique assumption of the fusion feature and regards the results of multiple different preferences of the fusion feature as samples of a certain probability distribution, thus transforming the learning problem of the fusion feature into the learning problem of the probability distribution of the fusion feature. Since the model needs to generate the best segmentation results for multiple different preference fusion feature samples during training, the model needs to fully extract and utilize the information of each modality. Therefore, this method can effectively alleviate the problem that the fusion result is vulnerable to modality preference interference. The present invention includes the following beneficial effects: proposing a spatial fusion method, the model can adaptively assign different weights to each pixel point of the feature according to the input image, giving full play to the advantages of the two modality features; proposing the idea of probability fusion features, regarding the fusion feature as a random variable, breaking the uniqueness of the fusion feature. The model needs to generate the best segmentation results for multiple different preference fusion feature samples during training, which forces it to fully extract and utilize the combined information of each modality. Therefore, this method can effectively alleviate the problem that the fusion result is vulnerable to modality preference interference.
[0102] On the other hand, an embodiment of the present invention provides an RGB-T semantic segmentation system, including: a first module for acquiring an RGB image and an infrared image of a scene to be recognized; a second module for respectively performing feature extraction on the RGB image and the infrared image through an encoder to obtain an RGB modality feature and an infrared modality feature; wherein both the RGB modality feature and the infrared modality feature include multi-level modality features; a third module for inputting the RGB modality feature and the infrared modality feature into a probability feature fusion module to calculate fusion feature preferences and obtain spatial fusion factors at each level; a fourth module for performing preference fusion on the RGB modality feature and the infrared modality feature based on the spatial fusion factors to obtain fusion features at each level; and a fifth module for inputting the fusion features at each level into a decoder for decoding processing to obtain a segmentation result.
[0103] The content of the method embodiments of the present invention is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method.
[0104] Another aspect of the embodiments of the present invention further provides an RGB-T semantic segmentation device, including a processor and a memory;
[0105] The memory is used to store a program;
[0106] The processor executes the program to implement the method as described above.
[0107] The content of the method embodiments of the present invention is applicable to the device embodiments of the present invention. The functions specifically implemented by the device embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method.
[0108] Another aspect of the embodiments of the present invention further provides a computer-readable storage medium storing a program, and the program is executed by a processor to implement the method as described above.
[0109] The content of the method embodiments of the present invention is applicable to the computer-readable storage medium embodiments of the present invention. The functions specifically implemented by the computer-readable storage medium embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method.
[0110] The embodiments of the present invention also disclose a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the method as described above.
[0111] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in reverse order. Additionally, the embodiments presented and described in the flowcharts of the present invention are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0112] Furthermore, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features thereof may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those of ordinary skill in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0113] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0114] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution apparatus, apparatus or device (such as a computer-based device, a device including a processor, or other devices that can fetch and execute instructions from the instruction execution apparatus, apparatus or device), or used in combination with these instruction execution apparatus, apparatus or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport a program for use by or in combination with an instruction execution apparatus, apparatus or device.
[0115] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation or, when necessary, other suitable processing, and then stored in a computer memory.
[0116] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution apparatus. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0117] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0118] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
[0119] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. An RGB-T semantic segmentation method, characterized in that, comprising: Obtain the RGB image and the infrared image of the scene to be recognized; Respectively perform feature extraction on the RGB image and the infrared image through an encoder to obtain RGB modality features and infrared modality features; wherein, both the RGB modality features and the infrared modality features include multi-level modality features; Wherein, the encoder includes two symmetric feature extraction networks, and the step of respectively performing feature extraction on the RGB image and the infrared image through the encoder to obtain RGB modality features and infrared modality features includes: Respectively perform low-level texture information extraction and high-level semantic classification information extraction on the RGB image and the infrared image through the two feature extraction networks to obtain multi-level RGB modality features and infrared modality features; Wherein, the feature extraction network is a multi-level structure, and each level includes bottleneck layers of different specifications; Input the RGB modality features and the infrared modality features into a probability feature fusion module to calculate fusion feature preferences, and obtain spatial fusion factors for each level; Wherein, the step of inputting the RGB modality features and the infrared modality features into a probability feature fusion module to calculate fusion feature preferences and obtain spatial fusion factors for each level includes: Perform intermediate fusion processing on the RGB modality features and the infrared modality features in a hierarchical correspondence manner to obtain multi-level intermediate fusion features; the intermediate fusion processing includes channel dimension connection, first convolution processing, batch normalization, and first activation processing; Based on the Gaussian distribution, determine the target samples for each level according to the intermediate fusion features; Perform preference calculation according to the target samples to obtain spatial fusion factors for each level; the preference calculation includes second convolution processing and second activation processing; Based on the spatial fusion factors, perform preference fusion on the RGB modality features and the infrared modality features to obtain fusion features for each level; Input the fusion features of each level into a decoder for decoding processing to obtain a segmentation result.
2. The RGB-T semantic segmentation method according to claim 1, characterized in that, In the step of obtaining the RGB image and the infrared image of the scene to be recognized, it further includes: Based on a preset resolution, adjust and unify the resolutions of the RGB image and the infrared image.
3. The RGB-T semantic segmentation method according to claim 1, characterized in that, It further includes: Construct a ResNet-50 network, remove the fully connected layer and the pooling layer in the ResNet-50 network, To obtain a feature extraction network.
4. The RGB-T semantic segmentation method according to claim 1, characterized in that, The step of performing preference fusion on the RGB modality features and the infrared modality features based on the spatial fusion factors to obtain fusion features for each level includes: Based on the spatial fusion factors of each level, set the fusion weight values of the RGB modality features and the infrared modality features at the corresponding level, and perform weighted fusion to obtain fusion features for each level; Among them, the expression of the fused feature is: Among them, represents the fusion feature, represents the spatial fusion factor, represents the RGB modality feature, represents the infrared modality feature, All are indices.
5. A method for RGB-T semantic segmentation according to claim 1, wherein, the decoder includes multiple decoding modules, and inputting the fused features of each level into the decoder for decoding processing to obtain a segmentation result includes: Inputting the fused features of each level into each layer of the multiple decoding modules in the multiple decoding modules for decoding processing to obtain a segmentation result; Among them, the input data of each decoding module in the multiple decoding modules includes the result of channel dimension connection between the output of the previous decoding module and the fused feature of the corresponding level.
6. An RGB-T semantic segmentation system, wherein, including: A first module for obtaining an RGB image and an infrared image of a scene to be recognized; A second module for respectively extracting features from the RGB image and the infrared image through an encoder to obtain RGB modality features and infrared modality features; among them, both the RGB modality features and the infrared modality features include multi-level modality features; Among them, the encoder includes two symmetric feature extraction networks, and respectively extracting features from the RGB image and the infrared image through the encoder to obtain RGB modality features and infrared modality features includes: Respectively extracting low-level texture information and high-level semantic classification information from the RGB image and the infrared image through the two feature extraction networks to obtain multi-level RGB modality features and infrared modality features; Among them, the feature extraction network is a multi-level structure, and each level includes bottleneck layers of different specifications; A third module for inputting the RGB modality features and the infrared modality features into a probability feature fusion module for calculating fused feature preferences to obtain spatial fusion factors of each level; Among them, inputting the RGB modality features and the infrared modality features into the probability feature fusion module for calculating fused feature preferences to obtain spatial fusion factors of each level includes: Performing intermediate fusion processing on the RGB modality features and the infrared modality features in a hierarchical correspondence manner to obtain multi-level intermediate fusion features; the intermediate fusion processing includes channel dimension connection, first convolution processing, batch normalization, and first activation processing; Based on a Gaussian distribution, determining target samples of each level according to the intermediate fusion features; Performing preference calculation according to the target samples to obtain spatial fusion factors of each level; the preference calculation includes second convolution processing and second activation processing; A fourth module for performing preference fusion on the RGB modality features and the infrared modality features based on the spatial fusion factors to obtain fused features of each level; A fifth module for inputting the fused features of each level into a decoder for decoding processing to obtain a segmentation result.
7. An RGB-T semantic segmentation device includes a processor and a memory; The memory is used for storing programs; The processor executes the program to implement the method according to any one of claims 1 to 5.
8. A computer-readable storage medium, wherein, The storage medium stores a program, and the program is executed by a processor to implement the method according to any one of claims 1 to 5.