RGB-T unregistered image saliency target detection method based on three decoders

Through the RGB-T significance object detection method based on the three-decoder, the problem of low detection accuracy in unregistered image processing and complex scenarios is solved, and efficient modal information fusion and detection accuracy are achieved.

CN120198640APending Publication Date: 2025-06-24SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510267028.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing RGB-T significance object detection methods have low detection accuracy in unregistered image processing and complex scenarios, and there is a gap between the manually registered data sets and the real environment, which affects the detection performance.

Method used

Using a three-decoder-based method, multi-level features are extracted through the feature encoding module, fusion registration module performs semantic mapping and feature registration, modal feature decoding module performs feature decoding, information flow weighted fusion module performs dynamic weighted fusion, and finally supervised and trained through the loss calculation module.

Benefits of technology

Effectively promote the fusion of information in different modes, suppress non-important information, and improve detection accuracy, especially in complex scenarios such as occlusion and multi-objective, which can achieve higher detection performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198640A_ABST
    Figure CN120198640A_ABST
Patent Text Reader

Abstract

The invention relates to a saliency target detection method of an RGB-T image based on cross-modal interaction of three decoders. The detection method structurally comprises a feature extraction module, a fusion complementary registration module, a fusion feature decoding module, a single-modal decoding module and an information flow weighted fusion module. The method comprises the steps that a SwinTransform feature extraction module extracts features of four levels from low to high from an input RGB image and an input T image, and RGB image feature information Fri and T image feature information Fti are obtained; the feature fusion complementary registration module carries out registration fusion on the feature information of the multi-stage RGB and T images to obtain Frti, the fusion feature decoder carries out aggregation on the fusion information to obtain SRT, and the single-mode decoding module carries out aggregation on the obtained RGB and T information to obtain SR and ST; and the three aggregation results are transmitted to an information flow weighted fusion module to obtain a final prediction map P. According to the method, a three-decoder architecture is adopted to fully mine bimodal complementary information, the TCINet network is supervised and trained by calculating the multi-level feature loss value, and the detection performance of a salient target and the robustness of a model in a complex scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly to a saliency object detection method for RGB-T unregistered images based on three decoders. Background Art

[0002] In recent years, with the great progress of deep learning technology and hardware conditions, the saliency object detection algorithm carried on drones has been applied to many industries. The saliency object detection algorithm can quickly segment salient objects and backgrounds in images or videos. As a basic vision algorithm, it has attracted much attention due to its practicality and importance in many fields. Currently, it has been widely applied to visual tasks that require rapid data processing, such as military target detection, topographic surveying, disaster detection, forest management, traffic monitoring, etc.

[0003] Since the height of the drone relative to the ground changes during flight, resulting in a large change in the size of the photographed target, the detection task is more complex than the usual SOD task. Moreover, during the flight of the drone, due to factors such as flight vibration, there are unregistered situations in the images captured by the device. Unregistration means that the positions of the targets in the RGB image and the thermal infrared image cannot be accurately matched on the image. This makes it difficult to establish the connection between the two modalities, thereby affecting the detection performance. Currently, there are some RGB-T saliency object detection methods, but most RGB-T SOD methods use manually registered datasets for training. However, in actual applications, there is a certain gap between the manually registered dataset and the real environment, resulting in the method being unable to effectively fuse the information of the two modalities in the real scene. This makes it difficult for the fused image to capture the key features of the target, while losing the local details of the detected target, thereby affecting the detection accuracy.

[0004] In addition, in the aspect of RGB-T SOD based on drones, due to the insufficient information interaction between different modalities, information loss occurs. And when the modal features are interacting, in the real scene, most methods do not consider the image registration problem, resulting in inconsistent modal features during feature interaction, leading to blurred edges of the detection result and loss of most spatial information. In the feature decoding stage, most researchers use a single encoder for feature decoding, extracting different-scale and different-dimension feature information from the RGB and T images respectively, and then inputting this feature information into a fusion method for feature fusion. Subsequently, the fused feature information is input into the decoder for feature decoding to generate the final prediction map. However, using a simple fusion strategy and a single decoder structure loses some useful information. In extremely challenging scenarios such as occlusion, extreme weather, and multiple targets, the target objects cannot be detected. Summary of the Invention

[0005] Objective of the Invention: The present invention provides a saliency object detection method for RGB-T unregistered images based on three decoders, aiming to solve the problems in the prior art such as the lack of research on unregistered image pairs and the low detection accuracy in complex scenarios. The present invention can well solve the problem of dynamically promoting the fusion between different modality information, suppressing unimportant information, and achieving high accuracy in challenging scenarios such as occlusion and multiple objects.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A saliency object detection method for RGB-T unregistered images based on three decoders, including a feature encoding module, a fusion registration module, a modality feature decoding module, an information flow weighted fusion module, and a loss calculation module;

[0008] The feature encoding module gradually extracts respective multi-level features of the input RGB image and T image through different Stages; the input RGB visible light image and thermal infrared image are processed step by step through feature extraction layers of different scales, so as to obtain multi-scale and multi-level feature representations, providing rich feature descriptions for subsequent feature fusion and object detection.

[0009] The fusion registration module is used to establish a semantic mapping between the RGB image features and T image feature information extracted by the above feature extraction module, generate fusion modality information, enhance the information association between the RGB image features and T image features, and effectively enhance the complementary information between the bimodal features through feature registration and semantic alignment guided by the attention mechanism, overcoming the problem of inconsistent features caused by the difference in imaging mechanisms, and laying a foundation for the subsequent feature decoding process;

[0010] The fusion feature decoding module is used to decode the multi-level fusion features generated by the above fusion registration module, and two single-modal decoding modules are respectively used to aggregate the multi-level RGB image features and T image features extracted by the above feature extraction module to supplement feature information;

[0011] The information flow weighted fusion module adopts an adaptive weighting mechanism to dynamically fuse the fusion feature flow output by the fusion feature decoding module and the RGB visible light feature flow and thermal infrared feature flow output by the two single-modal decoding modules to obtain the final prediction map;

[0012] The loss calculation module constructs a multi-task learning framework through the IOU loss function and binary cross-entropy loss function, and respectively conducts end-to-end supervised training on the information flow generated by the two single-modal decoders, the RGB feature flow, the fusion feature flow generated by the fusion feature decoder, and the final prediction value generated by the information flow weighted fusion module, so as to optimize the feature fusion weight and obtain the loss value;

[0013] The above-mentioned fusion registration module includes the Fusion_R module, the CBAM attention module, and the ASPP module. The Fusion_R module receives RGB visible light features and thermal infrared features of the same scale as input, and obtains more informative fusion features through feature fusion. Subsequently, the CBAM attention module effectively suppresses background interference and highlights the salient features of the target area through adaptive feature recalibration in the spatial and channel dimensions, achieving precise guidance and optimization of the fusion features. Finally, the ASPP module processes the attention-guided fusion features using multi-scale dilated convolutions to obtain registration fusion features with rich semantic information.

[0014] The modality feature decoding module mainly includes two parts: the single modality decoder SMD and the fusion modality decoder FFD. The two single modality decoders SMD respectively perform decoding operations on the RGB and T multi-level features extracted by the feature encoder to generate two single modality prediction maps, the RGB modality prediction map S R and the T modality prediction map S T . The fusion modality decoder FFD aggregates and decodes the multi-level fusion features generated by the fusion registration module to generate a fusion feature information stream;

[0015] The above-mentioned information stream weighted fusion module (WFM), whose input is the features decoded by the three decoding modules, performs weighted fusion on the three information streams, generates corresponding weight maps through the attention mechanism, filters out useless background information, adaptively selects the feature information beneficial to the task, and outputs the final result, thus completing the detection task.

[0016] A saliency object detection method for unregistered RGB-T images based on three decoders

[0017] Step 1: Respectively extract four-level features from low to high for the input RGB image and T image through the feature encoding module;

[0018] Step 2: Implement semantic alignment and information fusion on the multi-level features of the two modalities through the fusion registration module FCR to achieve precise registration of RGB visible light features and thermal infrared features, thereby obtaining fusion features F with rich complementary information and accurate semantic mapping relationships rti ;

[0019] Step 3: Perform decoding operations on the two single modality feature information in Step 1 through the single modality decoder SMD in the modality feature decoder module, and perform feature aggregation on the fusion feature information in Step 2 through the fusion modality decoder FFD to obtain prediction maps S R and S T , as well as prediction map S with similarity information RT .

[0020] Step 4: Dynamically adjust the weights of the non - passing feature maps through the Weighted Fusion Module (WFM) to effectively filter out unreliable information while retaining the useful features in the three - modality information flows. Perform a segmentation operation to generate the final saliency map P.

[0021] Step 5: Calculate the loss values corresponding to the three feature maps and the final prediction result to supervise the training of the entire network.

[0022] Furthermore, the specific operations of the Fusion and Registration Module (FCR) in Step 2 are as follows:

[0023] Step 2.1: Input the RGB feature map F ri and the thermal infrared feature map F ti of the same scale into the Fusion_R module. Through the CBS operation, respectively perform channel recombination and feature enhancement on the two - modality features. Subsequently, fuse F ri with the F ti processed by CBS, and at the same time fuse F ti with the F ri processed by CBS. Finally, perform pixel - level addition on the two fusion results to obtain the initial fusion feature F c of the fusion and registration module;

[0024] Step 2.2: Input the initial fusion feature F c into the CBAM attention module for feature guidance. Obtain the attention - guided feature through adaptive feature recalibration in the spatial and channel dimensions. Subsequently, obtain the enhanced feature F Ai through the CBG operation. Finally, perform multi - scale dilated convolution processing on the feature F Ai using the ASPP module. Stitch and fuse the feature maps obtained with different dilation rates (r = 1, 2, 3, 4), and finally output the registered fusion feature Fr ti with rich semantic information;

[0025] Furthermore, the modality feature decoding module in Step 3 runs the following program:

[0026] Step 3.1: The specific operations of the Fusion - based Feature Decoding Module (FFD) described in Step 3 are as follows:

[0027]

[0028] S RT = Fusion(Fusion(F ti , F rti ), Fusion(F ri , F rti )) i ∈ {1, 2, 3, 4}

[0029] Among them, F pre and F next represent the feature information of two adjacent levels. Fusion represents the fusion operation, Up represents the upsampling operation, and ResB represents the use of residual connections. S RT represents the fused feature information after decoding.

[0030] Step 3.2: The single-modal decoding module (SMD) described in Step 3 inputs the RGB and T features of different scales into two SMDs respectively for feature aggregation, and adopts the strategy of using high-level features to guide low-level features for aggregation. The final saliency maps S R and S T .

[0031] Furthermore, the specific operations of the weighted fusion module WFM in Step 4 are as follows:

[0032]

[0033] Step 4.1: Input the RGB feature map S R , the thermal infrared feature map S T , and the fused feature map S RT into the WFM module, where the fused feature map S RT is first upsampled by a factor of 2;

[0034] Step 4.2: Concatenate the upsampled S RT with S R , S T along the channel dimension, and perform feature extraction through convolution operations;

[0035] Step 4.3: Use the channel attention mechanism CA to enhance the features of S R and S T respectively, and obtain the corresponding weight maps M r and M t through convolution operations and the Sigmoid function; at the same time, perform channel attention guidance and convolution operations on the concatenated features to obtain the weight map M rt of the fused features;

[0036] Step 4.4: Perform weighted operations on the three weight maps with the corresponding feature maps respectively, and sum and fuse the weighted features to obtain the fused feature F f ;

[0037] Step 4.5: Perform CBG operations on the fused feature F f , and then perform feature optimization through residual blocks to obtain the final saliency prediction map S, where the introduction of the residual structure improves the generalization ability of the model.

[0038] Furthermore, in step 5, the BCE loss and IoU loss are used to supervise the network, and the final loss is as follows:

[0039]

[0040] Among them, Loss bce represents Binary Cross-Entropy Loss, that is, binary cross-entropy loss, and Loss IoU represents IoU loss, S i represents the results decoded by the three decoders, GT represents the ground truth map, and P represents the final prediction result.

[0041] Compared with the prior art, the advantages of the present invention are as follows:

[0042] For the RGB-T salient object detection problem, the present invention proposes a network called TCINet. By fully utilizing the differences and complementarities of the two-modal information, it models from different modal perspectives to enhance the connection between modalities.

[0043] The present invention proposes a cross-modal fusion and complementary registration module FCR, which applies the attention mechanism to guide the information fusion between the two modalities, establishes a semantic mapping between the two modalities, and enhances the association between the two modalities. It reduces the situation where the detected object appears blurred.

[0044] The present invention proposes a weighted fusion module WFM to fuse the three different data streams after decoding, so as to reduce the modal differences and improve the robustness of the method in the case of single-modal failure. Thus, the target can be clearly detected in complex scenarios such as occlusion, multi-target, and small target.

[0045] The present invention proposes a three-decoder architecture, and designs two different decoders for single-modal features and fused-modal features respectively, namely FFD and SMD. Among them, SMD is used for single-modal feature decoding, and FFD is used for fused-modal decoding, which greatly improves the generalization ability of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is a schematic structural diagram of the RGB-T image salient object detection method based on three decoders of the present invention;

[0047] Figure 2 is a schematic structural diagram of the fusion registration module FCR of the present invention;

[0048] Figure 3 is a schematic structural diagram of the fused-modal decoding module FFD of the present invention;

[0049] Figure 4 Schematic diagram of the single-modal decoding module SMD of the present invention;

[0050] Figure 5 Result diagram of quantitative comparison of the PR curve of the present invention;

[0051] Figure 6 Result diagram of quantitative comparison of the F-measure curve of the present invention;

[0052] Figure 7 Comparison diagram of the visual comparison experiment of the present invention; Detailed implementation manners

[0053] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant drawings. The drawings show relatively excellent implementation manners of the present application, but the implementation manners of the present application are not limited to only the implementation manners shown in the drawings. These implementation manners are provided to help understand the disclosure content of the present application.

[0054] The present invention proposes a three-decoder feature interaction network. This method simultaneously pays attention to the differences and complementary characteristics between RGB and T images, registers and fuses the RGB and T image features to obtain the fused RGB-T information, effectively solving the problems of insufficient image registration and fusion. Subsequently, the fused features are input into the fusion-modal decoder, and this part of the information contains the relationship mapping between RGB and T images. Subsequently, the extracted RGB and T features are input into the single-modal decoder for feature decoding to fully capture the differential information between the two types of information. To solve the subsequent problem of inputting the fused information and the two different differential information into the weighted fusion module to fully obtain the information in RGB, T, and RGB-T. Generate a final more accurate and clear predicted saliency map.

[0055] Figure 1 Schematic diagram of the saliency object detection method for RGB-T images based on three decoders of the present invention. The saliency object detection method for RGB-T images based on three decoders includes a feature extraction module, a fusion registration module, a feature decoder module, and a weighted fusion module.

[0056] The feature extraction module, as the encoder of the network, is used to extract multi-level feature information of the input RGB and T images.

[0057] Specifically, when implemented, two independent Swin Transformers are used as the backbone of the network to extract the features of RGB and thermal infrared images respectively. The multi-scale features extracted from the RGB image are represented as F from low-level features to high-level semantic representations r1 , F r2 , Fr3 , F r4 , whose resolutions are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image size respectively. The multi-scale features of the same thermal infrared image can be expressed as F t1 , F t2 , F t3 , F t4 .

[0058] Referring to Figure 1 and Figure 2 , the fusion registration module, that is, the FCR part in the figure, establishes mappings respectively between different levels of feature information of the RGB image and the T image for fusion. The fusion registration module includes the Fusion_R, CBAM, and ASPP modules. The CBAM attention mechanism is used to guide the fusion interaction of the two-modal information to suppress the noise in the background and highlight the foreground. The feature maps of the RGB and T images at the same scale are input into the FCR, and after convolution, they are respectively activated by the Sigmoid activation function to obtain the corresponding feature weight maps. Subsequently, we multiply the RGB features and the T features crosswise, perform a pixel-by-pixel addition operation on the two obtained feature maps, and then input them into the CBAM attention. To increase the receptive field, four parallel dilated convolutions with dilation rates of 1, 2, 3, and 4 are constructed, and the four groups of features are combined with the input features from the horizontal dimension. Then, 3×3 convolution is used for feature extraction to obtain the final fused features.

[0059] Continuing to refer to Figure 1 , Figure 3 , Figure 4 . The feature decoding module uses the three-decoder architecture designed by the present invention, which includes two single-modal decoders and one fusion-modal decoder. The inputs of the two single-modal decoders in the figure are the multi-level RGB features F ri (i = 1, 2, 3, 4) and the T features F ti (i = 1, 2, 3, 4) obtained by the two feature encoders. The features are adaptively enhanced through the feature enhancement unit FE, and the strategy of using high-level features to guide low-level features is adopted for feature aggregation. Finally, the saliency maps S R of the RGB modality and the saliency maps S T of the thermal infrared modality are output respectively; the fusion-modal decoder FFD receives the multi-scale fusion features Fr ti (i = 1, 2, 3, 4) output by the fusion registration module, uses the residual structure and the anti-pyramid architecture for feature decoding, and through complementary enhancement with the single-modal features, outputs the saliency prediction map S RT of the fusion modality. This design of the three decoders makes full use of the complementary information of different modalities and improves the detection performance of the model in complex scenarios.

[0060] Step 3: Input the three modal features into the feature decoding module, which contains two single-modal decoders SMD and one fused-modal decoder FFD, to process the feature information of different modalities respectively and perform decoding operations.

[0061] Reference Figure 1 , the weighted fusion module WFM takes the RGB saliency map S output by the single-modal decoder R , the thermal infrared saliency map S T and the fused saliency map S output by the fused-modal decoder RT as inputs, adaptively fuse the three saliency information through a dynamic weight allocation mechanism, and finally output a saliency prediction result with higher accuracy and robustness. The prediction result and the GT map are input into the loss calculation module to obtain the final loss value.

[0062] Based on the above saliency object detection method for RGB-T images based on three decoders, the present invention also provides a saliency object detection method for RGB-T images based on three decoders. The saliency object detection method for RGB-T images based on three decoders includes:

[0063] Step 1: Reference Figure 1 , two feature encoders extract features from the input RGB image and thermal infrared image T. Specifically, two independent Swin Transformers are used as the backbone networks to extract four different-scale features F ri (i = 0, 1, 2, 3) from the RGB image, and their resolutions are 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image respectively; at the same time, corresponding multi-scale features F ti (i = 0, 1, 2, 3) are extracted from the thermal infrared image. This step obtains the global information and multi-scale feature representation of the input image through a deep feature extraction network, laying a foundation for subsequent feature fusion and saliency detection.

[0064] Step 2: Reference Figure 1 and Figure 2 , fuse and register the RGB features and thermal infrared features. Specifically, the same-scale features F ri and F ti extracted in Step 1 are input into the fusion complementary registration module FCR. First, corresponding feature weight maps are obtained through convolution, batch normalization, and Sigmoid operations, and then feature interaction fusion is performed to obtain the initial fusion feature F c . Then, the CBAM attention mechanism is used to optimize the fusion feature at the channel level and spatial level to generate the attention-guided fusion feature F AiFinally, drawing on the idea of the ASSP structure, parallel dilated convolutions with different dilation rates are used to process the features, and the processed features are concatenated with the input features in the horizontal dimension. The final fused feature Fr is extracted through a convolution module. ti This module realizes the adaptive registration and effective fusion of features by dynamically establishing the connection between bimodal features.

[0065] Step 3.1: Refer to Figure 3 , and input the results of each layer in Step 2 into the fused modality decoder FFD, and use the residual structure and the anti-pyramid architecture for feature aggregation and decoding. Specifically, the FFD module takes the RGB feature F ri and the thermal infrared feature F ti as auxiliary features respectively, and performs complementary enhancement with the fused feature Fr ti . During the feature aggregation process, the input features are first upsampled by the Up operation, then feature-encoded through the residual block ResB, and the feature fusion operator is used to achieve the effective fusion of cross-modal features. To reduce the problem of feature dilution during information transmission, the present invention uses transposed convolution for the upsampling operation and reduces the risk of feature information loss through multi-level residual connections. Finally, the fused multi-modal feature S RT is obtained through the feature concatenation operation Concat. The design of this anti-pyramid structure can obtain a finer visual representation compared with the traditional multi-scale processing method, improving the feature retention ability and adaptability of the decoder.

[0066] Continue to refer to Figure 1 , Figure 3 and Figure 4 , the three-decoder module includes the fused modality decoder FFD and two single-modality decoders SMD. For the fused multi-scale features, the figure uses the method of high-level semantics guiding low-level features for feature aggregation, and three residual blocks ResB are used to perform feature encoding operations on the fused features. The features of the RGB and T images are used as auxiliary features respectively to supplement the fused RGB-T features. The overall architecture of the module adopts an inverted pyramid structure to obtain a finer fused feature data stream;

[0067] Step 3.2: Refer to Figure 4, in order to enhance the decoder's ability to capture fine-grained features under complex conditions, in the single-modal decoder SMD, the present invention designs a feature enhancement unit FE. This unit independently processes each channel using depthwise separable convolution and realizes adaptive feature selection and enhancement through element-wise multiplication and addition operations. Specifically, the depthwise convolution DW is used to preserve the unique information of each channel, and the CBG operation is used for feature enhancement. Subsequently, the features are respectively subjected to element-wise multiplication and addition operations to obtain S1 and S2. Finally, the enhanced feature enhanceF is obtained through feature concatenation and convolution operations. This structural design can not only highlight the key features of significant targets but also effectively preserve diverse feature details across channels, and is particularly suitable for dealing with challenges such as motion blur and low target-background contrast in the UAV scenario. In the single-modal decoder, the high-level features guide the aggregation of low-level features after being processed by the FE unit, and the processed feature maps are sequentially concatenated and decoded to finally obtain the saliency map S in the RGB modality R and the saliency map S in the thermal infrared modality T .

[0068] Step 4: Input the three decoded results processed in Step 3 into the weighted fusion module for weighted fusion to obtain the final result.

[0069] Reference Figure 1 , the weighted fusion module WFM is used to perform adaptive weight assignment and fusion on different modal features to obtain more accurate saliency detection results. Specifically, the WFM module first upsamples the fused feature map F rt by a factor of 2, and then concatenates the upsampled feature map with the RGB feature map F r and the thermal infrared feature map F t along the channel dimension. Then, the channel attention mechanism CA is used to enhance the single-modal features and the fused features respectively, and through convolution operations and the Sigmoid function, the corresponding weight maps M r , M t and M rt are obtained. This dynamic weight assignment mechanism can effectively handle the situation of single-modal failure, suppress the interference of unreliable information by adaptively adjusting the weights of different modalities. Finally, the three weight maps are respectively weighted-fused with the corresponding feature maps to obtain the feature F f , and it is optimized through the CBG operation and the residual structure, and the final saliency prediction map S is output, thereby improving the robustness and generalization ability of the model.

[0070] Step 5: Calculate the loss value of the network.

[0071] The loss function of the present invention consists of two single-modal feature losses, a fused-modal feature loss, and a final prediction result loss. For each feature S iFor (i = 1, 2, 3) and the final prediction result P, calculate their binary cross-entropy loss Loss with the ground truth map GT respectively. BCE and the intersection over union loss Loss IoU , and perform a weighted sum of all loss terms to obtain the overall loss function Loss. Thereby, optimize the network parameters and improve the detection accuracy and robustness of the model.

[0072]

[0073] Among them, Loss bce represents the Binary Cross-Entropy Loss, that is, the binary cross-entropy loss, Loss IoU represents the IoU loss, S i represents the results decoded by the three decoders, GT represents the ground truth map, P represents the final prediction result, and Loss represents the total loss value.

[0074] To verify the detection performance of the method of the present invention, the proposed TCINet method of the present invention is compared with 32 methods. Among them, the RGB SOD methods include CorrNet, PFSNet, TRACER, ACCoNet, ERPNet. The RGB-T(D) methods include MROS, MIDD, CGFNet, LSNet, DCNet, TAGF, TNet, RD3D+, RD3D, MGAI, EAEF, and CSRNet, DSNet, BTSNet, SPNet, MCFNet, GRNet, HAINet, BBSNet, CIRNet, and UTANet. The COD methods include: PFNet, SINetV2, FAPNet, SegMaR, BGNet, and BASNet. For fairness, all methods use default parameters and the same training set and test set.

[0075] The dataset used in this experiment is the extremely challenging UAV RGB-T2400 dataset from the perspective of drones, which contains 2400 sets of labeled images, each consisting of registered RGB visible light images and thermal infrared images. This dataset covers a variety of complex scenarios and challenging factors, including different lighting conditions (such as low illumination (LI), extreme low illumination (ELI), street light exposure (SLE), UAV light exposure (ULE), shadow interference (SI), sunlight exposure (SE), and exposure interference (EI)), target feature variations (such as small significant objects (SSO), multiple significant objects (MSOs), center bias (CB), out-of-view (OV), and scale variation (SV)), motion blur (caused by fast object movement (FOM) or fast UAV movement (FVM)), adverse weather (such as rain (R), snow (S), fog (F)), and various occlusion situations (such as tree occlusion (TO), plastic occlusion (PO), umbrella occlusion (UO), and glass occlusion (GO)). These diverse scenarios and challenges provide a comprehensive test benchmark for evaluating the performance of object detection algorithms, making the experimental results more valuable for practical applications and reference.

[0076] The network implementation details provided by the present invention are as follows: The proposed TCINet network model was implemented using PyTorch 1.11.0 and CUDA 11.3, and a computer equipped with a Xeon(R) Silver 4214R CPU and a single NVIDIA RTX 3090 24GB graphics card was used for training and testing. In the data preprocessing stage, all input images were uniformly resized to 384×384, and the parameters of the backbone network were initialized using the Swin Transformer model pre-trained on the ImageNet dataset. During the training process, the Adam optimizer was used, with a batch size of 10, an initial learning rate of 1e-4, and a total of 200 epochs of training. 1200 pairs of images in the UAV RGB-T 2400 were used as the training set to train the network, and the remaining 1200 pairs of images were used as the test set to evaluate the model performance.

[0077] Five widely used metrics in RGB-T saliency detection were adopted in this experiment to evaluate the detection results: the higher the Enhanced Alignment Measure (E-measure), the better, which is used to evaluate the overall alignment quality between the foreground and background of the image; the higher the Structural Similarity Measure (S-measure), the better, which is used to evaluate the structural similarity between the predicted map and the ground truth; the higher the F-measure, the better, which comprehensively considers Precision and Recall to evaluate the overall performance of the model; the lower the Mean Absolute Error (MAE), the better, which is used to measure the pixel-level difference between the predicted map and the ground truth; the larger the area under the Precision-Recall Curve (PR curve), the better, which intuitively shows the detection performance of the model at different thresholds. These evaluation metrics comprehensively reflect the detection effect of the model from different perspectives.

[0078] Quantitative experimental comparison: To verify the effectiveness of the method proposed in the present invention, the proposed method was compared with 32 state-of-the-art salient object detection methods on the UAV RGB-T2400 public dataset. Five evaluation metrics were used in the present invention to comparatively analyze the performance of each method. The experimental results show that the proposed TCINet in the present invention has achieved significant performance improvement on the UAVRGB-T 2400 dataset and has obvious technical advantages. Specifically, the method of the present invention has improved by 1.7% in the S-measure metric compared with the baseline method MROS of the state-of-the-art technology, and the F-measure metric has increased by 1.8%. At the same time, the MROS remains at a comparable level in the E-measure and MAE metrics.

[0079] To further verify the technical effect of the present invention, the present invention also conducted a performance comparison through the PR curve (Precision-Recall curve) and the F-measure curve. As Figure 5 and 6 shown, the proposed TCINet method in the present invention is superior to other methods in the state-of-the-art technology in both of these curve evaluations. The above experimental results fully prove that the method proposed in the present invention has achieved excellent detection performance while maintaining high precision and recall, and has significant technical effects.

[0080] Qualitative experimental comparison: To comprehensively verify the performance of the present invention in dealing with various complex scenarios, the present invention made a visual effect comparison between TCINet and 11 representative methods in the state-of-the-art technology. As Figure 7As shown, when the method of the present invention is dealing with challenging scenarios such as occlusion, small targets, multiple targets, and extreme weather, it can detect a saliency map that is clear, complete, and has a clear boundary, and this effect is significantly better than other existing methods. Specifically, in the five types of challenging scenarios of the UAV RGB-T2400 dataset, the present invention shows significant technical advantages in motion scenarios (fast moving targets and fast moving drones), complex illumination (extremely low illuminance and exposure interference), target features (multiple salient targets and small salient targets), occlusion scenarios (especially in terms of tree occlusion), and under adverse weather conditions. Compared with the baseline method MROS, it has achieved obvious improvements in various evaluation indicators, fully demonstrating the technical advancement and practical value of the present invention.

Claims

1. A method for salient object detection in RGB-T unregistered images based on three decoders, characterized by: include: Feature extraction module, fusion registration module, fusion feature decoding module, single modality decoding module, information flow weighted fusion module; The feature extraction module extracts the respective multi-level features of the input RGB image and T image step by step through different stages; The fusion registration module is used to establish a semantic mapping between the RGB image features extracted by the feature extraction module and the T image feature information, and enhance the information association between the RGB image features and the T image features; The fusion feature decoding module is used to decode the multi-level fusion features generated by the fusion registration module. The two single modality decoding modules are used to perform feature aggregation on the multi-level RGB image features and T image features extracted by the feature extraction module to supplement feature information. The information flow weighted fusion module is used to perform weighted fusion on the fused feature information flow obtained by the above-mentioned fusion feature decoding module, the RGB feature information flow obtained by the two single modality decoding modules, and the T feature information flow to obtain the final prediction value, and use IOU Loss and BCE Loss to perform supervised training on the RGB information flow, T information flow and fused information flow to obtain the loss value.

2. The method for salient object detection in RGB-T unregistered images based on three decoders according to claim 1, characterized in that: The above-mentioned fusion registration module includes Fusion_R, CBAM and ASPP, among which the input of Fusion_R is the RGB and T image features of the same scale, and the fusion features with richer feature information are obtained. The obtained results are input into CBAM to suppress the background information and highlight the foreground, and the fusion feature information after guided fusion is obtained. Finally, the fusion information is processed by ASPP to obtain the final registered fusion feature information.

3. The method for salient object detection in RGB-T unregistered images based on three decoders according to claim 1, characterized in that: The fusion feature decoding module (FFD) decodes the fusion features to obtain the aggregated mixed modal features, S RT Two single modality decoding modules (SMD) decode the four scales of RGB and T respectively to obtain a single modality prediction map S R and S T .

4. The method for salient object detection in RGB-T unregistered images based on three decoders according to claim 1, characterized in that: The above-mentioned information flow weighted fusion module (WFM) takes as input the features decoded by the three decoding modules, performs weighted fusion on the three information flows, generates a corresponding weight map through the attention mechanism, filters out useless background information, adaptively selects feature information that is beneficial to the task, and outputs the final result, thereby completing the detection task.

5. A method for salient object detection in RGB-T unregistered images based on three decoders as claimed in claim 1, characterized in that: Step 1: Use the feature extraction module to extract four levels of features F from low to high for the input RGB image and T image ri ,F ti ; Step 2: The features of the two modalities are fused and registered through the fusion registration module to obtain the fusion feature F with rich similarity information. rti ; Step 3: The differential feature information in step 1 and the similarity fusion feature information in step 2 are aggregated through the three decoder modules to obtain an information stream S with differential feature information. R and S T , and S with similarity information RT . Step 4: Dynamically adjust the weights assigned to different information flows through the information flow weighted fusion module, thereby effectively filtering out unreliable information while retaining the useful features in the three modal information flows. Perform segmentation operations to generate the final saliency map P. Step 5: Calculate the loss values ​​corresponding to the three information flows to supervise the entire network.

6. The method for salient object detection in RGB-T unregistered images based on three decoders according to claim 5, characterized in that: The specific method of step 2 is: Step 2.1: Input the feature maps of the RGB and T images of the same scale into Fusion_R for registration and fusion, and obtain the first step output F of the fusion registration module. Ai ; Step 2.2: The CBAM structure is used to further guide features, and the ASPP module is used to obtain the final output F of the fusion registration module. rti ,i=(1,2,3,4).

7. The method for salient object detection in RGB-T unregistered images based on three decoders according to claim 5, characterized in that: The specific method of step 3 is: Step 3.1: The fused modality decoding module (FFD) described in step 3 uses the features of RGB and T images as auxiliary features to supplement the fused RGBT features. The overall architecture of the module adopts an inverted pyramid structure to obtain a more refined fused feature data stream. The specific operations are as follows: S RT =Fusion(Fusion(F ti ,F rti ),Fusion(F ri ,F rti ))i∈{1,2,3,4} where F pre and F next Represents the feature information of two adjacent levels, Up represents the upsampling operation, and ResB represents the use of residual connection. RT Represents the decoded fusion feature information. Step 3.2: The single modality decoding module (SMD) described in step 3 inputs the RGB and T features of different scales into two SMDs for feature aggregation, and adopts the strategy of high-level features guiding low-level features for aggregation. The final saliency maps S of the two single modalities are obtained. R With S T .

8. The method for salient object detection in RGB-T images based on three decoders according to claim 5, characterized in that: The specific method of step 4 is: The three weight graphs M obtained by the information flow weighted fusion module r 、M t and M rt , the specific process of obtaining the weight map is as follows: Among them, σ represents Sigmoid and CA represents the channel attention mechanism. The weight map maps its own features into important feature information that is helpful for the task. Finally, the three types of effective information are added and merged, and the key features of the merged features are extracted to generate the final saliency map.

9. The method for detecting salient objects in RGB-T images based on three decoders according to claim 5, characterized in that: In step 5, BCE loss and IOU loss are used to supervise the network, and the final loss is as follows: Among them, Loss_bce represents Binary Cross-Entropy Loss, Loss_IOU represents IoU loss, S_i, i = (1, 2, 3) represents the results after decoding by three decoders, GT represents the true value map, and P represents the final prediction result.