RGB-T saliency target detection method based on query guide specific learning network

By combining query-guided specific learning networks and deep separable convolutions, the complexity and redundancy issues of multimodal feature fusion in RGB-T saliency target detection are resolved, achieving efficient multimodal information fusion and target detection, and improving the model's performance in complex environments.

CN120953559APending Publication Date: 2025-11-14TONGJI UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511229024.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing RGB-T salient target detection methods suffer from problems such as high encoder complexity, feature redundancy, and difficulty in cross-modal alignment during multimodal feature fusion. They also struggle to effectively utilize the complementary information from infrared and visible light images, resulting in poor detection performance.

Method used

We employ a vision state space block based on Vision Mamba and a selective scanning mechanism. By using a query-guided specific learning network, we dynamically guide the fusion of cross-modal information during the decoding stage. Combined with deep separable convolution and attention mechanisms, we achieve efficient fusion and interaction of multimodal features.

Benefits of technology

It improves the accuracy and efficiency of multimodal information fusion, enhances the model's target recognition accuracy and robustness in complex environments, reduces computational costs, and is suitable for edge devices with limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953559A_ABST
    Figure CN120953559A_ABST
Patent Text Reader

Abstract

The invention relates to an RGB-T saliency target detection method based on a query guide specific learning network, and the method comprises the steps: constructing a network model of an encoder-decoder architecture, and setting a visible light image saliency detection module, an infrared image saliency detection module and a bimodal saliency detection module according to the network model, visual features of the visible light RGB image are obtained through an encoder of the visible light image saliency detection module, and infrared features of the infrared thermal imaging image are obtained through the infrared image saliency detection module; in a bimodal information fusion module in the bimodal saliency detection module, generating an RGB modal feature, an infrared modal feature and a cross-modal fusion feature; and the three-feature cross block outputs the enhanced multi-modal fusion features. Compared with the prior art, the method has the advantages of high accuracy, high modal adaptability, high efficiency and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to train signal control systems, and more particularly to an RGB-T saliency target detection method based on a query-guided specific learning network. Background Technology

[0002] In today's information-saturated world, humans are bombarded with massive amounts of visual information every day. To process this information efficiently, the human visual system has evolved a saliency mechanism, enabling it to quickly filter out the most valuable parts of a scene and ignore irrelevant details, thus helping us better understand our surroundings and make decisions. With the rapid development of computer vision technology, there is a growing desire to endow computers with similar capabilities, enabling them to accurately locate and identify key targets in complex scenes. This would not only improve the efficiency of algorithms but also play a significant role in numerous practical applications, such as autonomous driving, security monitoring, and medical image analysis.

[0003] In past research, salient object detection has primarily relied on various manually designed prior knowledge, such as contrast, center prior, edge prior, and semantic prior. While these prior knowledge can guide algorithms to extract salient information to some extent, they often fall short when faced with complex and varied real-world scenarios. This is because these prior knowledge are mostly based on low-level visual features, such as color, brightness, and contrast, making it difficult to capture the essential features of salient objects and adapting to diverse scenes and targets.

[0004] In recent years, the rapid development of deep learning technology has brought revolutionary changes to the field of computer vision, especially neural network-based architectures, which can automatically learn complex feature representations from massive amounts of data, greatly improving the generalization ability and accuracy of models. With advancements in sensor technology, particularly the widespread application of thermal infrared sensors, we can now acquire not only traditional RGB images but also their corresponding thermal infrared images. This is crucial in many scenarios, such as pedestrian re-identification and object tracking under nighttime or low-light conditions. Thermal infrared images can compensate for the shortcomings of RGB images, helping to more accurately detect and identify targets.

[0005] In the RGB-T saliency target detection task, fully utilizing the complementary information of visible light (RGB) and infrared (thermal) images is crucial for improving detection performance. Existing methods mostly focus on multimodal feature fusion at the encoder stage, such as direct deep feature fusion, modality-assisted fusion, encoder-based fusion, and encoder-decoder-based alignment strategies. However, these methods have the following limitations: high encoder fusion complexity—complex feature alignment modules increase the number of model parameters, leading to higher computational costs, and they struggle to fully utilize the multimodal features of pre-trained encoders, potentially causing overfitting; modality feature uncertainty—pre-trained encoders have varying feature extraction capabilities for different modalities, resulting in uncertainty and variability in encoder output features, which traditional fusion methods cannot effectively handle; and lack of decoder-specific information—existing methods neglect the preservation and utilization of modality-specific information at the decoder stage, leading to feature redundancy and an inability to specifically learn task-related key information. Therefore, how to maintain model lightweighting while efficiently utilizing the modality-specific information of RGB and infrared images to improve saliency detection performance under model uncertainty is a technical problem that needs to be solved. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the existing technology by providing an RGB-T saliency target detection method based on a query-guided specific learning network. It uses a vision state space block based on Vision Mamba as the core processing unit, and combines a selective scanning mechanism with a state space model. This method can effectively capture long-range dependencies and adapt to the uncertainty of different modal features. By converting modal features into query information, it dynamically guides cross-modal information fusion during the decoding stage, thereby improving the accuracy and efficiency of multimodal information fusion.

[0007] The objective of this invention can be achieved through the following technical solutions:

[0008] According to one aspect of the present invention, an RGB-T saliency target detection method based on a query-guided specificity learning network is provided, the specific steps of which include:

[0009] S1. Acquire visible light RGB images and their corresponding infrared thermal imaging images;

[0010] S2. Construct a network model with an encoder-decoder architecture, and set up a visible light image saliency detection module, an infrared image saliency detection module, and a dual-modal saliency detection module according to the network model. The visible light RGB image is processed by the encoder of the visible light image saliency detection module to obtain visual features, and the infrared thermal image is processed by the infrared image saliency detection module to obtain infrared features.

[0011] S3. Input the visual features and infrared features into the dual-modal information fusion module in the dual-modal saliency detection module, perform layer normalization processing, depthwise separable convolution operation and three-branch attention calculation to generate RGB modal features, infrared modal features and cross-modal fusion features;

[0012] S4. Input the RGB modal features, infrared modal features, and cross-modal fusion features into the three-feature cross block in the dual-modal saliency detection module, and output the enhanced multimodal fusion features; input the multimodal fusion features into the fusion decoder, which consists of three cascaded query-guided decoding blocks, and outputs the final saliency target detection result image after step-by-step upsampling processing.

[0013] Furthermore, the network model of the encoder-decoder architecture is a visual state space block based on VisionMamba; each visual state space block includes an input layer, an output layer, and a hidden layer, and the specific steps for processing data include:

[0014] The input image is segmented into multiple image patches, and each image patch is mapped to a one-dimensional vector of length D to obtain the input vector. The input vector is then sequentially input into the visual state space block. The hidden layer sequentially performs the following processing on the input vector: linear transformation through a first linear projection layer; feature extraction through a depthwise separable convolutional layer; nonlinear activation through a first SiLU activation function layer; scanning expansion, merging, and selective state space modeling of the feature sequence through a selective scanning module to obtain intermediate features; layer normalization, second linear projection, and second SiLU activation function processing are sequentially applied to the intermediate features; the output of the second SiLU activation function and the output of the first linear projection layer are multiplied element-wise to obtain the core operation output; the core operation output is then subjected to a third linear projection and added element-wise to the original input vector to obtain the final feature output vector of the visual state space block.

[0015] Furthermore, the visible light image saliency detection module includes four cascaded visible light image encoder groups, four visible light image decoder groups, and a visible light image saliency detection output layer. The output of the first visible light image encoder serves as the input of the fourth visible light image decoder, the output of the second visible light image encoder serves as the input of the third visible light image decoder, the output of the third visible light image encoder serves as the input of the second visible light image decoder, and the output of the fourth visible light image encoder serves as the input of the first visible light image decoder. The output of the second visible light image decoder serves as the input of the first query response decoder in the dual-modal saliency detection module, the output of the third visible light image decoder serves as the input of the second query response decoder in the dual-modal saliency detection module, and the output of the fourth visible light image decoder serves as the input of the third query response decoder in the dual-modal saliency detection module.

[0016] Furthermore, both the visible light image encoder and the visible light image decoder include two visual state space blocks connected end-to-end. Each visual state space block specifically includes a first residual connection layer, a first normalization layer, a second residual connection layer, a first fully connected neural network layer, a first depthwise separable convolutional layer, a first selective scanning module, a second normalization layer, and a second fully connected neural network layer. The first residual connection layer adds the input feature map to the output of the other visual state space block. After passing through a fully connected neural network layer and an activation function layer, the second residual layer performs a Hada code product operation with the result obtained from the second normalization layer.

[0017] Further, the width and height of the output feature map of the first visible light image encoder are one-quarter of the initial visible light RGB image; the width and height of the output feature map of the second visible light image encoder are one-eighth of the initial visible light RGB image; the width and height of the output feature map of the third visible light image encoder are one-sixteenth of the initial visible light RGB image; the width and height of the output feature map of the fourth visible light image encoder are one-thirty-second of the initial visible light RGB image; the width and height of the output feature map of the first visible light image decoder are one-sixteenth of the initial visible light RGB image; the width and height of the output feature map of the second visible light image decoder are one-eighth of the initial visible light RGB image; the width and height of the output feature map of the third visible light image decoder are one-quarter of the initial visible light RGB image; and the width and height of the output feature map of the fourth visible light image decoder are consistent with the initial visible light RGB image.

[0018] Furthermore, the structure of the infrared image saliency detection module is consistent with that of the visible light image saliency detection module.

[0019] Further, the specific steps of the dual-modal saliency detection module in S3 include: inputting the visual features output by the fourth visible light image encoder and the infrared features output by the fourth infrared image encoder after normalization of the receiving layer into the depthwise separable convolutional layer; converting the visual features and infrared features into query matrices, feature key matrices and feature value matrices respectively according to the attention mechanism to obtain self-attention features based on visible light modality, cross-attention features based on visible light key-value pairs and infrared queries, and self-attention features based on infrared modality; and obtaining RGB modal features, infrared modal features and cross-modal fusion features through the hidden layer.

[0020] Furthermore, the three-feature cross block in S4 includes an input layer, a hidden layer, and an output layer; wherein the operation of the hidden layer includes: performing layer normalization on the cross-modal fusion features and RGB modal features respectively, and inputting them into the Swing Transformer network layer. Local window attention calculation is performed, and the result is added element-wise to the cross-modal fusion feature to obtain a first temporary feature. The first temporary feature is then input into the forward propagation network layer and added to itself to obtain the feature fused with RGB information. The cross-modal fusion feature and the infrared modal feature are respectively layer normalized and input into the Swing Transformer network layer C for local window attention calculation. The result is then added element-wise to the cross-modal fusion feature to obtain a second temporary feature. The second temporary feature is then input into the forward propagation network layer and added to itself to obtain the feature fused with infrared information. The enhanced multimodal fusion feature includes the feature fused with RGB information and the feature fused with infrared information.

[0021] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.

[0022] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] (1) Improve the accuracy and efficiency of multimodal feature fusion, thereby enhancing the accuracy of salient target detection: By adopting a dual-modal information fusion module and a three-feature cross-block structure, the deep fusion and interaction of visible light and infrared image features are realized in the decoding stage, effectively utilizing the complementary information of the two modes, solving the problems of feature redundancy and cross-modal alignment difficulties in the existing technology, thereby improving the accuracy and reliability of target recognition in complex environments.

[0025] (2) Enhance the model’s ability to adapt to complex scenarios and uncertainties, thereby improving the system’s robustness: The application of the visual state space block and selective scanning mechanism based on Vision Mamba enables the model to effectively model long-range dependencies and adaptively handle the feature uncertainties of different modalities, improving the model’s stable performance under complex conditions such as lighting changes, occlusion and multiple background interferences. Finally, in practical application scenarios such as night monitoring and target detection in severe weather, it shows stronger environmental adaptability and system robustness.

[0026] (3) Optimize the model computation structure and parameter efficiency to improve deployment feasibility and application breadth: By combining the multi-level feature fusion mechanism and lightweight components, such as depthwise separable convolution, in the encoder-decoder framework, the model complexity and computation cost are effectively controlled while ensuring detection performance, reducing hardware resource requirements and inference latency. This allows the technology of this invention to be deployed more efficiently on edge devices with limited computing resources, such as vehicle systems and embedded monitoring devices, thus broadening its practical application in real-time vision systems. Attached Figure Description

[0027] Figure 1 This is a data flow diagram for the RGB-T saliency target detection method based on a query-guided specificity learning network.

[0028] Figure 2 A flowchart for processing image features for visual state space blocks;

[0029] Figure 3 Flowchart for image feature processing by the dual-modal information fusion module;

[0030] Figure 4 Flowchart for processing image features using three-feature cross blocks;

[0031] Figure 5 This is a comparison chart of the results of this embodiment and existing saliency detection methods;

[0032] Figure 6 This is a comparison chart showing the ability of this embodiment to extract object structure with existing saliency detection methods. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0034] like Figure 1The diagram shown is a data flow diagram of an RGB-T saliency target detection method based on a query-guided specificity learning network. The specific steps include:

[0035] S1. Acquire visible light RGB images and their corresponding infrared thermal imaging images;

[0036] S2. Construct a network model with an encoder-decoder architecture, and set up a visible light image saliency detection module, an infrared image saliency detection module, and a dual-modal saliency detection module according to the network model. The visible light RGB image is processed by the encoder of the visible light image saliency detection module to obtain visual features, and the infrared thermal image is processed by the infrared image saliency detection module to obtain infrared features.

[0037] S3. Input the visual features and infrared features into the dual-modal information fusion module in the dual-modal saliency detection module, perform layer normalization processing, depthwise separable convolution operation and three-branch attention calculation to generate RGB modal features, infrared modal features and cross-modal fusion features;

[0038] S4. Input the RGB modal features, infrared modal features, and cross-modal fusion features into the three-feature cross block in the dual-modal saliency detection module, and output the enhanced multimodal fusion features; input the multimodal fusion features into the fusion decoder, which consists of three cascaded query-guided decoding blocks. After step-by-step upsampling processing, the final saliency target detection result map is output.

[0039] During training, Q original color object images and their corresponding infrared thermal images, along with true saliency detection label images, are selected to form the training set. For each object q in the training set, the color object image, infrared thermal image, and saliency detection label image are represented as Visual... q (i,j),Thermal q (i,j) and Out q (i,j). Each image has a width of Width and a height of Height. The coordinates i,j of each pixel in each image of object q satisfy 1≤i≤Width, 1≤j≤Height. Furthermore, during each training iteration, a batch of objects from the training set Q are extracted and trained simultaneously to fully utilize GPU memory and accelerate the training process.

[0040] The encoder-decoder architecture network model is based on Vision Mamba visual state space blocks. By encoding the input image into feature representations and then decoding them into the desired output, efficient feature extraction and complex pattern recognition are achieved. The resolution of the feature map obtained by the image encoder gradually decreases with increasing depth, while the resolution of the feature map obtained by the image decoder gradually increases with increasing depth, enabling the model to learn multi-scale features of the image.

[0041] like Figure 2 The diagram shows a flowchart of image feature processing using visual state space blocks. Each visual state space block includes an input layer, an output layer, and a hidden layer. For a color object image of input object q, Visual... q (i,j) and infrared thermal imaging images Thermal q (i,j), after dividing it into several smaller blocks, is mapped to a one-dimensional vector f of length D. in For each visual state space block, its input layer receives a vector f. in The output layer obtains vector f out The hidden layers of the visual state space block consist of two normalization layers Γ, three linear projection layers l, and one depthwise separable convolutional layer. Two SiLU activation function layers With 1 selectable scan module The selection scanning module comes from the Vision Mamba model and mainly includes scan expansion operations, merging, and selective state space models.

[0042] The specific steps for processing the data include: segmenting the input image into multiple image blocks and mapping each image block to a one-dimensional vector of length D to obtain the input vector; inputting the input vector sequentially into the visual state space block, and the hidden layer sequentially performing the following processing on the input vector: performing linear transformation through the first linear projection layer; performing feature extraction through the depthwise separable convolutional layer; performing nonlinear activation through the first SiLU activation function layer; performing scanning expansion, merging, and selective state space modeling on the feature sequence through a selective scanning module to obtain intermediate features; performing layer normalization, second linear projection, and second SiLU activation function processing on the intermediate features sequentially; multiplying the output of the second SiLU activation function and the output of the first linear projection layer element-wise to obtain the core operation output; and adding the core operation output element-wise with the original input vector after the third linear projection to obtain the final feature output vector of the visual state space block.

[0043] Because of the use of skip connections, the expression for the processing of each visual state space block is as follows:

[0044] f out =f in ⊕l(Θ(Γ(f in ))),

[0045] Let the input vector accepted by the core computation part Θ of this module be... The output vector is The calculation process is expressed as follows:

[0046]

[0047] In this module, ⊕ and ⊙ represent element-wise addition and multiplication of two matrices, respectively, and Θ is the core operation part.

[0048] The visible light image saliency detection module includes four cascaded visible light image encoder groups, four visible light image decoder groups, and a visible light image saliency detection output layer. The output of the first visible light image encoder serves as the input of the fourth visible light image decoder, the output of the second visible light image encoder serves as the input of the third visible light image decoder, the output of the third visible light image encoder serves as the input of the second visible light image decoder, and the output of the fourth visible light image encoder serves as the input of the first visible light image decoder. The output of the second visible light image decoder serves as the input of the first query response decoder in the bimodal saliency detection module, the output of the third visible light image decoder serves as the input of the second query response decoder in the bimodal saliency detection module, and the output of the fourth visible light image decoder serves as the input of the third query response decoder in the bimodal saliency detection module.

[0049] Both the visible light image encoder and the visible light image decoder include two visual state space blocks connected end-to-end. Each visual state space block specifically includes a first residual connection layer, a first normalization layer, a second residual connection layer, a first fully connected neural network layer, a first depthwise separable convolutional layer, a first selective scanning module, a second normalization layer, and a second fully connected neural network layer. The first residual connection layer adds the input feature map to the output of the other visual state space block. After the second residual layer passes through a fully connected neural network layer and an activation function layer, it performs a Hada code product operation with the result obtained from the second normalization layer.

[0050] For a visible light image encoder, its input accepts batches of color RGB images for training, and its output is the encoded visual features f. V , representing the low-level features of the image; for an infrared thermal image encoder, its input receives the corresponding batch of infrared thermal images used for training, and its output is the encoded infrared feature f. T , representing the temperature distribution of the image.

[0051] For the first-layer visible light image encoder, its input accepts a batch of visual images. q The vector f mapped from (i,j) in The output is a batch with a width of [value missing]. Height is The feature map has 3 channels, and the set of all output feature maps is denoted as VE1;

[0052] For the second-layer visible light image encoder, its input accepts a set of Batch feature maps VE1, and its output is a Batch width of [missing value]. Height is The feature map has 3 channels, and the set of all output feature maps is denoted as VE2;

[0053] For the third-layer visible light image encoder, its input accepts a batch of feature maps VE2, and its output is a batch with a width of [missing information]. Height is The feature map has 3 channels, and the set of all output feature maps is denoted as VE3;

[0054] For the fourth layer visible light image encoder, its input accepts a batch of feature maps VE3, and its output is a batch with a width of [missing information]. Height is A feature map with 3 channels is used, and the set of all output feature maps is denoted as f. V .

[0055] For a visible light image decoder, its input receives a set of encoded RGB feature maps f used for batch training. V Output RGB saliency test map V For an infrared thermal imaging decoder, its input receives the corresponding batch of encoded infrared thermal imaging feature maps f used for training. T Output infrared thermal imaging saliency detection map T .

[0056] For the first-layer visible light image decoder, its input receives a set of Batch encoded RGB feature maps f. V The output is a batch with a width of [value missing]. Height is The feature map has 3 channels, and the set of all output feature maps is denoted as VD1;

[0057] For the second-layer visible light image decoder, its input accepts a batch of feature maps VD1, and its output is a batch with a width of [missing information]. Height is The feature map has 3 channels, and the set of all output feature maps is denoted as VD2;

[0058] For the third-layer visible light image decoder, its input accepts a batch of feature maps VD2, and its output is a batch with a width of [missing information]. Height is The feature map has 3 channels, and the set of all output feature maps is denoted as VD3;

[0059] For the fourth layer of the visible light image decoder, its input accepts a batch of feature maps VD3, and its output is a batch of feature maps with a width of W, a height of H, and 3 channels. After passing through a fully connected layer, it outputs an RGB saliency detection map. V .

[0060] The structure of the infrared image saliency detection module is consistent with that of the visible light image saliency detection module. For the first-layer infrared image encoder, its input accepts a batch of thermal... q The vector f mapped from (i,j) in The output is a batch with a width of [value missing]. Height is The feature map with 1 channel is used, and the set of all output feature maps is denoted as TE1.

[0061] For the second-layer infrared image encoder, its input receives a set of Batch feature maps TE1, and its output is a Batch image with a width of [missing value]. Height is The feature map with 1 channel is denoted as TE2, and the set of all output feature maps is denoted as TE2.

[0062] For the third-layer infrared image encoder, its input accepts a batch of feature maps TE2, and its output is a batch with a width of [missing information]. Height is The feature map with 1 channel is denoted as TE3, and the set of all output feature maps is denoted as TE3.

[0063] For the fourth-layer infrared image encoder, its input accepts a batch of feature maps TE3, and its output is a batch with a width of [missing information]. Height is The feature map has 1 channel, and the set of all output feature maps is denoted as f. T .

[0064] For the first-layer infrared image decoder, its input receives a set f of Batch encoded infrared thermal imaging feature maps. T The output is a batch with a width of [value missing]. Height is The feature map with 1 channel is denoted as TD1, and the set of all output feature maps is denoted as TD1.

[0065] For the second-layer infrared image decoder, its input accepts a batch of feature map sets TD1, and its output is a batch with a width of [missing information]. Height is The feature map with 1 channel is denoted as TD2, and the set of all output feature maps is denoted as TD2.

[0066] For the third-layer infrared image decoder, its input accepts a batch of feature map sets TD2, and its output is a batch with a width of [missing information]. Height is The feature map with 1 channel is denoted as TD3, and the set of all output feature maps is denoted as TD3.

[0067] For the fourth layer infrared image decoder, its input receives a batch of feature map sets TD3, and its output is a batch of feature maps with a width of W, a height of H, and 1 channel. After passing through a fully connected layer, it outputs an RGB saliency detection map. T .

[0068] A dual-modal information fusion module is constructed. Addressing the current research focus on multimodal feature fusion at the encoder stage, most existing methods concentrate on this approach. This module is primarily used to fuse visual and thermal imaging features in deeper neural networks. Each module comprises an input layer, a hidden layer, and an output layer. The input layer receives a set of RGB feature maps f after layer normalization. V and infrared thermal imaging feature map set f T It is divided into three branches: skip connection branch (RGB feature map and infrared thermal imaging feature map are respectively represented as f) v and f t The initial feature fusion branch f′ is used, and the remaining vectors are input into the depthwise separable convolutional layer in the hidden layer. In addition, this module outputs three sets of feature maps, namely RGB feature f 1 RGB-infrared thermal imaging fusion feature f and infrared thermal imaging feature f 2 The hidden layer of the bimodal cross module consists of one depthwise separable convolutional layer. One three-branch attention mechanism layer, two normalization layers (Γ), and one forward propagation network layer.

[0069] like Figure 3 As shown, the specific steps of the bimodal saliency detection module in S3 include: inputting the visual features output by the fourth visible light image encoder and the infrared features output by the fourth infrared image encoder after normalization of the receiver layer into the depthwise separable convolutional layer; converting the visual features and infrared features into query matrices, feature key matrices and feature value matrices respectively according to the attention mechanism to obtain self-attention features based on visible light modality, cross-attention features based on visible light key-value pairs and infrared queries, and self-attention features based on infrared modality; and obtaining RGB modal features, infrared modal features and cross-modal fusion features through the hidden layer.

[0070] f V with f T After layer normalization, the input is fed into a depthwise separable convolutional layer. Based on the attention mechanism, the RGB visual features are converted into a feature query matrix Q. v eigenkey matrix K v With the eigenvalue matrix V v Convert infrared thermal imaging features into a feature query matrix Q. t eigenkey matrix K t With the eigenvalue matrix V t Expressed as a formula:

[0071]

[0072] Next, the attention matrix obtained from the above operations is added to the corresponding skip connection branch and fed into the feature mixing branch. This branch is based on a cross-attention structure, where the key and value matrices are derived from infrared thermal imaging features, and the query matrix is ​​derived from RGB visual features, to handle the relationship between the two modalities. Let... For attention calculation mechanisms, visual features after attention calculation Visual-infrared fusion feature f in infrared features The expression is:

[0073]

[0074] f in and After passing through other hidden layers, the final output is obtained: RGB feature f 1 RGB-infrared thermal imaging fusion feature f and infrared thermal imaging feature f 2 The calculation method is as follows:

[0075]

[0076] like Figure 4 The diagram shows the flowchart for image feature processing using a three-feature cross-block. The three-feature cross-block in S4 includes an input layer, a hidden layer, and an output layer. Each input layer receives RGB features f obtained from the bimodal cross-block. 1 RGB-infrared thermal imaging fusion feature f and infrared thermal imaging feature f 2 Each output layer outputs the fused features of the three. The operation of the hidden layer is divided into two steps. The first step is used to fuse RGB information, including a normalization layer Γ, a forward propagation network layer F, and a Swing Transformer network layer C. The second step is used to fuse thermal imaging information.

[0077] In the hidden layer, the cross-modal fusion features and RGB modal features are first normalized by the layer and then input into the SwingTransformer network layer. Local window attention calculation is performed, and the result is added element-wise to the cross-modal fusion feature to obtain the first temporary feature. The first temporary feature is then input into the forward propagation network layer and added to itself to obtain the feature fused with RGB information. The cross-modal fusion feature and the infrared modal feature are respectively layer normalized and input into the SwinTransformer network layer C for local window attention calculation. The result is then added element-wise to the cross-modal fusion feature to obtain the second temporary feature. The second temporary feature is then input into the forward propagation network layer and added to itself to obtain the feature fused with infrared information. The enhanced multimodal fusion feature includes the feature fused with RGB information and the feature fused with infrared information.

[0078] Let the RGB information feature be f temp1 The first step is to output out. part1 The calculation formula is as follows:

[0079]

[0080] The second step is used to fuse thermal imaging information, and the process is similar to the first step. Let the infrared thermal imaging information feature be f. temp2 The first step is to output out. part2 The calculation formula is as follows:

[0081]

[0082] A fusion decoder is constructed to perform subsequent processing on the information obtained from the three-feature cross-block. The fusion decoder consists of three query-guided decoder blocks (QDBs), and each QDB contains a different number of TFC structures. Here, the number of TFCs is set to [2, 2, 9].

[0083] During the training of the model used in this embodiment, each branch uses the sum of weighted binary cross-entropy (w-BCE) and weighted intersection-union ratio loss (w-IoU) as the loss function, expressed as:

[0084] Loss * =L w-IoU +L w-BCE ,

[0085] This invention sums the loss functions of the three branches and sets hyperparameters α = β = 0.25 and γ = 0.5 to assign different terms of loss. The total model training loss is expressed as:

[0086] Loss total=αLoss1+βLoss2+γLoss3

[0087] The overall training process of the model is as follows: Repeat S1 to S4 Epochs to obtain the neural network model and a total of Epoch × Q loss function values; then find the minimum loss function value and use its corresponding weight matrix and bias matrix as the optimal weight matrix and optimal bias matrix of the neural network. At this point, the model performs optimally.

[0088] In the testing phase of this embodiment, W original color object images and their corresponding infrared thermal imaging images, along with true saliency detection label images, are selected to form a test set. For each object w in the test set, the color object image and the infrared thermal imaging image are represented as Visual... w (i,j) and Thermal w (i,j). Let the width of each image be Width and the height be Height. The coordinates i,j of each pixel in each image of object w satisfy 1≤i≤Width, i≤j≤Height. To ensure that the resolution of the feature map is an integer during downsampling, the width Width and height Height of the input image must both be divisible by 2. Furthermore, during each training process, a batch of objects from the training set W are extracted for simultaneous testing to fully utilize GPU memory and accelerate the testing process. To comprehensively evaluate the generalization ability of the model proposed in this invention, the VT1000 dataset is used for evaluation, which includes 1000 pairs of RGB images—infrared thermal imaging—saliency detection result images, and there are no duplicate images from the training set.

[0089] Visualize color object images w The R-channel, G-channel, and B-channel components of (i,j) and Thermal w (i,j) are input into the optimal neural network model trained in step 1, and w is used to... best With b best Make predictions and obtain Visual w (i,j) and Thermal w The significance test result image corresponding to (i,j) is denoted as This represents the image composed of the pixel values ​​of all pixels with coordinates (i,j).

[0090] The method of this embodiment was verified using a Linux operating system, Python 3.7, and an Nvidia RTX 3090 GPU. The optimizer used was Adam, with parameters β1 set to 0.9, β2 set to 0.1, and the learning rate le set to 0.0001. Here, four commonly used metrics for significance evaluation were used to evaluate the invention: Mean Absolute Error (MAE), F-Measure, E-Measure, and S-Measure score. After testing using the method proposed in this invention, the MAE value was 0.0242. The average f-index was 0.8790, the average E-index was 0.9320, and the S-measure score was... α It is 0.9165. For example... Figure 5 The image shown is a comparison chart of the results of this embodiment and existing saliency detection methods; as shown... Figure 6 The figure shown is a comparison of the object structure extraction capabilities of this embodiment and existing saliency detection methods.

[0091] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0092] The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0093] Multiple components in the device are connected to an I / O interface, including: input units such as a keyboard, mouse, etc.; output units such as various types of displays, speakers, etc.; storage units such as disks, optical disks, etc.; and communication units such as network interface cards, modems, wireless transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks. The processing unit performs the various methods and processes described above, such as the method of the present invention. For example, in some embodiments, the method of the present invention may be implemented as a computer software program tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or the communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of the method of the present invention described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute the method of the present invention by any other suitable means (e.g., by means of firmware).

[0094] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0095] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0096] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0097] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A saliency target detection method based on a query-guided specificity learning network in RGB-T, characterized in that, The specific steps include: S1. Acquire visible light RGB images and their corresponding infrared thermal imaging images; S2. Construct a network model with an encoder-decoder architecture, and set up a visible light image saliency detection module, an infrared image saliency detection module, and a dual-modal saliency detection module according to the network model. The visible light RGB image is processed by the encoder of the visible light image saliency detection module to obtain visual features, and the infrared thermal image is processed by the infrared image saliency detection module to obtain infrared features. S3. Input the visual features and infrared features into the dual-modal information fusion module in the dual-modal saliency detection module, perform layer normalization processing, depthwise separable convolution operation and three-branch attention calculation to generate RGB modal features, infrared modal features and cross-modal fusion features; S4. Input the RGB modal features, infrared modal features, and cross-modal fusion features into the three-feature cross block in the dual-modal saliency detection module, and output the enhanced multimodal fusion features; input the multimodal fusion features into the fusion decoder, which consists of three cascaded query-guided decoding blocks, and outputs the final saliency target detection result image after step-by-step upsampling processing.

2. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 1, characterized in that, The network model of the encoder-decoder architecture is a visual state space block based on Vision Mamba; Each visual state space block includes an input layer, an output layer, and a hidden layer. The specific steps for processing the data include: The input image is segmented into multiple image patches, and each image patch is mapped to a one-dimensional vector of length D to obtain the input vector. The input vector is then sequentially input into the visual state space block. The hidden layer sequentially performs the following processing on the input vector: linear transformation through a first linear projection layer; feature extraction through a depthwise separable convolutional layer; nonlinear activation through a first SiLU activation function layer; scanning expansion, merging, and selective state space modeling of the feature sequence through a selective scanning module to obtain intermediate features; layer normalization, second linear projection, and second SiLU activation function processing are sequentially applied to the intermediate features; the output of the second SiLU activation function and the output of the first linear projection layer are multiplied element-wise to obtain the core operation output; the core operation output is then subjected to a third linear projection and added element-wise to the original input vector to obtain the final feature output vector of the visual state space block.

3. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 1, characterized in that, The visible light image saliency detection module includes four cascaded visible light image encoder groups, four visible light image decoder groups, and a visible light image saliency detection output layer. The output of the first visible light image encoder serves as the input of the fourth visible light image decoder, the output of the second visible light image encoder serves as the input of the third visible light image decoder, the output of the third visible light image encoder serves as the input of the second visible light image decoder, and the output of the fourth visible light image encoder serves as the input of the first visible light image decoder. The output of the second visible light image decoder serves as the input of the first query response decoder in the dual-modal saliency detection module, the output of the third visible light image decoder serves as the input of the second query response decoder in the dual-modal saliency detection module, and the output of the fourth visible light image decoder serves as the input of the third query response decoder in the dual-modal saliency detection module.

4. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 3, characterized in that, Both the visible light image encoder and the visible light image decoder include two visual state space blocks connected end-to-end. Each visual state space block specifically includes a first residual connection layer, a first normalization layer, a second residual connection layer, a first fully connected neural network layer, a first depthwise separable convolutional layer, a first selective scanning module, a second normalization layer, and a second fully connected neural network layer. The first residual connection layer adds the input feature map to the output of the other visual state space block. After the second residual layer passes through a fully connected neural network layer and an activation function layer, it performs a Hada code product operation with the result obtained from the second normalization layer.

5. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 3, characterized in that, The width and height of the output feature map of the first visible light image encoder are one-quarter of the initial visible light RGB image; the width and height of the output feature map of the second visible light image encoder are one-eighth of the initial visible light RGB image; the width and height of the output feature map of the third visible light image encoder are one-sixteenth of the initial visible light RGB image; the width and height of the output feature map of the fourth visible light image encoder are one-thirty-second of the initial visible light RGB image; the width and height of the output feature map of the first visible light image decoder are one-sixteenth of the initial visible light RGB image; the width and height of the output feature map of the second visible light image decoder are one-eighth of the initial visible light RGB image; the width and height of the output feature map of the third visible light image decoder are one-quarter of the initial visible light RGB image; and the width and height of the output feature map of the fourth visible light image decoder are the same as the initial visible light RGB image.

6. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 3, characterized in that, The structure of the infrared image saliency detection module is the same as that of the visible light image saliency detection module.

7. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 1, characterized in that, The specific steps of the dual-modal saliency detection module in S3 include: inputting the visual features output by the fourth visible light image encoder and the infrared features output by the fourth infrared image encoder after normalization of the receiving layer into the depthwise separable convolutional layer; converting the visual features and infrared features into a query matrix, a feature key matrix, and a feature value matrix respectively according to the attention mechanism to obtain self-attention features based on the visible light modality, cross-attention features based on visible light key-value pairs and infrared queries, and self-attention features based on the infrared modality. The hidden layer yields RGB modal features, infrared modal features, and cross-modal fusion features.

8. The RGB-T saliency target detection method based on a query-guided specificity learning network according to claim 1, characterized in that, The three-feature cross block in S4 includes an input layer, a hidden layer, and an output layer; the hidden layer operation includes: performing layer normalization on the cross-modal fusion features and RGB modal features respectively, and inputting them into the Swing Transformer network layer. Local window attention calculation is performed, and the result is added element-wise to the cross-modal fusion feature to obtain a first temporary feature. The first temporary feature is then input into the forward propagation network layer and added to itself to obtain the feature fused with RGB information. The cross-modal fusion feature and the infrared modal feature are respectively layer normalized and input into the Swing Transformer network layer C for local window attention calculation. The result is then added element-wise to the cross-modal fusion feature to obtain a second temporary feature. The second temporary feature is then input into the forward propagation network layer and added to itself to obtain the feature fused with infrared information. The enhanced multimodal fusion feature includes the feature fused with RGB information and the feature fused with infrared information.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.

Citation Information

Cited By

  • Composite insulator fault detection method based on YOLOV8

    CN121190491A

  • Salient target detection method, device and system and electronic equipment

    CN122023751A

  • A method, apparatus, system and electronic device for salient object detection

    CN122023751B