Multi-modal image saliency target detection method
By building a multi-layer feature extraction and trimodal fusion module in the neural network and combining the information of color visible light, infrared and depth images, the problem of insufficient information fusion and utilization in trimodal image detection is solved, and high-precision salient target detection is achieved.
Patent Information
- Application Number
- CN202510641971.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-09-26
AI Technical Summary
Existing multimodal image salient object detection models find it difficult to effectively fuse the geometric information of infrared and depth images with the texture information of color visible light images when given trimodal input, and fail to fully utilize the shared and unique information between the modalities, which affects detection accuracy.
A neural network structure consisting of three backbone networks is adopted to extract multi-layer features from color visible light, infrared and depth images respectively. The three-modal fusion module is used to deeply couple shared and complementary information, and the combined decoding module is used to independently construct prediction branches. The prediction branches and modal combination results are dynamically fused to achieve information complementarity and feature enhancement.
The accuracy and robustness of multimodal salient object detection are significantly improved, generating high-precision salient object images.
Smart Images

Figure CN120707813A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of salient object detection, and in particular to a method for detecting salient objects in multimodal images. Background Art
[0002] Salient object detection (SOD) aims to capture and segment the most attention-grabbing regions or objects in images or videos. As a crucial preprocessing step, SOD has been widely used in computer vision and image processing tasks, such as semantic segmentation, object tracking, image cropping, image retrieval, and image editing. In recent years, convolutional and deep neural networks, due to their powerful learning capabilities and excellent feature extraction, have pushed SOD performance to new heights. However, in some extreme scenarios (such as low illumination and cluttered scenes), it is often difficult to extract valuable information from color visible light images alone, which hinders the effectiveness of SOD models in complex conditions. Unlike color visible light images, thermal infrared cameras can capture thermal information that is less affected by environmental factors such as darkness and inclement weather, thereby revealing temperature differences, boundaries, and geometric shapes. Therefore, deploying thermal infrared cameras to collect thermal information and using a SOD model that combines color visible light and infrared images to record objects can improve the model's perception capabilities in a wider range of scenarios. Depth images contain information about the geometric structure of objects. In complex color scenes, the spatial information of different objects can better help models perceive the objects in the scene. Therefore, in recent years, research on combining color visible light images with infrared images and depth images for salient object detection has received increasing attention.
[0003] Currently, there are still several problems in multimodal image salient object detection:
[0004] First, how to achieve effective fusion of the three modalities? Existing models are designed based on the characteristics of two modalities. Cross-modal salient target detection is mainly aimed at the dual-modal situations of color visible light images and infrared images, and color visible light images and depth images. When the three modalities are input simultaneously, how to effectively utilize the geometric information shared by infrared images and depth images, as well as the rich texture information in color visible light images, to achieve information complementarity between different modalities is a key issue in the multimodal salient target detection task. Therefore, by analyzing the differences in effective information among the three modalities, designing a module that effectively utilizes modality-specific information to achieve feature complementarity to achieve the purpose of cross-modal fusion is crucial to improving the performance of convolutional / deep neural networks.
[0005] Second, how to make full use of the information of each modality and the information shared between modalities? The key to cross-modal tasks is to effectively explore the commonalities and differences between different modalities, and then use certain methods to utilize the commonalities between modalities and reduce the differences. Existing methods fuse modal features by designing a cross-modal fusion module, and use the fused features to upsample to obtain the predicted salient target image. Although these methods perform well in multimodal salient target detection tasks, they do not fully utilize the modality-specific information between the two modalities. Therefore, a combined decoding module is designed to independently predict salient target images for the features of different modalities and fused modalities through different branches, and then the prediction results of different branches are added to obtain the final predicted salient target image. This process can better utilize the shared information between modalities while paying attention to the characteristics of the modality itself, which plays a very important role in improving the accuracy of multimodal salient target detection. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a multimodal image salient object detection method, which can effectively solve the salient object detection problem of trimodal combination input images and has high salient object detection accuracy.
[0007] The technical solution adopted by the present invention to solve the above technical problems is: a multimodal image salient object detection method, characterized by comprising the following steps: first, constructing a training set containing several pairs of training images and building a neural network, wherein each pair of training images includes a color visible light image and its corresponding infrared image and depth image; second, inputting the several pairs of training images in the training set into the neural network for several rounds of training, and obtaining a neural network model after the training is completed; third, using the neural network model to predict test image pairs to obtain salient object images corresponding to the test image pairs, wherein the test image pairs include a color visible light image and its corresponding infrared image and depth image;
[0008] The neural network is mainly composed of a feature extraction module, a trimodal feature fusion module, and a combined decoding module. The feature extraction module includes a first backbone network for extracting feature information and scale information of a color visible light image, a second backbone network for extracting feature information and scale information of an infrared image, and a third backbone network for extracting feature information and scale information of a depth image. They all have a five-layer structure. The trimodal feature fusion module includes a first trimodal fusion module, a second trimodal fusion module, a third trimodal fusion module, a fourth trimodal fusion module, and a fifth trimodal fusion module. The combined decoding module includes a first prediction branch, a second prediction branch, a third prediction branch, a fourth prediction branch, a fifth prediction branch, and a sixth prediction branch.
[0009] The color visible light image passes through the first layer, second layer, third layer, fourth layer and fifth layer of the first backbone network in sequence, and the first layer, second layer, third layer, fourth layer and fifth layer of the first backbone network correspondingly output the first color visible light feature map FR1, the second color visible light feature map FR2, the third color visible light feature map FR3, the fourth color visible light feature map FR4 and the fifth color visible light feature map FR5; the infrared image passes through the first layer, second layer, third layer, fourth layer and fifth layer of the second backbone network in sequence, and the first layer, second layer, third layer, fourth layer and fifth layer of the second backbone network correspondingly output the first color visible light feature map FR1, the second color visible light feature map FR2, the third color visible light feature map FR3, the fourth color visible light feature map FR4 and the fifth color visible light feature map FR5. The first layer, the second layer, the third layer, the fourth layer, and the fifth layer correspondingly output the first infrared feature map FT1, the second infrared feature map FT2, the third infrared feature map FT3, the fourth infrared feature map FT4, and the fifth infrared feature map FT5; the depth image passes through the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the third backbone network in sequence, and the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the third backbone network correspondingly output the first depth feature map FD1, the second depth feature map FD2, the third depth feature map FD3, the fourth depth feature map FD4, and the fifth depth feature map FD5;
[0010] Input FR1, FT1, and FD1 into the first trimodal fusion module to obtain the first RD modal fusion feature map First RT modality fusion feature map First RDT modality fusion feature map Input FR2, FΤ2, and FD2 into the second trimodal fusion module to obtain the second RD modal fusion feature map Second RT modality fusion feature map Second RDT modality fusion feature map Input FR3, FT3, and FD3 into the third trimodal fusion module to obtain the third RD modal fusion feature map Third RT modality fusion feature map The third RDT modality fusion feature map Input FR4, FT4, and FD4 together into the fourth trimodal fusion module to obtain the fourth RD modal fusion feature map Fourth RT modality fusion feature map Fourth RDT modality fusion feature map Input FR5, FT5, and FD5 into the fifth trimodal fusion module to obtain the fifth RD modal fusion feature map Fifth RT modality fusion feature map Fifth RDT modality fusion feature map Among them, RD modality represents color visible light and depth dual modalities, RT modality represents color visible light and infrared dual modalities, and RDT modality represents color visible light, depth, and infrared triple modalities;
[0011] Input FR1, FR2, FR3, FR4, and FR5 into the first prediction branch to obtain the color visible light prediction feature map P R ; Input FT1, FΤ2, FT3, FT4, and FT5 together into the second prediction branch to obtain the infrared prediction feature map P T ; Input FD1, FD2, FD3, FD4, and FD5 into the third prediction branch to obtain the depth prediction feature map P D ;Will Input them together into the fourth prediction branch to obtain the RD modal prediction feature map P RD ;Will Input them together into the fifth prediction branch to obtain the RT modal prediction feature map P RT ;Will Input them together into the sixth prediction branch to obtain the RDT modal prediction feature map P RDT ;P R 、P T 、P D 、P RD 、P RT and P RDT Perform element-wise addition operation to obtain the salient target image S.
[0012] The first backbone network, the second backbone network, and the third backbone network all adopt the Res2Net-50 network.
[0013] The first trimodal fusion module, the second trimodal fusion module, the third trimodal fusion module, the fourth trimodal fusion module, and the fifth trimodal fusion module have the same structure, and are all composed of a first spatial attention layer, a second spatial attention layer, a first global average pooling layer, a second global average pooling layer, a first 1×1 convolutional layer, a second 1×1 convolutional layer, a third 1×1 convolutional layer, a fourth 1×1 convolutional layer, a first ReLU activation layer, a second ReLU activation layer, a third ReLU activation layer, a first Sigmoid activation layer, a second Sigmoid activation layer, a first 3×3 convolutional layer, and a first BN layer;
[0014] The implementation process of the j-th trimodal fusion module is as follows: the first input end of the j-th trimodal fusion module receives FR i , the second input terminal receives FT i 、The third input terminal receives FD i ; FR i With FD i Add elements to get the feature map As the intermediate output of the j-th trimodal fusion module; iWith FT i Add elements to get the feature map It is also used as the intermediate output of the j-th trimodal fusion module; FD i After the first spatial attention layer, the feature map is obtained FT i After the second spatial attention layer, the feature map is obtained Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will After the first global average pooling layer, the first 1×1 convolution layer, the first ReLU activation layer, the second 1×1 convolution layer, and the first Sigmoid activation layer, the feature map is obtained. Will After the second global average pooling layer, the third 1×1 convolution layer, the second ReLU activation layer, the fourth 1×1 convolution layer, and the second Sigmoid activation layer, the feature map is obtained. Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will After the first 3×3 convolution layer, the first BN layer, and the third ReLU activation layer, the feature map is obtained. The feature map output as the output end of the j-th trimodal fusion module; where j∈{1,2,3,4,5}, i∈{1,2,3,4,5}, and j corresponds to i one by one.
[0015] The first prediction branch, the second prediction branch, the third prediction branch, the fourth prediction branch, the fifth prediction branch, and the sixth prediction branch have the same structure, and are all composed of a first hollow residual pyramid attention module, a first cross-level feature fusion module, a second hollow residual pyramid attention module, a second cross-level feature fusion module, a third hollow residual pyramid attention module, a third cross-level feature fusion module, a fourth hollow residual pyramid attention module, and a fourth cross-level feature fusion module connected in sequence;
[0016] The implementation process of the k-th prediction branch is as follows: the first input end of the k-th prediction branch receives X1, the second input end receives X2, the third input end receives X3, the fourth input end receives X4, and the fifth input end receives X5; X1 is input to the first void residual pyramid attention module to obtain the feature map Will Together with X2, it is input into the first cross-level feature fusion module to obtain the feature map Will Input to the second hole residual pyramid attention module to obtain the feature map Will Together with X3, it is input into the second cross-level feature fusion module to obtain the feature map Will Input to the third hole residual pyramid attention module to obtain the feature map Will Together with X4, it is input into the third cross-level feature fusion module to obtain the feature map Will Input to the fourth hole residual pyramid attention module to obtain the feature map Will Together with X5, it is input into the fourth cross-level feature fusion module to obtain the feature map Y; where k∈{1,2,3,4,5,6}, in the first prediction branch, X1, X2, X3, X4, X5, Y correspond to FR1, FR2, FR3, FR4, FR5, P R In the second prediction branch, X1, X2, X3, X4, X5, and Y correspond to FT1, FT2, FT3, FT4, FT5, and P T In the third prediction branch, X1, X2, X3, X4, X5, and Y correspond to FD1, FD2, FD3, FD4, FD5, and P D In the fourth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RDIn the fifth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RT In the sixth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RDT .
[0017] The first atrous residual pyramid attention module, the second atrous residual pyramid attention module, the third atrous residual pyramid attention module, and the fourth atrous residual pyramid attention module have the same structure, and are all composed of a fifth 1×1 convolutional layer, a sixth 1×1 convolutional layer, a seventh 1×1 convolutional layer, an eighth 1×1 convolutional layer, a ninth 1×1 convolutional layer, a first atrous convolutional layer with a dilation rate of 3, a second atrous convolutional layer with a dilation rate of 5, a third atrous convolutional layer with a dilation rate of 7, a second 3×3 convolutional layer, a second BN layer, a fourth ReLU activation layer, a first channel attention layer, and a fifth ReLU activation layer;
[0018] The implementation process of the p-th void residual pyramid attention module is as follows: the input end of the p-th void residual pyramid attention module receives Will Input them into the fifth 1×1 convolution layer, the sixth 1×1 convolution layer, the seventh 1×1 convolution layer, the eighth 1×1 convolution layer, and the ninth 1×1 convolution layer respectively, and obtain the corresponding feature maps Feature Map Feature Map Feature Map Feature Map Will and Add elements to get the feature map Will Input to the first hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the second hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the third hole convolution layer to get the feature map Will Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the second 3×3 convolution layer, the second BN layer, the fourth ReLU activation layer, and the first channel attention layer. Will and Add elements to get the feature map Will Through the fifth ReLU activation layer, the feature map is obtained As the feature map outputted by the output of the pth hole residual pyramid attention module, where p∈{1,2,3,4}, q∈{1,2,3,4}, and p and q correspond one to one. For the first hole residual pyramid attention module, That is X1.
[0019] The first cross-level feature fusion module, the second cross-level feature fusion module, the third cross-level feature fusion module, and the fourth cross-level feature fusion module have the same structure, which are all composed of a third 3×3 convolutional layer, a third BN layer, a sixth ReLU activation layer, a third spatial attention layer, a second channel attention layer, a fourth 3×3 convolutional layer, a fourth BN layer, and a seventh ReLU activation layer;
[0020] The implementation process of the p-th cross-level feature fusion module is as follows: the first input end of the p-th cross-level feature fusion module receives X q+1 , the second input terminal receives X q+1 The feature map is obtained by sequentially passing through the third 3×3 convolution layer, the third BN layer, and the sixth ReLU activation layer. Will The feature map is obtained by sequentially passing through the third spatial attention layer and the second channel attention layer. Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the fourth 3×3 convolution layer, the fourth BN layer, and the seventh ReLU activation layer. Among them, p∈{1,2,3,4}, q∈{1,2,3,4}, and p and q correspond one to one. For the fourth cross-level feature fusion module, That is Y.
[0021] The training set construction process is as follows: at least 1000 pairs of original color visible light images and their corresponding original infrared images and original depth images are selected; then each original color visible light image and its corresponding original infrared image and original depth image are downsampled, and the size of the downsampled image is H×W; then each color visible light image of size H×W and its corresponding infrared image and depth image are formed into a pair of training image pairs; finally, at least 1000 pairs of training image pairs are formed into a training set.
[0022] The acquisition process of the neural network model is as follows: several pairs of training images in the training set are input into the neural network for training, and the loss function Loss is calculated before the end of each round of training to optimize the neural network, and the neural network model is obtained after a total of Num rounds of training; wherein, Ω={1, 2, 3, 4, 5, 6},L k represents the prediction loss on the k-th prediction branch, P k represents the predicted feature map output by the k-th prediction branch, G represents the label image, S represents the salient target image, represents the weighted binary cross entropy loss, represents the weighted intersection-over-union loss.
[0023] The acquisition process of the test image pair is as follows: arbitrarily select a pair of original color visible light images and their corresponding original infrared images and original depth images, and perform a downsampling operation, so that the size of the downsampled images is H×W; then, the color visible light images of size H×W and their corresponding infrared images and depth images constitute a test image pair.
[0024] Compared with the prior art, the advantages of the present invention are:
[0025] 1) The proposed method uses three backbone networks to extract multi-layer features from color, visible light, infrared, and depth images, respectively. A subsequent trimodal fusion module deeply couples shared and complementary information, and a combined decoding module extracts shared features alongside modality-specific features, achieving collaborative optimization. This structural neural network model significantly enhances cross-modal discrimination and robustness, enabling the generation of highly accurate salient object images.
[0026] 2) By analyzing the differences in effective information among the three modalities, the method of the present invention designs a three-modal fusion module that fully utilizes the unique information of each modality and realizes complementary fusion. It mines the common geometric information in depth images and infrared images, and combines it with the texture information in color visible light images, effectively improving the accuracy of salient target detection.
[0027] 3) Based on a combined decoding module, the proposed method independently constructs prediction branches for different modal combinations and adaptively weights and fuses the prediction branches with the modal combination results through a dynamic fusion strategy. This design not only deeply mines shared information between modalities, but also retains and enhances the unique characteristics of each modality, achieving collaborative optimization and significantly improving the accuracy and robustness of multimodal salient object detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0029] Figure 1 This is a block diagram of the overall implementation of the method of the present invention;
[0030] Figure 2 Schematic diagram of the composition structure of the neural network constructed in the method of the present invention;
[0031] Figure 3 Schematic diagram of the composition structure of the trimodal fusion module in the neural network constructed in the method of the present invention;
[0032] Figure 4 Schematic diagram of the composition structure of the prediction branch in the neural network constructed in the method of the present invention;
[0033] Figure 5 Schematic diagram of the composition structure of the hole residual pyramid attention module in the prediction branch of the neural network constructed in the method of the present invention;
[0034] Figure 6 Schematic diagram of the composition structure of the cross-level feature fusion module in the prediction branch of the neural network constructed in the method of the present invention. DETAILED DESCRIPTION
[0035] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0036] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments.
[0037] The present invention proposes a multimodal image salient object detection method, the overall implementation diagram of which is as follows: Figure 1As shown, it includes the following steps: first, constructing a training set containing several pairs of training image pairs and building a neural network, wherein each pair of training image pairs includes a color visible light image and its corresponding infrared image and depth image; second, inputting several pairs of training image pairs in the training set into the neural network for several rounds (such as 250 rounds) of training, and obtaining a neural network model after the training is completed; third, using the neural network model to predict the test image pairs, and predicting the corresponding salient target image of the test image pairs, wherein the test image pairs include a color visible light image and its corresponding infrared image and depth image.
[0038] In one specific embodiment, the training set is constructed by selecting at least 1,000 pairs of original color visible light images and their corresponding original infrared images and original depth images; then downsampling each original color visible light image and its corresponding original infrared image and original depth image to an image size of H×W; then forming a training image pair from each H×W color visible light image and its corresponding infrared image and depth image; and finally, forming the training set from the at least 1,000 training image pairs. 1,048 pairs of original color visible light images and their corresponding original infrared images and original depth images are selected, with H=W=224.
[0039] In a specific embodiment, the neural network is mainly composed of a feature extraction module, a trimodal feature fusion module, and a combined decoding module. Figure 2 As shown, the feature extraction module includes a first backbone network for extracting feature information and scale information of color visible light images, a second backbone network for extracting feature information and scale information of infrared images, and a third backbone network for extracting feature information and scale information of depth images. They all have a five-layer structure. The trimodal feature fusion module includes a first trimodal fusion module, a second trimodal fusion module, a third trimodal fusion module, a fourth trimodal fusion module, and a fifth trimodal fusion module. The combined decoding module includes a first prediction branch, a second prediction branch, a third prediction branch, a fourth prediction branch, a fifth prediction branch, and a sixth prediction branch. The six prediction branches are used for independent prediction of different modal combinations.
[0040] In a specific implementation, the color visible light image is sequentially passed through the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the first backbone network, and the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the first backbone network correspondingly output the first color visible light feature map FR1, the second color visible light feature map FR2, the third color visible light feature map FR3, the fourth color visible light feature map FR4, and the fifth color visible light feature map FR5; the infrared image is sequentially passed through the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the second backbone network, and the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the second backbone network The second layer, the third layer, the fourth layer, and the fifth layer correspondingly output the first infrared feature map FT1, the second infrared feature map FΤ2, the third infrared feature map FT3, the fourth infrared feature map FT4, and the fifth infrared feature map FT5; the depth image passes through the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the third backbone network in sequence, and the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the third backbone network correspondingly output the first depth feature map FD1, the second depth feature map FD2, the third depth feature map FD3, the fourth depth feature map FD4, and the fifth depth feature map FD5.
[0041] In the specific implementation, FR1, FT1, and FD1 are input into the first three-modal fusion module to obtain the first RD modal fusion feature map First RT modality fusion feature map First RDT modality fusion feature map Input FR2, FΤ2, and FD2 into the second trimodal fusion module to obtain the second RD modal fusion feature map Second RT modality fusion feature map Second RDT modality fusion feature map Input FR3, FT3, and FD3 into the third trimodal fusion module to obtain the third RD modal fusion feature map Third RT modality fusion feature map The third RDT modality fusion feature map Input FR4, FT4, and FD4 together into the fourth trimodal fusion module to obtain the fourth RD modal fusion feature map Fourth RT modality fusion feature map Fourth RDT modality fusion feature map Input FR5, FT5, and FD5 into the fifth trimodal fusion module to obtain the fifth RD modal fusion feature map Fifth RT modality fusion feature map Fifth RDT modality fusion feature map Among them, RD modality represents color visible light and depth modalities, RT modality represents color visible light and infrared modalities, and RDT modality represents color visible light, depth, and infrared modalities.
[0042] In specific implementation, FR1, FR2, FR3, FR4, and FR5 are input into the first prediction branch to obtain the color visible light prediction feature map P R ; Input FT1, FΤ2, FT3, FT4, and FT5 together into the second prediction branch to obtain the infrared prediction feature map P T ; Input FD1, FD2, FD3, FD4, and FD5 into the third prediction branch to obtain the depth prediction feature map P D ;Will Input them together into the fourth prediction branch to obtain the RD modal prediction feature map P RD ;Will Input them together into the fifth prediction branch to obtain the RT modal prediction feature map P RT ;Will Input them together into the sixth prediction branch to obtain the RDT modal prediction feature map P RDT ;P R 、P T 、P D 、P RD 、P RT and P RDT Perform element-wise addition operation to obtain the salient target image S.
[0043] Here, the size of the color visible light image, infrared image, and depth image is H×W×3. In this embodiment, H×W is 224×224, and the sizes of FR1, FT1, and FD1 are all The sizes of FR2, FΤ2 and FD2 are The sizes of FR3, FT3 and FD3 are The sizes of FR4, FT4 and FD4 are The sizes of FR5, FT5 and FD5 are The size of The size of The size of The size of The size of P R 、P T 、P D 、P RD 、P RT 、P RDT The sizes of and S are both H×W×1.
[0044] In a specific embodiment, the first backbone network, the second backbone network, and the third backbone network all use the Res2Net-50 network. The Res2Net-50 network is an existing network and is described in Gao, SH, Cheng, MM, Zhao, K., Zhang, XY, Yang, MH, Torr, P. (2019). Res2net: A new multi-scale backbone architecture. IEEE transactions on pattern analysis and machine intelligence, 43(2), 652-662.
[0045] In a specific embodiment, the first trimodal fusion module, the second trimodal fusion module, the third trimodal fusion module, the fourth trimodal fusion module, and the fifth trimodal fusion module have the same structure, except for different inputs and outputs, such as Figure 3 As shown in FIG, each of them consists of a first spatial attention layer, a second spatial attention layer, a first global average pooling layer, a second global average pooling layer, a first 1×1 convolution layer, a second 1×1 convolution layer, a third 1×1 convolution layer, a fourth 1×1 convolution layer, a first ReLU activation layer, a second ReLU activation layer, a third ReLU activation layer, a first Sigmoid activation layer, a second Sigmoid activation layer, a first 3×3 convolution layer, and a first BN (Batch Normalization) layer. The implementation process of the j-th trimodal fusion module is as follows: the first input end of the j-th trimodal fusion module receives FR i , the second input terminal receives FT i 、The third input terminal receives FD i ; FR i With FD i Add elements to get the feature map As the intermediate output of the j-th trimodal fusion module; i With FT i Add elements to get the feature map It is also used as the intermediate output of the j-th trimodal fusion module; FD i After the first spatial attention layer, the feature map is obtained FT i After the second spatial attention layer, the feature map is obtained Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will After the first global average pooling layer, the first 1×1 convolution layer, the first ReLU activation layer, the second 1×1 convolution layer, and the first Sigmoid activation layer, the feature map is obtained. Will After the second global average pooling layer, the third 1×1 convolution layer, the second ReLU activation layer, the fourth 1×1 convolution layer, and the second Sigmoid activation layer, the feature map is obtained. Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will After the first 3×3 convolution layer, the first BN layer, and the third ReLU activation layer, the feature map is obtained. The feature map outputted by the jth trimodal fusion module, where j∈{1,2,3,4,5} and i∈{1,2,3,4,5}, with a one-to-one correspondence between j and i. Element-wise addition, element-wise multiplication, and concatenation along channels are all common operations in neural networks; spatial attention layers are a common module in neural networks.
[0046] In a specific embodiment, the first prediction branch, the second prediction branch, the third prediction branch, the fourth prediction branch, the fifth prediction branch, and the sixth prediction branch have the same structure, except for different inputs and outputs, such as Figure 4As shown, they are composed of the first hollow residual pyramid attention module, the first cross-level feature fusion module, the second hollow residual pyramid attention module, the second cross-level feature fusion module, the third hollow residual pyramid attention module, the third cross-level feature fusion module, the fourth hollow residual pyramid attention module, and the fourth cross-level feature fusion module connected in sequence. The implementation process of the k-th prediction branch is: the first input end of the k-th prediction branch receives X1, the second input end receives X2, the third input end receives X3, the fourth input end receives X4, and the fifth input end receives X5; X1 is input to the first hollow residual pyramid attention module to obtain the feature map Will Together with X2, it is input into the first cross-level feature fusion module to obtain the feature map Will Input to the second hole residual pyramid attention module to obtain the feature map Will Together with X3, it is input into the second cross-level feature fusion module to obtain the feature map Will Input to the third hole residual pyramid attention module to obtain the feature map Will Together with X4, it is input into the third cross-level feature fusion module to obtain the feature map Will Input to the fourth hole residual pyramid attention module to obtain the feature map Will Together with X5, it is input into the fourth cross-level feature fusion module to obtain the feature map Y; where k∈{1,2,3,4,5,6}, in the first prediction branch, X1, X2, X3, X4, X5, Y correspond to FR1, FR2, FR3, FR4, FR5, P R In the second prediction branch, X1, X2, X3, X4, X5, and Y correspond to FT1, FT2, FT3, FT4, FT5, and P T In the third prediction branch, X1, X2, X3, X4, X5, and Y correspond to FD1, FD2, FD3, FD4, FD5, and P D In the fourth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RD In the fifth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RT In the sixth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RDT Taking the first prediction branch as an example, the implementation process is as follows: FR1 is input into the first hole residual pyramid attention module to obtain the feature map Will Together with FR2, it is input into the first cross-level feature fusion module to obtain the feature map Will Input to the second hole residual pyramid attention module to obtain the feature map Will Together with FR3, it is input into the second cross-level feature fusion module to obtain the feature map Will Input to the third hole residual pyramid attention module to obtain the feature map Will Together with FR4, it is input into the third cross-level feature fusion module to obtain the feature map Will Input to the fourth hole residual pyramid attention module to obtain the feature map Will Together with FR5, it is input into the fourth cross-level feature fusion module to obtain the feature map P R .
[0047] In a specific embodiment, the structures of the first atrous residual pyramid attention module, the second atrous residual pyramid attention module, the third atrous residual pyramid attention module, and the fourth atrous residual pyramid attention module are the same, except that the inputs and outputs are different, such as Figure 5 As shown in Figure 2, it consists of the fifth 1×1 convolution layer, the sixth 1×1 convolution layer, the seventh 1×1 convolution layer, the eighth 1×1 convolution layer, the ninth 1×1 convolution layer, the first dilated convolution layer with a dilation rate of 3, the second dilated convolution layer with a dilation rate of 5, the third dilated convolution layer with a dilation rate of 7, the second 3×3 convolution layer, the second BN layer, the fourth ReLU activation layer, the first channel attention layer, and the fifth ReLU activation layer. The implementation process of the p-th dilated residual pyramid attention module is as follows: the input end of the p-th dilated residual pyramid attention module receives Will Input them into the fifth 1×1 convolution layer, the sixth 1×1 convolution layer, the seventh 1×1 convolution layer, the eighth 1×1 convolution layer, and the ninth 1×1 convolution layer respectively, and obtain the corresponding feature maps Feature Map Feature Map Feature Map Feature Map Will and Add elements to get the feature map Will Input to the first hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the second hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the third hole convolution layer to get the feature map Will Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the second 3×3 convolution layer, the second BN layer, the fourth ReLU activation layer, and the first channel attention layer. Will and Add elements to get the feature map Will Through the fifth ReLU activation layer, the feature map is obtained As the feature map outputted by the output of the pth hole residual pyramid attention module, where p∈{1,2,3,4}, q∈{1,2,3,4}, and p and q correspond one to one. For the first hole residual pyramid attention module, That is X1. Here, the channel attention layer is a common module in the neural network. Taking the first hole residual pyramid attention module in the first prediction branch as an example, its implementation process is: input FR1 to the fifth 1×1 convolution layer, the sixth 1×1 convolution layer, the seventh 1×1 convolution layer, the eighth 1×1 convolution layer, and the ninth 1×1 convolution layer respectively, and the corresponding feature map is obtained. Will and Add elements to get the feature map Will Input to the first hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the second hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the third hole convolution layer to get the feature map Will Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the second 3×3 convolution layer, the second BN layer, the fourth ReLU activation layer, and the first channel attention layer. Will and Add elements to get the feature map Will Through the fifth ReLU activation layer, the feature map is obtained
[0048] In a specific embodiment, the first cross-level feature fusion module, the second cross-level feature fusion module, the third cross-level feature fusion module, and the fourth cross-level feature fusion module have the same structure, except for different inputs and outputs, such as Figure 6 As shown in Figure 2, each of them consists of a third 3×3 convolutional layer, a third BN layer, a sixth ReLU activation layer, a third spatial attention layer, a second channel attention layer, a fourth 3×3 convolutional layer, a fourth BN layer, and a seventh ReLU activation layer. The implementation process of the p-th cross-level feature fusion module is as follows: the first input end of the p-th cross-level feature fusion module receives X q+1 , the second input terminal receives X q+1 The feature map is obtained by sequentially passing through the third 3×3 convolution layer, the third BN layer, and the sixth ReLU activation layer. Will The feature map is obtained by sequentially passing through the third spatial attention layer and the second channel attention layer. Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the fourth 3×3 convolution layer, the fourth BN layer, and the seventh ReLU activation layer. Among them, p∈{1,2,3,4}, q∈{1,2,3,4}, and p and q correspond one to one. For the fourth cross-level feature fusion module, That is Y. Taking the first cross-level feature fusion module in the first prediction branch as an example, its implementation process is: FR2 passes through the third 3×3 convolution layer, the third BN layer, and the sixth ReLU activation layer in sequence to obtain the feature map Will The feature map is obtained by sequentially passing through the third spatial attention layer and the second channel attention layer. Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the fourth 3×3 convolution layer, the fourth BN layer, and the seventh ReLU activation layer.
[0049] In a specific embodiment, the process of obtaining the neural network model is as follows: several pairs of training images in the training set are input into the neural network for training, and before the end of each round of training, the loss function Loss is calculated to optimize the neural network, and the neural network model is obtained after a total of Num rounds (e.g., Num=250 rounds) of training; wherein, Ω={1, 2, 3, 4, 5, 6},L k represents the prediction loss on the k-th prediction branch, P k represents the predicted feature map output by the k-th prediction branch, G represents the label image, i.e., the true value map of the salient target, S represents the salient target image, represents the weighted binary cross entropy loss, which is described in, for example, R. Cong, Q. Lin, C. Zhang, C. Li, X. Cao, Q. Huang and Y. Zhao, “CIR-Net: Cross-modality interaction and refinement for RGB-D salient object detection,” IEEE Transactions on Image Processing, 31, 6800–6815, 2022. represents the weighted intersection-over-union loss, which is described in, for example, G. Máttyus, W. Luo and R. Urtasun, “Deeproadmapper: Extracting road topology from aerial images,” In Proceedings of the IEEE international conference on computer vision, pp. 3438–3446, 2017.
[0050] In a specific embodiment, the process of acquiring the test image pair is as follows: arbitrarily select a pair of original color visible light images and their corresponding original infrared images and original depth images, and perform a downsampling operation, where the size of the downsampled images is H×W; then, the color visible light images of size H×W and their corresponding infrared images and depth images constitute a test image pair.
[0051] In order to further illustrate the feasibility and effectiveness of the method of the present invention, the method of the present invention was tested.
[0052] In the experiment, the method of the present invention was tested on a dataset of salient object detection in color, visible light, depth, and infrared images (VDT2048 dataset). The VDT2048 dataset consists of a training set and a test set. The training set includes 1048 pairs of color, visible light, infrared, and depth images (i.e., 1048 training image pairs), and the test set includes 1000 pairs of color, visible light, infrared, and depth images (i.e., 1000 test image pairs).
[0053] In the experiment, four commonly used objective parameters were selected to evaluate the performance of the method of the present invention, namely S-measure, E-measure, F-measure, and Mean Absolute Error (MAE).
[0054] Table 1 shows the correlation between the salient target image and the label image obtained by the method of the present invention on the VDT2048 dataset.
[0055] Table 1 S-measure, E-measure, F-measure, and Mean Absolute Error (MAE) between the salient target image and the label image obtained by the method of the present invention on the VDT2048 dataset
[0056] Dataset S-measure E-measure F-measure MAE VDT2048 0.9012 0.9709 0.8519 0.0034
[0057] The results in Table 1 show that the proposed method achieves higher S-measure, E-measure, and F-measure, as well as lower MAE, on an existing trimodal salient object detection dataset. This indicates that the salient object images obtained using the proposed method are relatively close to the labeled images, and that the proposed method can effectively complete salient object detection tasks with inputs from different modalities.
[0058] The above are merely embodiments of the present invention and are not intended to limit the present invention. It will be apparent to those skilled in the art that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are intended to be included within the scope of the claims of the present invention.
Claims
1. A multimodal image salient object detection method, characterized by The following steps are involved: First, a training set containing several pairs of training images is constructed, and a neural network is built. Each training image pair includes a color visible light image and its corresponding infrared image and depth image. Second, the training image pairs in the training set are input into the neural network for several rounds of training. After the training, a neural network model is obtained. Third, the neural network model is used to predict the salient target image corresponding to the test image pair. The test image pair includes a color visible light image and its corresponding infrared image and depth image. The neural network is mainly composed of a feature extraction module, a trimodal feature fusion module, and a combined decoding module. The feature extraction module includes a first backbone network for extracting feature information and scale information of a color visible light image, a second backbone network for extracting feature information and scale information of an infrared image, and a third backbone network for extracting feature information and scale information of a depth image. They all have a five-layer structure. The trimodal feature fusion module includes a first trimodal fusion module, a second trimodal fusion module, a third trimodal fusion module, a fourth trimodal fusion module, and a fifth trimodal fusion module. The combined decoding module includes a first prediction branch, a second prediction branch, a third prediction branch, a fourth prediction branch, a fifth prediction branch, and a sixth prediction branch. The color visible light image passes through the first layer, second layer, third layer, fourth layer and fifth layer of the first backbone network in sequence, and the first layer, second layer, third layer, fourth layer and fifth layer of the first backbone network correspondingly output the first color visible light feature map FR1, the second color visible light feature map FR2, the third color visible light feature map FR3, the fourth color visible light feature map FR4 and the fifth color visible light feature map FR5; the infrared image passes through the first layer, second layer, third layer, fourth layer and fifth layer of the second backbone network in sequence, and the first layer, second layer, third layer, fourth layer and fifth layer of the second backbone network correspondingly output the first color visible light feature map FR1, the second color visible light feature map FR2, the third color visible light feature map FR3, the fourth color visible light feature map FR4 and the fifth color visible light feature map FR5. The first layer, the second layer, the third layer, the fourth layer, and the fifth layer correspondingly output the first infrared feature map FT1, the second infrared feature map FT2, the third infrared feature map FT3, the fourth infrared feature map FT4, and the fifth infrared feature map FT5; the depth image passes through the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the third backbone network in sequence, and the first layer, the second layer, the third layer, the fourth layer, and the fifth layer of the third backbone network correspondingly output the first depth feature map FD1, the second depth feature map FD2, the third depth feature map FD3, the fourth depth feature map FD4, and the fifth depth feature map FD5; Input FR1, FT1, and FD1 into the first trimodal fusion module to obtain the first RD modal fusion feature map First RT modality fusion feature map First RDT modality fusion feature map Input FR2, FΤ2, and FD2 into the second trimodal fusion module to obtain the second RD modal fusion feature map Second RT modality fusion feature map Second RDT modality fusion feature map Input FR3, FT3, and FD3 into the third trimodal fusion module to obtain the third RD modal fusion feature map Third RT modality fusion feature map The third RDT modality fusion feature map Input FR4, FT4, and FD4 together into the fourth trimodal fusion module to obtain the fourth RD modal fusion feature map Fourth RT modality fusion feature map Fourth RDT modality fusion feature map Input FR5, FT5, and FD5 into the fifth trimodal fusion module to obtain the fifth RD modal fusion feature map Fifth RT modality fusion feature map Fifth RDT modality fusion feature map Among them, RD modality represents color visible light and depth dual modalities, RT modality represents color visible light and infrared dual modalities, and RDT modality represents color visible light, depth, and infrared triple modalities; Input FR1, FR2, FR3, FR4, and FR5 into the first prediction branch to obtain the color visible light prediction feature map P R ; Input FT1, FΤ2, FT3, FT4, and FT5 together into the second prediction branch to obtain the infrared prediction feature map P T ; Input FD1, FD2, FD3, FD4, and FD5 into the third prediction branch to obtain the depth prediction feature map P D ;Will Input them together into the fourth prediction branch to obtain the RD modal prediction feature map P RD ;Will Input them together into the fifth prediction branch to obtain the RT modal prediction feature map P RT ;Will Input them together into the sixth prediction branch to obtain the RDT modal prediction feature map P RDT ;P R 、P T 、P D 、P RD 、P RT and P RDT Perform element-wise addition operation to obtain the salient target image S.
2. A multimodal image salient object detection method according to claim 1, characterized in that The first backbone network, the second backbone network, and the third backbone network all adopt the Res2Net-50 network.
3. The multimodal image salient object detection method according to claim 1, characterized in that The first trimodal fusion module, the second trimodal fusion module, the third trimodal fusion module, the fourth trimodal fusion module, and the fifth trimodal fusion module have the same structure, and are all composed of a first spatial attention layer, a second spatial attention layer, a first global average pooling layer, a second global average pooling layer, a first 1×1 convolutional layer, a second 1×1 convolutional layer, a third 1×1 convolutional layer, a fourth 1×1 convolutional layer, a first ReLU activation layer, a second ReLU activation layer, a third ReLU activation layer, a first Sigmoid activation layer, a second Sigmoid activation layer, a first 3×3 convolutional layer, and a first BN layer; The implementation process of the j-th trimodal fusion module is as follows: the first input end of the j-th trimodal fusion module receives FR i , the second input terminal receives FT i 、The third input terminal receives FD i ; FR i With FD i Add elements to get the feature map As the intermediate output of the j-th trimodal fusion module; i With FT i Add elements to get the feature map It is also used as the intermediate output of the j-th trimodal fusion module; FD i After the first spatial attention layer, the feature map is obtained FT i After the second spatial attention layer, the feature map is obtained Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map FT i 3 ;Will After the first global average pooling layer, the first 1×1 convolution layer, the first ReLU activation layer, the second 1×1 convolution layer, and the first Sigmoid activation layer, the feature map is obtained. Will After the second global average pooling layer, the third 1×1 convolution layer, the second ReLU activation layer, the fourth 1×1 convolution layer, and the second Sigmoid activation layer, the feature map is obtained. Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will After the first 3×3 convolution layer, the first BN layer, and the third ReLU activation layer, the feature map is obtained. The feature map output as the output end of the j-th trimodal fusion module; where j∈{1,2,3,4,5}, i∈{1,2,3,4,5}, and j corresponds to i one by one.
4. A multimodal image salient object detection method according to claim 3, characterized in that The first prediction branch, the second prediction branch, the third prediction branch, the fourth prediction branch, the fifth prediction branch, and the sixth prediction branch have the same structure, and are all composed of a first hollow residual pyramid attention module, a first cross-level feature fusion module, a second hollow residual pyramid attention module, a second cross-level feature fusion module, a third hollow residual pyramid attention module, a third cross-level feature fusion module, a fourth hollow residual pyramid attention module, and a fourth cross-level feature fusion module connected in sequence; The implementation process of the k-th prediction branch is as follows: the first input end of the k-th prediction branch receives X1, the second input end receives X2, the third input end receives X3, the fourth input end receives X4, and the fifth input end receives X5; X1 is input to the first void residual pyramid attention module to obtain the feature map Will Together with X2, it is input into the first cross-level feature fusion module to obtain the feature map Will Input to the second hole residual pyramid attention module to obtain the feature map Will Together with X3, it is input into the second cross-level feature fusion module to obtain the feature map Will Input to the third hole residual pyramid attention module to obtain the feature map Will Together with X4, it is input into the third cross-level feature fusion module to obtain the feature map Will Input to the fourth hole residual pyramid attention module to obtain the feature map Will Together with X5, it is input into the fourth cross-level feature fusion module to obtain the feature map Y; where k∈{1,2,3,4,5,6}, in the first prediction branch, X1, X2, X3, X4, X5, Y correspond to FR1, FR2, FR3, FR4, FR5, P R In the second prediction branch, X1, X2, X3, X4, X5, and Y correspond to FT1, FT2, FT3, FT4, FT5, and P T In the third prediction branch, X1, X2, X3, X4, X5, and Y correspond to FD1, FD2, FD3, FD4, FD5, and P D In the fourth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RD In the fifth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RT In the sixth prediction branch, X1, X2, X3, X4, X5, and Y correspond to P RDT .
5. A multimodal image salient object detection method according to claim 4, characterized in that The first atrous residual pyramid attention module, the second atrous residual pyramid attention module, the third atrous residual pyramid attention module, and the fourth atrous residual pyramid attention module have the same structure, and are all composed of a fifth 1×1 convolutional layer, a sixth 1×1 convolutional layer, a seventh 1×1 convolutional layer, an eighth 1×1 convolutional layer, a ninth 1×1 convolutional layer, a first atrous convolutional layer with a dilation rate of 3, a second atrous convolutional layer with a dilation rate of 5, a third atrous convolutional layer with a dilation rate of 7, a second 3×3 convolutional layer, a second BN layer, a fourth ReLU activation layer, a first channel attention layer, and a fifth ReLU activation layer; The implementation process of the p-th void residual pyramid attention module is as follows: the input end of the p-th void residual pyramid attention module receives Will Input them into the fifth 1×1 convolution layer, the sixth 1×1 convolution layer, the seventh 1×1 convolution layer, the eighth 1×1 convolution layer, and the ninth 1×1 convolution layer respectively, and obtain the corresponding feature maps Feature Map Feature Map Feature Map Feature Map Will and Add elements to get the feature map Will Input to the first hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the second hole convolution layer to obtain the feature map Will and Add elements to get the feature map Will Input to the third hole convolution layer to get the feature map Will Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the second 3×3 convolution layer, the second BN layer, the fourth ReLU activation layer, and the first channel attention layer. Will and Add elements to get the feature map Will Through the fifth ReLU activation layer, the feature map is obtained As the feature map outputted by the output of the pth hole residual pyramid attention module, where p∈{1,2,3,4}, q∈{1,2,3,4}, and p and q correspond one to one. For the first hole residual pyramid attention module, That is X1.
6. A multimodal image salient object detection method according to claim 4 or 5, characterized in that The first cross-level feature fusion module, the second cross-level feature fusion module, the third cross-level feature fusion module, and the fourth cross-level feature fusion module have the same structure, which are all composed of a third 3×3 convolutional layer, a third BN layer, a sixth ReLU activation layer, a third spatial attention layer, a second channel attention layer, a fourth 3×3 convolutional layer, a fourth BN layer, and a seventh ReLU activation layer; The implementation process of the p-th cross-level feature fusion module is as follows: the first input end of the p-th cross-level feature fusion module receives X q+1 , the second input terminal receives X q+1 The feature map is obtained by sequentially passing through the third 3×3 convolution layer, the third BN layer, and the sixth ReLU activation layer. Will The feature map is obtained by sequentially passing through the third spatial attention layer and the second channel attention layer. Will and Perform element multiplication to obtain the feature map Will and Add elements to get the feature map Will and Splice along the channel to get the feature map Will The feature map is obtained by sequentially passing through the fourth 3×3 convolution layer, the fourth BN layer, and the seventh ReLU activation layer. Among them, p∈{1,2,3,4}, q∈{1,2,3,4}, and p and q correspond one to one. For the fourth cross-level feature fusion module, That is Y.
7. The multimodal image salient object detection method according to claim 1, characterized in that The training set construction process is as follows: at least 1000 pairs of original color visible light images and their corresponding original infrared images and original depth images are selected; then each original color visible light image and its corresponding original infrared image and original depth image are downsampled, and the size of the downsampled image is H×W; then each color visible light image of size H×W and its corresponding infrared image and depth image are formed into a pair of training image pairs; finally, at least 1000 pairs of training image pairs are formed into a training set.
8. The multimodal image salient object detection method according to claim 1, characterized in that The acquisition process of the neural network model is as follows: several pairs of training images in the training set are input into the neural network for training, and the loss function Loss is calculated before the end of each round of training to optimize the neural network, and the neural network model is obtained after a total of Num rounds of training; wherein, Ω={1, 2, 3, 4, 5, 6},L k represents the prediction loss on the k-th prediction branch, P k represents the predicted feature map output by the k-th prediction branch, G represents the label image, S represents the salient target image, represents the weighted binary cross entropy loss, represents the weighted intersection-over-union loss.
9. The multimodal image salient object detection method according to claim 1, characterized in that The acquisition process of the test image pair is as follows: arbitrarily select a pair of original color visible light images and their corresponding original infrared images and original depth images, and perform a downsampling operation, so that the size of the downsampled images is H×W; then, the color visible light images of size H×W and their corresponding infrared images and depth images constitute a test image pair.
Citation Information
Cited By
Salient target detection method, device and system and electronic equipment
CN122023751A