A saliency object detection method based on lightweight neural network

By adopting lightweight neural network and multimodal feature fusion technology in RGB-T significance object detection, the problem of the existing model calculation cost and model size is solved, and efficient significance object detection in actual scenarios is achieved.

CN119741580BActive Publication Date: 2025-06-10数字宁波科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510246732.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-10
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing RGB-T significance target detection models have made breakthroughs in accuracy, but they usually require high computing costs and large model sizes, making it difficult to meet the needs of actual scenarios in terms of computing delay and bandwidth costs, such as autonomous driving and monitoring of the Internet of Things.

Method used

The significance object detection method based on lightweight neural network is adopted, and feature extraction is used to use the MobileViT-XS backbone network to perform multimodal feature fusion through high- and low-level lightweight fusion modules and two-stage decoders to reduce model size and calculation cost.

Benefits of technology

While maintaining the model size low, the prediction performance of significance target detection has been improved, which can effectively process color visible light and infrared images, and improve perception capabilities in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119741580B_ABST
    Figure CN119741580B_ABST
Patent Text Reader

Abstract

The present invention discloses a saliency object detection method based on a lightweight neural network. A lightweight neural network is built, which mainly consists of a high-level lightweight fusion module, a low-level lightweight fusion module, and a two-stage decoder. The lightweight neural network is composed of a feature extraction module, a high-level lightweight fusion module, a low-level lightweight fusion module, and a two-stage decoder: The feature extraction module uses a MobileViT-XS backbone network to extract features from color visible light images and infrared images; the high-level lightweight fusion module and the low-level lightweight fusion module use convolutions to extract local information and use an improved cross-attention operation to narrow the gap between the visible light and infrared modalities; the two-stage decoder densely connects three high-level features to generate a guiding feature, which guides the fusion of two low-level features to produce the final result. The present invention can effectively improve the accuracy of saliency object detection with low parameter and computational amounts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of salient object detection, and particularly to a salient object detection method based on a lightweight neural network. Background Art

[0002] Salient object detection aims to capture and segment prominent objects in images or videos. As an important preprocessing step, salient object detection has been widely applied to computer vision and image processing tasks, such as image segmentation, object tracking, image retrieval, and image quality assessment. In recent years, convolutional / deep neural networks have pushed the performance of salient object detection to a new height due to their powerful learning ability and excellent performance in feature extraction. However, in some extreme scenarios (such as low illumination and cluttered scenes), it is often difficult to extract valuable information from color visible light images alone, which affects the effectiveness of salient object detection models under complex conditions. Different from color visible light images, thermal infrared cameras can capture thermal information that is less affected by environmental factors such as darkness and bad weather, thereby reflecting temperature differences, boundaries, and geometric shapes. Therefore, deploying thermal infrared devices to collect thermal information and using a color visible light and infrared image salient object detection model to record objects can improve the perception ability of the model in more scenarios.

[0003] Color visible light and depth image salient object detection with auxiliary depth images, and color visible light and infrared image salient object detection with infrared images have been developed, using widely available depth sensors and infrared cameras as additional modality information. Depth information contains rich spatial structure and 3D layout information, but it is unreliable to provide useful information for salient object detection in some extreme environments (e.g., poor illumination and cluttered scenes). Since infrared images can reflect the thermal radiation of the object surface, they are naturally complementary to color visible light images in these extreme environments, and in recent years, the research on color visible light and infrared image salient object detection has received increasing attention.

[0004] On the other hand, existing RGB-T salient object detection models have made significant breakthroughs in accuracy, but these high-precision models usually require high computational costs and large model sizes to process multi-modal information, which makes them only deployable in facilities with high-performance graphics processing units and unable to meet the requirements of some practical scenarios in terms of computational latency and bandwidth costs, such as autonomous driving and the Internet of Things for monitoring.

[0005] At present, there are already some lightweight RGB-T saliency object detection models, which are mainly based on convolutional neural networks (CNNs). These methods usually adopt lightweight backbone networks to extract features from two modalities, and then perform multi-modal and multi-level feature fusion through various carefully designed modules. Although these methods have successfully compressed the model size and can be deployed on some resource-constrained devices, it is still difficult to achieve a balance between performance and complexity. Therefore, how to enhance the prediction performance of lightweight RGB-T saliency object detection models while maintaining a low model size has become an urgent problem to be solved. Compared with the CNN structure, Transformer can model long-range dependencies, which helps to locate salient objects and suppress noise. However, compared with CNNs, the Transformer architecture has a larger computational cost and a longer response time. Therefore, in the lightweight RGB-T SOD task, it is particularly crucial to combine the advantages of Transformer and CNN in the feature extraction and fusion stages while maintaining a small model size. Summary of the Invention

[0006] The present invention provides a saliency object detection method based on a lightweight neural network, which can effectively improve the accuracy of saliency object detection with a low number of parameters and computational complexity.

[0007] The present invention provides a saliency object detection method based on a lightweight neural network. First, a training set containing pairs of color visible light images and their corresponding infrared images is constructed, and a lightweight neural network is built. Secondly, the pairs of color visible light images and their corresponding infrared images in the training set are input into the neural network for multiple rounds of network training. After the network training is completed, a neural network training model is obtained. Thirdly, the neural network training model is used to predict the test image pairs, and a saliency object image of the test image pairs is obtained. The lightweight neural network consists of a feature extraction module, a high-level lightweight fusion module, a low-level lightweight fusion module, and a two-stage decoder:

[0008] The feature extraction module uses the MobileViT-XS backbone network to extract features from color visible light images and infrared images;

[0009] The high-level lightweight fusion module and the low-level lightweight fusion module use convolutions to extract local information in total, and use an improved cross-attention operation to narrow the gap between the visible light and infrared modalities;

[0010] The two-stage decoder densely connects three high-level features to generate a guiding feature, which guides the fusion of two low-level features to produce the final result.

[0011] Furthermore, it further includes:

[0012] The described feature extraction module includes two MobileViT-XS backbone networks; the first layer of the first MobileViT-XS backbone network receives a color visible light image, and the output feature map is denoted as , the second layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as , the third layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as , the fourth layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as , the fifth layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as ; the first layer of the second MobileViT-XS backbone network receives an infrared image, and the output feature map is denoted as , the second layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as , the third layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as , the fourth layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as , the fifth layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as ;

[0013] The described low-level lightweight fusion module includes two low-level lightweight fusion blocks with the same structure; the first low-level lightweight fusion block receives and , and the output feature map is denoted as ; the second low-level lightweight fusion block receives and , and the output feature map is denoted as ;

[0014] The described high-level lightweight fusion module includes three high-level lightweight fusion blocks with the same structure; the first high-level lightweight fusion block receives and , and the output feature map is denoted as ; the second high-level lightweight fusion block receives and , and the output feature map is denoted as ; the third high-level lightweight fusion block receives and The output feature map is denoted as ;

[0015] The two-stage decoder receives , , , and The output feature map is denoted as , , , and use as the final saliency target image.

[0016] Furthermore, it also includes:

[0017] The construction process of the training set is as follows: Select at least 1000 pairs of original color visible light images and their corresponding original infrared images; then perform downsampling operations on each original color visible light image and its corresponding original infrared image, and the size of the downsampled images is ; then all the color visible light images with the size of and their corresponding infrared images form the training set; where ;

[0018] The process of obtaining the neural network training model is as follows: Input each pair of color visible light images in the training set and their corresponding infrared images into the neural network for network training, and calculate the loss function before the end of each round of network training to optimize the neural network, and obtain the neural network training model after a total of 150 rounds of network training; where , represents the final saliency target image output by the neural network, represents the label image, represents the th rough saliency target image obtained in the neural network, represents the binary cross-entropy loss, represents the intersection over union loss.

[0019] Furthermore, it also includes:

[0020] The process of using the neural network training model to predict the saliency target image for the test image pair is as follows: Arbitrarily select a pair of original color visible light images and their corresponding original infrared images; then perform downsampling operations on this pair of original color visible light images and original infrared images, and the size of the downsampled images is , and use them as the test image pair; then input the test image pair into the trained neural network model to predict the corresponding saliency target image.

[0021] Furthermore, it further includes:

[0022] The structures of the three high-level lightweight fusion blocks are the same. For the th high-level lightweight fusion block ( ), the input end of the convolutional layer receives ( ), the input end of the Batch Normalization layer receives the feature map output from the output end of this convolutional layer, the input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer, the input end of the convolutional layer receives the feature map output from the output end of this ReLU function, and the output feature map is denoted as ;

[0023] Another convolutional layer's input end receives , the input end of the Batch Normalization layer receives the feature map output from the output end of this convolutional layer, the input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer, the input end of the convolutional layer receives the feature map output from the output end of this ReLU function, and the output feature map is denoted as .

[0024] Furthermore, it further includes:

[0025] The structures of the two low-level lightweight fusion blocks are the same. For the th low-level lightweight fusion block , the input end of the convolutional layer receives , the input end of the Batch Normalization layer receives the feature map output from the output end of this convolutional layer, the input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer, the input end of the convolutional layer receives the feature map output from the output end of this ReLU function, and the output feature map is denoted as ;

[0026] Another convolutional layer's input end receives , the input end of the Batch Normalization layer receives the feature map output from the output end of this The feature map output at the output end of the convolutional layer, and the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output at the output end of this ReLU function, and the output feature map is denoted as .

[0027] Furthermore, it also includes:

[0028] The second-stage decoder receives , , , and , performs an upsampling operation on , and the output feature map is channel-connected with . The obtained feature map is input to the input end of the convolutional layer. The input end of the BatchNormalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output at the output end of this ReLU function. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, and the output feature map is denoted as

[0029] Furthermore, it also includes:

[0030] Perform an unfolding operation on and , convert from the format to the format, where , and , and denote the obtained feature maps as and ;

[0031] An input end of the fully connected layer receives , and the input end of the ReLU function receives the feature map output at the output end of this fully connected layer, and the output feature map is denoted as ;

[0032] An input end of the fully connected layer receives , and the input end of the ReLU function receives the feature map output at the output end of this The feature map output by the fully connected layer, and the output feature map is denoted as ;

[0033] Perform a channel connection operation on and and denote the obtained feature map as ;

[0034] An The input end of the fully connected layer receives and the output feature map is denoted as ;

[0035] An The input end of the fully connected layer receives , the input end of the Softmax function receives the feature map output by this fully connected layer, and the output feature map is denoted as ;

[0036] Perform an element-wise multiplication operation on and , and perform a summation operation on the N dimension in the format, and denote the obtained feature map as ;

[0037] Perform an element-wise multiplication operation on and , The input end of the fully connected layer receives the result of this element-wise multiplication operation, and the output feature map is denoted as , perform an element-wise addition operation on and and denote the obtained feature map as ; An The input end of the fully connected layer receives , the input end of the SiLU function receives the feature map output by the output end of this fully connected layer, an The input end of the fully connected layer receives the feature map output by the output end of this SiLU function, and the output feature map is denoted as , perform an element-wise addition operation on and and denote the obtained feature map as ;

[0038] Perform an element-wise multiplication operation on and , The input end of the fully connected layer receives the result of this element-wise multiplication operation, and the output feature map is denoted as , perform an element-wise addition operation on and Perform an element-wise addition operation and denote the resulting feature map as , a The input end of a fully connected layer receives , and the input end of the SiLU function receives this The feature map output from the output end of the fully connected layer, a The input end of the fully connected layer receives the feature map output from the output end of this SiLU function, and the output feature map is denoted as , for and Perform an element-wise addition operation and denote the resulting feature map as ;

[0039] Repeat once the process of generating and , where the previously generated and will serve as the and in the next process, and the final output obtained is denoted as and ; and ;

[0040] Perform a folding operation on and , converting from the format to the format, and the resulting feature map is denoted as and , where , and has the same size as and ;

[0041] Perform a channel connection operation on and , and the resulting feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output from the output end of the convolutional layer, and the input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer. The output feature map is denoted as , for and Perform an element-wise addition operation, and the resulting feature map is denoted as ;

[0042] Perform a channel connection operation on and , and the resulting feature map is input to At the input end of the convolutional layer, the input end of the Batch Normalization layer receives this feature map output from the output end of the convolutional layer. The input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer, and the output feature map is denoted as , for and perform an element-wise addition operation, and the resulting feature map is denoted as ;

[0043] For and perform a channel concatenation operation, and the resulting feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives this feature map output from the output end of the convolutional layer. The input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer, and the output feature map is denoted as , which is the feature map output from the output end of the

[0044] Furthermore, it also includes:

[0045] For and perform an unfolding operation, converting from the format to the format, where , and , and denote the resulting feature maps as and ;

[0046] An input end of a fully connected layer receives , the input end of the ReLU function receives the feature map output from this fully connected layer, and the output feature map is denoted as ;

[0047] For and perform a channel concatenation operation, and denote the resulting feature map as ;

[0048] An input end of a fully connected layer receives , and the output feature map is denoted as ;

[0049] An input end of a fully connected layer receives The input end of the Softmax function receives this feature map output by the fully connected layer, and the output feature map is denoted as ;

[0050] Perform an element-wise multiplication operation on and , and perform a summation operation on the N dimension in the format, and denote the obtained feature map as ;

[0051] Perform an element-wise multiplication operation on and , The input end of the fully connected layer receives the result of this element-wise multiplication operation, and the output feature map is denoted as , perform an element-wise addition operation on and , and denote the obtained feature map as , an The input end of the fully connected layer receives , the input end of the SiLU function receives this feature map output by the output end of the fully connected layer, an The input end of the fully connected layer receives the feature map output by the output end of this SiLU function, and the feature map output by the output end is denoted as , perform an element-wise addition operation on and , and denote the obtained feature map as ;

[0052] Repeat the process of generating and from , where the generated in the previous time will be used as in the next process, and the final obtained output is denoted as ;

[0053] Perform a folding operation on , convert from the format to the format, and denote the obtained feature map as , where , and has the same size as ;

[0054] Perform a channel connection operation on , and , and input the obtained feature map into the input end of the convolutional layer, and the input end of the Batch Normalization layer receives this The feature map output at the output end of the convolutional layer, and the ReLU function receives at its input end the feature map output at the output end of this Batch Normalization layer. The output feature map is denoted as , which is the feature map output at the output end of the

[0055] Furthermore, it further includes:

[0056] Performing an upsampling operation on and , and connecting the output feature map to through channel connection operation. The obtained feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output at the output end of this ReLU function. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The output feature map is denoted as

[0057] Input to the input end of the convolutional layer, and perform an upsampling operation on the output feature. The obtained feature map is denoted as , which is the first feature map output at the output end of the second-stage decoder;

[0058] Performing an upsampling operation on , and connecting the output feature map to through channel connection operation. The obtained feature map is denoted as . Input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the The feature map output at the output end of the convolutional layer, and the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The output feature map is denoted as ;

[0059] Input into the input end of the convolutional layer, perform an upsampling operation on the output features, and the obtained feature map is denoted as , as the second feature map output at the output end of the second-stage decoder;

[0060] Perform an upsampling operation on , and perform a channel connection operation on the output feature map and . The obtained feature map is denoted as . Input into the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output at the output end of this ReLU function. The input end of the BatchNormalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The output feature map is denoted as

[0061] Input into the input end of the convolutional layer, perform an upsampling operation on the output features, and the obtained feature map is denoted as , as the third feature map output at the output end of the second-stage decoder;

[0062] Perform an upsampling operation on , and perform a channel connection operation on the output feature map and . The obtained feature map is denoted as . Input into the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, and the input end of the Batch Normalization layer receives the feature map output by the output end of the convolutional layer. The input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, performs an upsampling operation on the output features, and the obtained feature map is denoted as , which is used as the fourth feature map output by the output end of the second-stage decoder.

[0063] The beneficial effects of the present invention are as follows:

[0064] 1) The lightweight neural network constructed by the method of the present invention uses the MobileViT-XS backbone network for feature extraction, extracts features from color visible light images and infrared images; then uses a high-level lightweight fusion module and a low-level lightweight fusion module, uses convolution to summarize local information, and uses an improved cross-attention operation to narrow the gap between the visible light and infrared modalities; finally, uses a second-stage decoder to densely connect three high-level features to generate a guiding feature, which guides the fusion of two low-level features to produce the final result.

[0065] 2) The method of the present invention combines the respective advantages of the CNN and Transformer structures, designs low-level lightweight fusion modules and high-level lightweight fusion modules with different complexities, and fully captures global and local information with fewer parameters.

[0066] 3) The method of the present invention addresses the problem that it is difficult to fuse multi-modal features with high quality. A second-stage decoder is adopted in the constructed neural network. This decoder can mine high-level feature information and make full use of its semantic information, thereby obtaining better results and effectively improving the accuracy of significant object detection in color visible light and infrared images. Description of the Drawings

[0067] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0068] Figure 1 is the overall implementation framework diagram of the method of the present invention;

[0069] Figure 2 is the schematic diagram of the composition structure of the neural network built by the method of the present invention;

[0070] Figure 3Schematic diagram of the composition structure of the high-level lightweight fusion block in the neural network built for the method of the present invention;

[0071] Figure 4 Schematic diagram of the remaining composition structure of the high-level lightweight fusion block in the neural network built for the method of the present invention;

[0072] Figure 5 Schematic diagram of the composition structure of the low-level lightweight fusion block in the neural network built for the method of the present invention;

[0073] Figure 6 Schematic diagram of the composition structure of the two-stage decoder in the neural network built for the method of the present invention. Detailed implementation manners

[0074] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0075] The technical solutions provided by the embodiments of the present invention will be described in detail below with reference to the drawings.

[0076] A saliency object detection method based on a lightweight neural network proposed by the present invention, the overall implementation framework diagram is as Figure 1 shown. The method first constructs a training set containing pairs of color visible light images and their corresponding infrared images, and builds a lightweight neural network; secondly, inputs the pairs of color visible light images and their corresponding infrared images in the training set into the neural network for multiple rounds of network training, and after the network training is completed, a neural network training model is obtained; thirdly, uses the neural network training model to predict the test image pair, and predicts the saliency object image of the test image pair, as Figure 2 shown. The lightweight neural network mainly consists of a feature extraction module, a high-level lightweight fusion module, a low-level lightweight fusion module, and a two-stage decoder.

[0077] The feature extraction module uses the MobileViT-XS backbone network to extract features from color visible light images and infrared images;

[0078] The high-level lightweight fusion module and the low-level lightweight fusion module use convolution to extract local information and use an improved cross-attention operation to narrow the gap between the visible light and infrared modalities;

[0079] The two-stage decoder densely connects three high-level features to generate a guiding feature, which guides the fusion of two low-level features to produce the final result.

[0080] The feature extraction module includes two MobileViT-XS backbone networks; the first layer of the first MobileViT-XS backbone network receives a color visible light image, and the output feature map is denoted as , the second layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as , the third layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as , the fourth layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as , the fifth layer of the first MobileViT-XS backbone network receives , and the output feature map is denoted as ; the first layer of the second MobileViT-XS backbone network receives an infrared image, and the output feature map is denoted as , the second layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as , the third layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as , the fourth layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as , the fifth layer of the second MobileViT-XS backbone network receives , and the output feature map is denoted as ; among them, the MobileViT-XS backbone network is an existing structural framework, and its network structure has been made public, as recorded in the literature Mehta S, Rastegari M. Mobilevit: light-weight, general-purpose, and mobile-friendly vision transformer[J]. arXiv preprint arXiv:2110.02178, 2021.

[0081] The low-level lightweight fusion module includes two low-level lightweight fusion blocks with the same structure; the first low-level lightweight fusion block receives and , and the output feature map is denoted as ; The second low-level lightweight fusion block receives and , and the output feature map is denoted as .

[0082] The high-level lightweight fusion module includes three high-level lightweight fusion blocks with the same structure; the first high-level lightweight fusion block receives and , and the output feature map is denoted as ; The second high-level lightweight fusion block receives and , and the output feature map is denoted as ; The third high-level lightweight fusion block receives and , and the output feature map is denoted as .

[0083] The second-stage decoder receives , , , and , and the output feature map is denoted as , , , , and uses as the final saliency target image.

[0084] In a specific embodiment, the construction process of the training set is as follows: at least 1000 pairs of original color visible light images and their corresponding original infrared images are selected; then, each original color visible light image and its corresponding original infrared image are subjected to downsampling operations, and the size of the downsampled images is ; Then, all color visible light images with a size of and their corresponding infrared images are used to form the training set; where .

[0085] In a specific embodiment, the process of obtaining the neural network training model is as follows: each pair of color visible light images in the training set and their corresponding infrared images are input into the neural network for network training, and the loss function is calculated before the end of each round of network training to optimize the neural network. After a total of 150 rounds of network training, the neural network training model is obtained; , represents the final saliency target image output by the neural network, represents the label image, represents the th rough saliency target image obtained in the neural network, represents the binary cross-entropy (BCE) loss, Indicates the Intersection over Union (IoU) loss.

[0086] In a specific embodiment, the process of using a neural network training model to predict a test image pair and obtaining the saliency target image of the test image pair is as follows: Arbitrarily select a pair of original color visible light images and their corresponding original infrared images; then perform downsampling operations on the pair of original color visible light images and original infrared images, and the size of the downsampled images is , and use them as the test image pair; then input the test image pair into the trained neural network model to predict the corresponding saliency target image; where .

[0087] The structures of the three high-level lightweight fusion blocks are the same, only the inputs and outputs are different. In a specific embodiment, as Figure 3 and Figure 4 shown, for the th high-level lightweight fusion block ( ), the input end of the convolutional layer receives ( ), the input end of the Batch Normalization layer receives the feature map output by the output end of this convolutional layer, the input end of the ReLU function receives the feature map output by the output end of this Batch Normalization layer, the input end of the convolutional layer receives the feature map output by the output end of this ReLU function, and the output feature map is denoted as .

[0088] Another convolutional layer's input end receives , the input end of the Batch Normalization layer receives the feature map output by the output end of this convolutional layer, the input end of the ReLU function receives the feature map output by the output end of this Batch Normalization layer, the input end of the convolutional layer receives the feature map output by the output end of this ReLU function, and the output feature map is denoted as .

[0089] Perform unfolding operations on and , convert from the format to the format, where , and , and denote the obtained feature maps as and .

[0090] One The input end of the fully connected layer receives , and the input end of the ReLU function receives this feature map output by the fully connected layer, and the output feature map is denoted as .

[0091] A input end of the fully connected layer receives , and the input end of the ReLU function receives this feature map output by the fully connected layer, and the output feature map is denoted as .

[0092] Perform a channel connection operation (Concatenation) on and , and denote the resulting feature map as .

[0093] A input end of the fully connected layer receives , and the output feature map is denoted as .

[0094] A input end of the fully connected layer receives , and the input end of the Softmax function receives this feature map output by the fully connected layer, and the output feature map is denoted as .

[0095] Perform an element-wise multiplication operation on and , and perform a summation operation on the N dimension in the format, and denote the resulting feature map as .

[0096] Perform an element-wise multiplication operation on and . The input end of the fully connected layer receives the result of this element-wise multiplication operation, and the output feature map is denoted as . Perform an element-wise addition operation on and , and denote the resulting feature map as . A input end of the fully connected layer receives , and the input end of the SiLU function receives this feature map output by the output end of the fully connected layer. An input end of the fully connected layer receives the feature map output by the output end of this SiLU function, and the output feature map is denoted as . Perform an element-wise addition operation on and Perform an element-wise addition operation and denote the resulting feature map as .

[0097] Perform an element-wise multiplication operation on and . The input end of the fully connected layer receives the result of this element-wise multiplication operation, and the output feature map is denoted as . Perform an element-wise addition operation on and and , and denote the resulting feature map as . An input end of a fully connected layer receives . The input end of the SiLU function receives the feature map output from the output end of this fully connected layer. An input end of a fully connected layer receives the feature map output from the output end of this SiLU function, and the output feature map is denoted as . Perform an element-wise addition operation on and , and denote the resulting feature map as .

[0098] As shown in Figure 3 , the process of generating and from and will be repeated once, where the previously generated and will be used as and in the next process. The final output is denoted as and .

[0099] As shown in Figure 4 , perform a folding operation on and , convert from the format to the format, and denote the resulting feature map as and , where , and it has the same size as and .

[0100] Perform a channel connection operation on and . The resulting feature map is input to the input end of the convolution layer. The input end of the Batch Normalization layer receives this The feature map output at the output end of the convolutional layer is received at the input end of the ReLU function. The output feature map is denoted as . Perform and element-wise addition operation, and the resulting feature map is denoted as .

[0101] Perform and channel concatenation operation, and the resulting feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, and the output feature map is denoted as . Perform and element-wise addition operation, and the resulting feature map is denoted as .

[0102] Perform and channel concatenation operation, and the resulting feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output at the output end of this convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, and the output feature map is denoted as . is the feature map output at the output end of the th high-level lightweight fusion block.

[0103] Here, element-wise multiplication operation, element-wise addition operation, element-wise subtraction operation, and channel concatenation operation (Concatenation) are all conventional operations in neural networks; Figure 3 , Figure 4 in which the BN layer is the abbreviation of the Batch Normalization layer.

[0104] The structures of two low-level lightweight fusion blocks are the same, only the inputs and outputs are different. In a specific embodiment, as shown in Figure 5 , for the th low-level lightweight fusion block , the input end of the convolutional layer receives , and the input end of the BatchNormalization layer receives this The feature map output at the output end of the convolutional layer, and the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output at the output end of this ReLU function, and the output feature map is denoted as .

[0105] Another The input end of the convolutional layer receives , and the input end of the Batch Normalization layer receives the output of this feature map output at the output end of the convolutional layer. The input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output at the output end of this ReLU function, and the output feature map is denoted as .

[0106] Perform an unfolding operation on and , converting from the format to the format, where , and , and denote the obtained feature maps as and .

[0107] A fully connected layer's input end receives , and the input end of the ReLU function receives the feature map output by this fully connected layer. The output feature map is denoted as .

[0108] Perform a channel connection operation (Concatenation) on and , and denote the obtained feature map as .

[0109] A fully connected layer's input end receives , and the output feature map is denoted as .

[0110] A fully connected layer's input end receives , and the input end of the Softmax function receives the feature map output by this fully connected layer. The output feature map is denoted as .

[0111] Perform an operation on and Perform element-wise multiplication and The N dimensions in the format are summed and the resulting feature map is recorded as .

[0112] right and Perform element-wise multiplication, The input of the fully connected layer receives the result of the element-wise multiplication operation, and the output feature map is recorded as .right and Perform element addition operation and record the obtained feature map as .one The input of the fully connected layer receives The input of the SiLU function receives the The feature map output at the output of the fully connected layer is a The input end of the fully connected layer receives the feature map output by the output end of the SiLU function, and the output feature map is recorded as .right and Perform element addition operation and record the obtained feature map as .

[0113] like Figure 3 As shown, from and generate The process will be repeated once more, with the last generated This will be used as the next step in the The final output is recorded as .

[0114] like Figure 4 As shown, Folding operation is performed by Format converted to Format, the obtained feature map is recorded as ,in , and with The same size.

[0115] right , and Perform channel connection operation and input the obtained feature map into The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as . That is the feature map output from the output end of the th low-level lightweight fusion block.

[0116] Here, the element-wise multiplication operation, element-wise addition operation, element-wise subtraction operation, and channel concatenation operation (Concatenation) are all conventional operations in neural networks; Figure 5 In , the BN layer is the abbreviation of the Batch Normalization layer.

[0117] As Figure 6 shown, the second-stage decoder receives , , , and , performs an upsampling operation on , and the output feature map is concatenated with through a channel connection operation. The resulting feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output from the output end of this convolutional layer. The input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer. The input end of the convolutional layer receives the feature map output from the output end of this ReLU function. The input end of the Batch Normalization layer receives the feature map output from the output end of this .

[0118] Perform an upsampling operation on and . The output feature map is concatenated with through a channel connection operation. The resulting feature map is input to the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output from the output end of this convolutional layer. The input end of the ReLU function receives the feature map output from the output end of this Batch Normalization layer. The input end of the The feature map output at the output end of the convolutional layer, the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, and the output feature map is denoted as 。

[0119] Input into the input end of the convolutional layer, perform upsampling on the output features, and the obtained feature map is denoted as , as the first feature map output at the output end of the second-stage decoder.

[0120] Perform upsampling on , connect the output feature map with through channel connection, and the obtained feature map is denoted as 。Input into the input end of the convolutional layer, the input end of the Batch Normalization layer receives the feature map output at the output end of this 1×1 convolutional layer, the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, the input end of the convolutional layer receives the feature map output at the output end of this ReLU function, the input end of the BatchNormalization layer receives the feature map output at the output end of the convolutional layer, the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, and the output feature map is denoted as 。

[0121] Input into the input end of the convolutional layer, perform upsampling on the output features, and the obtained feature map is denoted as , as the second feature map output at the output end of the second-stage decoder.

[0122] Perform upsampling on , connect the output feature map with through channel connection, and the obtained feature map is denoted as 。Input into the input end of the convolutional layer, the input end of the Batch Normalization layer receives the feature map output at the output end of the convolutional layer, the input end of the ReLU function receives the feature map output at the output end of this Batch Normalization layer, The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, and the input end of the Batch Normalization layer receives the feature map output by the output end of the convolutional layer. The input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer, and the output feature map is denoted as .

[0123] Input into the input end of the convolutional layer, perform upsampling on the output features, and the resulting feature map is denoted as , which is used as the third feature map output by the output end of the second-stage decoder.

[0124] Perform upsampling on , and perform channel connection operation on the output feature map and , and the resulting feature map is denoted as . Input into the input end of the convolutional layer. The input end of the Batch Normalization layer receives the feature map output by the output end of the convolutional layer. The input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer, the input end of the convolutional layer receives the feature map output by the output end of the ReLU function. The input end of the Batch Normalization layer receives the feature map output by the output end of the convolutional layer. The input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. Perform upsampling on the output features, and the resulting feature map is denoted as , which is used as the fourth feature map output by the output end of the second-stage decoder.

[0125] To further illustrate the feasibility and effectiveness of the method of the present invention, experiments are conducted on the method of the present invention.

[0126] In this embodiment, the method of the present invention is used to test the VT5000 dataset. The training set in VT5000 includes 2500 pairs of color visible light images and infrared images. The test set in VT5000 has a total of 2500 pairs of color visible light images and infrared images, that is, 2500 pairs of test image pairs.

[0127] In this embodiment, four commonly used objective parameters are selected to evaluate the performance of the method of the present invention, which are S-measure, E-measure, F-measure, and Mean Absolute Error (MAE). Table 1 shows the correlation between the saliency target image and the label image obtained by using the method of the present invention on the VT5000 dataset.

[0128] Table 1 S-measure, E-measure, F-measure, and Mean Absolute Error (MAE) between the saliency target image and the label image obtained by using the method of the present invention on the VT5000 dataset

[0129]

[0130] From the results given in Table 1, it can be found that the method of the present invention has achieved relatively high S-measure, E-measure, and F-measure and relatively low MAE on the existing color visible light and infrared image saliency target detection datasets, and has relatively low parameter quantity and computational complexity. This shows that the saliency target image obtained by the method of the present invention is relatively close to the label image, and the method of the present invention can effectively complete the saliency target detection of color visible light and infrared images with low parameter quantity and computational complexity.

[0131] The above are only embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A salient target detection method based on a lightweight neural network, firstly constructing a training set including several pairs of color visible light images and corresponding infrared images, and building a lightweight neural network; secondly, inputting several pairs of color visible light images and corresponding infrared images in the training set into the neural network for several network trainings, and obtaining a neural network training model after the network training is completed; and again using the neural network training model to predict the test image pairs to obtain a salient target image, characterized in that: The lightweight neural network is composed of a feature extraction module, a high-level lightweight fusion module, a low-level lightweight fusion module, and a two-stage decoder: The feature extraction module uses the MobileViT-XS backbone network to extract features from color visible light images and infrared images; High-level lightweight fusion modules and low-level lightweight fusion modules use convolution to extract local information and use improved cross-attention operations to narrow the gap between visible and infrared modalities; The two-stage decoder densely connects the three high-level features to generate a guiding feature, which guides the fusion of the two low-level features to produce the final result; The feature extraction module includes two MobileViT-XS backbone networks; the first layer of the first MobileViT-XS backbone network receives a color visible light image, and the output feature map is recorded as , the second layer receiving of the first MobileViT-XS backbone network , the output feature map is recorded as , the third layer receiving of the first MobileViT-XS backbone network , the output feature map is recorded as , the fourth layer receiving of the first MobileViT-XS backbone network , the output feature map is recorded as , the fifth layer receiving of the first MobileViT-XS backbone network , the output feature map is recorded as ; The first layer of the second MobileViT-XS backbone network receives the infrared image, and the output feature map is recorded as , the second layer receiving of the second MobileViT-XS backbone network , the output feature map is recorded as , the third layer receiving of the second MobileViT-XS backbone network , the output feature map is recorded as , the fourth layer receiving of the second MobileViT-XS backbone network , the output feature map is recorded as , the fifth layer receiving of the second MobileViT-XS backbone network , the output feature map is recorded as ; The low-level lightweight fusion module includes two low-level lightweight fusion blocks with the same structure; the first low-level lightweight fusion block receives and , the output feature map is recorded as ; The second low-level lightweight fusion block receives and , the output feature map is recorded as ; The high-level lightweight fusion module includes three high-level lightweight fusion blocks with the same structure; the first high-level lightweight fusion block receives and , the output feature map is recorded as ; The second high-level lightweight fusion block receives and , the output feature map is recorded as ; The third high-level lightweight fusion block receives and , the output feature map is recorded as ; The two-stage decoder receives , , , and , the output feature map is recorded as , , , , and As the final salient target image; The structures of the three high-level lightweight fusion blocks are the same. High-level lightweight fusion blocks , The input of the convolutional layer receives ,in The input of the Batch Normalization layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, and the output feature map is recorded as ; another The input of the convolutional layer receives The input of the Batch Normalization layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, and the output feature map is recorded as ; right and To expand the Format converted to Format, where ,and , and the obtained feature map is recorded as and ; one The input of the fully connected layer receives , the input of the ReLU function receives the The feature map output by the fully connected layer is recorded as ; one The input of the fully connected layer receives , the input of the ReLU function receives the The feature map output by the fully connected layer is recorded as ; right and Perform channel connection operation and record the obtained feature map as ; one The input of the fully connected layer receives , the output feature map is recorded as ; one The input of the fully connected layer receives , the input of the Softmax function receives the The feature map output by the fully connected layer is recorded as ; right and Perform element-wise multiplication and The N dimensions in the format are summed and the resulting feature map is recorded as ; right and Perform element-wise multiplication, The input of the fully connected layer receives the result of the element-wise multiplication operation, and the output feature map is recorded as ,right and Perform element addition operation and record the obtained feature map as ;one The input of the fully connected layer receives The input of the SiLU function receives the The feature map output at the output of the fully connected layer is a The input end of the fully connected layer receives the feature map output by the output end of the SiLU function, and the output feature map is recorded as ,right and Perform element addition operation and record the obtained feature map as ; right and Perform element-wise multiplication, The input of the fully connected layer receives the result of the element-wise multiplication operation, and the output feature map is recorded as ,right and Perform element addition operation and record the obtained feature map as ,one The input of the fully connected layer receives The input of the SiLU function receives the The feature map output at the output of the fully connected layer is a The input end of the fully connected layer receives the feature map output by the output end of the SiLU function, and the output feature map is recorded as ,right and Perform element addition operation and record the obtained feature map as ; Repeat from and generate and The process of and This will be used as the next step in the and The final output is recorded as and ; right and Folding operation is performed by Format converted to Format, the obtained feature map is recorded as and ,in, , and with and The same size; right and Perform channel connection operation and input the obtained feature map into The input of the convolutional layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as ,right and Perform element addition operation, and the obtained feature map is recorded as ; right and Perform channel connection operation and input the obtained feature map into The input of the convolutional layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as ,right and Perform element addition operation, and the obtained feature map is recorded as ; right and Perform channel connection operation and input the obtained feature map into The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as , That is the The feature map output by the output of the high-level lightweight fusion block.

2. The method for detecting salient objects based on a lightweight neural network according to claim 1, characterized in that: The training set construction process is as follows: select at least 1000 pairs of original color visible light images and their corresponding original infrared images; then downsample each original color visible light image and its corresponding original infrared image, and the size of the downsampled image is ; Then all the sizes are The color visible light images and the corresponding infrared images constitute the training set; among them, ; The process of obtaining the neural network training model is as follows: each pair of color visible light images and their corresponding infrared images in the training set are input into the neural network for network training, and the loss function is calculated before each round of network training ends. To optimize the neural network, a neural network training model was obtained after a total of 150 rounds of network training; among them, , represents the final salient target image output by the neural network, represents the label image, Represents the first A coarse salient target image, represents the binary cross entropy loss, represents the intersection-and-combination loss.

3. The method for detecting salient objects based on a lightweight neural network according to claim 1, characterized in that: The process of using the neural network training model to predict the test image pair to obtain the salient target image is as follows: arbitrarily select a pair of original color visible light images and the corresponding original infrared images; then downsample the pair of original color visible light images and original infrared images, and the size of the downsampled image is , and used as a test image pair; the test image pair is then input into the trained neural network model to predict the corresponding salient target image.

4. The method for detecting salient objects based on a lightweight neural network according to claim 1, characterized in that: The structures of the two low-level lightweight fusion blocks are the same. Low-level lightweight fusion blocks , The input of the convolutional layer receives The input of the Batch Normalization layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, and the output feature map is recorded as ; another The input of the convolutional layer receives The input of the Batch Normalization layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, and the output feature map is recorded as .

5. The method for detecting salient objects based on a lightweight neural network according to claim 4, characterized in that: Two-stage decoder reception , , , and ,right Perform upsampling operation, the output feature map is Perform channel connection operation and input the obtained feature map into The input of the convolutional layer receives the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolution layer receives the feature map output by the output end of the ReLU function, and the input end of the Batch Normalization layer receives the feature map output by the ReLU function. The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as .

6. The method for detecting salient objects based on a lightweight neural network according to claim 5, characterized in that: Also includes: right and To expand the Format converted to Format, where ,and , and the obtained feature map is recorded as and ; one The input of the fully connected layer receives , the input of the ReLU function receives the The feature map output by the fully connected layer is recorded as ; right and Perform channel connection operation and record the obtained feature map as ; one The input of the fully connected layer receives , the output feature map is recorded as ; one The input of the fully connected layer receives , the input of the Softmax function receives the The feature map output by the fully connected layer is recorded as ; right and Perform element-wise multiplication and The N dimensions in the format are summed and the resulting feature map is recorded as ; right and Perform element-wise multiplication, The input of the fully connected layer receives the result of the element-wise multiplication operation, and the output feature map is recorded as ,right and Perform element addition operation and record the obtained feature map as ,one The input of the fully connected layer receives The input of the SiLU function receives the The feature map output at the output of the fully connected layer is a The input end of the fully connected layer receives the feature map output by the output end of the SiLU function, and the feature map output by the output end is recorded as ,right and Perform element addition operation and record the obtained feature map as ; Repeat from and generate The process of This will be used as the next step in the The final output is recorded as ; right Folding operation is performed by Format converted to Format, the obtained feature map is recorded as ,in , and with The same size; right , and Perform channel connection operation and input the obtained feature map into The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as , That is the The feature map output by the output of the low-level lightweight fusion block.

7. The method for detecting salient objects based on a lightweight neural network according to claim 6, characterized in that: Also includes: right and Perform upsampling operation, the output feature map is Perform channel connection operation and input the obtained feature map into The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolution layer receives the feature map output by the output end of the ReLU function, and the input end of the Batch Normalization layer receives the feature map output by the ReLU function. The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the BatchNormalization layer. The output feature map is recorded as ; Will Input to At the input end of the convolutional layer, the output features are upsampled, and the resulting feature map is recorded as , The first feature map output as the output of the two-stage decoder; right Perform upsampling operation, the output feature map is Perform channel connection operation, and the obtained feature map is recorded as ,Will Input to The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolution layer receives the feature map output by the output end of the ReLU function, and the input end of the BatchNormalization layer receives the feature map output by the output end of the ReLU function. The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as ; Will Input to At the input end of the convolutional layer, the output features are upsampled, and the resulting feature map is recorded as , The second feature map output as the output of the two-stage decoder; right Perform upsampling operation, the output feature map is Perform channel connection operation, and the obtained feature map is recorded as ,Will Input to The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolution layer receives the feature map output by the output end of the ReLU function, and the input end of the BatchNormalization layer receives the feature map output by the output end of the ReLU function. The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The output feature map is recorded as ; Will Input to At the input end of the convolutional layer, the output features are upsampled, and the resulting feature map is recorded as , The third feature map output as the output of the two-stage decoder; right Perform upsampling operation, the output feature map is Perform channel connection operation, and the obtained feature map is recorded as ,Will Input to The input of the convolutional layer and the input of the Batch Normalization layer receive the The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolution layer receives the feature map output by the output end of the ReLU function, and the input end of the BatchNormalization layer receives the feature map output by the output end of the ReLU function. The output end of the convolutional layer outputs the feature map, and the input end of the ReLU function receives the feature map output by the output end of the Batch Normalization layer. The input end of the convolutional layer receives the feature map output by the output end of the ReLU function, performs upsampling on the output features, and the obtained feature map is recorded as , The fourth feature map is output as the output of the two-stage decoder.

Citation Information

Patent Citations

  • Colorful visible light and infrared image saliency target detection method

    CN116863150A

  • Visible light-thermal infrared salient target detection method based on bidirectional alternating fusion strategy

    CN118038228A

  • Green fruit camouflage target detection method

    CN118154855A