Light field salient target detection method based on multi-modal feature fusion

By using Slice Interleaving Enhancement Module and Multimodal Feature Fusion Strategy built by Swin Transformer in the light field significant object detection, the problem of insufficient feature fusion in the existing technology is solved, and more efficient significant object detection performance is achieved, especially in complex scenarios.

CN119942289AActive Publication Date: 2025-05-06CHONGQING UNIV OF TECH

Patent Information

Application Number
CN202510102232.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The existing light field significant object detection methods are difficult to achieve effective feature fusion, and light field information is not fully mined, especially in scenes where complex backgrounds and lighting conditions change.

Method used

The multimodal feature fusion strategy based on Swin Transformer is adopted, which is a slice interleaving enhancement module, a high-level feature fusion module, a low-level cross attention module and a compact pyramid refinement module, and multimodal semantic information is fused locally and globally, and feature aggregation and decoding are performed.

Benefits of technology

It effectively enhances feature learning between focal slices, makes full use of multimodal information in light field images, improves the accuracy and robustness of significant object detection, and performs well in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942289A_ABST
    Figure CN119942289A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal feature fusion-based light field salient target detection method. The method comprises the following steps of: performing detection by using a trained light field salient target detection model to obtain a salient map of a light field multi-modal image; according to the model, features between focus slices are enhanced in a local mode through a slice interleaving enhancement module, enhanced focus stream features are obtained, feature fusion is carried out on high-level multi-mode semantic information in a local and global mode through a high-level feature fusion module, and a fused high-level feature map is output and obtained. The low-layer multi-mode space-channel features are fused in a cross enhancement mode through low-layer cross attention, a fused low-layer feature map is output, and finally, the features are aggregated through a compact pyramid refining module and decoded into an accurate saliency map. According to the method, the light field salient target detection model is constructed, so that the problems of insufficient feature learning between slices in a focus stack and insufficient multi-modal information feature fusion in a light field are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of light field image salient object detection, and in particular to a light field salient object detection method based on multimodal feature fusion. Background Art

[0002] Salient Object Detection (SOD) aims to locate and segment the most eye-catching areas in an image by simulating human vision. As an image preprocessing technique, it has been widely used in many visual tasks such as object detection, semantic segmentation, depth estimation, and image quality assessment. SOD methods based on red, green, and blue (RGB) are limited by the limited field of view of RGB images, and the detection effect is not ideal in scenes with complex backgrounds and changing lighting conditions. SOD methods based on RGB depth maps (RGB Depth, RGB-D) improve the detection performance in these complex environments, but it is very difficult to obtain high-quality depth maps. Unreliable depth maps (such as foreground and background at the same depth) may even reduce SOD performance. Unlike depth maps, light field cameras can provide spatial and angular information of light in the scene, and contain richer geometric representations of the scene, which is expected to improve SOD performance in complex scenes. The original light field information can be further converted into focus stacks and fully focused images for the light field salient object detection (LFSOD) task. At present, most light field salient object detection methods have difficulty in achieving effective feature fusion, and light field information has not been fully mined in the SOD task. Therefore, developing effective light field salient object detection methods has important practical significance.

[0003] In multimodal features, the focus stack focuses at different depths and blurs other areas, and the fully focused image can depict the entire scene. However, both images contain redundant information. All focus stack images depict the same scene, only focusing at different depths; non-salient targets are recorded in the fully focused image where all targets are clear. Some existing methods process focus stack images through attention mechanisms and ConvLSTM, but ConvLSTM treats all focus slices equally, cannot correctly distinguish foreground from background in some challenging scenes, and introduces huge computational complexity. Since multimodal features of different scales are rich in high-level semantic information and low-level spatial information. Other methods only fuse high-level semantic features to avoid background interference, but ignore the fact that low-level spatial information can help refine the contours of salient targets. Some methods use the same fusion strategy for all scale features, but since the different attributes of spatial details and semantic information play different roles in fusion, treating all features equally may reduce the accuracy of detection. The decoding stage uses a simple method of aggregating multi-scale information through element-by-element addition or concatenation operations. Although this can avoid the tedious task of designing a separate decoder, it is difficult to balance the differences in resolution, receptive field, and semantic information among feature maps of different scales, resulting in redundant background interference.

[0004] Therefore, how to utilize the complementary relationship between multimodal information of light fields and design an effective light field SOD method has become an urgent problem to be solved by those skilled in the art. Summary of the invention

[0005] In view of the above-mentioned deficiencies of the prior art, the present invention provides a light field salient target detection method based on multimodal feature fusion, so as to solve the problems of insufficient feature learning between slices in focal stack and insufficient feature fusion of multimodal information in light field.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A light field salient object detection method based on multimodal feature fusion comprises the following steps:

[0008] S1, obtaining a light field multimodal focal stack image and a fully focused image to be processed;

[0009] S2, using the light field multimodal focal stack image and the fully gathered image to be processed as inputs of the trained light field salient target detection model, and outputting light field salient target detection results of the light field multimodal focal stack image and the fully gathered image to be processed; the light field salient target detection model includes a slice interleaving enhancement module, a high-level feature fusion module, a low-level cross attention module, and a compact pyramid refinement module;

[0010] The training process of the light field salient object detection model includes:

[0011] S201, taking the light field multimodal focus stack image and the fully focused image as inputs of the light field salient target detection model as training samples, inputting them into the light field salient target detection model with Pvtv2 as the backbone network, performing feature extraction on the training samples, and respectively extracting four focal stack feature maps and fully focused feature maps of different scales;

[0012] S202, constructing the slice interleaving enhancement module based on the Swin-T module, taking the focus stack feature maps of the four different scales as the input of the slice interleaving enhancement module, enhancing the features between the focus slices in a local manner, and obtaining enhanced focus flow features;

[0013] S203, the four fully focused feature maps of different scales obtained by Pvtv2 are passed through the receptive field block to obtain an enhanced fully focused flow feature; the high-level features in the enhanced fully focused flow feature and the enhanced focus flow feature are used as inputs of the high-level feature fusion module, and the high-level multimodal semantic information is locally and globally feature-fused to obtain a fused high-level feature map as output;

[0014] S204, using the enhanced full focus flow features and the low-level features in the enhanced focus flow features as inputs of the low-level cross attention module, fusing the low-level multimodal spatial-channel features in a cross-enhancement manner, and outputting a fused low-level feature map;

[0015] S205, using the fused high-level feature map and the fused low-level feature map as inputs of the compact pyramid refinement module, aggregating multi-scale features from top to bottom, and outputting a final saliency map;

[0016] S206, using a total loss function composed of a slice interleaving enhancement loss function and a hybrid structure loss function to calculate the training loss of the light field salient object detection model, and updating the parameters of the light field salient object detection model with the goal of minimizing the total loss function;

[0017] S207 , repeating steps S201 to S206 until the light field salient object detection model converges or reaches a preset number of training times, and the training is completed.

[0018] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, in step S202, the slice interleaving enhancement module includes a cascaded Swin-T module, four parallel receptive field blocks and three upsampling units, wherein the Swin-T module includes a local window self-attention unit and a sliding window mask unit;

[0019] In the slice interleaving enhancement module, the focal stack feature map of each scale of the input is passed through four parallel receptive field blocks to obtain the corresponding focal slice feature map; then, the local window self-attention unit in the Swin-T module is used to divide the fourth-layer focal slice feature map into multiple local window blocks, and the multi-head self-attention mechanism is used to calculate the local similarity. Then, the sliding window mask unit in the Swin-T module is used to slide on the local feature map, and the global feature map is obtained through the mask operation; finally, the fourth-layer focal slice feature map is fused with the features obtained after passing through the Swin-T module through element-by-element multiplication operation, and the fused feature map is used as the input of the third upsampling unit. After the output of each upsampling unit is fused with the corresponding focal slice feature map, it is used as the input to the previous layer, and the focal slice feature maps of all levels are enhanced in turn to obtain the corresponding enhanced focal slice feature map.

[0020] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, the processing process of the slice interleaving enhancement module is:

[0021]

[0022] In the formula, RFB(·) represents the output after the receptive field block operation, represents the focal stack feature map after Pvtv2, represents the focal slice feature map output by the receptive field block, n represents the number of focal slices (n = 1, 2, ..., 12), i represents the number of layers of multi-scale features (i = 1, 2, 3, 4), represents the enhanced focal flow feature, Swin(·) represents the result after Swin-T module, U(·) represents bilinear upsampling, Represents element-wise multiplication.

[0023] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, in step S203, the processing process of the high-level feature fusion module includes:

[0024] In the first step of the high-level feature fusion module, the fourth-layer enhanced focal slice feature map in the slice interleaving enhancement module and the fourth-layer enhanced full-focus feature map obtained by processing the fourth receptive field block are feature fused to obtain the fourth-layer fusion total output; then, in the second step, the fourth-layer fusion global and local outputs and the third-layer enhanced focal slice feature map in the slice interleaving enhancement module and the third-layer enhanced full-focus feature map obtained by processing the third receptive field block are used as input to reconstruct the high-level feature fusion module, and the third-layer fusion total output is obtained as the total output of the high-level feature fusion module.

[0025] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, the processing process of the high-level feature fusion module is:

[0026] The first step of the process is expressed as:

[0027]

[0028]

[0029] Where A4 represents the fourth-layer similarity matrix, and They represent the global, local and total outputs of the fourth layer fusion, respectively. Softmax(·) represents the softmax operation, and mul(·) represents the matrix multiplication. represents the result after reshaping key K4, represents the result after reshaping the query Q4, γ represents the learnable parameter, conv 1×1 (·) represents 1×1 convolution, V4 represents the value generated by the full-focus feature guidance after the fourth layer enhancement, conv 3×3×3 (·) represents a 3×3×3 convolution, CS(·) represents a channel shuffle operation, cat(·) represents a concatenation operation, and X a (4) represents the fourth layer enhanced full focus feature map obtained by processing the fourth receptive field block, represents the focal slice feature map after the fourth layer enhancement in the slice interleaving enhancement module;

[0030] The second step process is expressed as:

[0031]

[0032] In the formula, and Represent the global and local outputs of the third layer fusion, P H represents the total output of the high-level feature fusion module, V3 represents the value generated by the full-focus feature guidance after the third layer enhancement, A3 represents the similarity matrix of the third layer, and X a (3) represents the fully focused feature map after the third layer enhancement, Represents the focal slice feature map after the third layer enhancement.

[0033] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, in step S204, the low-level cross attention module includes two groups of cascaded space-channel attention units and convolutional layers, and the two groups of space-channel attention units are connected in parallel, wherein the space-channel attention unit includes parallel orthogonal channel attention subunits and space attention subunits;

[0034] In the low-level cross-attention module, the low-level enhanced full focal flow features are copied along the batch dimension to align the focal flow, and the copied low-level full focal flow features and the low-level enhanced focal flow features are respectively used as the input of two groups of space-channel attention units, and the outputs are respectively output to generate their own orthogonal channel attention weights and spatial attention weights. The copied low-level full focal flow features and the low-level enhanced focal flow features are respectively multiplied with the orthogonal channel attention weights and spatial attention weights generated by each other, and the output results are spliced ​​and aggregated through the convolution layer to obtain the fused low-level feature map as the total output of the low-level cross-attention module.

[0035] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, the processing process of the low-level cross attention module is:

[0036]

[0037] P CA (i) = conv 3×3 (cat(T a (i),T fs (i)));

[0038] Where P CA (i) represents the output result of the low-level cross-attention module, T a (i) and T fs (i) represents the full focus feature and focus feature enhanced by dual attention, respectively. replicate(.) represents the replication operation along the batch dimension. represents the spatial attention weight, represents the orthogonal channel attention weight.

[0039] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, in step S205, the processing process of the compact pyramid refinement module is:

[0040] S final =conv 1×1 (CPR(CPR(P H ,P CA (2)),P CA (1)));

[0041] In the formula, S final represents the final saliency map, and CPR(·) represents the CPR refinement operation.

[0042] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, in step S206, the slice interleaving enhancement loss function is:

[0043]

[0044] Where, L E represents the slice interleaving enhancement loss function, L S represents the hybrid structure loss function, S fa (i) represents a rough saliency map, and G represents the ground truth.

[0045] In the above-mentioned light field salient object detection method based on multimodal feature fusion, as a preferred solution, the total loss function is:

[0046] L total =L E +L S (S final ,G);

[0047] Where, L total Represents the total loss function.

[0048] Compared with the prior art, the present invention has the following technical effects:

[0049] (1) The existing light field salient target detection method uses rigid ConvLSTM to process focal slice features, ignoring the unique structural information between focal slices. The present invention constructs a slice interleaving enhancement module based on Swin Transformer, which makes full use of the geometric relationship between focal slices to segment the foreground and background. The module uses the Swin-T module to learn the interleaving information of different focal slices in a local manner, and uses multi-head self-attention to calculate the local similarity of the local window block, so as to learn the focal interleaving information within the local window block of the whole group of slices, and uses a cyclic sliding window to model the global information. At the same time, the present invention follows the Swin Transformer to use masked multi-head self-attention calculation and then restore, so as to avoid the interference of information in different regions and the destruction of semantic information.

[0050] (2) The existing light field salient target detection methods do not adequately integrate multimodal information features, ignoring the unique structural characteristics of the two modal information. The present invention proposes a multimodal feature fusion strategy consisting of a high-level feature fusion module, a low-level cross-attention module, and a compact pyramid refinement module. Through the high-level feature fusion module and the low-level cross-attention module, the model can make more full use of the multimodal information in the light field image, and perform feature aggregation and decoding through the compact pyramid refinement module, thereby optimizing the computational efficiency. This strategy fully considers the complementary relationship between multimodal and multi-scale features, thereby enabling the network to pay more attention to the salient target information. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0052] Figure 1 The present invention discloses a flow chart of a light field salient target detection method based on multimodal feature fusion;

[0053] Figure 2 Schematic diagram of the overall architecture of the light field salient object detection model of the present invention;

[0054] Figure 3 is a schematic diagram of the spatial alignment characteristics of the focal slices of this embodiment;

[0055] Figure 4 Schematic diagram of Swin-T processing focus slice features in this embodiment;

[0056] Figure 5 is a schematic diagram of a high-level feature fusion module of this embodiment;

[0057] Figure 6 The figure shows the prediction results of the model of this embodiment and other methods in challenging scenarios. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but only represents selected embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work belong to the scope of protection of the present invention.

[0059] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0060] Existing salient object detection methods have insufficient performance when dealing with scenes such as complex backgrounds, changing lighting conditions, and occlusions. However, light field cameras can capture rich spatial and angular information, but most light field salient object detection methods fail to fully utilize this information and find it difficult to achieve effective feature fusion. For multimodal features, some existing methods process focal stack images through attention mechanisms and ConvLSTM, but ConvLSTM treats all focal slices equally, resulting in problems such as background interference and information loss.

[0061] In order to overcome the above problems, the present invention proposes a light field salient target detection method based on multimodal feature fusion. The method process is as follows: Figure 1 As shown, including:

[0062] S1, obtaining a light field multimodal focal stack image and a fully focused image to be processed;

[0063] S2, using the light field multimodal focal stack image and the fully gathered image to be processed as inputs of the trained light field salient target detection model, and outputting light field salient target detection results of the light field multimodal focal stack image and the fully gathered image to be processed; the light field salient target detection model includes a slice interleaving enhancement module, a high-level feature fusion module, a low-level cross attention module, and a compact pyramid refinement module;

[0064] The training process of the light field salient object detection model includes:

[0065] S201, taking the light field multimodal focus stack image and the fully focused image as inputs of the light field salient target detection model as training samples, inputting them into the light field salient target detection model with Pvtv2 as the backbone network, performing feature extraction on the training samples, and respectively extracting four focal stack feature maps and fully focused feature maps of different scales;

[0066] S202, constructing the slice interleaving enhancement module based on the Swin-T module, taking the focus stack feature maps of the four different scales as the input of the slice interleaving enhancement module, enhancing the features between the focus slices in a local manner, and obtaining enhanced focus flow features;

[0067] S203, the four fully focused feature maps of different scales obtained by Pvtv2 are passed through the receptive field block to obtain an enhanced fully focused flow feature; the high-level features in the enhanced fully focused flow feature and the enhanced focus flow feature are used as inputs of the high-level feature fusion module, and the high-level multimodal semantic information is locally and globally feature-fused to obtain a fused high-level feature map as output;

[0068] S204, using the enhanced full focus flow features and the low-level features in the enhanced focus flow features as inputs of the low-level cross attention module, fusing the low-level multimodal spatial-channel features in a cross-enhancement manner, and outputting a fused low-level feature map;

[0069] S205, using the fused high-level feature map and the fused low-level feature map as inputs of the compact pyramid refinement module, aggregating multi-scale features from top to bottom, and outputting a final saliency map;

[0070] S206, using a total loss function composed of a slice interleaving enhancement loss function and a hybrid structure loss function to calculate the training loss of the light field salient object detection model, and updating the parameters of the light field salient object detection model with the goal of minimizing the total loss function;

[0071] S207 , repeating steps S201 to S206 until the light field salient object detection model converges or reaches a preset number of training times, and the training is completed.

[0072] The present invention aims at the problem of insufficient multimodal feature fusion that has not been solved in existing light field salient target detection methods, and proposes a light field salient target detection method based on multimodal feature fusion. On the one hand, the slice interleaving enhancement module extracts multiple focal slice information in a local manner and enhances four-layer multi-scale information to solve the problem of insufficient feature learning between slices, thereby effectively enhancing the geometric features of the focal slices. On the other hand, a multimodal feature fusion strategy is constructed by a high-level feature fusion module, a low-level cross-attention module and a compact pyramid refinement module to solve the problem of insufficient multimodal feature fusion.

[0073] The light field salient object detection method based on multimodal feature fusion proposed by the present invention is described in more detail below.

[0074] 1. Image preprocessing

[0075] In the specific application implementation, the light field data used for training and testing the light field salient object detection model can be selected from the public light field data sets, namely DUTLF-FS, HFUT and LFSD. Among them, 1000 samples in DUTLF-FS and 100 samples in HFUT are selected for training, and the remaining samples are all used for testing. All light field samples consist of focus stack images, full focus images and corresponding truth maps, with a resolution of 256×256. The data is enhanced by random rotation, cropping and flipping and then input into the network for training.

[0076] 2. Light field salient object detection model

[0077] In the specific application implementation, in order to adapt to the task of salient target detection, Pvtv2 was selected as the backbone network of the light field salient target detection model to extract focal stack and full focus features respectively, and four feature maps of different scales were obtained, which were 1 / 4, 1 / 8, 1 / 16 and 1 / 32 of the input image respectively.

[0078] like Figure 2As shown, the overall architecture of the light field salient target detection model proposed in the method of the present invention includes a slice interweaving enhancement module (SIEM), a high-level feature fusion module (HFFM), a low-level cross attention module (Cross Attention Module, CrossA) and a compact pyramid refinement module (CPR). Among them, the slice interweaving enhancement module is used to enhance the focus feature maps of four different scales obtained by Pvtv2, the high-level feature fusion module is used to fuse the high-level multimodal semantic information in a local and global manner, the low-level cross attention module is used to realize the low-level multimodal space-channel feature fusion in a cross-enhancement manner, and the compact pyramid refinement module aggregates the fused multi-scale light field features and decodes them into an accurate salient map.

[0079] Next, each module of the light field salient object detection model is explained in detail.

[0080] 2.1 Slice interleaving enhancement module

[0081] In order to solve the problem of insufficient feature learning between light field focus slices, the present invention combines Figure 3 Based on the spatial alignment of the focal stack shown in the figure, a slice interleaving enhancement module (SIEM) is constructed based on Swin Transformer. This module uses Swin-Transformer Block (Swin-T) to learn the interleaving information of different focal slices in a local manner, such as Figure 4 As shown, the focal slice feature maps of different scales are enhanced.

[0082] In this embodiment, the four-layer dual-stream features of Pvtv2 are first passed through a lightweight receptive field block to enhance the distinguishing ability and robustness of multimodal features, and the channels are uniformly compressed to 32. The specific expression is:

[0083] X a (i) = RFB(F a (i));

[0084]

[0085] Where RFB(·) represents the output after the receptive field block operation, F a (i) and represents the full focus feature map and focal stack feature map after Pvtv2, X a(i) represents the enhanced fully focused feature map output by the receptive field block, where the set of 4 layers of enhanced fully focused feature maps is the enhanced fully focused flow feature. represents the focal slice feature map output by the receptive field block, n represents the number of focal slices (n = 1, 2, ..., 12), and i represents the number of layers of multi-scale features (i = 1, 2, 3, 4).

[0086] The slice interleaving enhancement module includes a cascaded Swin-T module, four parallel receptive field blocks and three upsampling units, wherein the Swin-T module includes a local window self-attention unit and a sliding window mask unit;

[0087] In the slice interleaving enhancement module, the focal stack feature map of each scale of the input is passed through four parallel receptive field blocks to obtain the corresponding focal slice features; then, the local window self-attention unit in the Swin-T module is used to divide the fourth-layer focal stack feature map into multiple local window blocks, and the multi-head self-attention mechanism is used to calculate the local similarity. Then, the sliding window mask unit in the Swin-T module is used to slide on the local feature map, and the global feature map is obtained through the mask operation; finally, the fourth-layer focal slice features are fused with the features obtained after passing through the Swin-T module through element-by-element multiplication operation, and the fused feature map is used as the input of the third upsampling unit. After the output of each upsampling unit is fused with the corresponding focal slice features, it is used as the input to the previous layer, and the focal slice features of all levels are enhanced in turn to obtain the corresponding enhanced focal slice features.

[0088] Specifically, the slice interleaving enhancement module is divided into three parts during the processing: the first part performs self-attention on the local window block to calculate the local similarity, the second part uses sliding window and mask operation to summarize the global information, and the third part uses the results after Swin-T to enhance the four-layer focus slice flow features respectively. For the first part, given the focus slice feature Where H, W and C represent the height, width and number of channels of the image. 12×N×C ), where N = H × W. The slice feature is decomposed into m local window blocks, such as Figure 4As shown in the red box, multi-head self-attention (MSA) is performed on the box to calculate the similarity, and the focus interlacing information in the entire set of slice frames can be learned. Since there is only local interaction within the window in the first part and the global information is not modeled, the operation of cyclic sliding window is used in the second part to model the global information. The window is slid 1 / 2 window size to the right and downward from the upper left corner, and the results after the movement are combined into a new image group, so that a window block with the same computational cost as the first image group is obtained. Since the global information of the recombined image group is chaotic, the present invention follows SwinTransformer to use masked MSA calculation and then restore. This operation can isolate the interference of information in different regions and avoid destroying semantic information. Please note that since the deepest layer features (when i=4) have the richest semantic features, the background interference information of the lower layer will have an adverse effect on the result, so the present invention only uses the features of the last layer. For the third part, the focus interlacing information obtained after the fourth layer features are passed through Swin-T is enhanced by upsampling operations on the focus flow features of all levels.

[0089] Therefore, the processing of the slice interleaving enhancement module is expressed as:

[0090]

[0091] In the formula, represents the enhanced focal flow feature, Swin(·) represents the result after Swin-T module, U(·) represents bilinear upsampling, Represents element-wise multiplication.

[0092] 2.2 High-level feature fusion module

[0093] The focus stack focuses on a specific depth and blurs other areas, while the fully focused image depicts the entire scene. The information of these two modalities is complementary. For the task of salient object detection, it is very critical to effectively fuse multimodal complementary information. High-resolution low-level features have rich spatial details, and low-resolution high-level features have rich semantic information. In view of these different characteristics, this embodiment proposes to use different methods to fuse features of different layers.

[0094] Convolution performs well in processing local features, but its receptive field is limited by the size of the convolution kernel, making it difficult to model the global situation. Compared with convolution, Transformer performs very well in processing global information and long-range dependencies in various tasks. Therefore, the construction of the High-Level Feature Fusion Module (HFFM) adopts a combination of Transformer and convolution. Transformer is used to fuse global semantic information to reduce inconsistency, and convolution is used to supplement local information to accurately locate salient targets.

[0095] like Figure 5 As shown, in this embodiment, the module is designed as two steps: in the first step of the high-level feature fusion module, the fourth-layer enhanced focus slice feature map in the slice interleaving enhancement module and the fourth-layer enhanced full-focus feature map obtained by processing the fourth receptive field block are feature fused to obtain the fourth-layer fusion total output; then, in the second step, the fourth-layer fusion global and local outputs and the third-layer enhanced focus slice feature map in the slice interleaving enhancement module and the third-layer enhanced full-focus feature map obtained by processing the third receptive field block are used as input to reconstruct the high-level feature fusion module, and the third-layer fusion total output is obtained as the total output of the high-level feature fusion module.

[0096] Specifically, for the first step, given the input fourth layer enhanced focal slice feature And the fourth layer of enhanced all-focus feature X a (4)∈R 1×32×8×8 , first reshape the focus feature to match the feature size, and then use 1×1 convolution to process the multimodal features separately to adjust the channel dimension to align the multimodal information. In the global part, 3×3 deep convolution is used to guide the generation of query Q4 and key K4, and the full focus feature is used to guide the generation of value V4. Q4 is reshaped into K4 was reshaped as and Used to construct the similarity matrix A4∈R 384×384 . A learnable parameter γ is introduced before the softmax operation to control and The magnitude of matrix multiplication. Multiply the matrix A4 and V4 to obtain the fused global features. Such a scheme can not only enhance the expression ability of the focal part but also achieve full fusion of multimodal information. In the local part, the focal features and the fully focused features are spliced ​​along the channel dimension, and then the channel shuffle operation is performed to promote the fusion of multimodal channel information. The fusion result is extracted through 3×3×3 convolution to extract local features. Finally, the local features and global features are aggregated through element-by-element addition operation. The first step can be described as:

[0097]

[0098] Where A4 represents the fourth-layer similarity matrix, and They represent the global, local and total outputs of the fourth layer fusion, respectively. Softmax(·) represents the softmax operation, and mul(·) represents the matrix multiplication. represents the result after reshaping key K4, represents the result after reshaping the query Q4, γ represents the learnable parameter, conv 1×1 (·) represents 1×1 convolution, V4 represents the value generated by the full-focus feature guidance after the fourth layer enhancement, conv 3×3×3 (·) represents a 3×3×3 convolution, CS(·) represents a channel shuffle operation, cat(·) represents a concatenation operation, and X a (4) represents the fourth layer enhanced full focus feature map obtained by processing the fourth receptive field block, represents the focal slice feature map after the fourth layer enhancement in the slice interleaving enhancement module;

[0099] For the second step, due to the lack of effective interaction between high-level multi-scale features and the differences and noise in features of different scales and modalities, it is very difficult to effectively fuse and avoid excessive computation. Full Focus Feature X a (3) and the output of the fourth layer As input, the fusion module is rebuilt to fully fuse multimodal and multi-scale information while ensuring the same amount of computation as the fourth layer. After the three inputs are spliced ​​along the channel in the local part, the subsequent processing is the same as the first step. In the global part, the focal flow is used to guide the generation of query Q3, the full focal flow is used to guide the generation of key K3, and the fourth layer output is used to guide the generation of value V3. The purpose of this is to optimize the output obtained from the previous layer in a progressive manner with light field information at different levels to accurately locate salient objects and reduce noise interference. Similar to the first step, the query Q3 and key K3 are used to generate the similarity matrix A3 of the third layer, which is used to optimize the value V3 generated by the upsampled upper layer output to obtain global information. Finally, the global and local features are aggregated by element-by-element addition to obtain the total output.

[0100] The second step can be described as:

[0101]

[0102] In the formula, and Represent the global and local outputs of the third layer fusion, P Hrepresents the total output of the high-level feature fusion module, V3 represents the value generated by the full-focus feature guidance after the third layer enhancement, A3 represents the similarity matrix of the third layer, and X a (3) represents the fully focused feature map after the third layer enhancement, Represents the focal slice feature map after the third layer enhancement.

[0103] 2.3 Low-level Cross-Attention Module

[0104] Low-level features (i=1,2) have rich spatial details. Many previous methods believe that low-level features contribute too little to the SOD task and will introduce background interference and additional computation. However, only using high-level features for salient target detection may lose accurate perception of target edges and detailed structures, resulting in blurred target edges in salient maps or inability to accurately detect small targets. The present invention has found in the research that low-level features can well suppress non-salient areas and provide more fine-grained details. For this reason, a cross-attention module (CrossA) is designed by combining spatial-channel attention with low-level features, such as Figure 2 As shown in the lower right corner, multimodal features enhance fine-grained details and reduce background interference in a coarse-to-fine manner by learning mutual spatial-channel complementary information.

[0105] In this embodiment, the low-level cross attention module includes two sets of cascaded space-channel attention units and convolutional layers, and the two sets of space-channel attention units are connected in parallel, wherein the space-channel attention unit includes parallel orthogonal channel attention subunits and space attention subunits;

[0106] In the low-level cross-attention module, the low-level enhanced full focal flow features are copied along the batch dimension to align the focal flow, and the copied low-level full focal flow features and the low-level enhanced focal flow features are respectively used as the input of two groups of space-channel attention units, and the outputs are respectively output to generate their own orthogonal channel attention weights and spatial attention weights. The copied low-level full focal flow features and the low-level enhanced focal flow features are respectively multiplied with the orthogonal channel attention weights and spatial attention weights generated by each other, and the output results are spliced ​​and aggregated through the convolution layer to obtain the fused low-level feature map as the total output of the low-level cross-attention module.

[0107] Specifically, given the low-level focal flow features and the full focusing flow feature X a (i)∈R 1 ×C×H×W, due to the difference between the focal flow and the fully focused flow in the batch dimension, in order to facilitate calculation, the fully focused flow features are replicated 12 times along the batch dimension to align the focal flow. Spatial information plays an important role in high-resolution low-level features. This operation can also increase the importance of the fully focused features with rich spatial information to be consistent with the focal features. Before fusion, the number of channels of the two modal features is reduced by half through 1×1 convolution for ease of calculation. For the feature maps from the focal flow and the fully focused flow, their respective orthogonal channel attention weights are generated. and the spatial attention weight K SA (i)∈R 1×H×W , where C * =C / 2. Compared with traditional channel attention, orthogonal channel attention avoids the disadvantage of global average pooling that easily loses low-frequency information, and can extract richer representations of each feature map. The present invention proposes cross-modal cross-enhancement, where the focus feature and the full focus feature learn the dual attention weights generated by each other respectively, and aggregate them in the form of series connection and convolution. This approach can enhance the common salient target information in different modalities and reduce inconsistency.

[0108] Therefore, the whole process can be described as:

[0109]

[0110] P CA (i) = conv 3×3 (cat(T a (i),T fs (i)));

[0111] Where P CA (i) represents the output result of the low-level cross-attention module, T a (i) and T fs (i) represents the full focus feature and focus feature enhanced by dual attention, respectively. replicate(.) represents the replication operation along the batch dimension. represents the spatial attention weight, represents the orthogonal channel attention weight.

[0112] 2.4 Compact Pyramid Refinement Module

[0113] In order to ensure the efficiency of fusing multi-scale information, the Compact Pyramid Refinement (CPR) module is used as the decoder module, which uses 1×1 and depth-wise separable convolution for multi-scale learning and aggregates multi-scale features from top to bottom. Figure 2 As shown in Figure 2, a total of two CPR modules are used, and finally a 1×1 convolution is used to predict the final saliency map S. final.

[0114] Therefore, the processing of this module is:

[0115] S final =conv 1×1 (CPR(CPR(P H ,P CA (2)),P CA (1)));

[0116] In the formula, S final represents the final saliency map, and CPR(·) represents the CPR refinement operation.

[0117] 3. Light field salient object detection model training loss

[0118] In a specific embodiment, during the training of the light field salient object detection model, in order to fully learn the enhanced multi-layer focal slice flow information, a hybrid structure loss consisting of a binary cross entropy loss and an intersection-over-union loss is introduced during the training process. The four-layer enhanced features are predicted to obtain a rough salient map, which is upsampled to the same resolution as the true value map. Therefore, the slice interleaving enhancement loss function L E It can be expressed as:

[0119]

[0120] Where, L E represents the slice interleaving enhancement loss function, L S represents the hybrid structure loss function, S fs (i) represents a rough saliency map, and G represents the ground truth.

[0121] The slice interleaving enhancement loss function and the hybrid structure loss function are combined into a total loss function to calculate the training loss of the light field salient object detection model. The parameters of the light field salient object detection model are updated with the goal of minimizing the total loss function. The total loss function L total It can be expressed as:

[0122] L total =L E +L S (S final ,G);

[0123] Where, L total Represents the total loss function.

[0124] 4. Experimental description

[0125] In order to better illustrate the advantages of the technical solution of the present invention, this embodiment discloses the following experiment.

[0126] 4.1 Experimental Datasets and Evaluation Metrics

[0127] To verify the effectiveness of the algorithm, the proposed network is evaluated on three extensive public light field datasets: DUTLF-FS, HFUT, and LFSD. DUTLF-FS contains 1462 light field samples of indoor and outdoor scenes, HFUT and LFSD contain 255 and 100 samples of challenging scenes, such as occlusion, small targets and multiple targets, etc., respectively, where all samples are composed of full focus images, multiple focus stack images (the number is uniformly set to 12) and corresponding truth maps. To ensure fairness and repeatability, following the public training settings, the training set consists of 1000 samples in DUTLF-FS and 100 samples in HFUT, and all the remaining samples are used as test sets.

[0128] This embodiment uses four commonly used significance detection indicators as evaluation indicators: mean absolute error (MAE) is the average of the absolute values ​​of the difference between the predicted value and the true value; F-measure (F β ) is defined as an adaptive value to measure precision and recall; S-measure (S) evaluates the structural similarity between the predicted value and the true value; E-measure (E β ) takes into account the local and global similarity between the predicted value and the true value. β , S and E β The larger the value, the better the model performance, and the smaller the MAE, the better.

[0129] 4.2 Experimental Implementation Details

[0130] The network proposed in this embodiment is implemented based on the PyTorch framework. The encoder is initialized with the pre-trained Pvtv2. In order to ensure the consistency of the image, the size of the full focus image, the focus stack image and the true value image are adjusted to 256×256, and random rotation, cropping and flipping are used for data enhancement. The model in this paper is trained and tested on a single NVIDIA RTX 4090 GPU, and the AdamW optimizer with a learning rate of 5e-5 and a weight decay of 1e-4 is selected to optimize the network. The training epoch is 150 and the batch size is 2. In Swin-T, the window size is set to 4 and m is 4.

[0131] 4.3 Experimental comparison

[0132] For comprehensive comparison, the network of this embodiment is compared with 4 state-of-the-art RGB-D and 11 light field salient object detection methods, namely BBSNet, JL-DCF, VST, SwinNet, MoLF, ERNet, SANet, DLGLRG, PANet, MEANet, STSA, GFRNet, LFTransNet (LFTrans), LF-Tracy and SAGNet on three public light field datasets (DUTLF-FS, HFUT, LFSD). In order to ensure the fairness of the comparison, all results are obtained from official releases or tested by officially released prediction maps or models, and all evaluation criteria are consistent. According to the existing comparison scheme, the RGB-D method uses a fully focused image and a depth map as input, and the light field method uses a fully focused image and a focal stack as input.

[0133] Table 1 Quantitative comparison of different methods on three light field datasets, using mean absolute error (MAE), adaptive F-measure (F β ), S-measure(S) and adaptive E-measure(F β ) as an evaluation indicator

[0134]

[0135] In Table 1, ↑(↓) means higher (lower) is better, the best results are marked in bold, and the second best results are marked with underlines.

[0136] Table 1 shows the quantitative comparison results. The network of this embodiment surpasses 15 state-of-the-art SOD methods on three datasets, demonstrating its strong competitiveness and generalization ability. The results show that RGB-D methods are significantly inferior to the method of this paper, mainly due to the presence of a large number of unreliable depth maps in the dataset, which limits their ability to effectively distinguish foreground and background in complex scenes. Comparison with light field methods shows that the network of this embodiment achieves highly competitive results. Specifically, the proposed model performs well on the HFUT dataset, which contains a variety of challenging scenes such as complex backgrounds, small objects, and occlusions, making it one of the most stringent datasets in the LFSOD task. In addition, it also achieved the top two rankings on the other two datasets. For example, compared with the state-of-the-art method SAGNet, the method of this embodiment performs better in detection performance. Specifically, SAGNet uses multimodal spatial attention to guide the fusion module and the implicit detail recovery module to achieve the best results on the large dataset DUTLF-FS, but its limitations become apparent when applied to challenging datasets such as HFUT and LFSD. In contrast, the method proposed in this example surpasses SAGNet in all challenging scenarios and achieves comparable performance on the DUTLF-FS dataset. The main reason for the success is the effectiveness of the multi-modal and multi-scale fusion strategy, which effectively avoids the interference of complex scenarios and accurately locates salient objects.

[0137] exist Figure 6 Figure 2 shows the visual comparison of salient maps generated by the proposed network and eight state-of-the-art light field salient object detection methods in different challenging scenarios, including small objects with occlusions (1st row), multiple objects (2nd and 6th rows), insufficient light (3rd and 7th rows), transparent objects (4th row), and complex edges (5th and 8th rows). Figure 6 It can be clearly seen that the results of the network in this embodiment are closer to the ground truth (GT) in these scenarios, showing strong robustness. Figure 6 In the case of multiple foreground objects in the second row and branches interfering with the foreground, other methods will be interfered by the branches and fail to completely separate the salient objects. However, the network proposed in this embodiment can not only selectively remove the background area (branches), but also more accurately capture the edge contour details of the salient objects, thanks to the powerful scene segmentation and detail recovery capabilities of the proposed model.

[0138] 4.4 Ablation Experiment

[0139] In order to verify the effectiveness of the fusion strategy, the present invention conducts an ablation experiment on the DUTLF-FS dataset, as shown in Table 2, and uses the mean absolute error (MAE), F-measure (F β), S-measure(S), and E-measure(E β ) as the evaluation index. The unmarked part indicates that element-by-element multiplication is used for replacement. No.1 is used as the baseline, and No.2 and No.5 represent the results of using two fusion modules on the corresponding layers. The comparison results show that the fusion modules (HFFM and CrossA) proposed in this paper significantly improve the detection performance, verifying their effectiveness.

[0140] Table 2 Ablation experiments of HFFM and CrossA fusion modules

[0141]

[0142] In Table 2, H indicates the use of HFFM module for fusion, C indicates the use of CrossA module, S1-S4 correspond to four layers, and the best results are marked in bold.

[0143] As can be seen from Table 2, No.2 to No.4 indicate that CrossA is gradually applied from the low layer (1st layer and 2nd layer) to all layers. No.5 to No.7 indicate that HFFM is gradually applied from the high layer (3rd layer and 4th layer) to all layers. No.8 indicates that HFFM is applied to the low layer and CrossA is applied to the high layer. No.9 is the fusion strategy proposed in the present invention, which applies CrossA to the low layer and HFFM to the high layer. The experimental results show that applying CrossA to the high layer and applying HFFM to the low layer will lead to performance degradation, while the strategy proposed in the present invention achieves the best results. This shows that when the spatial information of the low layer is introduced, the ability of HFFM to distinguish salient target information will be weakened, and CrossA is limited in its ability to model low-resolution deep semantic information. This further verifies that HFFM is more suitable for high-level feature fusion, while CrossA is more suitable for low-level feature fusion. The strategy proposed in the present invention can emphasize the integrity of salient targets by integrating multimodal semantic information, while enhancing the contour details of salient targets by using shared spatial and channel information, thereby generating high-quality saliency maps.

[0144] 5. Overview

[0145] In summary, the method of the present invention has the following technical advantages:

[0146] (1) The existing light field salient target detection method uses rigid ConvLSTM to process focal slice features, ignoring the unique structural information between focal slices. The present invention constructs a slice interleaving enhancement module based on Swin Transformer, which makes full use of the geometric relationship between focal slices to segment the foreground and background. The module uses the Swin-T module to learn the interleaving information of different focal slices in a local manner, and uses multi-head self-attention to calculate the local similarity of the local window block, so as to learn the focal interleaving information within the local window block of the whole group of slices, and uses a cyclic sliding window to model the global information. At the same time, the present invention follows the Swin Transformer to use masked multi-head self-attention calculation and then restore, so as to avoid the interference of information in different regions and the destruction of semantic information.

[0147] (2) The existing light field salient target detection methods do not adequately integrate multimodal information features, ignoring the unique structural characteristics of the two modal information. The present invention proposes a multimodal feature fusion strategy consisting of a high-level feature fusion module, a low-level cross-attention module, and a compact pyramid refinement module. Through the high-level feature fusion module and the low-level cross-attention module, the model can make more full use of the multimodal information in the light field image, and perform feature aggregation and decoding through the compact pyramid refinement module, thereby optimizing the computational efficiency. This strategy fully considers the complementary relationship between multimodal and multi-scale features, thereby enabling the network to pay more attention to the salient target information.

[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described with reference to the preferred embodiments of the present invention, it should be understood by those skilled in the art that various changes may be made in form and details without departing from the spirit and scope of the present invention as defined in the appended claims.

Claims

1. A light field salient object detection method based on multimodal feature fusion, characterized in that: The steps include: S1, obtaining a light field multimodal focal stack image and a fully focused image to be processed; S2, using the light field multimodal focal stack image and the fully gathered image to be processed as inputs of the trained light field salient target detection model, and outputting light field salient target detection results of the light field multimodal focal stack image and the fully gathered image to be processed; the light field salient target detection model includes a slice interleaving enhancement module, a high-level feature fusion module, a low-level cross attention module, and a compact pyramid refinement module; The training process of the light field salient object detection model includes: S201, taking the light field multimodal focus stack image and the fully focused image as inputs of the light field salient target detection model as training samples, inputting them into the light field salient target detection model with Pvtv2 as the backbone network, performing feature extraction on the training samples, and respectively extracting four focal stack feature maps and fully focused feature maps of different scales; S202, constructing the slice interleaving enhancement module based on the Swin-T module, taking the focus stack feature maps of the four different scales as the input of the slice interleaving enhancement module, enhancing the features between the focus slices in a local manner, and obtaining enhanced focus flow features; S203, the four fully focused feature maps of different scales obtained by Pvtv2 are passed through the receptive field block to obtain an enhanced fully focused flow feature; the high-level features in the enhanced fully focused flow feature and the enhanced focus flow feature are used as inputs of the high-level feature fusion module, and the high-level multimodal semantic information is locally and globally feature-fused to obtain a fused high-level feature map as output; S204, using the enhanced full focus flow features and the low-level features in the enhanced focus flow features as inputs of the low-level cross attention module, fusing the low-level multimodal spatial-channel features in a cross-enhancement manner, and outputting a fused low-level feature map; S205, using the fused high-level feature map and the fused low-level feature map as inputs of the compact pyramid refinement module, aggregating multi-scale features from top to bottom, and outputting a final saliency map; S206, using a total loss function composed of a slice interleaving enhancement loss function and a hybrid structure loss function to calculate the training loss of the light field salient object detection model, and updating the parameters of the light field salient object detection model with the goal of minimizing the total loss function; S207 , repeating steps S201 to S206 until the light field salient object detection model converges or reaches a preset number of training times, and the training is completed.

2. The light field salient object detection method based on multimodal feature fusion according to claim 1, characterized in that: In step S202, the slice interleaving enhancement module includes a cascaded Swin-T module, four parallel receptive field blocks and three upsampling units, wherein the Swin-T module includes a local window self-attention unit and a sliding window mask unit; In the slice interleaving enhancement module, the focal stack feature map of each scale of the input is passed through four parallel receptive field blocks to obtain the corresponding focal slice feature map; then, the local window self-attention unit in the Swin-T module is used to divide the fourth-layer focal slice feature map into multiple local window blocks, and the multi-head self-attention mechanism is used to calculate the local similarity. Then, the sliding window mask unit in the Swin-T module is used to slide on the local feature map, and the global feature map is obtained through the mask operation; finally, the fourth-layer focal slice feature map is fused with the features obtained after passing through the Swin-T module through element-by-element multiplication operation, and the fused feature map is used as the input of the third upsampling unit. After the output of each upsampling unit is fused with the corresponding focal slice feature map, it is used as the input to the previous layer, and the focal slice feature maps of all levels are enhanced in turn to obtain the corresponding enhanced focal slice feature map.

3. The light field salient object detection method based on multimodal feature fusion according to claim 2 is characterized in that: The processing process of the slice interleaving enhancement module is as follows: In the formula, RFB(·) represents the output after the receptive field block operation, represents the focal stack feature map after Pvtv2, represents the focal slice feature map output by the receptive field block, n represents the number of focal slices (n = 1, 2, ..., 12), i represents the number of layers of multi-scale features (i = 1, 2, 3, 4), represents the enhanced focal flow feature, Swin(·) represents the result after Swin-T module, U(·) represents bilinear upsampling, Represents element-wise multiplication.

4. The light field salient object detection method based on multimodal feature fusion according to claim 1, characterized in that: In step S203, the processing process of the high-level feature fusion module includes: In the first step of the high-level feature fusion module, the fourth-layer enhanced focal slice feature map in the slice interleaving enhancement module and the fourth-layer enhanced full-focus feature map obtained by processing the fourth receptive field block are feature fused to obtain the fourth-layer fusion total output; then, in the second step, the fourth-layer fusion global and local outputs and the third-layer enhanced focal slice feature map in the slice interleaving enhancement module and the third-layer enhanced full-focus feature map obtained by processing the third receptive field block are used as input to reconstruct the high-level feature fusion module, and the third-layer fusion total output is obtained as the total output of the high-level feature fusion module.

5. The light field salient object detection method based on multimodal feature fusion according to claim 4 is characterized in that: The processing process of the high-level feature fusion module is as follows: The first step of the process is expressed as: Where A4 represents the fourth-layer similarity matrix, and They represent the global, local and total outputs of the fourth layer fusion, respectively. Softmax(·) represents the softmax operation, and mul(·) represents the matrix multiplication. represents the result after reshaping key K4, represents the result after reshaping the query Q4, γ represents the learnable parameter, conv 1×1 (·) represents 1×1 convolution, V4 represents the value generated by the full-focus feature guidance after the fourth layer enhancement, conv 3×3×3 (·) represents a 3×3×3 convolution, CS(·) represents a channel shuffle operation, cat(·) represents a concatenation operation, and Xa(4) represents the fourth-layer enhanced all-focus feature map obtained by processing the fourth receptive field block. represents the focal slice feature map after the fourth layer enhancement in the slice interleaving enhancement module; The second step process is expressed as: In the formula, and Represent the global and local outputs of the third layer fusion, P H represents the total output of the high-level feature fusion module, V3 represents the value generated by the full-focus feature guidance after the third layer enhancement, A3 represents the similarity matrix of the third layer, and X a (3) represents the fully focused feature map after the third layer enhancement, Represents the focal slice feature map after the third layer enhancement.

6. The light field salient object detection method based on multimodal feature fusion according to claim 1, characterized in that: In step S204, the low-level cross attention module includes two groups of cascaded space-channel attention units and convolutional layers, and the two groups of space-channel attention units are connected in parallel, wherein the space-channel attention unit includes parallel orthogonal channel attention sub-units and space attention sub-units; In the low-level cross-attention module, the low-level enhanced full focal flow features are copied along the batch dimension to align the focal flow, and the copied low-level full focal flow features and the low-level enhanced focal flow features are respectively used as the input of two groups of space-channel attention units, and the outputs are respectively output to generate their own orthogonal channel attention weights and spatial attention weights. The copied low-level full focal flow features and the low-level enhanced focal flow features are respectively multiplied with the orthogonal channel attention weights and spatial attention weights generated by each other, and the output results are spliced ​​and aggregated through the convolution layer to obtain the fused low-level feature map as the total output of the low-level cross-attention module.

7. The light field salient object detection method based on multimodal feature fusion according to claim 6, characterized in that: The processing process of the low-level cross-attention module is: P CA (i)=conv 3×3 (cat(T a (i),T fs (i))); Where P CA (i) represents the output result of the low-level cross-attention module, T a (i) and T fs (i) represents the full focus feature and focus feature enhanced by dual attention, respectively. replicate(.) represents the replication operation along the batch dimension. represents the spatial attention weight, represents the orthogonal channel attention weight.

8. The light field salient object detection method based on multimodal feature fusion according to claim 1, characterized in that: In step S205, the processing process of the compact pyramid refinement module is as follows: S final =conv 1×1 (CPR(CPR(P H ,P CA (2)),P CA (1))); In the formula, S final represents the final saliency map, and CPR(·) represents the CPR refinement operation.

9. The light field salient object detection method based on multimodal feature fusion according to claim 1, characterized in that: In step S206, the slice interleaving enhancement loss function is: Where, L E represents the slice interleaving enhancement loss function, L S represents the hybrid structure loss function, S fs (i) represents a rough saliency map, and G represents the ground truth.

10. The light field salient object detection method based on multimodal feature fusion according to claim 9, characterized in that: The total loss function is: L total =L E +L S (S final ,G); Where, L total Represents the total loss function.

Citation Information

Patent Citations

  • Light field image salient target detection method based on learnable weight descriptor

    CN115546512A

  • Cross-modal feature fusion and asymptotic decoding saliency target detection method and device

    CN115908789A

  • Light field image saliency detection method and system based on multi-stage fusion

    CN117523347A

  • Efficient RGB-D saliency detection method using multiple information complementation

    CN118982655A

Cited By

  • Cross-modal common-attention light field salient target detection method

    CN120526169A

  • A cross-modal co-attention light field salient object detection method

    CN120526169B