Efficient and accurate integrated imaging light field three-dimensional salient target detection method

By using multi-layer visual state space coding and saliency-guided Mamba modules, combined with lens array reconstruction, the contradiction between accuracy and efficiency in three-dimensional salient target detection in light fields is resolved, achieving efficient and accurate salient target detection.

CN121305545APending Publication Date: 2026-01-09XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511531087.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing technologies struggle to improve computational efficiency while maintaining detection accuracy in 3D salient target detection using light fields, especially when dealing with complex lighting and background scenes.

Method used

A multi-layer visual state space coding module is used to extract features from four-dimensional light field data. Combined with a saliency-guided Mamba module and a selective scanning mechanism, global and local features are fused through a multi-directional scanning strategy to achieve fine reconstruction and spatial consistency of salient regions. Finally, optical reconstruction is performed by combining a lens array.

Benefits of technology

It achieves a balance between accuracy and efficiency in significant light field detection, improving the accuracy of 3D visual perception and the reliability of intelligent imaging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121305545A_ABST
    Figure CN121305545A_ABST
Patent Text Reader

Abstract

The invention provides an efficient and accurate integrated imaging light field three-dimensional salient target detection method. According to the method, feature extraction and fusion are carried out on four-dimensional light field data through a visual state space coding module, and three-dimensional saliency detection and optical reconstruction are achieved. Firstly, light field data of a three-dimensional scene are obtained through a lens array and recorded in a two-dimensional micro-image array. Then, the micro-image array is divided into sub-micro-image sequences, and the sub-micro-image sequences are input into a visual space state module for multi-scale feature coding and modeling; a salient region is modeled by a saliency-guided Mama module under a selective scanning mechanism, and the robustness and space consistency of features are enhanced through multi-path scanning. And after decoding by the hybrid attention module, the recovery and detail enhancement of the light field salient features are completed, and a high-quality salient micro-image array is generated. And finally, optical reconstruction is completed in combination with a lens array, a three-dimensional salient target with real depth features is obtained, and the balance between the detection precision and the calculation efficiency is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of salient object detection of computer vision tasks, more particularly, it is particularly related to an efficient and accurate integrated imaging light field three-dimensional salient object detection method. BACKGROUND

[0002] As an important direction of light field imaging, integrated imaging can obtain four-dimensional light field data containing spatial position information and depth clues at the same time by the regulation of microlens array, and realize three-dimensional information coding under non-coherent light conditions in the form of two-dimensional micro image array. This light field representation method based on micro image array provides a very potential implementation approach for naked eye three-dimensional display, and more importantly, its multi-view correlation characteristics construct a perfect geometric constraint system. The parallax information contained in each micro image can be mapped to the depth information of the spatial scene, and the continuous distribution of the target in space provides topological prior support for three-dimensional reconstruction. On this basis, saliency detection of micro image array is of great significance, which not only can improve the stability of depth estimation through salient region positioning, but also can provide accurate semantic segmentation boundary for target-oriented three-dimensional reconstruction, thereby enhancing the accuracy and reliability of three-dimensional scene analysis.

[0003] In the field of salient object detection, traditional CNN-based methods have structural defects such as limited local receptive field and insufficient feature modeling capability, and Transformer-based methods have low running efficiency due to the high computational complexity of self-attention mechanism, so it is difficult to realize efficient and accurate salient region extraction, and finally it performs poorly in spatial consistency and boundary detail recovery. Although some studies try to improve the architecture to preliminarily alleviate the contradiction between efficiency and accuracy of traditional methods, they mainly focus on two-dimensional images. However, light field data is prone to depth perception loss and edge detail blur during high-dimensional modeling, which makes existing light field three-dimensional salient object detection methods difficult to ensure detection accuracy while considering computational efficiency when dealing with complex lighting and background scenes. SUMMARY

[0004] This invention proposes a highly efficient and accurate integrated imaging light field method for 3D salient target detection. This method achieves a balance between accuracy and efficiency in saliency detection and optical reconstruction by rapidly extracting and integrating four-dimensional light field data. The acquired micro-image array is divided into sub-micro-image sequences according to the size of the micro-unit images, and input into an encoder composed of multi-layer visual spatial state modules for multi-scale feature extraction and fusion, generating a preliminary predicted salient micro-image array. A saliency-guided Mamba module dynamically adjusts the priority of the selective scanning mechanism, enhancing the feature response to salient regions. The optimized feature map is input into a decoder for feature reconstruction and integration, outputting a high-precision salient micro-image array. Finally, combined with a lens array, optical reconstruction is completed, generating a 3D salient target with true spatial depth, achieving accurate reconstruction and visualization of light field information in complex scenes.

[0005] The method comprises five processes: four-dimensional light field capture, light field data encoding, saliency-guided Mamba, saliency light field feature decoding, and saliency light field three-dimensional reconstruction. The specific flowchart is attached. Figure 1 As shown.

[0006] The four-dimensional light field capture process is achieved through a lens array, used to record object information of complex three-dimensional scenes in the form of a two-dimensional micro-image array. This micro-image array contains multi-view illumination information and scene geometric features, and can characterize the spatial and angular distribution characteristics of the light field. Further, the generated micro-image array is sequentially segmented according to the micro-unit image size and input into the encoder of the visual spatial state module to perform feature extraction and light field distribution modeling processing.

[0007] The optical field data encoding process is as follows: Figure 2 As shown. The visual state space module quickly incorporates input features. Normalization is performed, followed by two branches. The main branch sequentially passes through a linear layer, dimension reconstruction, depthwise separable convolution, SiLU activation function, selective scanning module, and normalization. Further, in the selective scanning module, the light field feature sequence after dimension reconstruction... The scan is unfolded into a one-dimensional sequence along four directions. Each sequence is then processed by the S6 module to capture long-range dependencies. Finally, the sequence information is reordered and aggregated along the original directions to extract high-dimensional light field features. The bypass branch contains only a linear layer and a SiLU activation function. The features from both branches are element-wise multiplied, passed through a linear layer, and then output. Finally, they are added to the original input via a residual connection to obtain the extracted light field features. Finally, the extracted light field features are input into the next layer of the visual state space module.

[0008] The saliency-guided Mamba process first divides the initially predicted saliency microimage array into sub-saliency microimage sequences. Subsequently, it was normalized and reconstructed using a linear layer. Then, each sub-microimage in the salient sub-microimage sequence is assigned an index representing its position in the salient microimage array. The saliency-guided selective scanning module scans all salient regions row by row, starting with the first salient sub-microimage array. The indices of all non-salient sub-microimage arrays are scanned sequentially, generating the complete scan path for the salient microimage array. The resulting features are then aggregated by the S6 module, and subsequently passed through a normalization layer and a linearization layer to output high-quality light field features. .

[0009] The decoding process for the significant optical field features is as follows: Figure 3 As shown, this process is based on the construction of a visual spatial state module. A hybrid attention module is introduced after the selective scanning module to form a visual spatial state decoder layer. Furthermore, the high-quality light field features obtained after processing by the saliency-guided selective scanning module are... Compared with the multi-scale coarse-grained light field features extracted by the encoder After aggregation, the features are used as input to the decoder. This structure enables hierarchical decoding and feature aggregation of salient light field features, ultimately restoring the decoded salient light field features into a salient micro-image array, thereby achieving high-efficiency and high-precision detection of salient light field targets.

[0010] The significant light field three-dimensional reconstruction process is shown in the attached figure. Figure 4 As shown, by using a lens array to reconstruct the angular and spatial information of the light field, a four-dimensional light field spatial reconstruction is achieved, thereby generating a three-dimensional salient object with depth features.

[0011] This invention proposes a highly efficient and accurate integrated imaging light field salient target detection method for improving the accuracy and computational efficiency of light field salient detection. The method employs a multi-layer visual state space coding module to extract features from four-dimensional light field data, maintaining spatial continuity through a selective scanning mechanism to achieve efficient modeling of salient features. A salientity-guided decoding module utilizes a multi-directional scanning strategy to fuse global and local features, achieving fine reconstruction and spatial consistency of salient regions. Finally, optical reconstruction of the salient micro-image array is performed using a lens array to generate three-dimensional salient targets with depth features, thus achieving a balance between accuracy and efficiency in light field salient detection and providing technical support for high-precision three-dimensional visual perception and intelligent imaging applications. Attached Figure Description

[0012] Appendix Figure 1 This is a flowchart of a highly efficient and accurate integrated imaging light field three-dimensional salient target detection method proposed in this invention.

[0013] Appendix Figure 2This is a schematic diagram of the encoder architecture based on the vision-space state module.

[0014] Appendix Figure 3 This is a schematic diagram of salient feature decoding.

[0015] Appendix Figure 4 This is a schematic diagram of a significant light field 3D reconstruction.

[0016] It should be understood that the above figures are only schematic and are not drawn to scale. Detailed Implementation

[0017] The following detailed description of a typical embodiment of the efficient and accurate integrated imaging light field three-dimensional salient target detection method proposed in this invention further illustrates the invention. It is necessary to point out that the following embodiments are only used for further illustrative purposes and should not be construed as limiting the scope of protection of this invention. Any non-essential improvements and adjustments made to this invention by those skilled in the art based on the above description still fall within the scope of protection of this invention.

[0018] This invention proposes an efficient and accurate integrated imaging light field three-dimensional salient target detection method, which specifically includes five processes: four-dimensional light field acquisition, light field data encoding, salientity-guided Mamba, salient feature decoding, and salient light field three-dimensional reconstruction. The specific process is shown in the attached figure. Figure 1 As shown.

[0019] The four-dimensional light field capture process is achieved through a lens array, used to record object information of complex three-dimensional scenes in the form of a two-dimensional micro-image array. This micro-image array contains multi-view illumination information and scene geometric features, and can characterize the spatial and angular distribution characteristics of the light field. Further, the generated micro-image array is sequentially segmented according to the micro-unit image size and input into the encoder of the visual spatial state module to perform feature extraction and light field distribution modeling processing.

[0020] The optical field data encoding process is as follows: Figure 2 As shown, a 2000×2000 micro-image array is divided into 8×8 sub-micro-image sequences, which are then input into the visual state space module. The sequences are normalized and then divided into two branches. The main branch sequentially passes through a linear layer, dimension reconstruction, depthwise separable convolution, SiLU activation function, selective scanning module, and normalization. Further, in the selective scanning module, the light field feature sequence after dimension reconstruction... The sequence is scanned in four directions and unfolded into a one-dimensional sequence of size 8×8. Each sequence is then processed by the S6 module to capture long-range dependencies. Finally, the sequence information is reordered and aggregated along the original directions to extract high-dimensional light field features. The bypass branch contains only a linear layer and a SiLU activation function. The features from both branches are element-wise multiplied, passed through a linear layer, and then output. Finally, they are added to the original input via a residual connection to obtain the extracted light field features. Finally, the extracted light field features are input into the next layer of the visual state space module.

[0021] The saliency-guided Mamba process first divides the initially predicted saliency microimage array of size 2000 × 2000 into sub-saliency microimage sequences of size 500 × 500. Subsequently, it was normalized and reconstructed using a linear layer. Then, each sub-microimage in the salient sub-microimage sequence is assigned an index representing its position in the salient microimage array. The saliency-guided selective scanning module scans all salient regions row by row, starting with the first salient sub-microimage array. The indices of all non-salient sub-microimage arrays are scanned sequentially, generating a complete scan path for the salient microimage array. Each path contains the scan sequence of 16 sub-microimages. The resulting features are then aggregated by the S6 module, followed by a normalization layer and a linearization layer to output high-quality light field features. .

[0022] The salient feature decoding process is as follows: Figure 3 As shown, this process is based on the construction of a visual spatial state module. A hybrid attention module is introduced after the selective scanning module to form a visual spatial state decoder layer. Furthermore, the high-quality light field features obtained after processing by the saliency-guided selective scanning module are... Compared with the multi-scale coarse-grained light field features extracted by the encoder After aggregation, the features are used as input to the decoder. This structure enables hierarchical decoding and feature aggregation of salient light field features, ultimately restoring the decoded salient light field features into a salient micro-image array, thereby achieving high-efficiency and high-precision detection of salient light field targets.

[0023] The significant light field three-dimensional reconstruction process is shown in the attached figure. Figure 4 As shown, by using a lens array to reconstruct the angular and spatial information of the light field, a four-dimensional light field spatial reconstruction is achieved, thereby generating a three-dimensional salient object with depth features.

[0024] This invention proposes a highly efficient and accurate integrated imaging light field salient target detection method for improving the accuracy and computational efficiency of light field salient detection. The method employs a multi-layer visual state space coding module to extract features from four-dimensional light field data, maintaining spatial continuity through a selective scanning mechanism to achieve efficient modeling of salient features. A salientity-guided decoding module utilizes a multi-directional scanning strategy to fuse global and local features, achieving fine reconstruction and spatial consistency of salient regions. Finally, optical reconstruction of the salient micro-image array is performed using a lens array to generate three-dimensional salient targets with depth features, thus achieving a balance between accuracy and efficiency in light field salient detection and providing technical support for high-precision three-dimensional visual perception and intelligent imaging applications.

Claims

1. A highly efficient and accurate integrated imaging light field three-dimensional salient target detection method, the core of which is that the method includes five processes: four-dimensional light field acquisition, light field data encoding, saliency-guided Mamba, salient light field feature decoding, and salient light field three-dimensional reconstruction. By rapidly extracting and integrating four-dimensional light field data, a balance between accuracy and efficiency in saliency detection and optical reconstruction is achieved. The acquired micro-image array is divided into sub-micro-image sequences according to the micro-unit image size, and input into an encoder composed of multi-layer visual spatial state modules for multi-scale feature extraction and fusion to generate a preliminary predicted salient micro-image array. The saliency-guided Mamba module dynamically adjusts the priority of the selective scanning mechanism to enhance the feature response to salient regions. The optimized feature map is input into the decoder for feature reconstruction and integration, outputting a high-precision salient micro-image array. Finally, combined with a lens array, optical reconstruction is completed to generate a three-dimensional salient target with real spatial depth, realizing accurate reconstruction and visualization of light field information in complex scenes.

2. The efficient and accurate integrated imaging light field three-dimensional salient target detection method according to claim 1, characterized in that, The four-dimensional light field capture process is achieved through a lens array, which is used to record the object information of complex three-dimensional scenes in the form of a two-dimensional micro-image array. The micro-image array contains multi-view illumination information and scene geometric features, which can characterize the spatial and angular distribution characteristics of the light field. Further, the generated micro-image array is sequentially segmented according to the micro-unit image size and input to the encoder of the visual spatial state module to perform feature extraction and light field distribution modeling processing.

3. The efficient and accurate integrated imaging light field three-dimensional salient target detection method according to claim 1, characterized in that, The light field data encoding process encodes the input features through a visual state space module. Normalization is performed, followed by two branches; the main branch sequentially passes through a linear layer, dimension reconstruction, depthwise separable convolution, SiLU activation function, selective scanning module, and normalization; further, in the selective scanning module, the light field feature sequence after dimension reconstruction is processed. The scan is unfolded into a one-dimensional sequence along four directions. Each sequence is then processed by the S6 module to capture long-range dependencies. Finally, the sequence information is reordered and aggregated along the original directions to extract high-dimensional light field features. The bypass branch contains only a linear layer and a SiLU activation function; the features from the two branches are element-wise multiplied, passed through a linear layer, and then output. Finally, they are added to the original input via a residual connection to obtain the extracted light field features. Finally, the extracted light field features are input into the next layer of the visual state space module.

4. The efficient and accurate integrated imaging light field three-dimensional salient target detection method according to claim 1, characterized in that, The saliency-guided Mamba process first divides the initially predicted saliency microimage array into sub-saliency microimage sequences. Subsequently, it was normalized and reconstructed using a linear layer. Then, each sub-microimage in the salient sub-microimage sequence is assigned an index, representing its position in the salient microimage array; the saliency-guided selective scanning module scans all salient regions row by row, starting from the first salient sub-microimage array; the indices of all non-salient sub-microimage arrays are scanned sequentially to generate the complete scan path of the salient microimage array; after passing through the S6 module, the scans are aggregated, and then passed through a normalization layer and a linear layer to output high-quality light field features. .

5. The efficient and accurate integrated imaging light field three-dimensional salient target detection method according to claim 1, characterized in that, The salient light field feature decoding process is based on a visual spatial state module. A hybrid attention module is introduced after the selective scanning module to form a visual spatial state decoder layer. Furthermore, the high-quality light field features obtained after processing by the saliency-guided selective scanning module are... Compared with the multi-scale coarse-grained light field features extracted by the encoder After aggregation, it serves as the input to the decoder; through this structure, hierarchical decoding and feature aggregation of salient light field features are realized, and finally the decoded salient light field features are restored into a salient micro-image array, thereby achieving high-efficiency and high-precision detection of salient light field targets.

6. The efficient and accurate integrated imaging light field three-dimensional salient target detection method according to claim 1, characterized in that, The salient light field 3D reconstruction process uses a lens array to reconstruct the angle and spatial position information of the light field, thereby realizing the spatial restoration of the four-dimensional light field and generating a salient 3D object with depth features.