A multi-dimensional feature extraction and fusion light field semantic segmentation method based on SS-SSM

By using a multi-dimensional feature extraction and fusion method based on SS-SSM, the problems of inefficient high-dimensional data processing, insufficient feature interaction, and high computational cost in light field semantic segmentation are solved, achieving efficient light field semantic segmentation and improving segmentation accuracy and model stability.

CN120747523BActive Publication Date: 2025-11-28HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511258733.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-28
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing light field semantic segmentation methods are inefficient in high-dimensional data processing, have insufficient feature interaction, high computational overhead, and poor adaptability to dynamic scenes, resulting in decreased segmentation accuracy.

Method used

A multi-dimensional feature extraction and fusion method based on SS-SSM is adopted. By preprocessing, extracting, fusing and stitching four-dimensional light field images, the efficient fusion of light field space-angle and bidirectional structural information is achieved by using the SSM-attention cross-fusion module.

Benefits of technology

It improves the efficiency and accuracy of light field semantic segmentation, reduces occlusion and misclassification problems, enhances the generalization and stability of the model, and reduces computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747523B_ABST
    Figure CN120747523B_ABST
Patent Text Reader

Abstract

The application provides a multi-dimensional feature extraction and fusion light field semantic segmentation method based on SS-SSM, comprising: acquiring a four-dimensional light field image; performing a preprocessing operation on the four-dimensional light field image to obtain preprocessed four-dimensional light field data information; performing a feature extraction operation on the four-dimensional light field data information to obtain feature matrix information; performing a fusion and splicing operation on the feature matrix information to obtain fused and spliced feature information; and determining the category prediction of each pixel based on the fused and spliced feature information. The method provided in the application uses SS-SSM to efficiently extract the space-angle and bidirectional structure information of the light field, enhances the target region recognition capability, and reduces the problems of occlusion and misclassification.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computer vision and light field image segmentation, and particularly relates to a light field multi-dimensional feature extraction and fusion semantic segmentation method based on SS-SSM. BACKGROUND

[0002] Light field (LF) imaging technology records the spatial and angular information of light rays, providing rich multi-dimensional data for scene understanding, and has wide application in automatic driving, robot navigation, computational photography, etc. Among them, light field semantic segmentation technology is the key to realize accurate target recognition, depth estimation and complex scene analysis. The mainstream model of current light field research is the 4D double-plane model of light field first proposed by Levoy et al. in 1996 . As shown in Figure 1 , (x, y) and (u, v) represent the intersection of light rays on the spatial plane and the angular plane respectively. Therefore, how to analyze the information in the 4D light field is the key to the light field semantic segmentation task.

[0003] The current light field semantic segmentation method still faces the following technical problems: (1) Low efficiency of high-dimensional data processing: directly expanding the 4D light field into a 1D sequence will cause dimension explosion, for example, the sequence length of a 9x9x512x512 light field after expansion is more than 2 million, which is difficult to process in real time. (2) Insufficient feature interaction: existing methods mostly use a single scanning path, which cannot fully fuse the spatial, angular and structural information in the light field. (3) High computational overhead: the feature fusion method based on attention mechanism (such as CrossViT) has a computational complexity of , which is difficult to extend to high-resolution light fields. (4) Poor adaptability to dynamic scenes: existing methods have low robustness in complex occlusion, light change and other scenes, resulting in a decrease in semantic segmentation accuracy. SUMMARY

[0004] In order to overcome the above technical defects, the present application provides a multi-dimensional feature extraction and fusion light field semantic segmentation method based on SS-SSM, in order to achieve the above purpose, the present application is realized according to the following technical scheme:

[0005] The present application provides a multi-dimensional feature extraction and fusion light field semantic segmentation method based on SS-SSM, comprising:

[0006] obtaining a four-dimensional light field image;

[0007] performing a preprocessing operation on the four-dimensional light field image to obtain preprocessed four-dimensional light field data information;

[0008] Feature extraction is performed on the four-dimensional light field data information to obtain feature matrix information;

[0009] The feature matrix information is fused and spliced to obtain fused and spliced feature information.

[0010] Based on the fused and spliced feature information, the category prediction of each pixel is determined.

[0011] Optionally, the four-dimensional light field image is obtained in one of the following ways:

[0012] Light field camera system or data set.

[0013] Optionally, the four-dimensional light field image is preprocessed to obtain preprocessed four-dimensional light field data information, which includes:

[0014] The four-dimensional light field image is standardized to obtain standardized four-dimensional light field data information, and the standardized four-dimensional light field data information is data information with pixel values in a first interval or a second interval.

[0015] The standardized four-dimensional light field data information is convoluted to obtain shallow initialization feature information.

[0016] Based on the shallow initialization feature information, the preprocessed four-dimensional light field data information is obtained.

[0017] Optionally, the preprocessed four-dimensional light field data information includes sub-aperture images, macro-pixel data, horizontal polar line images, and vertical polar line images. Based on the shallow initialization feature information, the preprocessed four-dimensional light field data information is obtained, which includes:

[0018] The shallow initialization feature information is divided to obtain divided first, second, third, and fourth shallow initialization feature information.

[0019] The first shallow initialization feature information is reshaped to fix the angle dimension to obtain sub-aperture images.

[0020] The second shallow initialization feature information is reshaped to fix the spatial dimension to obtain macro-pixel data.

[0021] The third shallow initialization feature information is reshaped to fix the height dimension and vertical viewing angle to obtain horizontal polar line images.

[0022] The fourth shallow initialization feature information is reshaped to fix the width dimension and horizontal viewing angle to obtain vertical polar line images.

[0023] Optionally, the feature matrix information includes remote space information, remote angle information, horizontal epipolar image structure information, and vertical epipolar image structure information, and the feature extraction operation on the four-dimensional light field data information to obtain the feature matrix information includes:

[0024] The sub-aperture image, the macro-pixel data, the horizontal epipolar image, and the vertical epipolar image are respectively input into four cascaded light field selective structure state space model modules for feature extraction, and sub-aperture depth feature information, macro-pixel depth feature information, horizontal epipolar depth feature information, and vertical epipolar depth feature information are correspondingly output;

[0025] The sub-aperture depth feature information, the macro-pixel depth feature information, the horizontal epipolar depth feature information, and the vertical epipolar depth feature information are respectively connected with the shallow layer initialization feature information to obtain remote space information, remote angle information, horizontal epipolar image structure information, and vertical epipolar image structure information.

[0026] Optionally, the fusion and splicing operation on the feature matrix information to obtain fused and spliced feature information includes:

[0027] The remote space information and the remote angle information are fused and spliced to obtain a remote angle-space splicing sequence.

[0028] The horizontal epipolar image structure information and the vertical epipolar image structure information are fused and spliced to obtain a horizontal-vertical structure splicing sequence.

[0029] The remote angle-space splicing sequence and the horizontal-vertical structure splicing sequence are fused and spliced to obtain the fused and spliced feature information.

[0030] Optionally, the fusion and splicing of the remote space information and the remote angle information to obtain a remote angle-space splicing sequence includes:

[0031] Based on the remote space information, a remote space projection matrix and a remote space hidden state corresponding to the remote space information are determined.

[0032] Based on the remote angle information, a remote angle projection matrix and a remote angle hidden state corresponding to the remote angle information are determined.

[0033] The remote space projection matrix and the remote angle hidden state are cross-fused to obtain remote angle-space cross-fusion information.

[0034] fuse the remote angle projection matrix and the remote spatial hidden state, and output remote spatial-angle cross fusion information;

[0035] splice the remote angle-space cross fusion information and the remote spatial-angle cross fusion information, and output a remote angle-space spliced sequence.

[0036] Optionally, the fusing and splicing of the horizontal structure information of the epipolar image and the vertical structure information of the epipolar image, and outputting a horizontal-vertical structure spliced sequence, comprises:

[0037] determining a horizontal structure projection matrix and a horizontal structure hidden state corresponding to the horizontal structure information of the epipolar image based on the horizontal structure information of the epipolar image;

[0038] determining a vertical structure projection matrix and a vertical structure hidden state corresponding to the vertical structure information of the epipolar image based on the vertical structure information of the epipolar image;

[0039] cross-fusing the horizontal structure projection matrix and the vertical structure hidden state, and outputting vertical-horizontal structure cross fusion information;

[0040] cross-fusing the vertical structure projection matrix and the horizontal structure hidden state, and outputting horizontal-vertical structure cross fusion information;

[0041] splicing the vertical-horizontal structure cross fusion information and the horizontal-vertical structure cross fusion information, and outputting a horizontal-vertical structure spliced sequence.

[0042] Optionally, the fusing and splicing of the remote angle-space spliced sequence and the horizontal-vertical structure spliced sequence, and outputting a fused and spliced feature information, comprises:

[0043] performing reverse sequence scanning on the remote angle-space spliced sequence, and obtaining a remote angle-space reverse sequence spliced sequence;

[0044] performing reverse sequence scanning on the horizontal-vertical structure spliced sequence, and obtaining a horizontal-vertical structure reverse sequence spliced sequence;

[0045] fusing the remote angle-space spliced sequence and the remote angle-space reverse sequence spliced sequence, and outputting a remote angle-space fused feature sequence;

[0046] fusing the horizontal-vertical structure spliced sequence and the horizontal-vertical structure reverse sequence spliced sequence, and outputting a horizontal-vertical fused feature sequence;

[0047] The remote angle-space fusion feature sequence and the horizontal-vertical fusion feature sequence are spliced to output fusion spliced feature information.

[0048] Optionally, the class prediction of each pixel is determined based on the fusion spliced feature information, including:

[0049] The fusion spliced feature information is input into four cascaded channel perception visual state space model decoders, and the fusion spliced feature information is progressively upsampled to output a channel perception four-cascaded decoding upsampled feature tensor.

[0050] The channel perception four-cascaded decoding upsampled feature tensor is input into a classifier to perform a classification operation, and the class prediction of each pixel is output.

[0051] The present application has the following beneficial effects:

[0052] The method provided in the present application uses the SS-SSM to efficiently extract the spatial-angle and bidirectional structure information of the light field, enhances the target region recognition capability, and reduces the occlusion and misclassification problems. The linear time computer mechanism of the SS-SSM greatly reduces the computing overhead compared with the Transformer method, improves the segmentation efficiency, and is suitable for large-scale light field data processing. Through the light field multi-dimensional feature learning and the SSM-attention cross fusion module, fine information integration is realized, and the model generalization and stability are improved.

[0053] In addition to the purposes, features and advantages described above, the present application has other purposes, features and advantages. The present application will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0054] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and the illustrative embodiments of the present application and their description serve the purpose of explaining the present application. The accompanying drawings should not be construed as an inappropriate limitation on the present application. In the drawings:

[0055] Figure 1 is a schematic diagram of a four-dimensional function of a light field mentioned in the background of the present application;

[0056] Figure 2 is a flowchart of a multi-dimensional feature extraction and fusion light field semantic segmentation method based on SS-SSM provided by an embodiment of the present application;

[0057] Figure 3 is a whole network structure diagram of the method provided by an embodiment of the present application;

[0058] Figure 4 is a schematic diagram of four different 2D expansions of 4D LF provided by an embodiment of the present application;

[0059] Figure 5 is an architecture diagram of an SSM-attention cross-fusion enhancement module provided by an embodiment of the present application;

[0060] Figure 6 is a visual result schematic diagram of the SOTA method on the UrbanLF-Syn data set provided by an experimental simulation stage of an embodiment of the present application;

[0061] Figure 7 is a visual result schematic diagram of the SOTA method on the UrbanLF-UrbanLF-Real data set provided by an experimental simulation stage of an embodiment of the present application; DETAILED DESCRIPTION

[0062] The embodiments of the present application are described in detail below with reference to the accompanying drawings, but the present application can be implemented in various different ways limited and covered by the claims.

[0063] Therefore, in order to solve the above problems, as Figure 2 indicated, the present application proposes a multi-dimensional feature extraction and fusion light field semantic segmentation method based on SS-SSM, which combines Figure 3 to describe the above method in detail, including:

[0064] Step S201: acquiring a four-dimensional light field image;

[0065] As Figure 3 indicated, it is the overall network architecture diagram of the method proposed by the present application, according to Figure 3 , to describe the method of the present application in detail:

[0066] Unlike traditional images that only record light intensity, light field (LF) imaging technology can capture the complete propagation information of light in a scene, providing rich multi-dimensional data for scene understanding, and providing a richer data basis for computer vision, three-dimensional reconstruction, camera refocusing, etc. The four-dimensional light field image (4D LF, 4D Light Field) used in the present application is the most common one at present, which can record the spatial and angular information of light. The core of the present application is to map the four-dimensional physical properties of light field (position and direction of light) to the pixel arrangement, block or channel of two-dimensional image, so as to facilitate storage, transmission and algorithm processing. The double-plane model is used to represent 4D LF, wherein represents the angular dimension, u is the horizontal viewing angle, v is the vertical viewing angle, represents the spatial dimension, h is the height viewing angle, and w is the width viewing angle. As Figure 4 indicated, different 2D expansion of 4D LF can be represented as: the same view under Sub-Aperture Image (SAI), each SAI image contains spatial position information of the image; the same position under Macro-Pixel Image (MacPI), the MacPI image can retain the detail difference of the same position under different viewing angles; the same angle and the same height (width) under Epipolar Plane Image (EPI), which can reflect the change trajectory of a row (column) of pixel points from multiple viewing angles, and the EPI image contains the disparity and depth information of the image. The above 2D representation can be represented as:

[0067] (1)

[0068] The core function of a light field camera system (such as a microlens array camera or a multi-camera array) is to directly capture light field data, and its hardware architecture is designed to capture the “position-direction” information of light. The data set is the storage and sharing carrier of the light field image: after standardizing and annotating the dispersed light field data, it provides support for algorithm research and application development. Therefore, the four-dimensional light field image can be obtained from the light field camera system or the data set.

[0069] Step S202: performing a preprocessing operation on the four-dimensional light field image to obtain preprocessed four-dimensional light field data information;

[0070] The preprocessed four-dimensional light field data information in the present application includes sub-aperture images, macro-pixel data, horizontal epipolar images, and vertical epipolar images.

[0071] The four-dimensional light field image contains spatial position and direction information, and the numerical values of different dimensions may have a coupling relationship. Standardization can reduce the scale difference between dimensions and make the features (such as light direction and spatial structure) more prominent. Therefore, the four-dimensional light field image data is standardized to obtain standardized four-dimensional light field data information. When the four-dimensional light field data information is standardized, the pixel value is normalized to a first interval or a second interval . Therefore, the four-dimensional light field data information can be understood as data information with pixel values in the first interval or data information with pixel values in the second interval.

[0072] The four-dimensional light field data information is convolved to obtain shallow layer initialization feature information The shallow layer initialization feature information is then divided into four data branches, preparing for the subsequent data residual connection after depth extraction. The four data branches are the first shallow layer initialization feature information, the second shallow layer initialization feature information, the third shallow layer initialization feature information, and the fourth shallow layer initialization feature information after division.

[0073] In computer science and data processing, the Reshape operation refers to the process of rearranging the dimensions or shape of data without changing its essential content. It is a fundamental operation in data preprocessing, neural network model construction, and tensor operations, widely used in image processing, natural language processing, and other fields. Therefore, the Reshape operation is used in this application to operate on the first shallow layer initialization feature information, the second shallow layer initialization feature information, the third shallow layer initialization feature information, and the fourth shallow layer initialization feature information. Specifically:

[0074] The first shallow layer initialization feature information is fixed in angle dimension through the Reshape operation. , obtaining a sub-aperture image for subsequent spatial feature extraction. The second shallow layer initialization feature information is fixed in space dimension through the Reshape operation. , obtaining macro-pixel data for angle information modeling. The third shallow layer initialization feature information is fixed in height dimension and vertical viewing angle , obtaining a horizontal epipolar image (EPI-H). The fourth shallow layer initialization feature information is fixed in width dimension and horizontal viewing angle , obtaining a vertical epipolar image (EPI-V) to extract bidirectional structural information of light field.

[0075] Step S203: Perform feature extraction operation on the four-dimensional light field data information to obtain feature matrix information.

[0076] The feature matrix information includes remote spatial information, remote angle information, epipolar image horizontal structure information, and epipolar image vertical structure information.

[0077] In this application, four cascaded light field selective structured state space model blocks (LFSS-SSM-Block) are used to extract depth features and model long-range dependencies for the sub-aperture image, macro-pixel data, horizontal epipolar image, and vertical epipolar image. The sub-aperture image, macro-pixel data, horizontal epipolar image, and vertical epipolar image are input into the four cascaded light field selective structured state space model blocks for feature extraction, and the corresponding sub-aperture depth feature information, macro-pixel depth feature information, horizontal epipolar depth feature information, and vertical epipolar depth feature information are output. That is, the sub-aperture image, macro-pixel data, horizontal epipolar image, and vertical epipolar image each need to be extracted four times for depth feature extraction.

[0078] Since the processing of each light field selective structured state space model block is the same, the subsequent cascaded blocks only input the depth features extracted by the previous block and perform the same operation. Therefore, to simplify understanding and avoid excessive repetition, this example uses the sub-aperture image input to illustrate the processing of the first light field selective structured state space model block, and subsequent blocks will not be repeated.

[0079] First, the sub-aperture image is normalized, then scanned in space using the SS2D Block (core submodule of LFSS-SSM-Block) to achieve efficient feature extraction, and then the feature information after spatial scanning is connected twice with learnable scale residual, and then a convolution attention operation is performed, to obtain the sub-aperture depth feature information. The process can also be represented as:

[0080] (2)

[0081] wherein, and represent the learnable scale parameters of the first and second residual connections, and represent the first and sub-aperture depth feature information (output after the second residual connection), , and represent channel attention, convolution, and normalization operations, respectively.

[0082] After obtaining the sub-aperture depth feature information, it is connected with the shallow initialization feature information to obtain the long-range spatial information The macro-pixel depth feature information is connected with the shallow initialization feature information in residual, the detail difference of the same position under different perspectives is captured to obtain remote angle information The horizontal epipolar depth feature information , the vertical epipolar depth feature information is connected with the shallow initialization feature information in residual, and then the horizontal epipolar image structure information , the vertical epipolar image structure information is obtained.

[0083] Step S204: performing fusion and splicing operation on the feature matrix information to obtain fused and spliced feature information.

[0084] In order to realize efficient fusion of light field space-angle and bidirectional structure information, and thus obtain fused and spliced feature information, the specific calculation process is as follows:

[0085] First, the remote space information and the remote angle information are fused to obtain a remote angle-space splicing sequence; the horizontal epipolar image structure information and the vertical epipolar image structure information are fused to obtain a horizontal-vertical structure splicing sequence, and then the remote angle-space splicing sequence and the horizontal-vertical structure splicing sequence are spliced to obtain the fused and spliced feature information.

[0086] The present application adopts the SSM attention cross-fusion enhancement module as shown in Figure 5 to complete the above fusion operation, and the SSM-attention cross-fusion enhancement module includes a cross-selective scanning module and a splicing-selective scanning module, and the calculation process in the cross-selective scanning module is as follows:

[0087] First, the remote space information is input, then it is shallowly processed through two linear projection operations, then the remotely processed space information is state space scanned, and the remote space projection matrix and the remote space hidden state corresponding to the remote space information are obtained through the state space model SSM (State Spatial Model). The state space model is a kind of linear time-invariant system, which maps the relationship between 1D sequence input and output through hidden state , including state equation and output equation:

[0088] (3)

[0089] (4)

[0090] wherein, is the current hidden state, is the previous hidden state, is the current input, A is the system state transition matrix, , B , C is the projection matrix, , .

[0091] The system state transition matrix A and the projection matrix B , C are obtained by performing SSM scanning on the input information after shallow processing, which ensures efficient fusion between spatial-angular information and bidirectional structure information, and improves the accuracy of semantic segmentation. The current hidden state is determined according to the previous hidden state and the current input. That is, is obtained by performing SSM scanning on the remote angular information after shallow processing. The remote angular hidden state is determined according to the previous remote angular hidden state and the shallow processed remote angular information of the current input.

[0092] Since cross fusion is required here to obtain the output, the remote angular projection matrix corresponding to the remote angular information and the remote angular hidden state are also determined according to the remote angular information, the horizontal structure projection matrix corresponding to the horizontal structure information of the epipolar image and the horizontal structure hidden state are determined according to the horizontal structure information of the epipolar image, and the vertical structure projection matrix corresponding to the vertical structure information of the epipolar image and the vertical structure hidden state are determined according to the vertical structure information of the epipolar image. Then cross fusion is performed according to formula (4), that is, the contents of formula (5) and (6):

[0093]

[0094] (5)

[0095] (6)

[0096] That is, the remote angular projection matrix and the remote spatial hidden state ​​The cross fusion is performed, and remote space-angle cross fusion information is output The remote space projection matrix and the remote angle hidden state The cross fusion is performed, and remote angle-space cross fusion information is output The horizontal structure projection matrix and the vertical structure hidden state The cross fusion is performed, and vertical-horizontal structure cross fusion information is output The vertical structure projection matrix and the horizontal structure hidden state The cross fusion is performed, and horizontal-vertical structure cross fusion information is output .

[0097] After obtaining the cross fusion information, it is input into the stitching selective scanning module after layer normalization and linear projection, and fusion processing is performed. The specific processing in the stitching scanning module is as follows:

[0098] Here, the fusion process of the remote space-angle cross fusion information and the remote angle-space cross fusion information is described as follows:

[0099] The remote space-angle cross fusion information and the remote angle-space cross fusion information are subjected to linear projection and convolution operation, and then are scanned in the state space, are expanded into 1D sequences and are spliced, so as to obtain a remote angle-space splicing sequence, as shown in the following formula:

[0100] (7)

[0101] In the formula, represents and the sequence after 1D dimension splicing, i.e., the remote angle-space splicing sequence.

[0102] It should be noted that, in order to facilitate the understanding of too many variables, the remote space-angle cross fusion information and the remote angle-space cross fusion information after linear projection and convolution operation in the cross selective scanning module after layer normalization and linear projection are still represented by and variables, and the representation of and variables in the following is the same, and details are not described here.

[0103] Then the remote angle-space concatenation sequence is reversely scanned to obtain a remote angle-space reverse concatenation sequence , as follows:

[0104] (8)

[0105] wherein, represents the sequence after reverse scanning, i.e. the remote angle-space reverse concatenation sequence .

[0106] Then the remote angle-space concatenation sequence and the remote angle-space reverse concatenation sequence are fused to obtain a remote angle-space fusion feature sequence .

[0107] (9)

[0108] Since the calculation process of fusion is the same, the calculation process of fusion of the vertical-horizontal structure cross fusion information and the horizontal-vertical structure cross fusion information will not be elaborated here. Only the process of fusion of the vertical-horizontal structure cross fusion information and the horizontal-vertical structure cross fusion information will be briefly described as follows:

[0109] The vertical-horizontal structure cross fusion information and the horizontal-vertical structure cross fusion information are concatenated to obtain a horizontal-vertical structure concatenation sequence , and then the horizontal-vertical structure concatenation sequence is reversely scanned to obtain a horizontal-vertical structure reverse concatenation sequence , and then the horizontal-vertical structure concatenation sequence and the horizontal-vertical structure reverse concatenation sequence are fused to obtain a horizontal-vertical fusion feature sequence .

[0110] Then the remote angle-space fusion feature sequence and the horizontal-vertical fusion feature sequence are concatenated by Concat to unify the data, and finally the output, i.e. the fusion concatenation feature information is obtained.

[0111] Step S205: determining the class prediction of each pixel based on the fusion concatenation feature information.

[0112] After obtaining the fusion splicing feature information , the fusion splicing feature information is input into four cascaded channel-aware visual state space model decoders (Channel-aware Visual SSM Decoder) for step-by-step up-sampling, and the channel-aware four-cascaded decoding up-sampling feature tensor is output. The specific process is as follows:

[0113] (10)

[0114] Then, the channel-aware four-cascaded decoding up-sampling feature tensor is input into a classifier for classification operation to generate a semantic segmentation result and output the class prediction of each pixel.

[0115] Experimental simulation stage

[0116] 1. Experimental data set: The effectiveness of the method is verified on the synthetic subset UrbanLF-Syn (250 scenes) and the real subset UrbanLF-Real (824 scenes) of the UrbanLF data set. The UrbanLF data set is the only available light field semantic segmentation benchmark to date, and each scene contains 81 regularly arranged 9 × 9 angular resolution sub-aperture images. The images in the data set depict urban road scenes, including complex occlusions and lighting conditions.

[0117] 2. Evaluation index: Acc, mAcc and mIoU are selected to evaluate the performance of the method from the overall accuracy of the segmentation result, the segmentation effect of each class, and the comprehensive effect of classification accuracy and segmentation precision.

[0118] Table 1: Quantitative experimental results of the method on the UrbanLF-Syn data set

[0119] ;

[0120] Table 2: Quantitative experimental results of the method on the UrbanLF-Real data set

[0121] ;

[0122] 3. Experimental details: The training and inference of the method proposed in this application use the Pytorch framework, and are performed on two RTX 4090 GPUs on an Ubuntu 20.04 system. In the training process, the AdamW optimizer is used, the initial learning rate is 6e-5, and the weight decay is 0.01. The model is trained for 500 epochs, and the batch size is 2. Table 1 and Table 2 are the quantitative experimental results of the method on the UrbanLF-Syn and UrbanLF-Real data sets, Figure 6 , Figure 7 respectively are the qualitative visualization results of the method on the UrbanLF-Syn and UrbanLF-Real data sets.

[0123] 4. Experimental analysis

[0124] (1) Experimental data comparison: Experiments on the UrbanLF-Syn dataset show that the mIoU of the method reaches 85.97%, which is 6.38% higher than the existing SOTA method.

[0125] (2) Computational complexity analysis: Compared with LF-IENet++, the computational efficiency is higher.

[0126] (3) Visual effect comparison: From the visualization results of the UrbanLF-Real dataset, the segmentation accuracy of the method in the occluded area and the complex edge target is significantly better than that of the traditional method.

[0127] In summary, the method proposed in this application uses SS-SSM to efficiently extract the spatial-angle and bidirectional structure information of the light field, enhances the target area recognition ability, and reduces the occlusion and misclassification problems. The linear time calculation mechanism of SS-SSM greatly reduces the computational overhead compared to the Transformer method, improves the segmentation efficiency, and is suitable for large-scale light field data processing. Through multi-dimensional feature learning of the light field and SSM-attention cross fusion module, fine information integration is realized, and the model generalization and stability are improved.

[0128] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-dimensional feature extraction and fusion method for light field semantic segmentation based on SS-SSM, characterized in that, include: Acquire a four-dimensional light field image; The four-dimensional light field image is preprocessed to obtain preprocessed four-dimensional light field data information, which includes sub-aperture image, macro-pixel data, horizontal epipolar image, and vertical epipolar image. The four-dimensional light field data information is subjected to feature extraction to obtain feature matrix information, which includes remote spatial information, remote angle information, epipolar image horizontal structure information, and epipolar image vertical structure information. The feature matrix information is fused and spliced ​​to obtain fused and spliced ​​feature information; Based on the fused and stitched feature information, the category prediction for each pixel is determined; The step of performing feature extraction on the four-dimensional light field data to obtain feature matrix information includes: After the sub-aperture image, the macro-pixel data, the horizontal epipolar image, and the vertical epipolar image are respectively input into four cascaded optical field selective structure state space model modules for feature extraction, the corresponding sub-aperture depth feature information, macro-pixel depth feature information, horizontal epipolar depth feature information, and vertical epipolar depth feature information are output. The sub-aperture depth feature information, the macro-pixel depth feature information, the horizontal epipolar depth feature information, and the vertical epipolar depth feature information are respectively residually concatenated with the shallow initialization feature information to obtain remote spatial information, remote angle information, epipolar image horizontal structure information, and epipolar image vertical structure information; the shallow initialization feature information is the feature information obtained after standardization and convolution processing of the four-dimensional light field image. The process of fusing and concatenating the feature matrix information to obtain fused and concatenated feature information includes: The remote spatial information and the remote angle information are fused and stitched together to output a remote angle-space stitched sequence; The horizontal structure information and the vertical structure information of the epipolar image are fused and stitched together to output a horizontal-vertical structure stitched sequence. The remote angle-space stitching sequence and the horizontal-vertical structure stitching sequence are fused and stitched together to output fused stitching feature information.

2. The method according to claim 1, characterized in that, The method for acquiring the four-dimensional light field image is one of the following: Light field camera system or dataset.

3. The method according to claim 1, characterized in that, The preprocessing operation on the four-dimensional light field image to obtain preprocessed four-dimensional light field data information includes: The four-dimensional light field image is standardized to obtain standardized four-dimensional light field data information, wherein the standardized four-dimensional light field data information is data information in the first interval or data information in the second interval of pixel values. The standardized four-dimensional light field data is convolved to obtain shallow initialization feature information. Based on the shallow initialization feature information, the preprocessed four-dimensional light field data information is obtained.

4. The method according to claim 3, characterized in that, The process of obtaining preprocessed four-dimensional light field data information based on the shallow initialization feature information includes: The shallow initialization feature information is divided to obtain the divided first shallow initialization feature information, second shallow initialization feature information, third shallow initialization feature information, and fourth shallow initialization feature information; The first shallow layer initialization feature information is used to fix the angle dimension through a Reshape operation to obtain a sub-aperture image; The second shallow layer initialization feature information is fixed in spatial dimension by a Reshape operation to obtain macro-pixel data; The third shallow layer initialization feature information is processed by a Reshape operation to fix the height dimension and vertical viewpoint, resulting in a horizontal epipolar image; The fourth shallow layer initialization feature information is processed by a Reshape operation to fix the width dimension and horizontal viewpoint, resulting in a vertical epipolar image.

5. The method according to claim 1, characterized in that, The process of fusing and stitching the remote spatial information and the remote angle information to output a remote angle-space stitched sequence includes: Based on the remote spatial information, determine the remote spatial projection matrix and remote spatial hiding state corresponding to the remote spatial information; Based on the remote angle information, determine the remote angle projection matrix and remote angle hiding state corresponding to the remote angle information; The remote spatial projection matrix and the remote angle hidden state are cross-fused to output remote angle-space cross-fusion information. The remote angle projection matrix and the remote spatial hidden state are cross-fused to output remote space-angle cross-fusion information. The remote angle-space cross-fusion information and the remote space-angle cross-fusion information are spliced ​​together to output a remote angle-space spliced ​​sequence.

6. The method according to claim 1, characterized in that, The process of fusing and stitching the horizontal and vertical structure information of the epipolar image to output a horizontal-vertical structure stitched sequence includes: Based on the horizontal structure information of the epipolar image, determine the horizontal structure projection matrix and the horizontal structure hidden state corresponding to the horizontal structure information of the epipolar image; Based on the vertical structure information of the epipolar image, determine the vertical structure projection matrix and the vertical structure hidden state corresponding to the vertical structure information of the epipolar image; The horizontal structure projection matrix and the vertical structure hidden state are cross-fused to output vertical-horizontal structure cross-fusion information. The vertical structure projection matrix and the horizontal structure hidden state are cross-fused to output horizontal-vertical structure cross-fusion information. The vertical-horizontal structure cross-fusion information and the horizontal-vertical structure cross-fusion information are spliced ​​together to output a horizontal-vertical structure splicing sequence.

7. The method according to claim 1, characterized in that, The process of fusing and splicing the remote angle-space stitching sequence and the horizontal-vertical structure stitching sequence to output fused stitching feature information includes: The remote angle-space stitching sequence is scanned in reverse order to obtain a remote angle-space reverse stitching sequence; The horizontal-vertical structure splicing sequence is scanned in reverse order to obtain the horizontal-vertical structure reverse splicing sequence; The remote angle-space stitched sequence and the remote angle-space reverse stitched sequence are fused to output a remote angle-space fused feature sequence; The horizontal-vertical structure splicing sequence and the horizontal-vertical structure reverse splicing sequence are fused to output a horizontal-vertical fused feature sequence; The remote angle-space fusion feature sequence and the horizontal-vertical fusion feature sequence are spliced ​​together to output the fused spliced ​​feature information.

8. The method according to claim 1 or 7, characterized in that, The step of determining the category prediction for each pixel based on the fused and stitched feature information includes: The fused and stitched feature information is input into four cascaded channel-aware visual state space model decoders, and the fused and stitched feature information is upsampled step by step to output the channel-aware four-cascaded decoded upsampled feature tensor. The channel-aware four-cascade decoding upsampled feature tensor is input into the classifier for classification, and the class prediction for each pixel is output.

Citation Information

Patent Citations

  • A light field image mosaic method based on super pixel

    CN109064410A

  • Light field camera semantic segmentation method and system and electronic equipment

    CN118552737A