A Method and System for Semantic Segmentation of Light Fields in Complex Urban Scenes Based on State Space Model

Through the Mamba model based on the state space model, the calculation complexity and information fusion problems of light field semantic segmentation in complex urban scenes are solved, and efficient and accurate light field semantic segmentation is achieved.

CN120047690BActive Publication Date: 2025-07-18HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510535732.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-18
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing technology of the light field semantic segmentation method in complex urban market scenes faces the problems of high computational complexity, insufficient long-range dependency modeling and difficulty in fusion of multimodal information, and is difficult to meet the needs of dynamic occlusion and lighting changes.

Method used

Using a state space model-based method, the 4D light field image is extracted and interactively enhanced by the Mamba model, and combined with the feature data of the central view, depth map and 4D light field image, cross-modal and cross-view are performed to realize the semantic segmentation of the light field.

Benefits of technology

It effectively solves the problems of high computational complexity, insufficient long-range dependency modeling and difficulty in fusion of multimodal information, and realizes efficient and accurate light field semantic segmentation of complex urban market scenarios, improving segmentation accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047690B_ABST
    Figure CN120047690B_ABST
Patent Text Reader

Abstract

The present invention discloses a complex urban scene light field semantic segmentation method and system based on a state space model. The method includes: preprocessing the central view, depth map, and 4D light field image of the target scene and respectively inputting them into the first branch, the second branch, and the third branch for multi-scale feature extraction; the feature extraction of the third branch is based on the first strategy; based on the second strategy, the feature data of each scale in the outputs of the three branches are fused to obtain a fused data set; according to the fused data set, a light field semantic segmentation map of the target scene is obtained; the first strategy includes: extracting spatial long-range context features through the Mamba model according to the 4D light field image to obtain a first result; performing bidirectional scanning according to the first result through the Mamba model and modeling the complementary relationship features between angles; the second strategy includes: performing interactive enhancement on the feature data from the three branches through the Mamba model, splicing them, and performing reverse scanning to obtain fused enhanced features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and in particular, to a method and system for semantic segmentation of light fields in complex urban scenes based on a state space model. Background Art

[0002] Semantic segmentation is a core task in computer vision, aiming to assign semantic labels to each pixel in an image, and is an important basis for autonomous driving, smart cities, and robot perception. In complex urban scenes, due to the dense distribution of targets such as buildings, vehicles, and pedestrians, as well as factors such as dynamic occlusion and lighting changes, traditional monocular image segmentation methods are difficult to meet the requirements.

[0003] 4D light fields record the direction and position information of light rays to form a regularly arranged dense viewpoint array and macro-pixel array, containing rich spatial-angular information, which can effectively solve the problems of occlusion and lighting interference in complex urban scenes. However, existing methods for light field semantic segmentation face problems such as high computational complexity, insufficient long-range dependence modeling, and difficulty in fusing dynamic information. In terms of high computational complexity, methods based on Transformer rely on self-attention mechanisms, and the computational complexity grows quadratically, making it difficult to efficiently process high-dimensional light field data; in terms of insufficient long-range dependence modeling, traditional convolutional neural networks are limited by local receptive fields and cannot fully extract the multi-dimensional spatial-angular information embedded in light fields; in terms of difficulty in fusing dynamic information, existing methods are difficult to effectively fuse the multi-view, macro-pixel, and multi-modal (such as depth information) features of light fields, resulting in limited segmentation accuracy in complex urban scenes. For example, some researchers have proposed multi-view fusion methods based on Transformer, but they cannot handle a large amount of viewpoint information simultaneously; other researchers have used CNNs to extract sub-aperture image features, but they have insufficient modeling ability for global spatial-angular relationships and are difficult to cope with dynamic occlusion and lighting changes in urban complex scenes.

[0004] In recent years, SSM (State Space Model) as an efficient sequence modeling method has gradually shown great potential in the fields of natural language processing and computer vision. SSM models long-range dependence relationships with linear complexity and can effectively process global information in high-dimensional data. As a new architecture based on SSM, the Mamba model introduces a selective scanning mechanism that can dynamically adjust model parameters according to the input, significantly improving the model's performance in long-sequence tasks and having high efficiency and robustness when processing high-dimensional data. However, existing research mainly focuses on monocular images or low-dimensional data, and the Mamba model has not been applied to the semantic segmentation task of 4D light field data.

[0005] Therefore, there is an urgent need for a new technical solution to solve the technical problem of how to overcome existing defects such as high computational complexity, insufficient long-range dependence modeling, and difficulty in multi-modal information fusion to perform light field semantic segmentation on complex urban scenes. Summary of the Invention

[0006] The present invention provides a method and system for light field semantic segmentation of complex urban scenes based on a state space model to solve the technical problem of how to overcome existing defects such as high computational complexity, insufficient long-range dependence modeling, and difficulty in multi-modal information fusion to perform light field semantic segmentation on complex urban scenes.

[0007] To achieve the above object, the present invention provides a method for light field semantic segmentation of complex urban scenes based on a state space model, including:

[0008] Obtain the central view, depth map, and 4D light field image of the target scene; preprocess the central view, depth map, and 4D light field image and then input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction.

[0009] The feature extraction of the first branch and the second branch is based on a visual state space block, and the feature extraction of the third branch is based on a first strategy; the first branch, the second branch, and the third branch all output feature data at a preset multi-scale.

[0010] Fuse the feature data at each scale in the outputs of the first branch, the second branch, and the third branch based on a second strategy to obtain a fused data set; decode the fused data set to obtain the light field semantic segmentation map of the target scene.

[0011] The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, extracting spatial long-range context features of the 2D sub-aperture slices through a Mamba model to obtain a first result; obtaining macro-pixel slices based on the first result, and performing two-way scanning and modeling the complementary relationship features between angles on the macro-pixel slices through the Mamba model.

[0012] The second strategy includes: enhancing the interaction of the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively, and splicing the first data, the second data, and the third data and then performing reverse scanning to obtain fused enhanced features.

[0013] Preferably, obtaining the central view, depth map, and 4D light field image of the target scene includes:

[0014] Obtain the 4D light field image .

[0015] According to the 4D light field image Obtain the central view .

[0016] According to the central view obtain the depth map .

[0017] Among them, represents a matrix with light field parameters of U and V , height of H , and width of W ; represents a matrix with height of H , width of W , and number of channels of 3; represents a matrix with height of H , and width of W .

[0018] Preferably, the preprocessing includes:

[0019] Normalize the central view, depth map, and 4D light field image:

[0020] The normalization of the 4D light field image includes:

[0021] ;

[0022] Among them, represents the 4D light field image after normalization; represents the 4D light field image; represents the mean value of the 4D light field image; represents the standard deviation of the 4D light field image.

[0023] The normalization of the central view and depth map includes:

[0024] ;

[0025] Among them, represents the central view or depth map after normalization; represents the central view or depth map; represents the mean value of the central view or depth map; represents the standard deviation of the central view or depth map.

[0026] Perform data augmentation on the normalized central view, depth map, and 4D light field image using random rotation and color jitter;

[0027] Segment the central view, depth map, and 4D light field image that have undergone the said normalization processing and the said data augmentation into local blocks through a ViT-style image chunking module, and map them to a high-dimensional feature space, including:

[0028] ;

[0029] Among them, represents the image data output by the image block module, , represents the height of , the width of , and the number of channels of matrix; ; represents the scale level; represents the cropping module.

[0030] Preferably, a 2D sub-aperture slice is obtained from the 4D light field image, and the spatial long-range context features of the 2D sub-aperture slice are extracted through the Mamba model. The first result includes:

[0031] The 4D light field image is decoupled into 2D sub-aperture slices through the Reshape operation .

[0032] The 2D sub-aperture slice is unfolded into a 1D sequence along the horizontal or vertical direction through the Mamba model and the spatial long-range context features are extracted to obtain an intermediate quantity :

[0033] ;

[0034] Among them, represents the layer normalization operation; represents the proportional parameter of the residual connection; represents the light field spatial selective scanning module;

[0035] The are respectively passed through layer normalization , convolution , channel attention and the residual connection with a proportional parameter of to obtain an intermediate quantity :

[0036] ;

[0037] The is reshaped through the Reshape operation to obtain the first result .

[0038] Preferably, a macro-pixel slice is obtained from the first result, and the complementary relationship features between angles are modeled by bidirectionally scanning the macro-pixel slice through the Mamba model, including:

[0039] The first result is reshaped through the Reshape operation to obtain the macro-pixel slice .

[0040] The macro-pixel slices are bidirectionally scanned by the Mamba model and the complementary relationship features between angles are modeled to obtain an intermediate quantity :

[0041] ;

[0042] Among them, represents the light field angle selective scanning module;

[0043] The are respectively passed through layer normalization , convolution , channel attention and a residual connection with a scaling parameter of to obtain an intermediate quantity :

[0044] ;

[0045] Through the Reshape operation, the is reshaped to obtain the output feature data in 4D representation .

[0046] Preferably, the feature data from the first branch, the second branch, and the third branch are interactively enhanced by the Mamba model to respectively obtain the first data, the second data, and the third data, including:

[0047] ;

[0048] ;

[0049] ;

[0050] ;

[0051] Among them, , represents the central view; represents the depth map; represents the 4D light field view; represents the discretized system transition matrix; represents the exponential function; represents the adjustment parameter of the time step; represents the system state transition matrix; represents the system projection matrix; represents the discretized system projection matrix; represents the time step t 's hidden state; represents the time step tThe hidden state of -1; Denote the time step t The input at; Denote the time step t The output of the central view at; Denote the cross-fusion system matrix; Denote the time step t The hidden state of the central view at; Denote the time step t The selective scan output of the depth map branch of the state at; Denote the central view system matrix; Denote the time step t The output of the depth view at; Denote the time step t The selective scan output of the light field image branch of the state at; Denote the time step t The hidden state of the depth map at; Denote the output; Denote the time step t The selective scan output of the state at.

[0052] ;

[0053] Wherein, And Are scale parameters; Denote the depth map system matrix at time step t; Denote the 4D light field image system matrix at time step t.

[0054] Preferably, after splicing the first data, the second data and the third data and performing reverse scanning, the fusion enhanced features obtained include:

[0055] ;

[0056] ;

[0057] ;

[0058] ;

[0059] ;

[0060] Wherein, Denote the image feature; Denote the depthwise separable convolution; Denote the linear projection layer; Denote the spliced image sequence; Denote the splicing operation; Represents the central view image features; Represents the depth view image features; Represents the light field view image features; Represents the dimension; Represents the sequence of image features after inverse scanning; Represents the inverse scanning operation; Represents the Mamba model; Represents the output enhanced image features; Represents the feature separation operation; Represents the sequence of stitched image features output by the selective scanning process of the Mamba model; Represents the sequence of inverse scanned image features output by the selective scanning process of the Mamba model.

[0061] Obtain the scanning features and After that, along the channel dimension, and Are concatenated to obtain the cross-modal and cross-view fusion enhanced first feature with a shape of ; Represents a height of , a width of , and the number of channels is 3× C k Matrix.

[0062] Adjust the shape of the fusion enhanced first feature to through a linear projection layer to obtain the fusion enhanced feature.

[0063] Preferably, decoding according to the fusion dataset to obtain the light field semantic segmentation map of the target scene includes:

[0064] Decoding according to the fusion enhanced features at each scale in the fusion dataset in combination with a channel-aware visual state space decoder, including:

[0065] Assume that the fusion dataset includes Q Fusion enhanced features at scales, that is, including Q Fusion enhanced features, where the scales of the 1st to the Q th fusion enhanced features decrease.

[0066] Enhance and upsample the Q th fusion enhanced feature to restore the scale to the scale of the Q -1th fusion enhanced feature, and splice it with the Q -1th fusion enhanced feature and input it into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs the first decoded data.

[0067] Upsample the first decoded data to restore the scale to the scale of the Q -2 fused enhanced features, and splice them with the Q -2 fused enhanced features and then input them into the channel-aware visual state space decoder for decoding. The channel-aware visual state space decoder outputs the second decoded data.

[0068] Continue to perform upsampling, feature splicing, and decoding based on the second decoded data until the scale of the decoded data is restored to the scale of the first fused enhanced feature. The channel-aware visual state space decoder outputs the decoded completed data.

[0069] After decoding is completed, upsample the decoded completed data to restore the scale to the scale of the central view to obtain the decoded image; input the decoded image into the Softmax classifier, and the Softmax classifier outputs the light field semantic segmentation map of the target scene.

[0070] The present invention also provides a light field semantic segmentation system for complex urban scenes based on a state space model for the method of the present invention. The system includes a first module, a second module, a third module, and a fourth module.

[0071] The first module is used to obtain the central view, depth map, and 4D light field image of the target scene.

[0072] The second module is used to preprocess the central view, depth map, and 4D light field image and then input them into the first branch, the second branch, and the third branch respectively for multi-scale feature extraction; the feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; the first branch, the second branch, and the third branch all output feature data at a preset multi-scale.

[0073] The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result; obtaining macro-pixel slices from the first result, and performing two-way scanning and modeling the complementary relationship features between angles on the macro-pixel slices through the Mamba model.

[0074] The third module is used to fuse the feature data at each scale in the outputs of the first branch, the second branch, and the third branch based on the second strategy to obtain a fused data set.

[0075] The second strategy includes: performing interactive enhancement on the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively, splicing the first data, the second data, and the third data and then performing reverse scanning to obtain the fused enhanced features.

[0076] The fourth module is used to decode according to the fused dataset to obtain the light field semantic segmentation map of the target scene.

[0077] The present invention has the following beneficial effects:

[0078] The method for light field semantic segmentation of complex urban scenes based on the state space model of the present invention can comprehensively extract multi-scale long-distance spatial angle information by performing multi-scale feature extraction on 4D light field images based on the first strategy; the extracted features are integrated with multi-scale modal information of the other two branches based on the second strategy, which promotes feature fusion and the enhancement of the central view, and realizes the adaptive fusion of light field multi-modal information. By decoupling the spatial and angular information of the light field and performing independent modeling respectively, the limitations in long-range dependence modeling of traditional methods are solved; the method of the present invention provides an efficient and accurate solution for dynamic occlusion, illumination change and multi-scale target segmentation in complex urban scenes. The method of the present invention overcomes the existing defects such as high computational complexity, insufficient long-range dependence modeling and difficult multi-modal information fusion, and can perform light field semantic segmentation on complex urban scenes.

[0079] The system for light field semantic segmentation of complex urban scenes based on the state space model of the present invention is used for the method of the present invention and has the same beneficial effects as the method of the present invention.

[0080] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0081] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0082] Figure 1 is a schematic flowchart of the method of the preferred embodiment of the present invention.

[0083] Figure 2 is a schematic flowchart of the light field semantic segmentation process of the preferred embodiment of the present invention.

[0084] Figure 3 is a schematic diagram of the first strategy of the preferred embodiment of the present invention.

[0085] Figure 4 is a schematic flowchart of cross-selective scanning based on the cross-selective scanning module of the preferred embodiment of the present invention.

[0086] Figure 5 is a schematic flowchart of feature scanning based on the connection selection scanning module of the preferred embodiment of the present invention. Detailed Embodiments

[0087] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.

[0088] See Figures 1 to 2 , in a preferred embodiment of the present invention, a method for semantic segmentation of light fields in complex urban scenes based on a state space model is provided, including:

[0089] P1. Obtain the central view, depth map, and 4D light field image of the target scene. P1 specifically includes:

[0090] Obtain the 4D light field image .

[0091] According to the 4D light field image obtain the central view .

[0092] According to the central view obtain the depth map .

[0093] Among them, represents a matrix with light field parameters of U and V , height of H , and width of W ; represents a matrix with height of H , width of W , and number of channels of 3; represents a matrix with height of H , width of W .

[0094] P2. After preprocessing the central view, depth map, and 4D light field image, input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction, where the feature extraction of the third branch is based on the first strategy.

[0095] In a preferred embodiment of the present invention, the preprocessing includes:

[0096] Normalize the central view, depth map, and 4D light field image:

[0097] The normalization processing of the 4D light field image includes:

[0098] ;

[0099] Among them, represents the normalized 4D light field image; represents the 4D light field image; represents the mean value of the 4D light field image; Represents the standard deviation of the 4D light field image.

[0100] The normalization processing of the central view and the depth map includes:

[0101] ;

[0102] Among them, Represents the central view or depth map after normalization processing; Represents the central view or depth map; Represents the mean value of the central view or depth map; Represents the standard deviation of the central view or depth map.

[0103] The normalized central view, depth map, and 4D light field image are subjected to data augmentation by random rotation (±15°) and color jitter (brightness, contrast ±10%).

[0104] The normalized and data-augmented central view, depth map, and 4D light field image are segmented into local blocks by an image patch module in the ViT style and mapped to a high-dimensional feature space, including:

[0105] ;

[0106] Among them, Represents the image data output by the image patch module, , Represents the height as and the width as with the number of channels being C k matrix; ; Represents the scale level; Represents the cropping module.

[0107] In the preferred embodiment of the present invention, when performing multi-scale feature extraction, it is preferably to perform feature extraction at 4 levels of scales, that is, the preset multi-scale includes 4 levels of scales. Therefore , .

[0108] In the preferred embodiment of the present invention, the feature extraction of the first branch and the second branch is based on the Visual State Space (VSS), and the feature extraction of the third branch is based on the first strategy; the first branch, the second branch, and the third branch all output feature data at the preset multi-scale.

[0109] In the preferred embodiment of the present invention, in the visual state space block, first, the input features are processed by layer normalization, linear projection, and depth wise convolution to extract local information and maintain spatial structure. Then, the selective scan 2D (SS2D) module is introduced to model the long-range spatial information in the features. Specifically, the SS2D module converts the shape The input features are flattened into 4 along four directions (top left to bottom right, bottom right to top left, top right to bottom left, bottom left to top right). The sequences are then transformed into four different selective scanning modules to extract long-range dependencies in multiple directions. Finally, the sequences are reshaped back to their original shape and summed to integrate multi-directional information, and finally the feature representation is enhanced through residual connections to effectively model the input data.

[0110] In a preferred embodiment of the present invention, a first strategy includes: obtaining 2D sub-aperture slices based on the 4D light field image, performing spatial long-range context feature extraction on the 2D sub-aperture slices through a Mamba model to obtain a first result; obtaining macro-pixel slices based on the first result, performing bidirectional scanning on the macro-pixel slices through a Mamba model and modeling complementary relationship features between angles.

[0111] In a preferred embodiment of the present invention, the first strategy dynamically adjusts the SSM parameters to preferentially select areas near the target point for sampling, thereby significantly improving the efficiency of path search and solving the shortcomings of traditional methods in long-range dependency modeling.

[0112] See also Figure 3 In a preferred embodiment of the present invention, a 2D sub-aperture slice is obtained according to the 4D light field image, and a spatial long-range context feature is extracted from the 2D sub-aperture slice by using a Mamba model, and the first result obtained includes:

[0113] Decoupling the 4D light field image into 2D sub-aperture slices via the Reshape operation .

[0114] The 2D sub-aperture slices are expanded into 1D sequences along the horizontal or vertical direction through the Mamba model and the spatial long-range context features are extracted to obtain the intermediate quantity :

[0115] ;

[0116] in, Representation layer normalization operation; Represents the scale parameter of the residual connection; represents the light field space selective scanning module, that is,Figure 3 The LFSpaSS module in

[0117] is used to extract the light field spatial feature sequence. are respectively passed through layer normalization , convolution , channel attention and a residual connection with a scaling parameter of to obtain an intermediate quantity :

[0118] ;

[0119] Through the Reshape operation, is reshaped to obtain the first result .

[0120] In the preferred embodiment of the present invention, according to the first result, macro-pixel slices are obtained, and the complementary relationship features between angles are modeled by bidirectionally scanning the macro-pixel slices through the Mamba model, including:

[0121] Through the Reshape operation, the first result is reshaped to obtain macro-pixel slices .

[0122] The complementary relationship features between angles are modeled by bidirectionally scanning the macro-pixel slices through the Mamba model to obtain an intermediate quantity :

[0123] ;

[0124] wherein, represents the light field angular selective scanning module, that is, Figure 3 the LFAngSS module in

[0125] is used to extract the light field angular feature sequence. are respectively passed through layer normalization , convolution , channel attention and a residual connection with a scaling parameter of to obtain an intermediate quantity :

[0126] ;

[0127] Through the Reshape operation, is reshaped to obtain the output feature data in 4D representation .

[0128] P3. For the feature data at each scale in the outputs of the first branch, the second branch, and the third branch, perform fusion based on the second strategy to obtain a fused dataset; decode according to the fused dataset to obtain the light field semantic segmentation map of the target scene.

[0129] See Figures 4 to 5 , in a preferred embodiment of the present invention, the second strategy includes: using the Mamba model to perform interactive enhancement on the feature data from the first branch, the second branch, and the third branch to respectively obtain first data, second data, and third data, and splicing the first data, the second data, and the third data and then performing reverse scanning to obtain fused enhanced features.

[0130] In a preferred embodiment of the present invention, using the Mamba model to perform interactive enhancement on the feature data from the first branch, the second branch, and the third branch to respectively obtain first data, second data, and third data includes:

[0131] See Figure 4 , use the feature data from the first branch, the second branch, and the third branch as inputs, and perform cross-selective scanning based on the cross-selective scanning module Cross SS to enhance the feature interaction between the central view, the depth map, and the 4D light field image:

[0132] ;

[0133] ;

[0134] ;

[0135] ;

[0136] Wherein, , represents the central view; represents the depth map; represents the 4D light field view; represents the discretized system transition matrix; represents the exponential function; represents the adjustment parameter of the time step; represents the system state transition matrix; represents the system projection matrix; represents the discretized system projection matrix; represents the hidden state at time step t; represents the hidden state at time step t-1; represents the time step t at the input; represents the time step t at the output of the central view; Represents the cross - fusion system matrix; Represents the time step t The hidden state of the central view at; Represents the time step t The selective scan output of the depth map branch of the state at; Represents the central view system matrix; Represents the time step t The output of the depth view at; Represents the time step t The selective scan output of the light field image branch of the state at; Represents the time step t The hidden state of the depth map at; Represents the output; Represents the time step t The selective scan output of the state at;

[0137] ;

[0138] Wherein, And Are scale parameters; Represents the depth map system matrix at time step t; Represents the 4D light field image system matrix at time step t.

[0139] In the preferred embodiment of the present invention, after splicing the first data, the second data and the third data and performing reverse scanning, the fusion - enhanced features obtained include:

[0140] See Figure 5 , Taking the first data, the second data and the third data as inputs, performing feature scanning based on the connection - selection scanning module Concat SS to optimize the feature fusion effect:

[0141] ;

[0142] ;

[0143] ;

[0144] ;

[0145] ;

[0146] Wherein, Represents the image feature; Represents the depth - separable convolution; Represents the linear projection layer; Represents the spliced image sequence; Represents the splicing operation; Represents the central view image features; Represents the depth view image features; Represents the light field view image features; Represents the dimension; Represents the sequence of image features after inverse scanning; Represents the inverse scanning operation; Represents the Mamba model; Represents the output enhanced image features; Represents the feature separation operation; Represents the sequence of stitched image features output after selective scanning by the Mamba model; Represents the sequence of inverse scanned image features output after selective scanning by the Mamba model.

[0147] Obtain the scanning features and After that, along the channel dimension, and are concatenated to obtain the cross-modal and cross-view fusion enhanced first feature with the shape of ; Represents the height of , the width of , and the number of channels is 3 × C k matrix.

[0148] The shape of the fusion enhanced first feature is adjusted to through a linear projection layer to obtain the fusion enhanced feature.

[0149] In the preferred embodiment of the present invention, decoding according to the fusion dataset to obtain the light field semantic segmentation map of the target scene includes:

[0150] Decoding according to the fusion enhanced features at each scale in the fusion dataset in combination with the Channel-Aware Visual State Space (CAVSS) decoder, including:

[0151] Assume that the fusion dataset includes Q fusion enhanced features at Q scales, that is, includes Q fusion enhanced features, where the scales of the 1st to the

[0152] th fusion enhanced features decrease. Q Enhance and upsample the Q -1st fusion enhanced feature to restore the scale to that of the Q-1 concatenated fusion enhanced features are input into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs the first decoded data.

[0153] The first decoded data is upsampled to restore the scale to that of the Q -2 fusion enhanced features, and concatenated with the Q -2 fusion enhanced features and then input into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs the second decoded data.

[0154] Based on the second decoded data, continue with upsampling, feature concatenation and decoding until the scale of the decoded data is restored to that of the 1st fusion enhanced feature, and the channel-aware visual state space decoder outputs the decoded completed data.

[0155] After decoding is completed, the decoded completed data is upsampled to restore the scale to that of the central view to obtain the decoded image; the decoded image is input into the Softmax classifier, and the Softmax classifier outputs the light field semantic segmentation map of the target scene.

[0156] In the preferred embodiment of the present invention, in the channel-aware visual state space decoder, the input features are first passed through the visual state space block to extract features; then the features are normalized by layer normalization; and then through convolution operations to further extract local feature information. The processed features are respectively input into maxpooling and avgpooling to generate multi-scale features through different pooling methods; then using the skip connection mechanism, the features of different scales are fused to enhance the multi-scale feature fusion effect.

[0157] In the channel-aware visual state space decoder, the multi-scale features are upsampled by combining the channel attention mechanism to enhance the expression ability of spatial-channel information. The CAVSS module restores the high-resolution segmentation result through the selective scanning mechanism, and combines the weighted combination of cross-entropy and Dice loss to optimize the segmentation accuracy. Based on the channel-aware visual state space decoder for data decoding, it can significantly improve the restoration ability of detailed features in complex urban scenes and solve the problem of segmentation accuracy fluctuation in dynamic occlusion and illumination change scenes.

[0158] The method for semantic segmentation of light fields in complex urban scenes based on the state space model of the present invention can comprehensively extract multi-scale long-distance spatial angular information by performing multi-scale feature extraction on 4D light field images based on the first strategy; the extracted features are integrated with the multi-scale modal information of the other two branches based on the second strategy, promoting feature fusion and the enhancement of the central view, and realizing the adaptive fusion of multi-modal information of light fields. By decoupling the spatial and angular information of light fields and performing independent modeling respectively, the limitations of traditional methods in long-range dependence modeling are solved; the method of the present invention provides an efficient and accurate solution for dynamic occlusion, illumination change, and multi-scale target segmentation in complex urban scenes. The method of the present invention overcomes the existing defects such as high computational complexity, insufficient long-range dependence modeling, and difficulty in multi-modal information fusion, and can perform semantic segmentation of light fields in complex urban scenes.

[0159] In a preferred embodiment of the present invention, a system for semantic segmentation of light fields in complex urban scenes based on the state space model is also provided for the method of the present invention. The system includes a first module, a second module, a third module, and a fourth module.

[0160] The first module is used to obtain the central view, depth map, and 4D light field image of the target scene.

[0161] The second module is used to preprocess the central view, depth map, and 4D light field image and then input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction; the feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; the first branch, the second branch, and the third branch all output feature data at a preset multi-scale.

[0162] The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result; obtaining macro-pixel slices from the first result, and performing two-way scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles.

[0163] The third module is used to fuse the feature data at each scale in the outputs of the first branch, the second branch, and the third branch based on the second strategy to obtain a fused data set.

[0164] The second strategy includes: performing interaction enhancement on the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively, and splicing the first data, the second data, and the third data and then performing reverse scanning to obtain fused enhanced features.

[0165] The fourth module is used to decode according to the fused data set to obtain the light field semantic segmentation map of the target scene.

[0166] The light field semantic segmentation system for complex urban scenes based on the state space model of the present invention, which is used for the method of the present invention, has the same beneficial effects as the method of the present invention.

[0167] Verification part:

[0168] The effectiveness of the method of the present invention is verified on the UrbanLF dataset. The UrbanLF dataset is the only available light field semantic segmentation benchmark to date. Each scene contains 81 regularly arranged sub-aperture images with a 9×9 angular resolution. The images in the dataset depict urban road scenes, including complex occlusion and lighting conditions. The statistical data of the dataset used in the verification are shown in Table 1, which details the number of categories, spatial resolution, angular resolution, and training set division method. Among them, UrbanLF-Real contains real-world data, which is characterized by greater noise and natural light changes compared to the synthetic UrbanLF-Syn dataset.

[0169] Table 1 Statistical data table of the dataset

[0170] ,

[0171] The method of the present invention is implemented using PyTorch in the Ubuntu environment. Training and inference are carried out on two RTX 4090 GPUs. During the training process, the AdamW optimizer with an initial learning rate of 6e-5 and a weight decay of 0.01 is used. The model is trained for 500 epochs with a batch size of 2. In addition, it is also compared with several state-of-the-art methods. The mainstream models based on LF and RGBD and the models based on the method of the present invention adopted in the preferred embodiment of the present invention are shown in Table 2. Among them, LFSSMam(S), LFSSMam(M), and LFSSMam(L) are models based on the method of the present invention, and S, M, and L represent the size of the model. The size of the model is determined by the parameters of the pre-trained VSS module in the model. The quantitative experimental data of all selected models on the UrbanLF-Syn dataset and the evaluation results regarding the mIoU, mAcc, and Acc metrics are shown in Tables 3 and 4, and the quantitative experimental data on the UrbanLF-Real dataset and the evaluation results regarding the mIoU, mAcc, and Acc metrics are shown in Tables 5 and 6.

[0172] Table 2 Table of mainstream models based on LF and RGBD and models based on the method of the present invention

[0173] ,

[0174] Table 3 Quantitative experimental data table of the model on the UrbanLF-Syn dataset

[0175] ,

[0176] Table 4 Evaluation Results of mIoU, mAcc and Acc Metrics of the Model on the UrbanLF-Syn Dataset

[0177] ,

[0178] Table 5 Quantitative Experimental Data Table of the Model on the UrbanLF-Real Dataset

[0179] ,

[0180] Table 6 Evaluation Results of mIoU, mAcc and Acc Metrics of the Model on the UrbanLF-Real Dataset

[0181] ,

[0182] Analysis based on Tables 2 to 6 shows that:

[0183] (1) Precision comparison: The mIoU index of LFSSMam-L reaches 89.66% (UrbanLF-Syn) and 86.54% (UrbanLF-Real), which is 6.97% better than the current best method.

[0184] (2) Robustness verification: In complex urban scenes, in the face of dynamic occlusion and illumination changes, the segmentation precision fluctuates less than 2%, which is significantly better than traditional methods.

[0185] (3) Visual qualitative verification: According to the experimental data of this experiment, corresponding visualization results can be obtained. From the visualization results, it can be seen that for objects with complex and delicate structures, such as bicycles and big trees, etc., the method of the present invention can generate more coherent semantic instances.

[0186] The method of the present invention realizes the efficient semantic segmentation of light field images in complex urban scenes through a selective state space model and a multi-modal fusion strategy, providing an innovative solution for scene understanding in fields such as smart cities and autonomous driving. Experiments show that this method has reached the leading level in the industry in terms of precision, efficiency and robustness.

[0187] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for semantic segmentation of light fields in complex urban scenes based on a state space model, characterized in that Including: Obtain the central view, depth map, and 4D light field image of the target scene; Preprocess the central view, depth map, and 4D light field image, and then input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction; The feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; the first branch, second branch, and third branch all output feature data at a preset multi-scale; Fuse the feature data at each scale in the outputs of the first branch, second branch, and third branch based on the second strategy to obtain a fused data set; Decode according to the fused data set to obtain the light field semantic segmentation map of the target scene; The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, and performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result; obtaining macro-pixel slices according to the first result, and performing two-way scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles; The second strategy includes: enhancing the interaction of the feature data from the first branch, second branch, and third branch through the Mamba model to obtain the first data, second data, and third data respectively, and splicing the first data, second data, and third data and then performing reverse scanning to obtain the fused enhanced feature.

2. The method for semantic segmentation of light fields in complex urban scenes based on the state space model according to claim 1, wherein The obtaining of the central view, depth map, and 4D light field image of the target scene includes: Obtain a 4D light field image ; According to the described 4D light field image obtain a central view ; According to the central view obtain a depth map ; Among them, represents a matrix with optical field parameters of U and V , height of H , and width of W ; represents a matrix with a height of H , a width of W , and 3 channels; represents a matrix with a height of H , and a width of W .

3. The method for semantic segmentation of light field in complex urban scenes based on the state space model according to claim 2, wherein The preprocessing includes: Performing normalization processing on the central view, depth map, and 4D light field image: The normalization processing of the 4D light field image includes: ; Among them, represents the 4D light field image after normalization processing; represents the 4D light field image; represents the mean value of the 4D light field image; represents the standard deviation of the 4D light field image; The normalization processing of the central view and depth map includes: ; Among them, represents the centered view or depth map after normalization processing; represents the centered view or depth map; represents the mean value of the centered view or depth map; represents the standard deviation of the centered view or depth map; Performing data augmentation on the normalized central view, depth map, and 4D light field image by using random rotation and color jitter; Segmenting the central view, depth map, and 4D light field image that have undergone the normalization processing and the data augmentation into local blocks through a ViT-style image chunking module, and mapping them to a high-dimensional feature space, including: ; Among them, represents the image data output by the image block module, , represents the height of , the width of , and the number of channels of matrix; ; represents the scale level; represents the cropping module.

4. The method for semantic segmentation of light field in complex urban scenes based on the state space model according to claim 3, wherein Obtaining 2D sub-aperture slices from the 4D light field image, and performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result includes: Decouple the 4D light field image into 2D sub-aperture slices through the Reshape operation ; The 2D sub-aperture slices are unfolded into 1D sequences along the horizontal or vertical direction through the Mamba model, and spatial long-range context features are extracted to obtain intermediate quantities : ; Among them, represents a layer normalization operation; represents the proportional parameter of the residual connection; represents the light field spatial selective scanning module; Respectively through layer normalization convolution channel attention and a residual connection with a scale parameter of to obtain an intermediate quantity :​ ; Reshape the to obtain the first result .

5. The method for semantic segmentation of light fields in complex urban scenes based on the state space model according to claim 4, wherein Obtaining macro-pixel slices according to the first result, and performing two-way scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles includes: Reshape the first result to obtain macro-pixel slices ; The macro-pixel slices are bidirectionally scanned through the Mamba model and the complementary relationship features between angles are modeled to obtain an intermediate quantity : ; Among them, represents the optical field angle selective scanning module; By layer normalization , convolution , channel attention and a residual connection with a scaling parameter of , an intermediate quantity is obtained: ; Reshape the to obtain the output feature data in 4D representation .

6. The method for semantic segmentation of light fields in complex urban scenes based on the state space model according to claim 5, characterized in that, The enhancing the interaction of the feature data from the first branch, second branch, and third branch through the Mamba model to obtain the first data, second data, and third data respectively includes: ; ; ; ; Among them, , represents the central view; represents the depth map; represents the 4D light field view; represents the discretized system transfer matrix; represents the exponential function; represents the adjustment parameter of the time step; represents the system state transition matrix; represents the system projection matrix; represents the discretized system projection matrix; represents the time step t of the hidden state; represents the time step t -1 of the hidden state; represents the time step t at the input; represents the time step t at the output of the central view; represents the cross-fusion system matrix; represents the time step t at the hidden state of the central view; represents the time step t at the selective scan output of the depth map branch of the state; represents the central view system matrix; represents the time step t at the output of the depth view; represents the time step t at the selective scan output of the light field image branch of the state; represents the time step t at the hidden state of the depth map; represents the output; represents the time step t at the selective scan output of the state; ; Among them, and are proportional parameters; represents the depth map system matrix at time step t; represents the 4D light field image system matrix at time step t.

7. The method for optical field semantic segmentation in complex urban scenes based on the state space model according to claim 6, wherein Splicing the first data, second data, and third data and then performing reverse scanning to obtain the fused enhanced feature includes: ; ; ; ; ; Among them, represents image features; represents depthwise separable convolution; represents a linear projection layer; represents the concatenated image sequence; represents the concatenation operation; represents the central view image features; represents the depth view image features; represents the light field view image features; represents dimension; represents the image feature sequence after inverse scanning; represents the inverse scanning operation; represents the Mamba model; represents the enhanced image features output; represents the feature separation operation; represents the concatenated image feature sequence output after selective scanning by the Mamba model; represents the inverse scanned image feature sequence output after selective scanning by the Mamba model; Obtain the scanned feature and After that, along the channel dimension, and are concatenated to obtain a cross-modal and cross-view fusion enhanced first feature with a shape of ; represents a matrix with a height of , a width of , and the number of channels is 3× C k ; Adjust the shape of the fused and enhanced first feature through a linear projection layer to obtain the fused and enhanced feature.

8. The method for semantic segmentation of light field in complex urban scenes based on the state space model according to claim 7, characterized in that, Decoding according to the fused data set to obtain the light field semantic segmentation map of the target scene includes: Decoding according to the fused enhanced features at each scale in the fused data set in combination with a channel-aware visual state space decoder, including: Suppose the fusion dataset includes Q fusion enhancement features at Q scales, that is, it includes Q fusion enhancement features, where the scales of the first to the Q th fusion enhancement features decrease; Enhance and upsample the Q th fusion-enhanced feature, restore the scale to that of the Q -1th fusion-enhanced feature, and splice it with the Q -1th fusion-enhanced feature and then input it into the channel-aware visual state space decoder for decoding. The channel-aware visual state space decoder outputs the first decoded data; Upsample the first decoded data to restore the scale to the scale of the Q -2 fused enhanced features, and splice them with the Q -2 fused enhanced features and then input them into the channel-aware visual state space decoder for decoding. The channel-aware visual state space decoder outputs the second decoded data; Continue to perform upsampling, feature splicing, and decoding based on the second decoded data until the scale of the decoded data is restored to the scale of the first fused enhanced feature, and the channel-aware visual state space decoder outputs the decoded completed data; After decoding is completed, upsample the decoded completed data to restore the scale to the scale of the central view to obtain a decoded image; input the decoded image into a Softmax classifier, and the Softmax classifier outputs the light field semantic segmentation map of the target scene.

9. A complex urban scene light field semantic segmentation system based on a state space model, for the method according to any one of claims 1 to 8, characterized in that, The system includes a first module, a second module, a third module, and a fourth module; The first module is used to obtain the central view, depth map, and 4D light field image of the target scene; The second module is used to preprocess the central view, depth map, and 4D light field image and then input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction; the feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; The first branch, second branch, and third branch all output feature data at a preset multi-scale; The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result; obtaining macro-pixel slices from the first result, and performing two-way scanning and modeling of the complementary relationship features between angles on the macro-pixel slices through the Mamba model; The third module is used to fuse the feature data at each scale in the outputs of the first branch, second branch, and third branch based on the second strategy to obtain a fused data set; The second strategy includes: enhancing the interaction of the feature data from the first branch, second branch, and third branch through the Mamba model to obtain the first data, second data, and third data respectively, splicing the first data, second data, and third data and then performing reverse scanning to obtain a fused enhanced feature; The fourth module is used to decode the fused data set to obtain the light field semantic segmentation map of the target scene.

Citation Information

Patent Citations

  • Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model

    CN119152205A

  • Video sequence segmentation method based on selective scanning visual state space model

    CN119206568A