Complex city scene light field semantic segmentation method and system based on state space model

By applying the light field semantic segmentation method based on state space model in complex urban market scenarios, the problems of high computational complexity, insufficient long-range dependency modeling and difficulty in fusion of multimodal information are solved, and efficient and accurate light field semantic segmentation effect is achieved.

CN120047690AActive Publication Date: 2025-05-27HUNAN UNIV

Patent Information

Application Number
CN202510535732.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-27
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the problems of high computational complexity, insufficient long-range dependency modeling and difficulty in multimodal information fusion in complex urban market scenarios, resulting in limited semantic segmentation accuracy of the light field.

Method used

Using a state space model-based method, multi-scale feature extraction and fusion are performed by acquiring the central view, depth map and 4D light field images. The specific steps include: extracting 2D sub-aperture slices and macropixel slices based on the 4D light field image, using the Mamba model for spatial long-range context feature extraction and angular complementary relationship modeling; interactive enhancement and reverse scanning of feature data from different branches, realizing cross-modal and cross-view fusion.

Benefits of technology

It has realized efficient and accurate solutions for semantic segmentation of light fields in complex urban markets, overcomes the problems of high computational complexity, insufficient long-range dependency modeling and difficulty in fusion of multimodal information, and significantly improves segmentation accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047690A_ABST
    Figure CN120047690A_ABST
Patent Text Reader

Abstract

The invention discloses a complex city scene light field semantic segmentation method and system based on a state space model. The method comprises the steps that a center view, a depth map and a 4D light field image of a target scene are preprocessed and then input into a first branch, a second branch and a third branch for multi-scale feature extraction; feature extraction of the third branch is based on a first strategy; fusing the feature data of each scale in the three branch outputs based on a second strategy to obtain a fused data set; obtaining a light field semantic segmentation map of the target scene according to the fused data set; the first strategy comprises the following steps: carrying out space long-range context feature extraction through a Mangbar model according to the 4D light field image to obtain a first result; according to the first result, bidirectional scanning is carried out through a Mangbar model, and complementary relation characteristics between angles are modeled; the second strategy comprises the steps of performing interaction enhancement on the feature data from the three branches through the Mangbar model, splicing the feature data and performing reverse scanning to obtain fusion enhancement features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and image processing, and particularly to a method and system for semantic segmentation of light fields in complex urban scenes based on a state space model. Background Art

[0002] Semantic segmentation is a core task in computer vision, aiming to assign semantic labels to each pixel in an image, and is an important basis for autonomous driving, smart cities, and robot perception. In complex urban scenes, due to factors such as the dense distribution of targets such as buildings, vehicles, and pedestrians, as well as dynamic occlusion and lighting changes, traditional monocular image segmentation methods are difficult to meet the requirements.

[0003] 4D light fields form a regularly arranged dense viewpoint array and macro-pixel array by recording the direction and position information of light rays, containing rich spatial-angular information, and can effectively solve the problems of occlusion and lighting interference in complex urban scenes. However, existing methods for light field semantic segmentation face problems such as high computational complexity, insufficient long-range dependence modeling, and difficulty in fusing dynamic information. In terms of high computational complexity, methods based on Transformer rely on self-attention mechanisms, and the computational complexity grows quadratically, making it difficult to efficiently process high-dimensional light field data; in terms of insufficient long-range dependence modeling, traditional convolutional neural networks are limited by local receptive fields and cannot fully extract the multi-dimensional spatial-angular information embedded in light fields; in terms of difficulty in fusing dynamic information, existing methods are difficult to effectively fuse the multi-view, macro-pixel, and multi-modal (such as depth information) features of light fields, resulting in limited segmentation accuracy in complex urban scenes. For example, some researchers have proposed a multi-view fusion method based on Transformer, but it cannot handle a large amount of viewpoint information at the same time; other researchers have used CNNs to extract sub-aperture image features, but have insufficient modeling ability for global spatial-angular relationships and are difficult to handle dynamic occlusion and lighting changes in urban complex scenes.

[0004] In recent years, SSM (State Space Model) as an efficient sequence modeling method has gradually shown great potential in the fields of natural language processing and computer vision. SSM models long-range dependence relationships with linear complexity and can effectively process global information in high-dimensional data. As a new architecture based on SSM, the Mamba model introduces a selective scanning mechanism that can dynamically adjust model parameters according to the input, significantly improving the model's performance in long-sequence tasks and having high efficiency and robustness when processing high-dimensional data. However, existing research mainly focuses on monocular images or low-dimensional data, and the Mamba model has not been applied to the semantic segmentation task of 4D light field data.

[0005] Therefore, there is an urgent need for a new technical solution to solve the technical problem of how to overcome the existing defects such as high computational complexity, insufficient long-range dependence modeling, and difficult multi-modal information fusion to perform light field semantic segmentation on complex urban scenes. Summary of the Invention

[0006] The present invention provides a method and system for light field semantic segmentation of complex urban scenes based on a state space model to solve the technical problem of how to overcome the existing defects such as high computational complexity, insufficient long-range dependence modeling, and difficult multi-modal information fusion to perform light field semantic segmentation on complex urban scenes.

[0007] To achieve the above object, the present invention provides a method for light field semantic segmentation of complex urban scenes based on a state space model, including: Obtain the central view, depth map, and 4D light field image of the target scene; preprocess the central view, depth map, and 4D light field image and then input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction.

[0008] The feature extraction of the first branch and the second branch is based on a visual state space block, and the feature extraction of the third branch is based on a first strategy; the first branch, the second branch, and the third branch all output feature data at a preset multi-scale.

[0009] Fuse the feature data at each scale in the outputs of the first branch, the second branch, and the third branch based on a second strategy to obtain a fused data set; decode the fused data set to obtain the light field semantic segmentation map of the target scene.

[0010] The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, extracting spatial long-range context features of the 2D sub-aperture slices through a Mamba model to obtain a first result; obtaining macro-pixel slices from the first result, and performing bidirectional scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles.

[0011] The second strategy includes: enhancing the interaction of the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively, splicing the first data, the second data, and the third data and then performing reverse scanning to obtain the fused enhanced features.

[0012] Preferably, obtaining the central view, depth map, and 4D light field image of the target scene includes: Obtain the 4D light field image .

[0013] According to the 4D light field image Obtain the central view .

[0014] According to the central view Obtain the depth map .

[0015] Among them, Indicates that the light field parameter is U And V , with a height of H , and a width of W Matrix; Indicates a matrix with a height of H , a width of W , and 3 channels; Indicates a matrix with a height of H , a width of W Matrix.

[0016] Preferably, the preprocessing includes: Normalize the central view, depth map, and 4D light field image: The normalization process of the 4D light field image includes: ; Among them, Indicates the 4D light field image after normalization; Indicates the 4D light field image; Indicates the mean value of the 4D light field image; Indicates the standard deviation of the 4D light field image.

[0017] The normalization process of the central view and depth map includes: ; Among them, Indicates the central view or depth map after normalization; Indicates the central view or depth map; Indicates the mean value of the central view or depth map; Indicates the standard deviation of the central view or depth map.

[0018] Perform data augmentation on the normalized central view, depth map, and 4D light field image using random rotation and color jitter; Segment the normalized central view, depth map, and 4D light field image into local blocks through an image patch module in the ViT style and map them to a high-dimensional feature space, including: ; Among them, Indicates the image data output by the image patch module, , Indicates a height of , a width of , and the number of channels is matrix; ; represents the scale level; represents the cropping module.

[0019] Preferably, a 2D sub-aperture slice is obtained from the 4D light field image, and the 2D sub-aperture slice is subjected to spatial long-range context feature extraction through the Mamba model, and the first result includes: The 4D light field image is decoupled into 2D sub-aperture slices through a Reshape operation .

[0020] The 2D sub-aperture slice is expanded into a 1D sequence along the horizontal or vertical direction through the Mamba model and the spatial long-range context features are extracted to obtain an intermediate quantity : ; wherein, represents a layer normalization operation; represents the proportional parameter of the residual connection; represents the light field spatial selective scanning module; will respectively pass through layer normalization , convolution , channel attention and the residual connection with a proportional parameter of to obtain an intermediate quantity : ; Through a Reshape operation, is reshaped to obtain the first result .

[0021] Preferably, a macro-pixel slice is obtained from the first result, and the macro-pixel slice is subjected to bidirectional scanning through the Mamba model and the complementary relationship features between angles are modeled, including: Through a Reshape operation, the first result is reshaped to obtain a macro-pixel slice .

[0022] The macro-pixel slice is subjected to bidirectional scanning through the Mamba model and the complementary relationship features between angles are modeled to obtain an intermediate quantity : ; wherein, represents the light field angle selective scanning module; will respectively pass through layer normalization , convolution , channel attention and the proportional parameter is The residual connection is used to obtain an intermediate quantity : ; Through the Reshape operation, is reshaped to obtain the output feature data in 4D representation .

[0023] Preferably, the feature data from the first branch, the second branch, and the third branch are interactively enhanced through the Mamba model to respectively obtain the first data, the second data, and the third data, including: ; ; ; ; Among them, , represents the central view; represents the depth map; represents the 4D light field view; represents the discretized system transition matrix; represents the exponential function; represents the adjustment parameter of the time step; represents the system state transition matrix; represents the system projection matrix; represents the discretized system projection matrix; represents the time step t of the hidden state; represents the time step t -1 of the hidden state; represents the time step t at the input; represents the time step t at the output of the central view; represents the cross-fusion system matrix; represents the time step t at the hidden state of the central view; represents the time step t at the selective scan output of the depth map branch of the state; represents the central view system matrix; represents the time step t at the output of the depth view; represents the time step t at the selective scan output of the light field image branch of the state; represents the time step t at the hidden state of the depth map; Indicates the output; Indicates the time step t Selective scan output of the state at...

[0024] ; Wherein, and Are proportionality parameters; Indicates the depth map system matrix at time step t; Indicates the 4D light field image system matrix at time step t.

[0025] Preferably, after splicing the first data, the second data and the third data and performing reverse scanning, the fused enhanced features obtained include: ; ; ; ; ; Wherein, Indicates the image feature; Indicates the depthwise separable convolution; Indicates the linear projection layer; Indicates the spliced image sequence; Indicates the splicing operation; Indicates the central view image feature; Indicates the depth view image feature; Indicates the light field view image feature; Indicates the dimension; Indicates the image feature sequence after inverse scanning; Indicates the inverse scanning operation; Indicates the Mamba model; Indicates the enhanced image feature of the output; Indicates the feature separation operation; Indicates the spliced image feature sequence output by selective scanning processing of the Mamba model; Indicates the inverse scanned image feature sequence output by selective scanning processing of the Mamba model.

[0026] Obtain the scanned feature and After that, along the channel dimension, and Are connected to obtain the cross-modal and cross-view fused enhanced first feature with the shape of ; Indicates the height is , and the width is , the number of channels is 3× C k matrix.

[0027] The shape of the fused and enhanced first feature is adjusted to through a linear projection layer to obtain the fused and enhanced feature.

[0028] Preferably, decoding according to the fused dataset to obtain the light field semantic segmentation map of the target scene includes: Decoding the fused and enhanced features at each scale in the fused dataset in combination with a channel-aware visual state space decoder, including: Assume that the fused dataset includes Q fused and enhanced features at Q scales, that is, including Q fused and enhanced features, where the scales of the 1st to the Q th fused and enhanced features decrease.

[0029] Enhance and upsample the Q th fused and enhanced feature, restore the scale to the scale of the Q -1th fused and enhanced feature, and splice it with the Q -1th fused and enhanced feature and then input it into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs the first decoded data.

[0030] Upsample the first decoded data, restore the scale to the scale of the Q -2th fused and enhanced feature, and splice it with the Q -2th fused and enhanced feature and then input it into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs the second decoded data.

[0031] Continue to upsample, feature splice and decode according to the second decoded data until the scale of the decoded data is restored to the scale of the 1st fused and enhanced feature, and the channel-aware visual state space decoder outputs the decoded completed data.

[0032] After decoding is completed, upsample the decoded completed data, restore the scale to the scale of the central view to obtain the decoded image; input the decoded image into the Softmax classifier, and the Softmax classifier outputs the light field semantic segmentation map of the target scene.

[0033] The present invention also provides a light field semantic segmentation system for complex urban scenes based on a state space model for the method of the present invention. The system includes a first module, a second module, a third module and a fourth module.

[0034] The first module is used to obtain the central view, depth map and 4D light field image of the target scene.

[0035] The second module is used to preprocess the central view, depth map, and 4D light field image and then input them into the first branch, second branch, and third branch respectively for multi-scale feature extraction. The feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy. The first branch, the second branch, and the third branch all output feature data at a preset multi-scale.

[0036] The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result; obtaining macro-pixel slices based on the first result, and performing two-way scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles.

[0037] The third module is used to fuse the feature data at each scale in the outputs of the first branch, the second branch, and the third branch based on the second strategy to obtain a fused data set.

[0038] The second strategy includes: enhancing the interaction of the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively, splicing the first data, the second data, and the third data and then performing reverse scanning to obtain fused enhanced features.

[0039] The fourth module is used to decode based on the fused data set to obtain the light field semantic segmentation map of the target scene.

[0040] The present invention has the following beneficial effects: The method for light field semantic segmentation of complex urban scenes based on the state space model of the present invention can comprehensively extract multi-scale long-distance spatial angle information by performing multi-scale feature extraction on the 4D light field image based on the first strategy. The extracted features are integrated with the multi-scale modal information of the other two branches based on the second strategy, which promotes feature fusion and the enhancement of the central view, and realizes the adaptive fusion of light field multi-modal information. By decoupling the spatial and angular information of the light field and performing independent modeling respectively, the limitations of traditional methods in long-range dependence modeling are solved. The method of the present invention provides an efficient and accurate solution for dynamic occlusion, illumination change, and multi-scale target segmentation in complex urban scenes. The method of the present invention overcomes the existing defects such as high computational complexity, insufficient long-range dependence modeling, and difficulty in multi-modal information fusion, and can perform light field semantic segmentation on complex urban scenes.

[0041] The system for light field semantic segmentation of complex urban scenes based on the state space model of the present invention, which is used for the method of the present invention, has the same beneficial effects as the method of the present invention.

[0042] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. Description of the Drawings

[0043] The drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and the description thereof are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings: Figure 1 is a schematic flowchart of the method of the preferred embodiment of the present invention.

[0044] Figure 2 is a schematic flowchart of the light field semantic segmentation process of the preferred embodiment of the present invention.

[0045] Figure 3 is a schematic diagram of the first strategy of the preferred embodiment of the present invention.

[0046] Figure 4 is a schematic flowchart of the cross-selective scanning based on the cross-selective scanning module of the preferred embodiment of the present invention.

[0047] Figure 5 is a schematic flowchart of the feature scanning based on the connection selection scanning module of the preferred embodiment of the present invention. Detailed Embodiments

[0048] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.

[0049] See Figures 1 to 2 , in the preferred embodiment of the present invention, a method for light field semantic segmentation of complex urban scenes based on a state space model is provided, including: P1. Obtain the central view, depth map and 4D light field image of the target scene. P1 specifically includes: Obtain the 4D light field image .

[0050] According to the 4D light field image obtain the central view .

[0051] According to the central view obtain the depth map .

[0052] Among them, represents a matrix with light field parameters of U and V , height of H , and width of W ; represents a height ofH , with a width of W , a matrix with 3 channels; represents a matrix with a height of H , and a width of W .

[0053] P2. After preprocessing the central view, depth map, and 4D light field image, they are respectively input into the first branch, second branch, and third branch for multi-scale feature extraction, where the feature extraction of the third branch is based on the first strategy.

[0054] In the preferred embodiment of the present invention, the preprocessing includes: Normalize the central view, depth map, and 4D light field image: The normalization process of the 4D light field image includes: ; Among them, represents the normalized 4D light field image; represents the 4D light field image; represents the mean value of the 4D light field image; represents the standard deviation of the 4D light field image.

[0055] The normalization process of the central view and depth map includes: ; Among them, represents the normalized central view or depth map; represents the central view or depth map; represents the mean value of the central view or depth map; represents the standard deviation of the central view or depth map.

[0056] Perform data augmentation on the normalized central view, depth map, and 4D light field image using random rotation (±15°) and color jitter (brightness, contrast ±10%).

[0057] Segment the normalized and data-augmented central view, depth map, and 4D light field image into local blocks through an image patch module in the ViT style and map them to a high-dimensional feature space, including: ; Among them, represents the image data output by the image patch module, , represents a matrix with a height of , a width of , and the number of channels is C k ; ; Indicates the scale level; Indicates the cropping module.

[0058] In the preferred embodiment of the present invention, when performing multi-scale feature extraction, it is preferred to perform feature extraction at 4 levels of scales, that is, the preset multi-scale includes 4 levels of scales. Therefore , .

[0059] In the preferred embodiment of the present invention, the feature extraction of the first branch and the second branch is based on the Visual State Space (VSS), and the feature extraction of the third branch is based on the first strategy; the first branch, the second branch, and the third branch all output feature data at the preset multi-scale.

[0060] In the preferred embodiment of the present invention, in the Visual State Space block, first, the input features are respectively processed by layer normalization, linear projection, and depthwise convolution to extract local information and maintain the spatial structure. Subsequently, the Selective Scan 2D (SS2D) module is introduced to model the long-range spatial information in the features. Specifically, the SS2D module flattens the input feature with a shape of into 4 sequences along four directions (from top left to bottom right, from bottom right to top left, from top right to bottom left, from bottom left to top right), and extracts multi-directional long-range dependencies through four different selective scan modules respectively. Finally, these sequences are reshaped back to the original shape and summed to integrate multi-directional information, and finally the feature representation is enhanced through residual connection, thus completing the effective modeling of the input data.

[0061] In the preferred embodiment of the present invention, the first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, extracting spatial long-range context features from the 2D sub-aperture slices through the Mamba model to obtain the first result; obtaining macro-pixel slices from the first result, and performing bidirectional scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles.

[0062] In the preferred embodiment of the present invention, the first strategy dynamically adjusts the SSM parameters, preferentially selects the area near the target point for sampling, significantly improves the path search efficiency, and solves the deficiencies of traditional methods in long-range dependence modeling.

[0063] See Figure 3 , in the preferred embodiment of the present invention, obtaining 2D sub-aperture slices from the 4D light field image, and extracting spatial long-range context features from the 2D sub-aperture slices through the Mamba model, the first result obtained includes: Decouple the 4D light field image into 2D sub-aperture slices through the Reshape operation 。

[0064] Unfold the 2D sub-aperture slices into 1D sequences along the horizontal or vertical direction through the Mamba model and extract the spatial long-range context features to obtain an intermediate quantity : ; Among them, represents the layer normalization operation; represents the proportional parameter of the residual connection; represents the light field spatial selective scanning module, that is, Figure 3 the LFSpaSS module in

[0065] will be respectively passed through layer normalization , convolution , channel attention and a residual connection with a proportional parameter of to obtain an intermediate quantity : ; Reshape through the Reshape operation to obtain the first result 。

[0066] In the preferred embodiment of the present invention, the macro-pixel slices are obtained according to the first result, and the bidirectional scanning of the macro-pixel slices through the Mamba model and the modeling of the complementary relationship features between angles include: Reshape the first result through the Reshape operation to obtain the macro-pixel slices 。

[0067] Perform bidirectional scanning on the macro-pixel slices through the Mamba model and model the complementary relationship features between angles to obtain an intermediate quantity : ; Among them, represents the light field angular selective scanning module, that is, Figure 3 the LFAngSS module in

[0068] will be respectively passed through layer normalization , convolution , channel attention and a residual connection with a proportional parameter of to obtain an intermediate quantity : ; Reshape the to obtain the output feature data in 4D representation .

[0069] P3. For each scale of the feature data in the outputs of the first branch, the second branch, and the third branch, fuse them based on the second strategy to obtain a fused dataset; decode according to the fused dataset to obtain the light field semantic segmentation map of the target scene.

[0070] See Figures 4 to 5 . In a preferred embodiment of the present invention, the second strategy includes: enhancing the interaction of the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively, and splicing the first data, the second data, and the third data and then performing a reverse scan to obtain the fused enhanced features.

[0071] In a preferred embodiment of the present invention, enhancing the interaction of the feature data from the first branch, the second branch, and the third branch through the Mamba model to obtain the first data, the second data, and the third data respectively includes: See Figure 4 . Take the feature data from the first branch, the second branch, and the third branch as inputs, and perform cross-selective scanning based on the Cross SS module to enhance the feature interaction between the central view, the depth map, and the 4D light field image: ; ; ; ; where , represents the central view; represents the depth map; represents the 4D light field view; represents the discretized system transition matrix; represents the exponential function; represents the adjustment parameter of the time step; represents the system state transition matrix; represents the system projection matrix; represents the discretized system projection matrix; represents the hidden state at time step t; represents the hidden state at time step t - 1; represents the time step t at the input; Represents the time step t Output of the central view at; Represents the cross-fusion system matrix; Represents the time step t Hidden state of the central view at; Represents the time step t Selective scan output of the depth map branch of the state at; Represents the central view system matrix; Represents the time step t Output of the depth view at; Represents the time step t Selective scan output of the light field image branch of the state at; Represents the time step t Hidden state of the depth map at; Represents the output; Represents the time step t Selective scan output of the state at.

[0072] ; Among them, and Are scale parameters; Represents the depth map system matrix at time step t; Represents the 4D light field image system matrix at time step t.

[0073] In the preferred embodiment of the present invention, after splicing the first data, the second data, and the third data and performing reverse scanning, the fusion enhanced features obtained include: See Figure 5 , taking the first data, the second data, and the third data as inputs, and performing feature scanning based on the connection selection scanning module Concat SS to optimize the feature fusion effect: ; ; ; ; ; Among them, Represents the image feature; Represents the depthwise separable convolution; Represents the linear projection layer; Represents the spliced image sequence; Represents the splicing operation; Represents the central view image feature; Represents the depth view image feature; Represents the light field view image feature; Represents the dimension; Represents the image feature sequence after inverse scanning; Represents the inverse scanning operation; Represents the Mamba model; Represents the enhanced image feature output; Represents the feature separation operation; Represents the spliced image feature sequence output after selective scanning by the Mamba model; Represents the inverse scanned image feature sequence output after selective scanning by the Mamba model.

[0074] Obtain the scanning feature and After that, along the channel dimension, and Are concatenated to obtain the cross-modal and cross-view fusion enhanced first feature with the shape of ; Represents the height of , the width of , and the number of channels is 3× C k Matrix.

[0075] Adjust the shape of the fusion enhanced first feature to through the linear projection layer to obtain the fusion enhanced feature.

[0076] In the preferred embodiment of the present invention, the light field semantic segmentation map of the target scene obtained by decoding according to the fusion dataset includes: Decoding according to the fusion enhanced features at each scale in the fusion dataset in combination with the Channel-Aware Visual State Space (CAVSS) decoder, including: Assume that the fusion dataset includes Q Fusion enhanced features at scales, that is, including Q Fusion enhanced features, where the scales of the 1st to the Q th fusion enhanced features decrease.

[0077] Enhance and upsample the Q th fusion enhanced feature, restore the scale to the scale of the Q -1th fusion enhanced feature, and splice it with the Q -1th fusion enhanced feature and then input it into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs the first decoded data.

[0078] Upsample the first decoded data to restore the scale to the Q- The scales of two fused and enhanced features, and are concatenated with the Q - two fused and enhanced features and then input into the channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs second decoded data.

[0079] Based on the second decoded data, continue to perform upsampling, feature concatenation and decoding until the scale of the decoded data is restored to the scale of the first fused and enhanced feature, and the channel-aware visual state space decoder outputs the decoded completed data.

[0080] After decoding is completed, upsample the decoded completed data to restore the scale to the scale of the central view to obtain a decoded image; input the decoded image into a Softmax classifier, and the Softmax classifier outputs a light field semantic segmentation map of the target scene.

[0081] In the preferred embodiment of the present invention, in the channel-aware visual state space decoder, the input features are first passed through a visual state space block to extract features; then the features are normalized through layer normalization; and then through convolution operations, local feature information is further extracted. The processed features are respectively input into maxpooling and avgpooling, and multi-scale features are generated through different pooling methods; then using a skip connection mechanism, features of different scales are fused to improve the multi-scale feature fusion effect.

[0082] In the channel-aware visual state space decoder, upsampling is performed on the multi-scale features in combination with the channel attention mechanism to enhance the expression ability of spatial-channel information. The CAVSS module restores the high-resolution segmentation result through a selective scanning mechanism, and optimizes the segmentation accuracy by combining the weighted combination of cross entropy and Dice loss. Based on the channel-aware visual state space decoder for data decoding, it can significantly improve the recovery ability of detail features in complex urban scenes and solve the problem of segmentation accuracy fluctuations in dynamic occlusion and illumination change scenes.

[0083] The method for semantic segmentation of light field in complex urban scenes based on the state space model of the present invention can comprehensively extract multi-scale long-distance spatial angular information by performing multi-scale feature extraction on 4D light field images based on the first strategy; the extracted features are integrated with multi-scale modal information of the other two branches based on the second strategy, which promotes feature fusion and enhancement of the central view, and realizes the adaptive fusion of multi-modal information of the light field. By decoupling the spatial and angular information of the light field and performing independent modeling respectively, the limitations of traditional methods in long-range dependence modeling are solved; the method of the present invention provides an efficient and accurate solution for dynamic occlusion, illumination change and multi-scale target segmentation in complex urban scenes. The method of the present invention overcomes the existing defects such as high computational complexity, insufficient long-range dependence modeling and difficulty in multi-modal information fusion, and can perform semantic segmentation of light field on complex urban scenes.

[0084] In a preferred embodiment of the present invention, a system for semantic segmentation of light field in complex urban scenes based on the state space model is also provided for the method of the present invention. The system includes a first module, a second module, a third module and a fourth module.

[0085] The first module is used to obtain the central view, depth map and 4D light field image of the target scene.

[0086] The second module is used to preprocess the central view, depth map and 4D light field image and then input them into the first branch, the second branch and the third branch respectively for multi-scale feature extraction; the feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; the first branch, the second branch and the third branch all output feature data at a preset multi-scale.

[0087] The first strategy includes: obtaining 2D sub-aperture slices from the 4D light field image, performing spatial long-range context feature extraction on the 2D sub-aperture slices through the Mamba model to obtain a first result; obtaining macro-pixel slices according to the first result, and performing bidirectional scanning on the macro-pixel slices through the Mamba model and modeling the complementary relationship features between angles.

[0088] The third module is used to fuse the feature data at each scale in the outputs of the first branch, the second branch and the third branch based on the second strategy to obtain a fused data set.

[0089] The second strategy includes: enhancing the interaction of the feature data from the first branch, the second branch and the third branch through the Mamba model to obtain the first data, the second data and the third data respectively, splicing the first data, the second data and the third data and then performing reverse scanning to obtain fused enhanced features.

[0090] The fourth module is used to decode according to the fused data set to obtain the light field semantic segmentation map of the target scene.

[0091] The light field semantic segmentation system for complex urban scenes based on the state space model of the present invention, which is used for the method of the present invention, has the same beneficial effects as the method of the present invention.

[0092] Verification part: The effectiveness of the method of the present invention is verified on the UrbanLF dataset. The UrbanLF dataset is the only available light field semantic segmentation benchmark to date. Each scene contains 81 regularly arranged sub-aperture images with a 9×9 angular resolution. The images in the dataset depict urban road scenes, including complex occlusion and lighting conditions. The statistical data of the dataset used in the verification are shown in Table 1, which details the number of categories, spatial resolution, angular resolution, and training set division method. Among them, UrbanLF-Real contains real-world data, which is characterized by greater noise and natural light variations compared to the synthetic UrbanLF-Syn dataset.

[0093] Table 1 Statistical data table of the dataset , The method of the present invention is implemented using PyTorch in the Ubuntu environment. Training and inference are performed on two RTX 4090 GPUs. During training, the AdamW optimizer with an initial learning rate of 6e-5 and a weight decay of 0.01 is used. The model is trained for 500 epochs with a batch size of 2. In addition, it is compared with several state-of-the-art methods. The mainstream models based on LF and RGBD and the models based on the method of the present invention adopted in the preferred embodiment of the present invention are shown in Table 2. Among them, LFSSMam(S), LFSSMam(M), and LFSSMam(L) are models based on the method of the present invention, and S, M, and L represent the sizes of the models. The size of the model is determined by the parameters of the pre-trained VSS module in the model. The quantitative experimental data of all selected models on the UrbanLF-Syn dataset and the evaluation results of the mIoU, mAcc, and Acc metrics are shown in Tables 3 and 4, and the quantitative experimental data of all selected models on the UrbanLF-Real dataset and the evaluation results of the mIoU, mAcc, and Acc metrics are shown in Tables 5 and 6.

[0094] Table 2 Table of mainstream models based on LF and RGBD and models based on the method of the present invention , Table 2 Quantitative experimental data table of the models on the UrbanLF-Syn dataset , Table 4 Evaluation result table of the mIoU, mAcc, and Acc metrics of the models on the UrbanLF-Syn dataset , Table 5 Quantitative experimental data table of the model on the UrbanLF-Real dataset , Table 6 Evaluation results table of mIoU, mAcc and Acc metrics of the model on the UrbanLF-Real dataset , Analysis based on Tables 2 to 6 shows that: (1) Precision comparison: The mIoU index of LFSSMam-L reaches 89.66% (UrbanLF-Syn) and 86.54% (UrbanLF-Real), which is 6.97% better than the current best method.

[0095] (2) Robustness verification: In complex urban scenes, in the face of dynamic occlusion and light changes, the segmentation precision fluctuation is less than 2%, which is significantly better than traditional methods.

[0096] (3) Visual qualitative verification: According to the experimental data of this experiment, corresponding visualization results can be obtained. From the visualization results, it can be seen that for objects with complex and delicate structures, such as bicycles and big trees, etc., the method of the present invention can generate more coherent semantic instances.

[0097] The method of the present invention realizes the efficient semantic segmentation of light field images in complex urban scenes through a selective state space model and a multi-modal fusion strategy, providing an innovative solution for scene understanding in fields such as smart cities and autonomous driving. Experiments show that this method reaches the leading level in the industry in terms of precision, efficiency and robustness.

[0098] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A light field semantic segmentation method for complex urban scenes based on a state space model, characterized in that: include: Obtain the central view, depth map and 4D light field image of the target scene; The central view, the depth map and the 4D light field image are pre-processed and input into the first branch, the second branch and the third branch respectively for multi-scale feature extraction; The feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; the first branch, the second branch and the third branch all output feature data at preset multi-scales; The feature data at each scale in the outputs of the first branch, the second branch, and the third branch are fused based on the second strategy to obtain a fused data set; Decoding the fused data set to obtain a light field semantic segmentation map of the target scene; The first strategy includes: obtaining a 2D sub-aperture slice according to the 4D light field image, and extracting spatial long-range context features of the 2D sub-aperture slice through a Mamba model to obtain a first result; obtaining a macro-pixel slice according to the first result, and bidirectionally scanning the macro-pixel slice through a Mamba model to model complementary relationship features between angles; The second strategy includes: interactively enhancing the feature data from the first branch, the second branch and the third branch through the Mamba model to obtain first data, second data and third data respectively, splicing the first data, the second data and the third data and performing reverse scanning to obtain fused enhanced features.

2. The method for semantic segmentation of complex urban scene light field based on state space model according to claim 1, characterized in that: The obtaining of the central view, the depth map and the 4D light field image of the target scene comprises: Acquire 4D light field images ; According to the 4D light field image Get the center view ; According to the central view Get the depth map ; in, The light field parameters are U and V , Gao Wei H , width is W Matrix of Indicates high H , width is W , a matrix with 3 channels; Indicates high H , width is W The matrix of .

3. The complex urban scene light field semantic segmentation method based on the state space model according to claim 2 is characterized in that: The pre-processing comprises: The center view, depth map and 4D light field image are normalized: The standardized processing of the 4D light field image includes: ; in, Represents the 4D light field image after normalization; Represents a 4D light field image; represents the mean of the 4D light field image; represents the standard deviation of the 4D light field image; The normalization process of the center view and the depth map includes: ; in, represents the normalized center view or depth map; represents the center view or depth map; represents the mean of the central view or depth map; represents the standard deviation of the center view or depth map; The normalized center view, depth map, and 4D light field image are augmented by random rotation and color jitter. The center view, depth map, and 4D light field image after the standardization and data enhancement are segmented into local blocks through a ViT-style image segmentation module and mapped to a high-dimensional feature space, including: ; in, Represents the image data output by the image segmentation module, , Indicates high , width is , the number of channels is Matrix of ; Indicates the scale level; Represents a cropping module.

4. The method for semantic segmentation of complex urban scene light field based on state space model according to claim 3, characterized in that: Obtaining a 2D sub-aperture slice according to the 4D light field image, and performing spatial long-range context feature extraction on the 2D sub-aperture slice through a Mamba model, to obtain a first result including: The 4D light field image is decoupled into 2D sub-aperture slices through the Reshape operation ; The 2D sub-aperture slices are expanded into 1D sequences along the horizontal or vertical direction through the Mamba model and the spatial long-range context features are extracted to obtain the intermediate quantity : ; in, Representation layer normalization operation; Represents the scale parameter of the residual connection; represents the light field spatial selective scanning module; Will By layer normalization ,convolution , channel attention and the scale parameter is The residual connection of : ; Through the Reshape operation Reshape to get the first result .

5. The complex urban scene light field semantic segmentation method based on the state space model according to claim 4 is characterized in that: A macro pixel slice is obtained according to the first result, and a bidirectional scanning is performed on the macro pixel slice through a Mamba model to model the complementary relationship characteristics between angles, including: The first result is transformed into Reshape to get macro pixel slices ; The macro pixel slices are scanned bidirectionally by the Mamba model and the complementary relationship characteristics between angles are modeled to obtain the intermediate quantity : ; in, represents a light field angle selective scanning module; Will By layer normalization ,convolution , channel attention and the scale parameter is The residual connection of : ; Through the Reshape operation Reshape to get the output feature data represented by 4D .

6. The complex urban scene light field semantic segmentation method based on the state space model according to claim 5 is characterized in that: The interactive enhancement of the feature data from the first branch, the second branch and the third branch by the Mamba model to obtain first data, second data and third data respectively includes: ; ; ; ; in, , Indicates the central view; represents a depth map; Represents a 4D light field view; represents the discretized system transfer matrix; represents the exponential function; represents the adjustment parameter of the time step; represents the system state transfer matrix; Represents the system projection matrix; represents the discretized system projection matrix; Indicates the time step t The hidden state of Indicates the time step t Hidden state of -1; Indicates the time step t Input at Indicates the time step t Output of the center view; represents the cross-fusion system matrix; Indicates the time step t The hidden state of the center view; Indicates the time step t Selective scan output of the depth map branch at state; Represents the central view system matrix; Indicates the time step t Output of the depth view at Indicates the time step t A selective scanning output of a light field image branch at a state; Indicates the time step t The hidden state of the depth map; Indicates output; Indicates the time step t Selective scan output of the state; ; in, and is the scale parameter; Represents the depth map system matrix at time step t; Represents the 4D light field image system matrix at time step t.

7. The complex urban scene light field semantic segmentation method based on the state space model according to claim 6 is characterized in that: The first data, the second data and the third data are spliced ​​and then reversely scanned to obtain fusion enhancement features including: ; ; ; ; ; in, Represents image features; represents depthwise separable convolution; Represents a linear projection layer; represents the spliced ​​image sequence; Represents a splicing operation; represents the central view image features; Represents the depth view image features; Representing light field view image features; Represents dimension; Represents the image feature sequence after inverse scanning; Indicates the inverse scan operation; represents the Mamba model; Represents the enhanced image features of the output; Represents feature separation operation; represents the feature sequence of the spliced ​​image output by the selective scanning process of the Mamba model; represents the inverse scan image feature sequence output by the Mamba model selective scanning process; Get scan features and Then, along the channel dimension and Connect them to get the shape The cross-modal and cross-view fusion enhances the first feature; Indicates high , width is , the number of channels is 3× C k Matrix of The shape of the fused enhanced first feature is adjusted to , and obtain the fusion enhancement feature.

8. The complex urban scene light field semantic segmentation method based on the state space model according to claim 7 is characterized in that: The light field semantic segmentation map of the target scene obtained by decoding the fused data set includes: Decoding is performed according to the fusion enhanced features at each scale in the fusion data set combined with a channel-aware visual state space decoder, including: Assume that the fused data set includes Q The fusion enhancement features at different scales include Q fusion enhancement features, of which the first to the Q The scale reduction of the fused enhanced features; The first Q The fusion enhancement features are enhanced and upsampled to restore the scale to the Q -1 scale of fused enhanced features and combined with Q -1 fused enhanced features are spliced ​​and input into a channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs first decoded data; The first decoded data is upsampled and the scale is restored to the Q -2 scales of fused enhanced features and combined with the Q -The two fused enhanced features are concatenated and input into a channel-aware visual state space decoder for decoding, and the channel-aware visual state space decoder outputs second decoded data; Continue upsampling, feature splicing and decoding according to the second decoded data until the scale of the decoded data is restored to the scale of the first fusion enhanced feature, and the channel-aware visual state space decoder outputs decoding completion data; After decoding is completed, the decoded data is upsampled and the scale is restored to the scale of the central view to obtain a decoded image; the decoded image is input into a Softmax classifier, and the Softmax classifier outputs a light field semantic segmentation map of the target scene.

9. A complex urban scene light field semantic segmentation system based on a state space model, used in the method according to any one of claims 1 to 8, characterized in that: The system comprises a first module, a second module, a third module and a fourth module; The first module is used to obtain a central view, a depth map and a 4D light field image of a target scene; The second module is used to pre-process the central view, the depth map and the 4D light field image and input them into the first branch, the second branch and the third branch respectively for multi-scale feature extraction; the feature extraction of the first branch and the second branch is based on the visual state space block, and the feature extraction of the third branch is based on the first strategy; The first branch, the second branch and the third branch all output feature data at preset multiple scales; The first strategy includes: obtaining a 2D sub-aperture slice according to the 4D light field image, and extracting spatial long-range context features of the 2D sub-aperture slice through a Mamba model to obtain a first result; obtaining a macro-pixel slice according to the first result, and bidirectionally scanning the macro-pixel slice through a Mamba model to model complementary relationship features between angles; The third module is used to fuse the feature data at each scale in the outputs of the first branch, the second branch and the third branch based on the second strategy to obtain a fused data set; The second strategy includes: interactively enhancing the feature data from the first branch, the second branch and the third branch through the Mamba model to obtain first data, second data and third data respectively, and concatenating the first data, the second data and the third data and performing reverse scanning to obtain fused enhanced features; The fourth module is used to decode the fused data set to obtain a light field semantic segmentation map of the target scene.

Citation Information

Patent Citations

  • Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model

    CN119152205A

  • Video sequence segmentation method based on selective scanning visual state space model

    CN119206568A

  • Light field image super-resolution reconstruction method and device based on state model

    CN119809940A

  • Multi-scale double-flow fusion real-time semantic segmentation method for road scene

    CN119888222A

  • Dynamic obstacle avoidance method based on real-time local grid map construction

    US20230161352A1

Cited By

  • Cooling plate topology generation model, method and device, medium and computer

    CN121118563A

  • Bridge crack segmentation method and system based on direction perception convolution and state space modeling

    CN122175875A