Method and device for predicting semantic occupancy of look-around camera based on time cross state space model and center voting, processor and medium

By constructing a temporal cross-state space model and a center voting module, fusing multi-frame 3D voxel features and improving the foreground object center prediction, the computational complexity and semantic ambiguity problems in the semantic occupancy prediction of surround view cameras are solved, achieving more efficient and accurate semantic occupancy prediction.

CN120635847APending Publication Date: 2025-09-12EAST CHINA UNIV OF SCI & TECH

Patent Information

Application Number
CN202510721549.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing methods suffer from high computational complexity and semantic ambiguity in semantic occupancy prediction based on surround-view cameras, resulting in the foreground object center often being predicted as an empty voxel.

Method used

The temporal cross-state space model and center voting method are adopted to fuse multi-frame 3D voxel features by constructing a cross-state space module, and the relative center voting module is used to improve the prediction of the foreground object center, combined with the deformable attention mechanism and semantic occupancy prediction head for refined processing.

Benefits of technology

The computational complexity is reduced, thereby improving the efficiency of the model. It explicitly votes on the instance center to which the voxel belongs, spontaneously learns to utilize other voxel features of the same instance, and updates the semantics of the vacant space inside the object from coarse to fine, thereby improving the accuracy of semantic occupancy prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635847A_ABST
    Figure CN120635847A_ABST
Patent Text Reader

Abstract

The invention relates to an all-around camera semantic occupation prediction method based on a time cross state space model and center voting. The method comprises the following steps: acquiring an image data set by using an all-around array composed of multiple vehicle-mounted cameras; performing feature extraction and coding by using an image backbone network to obtain multi-scale features of the multi-camera image; generating 3D voxel features of the current frame by adopting a deformable attention mechanism for the multi-scale features of the multi-camera images; fusing 3D voxel features of past multiple frames by using a cross state space module to obtain 3D voxel features fused with time information; performing up-sampling, feature splicing and feature decoding in a multi-scale manner by using a voxel backbone network and a semantic occupancy prediction head to obtain semantic occupancy prediction; a relative central voting module and a parallel fusion policy module are applied to the rough semantic occupancy prediction to generate a fine semantic occupancy prediction. The method has a more accurate semantic occupancy prediction effect on images collected by a look-around array formed by multiple cameras.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method, device, processor, and storage medium for predicting semantic occupancy of a surround view camera based on a time-crossed state-space model and center voting. Background Art

[0002] In recent years, image-based 3D semantic occupancy prediction has become a hot topic in 3D scene understanding for autonomous driving. Image-based 3D semantic occupancy prediction constructs a 3D scene as a grid and determines whether a grid is occupied and the specific semantic category of the occupied grid based on purely visual images. In autonomous driving, it provides more geometric detail than 3D object detection. Compared to the bounding box representation of 3D object detection, the ability to describe the detailed outlines of any obstacles in the scene is a key insight of voxel-based occupancy representation, which facilitates subsequent autonomous driving tasks.

[0003] Currently, there are many existing semantic occupancy prediction methods based on surround-view cameras. For example, the public document "FB-OCC: 3D Occupancy Prediction based on Forward-Backward View Transformation" proposes a semantic occupancy prediction method based on the interaction of 3D forward and backward view transformations; the public document "PanoOcc: Unified Occupancy Representation for Camera-based 3D Panoptic Segmentation" proposes a unified semantic occupancy representation method based on camera 3D panoptic segmentation; and the public document "OctreeOcc: Efficient and Multi-Granularity Occupancy Prediction Using OctreeQueries" proposes an efficient multi-granularity semantic occupancy prediction method based on octree structure queries.

[0004] Invention patent application number CN202410477015.0 proposes a method and apparatus for 3D semantic occupancy prediction driven by a large language model; invention patent application number CN202310316950.4 proposes a method, system, device, medium, and product for 3D semantic occupancy prediction. However, the technical solutions of these applications do not involve the construction of a cross-state space module and its application in the fusion of multi-frame 3D voxel features, nor do they involve the construction of a relative center voting module and its application in improving the problem of foreground object centers often being predicted as empty voxels.

[0005] However, existing technologies have the following problems: (1) Existing methods fuse temporal information based on an attention mechanism, but are limited by high complexity. (2) Existing methods are limited by semantic ambiguity, resulting in the center of foreground objects often being predicted as empty voxels. Therefore, the present invention provides a semantic occupancy prediction method for surround view cameras based on a temporal cross-state space model and center voting. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a surround view camera semantic occupancy prediction method based on a time-crossed state space model and center voting.

[0007] To achieve the above objectives, the method, device, processor, and storage medium for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting are as follows:

[0008] The method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting is characterized in that the method comprises the following steps:

[0009] Step S1: using a surround view array consisting of multiple onboard cameras to collect a real-world image dataset;

[0010] Step S2: preprocessing the image dataset and performing feature extraction and encoding using an image backbone network to obtain a multi-scale feature map of the multi-camera image;

[0011] Step S3: constructing an initialized learnable 3D query voxel, combining the 3D query voxel, the intrinsic parameter matrix of each camera in the surround view array, and the extrinsic parameter matrix relative to the ego vehicle, applying a deformable attention mechanism to the multi-scale feature map of the multi-camera image to generate the 3D voxel feature of the current frame;

[0012] Step S4: constructing a cross-state space module, and using the cross-state space module to fuse the 3D voxel features of multiple past frames based on the 3D voxel features of the current frame to obtain 3D voxel features fused with time information;

[0013] Step S5: constructing a semantic occupancy prediction head, performing upsampling, feature concatenation, and feature decoding on the 3D voxel features fused with temporal information using a voxel backbone network and the semantic occupancy prediction head in a multi-scale manner to obtain a preliminary semantic occupancy prediction;

[0014] Step S6: construct a relative center voting module and a parallel fusion strategy module, apply the relative center voting module and the parallel fusion strategy module to the preliminary semantic occupancy prediction to update the 3D voxel features of the fused time information, and use the semantic occupancy prediction head again to generate a final semantic occupancy prediction as the final result.

[0015] Preferably, the surround view array composed of multiple on-board cameras described in step S1 is specifically as follows: one or more RGB cameras are installed on the vehicle, and the layout is determined by the perspectives of these cameras observing the vehicle's surrounding environment. After the camera installation positions are determined, for each camera, its intrinsic parameter matrix and extrinsic parameter matrix relative to the vehicle are recorded.

[0016] Preferably, the image backbone network described in step S2 is any specific network, such as ResNet, ViT, SwinTransformer, InternImage or VoVNet.

[0017] Preferably, the step S3 specifically includes the following steps:

[0018] Step S31: Construct an initialized learnable 3D query voxel tensor As the basis for 3D spatial information aggregation;

[0019] Step S32: Based on the intrinsic and extrinsic parameters of the N cameras recorded in S1, the center coordinates of each voxel in the 3D voxel query in S31 are projected onto the two-dimensional coordinate system of the N camera images. If there is partial overlap between the multiple camera views, then some 3D voxels will be projected onto multiple camera views simultaneously to obtain voxel-pixel pairs.

[0020] Step S33: perform feature sampling around the two-dimensional coordinates projected in S32, and aggregate the two-dimensional features to 3D voxels through the deformable attention mechanism to generate the 3D voxel features of the current frame. Specifically:

[0021] The pth element of the 3D voxel feature of the current frame is calculated according to the following formula:

[0022]

[0023] Among them, F p and Q p They represent the pth element of the 3D voxel feature of the current frame and the pth element of the 3D voxel query, respectively. p It's Q p The corresponding 3D coordinates of It is a 3D to 2D projection function including internal and external parameters, X i is the feature map of the i-th camera, Indicated by q p The set of hit view indices obtained by projection, DefAttn represents the deformable attention calculation.

[0024] Preferably, the step S4 specifically includes the following steps:

[0025] Step S41: construct a cross-state space module, specifically: simultaneously convert the 3D voxel feature V of the current frame T described in S3 into T and the 3D voxel features V of the historical frame Ti T-i As the input of the cross state space module, construct the initialized learnable tensor and The V T-i As By performing MLP mapping on x, we can obtain By performing MLP mapping on B, we can obtain By the V T Perform MLP mapping to obtain where d state represents the state dimension of the cross-state space module;

[0026] Step S42: V is converted into T-i and V T Flatten into sequences of equal size Simultaneously input into the cross-state space module described in S41;

[0027] Step S43: reconstruct the output of the cross-state space module described in S41 into a 3D voxel feature fused with time information through tensor reshaping and interpolation upsampling

[0028] Step S44: storing the 3D voxel features fused with the time information described in S43 for input into the voxel backbone network and the semantic occupancy prediction head for feature decoding.

[0029] Preferably, the step S5 specifically includes the following steps:

[0030] Step S51: upsampling, feature concatenation, and feature decoding are performed on the 3D voxel features fused with the time information described in S4 using a voxel backbone network in a multi-scale manner to obtain high-resolution 3D voxel features;

[0031] Step S52: Use a two-layer MLP to construct a semantic occupancy prediction head, use the semantic occupancy prediction head to process the high-resolution 3D voxel features described in S51, and output a preliminary occupancy prediction value. and preliminary semantic prediction values

[0032] Step S53: Construct the semantic occupancy prediction head loss function:

[0033]

[0034] in, is the focal loss, For Lovasz's loss, is the scene category affinity loss, The semantic label loss is included in and geometric label loss

[0035] The method further comprises the following steps:

[0036] The semantic occupancy prediction head loss function and the center voting mechanism loss function described in S53 are used to optimize the occupancy prediction model.

[0037] Preferably, the step S6 specifically includes the following steps:

[0038] Step S61: Construct a relative center voting module, specifically:

[0039] The preliminary semantic occupancy prediction described in S5 is used as a category mask to filter the high-resolution 3D voxel features described in S51, retaining only the voxel features of the foreground object. The offsets of these voxels from the center voxel of their respective objects are predicted by the MLP Obtain the voting relationship between each voxel and the central voxel of the object where it is located, that is, the voxel-relative central voxel pair

[0040] Step S62: Construct a parallel fusion strategy module according to the following formula:

[0041]

[0042]

[0043]

[0044] Among them, GP stands for global pooling function and CrossAttn stands for cross attention calculation.

[0045] Step S63: applying the relative center voting module described in S61 and the parallel fusion strategy module described in S62 to the high-resolution 3D voxel features described in S51 to update the features;

[0046] Step S64: applying the semantic occupancy prediction head described in S52 to the updated high-resolution 3D voxel features described in S63 to generate a refined semantic occupancy prediction as a final result;

[0047] Step S65: Construct a central voting mechanism loss function, specifically:

[0048] The true center coordinates δ of the foreground object are obtained by the center coordinates of the 3D bounding box label cen, and the loss is calculated using the following formula:

[0049]

[0050] Among them, MSE stands for mean square error calculation function.

[0051] The device for semantic occupancy prediction of a surround view camera based on a time-crossed state-space model and center voting is characterized in that the device comprises:

[0052] a processor configured to execute computer-executable instructions;

[0053] A memory stores one or more computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting are implemented.

[0054] The processor for semantic occupancy prediction of a surround-view camera based on a temporal cross-state space model and center voting is characterized in that the processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the various steps of the method for semantic occupancy prediction of a surround-view camera based on a temporal cross-state space model and center voting are implemented.

[0055] The computer-readable storage medium is characterized in that a computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) The present invention extends the state-space model to support multi-input sequence interactions and performs temporal modeling in a cascade architecture, thereby reducing the computational complexity from quadratic to linear.

[0058] (2) The present invention enables the model to explicitly vote for the instance center to which a voxel belongs, and spontaneously learn to utilize other voxel features of the same instance to update the semantics of the internal space of the object from coarse to fine. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 It is a schematic diagram of the overall structure of the present invention;

[0060] Figure 2 Schematic diagram of the method flow of the present invention;

[0061] Figure 3 Schematic diagram of the cross state space module of the present invention;

[0062] Figure 4 This is a schematic diagram of the relative center voting module of the present invention;

[0063] Figure 5 This is a schematic diagram of the parallel fusion strategy module of the present invention;

[0064] Figure 6 This is the visualization result of the present invention on the Occ3D-nuScenes dataset. DETAILED DESCRIPTION

[0065] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.

[0066] The method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting of the present invention comprises the following steps:

[0067] Step S1: using a surround view array consisting of multiple onboard cameras to collect a real-world image dataset;

[0068] Step S2: preprocessing the image dataset and performing feature extraction and encoding using an image backbone network to obtain a multi-scale feature map of the multi-camera image;

[0069] Step S3: constructing an initialized learnable 3D query voxel, combining the 3D query voxel, the intrinsic parameter matrix of each camera in the surround view array, and the extrinsic parameter matrix relative to the ego vehicle, applying a deformable attention mechanism to the multi-scale feature map of the multi-camera image to generate the 3D voxel feature of the current frame;

[0070] Step S4: constructing a cross-state space module, and using the cross-state space module to fuse the 3D voxel features of multiple past frames based on the 3D voxel features of the current frame to obtain 3D voxel features fused with time information;

[0071] Step S5: constructing a semantic occupancy prediction head, performing upsampling, feature concatenation, and feature decoding on the 3D voxel features fused with temporal information using a voxel backbone network and the semantic occupancy prediction head in a multi-scale manner to obtain a preliminary semantic occupancy prediction;

[0072] Step S6: construct a relative center voting module and a parallel fusion strategy module, apply the relative center voting module and the parallel fusion strategy module to the preliminary semantic occupancy prediction to update the 3D voxel features of the fused time information, and use the semantic occupancy prediction head again to generate a final semantic occupancy prediction as the final result.

[0073] As a preferred embodiment of the present invention, the surround view array composed of multiple on-board cameras described in step S1 is specifically as follows: one or more RGB cameras are installed on the vehicle, and the layout is determined by the perspectives of these cameras observing the vehicle's surrounding environment. After the camera installation positions are determined, for each camera, its intrinsic parameter matrix and extrinsic parameter matrix relative to the vehicle are recorded.

[0074] As a preferred embodiment of the present invention, the image backbone network described in step S2 is selected from any specific network, such as ResNet, ViT, SwinTransformer, InternImage or VoVNet.

[0075] As a preferred embodiment of the present invention, step S3 specifically includes the following steps:

[0076] Step S31: Construct an initialized learnable 3D query voxel tensor As the basis for 3D spatial information aggregation;

[0077] Step S32: Based on the intrinsic and extrinsic parameters of the N cameras recorded in S1, the center coordinates of each voxel in the 3D voxel query in S31 are projected onto the two-dimensional coordinate system of the N camera images. If there is partial overlap between the multiple camera views, then some 3D voxels will be projected onto multiple camera views simultaneously to obtain voxel-pixel pairs.

[0078] Step S33: perform feature sampling around the two-dimensional coordinates projected in S32, and aggregate the two-dimensional features to 3D voxels through the deformable attention mechanism to generate the 3D voxel features of the current frame. Specifically:

[0079] The pth element of the 3D voxel feature of the current frame is calculated according to the following formula:

[0080]

[0081] Among them, F p and Q p They represent the pth element of the 3D voxel feature of the current frame and the pth element of the 3D voxel query, respectively. p It's Q p The corresponding 3D coordinates of It is a 3D to 2D projection function including internal and external parameters, X i is the feature map of the i-th camera, Indicated by q p The set of hit view indices obtained by projection, DefAttn represents the deformable attention calculation.

[0082] As a preferred embodiment of the present invention, step S4 specifically includes the following steps:

[0083] Step S41: construct a cross-state space module, specifically: simultaneously convert the 3D voxel feature V of the current frame T described in S3 into T and the 3D voxel features V of the historical frame Ti T-i As the input of the cross state space module, construct the initialized learnable tensor and The V T-i As By performing MLP mapping on x, we can obtain By performing MLP mapping on B, we can obtain By the V T Perform MLP mapping to obtain where d state represents the state dimension of the cross-state space module;

[0084] Step S42: V is converted into T-i and V T Flatten into sequences of equal size Simultaneously input into the cross-state space module described in S41;

[0085] Step S43: reconstruct the output of the cross-state space module described in S41 into a 3D voxel feature fused with time information through tensor reshaping and interpolation upsampling

[0086] Step S44: storing the 3D voxel features fused with the time information described in S43 for input into the voxel backbone network and the semantic occupancy prediction head for feature decoding.

[0087] As a preferred embodiment of the present invention, step S5 specifically includes the following steps:

[0088] Step S51: upsampling, feature concatenation, and feature decoding are performed on the 3D voxel features fused with the time information described in S4 using a voxel backbone network in a multi-scale manner to obtain high-resolution 3D voxel features;

[0089] Step S52: Use a two-layer MLP to construct a semantic occupancy prediction head, use the semantic occupancy prediction head to process the high-resolution 3D voxel features described in S51, and output a preliminary occupancy prediction value. and preliminary semantic prediction values

[0090] Step S53: Construct the semantic occupancy prediction head loss function:

[0091]

[0092] in, is the focal loss, For Lovasz's loss, is the scene category affinity loss, The semantic label loss is included in and geometric label loss

[0093] The method further comprises the following steps:

[0094] The semantic occupancy prediction head loss function and the center voting mechanism loss function described in S53 are used to optimize the occupancy prediction model.

[0095] As a preferred embodiment of the present invention, step S6 specifically includes the following steps:

[0096] Step S61: Construct a relative center voting module, specifically:

[0097] The preliminary semantic occupancy prediction described in S5 is used as a category mask to filter the high-resolution 3D voxel features described in S51, retaining only the voxel features of the foreground object. The offsets of these voxels from the center voxel of their respective objects are predicted by the MLP Obtain the voting relationship between each voxel and the central voxel of the object where it is located, that is, the voxel-relative central voxel pair

[0098] Step S62: Construct a parallel fusion strategy module according to the following formula:

[0099]

[0100]

[0101]

[0102] Among them, GP stands for global pooling function and CrossAttn stands for cross attention calculation.

[0103] Step S63: applying the relative center voting module described in S61 and the parallel fusion strategy module described in S62 to the high-resolution 3D voxel features described in S51 to update the features;

[0104] Step S64: applying the semantic occupancy prediction head described in S52 to the updated high-resolution 3D voxel features described in S63 to generate a refined semantic occupancy prediction as a final result;

[0105] Step S65: Construct a central voting mechanism loss function, specifically:

[0106] The true center coordinates δ of the foreground object are obtained by the center coordinates of the 3D bounding box label cen , and the loss is calculated using the following formula:

[0107]

[0108] Among them, MSE stands for mean square error calculation function.

[0109] The device for semantic occupancy prediction of a surround view camera based on a time-crossed state-space model and center voting of the present invention comprises:

[0110] a processor configured to execute computer-executable instructions;

[0111] A memory stores one or more computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting are implemented.

[0112] The processor of the surround-view camera semantic occupancy prediction based on the temporal cross-state space model and center voting of the present invention is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the various steps of the method for semantic occupancy prediction of the surround-view camera based on the temporal cross-state space model and center voting are implemented.

[0113] The computer-readable storage medium of the present invention stores a computer program thereon, which can be executed by a processor to implement the various steps of the above-mentioned method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting.

[0114] The examples are as follows:

[0115] This embodiment provides a semantic occupancy prediction method for surround view cameras based on a time-crossing state space model and center voting, such as Figure 1 As shown, the method includes: a cross-state space module, a relative center voting module, a parallel fusion strategy module, and a semantic occupancy prediction head module.

[0116] like Figure 2 As shown, the method includes the following steps:

[0117] Step 1: Use a surround view array consisting of multiple on-board cameras to collect a real-world image dataset. Specifically, one or more RGB cameras are installed on the car. These cameras are arranged based on the perspectives from which they can observe the car's surrounding environment. The perspectives of each camera can have some overlap or no overlap. After the camera installation positions are determined, for each camera, its intrinsic parameter matrix and the extrinsic parameter matrix relative to the vehicle must be recorded.

[0118] Step 2: Preprocess the image dataset and use the image backbone network ResNet-101-DCN to perform feature extraction and encoding to obtain a multi-scale feature map of the multi-camera image.

[0119] Step 3: Construct an initialized learnable 3D query voxel. Combine the 3D query voxel, the intrinsic parameter matrix of each camera in the surround view array, and the extrinsic parameter matrix relative to the ego vehicle, and apply a deformable attention mechanism to the multi-scale feature map of the multi-camera image to generate the 3D voxel features of the current frame.

[0120] Specifically, the step 3 includes the following steps:

[0121] Step 3.1: Construct an initialized learnable 3D query voxel tensor As the basis for 3D spatial information aggregation;

[0122] Step 3.2: Based on the intrinsic and extrinsic parameters of the N cameras recorded in step 1, project the center coordinates of each voxel in the 3D voxel query in step 3.1 onto the 2D coordinate system of the N camera images. If there is some overlap between the camera views, some 3D voxels will be projected onto multiple camera views simultaneously to obtain voxel-pixel pairs.

[0123] Step 3.3: Perform feature sampling around the two-dimensional coordinates projected in step 3.2, and aggregate the two-dimensional features to 3D voxels through the deformable attention mechanism to generate the 3D voxel features of the current frame. The calculation uses the following formula;

[0124]

[0125] Among them, F p and Q p They represent the pth element of the output 3D voxel feature of the current frame and the pth element of the 3D voxel query respectively. p It's Q p The corresponding 3D coordinates of It is a 3D to 2D projection function including internal and external parameters, X i is the feature map of the i-th camera. Indicated by q pThe set of hit view indices obtained by projection. DefAttn represents deformable attention calculation.

[0126] Step 4: construct a cross-state space module. Based on the 3D voxel features of the current frame, use the cross-state space module to fuse the 3D voxel features of multiple past frames to obtain 3D voxel features that fuse time information.

[0127] Specifically, if Figure 3 As shown, the step 4 includes the following steps:

[0128] Step 4.1: Construct a cross-state space module, specifically: at the same time, the 3D voxel feature V of the current frame T described in step 3 is T and the 3D voxel features V of the historical frame Ti T-i As the input of the cross state space module, construct the initialized learnable tensor and V T-i As By performing MLP mapping on x, we can obtain By performing MLP mapping on B, we can obtain By V T Perform MLP mapping to obtain where d state represents the state dimension of the cross-state space module;

[0129] Step 4.2: V is converted to ∑ T-i and V T Flatten into sequences of equal size At the same time, input the cross-state space model described in step 4.1;

[0130] Step 4.3: Reconstruct the output of the cross-state space model described in step 4.1 into a 3D voxel feature fused with temporal information by tensor reshaping and interpolation upsampling

[0131] Step 4.4: Store the 3D voxel features fused with the temporal information described in step 4.3 for input into the voxel backbone network and the semantic occupancy prediction head for feature decoding.

[0132] Step 5: Construct a semantic occupancy prediction head, and use a voxel backbone network and the semantic occupancy prediction head to perform upsampling, feature concatenation, and feature decoding on the 3D voxel features fused with temporal information in a multi-scale manner to obtain a rough semantic occupancy prediction.

[0133] Specifically, the step 5 includes the following steps:

[0134] Step 5.1: Upsampling, feature concatenation, and feature decoding the 3D voxel features fused with temporal information described in step 4 are performed using a voxel backbone network in a multi-scale manner to obtain high-resolution 3D voxel features.

[0135] Step 5.2: Use a two-layer MLP to construct a semantic occupancy prediction head, use the semantic occupancy prediction head to process the high-resolution 3D voxel features described in step 5.1, and output a rough occupancy prediction value and the rough semantic prediction value

[0136] Step 5.3: Construct the semantic occupancy prediction head loss function:

[0137]

[0138] in, is the focal loss, For Lovasz's loss, is the scene category affinity loss, The semantic label loss is included in and geometric label loss

[0139] Step 6: Construct a relative center voting module and a parallel fusion strategy module, apply the relative center voting module and the parallel fusion strategy module to the rough semantic occupancy prediction to update the 3D voxel features of the fused time information, and use the semantic occupancy prediction head again to generate a fine semantic occupancy prediction as the final result.

[0140] Specifically, the step 6 includes the following steps:

[0141] Step 6.1, as Figure 4 As shown, a relative center voting module is constructed: the coarse semantic occupancy prediction described in step 5 is used as a category mask to filter the high-resolution 3D voxel features described in step 5.1, retaining only the voxel features of the foreground target The MLP predicts the voxel distances Δx from the center voxel of their respective objects (including the voxel distances in the H, W, and L directions) i , Δy i , Δz i ) Obtain the voting relationship between each voxel and the central voxel of the object where it is located, that is, the voxel-relative central voxel pair

[0142] Step 6.2, such as Figure 5 As shown, build a parallel fusion strategy module:

[0143]

[0144]

[0145]

[0146] Among them, GP stands for global pooling function and CrossAttn stands for cross attention calculation.

[0147] Step 6.3, applying the relative center voting module described in step 6.1 and the parallel fusion strategy module described in step 6.2 to the high-resolution 3D voxel features described in step 5.1 to update the features;

[0148] Step 6.4: Apply the semantic occupancy prediction head described in step 5.2 to the updated high-resolution 3D voxel features described in step 6.3 to generate a refined semantic occupancy prediction as the final result.

[0149] Step 6.5: Construct the center voting mechanism loss function: obtain the true center coordinates of the foreground object δ through the center coordinates of the 3D bounding box label cen , and the loss is calculated using the following formula:

[0150]

[0151] Among them, MSE stands for mean square error calculation function.

[0152] In addition, the semantic occupancy prediction head loss function and the center voting mechanism loss function are used to train and optimize the 3D semantic occupancy prediction model. Specifically, the overall loss function is constructed according to the following formula

[0153]

[0154] In order to verify the semantic occupancy prediction effect of this embodiment, an experiment was conducted to verify the method provided in this embodiment. The specific contents are as follows:

[0155] (1) Experimental verification was conducted using the Occ3D-nuScenes dataset. Each frame of this data consists of six images captured by a six-camera array and the semantic occupancy ground truth labels of the 3D scene. In this experimental verification, occupancy predictions were performed for all semantic categories.

[0156] (2) The training process is as follows: Common data augmentation strategies are applied to each frame, including color transformation, flipping, rotation, and scaling in image and 3D space. The batch size is set to 4. The AdamW optimizer is used, and the initial learning rate is set to 2×10 -4 , weight decay is set to 0.05, and cosine annealing strategy is applied. The model is trained for 24 epochs.

[0157] (3) In order to verify the effectiveness of the proposed method, this experiment selected three common 3D semantic occupancy prediction methods based on pure image vision, and used these three 3D semantic occupancy prediction methods as a comparison group for comparative experiments.

[0158] Comparison group 1 is: the 3D semantic occupancy prediction method mentioned in the public document "FB-OCC: 3D Occupancy Prediction based on Forward-Backward View Transformation".

[0159] Comparison group 2 is: the 3D semantic occupancy prediction method mentioned in the public document "PanoOcc: Unified Occupancy Representation for Camera-based 3D Panoptic Segmentation".

[0160] Comparison group 3 is: the 3D semantic occupancy prediction method mentioned in the public document "OctreeOcc: Efficient and Multi-GranularityOccupancy Prediction Using Octree Queries".

[0161] Table 1 shows the quantitative results of comparative experiments on the Occ3D-nuScenes data, using mIoU as the evaluation method. The data in Table 1 demonstrates that the method provided in this embodiment achieves the best results. The quantitative experimental results demonstrate that the method provided in this embodiment outperforms other methods in terms of mIoU, with precisions of 4.26, 2.80, and 0.91 higher, respectively.

[0162] Table 1 Comparison of 3D semantic occupancy prediction results

[0163]

[0164]

[0165] The qualitative results of the experiment can be found in Figure 6 , Figure 6This is a comparison of the 3D semantic occupancy prediction proposed in this embodiment with other comparison groups and true labels. The first row is the 3D semantic occupancy prediction results of comparison group 1, the second row is the 3D semantic occupancy prediction results of comparison group 2, the third row is the 3D semantic occupancy prediction results of this embodiment, and the fourth row is the 3D semantic occupancy true value label. Overall, the surround view camera semantic occupancy prediction method based on the time-crossed state-space model and center voting proposed in this invention can more completely understand the overall structure of the scene and capture finer-grained geometric details, and can effectively implement 3D semantic occupancy prediction for surround view camera scenes.

[0166] This technical solution involves the construction of a cross-state space module and its use in the fusion of multi-frame 3D voxel features. It also proposes the construction of a center voting module and its use in improving the problem that the center of foreground objects is often predicted as an empty voxel.

[0167] First, in existing computer vision techniques, state-space models are used to update the features of 2D or 3D feature maps themselves through various scanning methods. However, there is no precedent for using state-space models for feature updates between two 2D or 3D feature maps. This solution proposes a cross-state-space model and further applies it to the interaction and fusion of multi-frame 3D voxel features. Second, existing solutions are limited by semantic ambiguity, resulting in the center of foreground objects often being predicted as empty voxels. This solution explicitly constructs a center voting module and uses it to improve this problem, which is novel, innovative, and has practical significance.

[0168] The specific implementation scheme of this embodiment can be found in the relevant descriptions in the above embodiments and will not be repeated here.

[0169] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0170] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.

[0171] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0172] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0173] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the corresponding program can be stored in a computer-readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0174] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0175] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0176] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0177] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it will be apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A semantic occupancy prediction method for surround view cameras based on a temporal cross-state space model and center voting, characterized in that: The following steps are involved: Step S1: using a surround view array consisting of multiple onboard cameras to collect a real-world image dataset; Step S2: preprocessing the image dataset and performing feature extraction and encoding using an image backbone network to obtain a multi-scale feature map of the multi-camera image; Step S3: constructing an initialized learnable 3D query voxel, combining the 3D query voxel, the intrinsic parameter matrix of each camera in the surround view array, and the extrinsic parameter matrix relative to the ego vehicle, applying a deformable attention mechanism to the multi-scale feature map of the multi-camera image to generate the 3D voxel feature of the current frame; Step S4: constructing a cross-state space module, and using the cross-state space module to fuse the 3D voxel features of multiple past frames based on the 3D voxel features of the current frame to obtain 3D voxel features fused with time information; Step S5: constructing a semantic occupancy prediction head, performing upsampling, feature concatenation, and feature decoding on the 3D voxel features fused with temporal information using a voxel backbone network and the semantic occupancy prediction head in a multi-scale manner to obtain a preliminary semantic occupancy prediction; Step S6: construct a relative center voting module and a parallel fusion strategy module, apply the relative center voting module and the parallel fusion strategy module to the preliminary semantic occupancy prediction to update the 3D voxel features of the fused time information, and use the semantic occupancy prediction head again to generate a final semantic occupancy prediction as the final result.

2. The method for semantic occupancy prediction of surround view cameras based on a time-crossed state space model and center voting according to claim 1, characterized in that: The surround view array formed by multiple onboard cameras described in step S1 is specifically as follows: one or more RGB cameras are installed on the vehicle, and the layout is determined by the perspective of these cameras observing the vehicle's surrounding environment. After the camera installation positions are determined, for each camera, its intrinsic parameter matrix and extrinsic parameter matrix relative to the vehicle are recorded.

3. The method for semantic occupancy prediction of surround view cameras based on a time-crossed state space model and center voting according to claim 1, characterized in that: The image backbone network described in step S2 can be any specific network, such as ResNet, ViT, SwinTransformer, InternImage or VoVNet.

4. The method for semantic occupancy prediction of surround view cameras based on a time-crossed state space model and center voting according to claim 1, characterized in that: The step S3 specifically includes the following steps: Step S31: Construct an initialized learnable 3D query voxel tensor As the basis for 3D spatial information aggregation; Step S32: Based on the intrinsic and extrinsic parameters of the N cameras recorded in S1, the center coordinates of each voxel in the 3D voxel query in S31 are projected onto the two-dimensional coordinate system of the N camera images. If there is partial overlap between the multiple camera views, then some 3D voxels will be projected onto multiple camera views simultaneously to obtain voxel-pixel pairs. Step S33: perform feature sampling around the two-dimensional coordinates projected in S32, and aggregate the two-dimensional features to 3D voxels through the deformable attention mechanism to generate the 3D voxel features of the current frame. Specifically: The pth element of the 3D voxel feature of the current frame is calculated according to the following formula: Among them, F p and Q p They represent the pth element of the 3D voxel feature of the current frame and the pth element of the 3D voxel query, respectively. p It's Q p The corresponding 3D coordinates of It is a 3D to 2D projection function including internal and external parameters, X i is the feature map of the i-th camera, Indicated by q p The set of hit view indices obtained by projection, DefAttn represents the deformable attention calculation.

5. The method for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting according to claim 1, characterized in that: The step S4 specifically includes the following steps: Step S41: construct a cross-state space module, specifically: simultaneously convert the 3D voxel feature V of the current frame T described in S3 into T and the 3D voxel features V of the historical frame Ti T-i As the input of the cross state space module, construct the initialized learnable tensor and The V T-i As By performing MLP mapping on x, we can get By performing MLP mapping on B, we can obtain By the V T Perform MLP mapping to obtain where d state represents the state dimension of the cross-state space module; Step S42: V is converted into T-i and V T Flatten to a sequence of equal size Simultaneously input into the cross-state space module described in S41; Step S43: reconstruct the output of the cross-state space module described in S41 into a 3D voxel feature fused with time information through tensor reshaping and interpolation upsampling Step S44: storing the 3D voxel features fused with the time information described in S43 for input into the voxel backbone network and the semantic occupancy prediction head for feature decoding.

6. The method for semantic occupancy prediction of surround view cameras based on a time-crossed state space model and center voting according to claim 1, characterized in that: The step S5 specifically includes the following steps: Step S51: upsampling, feature concatenation, and feature decoding are performed on the 3D voxel features fused with the time information described in S4 using a voxel backbone network in a multi-scale manner to obtain high-resolution 3D voxel features; Step S52: Use a two-layer MLP to construct a semantic occupancy prediction head, use the semantic occupancy prediction head to process the high-resolution 3D voxel features described in S51, and output a preliminary occupancy prediction value. and preliminary semantic prediction values Step S53: Construct the semantic occupancy prediction head loss function: in, is the focal loss, For Lovasz's loss, is the scene category affinity loss, The semantic label loss is included in and geometric label loss The method further comprises the following steps: The semantic occupancy prediction head loss function and the center voting mechanism loss function described in S53 are used to optimize the occupancy prediction model.

7. The method for semantic occupancy prediction of surround view cameras based on a time-crossed state space model and center voting according to claim 1, characterized in that: The step S6 specifically includes the following steps: Step S61: Construct a relative center voting module, specifically: The preliminary semantic occupancy prediction described in S5 is used as a category mask to filter the high-resolution 3D voxel features described in S51, retaining only the voxel features of the foreground object. The offsets of these voxels from the center voxel of their respective objects are predicted by the MLP Obtain the voting relationship between each voxel and the central voxel of the object where it is located, that is, the voxel-relative central voxel pair Step S62: Construct a parallel fusion strategy module according to the following formula: Among them, GP stands for global pooling function and CrossAttn stands for cross attention calculation. Step S63: applying the relative center voting module described in S61 and the parallel fusion strategy module described in S62 to the high-resolution 3D voxel features described in S51 to update the features; Step S64: applying the semantic occupancy prediction head described in S52 to the updated high-resolution 3D voxel features described in S63 to generate a refined semantic occupancy prediction as a final result; Step S65: Construct a central voting mechanism loss function, specifically: The true center coordinates δ of the foreground object are obtained by the center coordinates of the 3D bounding box label cen , and the loss is calculated using the following formula: Among them, MSE stands for mean square error calculation function.

8. A device for semantic occupancy prediction of surround view cameras based on a time-crossed state-space model and center voting, characterized in that: The device comprises: a processor configured to execute computer-executable instructions; A memory storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by the processor, the steps of the method for semantic occupancy prediction of a surround view camera based on a time-crossed state-space model and center voting are implemented as described in any one of claims 1 to 7.

9. A processor for semantic occupancy prediction of surround view cameras based on a time-crossing state-space model and center voting, characterized in that: The processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for semantic occupancy prediction of a surround view camera based on a time-crossed state-space model and center voting are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the method for semantic occupancy prediction of a surround view camera based on a time-crossed state-space model and center voting as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Three-dimensional semantic occupancy prediction method, system and device, medium and product

    CN116630912A

  • Three-dimensional semantic occupancy prediction method and device based on large language model driving

    CN118196796A

Cited By

  • Double-flow layered brain image registration method and system based on state space model

    CN121482114A