Occupancy prediction method, system and device for automatic driving and medium
Through the combination of a two-dimensional convolutional network and a scene memory gate module, the problem of low computing efficiency in the prior art is solved, and the three-dimensional occupation prediction with high precision and low parameter quantity is realized, meeting the real-time requirements of autonomous driving.
Patent Information
- Application Number
- CN202510508125.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
AI Technical Summary
In existing autonomous driving systems, 3D semantic occupation prediction technology ignores computing efficiency and resource limitations when pursuing high-precision prediction, resulting in a huge amount of model parameters and insufficient inference speed, making it difficult to meet real-time requirements.
A two-dimensional convolutional network is used to extract multi-view image features, convert them into three-dimensional voxel features and project them to generate bird's-eye view features, and dynamically fusion is combined with the scene memory gate module. Through multi-scale geometric coding and multi-level semantic coding, three-dimensional occupation prediction results are generated, and a lightweight encoder is used to reduce the complexity of the model.
With limited computing resources, high-precision three-dimensional occupation prediction is achieved, the number of model parameters is reduced, the inference speed is improved, and the real-time needs of autonomous driving are met.
Smart Images

Figure CN120339519A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving environment perception, and more specifically, to an occupancy prediction method, system, device, and medium for autonomous driving. Background Art
[0002] In an autonomous driving system, accurately perceiving and understanding the surrounding environment is the core prerequisite for achieving safe navigation and decision-making. In recent years, 3D semantic occupancy prediction technology has attracted much attention because it can reconstruct the fine-grained geometry and semantic information of the scene in the form of voxels, providing the vehicle with the ability to deeply understand occluded areas and complex scenes. However, while existing methods pursue high-precision prediction, they often neglect the computational efficiency and resource limitations in actual deployment, resulting in a large number of model parameters and insufficient inference speed, making it difficult to meet the real-time requirements of in-vehicle embedded platforms.
[0003] In the prior art, fusion methods based on temporal image sequences are widely used to improve occupancy prediction performance. Figure 1 (a) shows a common method of temporal feature aggregation, which repeatedly uses the backbone network to extract features from historical images and inputs these temporal features into a fusion module for aggregation. The fusion module usually consists of multiple convolutional layers; however, such methods have significant drawbacks: First, the repeated calculation of historical image features leads to a large amount of redundant overhead; Second, there is redundant aggregation of feature overlap regions in the temporal fusion module when processing long sequences, such as the repeated fusion of features between adjacent time windows; Third, the pose differences between early historical frames and the current scene caused by vehicle movement are not effectively processed, and the rough fusion method may introduce noise interference. Although methods such as GSD-OCC (as shown in Figure 1 (b)) propose a feature queue caching mechanism to avoid repeated feature extraction, its storage overhead grows linearly with the sequence length, and it does not solve the inefficiency problem of long-term feature fusion.
[0004] In addition, existing 3D occupancy prediction models generally adopt complex three-dimensional convolutional or Transformer architectures, such as OccFormer, SurroundOcc, etc. Although the prediction accuracy is improved, the number of model parameters and computational complexity increase significantly. Although lightweight methods such as SparseOcc, TPVFormer simplify the calculation through sparsification or perspective projection, their design of the temporal fusion module is still insufficient, making it difficult to balance efficiency and performance. Therefore, how to design an occupancy prediction model that takes into account high accuracy, low number of parameters, and real-time inference under limited computational resources has become an urgent technical problem to be solved. Summary of the Invention
[0005] The present invention provides an occupancy prediction method, system, device and medium for autonomous driving, aiming to solve the problems of redundant feature extraction, low long-term fusion efficiency and high model complexity in existing methods.
[0006] To achieve the above object, the first aspect of the present invention provides an occupancy prediction method for autonomous driving, including the following steps:
[0007] Obtain a multi-view image sequence at the current moment, and extract multi-view image features through a two-dimensional convolutional network;
[0008] Convert the multi-view image features into three-dimensional voxel features, and project them to generate bird's-eye view features at the current moment;
[0009] Input the bird's-eye view features at the current moment into the scene memory gate module, and perform dynamic fusion with the historical feature map to generate an updated historical feature map;
[0010] Convert the updated historical feature map into three-dimensional semantic features, and fuse them with the three-dimensional voxel features to generate enhanced three-dimensional features;
[0011] Perform multi-scale geometric encoding processing on the enhanced three-dimensional features to generate three-dimensional geometric features;
[0012] Perform multi-level semantic encoding processing on the updated historical feature map to generate two-dimensional semantic features;
[0013] Fuse the three-dimensional geometric features and the two-dimensional semantic features, and output a three-dimensional occupancy prediction result.
[0014] Further, the dynamic fusion operation of the scene memory gate includes:
[0015] Convert the historical feature map into a hidden state through a hidden block;
[0016] Calculate the value score of each pixel based on the hidden features and the bird's-eye view features at the current moment;
[0017] Generate a mask according to the value score, and filter out valuable historical features;
[0018] Fuse the filtered historical feature map and the bird's-eye view features at the current moment to generate an updated historical feature map.
[0019] Further, the method for converting the multi-view image features into three-dimensional voxel features includes:
[0020] Use a depth network to predict the depth distribution of each view image;
[0021] Extract image semantic features through a semantic network;
[0022] Generate pseudo 3D point cloud features by taking the outer product of the image semantic features and the depth distribution;
[0023] Convert the pseudo 3D point cloud features into three-dimensional voxel features through voxel pooling.
[0024] Further, the method for converting the updated historical feature map into three-dimensional semantic features includes:
[0025] Decompose the updated historical feature map into geometric features and semantic features;
[0026] Fuse the geometric features and semantic features through an outer product operation to generate three-dimensional semantic features.
[0027] Further, the multi-scale geometric encoding process uses a reparameterized three-dimensional convolutional network, including:
[0028] In the training stage, extract local geometric features through the first three-dimensional convolutional layer using a non-dilated small convolutional kernel;
[0029] In the training stage, stack the second three-dimensional convolutional layer, and the second three-dimensional convolutional layer uses a dilated small convolutional kernel to expand the receptive field;
[0030] In the inference stage, merge the parameters of the non-dilated small convolutional kernel and the dilated small convolutional kernel into a single equivalent large convolutional kernel to generate reparameterized three-dimensional geometric features.
[0031] Further, the method for multi-level semantic encoding processing includes:
[0032] Step S61: Perform a spatial downsampling operation on the updated historical feature map to generate a high-resolution low-dimensional feature and a low-resolution high-dimensional feature;
[0033] Step S62: Input the high-resolution low-dimensional feature and the low-resolution high-dimensional feature into an encoding block containing a variable receptive field weighted key-value module, and achieve cross-level fusion through the following sub-steps:
[0034] Step S621: In the spatial mixing branch, perform a Q-shift operation on the input feature to generate a spatial key feature, a spatial value feature, and a spatial gating weight;
[0035] Step S622: Based on the spatial key feature and the spatial value feature, calculate the global spatial attention feature through bidirectional weighted key-value attention, and its formula is:
[0036]
[0037] where, wkv tRepresents the output result of the bidirectional attention mechanism at position t. Bi-WKV is the bidirectional attention mechanism, K and V are the key-value pairs input to the bidirectional attention mechanism, t represents the target position index in the current feature sequence for which the attention weights are to be calculated, i represents the other positions in the feature sequence except the target position t, L represents the length of the sequence, w is a learnable scalar parameter that controls the attenuation intensity of adjacent position features, u is a learnable scalar parameter that enhances the weight of the current position, k i is the value of the key at position i, k t is the value of the key at position t, v i is the value of the value at position i, v t is the value of the value at position t;
[0038] Step S623: Multiply the global spatial attention feature and the spatial gating weight element-wise, and generate the spatial branch output feature through linear projection;
[0039] Step S624: In the channel mixing branch, perform channel projection on the input feature to generate the channel key feature and the channel gating weight, and generate the channel branch output feature through the SquaredReLU activation function;
[0040] Step S63: Add the spatial branch output feature and the channel branch output feature to obtain the fused cross-level semantic feature;
[0041] Step S64: Perform convolutional aggregation on the cross-level semantic feature and output the two-dimensional semantic feature.
[0042] Further, the method for fusing the three-dimensional geometric feature and the two-dimensional semantic feature and outputting the three-dimensional occupancy prediction result includes:
[0043] Add the three-dimensional geometric feature and the transformed three-dimensional semantic feature voxel-wise;
[0044] Upsample and classify the added feature through the prediction head to output the final three-dimensional occupancy prediction result.
[0045] To achieve the above object, the second aspect of the present invention provides an occupancy prediction system for autonomous driving, including the following modules:
[0046] The multi-view image acquisition module is used to acquire the multi-view image sequence at the current moment;
[0047] The two-dimensional feature extraction module is based on a two-dimensional convolutional network to extract features from the multi-view image and generate multi-view image features;
[0048] The three-dimensional voxel projection module is used to convert the multi-view image features into three-dimensional voxel features and project them to generate the bird's-eye view feature at the current moment;
[0049] A scene memory gate module for generating an updated historical feature map by dynamically fusing the current bird's-eye view features with the historical feature map;
[0050] A three-dimensional semantic conversion module for converting the updated historical feature map into three-dimensional semantic features and fusing them with the three-dimensional voxel features to generate enhanced three-dimensional features;
[0051] A multi-scale geometric encoding module for performing multi-scale geometric structure encoding on the enhanced three-dimensional features to generate three-dimensional geometric features;
[0052] A multi-level semantic encoding module for performing multi-level semantic feature extraction on the historical feature map to generate two-dimensional semantic features;
[0053] A cross-dimensional fusion module for fusing the three-dimensional geometric features with the two-dimensional semantic features and outputting a three-dimensional occupancy prediction result.
[0054] To achieve the above object, a third aspect of the present invention provides an electronic device, including a processor and a memory, where the processor is configured to implement the steps of the occupancy prediction method for autonomous driving when executing a computer program stored in the memory.
[0055] To achieve the above object, a fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program is configured to implement the steps of the occupancy prediction method for autonomous driving when run by a processor.
[0056] The beneficial effects of the present invention:
[0057] Compared with the prior art, a occupancy prediction method, system, device and medium for autonomous driving provided by the present invention first proposes a Scene Memory Gate (SMG) for the problem of repeated extraction of temporal features. It dynamically condenses historical multi-frame information into a single feature map, selectively retains valid historical features (such as occlusion area information) through a masking mechanism, and recursively updates the historical feature map, avoiding the feature extraction operation of repeatedly calling the backbone network for historical images in traditional methods. Secondly, for the problem of redundant long-term fusion, a single feature map fusion paradigm is adopted to replace the traditional multi-frame sequence aggregation. By dynamically weighted (adjusting the fusion ratio of historical features and current features based on the score of current feature uncertainty), the redundant calculation caused by overlapping time windows is eliminated, and the computational complexity of long-term sequence fusion is reduced. Finally, for the problem of high model complexity, a lightweight encoder is designed. A vector recursive weighted key-value (VRWKV) with linear complexity is used to replace the traditional Transformer or large kernel convolution. Through a spatial-channel double-branch structure, the balance between global perception and local details is achieved. Combined with a geometric encoder of reparameterized 3D convolution, while maintaining the multi-scale geometric modeling ability, the model parameters are compressed, achieving a double breakthrough in efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0059] Figure 1 It is a comparison diagram of an occupancy prediction temporal fusion method disclosed in an embodiment of the present invention.
[0060] Figure 2 It is an overall architecture diagram of a ReOcc disclosed in an embodiment of the present invention.
[0061] Figure 3 It is an overview diagram of a Scene Memory Gate disclosed in an embodiment of the present invention.
[0062] Figure 4 It is a detailed structure diagram of a scene memory gate disclosed in an embodiment of the present invention.
[0063] Figure 5 It is a comparison diagram of qualitative results with GSD-Occ disclosed in an embodiment of the present invention.
[0064] Figure 6 It is a qualitative result diagram on the Occ3D-NuScenes validation set disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0066] According to the embodiments of the present invention, it should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the following manufacturing method, in some cases, the steps shown or described can be executed in a different order than here.
[0067] Occupancy Prediction refers to a key perception task that provides a three-dimensional understanding of the surrounding environment for autonomous vehicles. It reconstructs the environment into a three-dimensional voxel representation with semantic labels by processing multi-view camera images, accurately identifying which regions in the space are occupied by objects and the categories of these objects. This representation form is more comprehensive than traditional two-dimensional perception or point clouds, can provide fine-grained scene understanding, and has advantages especially when dealing with occluded areas and distant objects, thus providing a more reliable environmental information basis for the vehicle's path planning and decision-making.
[0068] As Figure 1 (c) shows, the present invention provides an occupancy prediction method for autonomous driving, including the following steps:
[0069] Step S1, obtain a multi-view image sequence at the current moment, and extract multi-view image features through a two-dimensional convolutional network;
[0070] In this step, the autonomous vehicle obtains a multi-view image sequence at the current moment through a plurality of panoramic cameras mounted on it. The image can be expressed as I ∈ R B×T×N×3×h×w , where B represents the batch size, T represents the length of the time series, N represents the number of camera views, and h and w respectively represent the height and width resolutions of the image.
[0071] After obtaining the multi-view images, a two-dimensional convolutional neural network (CNN) is used as the backbone network to extract features from these images. This backbone network can be a commonly used image feature extraction network such as ResNet, and each view of the image is processed separately to output the corresponding two-dimensional image feature F. The two-dimensional image feature F contains visual information such as the texture, shape, and edges of the scene.
[0072] The two-dimensional convolutional network extracts features from the input image through multiple convolutional operations. Each layer of convolution can capture image features at different scales and levels of abstraction. Shallow layers capture low-level features such as edges and textures, while deeper layers capture high-level semantic features such as object shapes and category information.
[0073] Step S2: Convert the multi-view image features into three-dimensional voxel features and project them to generate the bird's-eye view features at the current moment;
[0074] In this step, the multi-view two-dimensional image features F extracted in step S1 are converted into voxel features V in three-dimensional space. The conversion can adopt the LSS (Lift-Splat-Shoot) method, which can effectively lift the two-dimensional image features into three-dimensional space.
[0075] Specifically, the multi-view two-dimensional image features F are first input into the depth network and the semantic network. The depth network is used to predict the depth distribution D of each view image, and the depth distribution D represents the possible depth values and their probabilities of each pixel point in the image. The semantic network further processes the multi-view two-dimensional image features F and outputs the semantically enhanced image features F'.
[0076] Next, an outer product operation is performed on the semantically enhanced image features F' and the depth distribution D to generate pseudo three-dimensional point cloud features. The outer product operation essentially "lifts" the two-dimensional features into three-dimensional space according to the predicted depth distribution to form a probabilistic point cloud representation. These point cloud features are then converted into regular three-dimensional voxel features V through voxel pooling operation. Voxel pooling divides the three-dimensional space into regular grids (voxels) and aggregates the point cloud features falling within the same voxel, thus forming a structured three-dimensional feature representation.
[0077] To reduce the computational complexity of subsequent processing, the three-dimensional voxel features V are projected onto the bird's-eye view (BEV) plane to generate the bird's-eye view features B at the current moment. The bird's-eye view is a two-dimensional representation looking down from above, which retains the distribution information of the three-dimensional scene on the horizontal plane and significantly reduces the amount of data to be processed.
[0078] It can be understood that the method of converting multi-view image features into three-dimensional voxel features includes:
[0079] Step S21: Use the depth network to predict the depth distribution of each view image;
[0080] Step S22: Extract image semantic features through the semantic network;
[0081] Step S23: Perform an outer product on the image semantic features and the depth distribution to generate pseudo 3D point cloud features;
[0082] Step S24: Convert the pseudo 3D point cloud features into three-dimensional voxel features through voxel pooling.
[0083] Through this step, the multi-view two-dimensional image features are successfully converted into a three-dimensional space representation, and the computational efficiency is improved through bird's-eye view projection.
[0084] Step S3: Input the bird's-eye view features at the current moment into the scene memory gate module, dynamically fuse them with the historical feature map, and generate an updated historical feature map;
[0085] In this step, the efficient fusion of the bird's-eye view features at the current moment and the historical features is achieved through the Scene Memory Gate (SMG) module. The recursive processing method is used for temporal feature fusion. Different from the traditional method that needs to store and process multiple frames of historical features, SMG only maintains and updates one historical feature map, greatly reducing the storage overhead and computational complexity.
[0086] Specifically, the bird's-eye view features B at the current moment T and the historical feature map h at the previous moment T-1 are input into the SMG module. Due to the movement of the vehicle, the historical feature map first needs to be spatially aligned according to the pose change of the vehicle to ensure consistency with the current coordinate system. Then, SMG converts the historical feature h T-1 into a hidden state S T-1 .
[0087] The core idea of SMG is to selectively retain the valuable information in the historical features while discarding the features from relatively far or highly uncertain regions. For this purpose, a convolutional layer and a sigmoid activation function are used to calculate the value score of each pixel in the historical features based on S T-1 . The same operation is also used to calculate the uncertainty score of the current bird's-eye view features B T .
[0088] Based on these two scores, a mask M T-1 is generated:
[0089] M T-1 = σ(f s (S T-1 )) + (1 - σ(f c (B T )))
[0090] where M T-1 represents the mask, f c (·) and f s(·) represents the convolution operation, which acts on the current BEV feature and the hidden state respectively. σ(·) represents the sigmoid activation function, and S T-1 represents the hidden state, which is generated by the historical scene feature through the hidden block, and B T represents the bird's-eye view feature at the current moment, which is obtained by converting the image feature. The first term σ(f s (S T-1 )) in the equation is used to retain the valuable information in the historical feature, while the second term (1 - σ(f c (B T ))) is used to enhance the weight of the historical feature in the highly uncertain regions (such as occlusion regions) in the current scene. This is because the current bird's-eye view feature B T generated based on the LSS method is usually sparse in the occlusion region, and introducing the historical feature can make up for this deficiency.
[0091] Next, the generated mask is used to filter out the valuable historical features:
[0092]
[0093] where, is the filtered historical feature map, which retains the valuable information; ⊙ is the element-wise multiplication (Hadamard product); h T-1 represents the historical scene feature map at time step T - 1, which is saved by the fusion result of the previous moment;
[0094] Finally, the filtered historical feature is added to the current bird's-eye view feature to generate the updated historical feature map:
[0095]
[0096] where, h T is the updated historical scene feature map, which is used for the iteration of the next moment; B T represents the bird's-eye view feature at the current moment, which is obtained by converting the image feature.
[0097] This fused feature map h T not only contains the current scene information, but also fuses the valuable features in the historical scene, especially the supplementary information for the current occlusion region. At the same time, h T will be saved as the historical feature map for the next moment, which is used for the subsequent recursive fusion process.
[0098] The recursive feature of SMG endows it with the characteristic of gradually forgetting long-term memory, which is particularly suitable for the autonomous driving scenario. Since vehicles usually move forward, the relevance of the features of early scenes to the reconstruction of the current scene will decrease over time. The long-term forgetting mechanism can automatically discard these early features that are no longer relevant, enabling the network to focus more on the current moment and recent temporal features.
[0099] It can be understood that the dynamic fusion operation of the scene memory gate includes:
[0100] Step S31: Convert the historical feature map into a hidden state through a hidden block;
[0101] Step S32: Calculate the value score of each pixel based on the hidden feature and the bird's-eye view feature at the current moment;
[0102] Step S33: Generate a mask according to the value score to filter out valuable historical features;
[0103] Step S34: Fuse the filtered historical feature map with the bird's-eye view feature at the current moment to generate an updated historical feature map.
[0104] Step S4: Convert the updated historical feature map into a three-dimensional semantic feature and fuse it with the three-dimensional voxel feature to generate an enhanced three-dimensional feature;
[0105] In this step, the updated historical feature map (bird's-eye view feature) obtained from Step S3 is converted into a three-dimensional semantic feature. This conversion process is implemented using the BVL (BEV-to-Volume Lifting) method. Specifically, the historical feature map is first decomposed into two components: geometric features and semantic features. Geometric features represent the spatial structure information of the scene, while semantic features contain the recognition information of object categories. Subsequently, these two types of features are fused through an outer product operation to restore the height information and generate a complete three-dimensional semantic feature representation.
[0106] This three-dimensional semantic feature is fused with the original three-dimensional voxel feature generated in Step S2. The fused feature is the enhanced three-dimensional feature, which contains both the direct visual information provided by the current image and the temporal semantic information accumulated in the historical feature map. This fusion mechanism enables the network to effectively handle occluded areas and uncertain areas in dynamic scenes, improving the accuracy and integrity of occupancy prediction. The enhanced three-dimensional feature will be used for geometric encoding and final occupancy prediction.
[0107] Step S5: Perform multi-scale geometric encoding processing on the enhanced three-dimensional feature to generate a three-dimensional geometric feature;
[0108] In this step, multi-scale geometric encoding is performed on the enhanced 3D features generated in step S4 to extract rich spatial geometric information. This process is implemented using a reparameterized 3D convolutional network, which features high efficiency and a large receptive field.
[0109] Specifically, the geometric encoder consists of two types of 3D convolutional layers: the first type uses non-dilated small convolutional kernels to extract local geometric features; the second type uses dilated small convolutional kernels to expand the receptive field through dilated convolution and capture context information over a larger range. This multi-scale structure can simultaneously focus on local details and the global spatial structure, improving the model's ability to understand complex 3D scenes.
[0110] During the training phase, these two types of convolutional layers operate independently and jointly learn the geometric features of the scene. In the inference phase, to improve computational efficiency, the system merges the parameters of these small convolutional kernels into a single equivalent large convolutional kernel, achieving parameter sharing and computational optimization through reparameterization techniques. This design significantly reduces the computational cost during inference while maintaining the model's expressive power. The resulting 3D geometric features contain important geometric information such as the spatial layout, shape, and boundaries of objects in the scene.
[0111] It can be understood that the multi-scale geometric encoding process using a reparameterized 3D convolutional network includes:
[0112] Step S51: During the training phase, use the first 3D convolutional layer with non-dilated small convolutional kernels to extract local geometric features;
[0113] Step S52: During the training phase, stack the second 3D convolutional layer, which uses dilated small convolutional kernels to expand the receptive field;
[0114] Step S53: During the inference phase, merge the parameters of the non-dilated small convolutional kernels and the dilated small convolutional kernels into a single equivalent large convolutional kernel to generate reparameterized 3D geometric features.
[0115] Step S6: Perform multi-level semantic encoding on the updated historical feature map to generate 2D semantic features;
[0116] In this step, a lightweight encoder is used to perform semantic encoding on the updated historical feature map to extract rich 2D semantic features. Specifically, first, spatial downsampling is performed on the historical feature map to generate high-resolution low-dimensional features and low-resolution high-dimensional features. Then, these two types of features are input into an encoding block containing a VRKWV module, which is processed through two paths: a spatial mixing branch and a channel mixing branch.
[0117] In the spatial mixing branch, a Q-shift operation is performed on the input features to generate spatial key features, spatial value features, and spatial gating weights. Then, a global spatial attention feature is calculated using a bidirectional weighted key-value attention mechanism, which has a linear complexity and can efficiently capture long-range dependencies in the sequence. After multiplying the calculated global attention feature by the spatial gating weights, a linear projection is performed to generate the output of the spatial branch.
[0118] Meanwhile, in the channel mixing branch, the input features are projected in channels to generate channel key features and channel gating weights, and the output feature of the channel branch is generated through the SquaredReLU activation function. Finally, the output features of the spatial branch and the channel branch are added together and aggregated through a convolutional layer to output the final two-dimensional semantic features. The design of the lightweight encoder greatly reduces the number of model parameters while maintaining a strong semantic feature extraction ability.
[0119] Step S7: Fuse the three-dimensional geometric features and the two-dimensional semantic features to output a three-dimensional occupancy prediction result.
[0120] In this step, it is necessary to effectively fuse the three-dimensional geometric features obtained in the previous steps with the two-dimensional semantic features to generate the final three-dimensional occupancy prediction result. First, the two-dimensional semantic features are restored to three-dimensional semantic features through a BEV-to-three-dimensional conversion method. During this conversion process, the two-dimensional semantic features are decomposed into semantic information and geometric information, and then they are recombined into a complete three-dimensional semantic feature representation through an outer product operation.
[0121] Next, the converted three-dimensional semantic features are added to the three-dimensional geometric features obtained in the previous steps on a voxel-by-voxel basis to achieve a deep fusion of semantic information and geometric information. This fusion method can retain the accuracy of the geometric structure while endowing the structure with rich semantic information.
[0122] Finally, the fused features are input into a prediction head network, which maps the features to a predefined semantic category space through a series of upsampling and classification operations. What the prediction head outputs is a voxel grid, where each voxel contains the probability of "occupied" and the corresponding semantic label. In this way, the three-dimensional occupancy prediction of the current scene is completed, providing detailed environmental perception information for autonomous driving vehicles.
[0123] Through the collaborative work of the above seven steps, this method can efficiently utilize temporal information and multi-modal features, and achieve accurate three-dimensional occupancy prediction while maintaining a low computational complexity, meeting the dual requirements of real-time and accuracy for autonomous driving.
[0124] Figure 2The overall architecture of ReOcc is shown, and the double-branch architecture of GSD-OCC is used as the baseline model. First, the backbone network extracts features from the input multi-view images and converts the 2D features into 3D voxel features through the LSS method. To reduce the computational cost, the 3D voxel features are further converted into Bird's Eye View (BEV) features for semantic reasoning.
[0125] During the temporal fusion process, a Scene Memory Gate (SMG) is introduced. This module aggregates the current BEV features with the historical feature maps to generate new fused features, which are stored as the historical feature maps for the next time step. At the same time, the BVL method is used to restore the fused BEV features to 3D features and combine them with the initial 3D voxel features to generate enhanced 3D voxel features.
[0126] Subsequently, the enhanced 3D voxel features are processed by a 3D geometric encoder and a 2D semantic encoder respectively to generate 3D geometric features and 2D semantic features. Finally, the BVL method is used to restore the 2D semantic features to 3D semantic features and combine them with the 3D geometric features to generate the final occupancy prediction result. In the entire network architecture, the 3D geometric encoder design of GSD-OCC is retained, and a new lightweight encoder is proposed to efficiently process 2D semantic features, thus significantly reducing the number of model parameters while ensuring performance.
[0127] Figure 3 The left part shows the iterative process of SMG, while the right part details its architecture. For clarity, the normalization layer is omitted.
[0128] Although the scene memory gate performs well in achieving temporal feature fusion, compared with convolution-based methods, it introduces additional computational cost and parameters under the premise of fusing two frames. Therefore, this embodiment designs a simpler and more efficient 2D encoder to work with SMG to balance computational speed, network scale, and performance.
[0129] The architecture of the lightweight encoder in this embodiment is as shown in Figure 4 the left. This lightweight encoder consists of only two encoding modules, and a downsampling layer is added to obtain high-dimensional semantic features. The high-dimensional features and low-dimensional features are aggregated through a single convolutional layer. For the encoding module, the purpose of the present invention is to seek a network structure that can both have strong reasoning ability and fast reasoning speed. Therefore, this embodiment tries models such as RWKV and Mamba, which have excellent reasoning ability and linear complexity. Finally, RWKV is adopted in the encoder module. The relevant experimental process will be elaborated in detail in the ablation experiment.
[0130] Although the scene memory gate achieves efficient temporal feature fusion, compared with convolution-based methods, it introduces additional computational complexity and parameters when fusing two frames of images. Therefore, in this embodiment, a simpler and more efficient two-dimensional encoder, namely the Light Encoder, is designed to work in cooperation with the SMG, so as to achieve a balance among computational speed, network size, and performance. The architecture of the Light Encoder is as shown in Figure 4 the left figure. This encoder consists of two encoding blocks. In this embodiment, a downsampling layer is inserted to obtain high-dimensional semantic features, and the high-dimensional features and low-dimensional features are merged through a single convolutional layer.
[0131] The objective of the present invention is to find an encoding block structure that has both powerful reasoning ability and fast reasoning speed. Therefore, VRWKV is adopted as the encoding block. The architecture of VRWKV is as shown in Figure 4 the right figure. The input is first fed into the spatial mixing module. After the Q-shift operation, the input is converted into r s , k s , v s , where r s is the spatial gating weight, k s is the key vector, and v s is the value vector. Then, the global attention result is obtained through the bidirectional attention mechanism, which has a linear complexity, and its calculation can be expressed as:
[0132]
[0133] where wkv t represents the output result of the bidirectional attention mechanism at position t, Bi-WKV is the bidirectional attention mechanism, K and V are the key and value input to the bidirectional attention mechanism, t represents the target position index of the attention weight to be calculated in the current feature sequence, i represents other positions in the feature sequence except the target position t, L represents the length of the sequence, w is a learnable scalar parameter that controls the attenuation intensity of adjacent position features, u is a learnable scalar parameter that enhances the weight of the current position, k i is the value of the key at position i, k t is the value of the key at position t, v i is the value of the value at position i, and v t is the value of the value at position t;
[0134] wkv t is multiplied by σ(r s ) and linearly projected to generate the output o s of the spatial mixing module:
[0135] o s= F l (σ(r s )) ⊙ wkv)
[0136] where, o s is the output feature of the Spatial Mixing Module; F l (·) represents a fully connected linear projection for transforming the feature representation; σ(·) represents the sigmoid activation function that controls the weight range of the features between (0, 1); ⊙ represents the Hadamard product (element-wise multiplication) used to combine the gating weights with the attention features in the calculation, and r s represents the spatial gating weight for controlling the feature information at different positions. The calculation process of the channel mixing module is similar, and the formula is expressed as:
[0137] v c = F l (SquaredReLU(k c ))
[0138] O c = F l (σ(r c ) ⊙ v c )
[0139] where, v c represents the intermediate feature of the Channel Mixing Module, which is used for the final calculation after non-linear transformation; SquaredReLU(·) represents the Squared Rectified Linear Unit, which is a non-linear activation function, and k c represents the channel key feature for calculating the feature interaction in the channel dimension; O c represents the final output feature of the Channel Mixing Module; σ(r c ) represents the channel gating weight, which is used to adjust the weight of the channel information after sigmoid activation.
[0140] To more clearly illustrate the effect of the method of the present invention, the following experiments will be conducted:
[0141] 1. Datasets and experimental settings:
[0142] This experiment was carried out based on the NuScenes dataset. As an authoritative benchmark in the field of autonomous driving, this dataset contains 1,000 driving scenarios (700 for training / 150 for validation) collected in the Boston and Singapore regions. Each scenario lasts for 20 seconds and covers diverse driving environments such as urban roads, complex weather, and day-night variations. The dataset provides multi-modal sensor data from 6 surround-view cameras, lidar, etc. Among them, the occupancy annotations are from the Occ3D-NuScenes benchmark. The scene range is [-40m, 40m] × [-40m, 40m] × [-1m, 5.4m], the voxel resolution is 0.4m, and it contains 16 types of objects and the "empty" class label.
[0143] 2. Adopt a two-dimensional evaluation metric:
[0144] The mIoU (mean intersection over union) and RayIoU (Ray-1level mIoU) were measured to evaluate the prediction performance of the model of the present invention. In addition, the number of model parameters, frames per second (FPS), runtime memory size, and storage size of time-series data were also measured.
[0145] 3. Implementation details:
[0146] ResNet-50 was used as the image backbone network, and the resolution of the input image was set to 704×256. Two versions of the model were proposed. The feature dimensions of ReOcc-B and ReOcc-S were set to 128 and 64 respectively. The model was trained for 24 epochs using 8 NVIDIA 4090 GPUs with a batch size of 3. The Adam optimizer was adopted, and the learning rate was set to 1×10 -4 , with a weight decay of 0.01. Consistent with previous studies, the inference speed of the model was tested on a single NVIDIA A100 GPU with a batch size of 1. The historical features were saved using the'savez_compressed' function of the NumPy package, and their storage sizes were measured. See Tables 1 and 2 for details:
[0147] Table 1 Comparison results of SOTA methods on the Occ3D-NuScenes validation set
[0148]
[0149] Table 1 Evaluation of the 3D occupancy prediction performance on the Occ3D-nuScenes dataset. If a metric is followed by "-", it means that no relevant data is provided for that metric.
[0150] Table 2 Fine-grained comparison of real-time models
[0151]
[0152] Table 2 presents a detailed comparison between the proposed method and state-of-the-art (SOTA) 3D occupancy prediction methods on the Occ3D-nuScenes dataset. If a metric is followed by "-", it means that the data for that metric was not provided; if it is followed by "*", it means that the metric was self-measured.
[0153] 4. Performance Comparison:
[0154] In Table 1, ReOcc was compared with the most recent state-of-the-art (SOTA) methods on Occ3D-nuScenes. When trained with visibility masks, ReOcc-B achieved an mIoU of 39.9 and a RayIoU of 31.5 on the nuScenes validation dataset, with an inference speed of 23.4 FPS. Without using masks, ReOcc-B had an mIoU of 31.9 and a RayIoU of 38.2. ReOcc-S with reduced feature dimensions had a 1.7% decrease in mIoU and a 0.3% decrease in RayIoU when using masks; without using masks, the mIoU decreased by 0.9% and the RayIoU decreased by 1.0%. However, compared to ReOcc-B, ReOcc-S had a 12.8% increase in FPS.
[0155] To demonstrate the efficiency of the proposed model in this embodiment, a more detailed comparison was made with other real-time inference models, as shown in Table 2. Under the condition of aggregating only one feature map, the proposed model was comparable to GSD-OCC (16f) and Panoptic-FlashOcc (8f) in terms of prediction performance, and had 1.8 mIoU and 4.2 RayIoU higher than SparseOcc (8f). In terms of inference speed, ReOcc-B and ReOcc-S were 6.1 FPS and 9.1 FPS higher than SparseOcc (8f), 3.4 FPS and 6.4 FPS higher than GSD-OCC (16f), and 2.4 FPS and 5.4 FPS higher than GSD-OCC (2f), respectively.
[0156] However, they were still about 10 FPS slower than the Panoptic-FlashOcc series. Although the inference speed of Panoptic-FlashOcc is very fast, it still adopts the Figure 1 temporal fusion structure shown in (a), where Figure 1 (a) shows the architecture adopted by many existing methods, where the backbone network extracts all temporal features and feeds these features into the fusion module for aggregation; Figure 1(b) shows the architecture adopted by GSD-OCC, which stores the extracted historical features in a queue and retrieves them from the queue before the temporal fusion process; Figure 1 (c) represents the method proposed by the present invention, which condenses historical features through scene memory gates and passes the condensed historical feature map to the next iteration process. The test of its reasoning speed does not include the time required to extract features from historical images. Therefore, Panoptic-FlashOcc still cannot avoid redundant calculations in the temporal fusion process during the training phase. The present invention adds the time for extracting historical features and remeasures its reasoning speed. As shown in Panoptic-FlashOcc(2f)*, Panoptic-FlashOcc(8f)*, its FPS drops significantly. This shows that the training efficiency of the model is not high. In contrast, the model of the present invention maintains a high reasoning speed during training, thereby saving more training resources. The model of the present invention has the least number of parameters while maintaining good performance and reasoning speed. The number of parameters of ReOcc-S is 30.5% lower than that of SparseOcc(8f), 51.4% lower than that of GSD-Occ(16f), and 27.3% lower than that of Panoptic-Occ(8f).
[0157] In addition, since both GSD-Occ and Panoptic-FlashOcc use long-term temporal feature fusion, they need to store multiple historical feature maps. Their sizes were measured and it was found that the temporal feature map size of GSD-Occ (16f) is 32.3M, while that of Panoptic-FlashOcc is 55.0M. Thanks to the advantages of the SMG structure, ReOcc only needs to store one feature map, which is about one-tenth of their size. The temporal feature map size of ReOcc-S is only 1.9M, even smaller than GSD-Occ (2f) and Panoptic-FlashOcc (2f).
[0158] Figure 5 Qualitative comparison shows that ReOcc can accurately reconstruct occluded areas (such as guardrail extension structures) in complex scenes and avoid the ghosting artifact problem of GSD-OCC, verifying the long-range modeling capability of the scene memory gate.
[0159] 5. Ablation analysis:
[0160] Ablation experiments are conducted on the nuScenes dataset, and all experiments are based on ReOcc-B and trained using visibility masks.
[0161] Ablation experiments of different components:
[0162] Table 3 shows the effectiveness of the Scene Memory Gate (SMG) and the lightweight encoder. The baseline is based on the GSD-Occ structure with 2-frame temporal fusion. After replacing the temporal fusion method of the baseline with SMG, there are only slight changes in the number of model parameters and the inference speed, but the mIoU increases by 7.0% and the RayIoU increases by 6.0%. This highlights the effectiveness of SMG in retaining fused historical features. After replacing the 2D encoder of the baseline with a lightweight encoder, the mIoU and RayIoU remain almost unchanged, and the number of parameters decreases by 16.8%, which proves the efficiency of the lightweight encoder. When integrating these two modules simultaneously, compared with the baseline, the inference speed of the model decreases slightly, but its performance improves by 8.4% in mIoU and 6.3% in RayIoU, and the number of parameters decreases by 16.5%.
[0163] Ablation experiments of different components in Table 3
[0164]
[0165] Ablation experiments of the lightweight encoder:
[0166] To find a more efficient encoding block network structure, the present invention conducted experiments with VRWKV, VMamba, and Resnet networks, and the results are shown in Table 4. Compared with ResNet, VMamba has a higher mIoU by 0.20, a lower RayIoU by 0.03, 0.8M fewer parameters, but a lower running speed by 1.5 FPS. While VRWKV has a higher mIoU by 0.28, a lower RayIoU by 0.3, 0.4M fewer parameters, and a lower running speed by 0.8 FPS. Considering that mIoU is the main evaluation metric, the present invention selects the VRWKV network as the basic encoding module of the lightweight encoder. The present invention also conducted ablation experiments on the number of encoding modules. When the number of modules increases from 1 to 2, the mIoU increases by 0.3 and the RayIoU increases by 0.03, but the FPS decreases by 0.8 and the number of parameters increases by 2.34M. When the number of modules increases from 2 to 3, the mIoU increases by 0.23 and the RayIoU increases by 0.12, but the FPS decreases by 1.4 and the number of parameters increases by 5.77M. To balance the performance, size, and inference speed of the model, this paper selects to set the number of encoding blocks to 2.
[0167] Ablation experiments of the lightweight encoder in Table 4
[0168]
[0169] Ablation experiments of the Scene Memory Gate:
[0170] To further verify the effectiveness and efficiency of the Scene Memory Gate Temporal Fusion method compared to traditional methods, the present invention compared it with the convolutional fusion method with 2-frame and 16-frame temporal inputs. The results are shown in Table 5. Although the inference speed of SMG is slightly lower than that of the convolutional method, it is 0.89 higher than the convolutional method (16f) in terms of mIoU, 0.41 higher in terms of RayIoU, and only requires 3.53M of storage space for temporal features. This indicates that SMG has a more efficient temporal feature fusion ability compared to the fusion method with a longer time series.
[0171] Table 5 Ablation Experiment of Scene Memory Gate
[0172]
[0173] Table 6 details the IoU performance of each category. ReOcc has significant improvements in the detection of small objects such as "traffic cones" (+3.2%) and the reconstruction of complex structures such as "vegetation" (+2.8%). Figure 6 The visualization results show that the model maintains stable perception performance in challenging scenarios such as at night, in rain and fog.
[0174] Table 6 IoU Performance of Each Category
[0175]
[0176] This embodiment proposes the RecurrentOcc framework, which realizes the efficient distillation and dynamic fusion of historical features through an innovative scene memory gate. Combined with the lightweight encoder design, it achieves an SOTA performance of 39.9 mIoU with 59.1M parameters on the Occ3D-NuScenes benchmark, and a real-time inference speed of 23.4 FPS meets the stringent deployment requirements of autonomous driving. The ablation experiment confirms that the SMG module reduces the storage consumption by 91% compared to the traditional temporal fusion method, and the lightweight encoder realizes the co-optimization of parameter efficiency and inference speed.
[0177] According to another aspect of the embodiments of the present application, there is also provided an occupancy prediction system for autonomous driving, including the following modules:
[0178] A multi-view image acquisition module for acquiring a multi-view image sequence at the current moment;
[0179] A two-dimensional feature extraction module for extracting features from the multi-view image based on a two-dimensional convolutional network to generate multi-view image features;
[0180] A three-dimensional voxel projection module for converting the multi-view image features into three-dimensional voxel features and projecting them to generate the bird's-eye view features at the current moment;
[0181] A scene memory gate module for generating an updated historical feature map by dynamically fusing the current bird's-eye view features and the historical feature map;
[0182] A three-dimensional semantic conversion module for converting the updated historical feature map into three-dimensional semantic features and fusing them with the three-dimensional voxel features to generate enhanced three-dimensional features;
[0183] A multi-scale geometric encoding module for performing multi-scale geometric structure encoding on the enhanced three-dimensional features to generate three-dimensional geometric features;
[0184] A multi-level semantic encoding module for extracting multi-level semantic features from the historical feature map to generate two-dimensional semantic features;
[0185] A cross-dimensional fusion module for fusing the three-dimensional geometric features and the two-dimensional semantic features to output a three-dimensional occupancy prediction result.
[0186] The "Recurrent Occupancy Prediction Network (RecurrentOcc)" proposed by the present invention significantly improves the computational efficiency and performance of the 3D occupancy prediction task through an innovative temporal fusion paradigm and an efficient network structure design, and solves three main problems existing in traditional methods. First, there is a redundant calculation problem in traditional temporal image feature extraction methods, that is, historical image features are repeatedly extracted, which not only increases the computational overhead but also leads to a decline in the performance of the model. To solve this problem, RecurrentOcc eliminates redundant calculations and reduces the computational burden of the model by introducing the concept of "historical feature map" and condensing the features of multiple temporal images into a single feature map.
[0187] Secondly, there is a problem of redundant fusion in traditional temporal feature fusion methods. When fusing the features of multiple temporal images, some features will be repeatedly aggregated, resulting in low computational efficiency. Through the design of the "Scene Memory Gate" (SMG), valuable information in the historical scene features is selectively retained, avoiding interference from invalid or outdated information. SMG recursively updates the historical feature map and only fuses one historical feature map with the current scene features, thus avoiding redundant temporal feature fusion and at the same time enhancing the expression ability of temporal features.
[0188] Finally, existing long-term historical feature fusion methods often ignore the significant changes in scene information over long time intervals, resulting in inappropriate fusion of features at different timestamps, which in turn affects the accuracy of prediction. The Scene Memory Gate solves this problem by adaptively and selectively fusing historical feature maps and current features. By considering the correlation and spatio-temporal consistency of historical information, SMG intelligently discards no-longer-useful information, ensuring that the network pays more attention to the current moment and recent temporal features during fusion, thus effectively avoiding the interference of early timestamp information.
[0189] In addition to the innovation in temporal fusion, a lightweight encoder architecture was designed to reduce the number of model parameters and improve the inference speed. By introducing the RWKV module with linear complexity, this lightweight encoder significantly reduces the computational complexity while maintaining high inference ability and accuracy. Finally, the experimental results on the Occ3D-NuScenes dataset show that although the model has only 59.1M parameters, the inference speed reaches 23.4FPS, and at the same time, the best performance of 39.9 mIoU is achieved.
[0190] In summary, the present invention successfully solves the problems of redundant calculation, redundant fusion, and rough fusion of historical features in traditional methods by proposing a new temporal fusion paradigm and an efficient encoder design. It not only improves the performance of occupancy prediction but also significantly reduces the computational amount of the model, ensuring the requirements of real-time inference, and has broad application prospects, especially in fields such as autonomous driving that have high requirements for real-time performance and computational efficiency.
[0191] According to another aspect of the embodiments of the present application, an electronic device is also provided, including a processor and a memory. The processor is used to implement the steps of the method when executing the computer program stored in the memory.
[0192] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0193] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of units or modules can be in an electrical or other form.
[0194] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0195] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0196] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An occupancy prediction method for autonomous driving, characterized in that, It includes the following steps: Obtain the multi-view image sequence at the current moment, and extract the multi-view image features in the multi-view image sequence through a two-dimensional convolutional network; Convert the multi-view image features into three-dimensional voxel features, and project to generate the bird's-eye view feature at the current moment; Input the bird's-eye view feature at the current moment into the scene memory gate module, and perform dynamic fusion with the historical feature map to generate an updated historical feature map; Convert the updated historical feature map into three-dimensional semantic features, and fuse with the three-dimensional voxel features to generate enhanced three-dimensional features; Perform multi-scale geometric encoding processing on the enhanced three-dimensional features to generate three-dimensional geometric features; Perform multi-level semantic encoding processing on the updated historical feature map to generate two-dimensional semantic features; Fuse the three-dimensional geometric features and the two-dimensional semantic features, and output the three-dimensional occupancy prediction result.
2. The occupancy prediction method for autonomous driving according to claim 1, characterized in that, The dynamic fusion operation of the scene memory gate module includes: Convert the historical feature map into hidden features through a hidden block; Calculate the value score of each pixel based on the hidden features and the bird's-eye view feature at the current moment; Generate a mask according to the value score, and filter out the valuable historical feature maps; Fuse the filtered valuable historical feature maps with the bird's-eye view feature at the current moment to generate an updated historical feature map.
3. The occupancy prediction method for autonomous driving according to claim 1, wherein The method for converting the multi-view image features into three-dimensional voxel features includes: Use a depth network to predict the depth distribution of each multi-view image feature; Extract the image semantic features through a semantic network; Perform an outer product of the image semantic features and the depth distribution to generate pseudo 3D point cloud features; Convert the pseudo 3D point cloud features into three-dimensional voxel features through voxel pooling.
4. The occupancy prediction method for autonomous driving according to claim 1, wherein The method for converting the updated historical feature map into three-dimensional semantic features includes: Decompose the updated historical feature map into geometric features and semantic features; Fuse the geometric features and the semantic features through an outer product operation to generate three-dimensional semantic features.
5. The occupancy prediction method for autonomous driving according to claim 1, characterized in that The multi-scale geometric encoding processing uses a reparameterized three-dimensional convolutional network, including: In the training stage, extract local geometric features through the first three-dimensional convolutional layer using a non-dilated small convolutional kernel; In the training stage, stack the second three-dimensional convolutional layer, and the second three-dimensional convolutional layer uses a dilated small convolutional kernel to expand the receptive field; In the inference stage, merge the parameters of the non-dilated small convolutional kernel and the dilated small convolutional kernel into a single equivalent large convolutional kernel to generate the reparameterized three-dimensional geometric features.
6. The occupancy prediction method for autonomous driving according to claim 1, characterized in that, The method for multi-level semantic encoding processing includes: Step S61: Perform a spatial downsampling operation on the updated historical feature map to generate a high-resolution low-dimensional feature and a low-resolution high-dimensional feature; Step S62: Input the high-resolution low-dimensional feature and the low-resolution high-dimensional feature into an encoding block containing a variable receptive field weighted key-value module, and perform cross-level fusion through the following sub-steps: Step S621: In the spatial mixing branch, perform a Q-shift operation on the input feature to generate a spatial key feature, a spatial value feature, and a spatial gating weight; Step S622: Based on the spatial key feature and the spatial value feature, calculate the global spatial attention feature through bidirectional weighted key-value attention, and its formula is: Among them, wkv t represents the output result of the bidirectional attention mechanism at position t. Bi-WKV is the bidirectional attention mechanism, K and V are the key-value inputs to the bidirectional attention mechanism, t represents the target position index of the attention weight to be calculated in the current feature sequence, i represents the other position indexes in the feature sequence except the target position t, L represents the length of the sequence, w is a learnable scalar parameter that controls the attenuation intensity of adjacent position features, u is a learnable scalar parameter that enhances the weight of the current position, k i is the value of the key at position i, k t is the value of the key at position t, v i is the value of the value at position i, v t is the value of the value at position t; Step S623: Multiply the global spatial attention feature element-wise with the spatial gating weight, and generate the spatial branch output feature through linear projection; Step S624: In the channel mixing branch, perform channel projection on the input feature to generate the channel key feature and the channel gating weight, and generate the channel branch output feature through the SquaredReLU activation function; Step S63: Add the spatial branch output feature and the channel branch output feature to obtain the fused cross-level semantic feature; Step S64: Perform convolutional aggregation on the cross-level semantic feature to output the two-dimensional semantic feature.
7. The occupancy prediction method for autonomous driving according to claim 1, characterized in that, The method for fusing the three-dimensional geometric feature and the two-dimensional semantic feature and outputting the three-dimensional occupancy prediction result includes: Performing voxel-wise addition of the three-dimensional geometric feature and the transformed three-dimensional semantic feature; Performing upsampling and classification on the added feature through a prediction head to output the final three-dimensional occupancy prediction result.
8. An occupancy prediction system for autonomous driving, characterized in that, Including the following modules: A multi-view image acquisition module for acquiring a multi-view image sequence at the current moment; A two-dimensional feature extraction module for extracting features from the multi-view image based on a two-dimensional convolutional network to generate multi-view image features; A three-dimensional voxel projection module for converting the multi-view image features into three-dimensional voxel features and projecting to generate the bird's-eye view feature at the current moment; A scene memory gate module for generating an updated historical feature map by dynamically fusing the current bird's-eye view feature and the historical feature map; A three-dimensional semantic conversion module for converting the updated historical feature map into a three-dimensional semantic feature and fusing it with the three-dimensional voxel feature to generate an enhanced three-dimensional feature; A multi-scale geometric encoding module for performing multi-scale geometric structure encoding on the enhanced three-dimensional feature to generate a three-dimensional geometric feature; A multi-level semantic encoding module for extracting multi-level semantic features from the historical feature map to generate two-dimensional semantic features; A cross-dimensional fusion module for fusing the three-dimensional geometric feature and the two-dimensional semantic feature to output the three-dimensional occupancy prediction result.
9. An electronic device, characterized in that, Including a processor and a memory, the processor is used to implement the steps of the occupancy prediction method for autonomous driving according to any one of claims 1 to 7 when executing the computer program stored in the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it executes the steps of the occupancy prediction method for autonomous driving according to any one of claims 1 to 7.
Citation Information
Cited By
End-to-end automatic driving method based on linear time complexity
CN120763872A
BEV space construction method, automatic driving system, equipment and medium
CN121074319A
Semantic occupancy prediction method and system based on bidirectional multi-modal residual fusion
CN121304982A
Enhanced semantic occupancy dataset construction method for nuScenes dataset
CN122618383B