Grid occupation prediction method combining city-level neural radiation field prior and time domain enhancement
By combining the city-level neural radiation field prior and time domain enhancement occupancy raster prediction methods, the accuracy and timeliness of image recovery and three-dimensional object detection under severe weather conditions are solved, stable image recovery and high-precision occupancy prediction are achieved, and the perception ability of the autonomous driving system is enhanced.
Patent Information
- Application Number
- CN202510537118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art has problems of insufficient accuracy and timeliness in image recovery and three-dimensional object detection under severe weather conditions, and it is difficult to effectively use time information and multimodal data to accurately predict occupation.
Combining the occupancy grid prediction method of urban-level neural radiation field prior and time domain enhancement, by forming urban-level neural radiation field, a prior three-dimensional voxel characteristics are extracted, the weather image recovery network is used to process real-time images and convert them into three-dimensional voxel characteristics, and the multi-frame voxel characteristics are fused for occupancy prediction.
Provide stable image recovery and occupation prediction under severe weather conditions, enhance the perceptual robustness of the autonomous driving system, build high-precision vector maps in real time, and improve the perception ability and adaptability to complex environments.
Smart Images

Figure CN120426982A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving perception technology, and specifically relates to an occupancy grid prediction method that combines city-level neural radiation field priors with time domain enhancement. Background Art
[0002] Deepening semantic understanding in the process of accurately building vector maps can lay a solid foundation for autonomous driving technology, ensuring its effectiveness and credibility in the real world. With technological advancements, it is becoming increasingly important for autonomous driving systems to demonstrate strong environmental recognition and response capabilities in changing traffic conditions. A deep understanding of the road environment, particularly the ability to accurately build semantic vector maps, is crucial to the successful deployment and smooth operation of autonomous driving technology.
[0003] For multi-degraded image restoration, deep learning models, particularly convolutional neural networks, can learn the statistical properties of natural images from large-scale data, thereby implicitly acquiring prior knowledge about the image. Multi-stage architectures progressively restore images through multiple stages, with each stage optimizing for specific degradation types or features. Some network architectures that do not require nonlinear activation functions simplify the model structure by replacing or removing these functions while maintaining or improving performance. Cross-stage feature fusion mechanisms help propagate multi-scale contextual features from earlier stages to later stages, enhancing feature richness. While these techniques employ a common architecture to address diverse degradation problems, most require the manual selection of different pre-trained models to address specific degradation scenarios. This approach is unsuitable for practical applications, as the specific degradation patterns are often unpredictable. Furthermore, while it is possible to optimize these general image restoration techniques holistically, they often overlook induced biases caused by adverse weather conditions, such as texture degradation, color distortion, and contrast reduction caused by scattering from atmospheric particles. These issues create occlusion artifacts, limiting their performance in adverse weather conditions.
[0004] In terms of temporal modeling in 3D object detection, early 3D object detection methods mainly focused on processing single-frame images or point clouds, lacking consideration of the temporal dimension. As research deepened, researchers began to explore the role of temporal information in 3D object detection. Models such as BEVFormer began to attempt to incorporate temporal information into the model. Furthermore, some technologies further improved the accuracy of 3D object detection by combining multi-view, multimodal data, and temporal information. Self-supervised learning methods were proposed to explore methods for learning 3D occupancy using only video sequences, reducing dependence on large amounts of labeled data. Time series analysis technology enhances the analysis capabilities of time series data through sparse object queries and adaptive spatiotemporal sampling modules. Although temporal modeling technology in 3D object detection has made significant progress, it still faces challenges in effectively utilizing temporal information, handling dynamic environments, reducing computational complexity, fusing multimodal data, and improving generalization and real-time performance.
[0005] Among the technologies related to 3D occupancy prediction, point-based methods process point cloud data to predict 3D occupancy; camera-based methods use video sequences or monocular camera data to learn 3D occupancy; and multimodal data fusion combines LiDAR and camera data to improve the accuracy of occupancy prediction. Although these 3D occupancy prediction methods have achieved promising results, challenges remain in handling large-scale sparse data, anisotropy, spatial support, time series data, and multimodal data fusion. Therefore, how to effectively utilize temporal information to improve prediction accuracy and robustness is a potential improvement direction in current research.
[0006] It can be seen that there is an urgent need to provide a time-domain enhanced occupancy prediction method for multiple degraded images and three-dimensional space that can simultaneously meet the requirements of accuracy and timeliness. Summary of the Invention
[0007] The present invention is proposed based on the above-mentioned needs of the prior art. The technical problem to be solved by the present invention is to provide an occupancy grid prediction method that combines city-level neural radiation field prior and time domain enhancement to perform accurate and real-time scene prediction based on multiple degraded images.
[0008] In order to solve the above problems, the technical solutions provided by the present invention include:
[0009] Provided is an occupancy grid prediction method that combines city-level neural radiation field priors with time-domain enhancement, comprising: forming a city-level neural radiation field based on a historical data set; extracting prior three-dimensional voxel features from the city-level neural radiation field using a ray marching algorithm; processing real-time vehicle images based on a trained weather image restoration network to obtain high-visual-quality images, and converting them into three-dimensional voxel features; fusing three-dimensional voxel features with prior three-dimensional voxel features to obtain enhanced voxel features; processing the enhanced voxel features of multiple consecutive frames through a trained occupancy prediction network, and obtaining an occupancy grid map of the predicted area after occupancy prediction; wherein the weather image restoration network comprises: a first feature embedding layer, wherein the input image is subjected to the first feature embedding layer to obtain image features F i ; The first branch, the second branch and the third branch are connected in parallel; the first branch will F i Average pooling and maximum pooling are performed separately for local pixel level processing, and the two results are connected and processed by sigmoid function to obtain The second branch will be F i Perform vertical and horizontal strip pooling for strip-level processing respectively, and fuse the two results to perform sigmoid function processing to obtain The third branch is F i Perform convolution processing to obtain The convolutional layer with residual connection will and Connect them and combine with F i Generate multi-scale attention features; the global distribution-level attention mechanism module divides the multi-scale attention features into two features, processes one feature through the normalization layer, processes the other feature through the convolution layer, and integrates the two processed results; the second feature embedding layer, the integration result is passed through the second feature embedding layer to form a high visual quality image.
[0010] Preferably, it is characterized in that the historical data set includes first data, second data and third data, the first data represents the video collected when the vehicle was previously driving on the city road, the second data represents the position of the camera, and the third data represents the camera direction; the city-level neural radiation field is formed based on the historical data set, including: pixel-level enhancement of the RGB sequence of each frame image in the first data to form a video identifier; based on the video identifier and the corresponding second data and third data, the feature data of each point in the short-range scene is obtained, and the feature data includes density, color and semantic features; based on the video identifier and the third data, the feature data of each point in the sky scene is obtained; based on the feature data of each point in the process scene and the sky scene, the information of each point in the three-dimensional space is obtained, thereby forming a city-level neural radiation field.
[0011] Preferably, the method of obtaining the feature data of each point in the close-range scene based on the video identifier, the corresponding second data and the third data includes: the input data is the position x of the camera in the three-dimensional space i , camera direction d i , video identifier vid i , the output data is the density of each point in the three-dimensional space of the scene color Semantic feature f i sur : Among them, the characteristics of the point The density of the sum point is expressed as: H(·) represents multi-resolution hash grid, MLP(·) represents linear perceptron; the predicted color of the point Expressed as: Φ color (·) represents the MLP output layer for predicting color, γ(d i ) represents the direction corresponding to d i Spherical harmonic coding, V(vid i ) represents the video identifier corresponding to vid i Cross-video embedding of illumination conditions; semantic features of points f i sur Expressed as: Φ feat (·) represents the MLP output layer for predicting semantic features.
[0012] Preferably, obtaining the characteristic data of each point in the sky scene based on the video identifier and the third data includes: the input data is the camera direction d i and the video identifier vid i , the output data is the color of each point in the sky scene and semantic features f i sky , expressed as: where Φ sky (·) represents the MLP output layer for predicting the sky.
[0013] Preferably, the prior 3D voxel features are extracted in the city-level neural radiation field by a ray marching algorithm, including: for each ray, the ray is moved along d j Direction sampling M points p m , used to determine the intersection point p of the ray with the surface of the object in the neural radiation field n , by finding the cumulative transmittance α m and opacity T m The first point exceeding the threshold determines the surface point. The threshold is 0.5, which is expressed as: Among them, α m represents the cumulative transmittance, T m Indicates per-segment opacity.
[0014] Preferably, the training process of the occupancy prediction network includes: obtaining a training set; inputting the multi-view image of each frame in the training set into a multi-weather image restoration network, and obtaining a high-visual-quality image without degradation effects after processing; extracting features of the two-dimensional high-visual-quality image with an image encoder, and then converting the two-dimensional features into three-dimensional voxel features; fusing the obtained voxel features with the prior voxel features of the urban-level neural radiation field to obtain enhanced voxel features; performing a first process of time-domain enhancement and a second process of occupancy prediction on the enhanced voxel features, and obtaining first information after the first process; obtaining second information after the second process; and obtaining the final occupancy result of the lost frame based on the first information and the second information.
[0015] Preferably, the second process includes randomly discarding the features of a certain frame, and reconstructing the features of the lost frame based on the voxel features other than the lost frame through the first decoder to obtain a first pseudo feature; reconstructing the features of the lost frame based on the features of two adjacent frames of the lost frame through the second decoder to obtain a second pseudo feature; the first pseudo feature and the second pseudo feature are used as the second information, and the second information is input into the occupancy prediction head shared with the first process to predict the final occupancy result of the lost frame.
[0016] Preferably, the processing of the first branch includes: i Two types of feature maps are generated using average pooling and maximum pooling operations along the channel axis. and And connected, the spatial attention feature is generated through the convolution layer of the sigmoid function, which is expressed as: Conc() represents the connection operation, and Conv() represents the convolution operation.
[0017] Preferably, the processing of the second branch includes pooling the feature F by strip pooling operations in both horizontal and vertical directions. m,n,c Project to the horizontal and vertical directions to obtain horizontal features and vertical features Expressed as: Horizontal and vertical features are fused separately through convolutional layers and replicated to expand the generated features. and Expressed as: Among them, Conv 1×3 () is a 1×3 convolutional layer, Conv 3×1() is a 3×1 convolution layer, H is the height, W is the width, m, n, c represent the values of height, width and number of channels respectively, E() is the expansion operation; by adding and fusing two features, a convolution operation with a sigmoid function is performed and combined with the original feature F i Multiply element by element to generate features Expressed as:
[0018] Preferably, the processing performed by the convolutional layer with residual connection is expressed as: It is the output result of the convolution layer with residual connection, Conv() represents the convolution operation, and Conc() represents the connection operation.
[0019] Preferably, the processing of the global distribution level attention mechanism module includes: Split along the channel dimension through the convolution layer into and Processed by instance normalization layer Processed by convolutional layers Finally, the integrated features are obtained Expressed as: Among them, IN() is the normalization function, S() is the Split segmentation operation, Conv() represents the convolution operation, and Conc() represents the connection operation.
[0020] Compared with existing technologies, this invention can provide stable image restoration and occupancy prediction under adverse weather conditions, enhancing the perceptual robustness of autonomous driving. It also constructs high-precision vector maps in real time, improving the autonomous driving system's precise understanding of the road environment. Furthermore, by integrating city-level prior information and temporal enhancement technology, this invention provides a robust occupancy grid full-scene understanding method that operates stably under various adverse weather conditions. This method can process multiple degraded images and accurately and real-timely predict occupancy in three dimensions using prior information and temporal information, significantly improving the perception and adaptability of autonomous driving systems in complex environments. Furthermore, this method can construct high-precision vector maps in real time and provide accurate occupancy predictions in dynamic scenes. By integrating multi-view images, city-level static prior information, and temporal information, this invention enhances the in-depth understanding of the road environment, providing a solid foundation for the successful deployment and smooth operation of autonomous driving technology. In summary, this invention provides a robust occupancy grid full-scene understanding method that combines city-level neural radiance field priors with temporal enhancement. This method can handle multiple degraded images and accurately and real-timely predict occupancy in 3D space using prior information and temporal information. Furthermore, neural radiance field and occupancy grid prediction belong to two distinct technical fields, with significant differences in their implementation. Neural radiance field is a high-fidelity 3D reconstruction technology based on implicit neural representations in computer graphics, dedicated to achieving realistic new perspective synthesis through differentiable rendering. Occupancy grid prediction is a method for dynamic obstacle prediction in the field of autonomous driving perception. Neural radiance field models continuous radiance fields, pursuing sub-pixel optical accuracy. It relies on dense perspective input and long-term optimization, constructing an implicit representation of the 3D spatial information of static scenes. Occupancy grids use discrete voxels as units, perform single-frame inference, emphasize real-time performance and robustness with dynamic updates, and explicitly represent occupancy prediction probabilities. It is precisely this comprehensive gap, from theoretical foundations to application scenarios, that makes the integration of neural radiance field into occupancy grid prediction a highly challenging cross-disciplinary endeavor. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0022] Figure 1 This is a flowchart of the steps of an occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement in an embodiment of the present invention;
[0023] Figure 2 Schematic diagram of a weather image restoration network architecture in an embodiment of the present invention;
[0024] Figure 3 A flowchart of the data processing process and prior data acquisition process of the city-level neural radiation field in an embodiment of the present invention;
[0025] Figure 4 This is a schematic diagram of the architecture of the prediction network training process in an embodiment of the present invention;
[0026] Figure 5 Schematic diagram of the architecture of the occupancy prediction network prediction process in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0028] In the description of the embodiments of the present invention, it should be noted that, unless otherwise specified or limited, the term "connected" should be understood in a broad sense. For example, it can mean a fixed connection, a detachable connection, or an integral connection. It can be a mechanical connection, an electrical connection, a direct connection, or an indirect connection through an intermediate medium. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0029] The terms "top," "bottom," "above," "below," and "on" used throughout the description refer to relative positions of components of a device, such as the relative positions of top and bottom substrates within a device. It will be understood that devices are multifunctional regardless of their orientation in space.
[0030] To facilitate understanding of the embodiments of the present invention, specific embodiments will be further explained below with reference to the accompanying drawings. The embodiments do not limit the embodiments of the present invention.
[0031] This embodiment provides an occupancy grid prediction method that combines city-level neural radiation field priors with time domain enhancement, such as Figure 1-Figure 5 shown.
[0032] The method comprises:
[0033] Get historical datasets.
[0034] The historical dataset includes videos and camera placement data collected while the vehicle was previously traveling on urban roads, where the camera placement data includes the camera's position and camera angle. The videos collected while the vehicle was previously traveling on urban roads serve as the first data, the camera's position serves as the second data, and the camera angle serves as the third data.
[0035] Forming a city-level neural radiation field based on historical data sets, including:
[0036] Perform pixel-level enhancement on the RGB sequence of each frame of the image in the first data to form a video identifier.
[0037] The first data is obtained from an RGB image sequence, which is a series of color images arranged in time order. Each image contains information about the red, green, and blue color channels of the scene. A pre-trained vision model is used to perform pixel-level enhancement on the RGB image sequence to generate a video identifier.
[0038] Specifically, a deep inverse nonlinear optimization model is used to capture high-level semantic features of image content. These features help understand the types and attributes of different objects in the scene. A Transformer-based image segmentation model is used to generate high-quality semantic segmentation masks that clearly distinguish different objects from the background in the image. These semantic features and semantic segmentation masks together form the video identifier corresponding to the RGB image, providing a deeper understanding of the objects in the scene, which makes the subsequently constructed 3D scene representation richer and more accurate.
[0039] When building a city-level neural radiation field, we want the model to focus more on static features in the scene (such as buildings, roads, and trees). These static features are particularly important for understanding the long-term structure of the scene. Therefore, we use image segmentation masks generated by a Transformer-based image segmentation model to identify and exclude pixels belonging to moving objects (such as pedestrians and vehicles), allowing the model to focus more on static features.
[0040] Divide the city into several blocks, each of which covers an area of 1km 2 , each tile is independently optimized using the video and camera pose data within that tile, forming smaller, independent sub-radiance fields. Because the scene and sky areas near the camera are subject to different lighting conditions, each sub-radiance field is constructed for two scenes: the near-field scene and the sky scene. A multi-resolution hash grid and linear perceptron are used to encode the implicit spatial features of the near-field and sky scenes for each tile.
[0041] Feature data of each point in the short-range scene is obtained based on the video identifier, the corresponding second data, and the third data, where the feature data includes density, color, and semantic features.
[0042] Specifically: the input data is the position of the camera in three-dimensional space χ i , camera direction d i , video identifier vid i , the output data is the density of each point in the three-dimensional space of the scene color Semantic feature f i sur .
[0043] Among them, the characteristics of the point and the density of the points Expressed as:
[0044]
[0045] H(·) stands for multi-resolution hashing grid and MLP(·) stands for linear perceptron.
[0046] Predicted color of the point Expressed as:
[0047]
[0048] Φ color (·) represents the MLP output layer for predicting color, γ(d i ) represents the direction corresponding to d i Spherical harmonic coding, V(vid i ) represents the video identifier corresponding to vid i Cross-video embedding of lighting conditions.
[0049] The semantic feature f of the point i sur Expressed as:
[0050]
[0051] Φ feat (·) represents the MLP output layer for predicting semantic features.
[0052] In summary, the input and output relationship of the short-range scenario in each section is:
[0053]
[0054] Feature data of each point in the sky scene is obtained based on the video identifier and the third data.
[0055] Since the point information is independent of the camera position, the input data is the camera direction d i and the video identifier vid i , the output data is the color of each point in the sky scene and semantic features f i sky .
[0056]
[0057] where Φ sky (·) represents the MLP output layer for predicting the sky.
[0058] Based on the feature data of each point in the process scene and the sky scene, the information of each point in the three-dimensional space is obtained, thereby forming a city-level neural radiation field.
[0059] After the above processing, all sub-radiation fields of the city-level neural radiation field are constructed, that is, the information of each point in the three-dimensional space of the short-range scene and the sky scene of each plate of the city (including density, color, and semantic features). This information will be stored in the hash grid and multi-layer perceptron.
[0060] A priori three-dimensional voxel features are extracted in the city-level neural radiation field through a ray marching algorithm, including: sampling points along the ray direction, determining the surface points between the ray and the objects in the city-level neural radiation field by accumulating transmittance and opacity, and matching the voxel features in the city-level neural radiation field based on the surface points as the priori three-dimensional voxel features.
[0061] The algorithm works by sending rays from the training cameras and sampling multiple points along each ray to identify occupied voxels and capture their associated features.
[0062] Specifically include:
[0063] For each ray, the algorithm follows the ray d j Direction sampling M points p m These points are used to determine the intersection point p of the ray with the surface of the object in the neural radiation field n By finding the cumulative transmittance α m and opacity T m The first point that exceeds the threshold is used to determine the surface point. The threshold used here is 0.5, which means that when the cumulative opacity along the ray exceeds 50%, it is considered that the surface has been reached, expressed as:
[0064]
[0065] Among them, α m represents the cumulative transmittance, T m Indicates per-segment opacity.
[0066] After identifying the object's surface points, semantic features are matched against them using the hash grid of the neural radiance field and a multilayer perceptron. Surface points collected from all training views are downsampled on a voxel basis to minimize the number of points, and features are averaged within each voxel. This ensures that each voxel incorporates feature information from multiple views, forming a rich prior.
[0067] The above process includes: constructing static prior information of the city based on the video and camera placement posture data (camera position, camera angle, etc.) collected when the vehicle was driving on the city road, and applying it to a variety of cutting-edge real-time perception models to improve perception accuracy. A city-level neural radiation field construction module is designed, which inputs the previous video and camera placement posture data to generate a three-dimensional continuous volume representation at the city level, which contains the color and semantic feature information of each point in the three-dimensional space. After rendering, the pixel-level color and semantic information of the image under the new two-dimensional perspective can be obtained. A prior data extraction module is designed to convert the unstructured prior information in the city-level neural radiation field into structured information for use by other online perception models. A prior fusion and integration module is designed to apply the prior information to other online perception models.
[0068] Get real-time images of vehicles.
[0069] Real-time images are processed based on the trained weather image restoration network to obtain high visual quality images and convert them into three-dimensional voxel features.
[0070] The weather image restoration network is a convolutional neural network-based encoder-decoder network that exploits general knowledge covering multiple degradation types derived from inductive biases under various severe weather conditions.
[0071] Pre-training the weather image restoration network involves inputting images from a training set with degradation effects caused by severe weather conditions such as rain, snow, fog, and haze into the weather image restoration network. The input images are sequentially processed through two first feature embedding layers, multiple multi-scale hierarchical attention modules, and two second feature embedding layers to obtain images with the degradation effects removed. The difference between the images with the degradation effects removed and the images without the degradation effects is then compared, and the weather image restoration network is debugged by minimizing a loss function. This allows the weather image restoration network to process high-visual-quality images after inputting them into the real-time network.
[0072] The first and second feature embedding layers are both convolutional layers with three residual blocks, acting as encoders and decoders, respectively. The first feature embedding layer performs feature extraction, feature space compression, and feature enhancement, acting as an encoder. The second feature embedding layer performs feature refinement and upscaling, feature integration, and image reconstruction, acting as a decoder.
[0073] The input image of the weather image restoration network is processed by the initial feature embedding layer to obtain F i , F i ∈R H×W×C , where H, W, and C represent the height, width, and number of channels respectively. After passing through the convolutional layer, it is processed by three parallel branches F i , and obtain different image features.
[0074] The multi-scale hierarchical attention module can handle occlusion and scattering artifacts caused by various weather conditions. The constructed weather image restoration network efficiently addresses image restoration problems in various adverse weather conditions, enhancing the robustness of the entire model. Specifically, the multi-scale hierarchical attention module consists of a first branch, a second branch, and a third branch connected in parallel, a convolutional layer with a residual connection connected in series with the three branches, and a global hierarchical attention mechanism module.
[0075] The multi-scale hierarchical attention module consists of three parts: local pixel-level attention, global strip-level attention, and global distribution-level attention, capable of simultaneously handling occlusion and scattering artifacts. The local pixel-level attention mechanism corresponds to the first branch and focuses on capturing local spatial features to effectively handle occlusions caused by short-range degradation patterns. The global strip-level attention mechanism corresponds to the second branch and extracts global spatial features through horizontal and vertical strip pooling operations to effectively handle long-range degradation patterns of various directions and sizes. The global distribution-level attention mechanism corresponds to the global hierarchical attention mechanism module, which aims to capture changes in the distribution of atmospheric particles and adaptively adjust the feature distribution of the degraded image through instance normalization to address color distortion and contrast reduction caused by atmospheric particle scattering.
[0076] The first branch, after convolution operation on the image, captures local spatial features through local pixel-level attention mechanism to deal with short-range degradation. This mechanism uses average pooling and maximum pooling operations along the channel axis based on feature F i Generate two types of feature maps and And connect them together and use a convolution layer with a sigmoid function to generate spatial attention features Expressed as: Among them, Conv() represents the convolution operation, and Conc() represents the connection operation.
[0077] The second branch extracts global spatial features through a global strip-level attention mechanism after convolution operation on the image. Since occlusions such as rain and snow may cause image degradation in different directions, this mechanism uses strip pooling operations in both horizontal and vertical directions to pool the features F. m,n,c Projecting to the horizontal and vertical directions is helpful for dealing with long-distance degradation in various directions. The image features in the horizontal and vertical directions are recorded as and Then 1×3 and 3×1 convolutional layers are used to fuse horizontal and vertical features, and a copy operation is used to expand their sizes to generate related features. and m, n, c represent the values of height, width and number of channels respectively. Finally, the two features are fused by addition, and a convolution operation with a sigmoid function is performed and combined with the original feature F i Multiply element by element to generate features
[0078]
[0079] Among them, Conv 1×3 () is a 1×3 convolutional layer, Conv 3×1 () is a 3×1 convolutional layer, H is the height, W is the width, m, n, c represent the values of height, width and number of channels respectively, and E() is the expansion operation.
[0080] The third branch directly performs convolution operations on the image to obtain image features.
[0081]
[0082] Use convolutional layers with residual connections to Concatenate and add the original input features F i , generating multi-scale attention features The purpose of this step is to fuse features of different scales to better handle non-uniform degradation patterns, that is, the inconsistent degradation phenomena that occur in images under bad weather conditions.
[0083]
[0084] A global distribution-level attention mechanism is used to deal with the color distortion and contrast attenuation problems caused by atmospheric particle scattering. This mechanism is designed to capture the changes in the distribution of atmospheric particles, which is crucial for restoring the color and contrast of images under adverse weather conditions. The mechanism first inputs F i Split along the channel dimension through the convolution layer into and Processed by the instance normalization layer (IN) Processed by convolutional layers Finally, the integrated features are obtained
[0085]
[0086] Among them, S() is the Split operation.
[0087] Feature Map After processing through the last two feature embedding layers, a high visual quality image is obtained after removing the degradation effect.
[0088] Apply the Charbonnier loss function L char Measures the difference between the restored image O output by the model in the spatial domain and the real image G. By minimizing this loss, the model is encouraged to produce outputs that are closer to the real image at the pixel level.
[0089]
[0090] Here, ∈ is a small positive number used to control the smoothness and robustness of the loss function.
[0091] Apply the Fourier transform loss function L FFT Comparing the difference between the restored image O and the real image G in the frequency domain helps the model capture the characteristics of the real image in terms of frequency components, which is particularly effective when dealing with degradations such as blur that have obvious characteristics in the frequency domain.
[0092] L FFT =‖F(O)-F(G)‖1
[0093] The total loss is L char and L FFT The weighted sum of is used to optimize the model as a whole. By adjusting the values of parameters α and β, the influence of Charbonnier loss and Fourier transform loss on model training can be balanced to achieve good image restoration effect in both spatial and frequency domains. The total loss is expressed as: L = α1L char +β1L FF , where α1 and β1 are L char and L FF The corresponding weight parameter.
[0094] An image editor is used to extract a two-dimensional high visual quality image and convert the two-dimensional high visual quality image into a three-dimensional voxel feature.
[0095] An image encoder is used to extract features from 2D high visual quality images, which are then converted into 3D voxel features through a 2D-3D viewpoint conversion module.
[0096] The real-time image of the vehicle is input into the trained weather image restoration network for processing, and also passes through two first feature embedding layers, the parallel first branch, the second branch and the third branch, the convolution layer with residual connection, the global distribution level attention mechanism module and the two feature embedding layers.
[0097] The 3D voxel features and the prior voxel features are fused to obtain enhanced voxel features.
[0098] The enhanced voxel features of multiple consecutive frames are processed by the trained occupancy prediction network, and after occupancy prediction, the occupancy grid map of the predicted area is obtained.
[0099] The training process of the occupancy prediction network includes:
[0100] Get the training set.
[0101] The multi-view images of each frame in the training set are input into the multi-weather image restoration network, and high visual quality images without degradation effects are obtained after processing.
[0102] An image encoder is used to extract features from 2D images with high visual quality, and then a 2D-3D viewpoint conversion module is used to convert the 2D features into voxel features.
[0103] The obtained voxel features are fused with the prior voxel features of the city-level neural radiation field to obtain enhanced voxel features.
[0104] The enhanced voxel features undergo a first process of temporal enhancement and a second process of occupancy prediction. The first process yields first information, followed by a second process of occupancy prediction. The final occupancy result for the lost frame is obtained based on the first and second information. This trains the model's ability to handle dynamic scenes and temporal changes.
[0105] Specifically, the first process includes temporal fusion and fusion features. The enhanced voxel feature is represented as F t-N ,…,F t-k-1 ,F t-k ,F t-k+1 ,…,F t By integrating temporal information through temporal fusion, the model can capture and understand the changes of the scene over time, enhance the model's understanding of dynamic scenes, and more accurately predict the position and movement of objects. The integrated information, i.e., the first information, is input into the occupancy prediction head shared with the second process to predict the final occupancy result.
[0106] The second process includes randomly discarding the features of a certain frame, and reconstructing the features of the lost frame based on the voxel features other than the lost frame through the first decoder to obtain a first pseudo feature; reconstructing the features of the lost frame based on the features of two adjacent frames of the lost frame through the second decoder to obtain a second pseudo feature; the first pseudo feature and the second pseudo feature are used as second information, and the second information is input into the occupancy prediction head shared with the first process to predict the final occupancy result of the lost frame, thereby training the model's ability to handle dynamic scenes and time changes.
[0107] The second process is to enhance the voxel features F t-N ,…,F t-k-1 ,F t-k ,F t-k+1 ,…,F t , randomly discard the feature F of a certain frame t-k A first decoder is designed based on the other voxel features F except the lost frame. t-N ,…,F t-k-1 ,F t-k+1 ,…,F t Reconstruct the features of the lost frame to obtain pseudo features Perceive the long-term motion information of the object; design a second decoder based on the features F of the two adjacent frames of the lost frame t-k-1 ,F t-k+1 Reconstruct the features of the lost frame to obtain pseudo features To create a detailed spatial representation. and The input is the occupancy prediction head shared with the occupancy prediction branch, which predicts the final occupancy result of the lost frame and trains the model's ability to handle dynamic scenes and time changes.
[0108] The first decoder undergoes three downsampling steps, each of which generates three-dimensional voxel features at different scales. These features contain different levels of spatial information. Smaller-scale feature maps contain more detailed information, while larger-scale feature maps contain more extensive contextual information. These three different-scale features are upsampled to the same scale and fused using a three-dimensional convolutional layer. By fusing these features, the model can simultaneously utilize both detailed and contextual information to capture the characteristics of objects of different sizes. The short-term temporal decoder consists of a three-dimensional convolutional layer and a ReLU activation function.
[0109] The total loss function during the training phase is composed of the loss L occupied by the prediction branch occ And the corresponding loss L for the long-term and short-term time decoders in the temporal enhancement branch l and L s Composition, expressed as:
[0110] L=α2L occ+β2L l +γL s
[0111] Among them, α2, β2 and γ are L occ 、L l and L s The corresponding weight parameter.
[0112] The present invention provides a robust occupancy grid full-scene construction method that can work stably under a variety of severe weather conditions by combining city-level neural radiation field priors and time domain enhancement technology. This method can process multiple degraded images and use prior information and time information in three-dimensional space to accurately and in real time predict occupancy, significantly improving the perception and adaptability of autonomous driving systems in complex environments. In addition, it can construct high-precision vector maps in real time and provide accurate occupancy predictions in dynamic scenes. By fusing multi-perspective images, city-level static prior information, and time domain information, the present invention enhances the ability to deeply understand the road environment and provides a solid foundation for the successful deployment and smooth operation of autonomous driving technology.
[0113] In summary, the present invention provides an occupancy grid prediction method that combines city-level neural radiation field priors with time domain enhancement, which can process multiple degraded images and perform occupancy prediction accurately and in real time in three-dimensional space using prior information and time information.
[0114] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for occupancy grid prediction combining city-level neural radiation field prior and time domain enhancement, characterized in that: include: Forming a city-level neural radiation field based on historical data sets; Extracting prior 3D voxel features in city-level neural radiation fields through ray marching algorithms; Process the real-time image of the vehicle based on the trained weather image restoration network to obtain high-visual-quality images and convert them into 3D voxel features; Fusing 3D voxel features and prior 3D voxel features to obtain enhanced voxel features; The enhanced voxel features of multiple consecutive frames are processed by the trained occupancy prediction network, and after occupancy prediction, an occupancy grid map of the predicted area is obtained; Wherein, the weather image restoration network includes: The first feature embedding layer, the input image passes through the first feature embedding layer to obtain the image feature F i ; The first branch, the second branch and the third branch are connected in parallel; the first branch is F i Average pooling and maximum pooling are performed separately for local pixel level processing, and the two results are connected and processed by sigmoid function to obtain The second branch will be F i Perform vertical and horizontal strip pooling for strip-level processing respectively, and fuse the two results to perform sigmoid function processing to obtain The third branch is F i Perform convolution to get F i C ; The convolutional layer with residual connection will and Connect them and combine with F i Generate multi-scale attention features; The global distribution-level attention mechanism module splits the multi-scale attention feature into two features, processes one feature through a normalization layer, processes the other feature through a convolutional layer, and integrates the two processed results; The second feature embedding layer, the integration results are passed through the second feature embedding layer to form a high visual quality image.
2. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 1 is characterized in that: The historical data set includes first data, second data, and third data, wherein the first data represents a video previously collected when the vehicle was traveling on an urban road, the second data represents a position of a camera, and the third data represents a camera direction; The method of forming a city-level neural radiation field based on a historical data set includes: performing pixel-level enhancement on the RGB sequence of each frame image in the first data to form a video identifier; obtaining feature data of each point in the short-range scene based on the video identifier and the corresponding second data and third data, the feature data including density, color and semantic features; obtaining feature data of each point in the sky scene based on the video identifier and the third data; obtaining information of each point in the three-dimensional space based on the feature data of each point in the process scene and the sky scene, thereby forming a city-level neural radiation field.
3. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 2 is characterized in that: The method of obtaining feature data of each point in the close-range scene based on the video identifier, the corresponding second data and the third data includes: the input data is the position x of the camera in the three-dimensional space i , camera direction d i , video identifier vid i , the output data is the density of each point in the three-dimensional space of the scene color Semantic feature f i sur : Among them, the characteristics of the point The density of the sum point is expressed as: H(·) stands for multiresolution hashing grid, and MLP(·) stands for linear perceptron; Predicted color of the point Expressed as: Φ color (·) represents the MLP output layer for predicting color, γ(d i ) represents the direction corresponding to d i Spherical harmonic coding, V(vid i ) represents the video identifier corresponding to vid i Cross-video embedding of lighting conditions; The semantic feature f of the point i sur Expressed as: Φ feat (·) represents the MLP output layer for predicting semantic features.
4. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 2 is characterized in that: The characteristic data of each point in the sky scene is obtained based on the video identifier and the third data, including: the input data is the camera direction d i and the video identifier vid i , the output data is the color of each point in the sky scene and semantic features f i sky , expressed as: where Φ sky (·) represents the MLP output layer for predicting the sky.
5. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 1 is characterized in that: The prior 3D voxel features are extracted in the city-level neural radiation field by ray marching algorithm, including: for each ray, the ray is moved along d j Direction sampling M points p m , used to determine the intersection point p of the ray with the surface of the object in the neural radiation field n , by finding the cumulative transmittance α m and opacity T m The first point exceeding the threshold determines the surface point. The threshold is 0.5, which is expressed as: Among them, α m represents the cumulative transmittance, T m Indicates per-segment opacity.
6. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 1 is characterized in that: The training process of the occupancy prediction network includes: Get the training set; The multi-view images of each frame in the training set are input into the multi-weather image restoration network, and high-visual-quality images without degradation effects are obtained after processing; Use an image encoder to extract features from a 2D high-visual-quality image and then convert the 2D features into 3D voxel features; The obtained voxel features are fused with the prior voxel features of the city-level neural radiation field to obtain enhanced voxel features; A first process of temporal enhancement and a second process of occupancy prediction are performed on the enhanced voxel features, first information is obtained after the first process; second information is obtained after the second process; and a final occupancy result of the lost frame is obtained based on the first information and the second information.
7. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 6 is characterized in that: The second process includes randomly discarding the features of a certain frame, and reconstructing the features of the lost frame based on the voxel features other than the lost frame through the first decoder to obtain a first pseudo feature; reconstructing the features of the lost frame based on the features of two adjacent frames of the lost frame through the second decoder to obtain a second pseudo feature; the first pseudo feature and the second pseudo feature are used as second information, and the second information is input into the occupancy prediction head shared with the first process to predict the final occupancy result of the lost frame.
8. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 1 is characterized in that: The processing of the first branch includes: i Two types of feature maps are generated using average pooling and maximum pooling operations along the channel axis. and And connected, the spatial attention feature is generated through the convolution layer of the sigmoid function, which is expressed as: Conc() represents the connection operation, and Conv() represents the convolution operation.
9. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 1 is characterized in that: The processing of the second branch includes: pooling the feature F through strip pooling operations in both horizontal and vertical directions. m,n,c Project to the horizontal and vertical directions to obtain horizontal features and vertical features Expressed as: Horizontal and vertical features are fused separately through convolutional layers and replicated to expand the generated features. and Expressed as: Among them, Conv 1×3 () is a 1×3 convolutional layer, Conv 3×1 () is a 3×1 convolutional layer, H is the height, W is the width, m, n, c represent the values of height, width and number of channels respectively, and E() is the expansion operation; By adding and fusing the two features, a convolution operation with a sigmoid function is performed and combined with the original feature F i Multiply element by element to generate features Expressed as:
10. The occupancy grid prediction method combining city-level neural radiation field prior and time domain enhancement according to claim 1 is characterized in that: The processing performed by the convolutional layer with residual connection is expressed as: It is the output result of the convolution layer with residual connection, Conv() represents the convolution operation, and Conc() represents the connection operation.
Citation Information
Cited By
Model training method, indoor scene occupation prediction method, equipment and medium
CN120635679A