Vehicle scene reconstruction method and device, electronic equipment and storage medium

By identifying the severity of severe weather, dynamically adjusting the number of millimeter-wave radar data fusion frames, and performing radar data enhancement and image de-degradation processing, the problem of data quality degradation in lidar perception under severe weather conditions was solved, achieving stable and accurate 3D scene reconstruction.

CN122089933APending Publication Date: 2026-05-26CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD
Filing Date
2025-12-24
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing environmental perception solutions based on lidar suffer from severe degradation of raw perception data quality under adverse weather conditions such as fog and rain due to the limitations of their physical mechanisms. This results in insufficient accuracy and robustness in environmental perception and 3D scene reconstruction.

Method used

By acquiring image data, millimeter-wave radar data, and inertial measurement unit data, the severity of severe weather is identified, the number of millimeter-wave radar data fusion frames is dynamically adjusted, radar data enhancement processing and image de-degradation processing are performed, enhanced images and enhanced radar data are generated, and finally, three-dimensional scene reconstruction data of the vehicle's surrounding environment is generated.

Benefits of technology

It outputs stable and accurate 3D scene reconstruction results under severe weather conditions, ensuring high quality and strong complementarity of input information, and improving the system's robustness and accuracy under severe weather conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089933A_ABST
    Figure CN122089933A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle scene reconstruction method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring image data acquired by a vehicle, multi-frame millimeter wave radar data and inertial measurement unit data; when it is recognized that the weather state of the environment where the vehicle is located is severe weather, the severity level of the severe weather is determined; according to the severity level, performing fusion processing on the multiple frames of millimeter wave radar data to obtain accumulated radar data; performing enhancement processing on the accumulated radar data to generate enhanced radar data; performing dedegradation processing on the image data according to the enhanced radar data to generate an enhanced image; and generating three-dimensional scene reconstruction data of the surrounding environment of the vehicle according to the enhanced image, the enhanced radar data and the inertial measurement unit data. Therefore, a stable and accurate three-dimensional scene reconstruction result can be output in severe weather.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent driving technology, and in particular to a vehicle scene reconstruction method, device, electronic device and storage medium. Background Technology

[0002] With the development of autonomous driving technology, the ability of vehicles to accurately perceive their surroundings and reconstruct 3D scenes has become crucial for achieving advanced driver assistance systems (ADAS). In this field, environmental perception systems based on multi-sensor fusion, especially those integrating cameras and LiDAR, have become the mainstream technological approach. These systems achieve multi-dimensional complementary perception of the environment by aligning data from different sensors at different timestamps.

[0003] Existing technical solutions mainly rely on the fusion of cameras and LiDAR. Specifically, the system captures rich texture and color information through cameras, while using LiDAR to obtain precise spatial depth information; the two types of data are fused through timestamp alignment technology, and finally, a fusion algorithm is used to complete the 3D reconstruction and understanding of the environment.

[0004] However, the aforementioned environmental perception scheme based on lidar suffers from severe degradation of the quality of the original perception data under adverse weather conditions such as fog and rain due to the limitations of its physical mechanism. This ultimately leads to a serious lack of accuracy and robustness in subsequent environmental perception and 3D scene reconstruction. Summary of the Invention

[0005] This application provides a vehicle scene reconstruction method, device, electronic device, and storage medium to solve the problem that in the existing environmental perception scheme based on LiDAR, the quality of the original perception data is severely degraded due to the limitations of its physical mechanism in adverse weather conditions such as fog and rain, which ultimately leads to a serious lack of accuracy and robustness in subsequent environmental perception and 3D scene reconstruction.

[0006] Firstly, this application provides a vehicle scene reconstruction method, including: Acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle; If the weather conditions of the environment in which the vehicle is located are identified as severe weather, the severity level of the severe weather shall be determined. Based on the severity level, multiple frames of the millimeter-wave radar data are fused to obtain cumulative radar data. The accumulated radar data is enhanced to generate enhanced radar data; The image data is de-degraded based on the enhanced radar data to generate an enhanced image; Based on the enhanced image, the enhanced radar data, and the inertial measurement unit data, three-dimensional scene reconstruction data of the vehicle's surrounding environment is generated.

[0007] In one possible implementation, determining the severity level of the severe weather includes: Determine the weather type of the severe weather; Determine the assessment strategy based on the weather type; The severity level of the severe weather is obtained by analyzing the image data based on the assessment strategy.

[0008] In one possible implementation, the weather type is foggy, and the analysis of the image data based on the assessment strategy to obtain the severity level of the severe weather includes: Estimate the transmittance map based on the image data; From the image data and the transmittance map, multiple physical features for characterizing fog concentration are extracted; A weighted summation operation is performed on multiple physical characteristics to obtain a comprehensive score for fog concentration. The severity level is determined based on the predefined threshold range in which the comprehensive fog concentration score falls.

[0009] In one possible implementation, the step of fusing multiple frames of millimeter-wave radar data according to the severity level to obtain accumulated radar data includes: Based on the severity level, the total number of frames of millimeter-wave radar data to be fused is determined, wherein the higher the severity level, the more frames are fused. The millimeter-wave radar data from multiple frames is timestamped to obtain a millimeter-wave radar data sequence. Based on the total number of frames, select consecutive multiple frames of data with the current frame as the core from the millimeter-wave radar data sequence; The accumulated radar data is obtained by fusing the selected consecutive frames of data.

[0010] In one possible implementation, the enhancement processing of the accumulated radar data to generate enhanced radar data includes: The accumulated radar data is converted into a radar depth image; Using the radar depth image as a guiding condition, a pre-trained generative model is used to iteratively denoise from a random noise state to generate a target depth image with a higher spatial point cloud density than the radar depth image. The target depth image is converted into three-dimensional point cloud data, and the three-dimensional point cloud data is used as the enhanced radar data.

[0011] In one possible implementation, the step of performing de-degradation processing on the image data based on the enhanced radar data to generate an enhanced image includes: Spatially align the enhanced radar data with the image data; Using the aligned enhanced radar data as a condition, the image data is iteratively denoised using a pre-trained image generation model to generate the enhanced image.

[0012] In one possible implementation, generating three-dimensional scene reconstruction data of the vehicle's surrounding environment based on the enhanced image, the enhanced radar data, and the inertial measurement unit data includes: The inertial measurement unit data is pre-integrated to obtain vehicle motion information; Joint state estimation is performed on the enhanced image, the enhanced radar data, and the vehicle motion information to obtain an estimation result that includes the vehicle pose state and surrounding environment data. Based on the estimation results, the 3D scene reconstruction data is constructed.

[0013] Secondly, this application provides a vehicle scene reconstruction device, comprising: The acquisition module is used to acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle. The determination module is used to determine the severity level of the severe weather when the weather conditions of the environment in which the vehicle is located are identified as severe weather. The fusion module is used to fuse multiple frames of millimeter-wave radar data according to the severity level to obtain cumulative radar data. An enhancement module is used to enhance the accumulated radar data to generate enhanced radar data; The processing module is used to perform de-degradation processing on the image data based on the enhanced radar data to generate an enhanced image; The generation module is used to generate three-dimensional scene reconstruction data of the vehicle's surrounding environment based on the enhanced image, the enhanced radar data, and the inertial measurement unit data.

[0014] In one possible implementation, the determining module is specifically used for: Determine the weather type of the severe weather; Determine the assessment strategy based on the weather type; The severity level of the severe weather is obtained by analyzing the image data based on the assessment strategy.

[0015] In one possible implementation, the weather type is foggy, and the determining module is further configured to: Estimate the transmittance map based on the image data; From the image data and the transmittance map, multiple physical features for characterizing fog concentration are extracted; A weighted summation operation is performed on multiple physical characteristics to obtain a comprehensive score for fog concentration. The severity level is determined based on the predefined threshold range in which the comprehensive fog concentration score falls.

[0016] In one possible implementation, the fusion module is specifically used for: Based on the severity level, the total number of frames of millimeter-wave radar data to be fused is determined, wherein the higher the severity level, the more frames are fused. The millimeter-wave radar data from multiple frames is timestamped to obtain a millimeter-wave radar data sequence. Based on the total number of frames, select consecutive multiple frames of data with the current frame as the core from the millimeter-wave radar data sequence; The accumulated radar data is obtained by fusing the selected consecutive frames of data.

[0017] In one possible implementation, the enhancement module is specifically used for: The accumulated radar data is converted into a radar depth image; Using the radar depth image as a guiding condition, a pre-trained generative model is used to iteratively denoise from a random noise state to generate a target depth image with a higher spatial point cloud density than the radar depth image. The target depth image is converted into three-dimensional point cloud data, and the three-dimensional point cloud data is used as the enhanced radar data.

[0018] In one possible implementation, the processing module is specifically used for: Spatially align the enhanced radar data with the image data; Using the aligned enhanced radar data as a condition, the image data is iteratively denoised using a pre-trained image generation model to generate the enhanced image.

[0019] In one possible implementation, the generation module is specifically used for: The inertial measurement unit data is pre-integrated to obtain vehicle motion information; Joint state estimation is performed on the enhanced image, the enhanced radar data, and the vehicle motion information to obtain an estimation result that includes the vehicle pose state and surrounding environment data. Based on the estimation results, the 3D scene reconstruction data is constructed.

[0020] Thirdly, this application provides an apparatus comprising: a processor and a memory, the processor being configured to execute a vehicle scene reconstruction program stored in the memory to implement the vehicle scene reconstruction method described in any one of the first aspects.

[0021] Fourthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the vehicle scene reconstruction method described in any one aspect.

[0022] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application firstly acquires millimeter-wave radar data and image data that are less affected by weather, ensuring the availability and stability of basic data from the source of sensing; then, according to the identified severe weather level, the number of fusion frames of millimeter-wave radar data is dynamically adjusted, thereby specifically compensating for the sparsity of millimeter-wave data that may be caused by severe weather; then, by enhancing the accumulated radar data and using the enhanced radar data obtained after enhancement as a condition to perform de-degradation processing on the image data, the depth recovery and improvement of data quality are completed in both radar and visual modalities, respectively. Among them, radar data enhancement makes up for its inherent density deficiency, while cross-modal guided image de-degradation ensures the availability of visual information under severe conditions; finally, the fused and enhanced multi-source data is reconstructed, ensuring the high quality and strong complementarity of the input information itself, thereby outputting stable and accurate three-dimensional scene reconstruction results even under severe weather conditions. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0026] Figure 1A flowchart illustrating an embodiment of a vehicle scene reconstruction method provided in this application; Figure 2 A flowchart illustrating an embodiment of another vehicle scene reconstruction method provided in this application; Figure 3 A flowchart illustrating another embodiment of a vehicle scene reconstruction method provided in this application; Figure 4 A flowchart illustrating another embodiment of the vehicle scene reconstruction method provided in this application; Figure 5 A flowchart for foggy environment perception and classification is provided in an embodiment of this application; Figure 6 A block diagram illustrating an embodiment of a vehicle scene reconstruction device provided in this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0029] To address the technical problem that existing environmental perception solutions based on lidar suffer from severe degradation of raw perception data quality under adverse weather conditions such as fog and rain due to the limitations of their physical mechanisms, ultimately leading to insufficient accuracy and robustness in subsequent environmental perception and 3D scene reconstruction, this application provides a vehicle scene reconstruction method that can output stable and accurate 3D scene reconstruction results even under adverse weather conditions.

[0030] Figure 1 This is a flowchart illustrating an embodiment of a vehicle scene reconstruction method provided in this application. Figure 1 As shown, the method includes the following steps: Step 101: Acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle.

[0031] Image data: refers to a two-dimensional pixel array containing red, green, and blue color information collected by an in-vehicle camera.

[0032] Millimeter-wave radar data refers to a point cloud sequence collected by radar sensors operating in the millimeter-wave band, containing target range, radial velocity, azimuth angle, and reflection intensity.

[0033] Inertial Measurement Unit (IMU) data refers to time-series data collected by the IMU, which includes information on the angular velocity and acceleration of the vehicle in the three axes.

[0034] In this embodiment, the system synchronously triggers and receives raw data streams from cameras, millimeter-wave radar, and IMUs. These data streams are all precisely timestamped and are collected and cached via the vehicle bus, providing a synchronous and aligned multimodal data foundation for subsequent fusion processing.

[0035] Step 102: If the weather conditions of the environment in which the vehicle is located are identified as severe weather, determine the severity level of the severe weather.

[0036] Weather conditions: refers to the weather category of the vehicle's environment, such as sunny, foggy, rainy, or sandstorm.

[0037] Severity level: refers to the quantitative classification of the impact of severe weather (such as fog) on ​​the sensor's sensing ability. For example, it can be divided into light fog, dense fog and heavy fog, or light rain, moderate rain and heavy rain, etc.

[0038] In this embodiment, firstly, the weather type (e.g., fog, rain, or sandstorm) of the current weather state is identified based on image data. When severe weather is identified, a corresponding assessment strategy is selected and executed based on the specific weather type to determine its severity level. For example, for different weather types such as fog and rain, assessment strategies based on different physical models or feature systems can be used for quantitative analysis. This type-based differentiated assessment provides a more accurate and targeted adaptive control basis for subsequent processing.

[0039] In applications, the vehicle network module can also call the online weather application programming interface service to directly obtain real-time weather information of the area where the vehicle is located, which can be used as a basis for judging the weather condition and severity.

[0040] Step 103: Based on the severity level, fuse multiple frames of millimeter-wave radar data to obtain cumulative radar data.

[0041] Fusion processing: refers to the process of integrating multiple data frames that describe the same scene and were collected at different times into a single frame of more complete data through spatiotemporal alignment and merging.

[0042] Cumulative radar data: refers to the radar data set with improved point cloud density obtained after multi-frame fusion processing.

[0043] In this embodiment, a corresponding data fusion strategy is determined based on the severity level, and the timestamp-aligned multi-frame millimeter-wave radar data is processed based on this strategy. Specifically, by adjusting the number of data frames participating in the fusion, targeted compensation for data quality under different weather severity conditions is achieved, ultimately obtaining cumulative radar data with improved information density.

[0044] Step 104: Enhance the accumulated radar data to generate enhanced radar data.

[0045] Enhancement processing: refers to improving the quality, integrity, or information density of data through algorithms; in this step, it specifically refers to densifying the radar point cloud.

[0046] Enhanced radar data: refers to a three-dimensional point cloud obtained after enhancement processing, which has a much higher spatial point cloud density than the original data.

[0047] In this embodiment, the accumulated radar data is first converted from a 3D point cloud format to a 2D radar depth image via spherical coordinate projection. Then, this radar depth image is input into a pre-trained diffusion model for densification. The model is trained as follows: acquiring paired millimeter-wave radar depth images and lidar depth images; adding noise to the lidar depth image through a forward pass; and training a conditional denoising network that predicts the added noise using the millimeter-wave radar depth image and noise time steps as input. During inference, the model uses the radar depth image as a guiding condition, iteratively performing backward denoising starting from random noise to generate a high-resolution target depth image. Finally, this image is back-projected back into a 3D point cloud, yielding the enhanced radar data.

[0048] In applications, the enhancement process can be implemented using other generative methods, such as variational autoencoders or normalized flow models, to densify and enhance millimeter-wave radar data.

[0049] Step 105: Perform de-degradation processing on the image data based on the enhanced radar data to generate an enhanced image.

[0050] De-degradation: refers to reversing the quality degradation process of an image caused by external environment (such as fog or rain), restoring its sharpness, contrast and color fidelity.

[0051] Enhanced image: refers to an image whose visual quality is significantly restored after de-degradation processing.

[0052] In this embodiment, the enhanced radar data is first projected onto the image coordinate system using the extrinsic parameter matrices of the camera and radar to achieve spatial alignment. Then, the aligned enhanced radar data is used as conditional information and input into another pre-trained image diffusion model. This model is trained by: acquiring paired foggy and clear images; adding noise to the clear image through a forward pass; and training a conditional denoising network that predicts noise based on the aligned radar data and the noise time step. During inference, the model iteratively denoises the degraded foggy image using the aligned enhanced radar data as a condition, gradually removing weather degradation effects and ultimately generating an enhanced image with restored clarity.

[0053] In applications, the de-degradation processing can also be implemented using image restoration algorithms based on physical models, such as dehazing methods based on dark channel priors or color lines. These methods directly perform physical inversion operations on the degraded image by establishing the inverse process of an atmospheric scattering model, thereby recovering a clear image.

[0054] Step 106: Generate three-dimensional scene reconstruction data of the vehicle's surrounding environment based on the enhanced image, the enhanced radar data, and the inertial measurement unit data.

[0055] 3D scene reconstruction data: refers to a digital model that can represent the three-dimensional geometric structure and semantic information of the environment surrounding a vehicle, usually in the form of point cloud, mesh or voxel map.

[0056] In this embodiment, the system processes inertial measurement unit data to obtain vehicle motion information, and then deeply fuses the processed enhanced image, enhanced radar data, and vehicle motion information to generate 3D scene reconstruction data of the vehicle's surrounding environment. In a preferred embodiment, a joint state estimation method is used to simultaneously solve for the vehicle pose and environmental structure within a unified framework, thereby constructing an accurate and reliable 3D scene model.

[0057] The technical solution provided in this application first acquires millimeter-wave radar data and image data that are less affected by weather, ensuring the availability and stability of basic data from the sensing source. Next, based on the identified severity of the weather, the number of fused frames of the millimeter-wave radar data is dynamically adjusted to specifically compensate for the sparsity of millimeter-wave data that may be caused by severe weather. Then, by enhancing the accumulated radar data and using the enhanced radar data obtained after enhancement as a condition for de-degradation processing of the image data, the depth recovery and improvement of data quality are completed in both the radar and visual modalities. Radar data enhancement compensates for its inherent density deficiency, while cross-modal guided image de-degradation ensures the availability of visual information under severe conditions. Finally, the fused and enhanced multi-source data is used for reconstruction, ensuring the high quality and strong complementarity of the input information itself, thereby outputting stable and accurate 3D scene reconstruction results even under severe weather conditions.

[0058] Figure 2 A flowchart illustrating an embodiment of another vehicle scene reconstruction method provided in this application. Figure 2 The process shown is in Figure 1 Based on the illustrated process, the following steps are included: Step 201: Determine the weather type of the severe weather.

[0059] Weather type: refers to the classification of the meteorological conditions in which the vehicle is located, including but not limited to weather categories that affect the sensor's detection performance, such as fog, rain, sandstorms, and thunderstorms.

[0060] In this embodiment, the image data is input into a pre-trained weather classification model to identify specific weather types. This model preferably employs a lightweight network structure (such as a classifier based on the YOLOv11 framework) and is trained using a large-scale real-world driving scenario dataset (e.g., containing over 100,000 images labeled with categories such as sunny, foggy, rainy, cloudy, sandstorm, and thunderstorm) to ensure accuracy and real-time performance in practical applications. Its output is the type of severe weather currently affecting the environment.

[0061] Step 202: Determine the assessment strategy based on the weather type.

[0062] Assessment strategy: refers to a set of analytical rules, models, or algorithms developed for a specific weather type to quantify its severity.

[0063] In this embodiment, the system pre-stores the mapping relationship between different weather types and assessment strategies. For example, when identified as "foggy weather," an assessment strategy based on an atmospheric scattering physics model is invoked; when identified as "rainy weather," an assessment strategy based on rain line visual feature analysis may be invoked. Based on the output of step 201, this step selects and activates the dedicated assessment strategy that best matches the weather type from the strategy library, preparing for the next step of precise analysis.

[0064] Step 203: Analyze the image data based on the assessment strategy to obtain the severity level of the severe weather.

[0065] Severity level: refers to the discrete or continuous classification of the degree of impact of specific severe weather, such as dividing fog into three levels: light fog, dense fog, and heavy fog.

[0066] In this embodiment of the application, the evaluation strategy determined in step 202 is executed to process the image data and output a level.

[0067] In an optional implementation, the weather type is foggy, and step 203 may include the following steps: estimating a transmittance map based on the image data; extracting multiple physical features for characterizing fog concentration from the image data and the transmittance map; performing a weighted summation operation on the multiple physical features to obtain a comprehensive fog concentration score; and determining the severity level according to the predefined threshold range in which the comprehensive fog concentration score is located.

[0068] Transmittance map: refers to a two-dimensional matrix with the same size as the input image, where each pixel value represents the proportion of light that reaches the camera without being scattered by atmospheric particles at that location. Its value range is usually [0,1], and the lower the value, the higher the fog concentration at that location.

[0069] Physical characteristics: refer to physical quantities that can quantify the fog concentration level, calculated from images and transmittance maps, including but not limited to statistical characteristics and structural characteristics.

[0070] Fog Concentration Overall Score: This refers to a scalar value calculated by integrating multiple physical features, used to comprehensively characterize the fog concentration level presented in the image.

[0071] Predefined threshold range: refers to a numerical range predefined based on experiments or experience, used to map the continuous fog concentration comprehensive score to discrete severity levels.

[0072] This solution provides a specific evaluation strategy based on an atmospheric scattering physics model for foggy weather. First, a transmittance map of the image is estimated based on the dark channel prior principle. This map effectively reflects the fog concentration distribution at various points in the scene. Next, several key physical features characterizing fog concentration are extracted from the original image data and the estimated transmittance map. These features include, but are not limited to: the mean transmittance reflecting overall visibility, the standard deviation of transmittance reflecting scene depth complexity, the mean image gradient intensity reflecting edge and texture sharpness, the mean image saturation reflecting color fidelity, and the proportion of high values ​​in the dark channel reflecting the extent affected by fog. Then, these physical features are weighted and summed using preset weighting coefficients to calculate a comprehensive fog concentration score. Finally, by comparing this score with a predefined threshold range, the severity of fog is divided into several distinct levels, such as light fog, dense fog, and heavy fog. This scheme achieves accurate, objective, and robust quantitative assessment of fog concentration by comprehensively considering the impact of fog on multiple dimensions of images. It overcomes the limitations of single feature assessment and provides a highly reliable basis for subsequent targeted data processing strategies, significantly improving the overall perception robustness of the system in complex foggy environments.

[0073] Figure 2 The process shown, through the classification guidance strategy selection mechanism in steps 201 and 202, enables the system to call the most effective dedicated analysis model for the physical causes of different weather conditions such as fog and rain, avoiding the limitations of a single evaluation model in complex weather scenarios; finally, through the targeted quantitative analysis in step 203, it provides accurate and reliable adaptive control signals for subsequent sensor data fusion and processing, thereby improving the robustness and accuracy of the entire sensing system under variable and severe weather conditions from the source.

[0074] Figure 3 A flowchart illustrating another embodiment of a vehicle scene reconstruction method provided in this application. Figure 3 The process shown is in Figure 1 Based on the illustrated process, the following steps are included: Step 301: Determine the total number of frames of the millimeter-wave radar data to be fused based on the severity level, wherein the higher the severity level, the more frames are fused.

[0075] Total number of frames: refers to the number of millimeter-wave radar data frames that need to be fused to compensate for the degradation of sensing data.

[0076] In this embodiment, the system pre-defines a mapping relationship between different severity levels and the total number of frames. For example, when the fog is determined to be light, the total number of frames is set to 2; when it is determined to be dense fog, the total number of frames is set to 3; and when it is determined to be heavy fog, the total number of frames is set to 5. By establishing a positive correlation between "higher severity level, more frames" to achieve adaptive compensation for different degrees of data sparsity problems.

[0077] Step 302: Perform timestamp alignment processing on the multiple frames of millimeter-wave radar data to obtain a millimeter-wave radar data sequence.

[0078] Timestamp alignment: refers to the process of converting the timestamps of each frame of data to the same time base.

[0079] Millimeter-wave radar data sequence: refers to a set of radar data that has a continuous temporal relationship after being time-aligned.

[0080] In this embodiment, due to the difference in sensor acquisition frequency, the millimeter-wave radar data needs to be interpolated or resampled to synchronize it with the camera frame rate, forming a data sequence with strictly aligned timestamps, thus establishing a temporal consistency foundation for subsequent inter-frame fusion.

[0081] Step 303: Based on the total number of frames, select a series of consecutive frames of data with the current frame as the core from the millimeter-wave radar data sequence.

[0082] Current frame: refers to the latest acquired millimeter-wave radar data frame that serves as the fusion benchmark.

[0083] Continuous multi-frame data: refers to several adjacent frames of radar data in a time series.

[0084] In this embodiment of the application, based on the current frame, consecutive frames adjacent to each other are symmetrically selected from the aligned data sequence according to the total number of frames determined in step 301. For example, when the total number of frames is 3, the current frame and one frame before and after it are selected; when the total number of frames is 5, the current frame and two frames before and after it are selected.

[0085] Step 304: Fuse the selected consecutive frames of data to obtain the accumulated radar data.

[0086] Cumulative radar data: refers to the radar data set with increased information density obtained after multi-frame fusion processing.

[0087] In this embodiment, the selected multi-frame data is unified to the current frame coordinate system through coordinate transformation, and the spatially overlapping point clouds are weighted averaged or probabilistically fused to eliminate random noise, enhance stable target signals, and finally generate cumulative radar data with significantly improved point cloud density.

[0088] Figure 3The illustrated process establishes an adaptive mapping relationship between weather severity and the number of fused frames, enabling a targeted data compensation strategy. When severe weather leads to sparse single-frame radar data, intelligently increasing the number of fused frames effectively improves point cloud density. Simultaneously, the timestamp-aligned serialization processing and the selection strategy centered on the current frame ensure the spatiotemporal consistency of the fused data, providing high-quality sensing input for subsequent processing and significantly enhancing the system's sensing reliability under severe weather conditions.

[0089] Based on the information-enhanced cumulative radar data obtained in the aforementioned steps, in order to overcome the inherent sparsity of millimeter-wave radar and further improve data quality, this embodiment achieves depth enhancement of radar data through the following steps: converting the cumulative radar data into a radar depth image; using the radar depth image as a guiding condition, iteratively denoising from a random noise state using a pre-trained generative model to generate a target depth image with a higher spatial point cloud density than the radar depth image; converting the target depth image into three-dimensional point cloud data, and using the three-dimensional point cloud data as the enhanced radar data.

[0090] Radar depth image: refers to a two-dimensional image obtained by converting a three-dimensional radar point cloud through spherical projection, where each pixel value represents the distance information of the corresponding spatial point.

[0091] Guiding conditions: These refer to the auxiliary information provided to the model during the generation process, used to control and guide the direction and content of the generation process.

[0092] Pre-trained generative models: These are neural network models that have been pre-trained based on a large amount of paired data and are capable of generating target data according to conditional inputs.

[0093] Random noise state: refers to the Gaussian distributed random tensor that serves as the starting point of the generation process.

[0094] Iterative denoising: refers to the process of gradually removing noise and restoring data features through multiple time steps.

[0095] Target depth image: refers to the depth image obtained after processing by a generative model that is superior to the input in terms of spatial detail and integrity.

[0096] 3D point cloud data: refers to the set of 3D spatial points obtained by reconstructing a 2D depth image through coordinate back projection.

[0097] In this embodiment, the accumulated radar data obtained in the preceding steps is first converted into a radar depth image through spherical coordinate projection, completing the format conversion from a 3D point cloud to a 2D image domain. Subsequently, this radar depth image is used as a guiding condition and input into a pre-trained generative model. Starting from a state of random noise, the model gradually generates a target depth image that is significantly superior to the input in terms of spatial point cloud density and detail integrity through an iterative denoising process involving multiple time steps. This generative model is trained as follows: using paired millimeter-wave radar depth images and lidar depth images as training data, with the former as the condition and the latter as the learning objective, the model's ability to reconstruct high-quality depth images from noise is trained. Finally, the obtained target depth image is converted into 3D point cloud data through coordinate back-projection, which is the required enhanced radar data.

[0098] This solution transforms the radar data augmentation problem into a condition generation task. By leveraging the powerful feature learning capabilities of deep learning, it effectively recovers the spatial details lost due to weather effects and sensor physical limitations, providing an unprecedented data foundation for subsequent perception and reconstruction tasks and fundamentally improving the system's perception capabilities under adverse weather conditions.

[0099] Based on the radar data with spatial information enhancement, in order to synergistically improve the quality of visual perception, this embodiment achieves image de-degradation processing through the following steps: spatially aligning the enhanced radar data with the image data; using the aligned enhanced radar data as a condition, iteratively denoising the image data through a pre-trained image generation model to generate the enhanced image.

[0100] Spatial alignment: refers to the process of unifying data from different sensors into the same spatial coordinate system through coordinate transformation, establishing a precise correspondence between radar point clouds and image pixels.

[0101] Image generation model: refers to a deep learning network that can reconstruct degraded images based on conditional input. In this embodiment, it specifically refers to an image restoration network based on a diffusion model.

[0102] Iterative denoising: refers to a sequential computation process that gradually restores image quality through multiple rounds of feature extraction and noise removal.

[0103] Enhanced images: refers to image data that has undergone de-degradation processing and has significantly improved visual quality.

[0104] In this embodiment, enhanced radar data is first projected onto the image coordinate system using sensor calibration parameters to achieve pixel-level spatial alignment between the radar point cloud and the image, establishing an accurate cross-modal correspondence. Subsequently, the aligned enhanced radar data is used as geometric condition input to a pre-trained image generation model. This model processes foggy, degraded images, gradually restoring image details through multi-step iterative denoising. Specifically, in each denoising step, the model simultaneously considers the current noisy image, radar geometric conditions (providing scene depth and structural information), and time step information, predicting noise components and updating the image state, ultimately generating a sharpened enhanced image. The model is trained using real-degraded image pairs as supervisory signals, with aligned radar data as a condition, enabling the model to learn the ability to restore images under geometric constraints, ensuring that the restoration result maintains both visual realism and spatial structural consistency.

[0105] This embodiment utilizes enhanced radar data as a geometric prior to guide the image restoration process, effectively improving visual perception quality through a cross-modal collaborative mechanism. This method overcomes the structural distortion problem that single visual algorithms are prone to in adverse weather conditions, while ensuring the spatial consistency between the restored image and the real scene. It provides high-quality visual input for subsequent multi-sensor fusion perception, significantly improving the overall perception reliability of the system under complex weather conditions.

[0106] Figure 4 A flowchart illustrating another embodiment of the vehicle scene reconstruction method provided in this application. Figure 4 The process shown is in Figure 1 Based on the illustrated process, the following steps are included: Step 401: Perform pre-integration processing on the inertial measurement unit data to obtain vehicle motion information.

[0107] Pre-integration processing: refers to the process in an inertial navigation system of integrating the angular velocity and acceleration data measured by the IMU in a local coordinate system to obtain the relative motion increment.

[0108] Vehicle motion information: refers to incremental information obtained through pre-integration calculation that describes the relative pose changes of a vehicle over continuous time intervals.

[0109] In this embodiment, the acquired IMU angular velocity and acceleration data are pre-integrated. This method integrates the IMU measurements in a local coordinate system to obtain the relative motion transformation (including changes in position, velocity, and attitude) between adjacent keyframes, avoiding the problem of repeated integration due to state updates during optimization and significantly improving the computational efficiency of subsequent state estimation.

[0110] Step 402: Perform joint state estimation on the enhanced image, the enhanced radar data, and the vehicle motion information to obtain an estimation result that includes the vehicle pose state and surrounding environment data.

[0111] Joint state estimation: refers to the process of simultaneously solving for the vehicle's own state and the environmental characteristic state within a unified optimization framework.

[0112] Vehicle pose status: refers to the vehicle's position and attitude information in the global coordinate system.

[0113] Surrounding environment data: refers to the three-dimensional spatial information of the characteristics of the environment surrounding the vehicle.

[0114] In this embodiment, visual features from the enhanced image, 3D point clouds from the enhanced radar data, and motion information obtained from IMU pre-integration are jointly input into a tightly coupled fusion framework for joint optimization. This process constructs a joint objective function for visual reprojection error, radar point cloud matching error, and IMU pre-integration error, simultaneously solving for the vehicle's precise pose state and the 3D coordinates of environmental features within a unified nonlinear optimization problem, thus achieving simultaneous vehicle self-localization and environmental mapping.

[0115] Step 403: Based on the estimation results, construct the three-dimensional scene reconstruction data.

[0116] In this embodiment, based on the vehicle pose and three-dimensional coordinates of environmental features obtained from joint state estimation, a complete and consistent three-dimensional model of the vehicle's surrounding environment is constructed using algorithms such as point cloud stitching and surface reconstruction. This model can be represented as a dense point cloud, a mesh map, or a voxel map, accurately reflecting the three-dimensional geometric structure of the scene and providing reliable environmental perception results for subsequent advanced decision-making tasks such as path planning and obstacle avoidance.

[0117] Figure 4 The illustrated process improves the efficiency of IMU data utilization through pre-integration processing and fully leverages the complementary characteristics of multi-source perception information through tightly coupled joint state estimation, ultimately achieving high-precision synchronous estimation of vehicle pose and environmental structure. This deep fusion scheme effectively overcomes the limitations of single sensors in adverse weather conditions, ensuring the generation of accurate and reliable 3D environmental models even under complex meteorological conditions, providing a stable and reliable perception foundation for intelligent driving systems.

[0118] The following provides a comprehensive embodiment that systematically illustrates the complete application process of the method described in this application in the reconstruction of severe vehicle scenarios.

[0119] This application proposes a vehicle adverse scene reconstruction method based on multi-sensor fusion. When the vehicle is in adverse scenes such as fog or rain, the method uses data collected by the vehicle's own millimeter-wave radar, camera, and inertial measurement unit (IMU) as system input, and achieves accurate environmental perception and reconstruction through the following steps: 1. Weather Scene Recognition and Classification First, image data is used to identify the weather conditions of the vehicle's environment. This embodiment employs a weather classification model trained on the YOLOv11 framework, a mature and lightweight object detection and classification model. The training dataset contains six categories: sunny, foggy, rainy, cloudy, sandstorm, and thunderstorm. Each category contains 100,000 real images collected during actual vehicle operation. The training, test, and validation sets are divided in a 7:2:1 ratio. Through hyperparameter optimization, the model achieves an overall recognition accuracy of over 96% and a recall rate exceeding 95%. Specifically, the recall rate for severe weather (such as rain and fog) is no less than 99.9%, aiming to minimize the perceptual risks caused by weather misjudgments.

[0120] 2. Fog severity classification and dynamic accumulation of millimeter-wave radar data When the system determines that the current weather is foggy, it further refines the fog concentration based on a physical model. The specific algorithm logic is as follows: The first step is to acquire the image and estimate its global transmittance map, thereby triggering a road safety threshold alarm. The approximate calculation method can be obtained using equation (1): (1); in, Represents spatial pixel coordinates, usually referring to the position (coordinate vector) of a center pixel. Transmittance at pixel coordinate x represents the proportion of light that reaches the camera, and is used to map pixel intensity to visibility or for distance estimation. This is a conservative scaling factor, which controls the intensity of dehazing; here it is set to 0.95. The normalized dark channel value, ranging from 0 to 1, can be obtained using equation (2): (2); in, For adjacent windows (size is) (a square structural unit), y refers to the position of other pixels traversed within the neighborhood of the center pixel x. This corresponds to the pixel values ​​of the RGB three channels (usually integers in the range of 0 to 255, representing the brightness intensity of the red, green, and blue channels). The dark channel reflects the local minimum color intensity. Fog scattering causes a decrease in contrast, and the dark channel value increases, so it can be used to estimate the brightness shift caused by atmospheric scattering.

[0121] The second step involves extracting the following five fog density assessment features from the image and transmittance map: including the mean transmittance, which represents overall sharpness. The standard deviation of transmittance represents the depth of the scene, i.e., the structural diversity. Normalized gradient intensity, representing texture or edge information. Mean saturation, representing color information And the proportion of high values ​​in the dark channel, representing the proportion of bright spots and highlights. ; Mean transmittance It can be obtained using equation (3): (3); in, This represents the total number of pixels in the image. This represents the set of image pixels; the denser the fog, the more likely it is to cause fog. The lower; Standard deviation of transmittance It can be obtained using equation (4): (4); in The transmittance map represents the range of values ​​(0 to 1), with large depth variations. The larger; Normalized gradient strength It can be obtained using equation (5): (5); in, and The kernels are 3×3 Sobel convolutions, corresponding to horizontal and vertical edge detection, respectively. For discrete convolution operations, the scattering effect of fog reduces the edge sharpness and local contrast of the image (i.e., causes high-frequency information attenuation). Therefore, the values ​​calculated from the image... It will get smaller; It is a grayscale image, which can be obtained using equation (6): (6); Mean saturation It can be obtained using equation (7): (7); in, This is the saturation channel, with a value range of 0 to 255. Saturation decreases in foggy weather. The value decreases.

[0122] High value ratio in dark channel It can be obtained using equation (8): (8); In particular, the proportion increases as the dark passage is raised during periods of dense fog.

[0123] The third step is to calculate the comprehensive score of fog concentration based on the above-mentioned features using weighted calculation. See equation (9), and the fog level is divided into the following intervals: (9); When the score is between 0 and 0.35, the scene is judged as light fog; when the score is between 0.36 and 0.65, the scene is judged as dense fog; when the score is between 0.66 and 1, the scene is judged as heavy fog.

[0124] In applications, end-to-end deep learning models can also be used to directly determine the degree of fog. For example, a convolutional neural network can be trained, taking a foggy image as input, to directly regress a fog concentration score or output a classification result.

[0125] The fourth step is to dynamically accumulate millimeter-wave radar data frames based on the fog level. In this embodiment, the timestamp-aligned millimeter-wave radar data is accumulated based on the camera frame rate as follows: Light fog scene: the previous frame and the next frame are fused, for a total of 2 frames; Dense fog scene: the previous frame, the current frame, and the next frame are fused, for a total of 3 frames; Heavy fog scene: the previous two frames, the current frame, and the next two frames are fused, for a total of 5 frames.

[0126] 3. Data densification of millimeter-wave radar This step utilizes a diffusion model to enhance the density of millimeter-wave radar data. First, the accumulated millimeter-wave radar data is converted into a depth image format through spherical projection. The projection calculation is shown in equation (10): (10); Where (u, v) are the pixel coordinates in the depth map after projection. () represents the three-dimensional Cartesian coordinates of the radar point. Let be the azimuth of the point. The pitch angle of point A and For the width and height of the image, and Indicates the range of azimuth and elevation angles. This represents the geometric distance from the point to the origin. This represents the floor function.

[0127] During the training phase, the aligned lidar depth map is used as the ground truth, and noise is added to the lidar depth map through a forward diffusion process. The noise scheduling is shown in equation (11): ; (11) in, X represents the initial LiDAR depth map, which serves as the ground truth for training; Xt represents the noisy depth map at time step t. It is a cleaner image obtained after denoising at time step t; q() represents the distribution of the forward diffusion process; This represents a predefined noise schedule, a value between (0,1) that increases with time step t, controlling the amount of noise added at each step. N() represents a Gaussian distribution (normal distribution); I is the identity matrix, representing the covariance matrix of the noise, meaning that the noise is independent and has the same variance in each dimension (in the case of multidimensionality, each component is independent and has a variance of 1). It is the cumulative noise product, defined as: , It represents the initial image. The remaining proportion at time step t.

[0128] The training conditional denoising network (such as UNet) uses the millimeter-wave radar depth map and the noise time step as conditions to predict the added noise. The loss function is the mean square error, as shown in Equation (12): (12); in, It is a loss function; It is the expected value, representing the result at time step t and initial image. ,noise Find the average; It is real noise; It is the noise predicted by a denoising network (such as UNet), whose parameters are θ; t is the noisy image; t is the current time step, which is usually encoded and input into the network; c is the conditional information, which here refers to the millimeter-wave radar depth map, used to guide the denoising process.

[0129] In the inference stage, starting from random noise, reverse denoising is performed in combination with radar depth map conditions to generate a high-resolution depth image. The denoising process is shown in equation (13): (13); in, () represents the distribution of the reverse (generation) process, defined by the learned neural network parameters θ; This represents a cleaner image obtained after denoising at time step t. The mean of the reverse process of network prediction can be represented by... Calculated; represents the variance of the reverse process of network prediction, which is usually a fixed value; c represents the conditional information, namely the millimeter-wave radar depth map, which is also used to guide the denoising at each step in the inference stage.

[0130] Finally, the generated target depth image is back-projected into a 3D point cloud (i.e., each pixel corresponds to a 3D point, with coordinates calculated from distance, angle, and height), which serves as enhanced radar data. The complete flowchart of this process is shown below. Figure 5 As shown.

[0131] 4. Image data de-degradation processing Image de-degradation processing employs a diffusion model framework similar to radar densification: During the training phase, paired real clear images and foggy images were used, with the densed millimeter-wave radar point cloud data as conditions, to train the conditional denoising network to learn how to recover clear images from noise. During the inference phase, foggy images and aligned enhanced radar data are input into the pre-trained model, and a clear image after defogging is generated through iterative denoising.

[0132] 5. Environmental Perception and Reconstruction Finally, the system performs multi-sensor fusion and scene reconstruction: First, the IMU data is pre-integrated to obtain vehicle motion information. Then, a tightly coupled fusion framework (such as FAST-LIVO2) is used to simultaneously process the enhanced image, enhanced radar data, and IMU motion information to perform joint state estimation. Finally, a three-dimensional scene reconstruction result of the vehicle's surrounding environment is constructed, which can be directly used for subsequent tasks such as environmental semantic recognition, path planning, and control.

[0133] This comprehensive embodiment fully demonstrates the entire process of this application from weather recognition to 3D reconstruction. Through the collaborative processing and enhancement of multi-sensor data, it effectively improves the vehicle's environmental perception capability in adverse weather conditions.

[0134] In another alternative implementation of this application, the system architecture and data processing flow remain unchanged, but different sensing sources and implementation methods are used in some technical modules compared to the main solution, as follows: Weather perception and severity classification module: This alternative solution incorporates LiDAR data into the perception decision-making system. The system can independently determine weather conditions based on weather-induced degradation features of LiDAR point clouds (such as decreased point cloud density and shortened detection range), or fuse these features with image data and input them into the classification model to improve judgment accuracy. When classifying fog severity, the system can jointly analyze the visual features of the image and the physical statistical features of the LiDAR point cloud (such as the number of effective points and maximum detection range) and make a comprehensive judgment through feature-level fusion.

[0135] Data Projection Module: This solution employs layered bird's-eye view projection when preprocessing millimeter-wave radar data. This method divides the 3D point cloud into multiple layers based on height ranges and projects each layer onto a 2D plane, generating a set of BEV (Bird's Eye View) images. This representation method more completely preserves the vertical structure information of the scene, providing richer spatial context for subsequent processing.

[0136] Image De-degradation Module: This solution employs a generative adversarial network (GAN) for image dehazing. This network learns a direct end-to-end mapping from a foggy image to a clear image through adversarial training between the generator and discriminator. During inference, the foggy image is input into the trained generator, which directly outputs the dehazed, enhanced image.

[0137] This alternative embodiment demonstrates that the core objective of this invention can also be achieved by introducing lidar as an auxiliary sensing source and adjusting the data representation and processing algorithms accordingly. This further illustrates that the "multi-sensor adaptive fusion based on weather levels" system architecture proposed in this invention possesses good flexibility and scalability.

[0138] Figure 6 This is a block diagram illustrating an embodiment of a vehicle scene reconstruction device provided in this application. Figure 6 As shown, the device includes: The acquisition module 61 is used to acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle. The determination module 62 is used to determine the severity level of the severe weather when the weather condition of the environment in which the vehicle is located is identified as severe weather. The fusion module 63 is used to perform fusion processing on multiple frames of the millimeter-wave radar data according to the severity level to obtain cumulative radar data; Enhancement module 64 is used to enhance the accumulated radar data to generate enhanced radar data; Processing module 65 is used to perform de-degradation processing on the image data based on the enhanced radar data to generate an enhanced image; The generation module 66 is used to generate three-dimensional scene reconstruction data of the vehicle's surrounding environment based on the enhanced image, the enhanced radar data, and the inertial measurement unit data.

[0139] In one possible implementation, the determining module is specifically used for: Determine the weather type of the severe weather; Determine the assessment strategy based on the weather type; The severity level of the severe weather is obtained by analyzing the image data based on the assessment strategy.

[0140] In one possible implementation, the weather type is foggy, and the determining module is further configured to: Estimate the transmittance map based on the image data; From the image data and the transmittance map, multiple physical features for characterizing fog concentration are extracted; A weighted summation operation is performed on multiple physical characteristics to obtain a comprehensive score for fog concentration. The severity level is determined based on the predefined threshold range in which the comprehensive fog concentration score falls.

[0141] In one possible implementation, the fusion module is specifically used for: Based on the severity level, the total number of frames of millimeter-wave radar data to be fused is determined, wherein the higher the severity level, the more frames are fused. The millimeter-wave radar data from multiple frames is timestamped to obtain a millimeter-wave radar data sequence. Based on the total number of frames, select consecutive multiple frames of data with the current frame as the core from the millimeter-wave radar data sequence; The accumulated radar data is obtained by fusing the selected consecutive frames of data.

[0142] In one possible implementation, the enhancement module is specifically used for: The accumulated radar data is converted into a radar depth image; Using the radar depth image as a guiding condition, a pre-trained generative model is used to iteratively denoise from a random noise state to generate a target depth image with a higher spatial point cloud density than the radar depth image. The target depth image is converted into three-dimensional point cloud data, and the three-dimensional point cloud data is used as the enhanced radar data.

[0143] In one possible implementation, the processing module is specifically used for: Spatially align the enhanced radar data with the image data; Using the aligned enhanced radar data as a condition, the image data is iteratively denoised using a pre-trained image generation model to generate the enhanced image.

[0144] In one possible implementation, the generation module is specifically used for: The inertial measurement unit data is pre-integrated to obtain vehicle motion information; Joint state estimation is performed on the enhanced image, the enhanced radar data, and the vehicle motion information to obtain an estimation result that includes the vehicle pose state and surrounding environment data. Based on the estimation results, the 3D scene reconstruction data is constructed.

[0145] like Figure 7 As shown in the figure, this application provides a device including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114. Memory 113 is used to store computer programs; In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the vehicle scene reconstruction method provided in any of the foregoing method embodiments, including: Acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle; If the weather conditions of the environment in which the vehicle is located are identified as severe weather, the severity level of the severe weather shall be determined. Based on the severity level, multiple frames of the millimeter-wave radar data are fused to obtain cumulative radar data. The accumulated radar data is enhanced to generate enhanced radar data; The image data is de-degraded based on the enhanced radar data to generate an enhanced image; Based on the enhanced image, the enhanced radar data, and the inertial measurement unit data, three-dimensional scene reconstruction data of the vehicle's surrounding environment is generated.

[0146] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the vehicle scene reconstruction method provided in any of the foregoing method embodiments.

[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0148] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0149] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0150] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A vehicle scene reconstruction method, characterized in that, The method includes: Acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle; If the weather conditions of the environment in which the vehicle is located are identified as severe weather, the severity level of the severe weather shall be determined. Based on the severity level, multiple frames of the millimeter-wave radar data are fused to obtain cumulative radar data. The accumulated radar data is enhanced to generate enhanced radar data; The image data is de-degraded based on the enhanced radar data to generate an enhanced image; Based on the enhanced image, the enhanced radar data, and the inertial measurement unit data, three-dimensional scene reconstruction data of the vehicle's surrounding environment is generated.

2. The method according to claim 1, characterized in that, Determining the severity level of the severe weather includes: Determine the weather type of the severe weather; Determine the assessment strategy based on the weather type; The severity level of the severe weather is obtained by analyzing the image data based on the assessment strategy.

3. The method according to claim 2, characterized in that, The weather type is foggy. The analysis of the image data based on the assessment strategy to obtain the severity level of the severe weather includes: Estimate the transmittance map based on the image data; From the image data and the transmittance map, multiple physical features for characterizing fog concentration are extracted; A weighted summation operation is performed on multiple physical characteristics to obtain a comprehensive score for fog concentration. The severity level is determined based on the predefined threshold range in which the comprehensive fog concentration score falls.

4. The method according to claim 1, characterized in that, The step of fusing multiple frames of millimeter-wave radar data according to the severity level to obtain cumulative radar data includes: Based on the severity level, the total number of frames of millimeter-wave radar data to be fused is determined, wherein the higher the severity level, the more frames are fused. The millimeter-wave radar data from multiple frames is timestamped to obtain a millimeter-wave radar data sequence. Based on the total number of frames, select consecutive multiple frames of data with the current frame as the core from the millimeter-wave radar data sequence; The accumulated radar data is obtained by fusing the selected consecutive frames of data.

5. The method according to claim 1, characterized in that, The step of enhancing the accumulated radar data to generate enhanced radar data includes: The accumulated radar data is converted into a radar depth image; Using the radar depth image as a guiding condition, a pre-trained generative model is used to iteratively denoise from a random noise state to generate a target depth image with a higher spatial point cloud density than the radar depth image. The target depth image is converted into three-dimensional point cloud data, and the three-dimensional point cloud data is used as the enhanced radar data.

6. The method according to claim 1, characterized in that, The step of performing de-degradation processing on the image data based on the enhanced radar data to generate an enhanced image includes: Spatially align the enhanced radar data with the image data; Using the aligned enhanced radar data as a condition, the image data is iteratively denoised using a pre-trained image generation model to generate the enhanced image.

7. The method according to claim 1, characterized in that, The step of generating three-dimensional scene reconstruction data of the vehicle's surrounding environment based on the enhanced image, the enhanced radar data, and the inertial measurement unit data includes: The inertial measurement unit data is pre-integrated to obtain vehicle motion information; Joint state estimation is performed on the enhanced image, the enhanced radar data, and the vehicle motion information to obtain an estimation result that includes the vehicle pose state and surrounding environment data. Based on the estimation results, the 3D scene reconstruction data is constructed.

8. A vehicle scene reconstruction device, characterized in that, The device includes: The acquisition module is used to acquire image data, multi-frame millimeter-wave radar data, and inertial measurement unit data collected by the vehicle. The determination module is used to determine the severity level of the severe weather when the weather conditions of the environment in which the vehicle is located are identified as severe weather. The fusion module is used to fuse multiple frames of millimeter-wave radar data according to the severity level to obtain cumulative radar data. An enhancement module is used to enhance the accumulated radar data to generate enhanced radar data; The processing module is used to perform de-degradation processing on the image data based on the enhanced radar data to generate an enhanced image; The generation module is used to generate three-dimensional scene reconstruction data of the vehicle's surrounding environment based on the enhanced image, the enhanced radar data, and the inertial measurement unit data.

9. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a vehicle scene reconstruction program stored in the memory to implement the vehicle scene reconstruction method according to any one of claims 1-7.

10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the vehicle scene reconstruction method according to any one of claims 1-7.