Static scene model training method and device, equipment, medium and product

By preprocessing and rendering the multi-frame original point cloud and perimeter-view images, combined with lighting information analysis and multi-layer perceptron module inference, the problem of the existing technology being unable to truly restore scenes under different lighting conditions is solved, a high-fidelity and diverse simulation environment is achieved, and the effect of autonomous driving algorithm testing is improved.

CN120014147APending Publication Date: 2025-05-16CHONGQING CHANGAN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510095106.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art cannot truly restore scenes under different lighting conditions, resulting in a lack of authenticity and diversity in the simulation environment in the test of autonomous driving algorithms.

Method used

By obtaining multi-frame original point clouds and peripheral visual images in different scenes, combining the internal and external parameters of the sensor, the point clouds are preprocessed and rendered, and lighting information is used to infer scene color and transparency information under different lighting conditions by using lighting information analysis and multi-layer perceptron module, and then the parameters of the static scene model are updated.

Benefits of technology

High-fidelity reconstruction of scenes under different lighting conditions is achieved, which enhances the authenticity and diversity of the simulation environment, thereby improving the effectiveness of autonomous driving algorithm testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014147A_ABST
    Figure CN120014147A_ABST
Patent Text Reader

Abstract

The invention provides a static scene model training method and apparatus, a device, a medium and a product. The training method comprises the steps of obtaining multiple frames of original point clouds and original panoramic images of a current vehicle in different scenes and internal and external parameters of a sensor corresponding to the original panoramic images; on the basis of the original panoramic image and internal and external parameters of the sensor, preprocessing the multiple frames of original point clouds to obtain multiple frames of static scene point clouds; analyzing illumination information obtained after illumination estimation of the original panoramic image to obtain illumination color distribution and transparency information of a scene in the panoramic image; rendering the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of panoramic images; and updating the parameters of the static scene model by using the loss between the multi-frame panoramic image and the original panoramic image until the static scene model converges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and specifically to a training method, device, equipment, medium and product for a static scene model. Background Art

[0002] In recent years, autonomous driving technology has become a research hotspot in the automotive field, and autonomous driving algorithm testing is particularly important. Before the algorithm is deployed in a real vehicle system, a large number of reliability tests are usually carried out in a simulation environment, and the authenticity and diversity of the simulation environment become very critical. In related technologies, the input two-dimensional image is subjected to illumination analysis to determine its main illumination information and ambient illumination information, and then the main illumination of the three-dimensional scene is adjusted according to the main illumination information. After the illumination effect of the three-dimensional scene is adjusted to be consistent with the illumination effect of the two-dimensional image, the three-dimensional scene is illuminated and rendered. Due to the single illumination, it is impossible to truly restore scenes under different illumination conditions. Summary of the invention

[0003] The purpose of this application is to provide a training method, device, equipment, medium and product for a static scene model to solve the problem that the existing technology cannot truly restore scenes under different lighting conditions.

[0004] The technical solutions adopted in this application are as follows:

[0005] A training method for a static scene model, the training method comprising: obtaining multiple frames of original point clouds, original surrounding images, and internal and external parameters of sensors corresponding to the original surrounding images of a current vehicle in different scenes; preprocessing the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensors to obtain multiple frames of static scene point clouds; analyzing illumination information obtained after illumination estimation of the original surrounding images to obtain illumination color distribution and transparency information of the scene in the surrounding images; rendering the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of surrounding images; and updating the parameters of the static scene model using the loss between the multiple frames of surrounding images and the original surrounding images until the static scene model converges.

[0006] According to the above technical means, on the one hand, by acquiring multiple frames of original point clouds and original surrounding images, as well as the corresponding internal and external parameters of the sensor, the static scene can be accurately reconstructed; on the other hand, by analyzing the illumination information obtained after the illumination estimation of the original surrounding image, the illumination color distribution and transparency information of the scene in the surrounding image can be obtained, which can render the scene more realistically, making the generated surrounding image closer to the real scene.

[0007] Furthermore, based on the original surrounding image and the internal and external parameters of the sensor, the multiple frames of original point clouds are preprocessed to obtain multiple frames of static scene point clouds, including: removing dynamic point clouds in the multiple frames of original point clouds to obtain multiple frames of static point clouds; based on the original surrounding image and the internal and external parameters of the sensor, the multiple frames of static point clouds are colorized to obtain multiple frames of colored static point clouds; and the coordinate system of the multiple frames of colored static point clouds is converted into a world coordinate system to obtain the multiple frames of static scene point clouds.

[0008] According to the above technical means, on the one hand, by removing dynamic point clouds in multi-frame original point clouds, such as moving vehicles, pedestrians, etc., the accuracy of multi-frame static point clouds can be improved; on the other hand, by coloring multi-frame static point clouds according to the original surrounding images and the internal and external parameters of the sensor, the multi-frame colored static point clouds can more intuitively reflect the colors of objects in the scene; on the other hand, by converting the coordinate system of multi-frame colored static point clouds into the world coordinate system, the coordinates of point clouds in different frames can be unified, which is conducive to the splicing and fusion of multi-frame static scene point clouds.

[0009] Furthermore, the removing of dynamic point clouds from the multi-frame original point clouds to obtain multi-frame static point clouds includes: compensating the multi-frame original point clouds to obtain multi-frame compensated point clouds; performing three-dimensional target detection on the multi-frame compensated point clouds to obtain dynamic point clouds in the multi-frame compensated point clouds; and removing the dynamic point clouds from the multi-frame compensated point clouds to obtain the multi-frame static point clouds.

[0010] According to the above technical means, on the one hand, by compensating the multi-frame original point cloud, the problem of real position deviation caused by the current vehicle movement can be eliminated; on the other hand, by removing dynamic point clouds in the multi-frame original point cloud, such as moving vehicles, pedestrians, etc., the accuracy of the multi-frame static point cloud can be improved.

[0011] Furthermore, the coloring of the multiple frames of static point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of colored static point clouds includes: determining the posture information of the current vehicle based on the positioning data, acceleration and angular velocity of the current vehicle; de-distorting the original surrounding images based on the internal and external parameters of the sensor to obtain de-distorted surrounding images; coloring the multiple frames of static point clouds based on the posture information of the current vehicle, the internal and external parameters of the sensor and the de-distorted surrounding images to obtain the multiple frames of colored static point clouds.

[0012] According to the above technical means, on the one hand, by combining the current vehicle posture information and the internal and external parameters of the sensor, the pixel points in the original surround image can be accurately matched with the static point cloud; on the other hand, the colored static point cloud can more intuitively display the three-dimensional structure and color characteristics of objects in the environmental scene, allowing users to more easily understand the three-dimensional coordinate information and rich color information expressed by the point cloud.

[0013] Furthermore, the coloring of the multi-frame static point clouds based on the current vehicle's posture information, the internal and external parameters of the sensor, and the de-distorted surrounding image to obtain the multi-frame colored static point clouds includes: projecting the multi-frame static point clouds onto the de-distorted surrounding image based on the current vehicle's posture information and the internal and external parameters of the sensor to obtain colors corresponding to the multi-frame static point clouds; coloring the multi-frame static point clouds using the colors corresponding to the multi-frame static point clouds to obtain the multi-frame colored static point clouds.

[0014] According to the above technical means, on the one hand, by combining the current vehicle posture information and the internal and external parameters of the sensor, the pixel points in the original surround image can be accurately matched with the static point cloud; on the other hand, the colored static point cloud can more intuitively display the three-dimensional structure and color characteristics of objects in the environmental scene, allowing users to more easily understand the three-dimensional coordinate information and rich color information expressed by the point cloud.

[0015] Furthermore, the coordinate system of the multi-frame colored static point cloud is converted into a world coordinate system to obtain the multi-frame static scene point cloud, including: for the multi-frame colored static point cloud within a period of time, obtaining the posture information of the current vehicle corresponding to the multi-frame colored static point cloud; based on the posture information of the current vehicle, the coordinate system of the multi-frame colored static point cloud is converted into a world coordinate system to obtain the multi-frame static scene point cloud.

[0016] According to the above technical means, by converting the coordinate system of the multi-frame colored static point cloud to the world coordinate system, the data inconsistency problem caused by different coordinate systems is avoided. The multi-frame static scene point cloud obtained after the coordinate system conversion can more intuitively display the environmental scene structure, including roads, buildings, obstacles, etc., which helps the vehicle to better understand the surrounding environment scene.

[0017] Furthermore, the illumination information obtained after the illumination estimation of the original surrounding image is analyzed to obtain the illumination color distribution and transparency information of the scene in the surrounding image, including: performing illumination estimation on the original surrounding image to obtain the illumination information of the surrounding image; and analyzing the illumination information of the surrounding image to obtain the illumination color distribution and transparency information of the scene in the surrounding image.

[0018] According to the above technical means, on the one hand, by performing illumination estimation on the original surrounding image, the illumination information in the original surrounding image can be fully obtained, including illumination intensity, direction, color, etc.; on the other hand, by analyzing the illumination information of the surrounding image, the illumination color distribution and transparency information obtained can reflect the color of the light source in the scene, as well as the translucency or degree of occlusion of the object.

[0019] A training device for a static scene model, the training device comprising:

[0020] An input module, used to obtain multiple frames of original point clouds, original surrounding images, and internal and external parameters of sensors corresponding to the original surrounding images of the current vehicle in different scenes;

[0021] A data preprocessing module, used for preprocessing the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of static scene point clouds;

[0022] A multi-layer perceptron module is used to analyze the illumination information obtained after the illumination estimation of the original surrounding image, and obtain the illumination color distribution and transparency information of the scene in the surrounding image;

[0023] A three-dimensional Gaussian module, used for rendering the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of panoramic images;

[0024] The output module is used to update the parameters of the static scene model by using the loss between the multiple frames of surrounding images and the original surrounding images until the static scene model converges.

[0025] An electronic device comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and the processor implements the steps in the above method when executing the program.

[0026] A computer-readable storage medium stores a computer program, which implements the steps in the above method when executed by a processor.

[0027] A computer program product, comprising a computer program or instructions, wherein the computer program or instructions implement the steps in the above method when executed by a computer. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A schematic diagram of the implementation flow of a static scene model training method proposed in an embodiment of the present application;

[0029] Figure 2A schematic diagram of the overall framework of a high-fidelity static scene reconstruction method with controllable illumination based on a neural network proposed in an embodiment of the present application;

[0030] Figure 3 A schematic diagram of the implementation process of a neural network-based high-fidelity static scene reconstruction method with controllable illumination proposed in an embodiment of the present application;

[0031] Figure 4 A schematic diagram of a data preprocessing process proposed in an embodiment of the present application;

[0032] Figure 5 A schematic diagram of a method for generating a panoramic picture with six viewing angles proposed in an embodiment of the present application;

[0033] Figure 6 A schematic diagram of the use of a multi-layer perceptron module and a 3DGS reconstruction module for rendering a scene image proposed in an embodiment of the present application;

[0034] Figure 7 A schematic diagram of the structure of a static scene model training device proposed in an embodiment of the present application;

[0035] Figure 8 A schematic diagram of a hardware entity of an electronic device proposed in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will describe the implementation methods of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustrating the present application, not for limiting the scope of protection of the present application.

[0037] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application, and thus the drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed at will, and the component layout may also be more complicated.

[0038] The present application embodiment proposes a training method for a static scene model, such as Figure 1 As shown, the training method of the static scene model includes the following steps S101 to S105, wherein:

[0039] Step S101, obtaining multiple frames of original point clouds, original surrounding images, and internal and external parameters of sensors corresponding to the original surrounding images of the current vehicle in different scenes;

[0040] Here, different scenes can be scenes such as urban roads, rural roads, highways, parking lots, etc., which can cover various road types and surrounding environments that the current vehicle may encounter. The current vehicle can be any type of vehicle equipped with sensors (e.g., lidar, camera), such as a car, truck, bus, etc. The multi-frame original point cloud of the current vehicle is usually obtained by lidar. The original surrounding image can be an all-round image around the current vehicle, usually obtained by multiple cameras or cameras.

[0041] In some embodiments, a laser radar is used to collect point clouds of the vehicle at different time points to obtain multiple frames of original point clouds of the current vehicle, each frame of original point cloud contains location information (x, y, z), and these point clouds represent the three-dimensional shape and structure of the environment surrounding the current vehicle.

[0042] In some embodiments, the original surrounding image of the current vehicle is obtained by collecting the front image, the front left image, the front right image, the rear image, the rear left image, and the rear right image of the current vehicle through the six surrounding cameras on the current vehicle. In this way, the original surrounding image provides an all-round view of the surroundings of the current vehicle.

[0043] Here, the intrinsic and extrinsic parameters of the sensor can be used to align the image and point cloud data to the same coordinate system. The intrinsic parameters of the sensor are parameters related to the characteristics of the sensor itself, such as the focal length, pixel size, distortion parameters, etc. of the camera, and the scanning frequency, angular resolution, etc. of the lidar. The extrinsic parameters of the sensor can be parameters that describe the position and orientation of the sensor in the vehicle coordinate system, including rotation matrix and translation matrix.

[0044] In some embodiments, the intrinsic parameters of the sensor are usually calibrated before the sensor leaves the factory and stored in the internal memory of the sensor. The intrinsic parameters of the sensor can be obtained by reading the configuration file of the sensor or using a special calibration tool.

[0045] In some implementations, the acquisition of the extrinsic parameters of the sensor usually needs to be achieved through a calibration process. The calibration process can use a special calibration plate or calibration site to solve the extrinsic parameters of the sensor by measuring the relative position relationship between the sensor and the calibration plate or site.

[0046] Step S102, preprocessing the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of static scene point clouds;

[0047] Here, the static scene point cloud can represent the three-dimensional shape and structure of the static environment around the current vehicle.

[0048] In some embodiments, the multiple frames of original point clouds are preprocessed using the original surround images and the internal and external parameters of the sensor to remove dynamic objects (e.g., other vehicles, pedestrians, etc.) in the multiple frames of original point clouds and retain static objects (e.g., roads, buildings, etc.) to obtain multiple frames of static scene point clouds.

[0049] Step S103, analyzing the illumination information obtained after the illumination estimation of the original surrounding image, and obtaining the illumination color distribution and transparency information of the scene in the surrounding image;

[0050] Here, illumination estimation is an important task in computer vision, which aims to infer the lighting conditions of the scene from the image. This usually includes estimating the location, intensity, color and other properties of the light source. For the original panoramic image, illumination estimation can help us understand the illumination conditions of different areas in the original image, and provide a basis for the subsequent color distribution and transparency information analysis. The illumination color distribution can be the distribution of illumination colors in different areas of the image. Transparency information can be the ability of objects in the image to transmit light, which can help us understand the material and occlusion relationship of objects in the scene.

[0051] In some embodiments, based on the illumination estimation, analyzing the color distribution in the surrounding image may include: converting the image from a red green blue (RGB) color space to a color space that is more suitable for color distribution analysis, which can more intuitively represent color attributes such as brightness, hue, and saturation; calculating the histogram of each color space in the image to understand the distribution of different colors in the image, and the histogram can display the frequency of occurrence of each color value in the image, thereby helping us identify the main illumination color; using a clustering algorithm (such as K-means) or a segmentation algorithm to cluster or segment the colors in the image to identify areas with similar illumination colors, which helps to more accurately understand the illumination color distribution in the scene.

[0052] In some embodiments, based on the illumination estimation, analyzing the transparency information in the surrounding image may include: segmenting the image into different regions or objects to analyze the transparency of each region or object; estimating the transparency of each region or object using an image processing algorithm (e.g., a physically based rendering algorithm or a deep learning algorithm); and correcting the image based on the estimated transparency information to remove color distortion and shadows caused by illumination and material.

[0053] Step S104, rendering the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of panoramic images;

[0054] Here, the illumination color distribution represents the intensity and color distribution of illumination in different areas of the scene. The illumination color distribution is crucial for simulating the illumination effects in real scenes, because it can affect the brightness, color, and shadows of the image. Transparency information represents the ability of objects in the scene to transmit light. Transparency information is very important for generating images with a sense of depth and realism, because it can affect the occlusion relationship between objects and the way light propagates.

[0055] In some embodiments, each point in the multi-frame static scene point cloud is colored according to the illumination color distribution information. The coloring process needs to consider the position, direction and color of the light source, as well as the material and reflection properties of the object surface. According to the transparency information, a transparency effect is added to each point cloud in the multi-frame static scene point cloud, so that some point clouds have different degrees of transparency when rendered. Transparency processing can simulate the occlusion relationship between objects and the penetration effect of light, and enhance the realism of the image.

[0056] In some embodiments, in order to generate a multi-frame panoramic image, it is necessary to change the rendering perspective or camera position. This can be achieved by rotating, translating or zooming the camera to generate images of the scene from different angles. Then, the multi-frame static scene point cloud after coloring and transparency processing is projected onto a two-dimensional image plane to form a multi-frame panoramic image.

[0057] Step S105 , using the loss between the multiple frames of surrounding images and the original surrounding images, the parameters of the static scene model are updated until the static scene model converges.

[0058] Here, the loss can be a metric that measures the difference between the multi-frame panoramic image and the original panoramic image.

[0059] In some embodiments, features are extracted from the multi-frame surrounding view images and the original surrounding view images respectively. These features may be pixel values, edges, textures, shapes, etc. The purpose of feature extraction is to convert the image into a representation form that is easier to compare. The similarity between the multi-frame surrounding view images and the original surrounding view images in the feature space is calculated. The similarity may be measured by calculating the distance between the feature vectors. Based on the result of the similarity calculation, a loss function is defined to quantify the difference between the multi-frame surrounding view images and the original surrounding view images. The loss function may be mean square error, peak signal-to-noise ratio, structural similarity, etc.

[0060] In some embodiments, after calculating the loss between the multi-frame panoramic image and the original panoramic image, an optimization algorithm can be used to update the parameters of the static scene model to reduce the loss and improve the performance of the static scene model. For example, the gradient of the loss function with respect to the static scene model parameters is calculated. The gradient indicates the rate of change of the loss function in the parameter space, so the gradient indicates how to adjust the parameters to reduce the loss; based on the gradient information, an optimization algorithm (e.g., gradient descent, stochastic gradient descent, etc.) is used to adjust the parameters of the static scene model. The purpose of parameter adjustment is to minimize or approach the minimum loss; repeat the above gradient calculation and parameter adjustment process until the static scene model converges. Convergence means that the loss no longer decreases significantly, that is, the loss value reaches an acceptably low level, indicating that the static scene model is able to generate a rendered image that is close enough to the real world.

[0061] In the embodiments of the present application, on the one hand, by acquiring multiple frames of original point clouds and original surrounding images, as well as corresponding internal and external parameters of the sensor, a static scene can be accurately reconstructed; on the other hand, by analyzing the illumination information obtained after the illumination estimation of the original surrounding image, the illumination color distribution and transparency information of the scene in the surrounding image can be obtained, so that the scene can be rendered more realistically, making the generated surrounding image closer to the real scene.

[0062] In some embodiments, the implementation of step S102 "preprocessing the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of static scene point clouds" may include the following steps S111 to S113, wherein:

[0063] Step S111, removing dynamic point clouds from the multiple frames of original point clouds to obtain multiple frames of static point clouds;

[0064] Here, in the multi-frame original point cloud, dynamic point clouds generated by moving objects (e.g., vehicles, pedestrians, etc.) may be included. These dynamic point clouds are unnecessary for constructing a static scene model.

[0065] In some embodiments, first, each frame of the original point cloud is segmented to distinguish different objects or regions; this can be achieved through clustering algorithms, methods based on geometric features, or deep learning methods; then, the point clouds belonging to dynamic objects (e.g., other vehicles, pedestrians, etc.) are identified by analyzing the motion patterns of the point clouds (e.g., speed, acceleration, trajectory, etc.); once the dynamic objects are identified, their corresponding point clouds can be removed from the original point clouds. In this way, the remaining point clouds are mainly composed of static objects (e.g., buildings, trees, roads, etc.), forming static point clouds.

[0066] Step S112, coloring the multiple frames of static point clouds based on the original panoramic images and the internal and external parameters of the sensor to obtain multiple frames of colored static point clouds;

[0067] Here, the original point cloud usually only contains spatial position information but not color information. In order to make the point cloud more intuitive and easier to understand, it needs to be colored.

[0068] In some embodiments, the color information in the original surrounding image can be mapped to the corresponding point cloud according to the original surrounding image and the internal and external parameters of the sensor, which usually involves the conversion of image coordinates to point cloud coordinates and the assignment of color values.

[0069] In some embodiments, first, the original panoramic image needs to be registered with the corresponding static point cloud. Since the image and point cloud are usually acquired in different coordinate systems (the image is in the pixel coordinate system, and the point cloud is in the three-dimensional space coordinate system), it is necessary to convert them to the same coordinate system through the internal and external parameters of the sensor for subsequent color mapping; then, after the registration is completed, the point cloud can be colored based on the color information in the image, which usually involves mapping the color value of each pixel in the image to the corresponding point cloud; since this process is performed on multiple frames of static point clouds, the above-mentioned color mapping processing needs to be performed on each frame of point clouds. After the above-mentioned processing steps, multiple frames of colored static point clouds can be obtained, and the multi-frame colored static point clouds not only contain the spatial position information of the object, but also contain rich color information.

[0070] Step S113, converting the coordinate system of the multi-frame colored static point cloud into a world coordinate system to obtain the multi-frame static scene point cloud.

[0071] In some embodiments, after obtaining the static point cloud after multiple frames are colored, it is necessary to convert its coordinate system into the world coordinate system. This is because the point clouds of different frames may be acquired in different local coordinate systems, and in order to build a globally consistent static scene model, these point cloud data need to be unified into the same world coordinate system. Coordinate system conversion usually involves rotation and translation transformations, and these transformation parameters can be determined by the internal and external parameters of the sensor.

[0072] In some embodiments, in order to transform the colored static point cloud of multiple frames from the local coordinate system to the world coordinate system, it is necessary to apply the internal and external parameters of the sensor to calculate the transformation matrix from the sensor coordinate system to the world coordinate system; then, using the transformation matrix, each frame of the colored point cloud can be transformed from its original local coordinate system to the world coordinate system to obtain each frame of the static scene point cloud. This transformation process involves the transformation of the three-dimensional coordinates of the point, which usually includes rotation and translation operations.

[0073] It should be noted that since the above coordinate system conversion process is performed on the static point cloud after multi-frame coloring, it is necessary to perform the above coordinate conversion on the static point cloud after each frame coloring. The converted multi-frame static scene point cloud will share the same world coordinate system, so it can be integrated into a globally consistent static scene point cloud.

[0074] In the embodiments of the present application, on the one hand, by removing dynamic point clouds in multi-frame original point clouds, such as moving vehicles, pedestrians, etc., the accuracy of multi-frame static point clouds can be improved; on the other hand, by coloring multi-frame static point clouds according to the original surrounding images and internal and external parameters of the sensor, the multi-frame colored static point clouds can more intuitively reflect the colors of objects in the scene; on the other hand, by converting the coordinate system of multi-frame colored static point clouds into the world coordinate system, the coordinates of point clouds in different frames can be unified, which is conducive to the splicing and fusion of multi-frame static scene point clouds.

[0075] In some embodiments, the implementation of step S111 “removing dynamic point clouds from the multiple frames of original point clouds to obtain multiple frames of static point clouds” may include the following steps S121 to S123, wherein:

[0076] Step S121, compensating the multiple frames of original point clouds to obtain multiple frames of compensated point clouds;

[0077] In some embodiments, since the current vehicle will constantly change its position and posture during the motion process, the collected multi-frame original point cloud data may have a real position offset problem. In order to eliminate the real position offset problem caused by the current vehicle motion, motion compensation is performed on the original multi-frame original point cloud to obtain a multi-frame compensated point cloud.

[0078] Step S122, performing three-dimensional target detection on the multi-frame compensated point cloud to obtain a dynamic point cloud in the multi-frame compensated point cloud;

[0079] Here, three-dimensional target detection may be to identify and locate target objects of interest in three-dimensional space, such as vehicles, pedestrians, etc.

[0080] In some embodiments, a three-dimensional target detection model can be used to detect moving targets in the point cloud after multi-frame compensation to identify the dynamic point cloud (i.e., the point cloud of moving objects) therein, so as to subsequently remove moving targets or movable static targets to ensure that each frame of point cloud only contains static scene point clouds.

[0081] Step S123: removing the dynamic point cloud from the multi-frame compensated point cloud to obtain the multi-frame static point cloud.

[0082] In some embodiments, after the dynamic point cloud is identified, it is necessary to remove the dynamic point cloud from the multi-frame compensated point cloud to obtain a multi-frame static point cloud. This step usually involves operations such as segmentation and filtering of the point cloud to ensure that the integrity and accuracy of the static point cloud are retained while removing the dynamic point cloud. After removing the dynamic point cloud, the obtained multi-frame static point cloud will only contain information of static scenes, such as buildings, roads, trees, etc.

[0083] In the embodiments of the present application, on the one hand, by compensating the multi-frame original point cloud, the problem of real position deviation caused by the current vehicle movement can be eliminated; on the other hand, by removing dynamic point clouds in the multi-frame original point cloud, such as moving vehicles, pedestrians, etc., the accuracy of the multi-frame static point cloud can be improved.

[0084] In some embodiments, the implementation of step S112 "coloring the multiple frames of static point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of colored static point clouds" may include the following steps S131 to S133, wherein:

[0085] Step S131, determining the position information of the current vehicle based on the positioning data, acceleration and angular velocity of the current vehicle;

[0086] Here, in three-dimensional space, the pose information of an object is usually determined by its position (coordinates) and attitude (orientation). For a vehicle, the pose information describes the exact position (e.g., positioning data) and orientation (e.g., yaw angle, pitch angle, roll angle) of the vehicle in three-dimensional space. The positioning data of the current vehicle is collected by the Global Positioning System (GPS). The acceleration and angular velocity of the current vehicle are collected by the Inertial Measurement Unit (IMU).

[0087] In some embodiments, by combining the positioning data, acceleration and angular velocity of the current vehicle, the position information (including positioning data and orientation) of the current vehicle can be determined, thereby determining the precise position and posture of the current vehicle at each time point.

[0088] Step S132, dedistorting the original surrounding image based on the internal and external parameters of the sensor to obtain a dedistorted surrounding image;

[0089] Here, de-distortion can be to correct the image distortion caused by lens distortion (e.g., barrel distortion or pincushion distortion). This is necessary because the distortion affects the true shape and position of objects in the image, thereby affecting the accuracy of subsequent color mapping.

[0090] In some embodiments, a distortion correction algorithm may be used to transform the original panoramic image according to the distortion coefficient in the intrinsic and extrinsic parameters of the camera to eliminate the image distortion caused by the camera lens distortion. After de-distortion, the de-distorted panoramic image obtained more accurately reflects the scene captured by the camera, reducing the impact of distortion on image quality.

[0091] Step S133, coloring the multiple frames of static point clouds based on the current vehicle's position information, the internal and external parameters of the sensor, and the dedistorted surrounding image to obtain the multiple frames of colored static point clouds.

[0092] Here, the dedistorted surround image is used as a color source to provide color information for the multi-frame static point cloud. To ensure the accuracy of color mapping, the current vehicle pose information and the internal and external parameters of the sensor need to be considered. These parameters help determine the exact correspondence between the image and the point cloud.

[0093] In some embodiments, using the current vehicle's posture information and the sensor's internal and external parameters, each point cloud in the multi-frame static point cloud is projected onto the dedistorted panoramic image to find the pixel points corresponding to each point cloud; for each corresponding pixel point found in the static point cloud, its color information is extracted; and the extracted color information is assigned to the corresponding point cloud, thereby coloring the point cloud. After performing the above processing on all point clouds, a multi-frame colored static point cloud is obtained. The multi-frame colored static point cloud contains not only three-dimensional coordinate information, but also rich color information.

[0094] In the embodiment of the present application, on the one hand, by combining the current vehicle posture information and the internal and external parameters of the sensor, the pixel points in the original surrounding image can be accurately matched with the static point cloud; on the other hand, the colored static point cloud can more intuitively display the three-dimensional structure and color characteristics of objects in the environmental scene, so that users can more easily understand the three-dimensional coordinate information and rich color information expressed by the point cloud.

[0095] In some embodiments, the implementation of step S133 "coloring the multiple frames of static point clouds based on the current vehicle's posture information, the internal and external parameters of the sensor, and the dedistorted surrounding image to obtain the multiple frames of colored static point clouds" may include the following steps S141 and S142, wherein:

[0096] Step S141, projecting the multiple frames of static point clouds onto the dedistorted surrounding image based on the current vehicle posture information and the internal and external parameters of the sensor to obtain colors corresponding to the multiple frames of static point clouds;

[0097] In some embodiments, each point cloud in multiple frames of static point clouds is projected onto the dedistorted surrounding image using the current vehicle's posture information and the sensor's internal and external parameters, and the pixel points corresponding to each point cloud are found; for each corresponding pixel point found in the static point cloud, its color information is extracted.

[0098] Step S142, coloring the multiple frames of static point clouds using the colors corresponding to the multiple frames of static point clouds to obtain the multiple frames of colored static point clouds.

[0099] In some embodiments, the extracted color information is assigned to the corresponding point cloud, thereby completing the point cloud coloring process. The coloring process requires traversing all point clouds to ensure that each point cloud is assigned the correct color. After the above processing, the obtained colored multi-frame static point cloud contains not only three-dimensional coordinate information, but also rich color information.

[0100] In the embodiment of the present application, on the one hand, by combining the current vehicle posture information and the internal and external parameters of the sensor, the pixel points in the original surrounding image can be accurately matched with the static point cloud; on the other hand, the colored static point cloud can more intuitively display the three-dimensional structure and color characteristics of objects in the environmental scene, so that users can more easily understand the three-dimensional coordinate information and rich color information expressed by the point cloud.

[0101] In some embodiments, the implementation of step S113 "converting the coordinate system of the multi-frame colored static point cloud to the world coordinate system to obtain the multi-frame static scene point cloud" may include the following steps S151 and S152, wherein:

[0102] Step S151, for the multiple frames of colored static point clouds within a period of time, obtaining the position and posture information of the current vehicle corresponding to the multiple frames of colored static point clouds;

[0103] Here, the multi-frame colored static point cloud may be a multi-frame static point cloud that has been colored, and each frame of the static point cloud represents environmental information captured by the current vehicle at a certain moment.

[0104] In some implementations, first, a time range is determined, which covers all the colored static point clouds that need to be processed; then, for each colored static point cloud frame within the time range, the corresponding position information of the current vehicle is obtained. With the correct position information, point clouds captured at different time points can be aligned to the same world coordinate system.

[0105] Step S152, based on the position information of the current vehicle, convert the coordinate system of the multiple frames of colored static point clouds into a world coordinate system to obtain multiple frames of static scene point clouds.

[0106] Here, the world coordinate system can be a fixed, global reference coordinate system used to uniformly represent point clouds captured at different times and locations.

[0107] In some embodiments, the static point cloud after each frame of coloring is converted from its respective local coordinate system to the world coordinate system using the acquired position information of the current vehicle, thereby obtaining a multi-frame point cloud representing the static scene. This process involves complex coordinate transformation calculations, including operations such as translation and rotation. Among them, according to the translation information (i.e., position) of the current vehicle, each point cloud is moved in the corresponding direction so that its position in the world coordinate system matches the actual position of the current vehicle. According to the rotation information (i.e., orientation) of the current vehicle, each point cloud is rotated around the corresponding axis so that its orientation in the world coordinate system matches the actual orientation of the current vehicle.

[0108] In an embodiment of the present application, by converting the coordinate system of the multi-frame colored static point cloud to the world coordinate system, the data inconsistency problem caused by different coordinate systems is avoided. The multi-frame static scene point cloud obtained after the coordinate system conversion can more intuitively display the environmental scene structure, including roads, buildings, obstacles, etc., which helps the vehicle to better understand the surrounding environment scene.

[0109] In some embodiments, the implementation of step S103 of "analyzing the illumination information obtained after the illumination estimation of the original surrounding image to obtain the illumination color distribution and transparency information of the scene in the surrounding image" may include the following steps S161 and S162, wherein:

[0110] Step S161, performing illumination estimation on the original surrounding image to obtain illumination information of the surrounding image;

[0111] In some embodiments, illumination estimation is performed on the original surrounding image, and illumination estimation generally involves analyzing the properties of the color, brightness, and contrast of the pixels in the original surrounding image. By analyzing these properties, illumination information of the original surrounding image can be inferred, and the illumination information describes the conditions and characteristics of the illumination in the original surrounding image, and the illumination information may include parameters such as the intensity, direction, and color of the illumination.

[0112] Step S162: Analyze the illumination information of the surrounding image to obtain illumination color distribution and transparency information of the scene in the surrounding image.

[0113] In some embodiments, the surrounding image is further analyzed based on the illumination information obtained by the illumination estimation. The purpose of the analysis is to extract the illumination color distribution and transparency information of the scene in the surrounding image. The illumination color distribution can be obtained by clustering or segmenting the color features of different regions in the surrounding image. The transparency information can be estimated by analyzing the brightness changes and color mixing degree of pixels in the surrounding image. For example, a translucent object may cause the color of the background light to mix with the color of the foreground object, thereby forming a unique color feature.

[0114] In the embodiments of the present application, on the one hand, by performing illumination estimation on the original surrounding image, the illumination information in the original surrounding image can be comprehensively acquired, including illumination intensity, direction, color, etc.; on the other hand, by analyzing the illumination information of the surrounding image, the illumination color distribution and transparency information obtained can reflect the color of the light source in the scene, as well as the translucency or degree of occlusion of the object.

[0115] At present, autonomous driving technology has become a research hotspot in the automotive field, and autonomous driving algorithm testing is particularly important. Before the algorithm is deployed in the actual vehicle system, a large number of reliability tests are usually carried out in a simulation environment, and the authenticity and diversity of the simulation environment become very critical. How to ensure the authenticity and diversity of the simulation environment has always been a difficult problem in the industry.

[0116] Generally, there are two sources of simulation environment: one is the environment that comes with the virtual simulation software, which can make the environment diverse and rich through specific parameters and settings, but its scene elements are relatively simple and the textures are not detailed enough, which is very different from the actual road scene. Its unreality makes it difficult for the algorithm to perform consistently in the simulation environment and in the actual scene; the other is the environment built by professional engineers through scene reconstruction algorithms based on real vehicle collection data. It can better restore the real driving scene, making the algorithm simulation effect more consistent between the simulation environment and the actual scene, but the environment of a single real vehicle collection is single, and the weather, lighting, etc. are uncontrollable. If the same scene is collected by real vehicles multiple times, it will be costly and time-consuming.

[0117] In the related technology, the image collected by the on-board camera is used as the main input. The features of the visual image are first extracted through a convolutional neural network. At the same time, based on the differentiable homography transformation, the image features are used to construct a matching cost space. The matching cost space is regularized using a 3D convolutional neural network to predict the depth map. The image is semantically classified in combination with the semantic segmentation results of the two-dimensional image. Dynamic targets are removed in the multi-view fusion stage to reconstruct the static background. The optional input lidar point cloud data is used to predict the depth map for multi-view fusion. Figure 1The above-mentioned prediction of the depth map through a neural network is prone to introduce errors, resulting in inaccurate relative positions of the reconstructed scene elements, making it difficult to achieve high fidelity; and the color comes from the image at the corresponding moment, with a single illumination.

[0118] In the related art, a two-dimensional image is first input, and then the two-dimensional image is analyzed for illumination to determine its main illumination information and ambient illumination information, and finally the main illumination of the three-dimensional scene is adjusted according to the main illumination information, so that the illumination effect of the three-dimensional scene is adjusted to be consistent with the illumination effect of the two-dimensional image, and then the three-dimensional scene is rendered for illumination. This method renders the three-dimensional scene based on the color of the two-dimensional image, thereby restoring the real illumination environment, and also lacks the diversity of the illumination environment.

[0119] In view of the shortcomings of inaccurate relative positions of reconstructed static elements and single lighting in related technologies, the embodiments of the present application propose a high-fidelity static scene reconstruction method with controllable lighting based on a neural network. On the one hand, the optimized posture is used to splice all processed static scene point clouds as the input of scene reconstruction to ensure the accuracy of the relative positions of static elements; on the other hand, different lighting characteristic coefficients are input, and after being processed by the neural network model and algorithm, multiple sets of lighting conditions are inferred to enrich the lighting diversity of the reconstructed scene and make up for the defect of single lighting. Therefore, for the same scene, the embodiments of the present application only need to control the laser radar point cloud unchanged and change the lighting characteristic coefficient, so as to output multiple sets of scene images under different lighting conditions, so as to ensure the realistic restoration of the scene while enriching the lighting diversity of the constructed simulation environment, thereby improving the effect of autonomous driving algorithm testing.

[0120] The embodiment of the present application proposes a high-fidelity static scene reconstruction method with controllable illumination based on a neural network. By changing different illumination characteristic coefficients and combining pre-processed static scene point cloud data, the same scene can be reconstructed under different illumination conditions, and the scene under different weather, time, and illumination conditions can be restored with high fidelity. Figure 2 As shown in the figure, the static scene reconstruction model mainly includes the following modules:

[0121] The input module 210 is used to collect the original data of the vehicle, including camera images of six perspectives (directly in front of the front of the vehicle, the front left side, the front right side, directly behind the rear of the vehicle, the rear left side, and the rear right side), the original point cloud collected by the lidar point cloud, the internal and external parameters of the sensor, etc.

[0122] The data preprocessing module 220 is used to improve the shortcomings of using Structure for Motion (SfM) points as input. Specifically, a method of combining two-dimensional images and three-dimensional point clouds and decoupling dynamic and static is used to preprocess the original data. It should be noted that the embodiment of the present application reconstructs static scenes based on a three-dimensional Gaussian Splatting (3DGS) module as a framework. The 3DGS module uses SfM points as input for rendering, and the SfM points estimate the position of 3D points based on a sparse corresponding set of multiple images and image features. Since two-dimensional images lack depth information, position estimation is prone to accumulated errors, resulting in a lack of integrity and robustness in scene reconstruction.

[0123] The data preprocessing module in the embodiment of the present application performs fine distortion removal on the collected original image from the source. By accurately applying the camera's internal parameters and distortion coefficients, the data preprocessing module ensures that each frame of the image can be presented in the best state, not only eliminating the impact of distortion on image quality, but also providing accurate pixel correspondence for the subsequent point cloud coloring step, so that the reconstructed scene can be closer to the real world in color and detail.

[0124] In order to overcome the distortion problem of LiDAR point cloud data caused by the movement of the collection vehicle, advanced self-vehicle motion compensation technology is introduced. By integrating real-time data from high-precision sensors such as GPS and IMU, the trajectory and posture changes of the collection vehicle during movement can be accurately calculated, and each frame of point cloud data can be accurately corrected accordingly. This process not only eliminates motion distortion in the point cloud, but also ensures the consistency and continuity between different frames of point cloud data, providing a reliable data source for subsequent point cloud stitching and scene reconstruction.

[0125] In terms of eliminating dynamic point clouds, the trained 3D object detection model is used to accurately identify dynamic objects (such as pedestrians, vehicles, etc.) in the point cloud and effectively separate them from the static scene, which not only avoids the interference of dynamic objects on the reconstruction of static scenes, but also improves the accuracy and stability of scene reconstruction. At the same time, through the decoupling of dynamic and static processing, the data preprocessing module further enhances the flexibility and scalability of scene reconstruction.

[0126] After acquiring static point cloud data, the data preprocessing module uses the internal and external parameters of the sensor and the optimized posture information to convert multiple frames of static point cloud data into a unified world coordinate system and integrate them into a complete scene point cloud. This process not only eliminates the blind spot problem caused by viewing angle limitations, but also significantly improves the density and coverage of the scene point cloud.

[0127] In addition to the fine processing of static scene point cloud data, the data preprocessing module in the embodiment of the present application also fully considers the impact of lighting on the scene reconstruction effect. Through the lighting estimation algorithm, the data preprocessing module can analyze the collected image information and accurately extract the lighting data of the current scene. These lighting data include multiple key parameters such as lighting position, brightness, color temperature, color rendering, light intensity, and luminous flux, which together constitute a complete description of the scene lighting environment. In the three-dimensional reconstruction process, these lighting data are introduced as control variables of the neural network to guide the lighting rendering step.

[0128] The multi-layer perceptron module 230 is used to intelligently infer the corresponding scene color distribution and the display opacity of each part of the input various lighting characteristic coefficients (such as lighting position, brightness, color temperature, color rendering, light intensity, luminous flux, etc.) through deep learning.

[0129] Specifically, the multi-layer perceptron module first receives a series of illumination characteristic coefficients from the illumination estimation algorithm or user input as input. These coefficients accurately describe the current or expected illumination environment and are key parameters that are indispensable in the rendering process. Subsequently, multiple hidden layers inside the multi-layer perceptron module process the input data layer by layer through nonlinear activation functions (such as ReLU, Sigmoid, etc.), and gradually extract the complex relationship between illumination characteristics and scene color and opacity. In this process, each layer of neurons performs weighted summation and activation operations based on the output of the previous layer to capture different features in the input data. As the data is transmitted layer by layer, the multi-layer perceptron module gradually builds a mapping relationship from illumination characteristics to scene visual representation (including color and opacity). Finally, the multi-layer perceptron module outputs multiple sets of scene color and opacity data corresponding to different illumination conditions. The 3DGS reconstruction module uses the scene color and opacity data to infer and render static scene point clouds. By combining the geometric information of the point cloud and the scene color and opacity data output by the multi-layer perceptron module, the 3DGS reconstruction module can simulate the propagation and interaction of light in the scene, thereby generating high-fidelity static scene images under different lighting conditions.

[0130] The three-dimensional Gaussian module (i.e., 3DGS reconstruction module) 240 is used to reconstruct and render the static environment, and inputs the spliced ​​static scene point cloud and the scene color and opacity learned by the multi-layer perceptron module to predict the camera images of the six perspectives corresponding to the current scene. During the training process, the camera images of the six perspectives and the dedistorted images obtained by the preprocessing module are used to calculate the loss, and the parameters of the multi-layer perceptron module and the 3DGS reconstruction module are continuously updated through the back propagation algorithm, so that the reconstructed scene and color are restored more realistically.

[0131] The output module 250 is used to output camera images of six viewing angles of a scene at a certain fixed frequency, and the specific frequency is the same as the input image and lidar point cloud frequency. For example, if the training data is 10 Hz, the output image is 10 frames per second.

[0132] The present application embodiment proposes a high-fidelity static scene reconstruction method with controllable illumination based on a neural network, such as Figure 3 As shown, it includes the following steps S301 to S307, wherein:

[0133] Step S301, installation and layout of sensors on the vehicle;

[0134] The embodiment of the present application requires that the collection vehicle has good collection equipment (six panoramic cameras), which are located directly in front of the front of the vehicle, the front left side, the front right side, the rear of the vehicle, the rear left side, and the rear right side, so as to facilitate better capture of the surrounding panoramic view. At least one high-frequency laser radar is used to capture 360° panoramic point clouds in real time, such as installing a Pandar128 laser radar on the top of the collection vehicle; a GPS and IMU are used for better real-time positioning, and the installation position can also be on the top of the collection vehicle. Before collecting data on the actual vehicle, it is necessary to determine the coordinate system of each sensor device (laser radar, camera), and calibrate and verify the internal and external parameters of the sensor to ensure that the collected data is more accurate.

[0135] Step S302, collecting raw data of vehicles in different scenarios;

[0136] The embodiment of the present application requires multiple acquisitions of at least one scene of the vehicle to obtain real panoramic data of the same scene under different lighting conditions, such as once in the morning, noon, evening, sunny day, and rainy day, to verify the controllable lighting function of the embodiment of the present application. Other scenes are normally acquired to obtain a large amount of raw data for training to achieve controllable lighting. The collected raw data is saved in units of one minute as a file with the suffix dat.

[0137] Step S303, preprocessing the original data to obtain a static scene point cloud;

[0138] In order to obtain a complete and robust initial static scene point cloud, the present application embodiment proposes a data preprocessing process, such as Figure 4 As shown, it includes the following steps S401 to S407, wherein:

[0139] Step S401 , performing point cloud compensation on the original point cloud 401 to obtain a compensated point cloud 402 , so as to compensate for the problem that the real position of the collected point cloud is offset due to the movement of the vehicle.

[0140] Step S402 , using a 3D target detection model to perform 3D target detection on the compensated point cloud 402 to obtain a motion point cloud 403 .

[0141] Step S403 , removing the moving point cloud 403 in the compensated point cloud 402 to obtain the static point cloud 404 , ensuring that each frame of point cloud only contains the static point cloud.

[0142] Step S404, optimize the posture of GPS and IMU positioning data 405 (vehicle positioning data collected by GPS and vehicle acceleration and angular velocity measured by IMU) through a posture optimization algorithm to obtain vehicle posture information 406, so as to better obtain the precise positioning of each frame of point cloud.

[0143] Step S405, dedistort the camera images 408 of six viewing angles by using the intrinsic and extrinsic parameters 407 and the distortion coefficients of the sensor to obtain dedistorted images 409. In order to make the 3DGS reconstruction module better learn the original color of the static scene point cloud.

[0144] Step S406 , using the internal and external parameters 407 and the position and posture information 406 of the sensor, the static point cloud 404 is projected onto the dedistorted image 409 to colorize the static point cloud 404 , thereby obtaining a colored static point cloud 410 .

[0145] Step S407 , the coordinate system of the colored static point cloud 410 is converted to the world coordinate system (ie, multiple frames of static point cloud scene are stitched together) to form a complete colored static scene point cloud 411 , and save it as a file with a suffix of ply.

[0146] It should be noted that in order to train the multi-layer perceptron module and infer the illumination characteristic coefficients into the ability of attributes such as RGB color coefficients and opacity, the embodiment of the present application needs to perform illumination estimation based on the timestamp and the corresponding image as well as weather information. For example, an illumination estimation algorithm based on image decomposition is used to obtain illumination information from the dedistorted image, extract illumination characteristic coefficients such as illumination position, brightness, color temperature, light intensity, and luminous flux, and store them in a certain file format such as txt text as part of the input of the 3DGS reconstruction module. Static scene point clouds and dedistorted images are both in units of clips. The data of a single clip is all data within 20 seconds, including camera images of six perspectives, corresponding point cloud data, and corresponding sensor parameters. The frequency is 10hz, and the final training data is 1000clips and above.

[0147] Step S304, building a multi-layer perceptron module;

[0148] The multilayer perceptron module includes an input layer, a hidden layer, and an output layer. The input of the input layer is a plurality of illumination characteristic coefficients. According to the experimental data, if only the position (x, y, z) and brightness (L) of the illumination are processed, the dimension of the input layer should be set to 4. In the neural network framework (e.g., TensorFlow, PyTorch, etc.), this can be achieved by defining the number of nodes in the input layer as 4. The number of hidden layers and the number of nodes in each layer (also called the number of neurons) can be determined by experiments and tuning. In general, more hidden layers and nodes can increase the complexity of the multilayer perceptron module, thereby improving the fitting ability of the multilayer perceptron module, but may also lead to overfitting and increased training time. The number of layers can be tried from 1 layer and gradually increased to observe the changes in the performance of the multilayer perceptron module. The number of nodes in each layer can be selected according to the dimensions of the input and output and the complexity of the problem. In general, the number of nodes in each layer can be set to a multiple of the number of nodes in the input layer or a multiple of the number of nodes in the output layer. In the multilayer perceptron module, the activation function (e.g., ReLU, Sigmoid, and Tanh, etc.) is crucial for introducing nonlinearity. For hidden layers, ReLU is usually a good choice because it can speed up the training process and reduce the problem of gradient disappearance. For the output layer, since the output is the color coefficient (i.e., spherical harmonic function coefficient) and opacity, these values ​​may need to be within a certain range (such as between 0 and 1), so you may need to use Sigmoid or Tanh activation functions to ensure that the output value is in a suitable range. However, if the color coefficient does not need to be restricted to a specific range, ReLU or linear activation functions can also be considered. The output layer needs to output color coefficients (represented by 3rd-order spherical harmonic functions, a total of 27 coefficients) and opacity (1 value), so the dimension of the output layer should be set to 28.

[0149] Step S305, building a 3DGS reconstruction module;

[0150] The present application embodiment proposes a method for generating a panoramic picture with six viewing angles, such as Figure 5As shown, the static scene point cloud 411 is initialized to obtain a 3D Gaussian sphere 501 (with size and orientation) corresponding to each point cloud. During the training process, the number, size and orientation of the ellipses will change. In order to be closer to the original image, the 3D Gaussian sphere 502 is projected (Projection) on the original six-view camera image 502. The 3D Gaussian sphere 501 is converted into a panoramic image 504 of six perspectives according to the spherical harmonic coefficients and opacity 503 through differentiable tile rasterizer. The loss between the panoramic image 504 of six perspectives and the original six-view camera image 502 is calculated, and the loss is used to update (Adaptive Density Control) the 3D Gaussian sphere 501, so that the final output panoramic image of six perspectives is closer to the original camera image of six perspectives.

[0151] Step S306, training the multi-layer perceptron module and the 3DGS reconstruction module;

[0152] The multi-layer perceptron module and the 3DGS reconstruction module are connected in sequence, and the illumination estimation algorithm is used to estimate the illumination of the dedistorted image, and the illumination characteristic coefficients are obtained. The illumination characteristic coefficients are input into the multi-layer perceptron module, and the color coefficients and opacity are output. The color coefficients and opacity and the point cloud data output by the data preprocessing module are used as inputs into the 3DGS reconstruction module, and multi-frame multi-view images are forward reconstructed and rendered. The multi-frame multi-view images and the dedistorted images are used to calculate the loss, and the coefficients in the multi-layer perceptron module and the 3DGS reconstruction module are updated in the form of back propagation. After a large amount of data training, when the loss converges to a certain threshold, the static scene model completes iterative optimization and saves the structure and coefficients of the multi-layer perceptron module and the 3DGS reconstruction module, so that the entire static scene model has good scene reconstruction and rendering capabilities.

[0153] Step S307, using the multi-layer perceptron module and the 3DGS reconstruction module to render the scene image.

[0154] The use process of the multi-layer perceptron module and the 3DGS reconstruction module to render the scene image, such as Figure 6 As shown, the static scene point cloud 411 and the illumination characteristic coefficient 601 output by the data preprocessing module are input into the three-dimensional Gaussian module 240 via the color coefficient and opacity output by the multi-layer perceptron module 230 to render a panoramic picture 602 of the same scene under different illumination conditions.

[0155] The embodiment of the present application proposes a high-fidelity static scene reconstruction method with controllable illumination based on a neural network. On the one hand, a good initial static scene point cloud is obtained through a data preprocessing module to replace the input of the original 3DGS reconstruction module, thereby compensating for the incompleteness of the SfM points and reducing the estimated cumulative error. On the other hand, a multi-layer perceptron module is used to increase the controllable function of the illumination coefficient, enrich the illumination conditions of the static scene point cloud, and further compensate for the deficiency of the single illumination of the existing scene reconstruction.

[0156] The present application embodiment proposes a training device for a static scene model, such as Figure 7 As shown, the training device 700 of the static scene model includes:

[0157] An input module 210 is used to obtain multiple frames of original point clouds, original surrounding images, and internal and external parameters of sensors corresponding to the original surrounding images of the current vehicle in different scenes;

[0158] A data preprocessing module 220 is used to preprocess the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of static scene point clouds;

[0159] The multi-layer perceptron module 230 is used to analyze the illumination information obtained after the illumination estimation of the original surrounding image, and obtain the illumination color distribution and transparency information of the scene in the surrounding image;

[0160] A three-dimensional Gaussian module 240 is used to render the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of panoramic images;

[0161] The output module 250 is used to update the parameters of the static scene model by using the loss between the multiple frames of surrounding images and the original surrounding images until the static scene model converges.

[0162] In some embodiments, the data preprocessing module 220 includes: a removal unit, used to remove the dynamic point cloud in the multiple frames of original point cloud to obtain multiple frames of static point cloud; a coloring unit, used to color the multiple frames of static point cloud based on the original surrounding image and the internal and external parameters of the sensor to obtain multiple frames of colored static point cloud; a conversion unit, used to convert the coordinate system of the multiple frames of colored static point cloud into the world coordinate system to obtain the multiple frames of static scene point cloud.

[0163] In some embodiments, the removal unit includes: a compensation subunit, used to compensate the multi-frame original point cloud to obtain a multi-frame compensated point cloud; a detection subunit, used to perform three-dimensional target detection on the multi-frame compensated point cloud to obtain a dynamic point cloud in the multi-frame compensated point cloud; and a removal subunit, used to remove the dynamic point cloud from the multi-frame compensated point cloud to obtain the multi-frame static point cloud.

[0164] In some embodiments, the coloring unit includes: a determination subunit, used to determine the posture information of the current vehicle based on the positioning data, acceleration and angular velocity of the current vehicle; a dedistortion subunit, used to dedistort the original surrounding image based on the internal and external parameters of the sensor to obtain a dedistorted surrounding image; a coloring subunit, used to color the multiple frames of static point clouds based on the posture information of the current vehicle, the internal and external parameters of the sensor and the dedistorted surrounding image to obtain the multiple frames of colored static point clouds.

[0165] In some embodiments, a coloring subunit is used to project the multi-frame static point cloud onto the dedistorted surrounding image based on the current vehicle's posture information and the internal and external parameters of the sensor to obtain colors corresponding to the multi-frame static point cloud; and to color the multi-frame static point cloud using the colors corresponding to the multi-frame static point cloud to obtain the multi-frame colored static point cloud.

[0166] In some embodiments, the conversion unit includes: an acquisition subunit, used to obtain the posture information of the current vehicle corresponding to the multi-frame colored static point cloud within a period of time; a conversion subunit, used to convert the coordinate system of the multi-frame colored static point cloud into the world coordinate system based on the posture information of the current vehicle, so as to obtain the multi-frame static scene point cloud.

[0167] In some embodiments, the multi-layer perceptron module 230 includes: an illumination estimation unit, which is used to perform illumination estimation on the original surrounding image to obtain illumination information of the surrounding image; and an analysis unit, which is used to analyze the illumination information of the surrounding image to obtain illumination color distribution and transparency information of the scene in the surrounding image.

[0168] The description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiment of the present application can be used to execute the method described in the above method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.

[0169] It should be noted that in the embodiment of the present application, if the above method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software, and firmware.

[0170] An embodiment of the present application also provides an electronic device, including a memory and a processor of a server or a memory and a processor of a user terminal, wherein the memory stores a computer program that can be executed on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.

[0171] The embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, some or all of the steps in the above method are implemented. The computer-readable storage medium can be transient or non-transient.

[0172] An embodiment of the present application also provides a computer program, including a computer-readable code. When the computer-readable code runs in a server or a user terminal, a processor in the server or the user terminal executes some or all of the steps for implementing the above method.

[0173] The present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented specifically by hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium, and in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0174] It should be noted here that the description of the various embodiments above tends to emphasize the differences between the various embodiments, and the same or similar aspects can be referenced to each other. The description of the above device, storage medium, computer program and computer program product embodiments is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of this application, please refer to the description of the method embodiment of this application for understanding.

[0175] The present application embodiment provides an electronic device, such as Figure 8 As shown, the hardware entity of the electronic device 800 includes: a processor 801, a communication interface 802 and a memory 803, wherein: the processor 801 generally controls the overall operation of the electronic device 800. The communication interface 802 can enable the electronic device to communicate with other terminals or servers through a network. The memory 803 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or processed by the processor 801 and each module in the electronic device 800 (for example, image data, audio data, voice communication data and video communication data), which can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM). Data transmission can be carried out between the processor 801, the communication interface 802 and the memory 803 through the bus 804.

[0176] The above embodiments are only preferred embodiments for fully illustrating the present application, and the protection scope of the present application is not limited thereto. Any equivalent substitution or change made by a person skilled in the art based on the present application is within the protection scope of the present application.

Claims

1. A training method for a static scene model, characterized in that: The training method comprises: Acquire multiple frames of original point clouds, original surround images, and internal and external parameters of sensors corresponding to the original surround images of the current vehicle in different scenes; Based on the original surrounding images and the internal and external parameters of the sensor, preprocessing the multiple frames of original point clouds to obtain multiple frames of static scene point clouds; Analyzing the illumination information obtained after the illumination estimation of the original surrounding image to obtain the illumination color distribution and transparency information of the scene in the surrounding image; Rendering the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of panoramic images; The parameters of the static scene model are updated by using the loss between the multiple frames of surrounding images and the original surrounding images until the static scene model converges.

2. The training method according to claim 1, characterized in that: The preprocessing of the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of static scene point clouds includes: Removing dynamic point clouds from the multiple frames of original point clouds to obtain multiple frames of static point clouds; Coloring the multiple frames of static point clouds based on the original panoramic images and the internal and external parameters of the sensor to obtain multiple frames of colored static point clouds; The coordinate system of the multi-frame colored static point cloud is converted into a world coordinate system to obtain the multi-frame static scene point cloud.

3. The training method according to claim 2, characterized in that: The removing of dynamic point clouds from the multiple frames of original point clouds to obtain multiple frames of static point clouds includes: Compensating the multiple frames of original point clouds to obtain multiple frames of compensated point clouds; Performing three-dimensional target detection on the multi-frame compensated point cloud to obtain a dynamic point cloud in the multi-frame compensated point cloud; The dynamic point cloud is removed from the multi-frame compensated point cloud to obtain the multi-frame static point cloud.

4. The training method according to claim 2, characterized in that: The coloring of the multiple frames of static point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of colored static point clouds includes: Determine the position information of the current vehicle based on the positioning data, acceleration and angular velocity of the current vehicle; Based on the internal and external parameters of the sensor, the original surrounding image is dedistorted to obtain a dedistorted surrounding image; Based on the position information of the current vehicle, the internal and external parameters of the sensor, and the dedistorted surrounding image, the multiple frames of static point clouds are colored to obtain the multiple frames of colored static point clouds.

5. The training method according to claim 4, characterized in that: The coloring of the multiple frames of static point clouds based on the current vehicle's position information, the internal and external parameters of the sensor, and the dedistorted surrounding image to obtain the multiple frames of colored static point clouds includes: Based on the current vehicle posture information and the internal and external parameters of the sensor, the multiple frames of static point clouds are projected onto the dedistorted surrounding image to obtain colors corresponding to the multiple frames of static point clouds; The multiple frames of static point clouds are colored using the colors corresponding to the multiple frames of static point clouds to obtain the multiple frames of colored static point clouds.

6. The training method according to claim 2, characterized in that: The step of converting the coordinate system of the colored static point cloud of the multiple frames into a world coordinate system to obtain the static scene point cloud of the multiple frames includes: For the multiple frames of colored static point clouds within a period of time, obtaining the position and posture information of the current vehicle corresponding to the multiple frames of colored static point clouds; Based on the position information of the current vehicle, the coordinate system of the multi-frame colored static point cloud is converted into a world coordinate system to obtain the multi-frame static scene point cloud.

7. The training method according to any one of claims 1 to 6, characterized in that: The illumination information obtained after the illumination estimation of the original surrounding image is analyzed to obtain the illumination color distribution and transparency information of the scene in the surrounding image, including: Performing illumination estimation on the original surrounding image to obtain illumination information of the surrounding image; The illumination information of the surrounding image is analyzed to obtain illumination color distribution and transparency information of the scene in the surrounding image.

8. A training device for a static scene model, characterized in that: The training device comprises: An input module, used to obtain multiple frames of original point clouds, original surrounding images, and internal and external parameters of sensors corresponding to the original surrounding images of the current vehicle in different scenes; A data preprocessing module, used for preprocessing the multiple frames of original point clouds based on the original surrounding images and the internal and external parameters of the sensor to obtain multiple frames of static scene point clouds; A multi-layer perceptron module is used to analyze the illumination information obtained after the illumination estimation of the original surrounding image, and obtain the illumination color distribution and transparency information of the scene in the surrounding image; A three-dimensional Gaussian module, used for rendering the multiple frames of static scene point clouds based on the illumination color distribution and the transparency information to obtain multiple frames of panoramic images; The output module is used to update the parameters of the static scene model by using the loss between the multiple frames of surrounding images and the original surrounding images until the static scene model converges.

9. An electronic device, comprising a memory and a processor, characterized in that: The memory stores a computer program that can be run on the processor, and when the processor executes the program, the steps in the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Navigation method and system based on vehicle and road cloud and digital twinning technology and medium

    CN120164341A