Three-dimensional reconstruction method for vehicle driving scene, electronic equipment and medium

By setting a virtual camera on the vehicle and training the three-dimensional Gaussian sputtering model with on-board camera data, the problem of image quality degradation caused by the deviation of the rendering perspective and the training perspective in the prior art is solved, and high-quality bird's-eye view rendering is achieved.

CN120070714APending Publication Date: 2025-05-30安徽蔚来智驾科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510142782.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction method based on Gaussian sputtering leads to a degradation in the rendering view angle and the training view angle, especially when the aerial view view camera is not set, it is difficult to obtain a high-quality aerial view.

Method used

By setting a virtual camera on the vehicle, a bird's-eye view from the virtual camera perspective is generated, and the three-dimensional Gaussian sputtering model is trained in combination with the on-board camera data to realize image rendering at multiple perspectives, including bird's-eye viewing.

Benefits of technology

Without setting up aerial view view camera, the rendering quality of the aerial view in the three-dimensional reconstruction of the vehicle driving scene is improved, ensuring high-quality output of the rendered image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070714A_ABST
    Figure CN120070714A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, particularly provides a three-dimensional reconstruction method of a vehicle driving scene, electronic equipment and a medium, and aims to solve the problem of improving aerial view rendering quality in three-dimensional reconstruction of the vehicle driving scene. The method provided by the invention comprises the following steps: rendering by adopting a three-dimensional Gaussian sputtering model of a vehicle driving scene to obtain a rendered image under multiple view angles, and obtaining a reconstructed driving scene according to the rendered image; wherein a virtual camera with a bird's-eye view angle and a second camera parameter of the virtual camera at a moment t are set, and the moment t is any one of a plurality of moments; according to the vehicle-mounted camera data, generating a bird's-eye view under the visual angle of the virtual camera at the t moment, and according to the bird's-eye view and the second camera parameter, obtaining virtual camera data at the t moment; and training according to all the vehicle-mounted camera data and the virtual camera data of the virtual camera at the multiple moments to obtain a three-dimensional Gaussian sputtering model. Through the method, the rendering quality of the aerial view in three-dimensional reconstruction of the vehicle driving scene can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and specifically relates to a three-dimensional reconstruction method, an electronic device, and a medium for a vehicle driving scenario. Background Art

[0002] Vehicles are usually equipped with sensors such as cameras. The cameras can collect environmental images of the vehicle's environment. These environmental images can describe the vehicle's driving scenario, but these images cannot accurately describe the vehicle's driving scenario in three-dimensional space. In response to this, a three-dimensional reconstruction (3D Reconstruction) method can be used to obtain the driving scenario information of the vehicle in three-dimensional space, and currently the most favored three-dimensional reconstruction method is the method based on Gaussian Splatting.

[0003] Specifically, the method based on Gaussian Splatting mainly uses the original images collected by multiple cameras with different perspectives on the vehicle to reconstruct the driving scenario into a three-dimensional Gaussian splatting model. This three-dimensional Gaussian splatting model contains the Gaussian distribution of each point in the driving scenario. The Gaussian distribution of each point can be understood as a splat point with three-dimensional space attributes in this model. The three-dimensional space attributes include attributes such as the position, direction, size, and color of the splat point in three-dimensional space. Therefore, the three-dimensional Gaussian splatting model can accurately describe or reflect the three-dimensional structure of the driving scenario. After obtaining the three-dimensional Gaussian splatting model, the three-dimensional Gaussian splatting model can be rasterized and rendered to multiple different image perspectives (hereinafter described as rendering perspectives) to obtain rendered images at each rendering perspective. These rendered images combined can accurately describe the vehicle's driving scenario in three-dimensional space.

[0004] However, the current method based on Gaussian Splatting has the following disadvantages: If the rendering perspective deviates greatly from the training perspective, the quality of the rendered image at this rendering perspective will be relatively poor, and the training perspective is the image perspective of the above-mentioned original images. For example, the original images include the images collected by the front-view camera, rear-view camera, left-view camera, and right-view camera on the vehicle. Then the training perspectives include the front-view perspective, rear-view perspective, left-view perspective, and right-view perspective. However, if the rendering perspective is the bird's-eye view perspective, since the bird's-eye view perspective has a large gap with each training perspective, therefore, when rendering the three-dimensional Gaussian splatting model to the bird's-eye view perspective, a high-quality bird's-eye view image cannot be obtained, such as the road surface image in the bird's-eye view perspective. To solve this problem, a camera with a bird's-eye view perspective can be set on the vehicle again, but this will undoubtedly increase the costs of camera procurement and installation, data collection and calibration, etc.

[0005] Correspondingly, a new technical solution is needed in this field to solve the above problems. Summary of the Invention

[0006] To overcome the above - mentioned defects, the present application is proposed to solve or at least partially solve the following technical problems: in the case where there is no bird's - eye view camera on the vehicle, how to improve the rendering quality of the bird's - eye view (such as the road surface image from the bird's - eye view) in the 3D reconstruction of the vehicle driving scene.

[0007] In a first aspect, a 3D reconstruction method for a vehicle driving scene is provided, including:

[0008] Using a 3D Gaussian sputtering model of the vehicle driving scene, performing image rendering of the driving scene from multiple perspectives to obtain the rendered images from multiple perspectives, where the perspectives include a bird's - eye view;

[0009] According to the rendered images from multiple perspectives, obtaining the reconstructed driving scene;

[0010] Among them, the 3D Gaussian sputtering model is trained in the following way:

[0011] Setting an initial model of the 3D Gaussian sputtering model;

[0012] Obtaining on - vehicle camera data of each on - vehicle camera on the vehicle at a continuous plurality of moments, where the on - vehicle camera data includes on - vehicle images collected for the driving scene and first camera parameters;

[0013] Setting the virtual camera of the vehicle, where the perspective of the virtual camera is different from that of the on - vehicle camera, the perspective of the virtual camera is a bird's - eye view, and setting the second camera parameters of the virtual camera at time t, where time t is any one of the plurality of moments;

[0014] Generating a bird's - eye view at time t from the perspective of the virtual camera according to the on - vehicle camera data, and obtaining the virtual camera data at time t according to the bird's - eye view and the second camera parameters;

[0015] Training the initial model according to all the on - vehicle camera data and the virtual camera data of the virtual camera at multiple moments to obtain the final model of the 3D Gaussian sputtering model.

[0016] In a technical solution of the above 3D reconstruction method, the generating a bird's - eye view at time t from the perspective of the virtual camera according to the on - vehicle camera data includes:

[0017] Obtaining the first spatial position of the on - vehicle camera corresponding to the on - vehicle camera data in the world coordinate system according to the first camera parameters in the on - vehicle camera data, and obtaining the second spatial position of the virtual camera at time t in the world coordinate system according to the second camera parameters;

[0018] Obtain multiple vehicle-mounted camera data closest to the second spatial position from all vehicle-mounted camera data according to the first spatial position corresponding to the vehicle-mounted camera data as multiple neighboring data;

[0019] Generate a bird's-eye view at the virtual camera's perspective at time t according to all the neighboring data.

[0020] In a technical solution of the above three-dimensional reconstruction method, the generating a bird's-eye view at the virtual camera's perspective at time t according to the vehicle-mounted camera data includes:

[0021] Initialize a first bird's-eye view at the virtual camera's perspective, and set the corresponding relationship between the pixel coordinate system of the first bird's-eye view and the vehicle body coordinate system of the vehicle. The corresponding relationship includes that the u-axis and v-axis of the pixel coordinate system are respectively in the same direction as the Y-axis and X-axis of the vehicle body coordinate system, and the origin of the vehicle body coordinate system is orthogonally projected onto the midpoint of the first bird's-eye view; wherein, the imaging plane size of the virtual camera is S×S, and the image acquisition range of the virtual camera in the world coordinate system is L×L, the unit of S is pixel, and the unit of L is meter;

[0022] According to the corresponding relationship, project each pixel point in the first bird's-eye view from the pixel coordinate system to the vehicle body coordinate system, and obtain the first projection point P with a height of 0 corresponding to each pixel point in the vehicle body coordinate system v , P v =(x v , y v , 0);

[0023] Project the first projection point P v from the vehicle body coordinate system to the vehicle-mounted image of the vehicle-mounted camera data, obtain the second projection point corresponding to the first projection point P v in the vehicle-mounted image, and obtain the pixel value of the second projection point in the vehicle-mounted image;

[0024] Update the pixel value of each pixel point in the first bird's-eye view respectively according to the pixel value of the second projection point corresponding to each pixel point in the first bird's-eye view, and obtain a second bird's-eye view.

[0025] In a technical solution of the above three-dimensional reconstruction method, the respectively updating the pixel value of each pixel point in the first bird's-eye view includes:

[0026] If the semantics of the second projection point in the vehicle-mounted image is the road surface, then update the pixel value of the pixel point corresponding to the second projection point according to the pixel value of the second projection point.

[0027] In a technical solution of the above three-dimensional reconstruction method, generating a bird's-eye view at the virtual camera's perspective at time t according to the vehicle-mounted camera data further includes:

[0028] Using a neural network to predict the road surface height z corresponding to each pixel point in the second bird's-eye view in the vehicle body coordinate system v ;

[0029] According to the first projection point P corresponding to each pixel point v and the road surface height z v , obtaining the third projection point P corresponding to each pixel point in the vehicle body coordinate system v ′, P v ′=(x v , y v , z v );

[0030] Projecting the third projection point P v ′ from the vehicle body coordinate system onto the vehicle-mounted image of the vehicle-mounted camera data to obtain the fourth projection point corresponding to the third projection point P v ′ in the vehicle-mounted image, and obtaining the pixel value of the fourth projection point in the vehicle-mounted image;

[0031] According to the pixel values of the fourth projection points corresponding to the pixel points in the second bird's-eye view, respectively update the pixel values of the pixel points in the second bird's-eye view to obtain a third bird's-eye view.

[0032] In a technical solution of the above three-dimensional reconstruction method, respectively updating the pixel values of the pixel points in the second bird's-eye view includes:

[0033] If the semantics of the fourth projection point in the vehicle-mounted image is the road surface, then update the pixel value of the pixel point corresponding to the fourth projection point according to the pixel value of the fourth projection point.

[0034] In a technical solution of the above three-dimensional reconstruction method, the method further includes:

[0035] If the semantics of the fourth projection point in the vehicle-mounted image is the road surface, then obtain the prior road surface height corresponding to the fourth projection point in the vehicle body coordinate system, and update the network parameters of the neural network according to the height difference between the prior road surface height and the road surface height z v ;

[0036] In a technical solution of the above three-dimensional reconstruction method, generating a bird's-eye view at the virtual camera's perspective at time t according to all the neighboring data includes:

[0037] Obtain the camera type of the in-vehicle camera corresponding to the neighboring data, and obtain the pixel parameters corresponding to the camera type, where the pixel parameters include C A and C b ,

[0038] Obtain the first in-vehicle image in the neighboring data, and transform the pixel value I i of the i-th pixel point in the first in-vehicle image into a pixel value I ci , to obtain a second in-vehicle image, I ci = C A I i + C b , where i = 1, …, n, and n is the total number of pixel points in the in-vehicle image;

[0039] Replace the first in-vehicle image with the second in-vehicle image to obtain updated neighboring data, and generate the bird's-eye view according to the updated neighboring data;

[0040] Among them,

[0041] In the second in-vehicle images of all neighboring data, the pixel points on each second in-vehicle image that can be projected to the same spatial position in the world coordinate system have similar pixel values.

[0042] In a technical solution of the above three-dimensional reconstruction method, the method further includes:

[0043] Obtain the pixel value differences between the pixel points on each second in-vehicle image that can be projected to the same spatial position in the world coordinate system, and update the pixel parameters according to the pixel value differences.

[0044] In a technical solution of the above three-dimensional reconstruction method, the setting of the second camera parameters of the virtual camera at time t includes:

[0045] Randomly select the in-vehicle camera data of an in-vehicle camera at time t, and obtain the external parameter T w2v at time t according to the first camera parameter in the in-vehicle camera data, T w2v represents the external parameter from the world coordinate system to the vehicle body coordinate system of the vehicle;

[0046] According to the external parameter T w2v and the preset external parameter of the virtual camera, obtain the camera external parameter of the virtual camera at time t. The preset external parameter represents the external parameter from the vehicle body coordinate system to the camera coordinate system of the virtual camera;

[0047] According to the camera external parameter and the preset camera internal parameters K of the virtual camera b , and obtain the second camera parameters of the virtual camera at time t.

[0048] In a technical solution of the above three-dimensional reconstruction method, the training of the initial model includes:

[0049] alternately training the initial model using a first training mode and a second training mode;

[0050] wherein,

[0051] the first training mode is: training using the vehicle-mounted camera data;

[0052] the second training mode is: training using the virtual camera data.

[0053] In a second aspect, there is provided an electronic device, which includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the method described in any one of the technical solutions provided in the first aspect above is implemented.

[0054] In a third aspect, there is provided a computer-readable storage medium, which stores multiple program codes, and the program codes are adapted to be loaded and run by a processor to execute the method described in any one of the technical solutions provided in the first aspect above.

[0055] Solution 1. A three-dimensional reconstruction method for a vehicle driving scenario, characterized in that the method includes:

[0056] Using a three-dimensional Gaussian sputtering model of the vehicle driving scenario, rendering images of the driving scenario from multiple perspectives to obtain the rendered images from multiple perspectives, where the perspectives include a bird's-eye view;

[0057] Obtaining the reconstructed driving scenario according to the rendered images from multiple perspectives;

[0058] wherein, the three-dimensional Gaussian sputtering model is trained through the following method:

[0059] Setting an initial model of the three-dimensional Gaussian sputtering model;

[0060] Obtaining vehicle-mounted camera data of each vehicle-mounted camera on the vehicle at a plurality of consecutive times, where the vehicle-mounted camera data includes vehicle-mounted images collected from the driving scenario and first camera parameters;

[0061] Set a virtual camera of the vehicle, where the virtual camera has a different viewing angle from the on-vehicle camera. The viewing angle of the virtual camera is a bird's-eye view angle, and set second camera parameters of the virtual camera at time t, where t is any one of the multiple moments;

[0062] Generate a bird's-eye view image at time t from the perspective of the virtual camera based on the on-vehicle camera data, and obtain virtual camera data at time t according to the bird's-eye view image and the second camera parameters;

[0063] Train the initial model based on all the on-vehicle camera data and the virtual camera data of the virtual camera at multiple moments to obtain the final model of the three-dimensional Gaussian sputtering model.

[0064] Solution 2. According to the method described in Solution 1, it is characterized in that the generating a bird's-eye view image at time t from the perspective of the virtual camera based on the on-vehicle camera data includes:

[0065] Obtain a first spatial position of the on-vehicle camera corresponding to the on-vehicle camera data in the world coordinate system according to the first camera parameters in the on-vehicle camera data, and obtain a second spatial position of the virtual camera at time t in the world coordinate system according to the second camera parameters;

[0066] Obtain multiple on-vehicle camera data closest to the second spatial position from all the on-vehicle camera data as multiple neighboring data according to the first spatial position corresponding to the on-vehicle camera data;

[0067] Generate a bird's-eye view image at time t from the perspective of the virtual camera based on all the neighboring data.

[0068] Solution 3. According to the method described in Solution 1 or 2, it is characterized in that the generating a bird's-eye view image at time t from the perspective of the virtual camera based on the on-vehicle camera data includes:

[0069] Initialize a first bird's-eye view image from the perspective of the virtual camera, and set the corresponding relationship between the pixel coordinate system of the first bird's-eye view image and the vehicle body coordinate system. The corresponding relationship includes that the u-axis and v-axis of the pixel coordinate system are respectively in the same direction as the Y-axis and X-axis of the vehicle body coordinate system, and the origin of the vehicle body coordinate system is orthogonally projected onto the midpoint of the first bird's-eye view image; wherein, the imaging plane size of the virtual camera is S×S, and the image acquisition range of the virtual camera in the world coordinate system is L×L, the unit of S is pixel, and the unit of L is meter;

[0070] Project each pixel point in the first bird's-eye view image to the vehicle body coordinate system according to the corresponding relationship to obtain a first projection point P with a height of 0 in the vehicle body coordinate system corresponding to each pixel point v , Pv = (x v , y v , 0);

[0071] Project the first projection point P v from the vehicle body coordinate system onto the vehicle-mounted image of the vehicle-mounted camera data, to obtain the second projection point corresponding to the first projection point P v in the vehicle-mounted image, and obtain the pixel value of the second projection point in the vehicle-mounted image;

[0072] According to the pixel values of the second projection points corresponding to the respective pixel points in the first bird's-eye view, update the pixel values of the respective pixel points in the first bird's-eye view respectively, to obtain a second bird's-eye view.

[0073] Solution 4. According to the method described in Solution 3, it is characterized in that the step of respectively updating the pixel values of the respective pixel points in the first bird's-eye view includes:

[0074] If the semantics of the second projection point in the vehicle-mounted image is a road surface, then update the pixel value of the pixel point corresponding to the second projection point according to the pixel value of the second projection point.

[0075] Solution 5. According to the method described in Solution 3, it is characterized in that the step of generating a bird's-eye view at time t from the vehicle-mounted camera data in the virtual camera view further includes:

[0076] Using a neural network to predict the road surface height z corresponding to each pixel point in the second bird's-eye view in the vehicle body coordinate system v ;

[0077] According to the first projection point P corresponding to each pixel point v and the road surface height z v , obtain the third projection point P v ' corresponding to each pixel point in the vehicle body coordinate system, P v ' = (x v , y v , z v );

[0078] Project the third projection point P v ' from the vehicle body coordinate system onto the vehicle-mounted image of the vehicle-mounted camera data, to obtain the fourth projection point corresponding to the third projection point P v ' in the vehicle-mounted image, and obtain the pixel value of the fourth projection point in the vehicle-mounted image;

[0079] According to the pixel values of the fourth projection points corresponding to the respective pixel points in the second bird's-eye view, update the pixel values of the respective pixel points in the second bird's-eye view respectively, to obtain a third bird's-eye view.

[0080] Solution 6. The method according to Solution 5, characterized in that the step of separately updating the pixel values of the respective pixel points in the second bird's-eye view includes:

[0081] If the semantic of the fourth projection point in the vehicle-mounted image is a road surface, then according to the pixel value of the fourth projection point, update the pixel value of the pixel point corresponding to the fourth projection point.

[0082] Solution 7. The method according to Solution 5, characterized in that the method further includes:

[0083] If the semantic of the fourth projection point in the vehicle-mounted image is a road surface, then obtain the prior road surface height corresponding to the fourth projection point in the vehicle body coordinate system, and update the network parameters of the neural network according to the height difference between the prior road surface height and the road surface height z v therebetween.

[0084] Solution 8. The method according to Solution 2, characterized in that the step of generating a bird's-eye view at time t from all the neighboring data in the perspective of the virtual camera includes:

[0085] Obtain the camera type of the vehicle-mounted camera corresponding to the neighboring data, and obtain the pixel parameters corresponding to the camera type, where the pixel parameters include C A and C b ,

[0086] Obtain the first vehicle-mounted image in the neighboring data, and transform the pixel value I i of the i-th pixel point in the first vehicle-mounted image into a pixel value I ci according to the pixel parameters, to obtain a second vehicle-mounted image, I ci =C A I i +C b , where i = 1,..., n, and n is the total number of pixel points in the vehicle-mounted image;

[0087] Replace the first vehicle-mounted image with the second vehicle-mounted image to obtain updated neighboring data, and generate the bird's-eye view according to the updated neighboring data;

[0088] wherein,

[0089] In the second vehicle-mounted images of all the neighboring data, the pixel points that can be projected to the same spatial position in the world coordinate system on each second vehicle-mounted image have similar pixel values.

[0090] Solution 9. The method according to Solution 8, characterized in that the method further includes:

[0091] Obtain the pixel value differences between the pixel points on each of the second vehicle images that can be projected to the same spatial position in the world coordinate system, and update the pixel parameters according to the pixel value differences.

[0092] Solution 10. The method according to Solution 1, characterized in that setting the second camera parameters of the virtual camera at time t includes:

[0093] Randomly select the vehicle camera data of a vehicle camera at time t, and obtain the external parameter T at time t according to the first camera parameter in the vehicle camera data w2v , T w2v represents the external parameter from the world coordinate system to the vehicle body coordinate system of the vehicle;

[0094] According to the external parameter T w2v and the preset external parameter of the virtual camera obtain the camera external parameter of the virtual camera at time t The preset external parameter represents the external parameter from the vehicle body coordinate system to the camera coordinate system of the virtual camera;

[0095] According to the camera external parameter and the preset camera internal parameter K of the virtual camera b , obtain the second camera parameters of the virtual camera at time t.

[0096] Solution 11. The method according to Solution 1, characterized in that training the initial model includes:

[0097] Alternately train the initial model using a first training mode and a second training mode;

[0098] wherein,

[0099] The first training mode is: training using the vehicle camera data;

[0100] The second training mode is: training using the virtual camera data.

[0101] Solution 12. An electronic device, characterized in that it includes:

[0102] At least one processor;

[0103] And a memory communicatively connected to the at least one processor;

[0104] Wherein, a computer program is stored in the memory, and when the computer program is executed by the at least one processor, the three-dimensional reconstruction method for the vehicle driving scenario described in any one of Solutions 1 to 11 is implemented.

[0105] Solution 13. A computer-readable storage medium storing multiple program codes, characterized in that the program codes are adapted to be loaded and run by a processor to execute the three-dimensional reconstruction method of the vehicle driving scenario described in any one of Solutions 1 to 11.

[0106] One or more of the above technical solutions of the present application have at least one or more of the following beneficial effects:

[0107] In a technical solution of implementing the three-dimensional reconstruction method of the vehicle driving scenario provided by the present application, a three-dimensional Gaussian sputtering model of the vehicle driving scenario can be used to perform image rendering of the driving scenario from multiple perspectives to obtain rendered images from multiple perspectives, and the perspectives include a bird's-eye view perspective; according to the rendered images from multiple perspectives, the reconstructed driving scenario is obtained; wherein, the three-dimensional Gaussian sputtering model is trained in the following manner: setting an initial model of the three-dimensional Gaussian sputtering model; obtaining on-vehicle camera data of each on-vehicle camera on the vehicle at a continuous plurality of moments, the on-vehicle camera data including on-vehicle images collected from the driving scenario and first camera parameters; setting a virtual camera of the vehicle, the virtual camera having a different perspective from that of the on-vehicle camera, the perspective of the virtual camera being the bird's-eye view perspective, and setting second camera parameters of the virtual camera at the t-th moment, the t-th moment being any one of the plurality of moments; generating a bird's-eye view image at the virtual camera perspective at the t-th moment according to the on-vehicle camera data, and obtaining virtual camera data at the t-th moment according to the bird's-eye view image and the second camera parameters; training the initial model according to all the on-vehicle camera data and the virtual camera data of the virtual camera at multiple moments to obtain a final model of the three-dimensional Gaussian sputtering model.

[0108] In the above implementation, the on-vehicle cameras are real cameras in the physical world, and the on-vehicle images are also real images obtained by these cameras collecting images of the driving scenario. For example, the on-vehicle cameras may include a front-view camera, a rear-view camera, a left-view camera, a right-view camera, etc. on the vehicle. The virtual camera is not a real camera in the physical world. It is a simulated camera. The bird's-eye view image at the virtual camera perspective is generated using the on-vehicle camera data, and the bird's-eye view image can be understood as an image of the driving scenario collected by the virtual camera simulated using the on-vehicle images in the on-vehicle camera data.

[0109] Training the three-dimensional Gaussian sputtering model according to both the on-vehicle camera data and the virtual camera data can not only enable the three-dimensional Gaussian sputtering model to learn the three-dimensional information of the on-vehicle camera perspective, but also enable it to learn the three-dimensional information of the bird's-eye view perspective. In this way, when using the three-dimensional Gaussian sputtering model to perform image rendering of the vehicle driving scenario, even if the rendering perspective is the bird's-eye view perspective, a high-quality bird's-eye view image can still be rendered. That is, in the case where there is no bird's-eye view perspective camera on the vehicle, the rendering quality of the bird's-eye view image in the three-dimensional reconstruction of the vehicle driving scenario can still be effectively improved. Description of the Drawings

[0110] Referring to the accompanying drawings, the disclosure of the present application will become more readily understandable. It is easily understood by those skilled in the art that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present application. Among them:

[0111] Figure 1 is a schematic diagram of the main steps of a three-dimensional reconstruction method for a vehicle driving scenario according to an embodiment of the present application;

[0112] Figure 2 is an imaging schematic diagram of a virtual camera according to an embodiment of the present application Figure 1 ;

[0113] Figure 3 is an imaging schematic diagram of a virtual camera according to an embodiment of the present application Figure 2 ;

[0114] Figure 4 is a schematic diagram of the main steps of generating a bird's-eye view at time t according to an embodiment of the present application Figure 1 ;

[0115] Figure 5 is a schematic diagram of the neighboring area and neighboring data of a virtual camera according to an embodiment of the present application;

[0116] Figure 6 is a schematic diagram of the main steps of generating a bird's-eye view at time t according to an embodiment of the present application Figure 2 ;

[0117] Figure 7 is a schematic diagram of the first, second, and third bird's-eye views according to an embodiment of the present application;

[0118] Figure 8 is a schematic diagram of the overall process of a three-dimensional reconstruction method for a vehicle driving scenario according to an embodiment of the present application;

[0119] Figure 9 is a schematic diagram of the main structure of an electronic device according to an embodiment of the present application.

[0120] Reference numerals:

[0121] 11: Memory; 12: Processor. Detailed implementation manners

[0122] The following describes some implementation manners of the present application with reference to the accompanying drawings. It should be understood by those skilled in the art that these implementation manners are only used to explain the technical principle of the present application and are not intended to limit the protection scope of the present application.

[0123] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various appropriate sensors, communication ports, memory, and may also include a software part, such as program code, or may be a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other appropriate processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A computer-readable storage medium includes any appropriate medium that can store program code, such as magnetic disks, hard disks, optical disks, flash memories, read-only memories, random access memories, and so on.

[0124] In each embodiment of this application, the relevant user personal information that may be involved is strictly in accordance with the requirements of laws and regulations, following the principles of legality, legitimacy, and necessity, and for a reasonable purpose based on the business scenario, to process the personal information actively provided by the user during the use of the product / service, or generated due to the use of the product / service, and the personal information obtained with the user's authorization.

[0125] The user personal information processed by this application will vary depending on the specific product / service scenario. It is subject to the specific scenario of the user's use of the product / service and may involve the user's account information, device information, driving information, vehicle information, or other relevant information. This application will treat the user's personal information and its processing with a high degree of diligence.

[0126] This application attaches great importance to the security of user personal information and has taken security protection measures that meet industry standards and are reasonable and feasible to protect the user's information and prevent personal information from being accessed, publicly disclosed, used, modified, damaged, or lost without authorization.

[0127] The embodiments of the three-dimensional reconstruction method for the vehicle driving scenario provided by this application will be described below.

[0128] In some embodiments according to the present application, a three-dimensional Gaussian Splatting model of a vehicle driving scenario can be obtained. The three-dimensional Gaussian Splatting model contains the Gaussian distribution of each point in the driving scenario. The Gaussian distribution of each point can be understood as a splat point with three-dimensional space attributes, and the three-dimensional space attributes include attributes such as the position, direction, size, and color of the splat point in three-dimensional space. Then, using the three-dimensional Gaussian Splatting model of the vehicle driving scenario, image rendering of the driving scenario is performed from multiple perspectives to obtain rendered images from multiple perspectives. The multiple perspectives at least include the Bird's Eye View perspective. The reconstructed driving scenario is obtained based on the rendered images from multiple perspectives. Specifically, the rendered images from all perspectives can be combined to obtain the reconstructed driving scenario, that is, the reconstructed driving scenario is represented by the rendered images. In addition, when performing image rendering, the rasterization method in the three-dimensional Gaussian Splatting method is used to render the three-dimensional Gaussian Splatting model to each perspective to obtain the rendered images at each perspective. The rasterization method is a conventional method in the three-dimensional Gaussian Splatting method, and the specific principle of the rasterization method will not be elaborated in this embodiment.

[0129] Refer to the appendix Figure 1 , in the embodiments of the present application, it can be obtained through Figure 1 The following steps S101 to step S105 shown in the figure are used to train and obtain a three-dimensional Gaussian Splatting model of a vehicle driving scenario.

[0130] Step S101: Set an initial model of the three-dimensional Gaussian Splatting model.

[0131] In this embodiment, a point cloud map of the vehicle driving scenario can be obtained. Each point cloud in the point cloud map is used as each point (hereinafter described as a scene point) in the vehicle driving scenario, and an initial Gaussian distribution is set for each scene point. The initial Gaussian distributions of all scene points are combined to form an initial model of the three-dimensional Gaussian Splatting model.

[0132] Step S102: Obtain the on-vehicle camera data of each on-vehicle camera on the vehicle at a continuous plurality of moments. The on-vehicle camera data includes the on-vehicle images collected by the on-vehicle camera for the driving scenario and the first camera parameters. The perspectives of each on-vehicle camera are different, and none of them is the bird's-eye view perspective. For example, a front-view camera, a rear-view camera, a left-view camera, and a right-view camera are provided on the vehicle. The first camera parameters include the intrinsic parameters and extrinsic parameters of the on-vehicle camera. The intrinsic / extrinsic parameters of the camera are both conventional parameters in the camera technical field.

[0133] The camera intrinsic parameters can be expressed as f x 、f yThe value of can be the focal length of the vehicle-mounted camera, c x The value of can be half of the size of the imaging plane of the vehicle-mounted camera on the u-axis, c y The value of can be half of the size of the imaging plane on the v-axis. The u-axis and v-axis are the u-axis and v-axis in the pixel coordinate system respectively. The camera internal parameters are conventional parameters in the field of camera technology. For the derivation process of the above camera internal parameter values, this embodiment will not elaborate.

[0134] The camera external parameters are the external parameters from the camera coordinate system of the vehicle-mounted camera to the world coordinate system. The camera external parameters include a rotation matrix and a translation matrix. The rotation matrix can represent the attitude of the optical center of the vehicle-mounted camera in the world coordinate system, and the translation matrix can represent the spatial position of the optical center of the vehicle-mounted camera in the world coordinate system. Therefore, the camera external parameters can represent the pose of the optical center of the vehicle-mounted camera in the world coordinate system. In this embodiment, the camera internal parameters of the vehicle-mounted camera and the camera external parameters of the vehicle-mounted camera at each moment can be obtained from the parameter information of the vehicle-mounted camera, and then the first camera parameters of the vehicle-mounted camera at each moment can be obtained.

[0135] Step S103: Set a virtual camera for the vehicle. The virtual camera has a different viewing angle from the vehicle-mounted camera. The viewing angle of the virtual camera is a bird's-eye view angle, and set the second camera parameters of the virtual camera at time t, where t is any one of multiple moments.

[0136] The virtual camera is not a real camera in the physical world. It is a simulated camera. The second camera parameters also include the camera internal parameters and camera external parameters of the virtual camera. The camera external parameters are the external parameters from the camera coordinate system of the virtual camera to the world coordinate system. The camera external parameters include a rotation matrix and a translation matrix. The rotation matrix can represent the attitude of the optical center of the virtual camera in the world coordinate system, and the translation matrix can represent the spatial position of the optical center of the virtual camera in the world coordinate system. Therefore, the camera external parameters can represent the pose of the optical center of the virtual camera in the world coordinate system.

[0137] The meanings of the camera internal parameters of the virtual camera and the vehicle-mounted camera are the same. When setting the camera internal parameters of the virtual camera, they can be set according to the preset size and preset focal length of the imaging plane of the virtual camera.

[0138] In this embodiment, whether it is an in-vehicle camera or a virtual camera, as long as it is a camera of the current vehicle, the relative pose between the camera and the vehicle remains fixed, that is, the pose of the camera in the world coordinate system changes with the change of the pose of the vehicle in the world coordinate system. Therefore, when setting the extrinsic camera parameters of the virtual camera at time t, the vehicle pose of the vehicle in the world coordinate system at time t can be obtained, and the relative pose between the virtual camera and the vehicle can be set. According to this relative pose and the vehicle pose, the camera pose of the virtual camera at time t can be obtained. The rotation matrix in the extrinsic camera parameters can be obtained according to the attitude in the camera pose, and the translation matrix in the extrinsic camera parameters can be obtained according to the position in the camera pose.

[0139] Step S104: Generate a bird's-eye view under the perspective of the virtual camera according to the in-vehicle camera data, and obtain the virtual camera data at time t according to the bird's-eye view and the second camera parameters. The virtual camera data consists of the bird's-eye view at time t and the second parameters at time t.

[0140] The bird's-eye view can be understood as an image of the driving scene collected by the virtual camera simulated using the in-vehicle images in the in-vehicle camera data.

[0141] When generating the bird's-eye view, the projection points of each pixel point on the imaging plane of the virtual camera on the in-vehicle image can be obtained, and the pixel value of the projection point on the in-vehicle image can be used as the pixel value of the corresponding pixel point on the imaging plane. Finally, according to the pixel values of each pixel point on the imaging plane, the imaging image of the virtual camera is generated, and this imaging image is the bird's-eye view under the perspective of the virtual camera. If a pixel point can be projected onto multiple in-vehicle images, then all the projection points corresponding to this pixel point can be obtained, the average value of the pixel values of all the projection points can be obtained, and this average value can be used as the pixel value of this pixel point on the imaging plane.

[0142] Step S105: Train the initial model according to all the in-vehicle camera data and the virtual camera data of the virtual camera at multiple times to obtain the final model of the three-dimensional Gaussian sputtering model.

[0143] Specifically, a conventional training method for a three-dimensional Gaussian sputtering model can be adopted to train the initial model. For example, when training based on in-vehicle camera data, the initial model can be used to perform image rendering according to the first camera parameters in the in-vehicle camera data to obtain a first rendered image. The first camera parameters represent the viewing angle of the first rendered image. Then, the pixel value difference between the in-vehicle image in the in-vehicle camera data and the first rendered image is obtained, and the pixel value difference is used as the model loss. Each Gaussian distribution in the model is updated according to the model loss, that is, the three-dimensional space attributes such as the position, direction, size, and color of each sputtering point are updated. Similarly, when training based on virtual camera data, the model can also be used to perform image rendering according to the second camera parameters to obtain a second rendered image. The second camera parameters represent the viewing angle of the second rendered image. Then, the pixel value difference between the bird's-eye view and the second rendered image is obtained, and the pixel value difference is used as the model loss. Each Gaussian distribution in the model is updated according to the model loss. By training the initial model of the three-dimensional Gaussian sputtering model, the three-dimensional space attributes of each sputtering point in the model can reach the optimal state, so that the accuracy of the rendered image can be guaranteed when performing image rendering according to the final model.

[0144] In some embodiments, when training the initial model, the first training mode and the second training mode can be alternately adopted to train the initial model. The first training mode is to train using in-vehicle camera data, and the second training mode is to train using virtual camera data. In some embodiments, the initial model can also be trained multiple times using the first training mode first, and then the first and second training modes can be alternately adopted to train the initial model.

[0145] Training the three-dimensional Gaussian sputtering model based on the method described in the above steps S101 to S105 can not only enable the three-dimensional Gaussian sputtering model to learn the three-dimensional information from the perspective of the in-vehicle camera, but also enable it to learn the three-dimensional information from the bird's-eye view. In this way, when using the three-dimensional Gaussian sputtering model to perform image rendering on the vehicle driving scene, even if the rendering perspective is the bird's-eye view, a high-quality bird's-eye view can still be rendered. That is, in the case where there is no bird's-eye view camera installed on the vehicle, the rendering quality of the bird's-eye view in the three-dimensional reconstruction of the vehicle driving scene can still be effectively improved.

[0146] Next, the embodiments of the three-dimensional reconstruction method for the vehicle driving scene provided in the present application will be further described, specifically for the above steps S103 to S105.

[0147] First, an explanation of step S103 will be given.

[0148] According to the description of the foregoing step S102, the camera internal parameters can be expressed as f x 、f yThe value of can be the focal length of the camera, c x The value of can be half of the size of the imaging plane of the camera on the u-axis, c y The value of can be half of the size of the imaging plane on the v-axis. According to the description of the foregoing step S103, when setting the internal parameters of the virtual camera, it can be set according to the preset size and preset focal length of the imaging plane of the virtual camera. For example Figure 3 As shown, in some embodiments of the foregoing step S103, the size of the imaging plane of the virtual camera is S×S, and the unit of S is pixels.

[0149] The following combines the Figure 2 As shown in the pinhole imaging model, the derivation process of the preset focal length f of the virtual camera is described. The image acquisition range of the virtual camera in the world coordinate system is L×L, and the unit of L is meters. For example Figure 2 As shown, Pc represents a road surface point in the camera coordinate system, Zc represents the distance from the road surface point Pc to the optical center of the virtual camera, f represents the focal length, and x c represents the coordinate of the road surface point Pc on the X-axis in the camera coordinate system, and x c =L, and x z represents the coordinate of the pixel point corresponding to the road surface point Pc on the u-axis of the virtual camera imaging plane, and x z =S. According to the principle of similar triangles, it can be obtained that Furthermore, it can be obtained that

[0150] In some embodiments of the foregoing step S103, the pose of the vehicle at time t can be obtained by using the first camera parameters of the vehicle-mounted camera, and then the external parameters of the virtual camera can be set. Specifically, in this embodiment, the second camera parameters of the virtual camera at time t can be set through the following steps S1031 to S1033.

[0151] Step S1031: Randomly select a piece of vehicle-mounted camera data at time t, and obtain the external parameter T at time t according to the first camera parameter in the vehicle-mounted camera data w2v .

[0152] T w2v represents the external parameter from the world coordinate system to the vehicle body coordinate system of the vehicle, and T w2v can also represent the pose of the vehicle at time t. Specifically, when obtaining the external parameter T w2v , the external parameter T of the first camera parameter can be obtained c2w (that is, the external parameter from the camera coordinate system of the vehicle-mounted camera to the world coordinate system), and then the external parameter T between the vehicle body coordinate system and the camera coordinate system of the vehicle-mounted camera can be obtained v2c , and according to the external parameter T v2c and the camera external parameter T c2w T is obtainedw2v , T w2v = T c2w -1 T v2c -1 .

[0153] Step S1032: Obtain the external camera parameters of the virtual camera at time t according to the external parameter T w2v and the preset external parameters of the virtual camera The preset external parameters represent the external parameters from the vehicle body coordinate system to the camera coordinate system of the virtual camera.

[0154] Step S1033: Obtain the second camera parameters of the virtual camera at time t according to the camera external parameters and the preset camera internal parameter K of the virtual camera b .

[0155] Based on the method described in the above steps S1031 to S1033, the external camera parameters of the virtual camera can be obtained by using the first camera parameters of the vehicle-mounted camera Combined with the preset camera internal parameter K of the virtual camera b , the second camera parameters of the virtual camera at time t can be obtained.

[0156] Figure 3 In some embodiments of the above step S1032, the derivation process of the preset external parameters with respect to the imaging plane shown in the following figure is described. Among them, the size of the imaging plane of the virtual camera is S×S, the unit of S is pixel, the image acquisition range of the virtual camera in the world coordinate system is L×L, the unit of L is meter, and the corresponding relationship between the pixel coordinate system of the imaging plane and the vehicle body coordinate system of the vehicle is: the u-axis and v-axis of the pixel coordinate system are respectively in the same direction as the Y-axis and X-axis of the vehicle body coordinate system, and the origin of the vehicle body coordinate system is orthogonally projected onto the midpoint of the imaging plane.

[0157] First, set a point P with a height of 0 in the vehicle body coordinate system. The coordinates of point P in the vehicle body coordinate system are (x v , y v , 0). Then, according to the above corresponding relationship, the coordinates of point P in the pixel coordinate system can be obtained as

[0158] According to pixel coordinate normalization, the coordinates (x v , y v , 0) and the coordinates

[0159] satisfy the following projection relationship:

[0160] ​​Among them, Norm represents normalizing the coordinates of pixel points on the imaging plane to the plane where Z = 1, that is, Norm represents the depth normalization operation, and Z represents the Z-axis in the vehicle body coordinate system; and are respectively the rotation matrix and the translation matrix in the preset external parameters ; K b is the preset camera internal parameter of the virtual camera. By solving the equation of the above projection relationship,

[0161] II. Explanation of step S104.

[0162] When generating a bird's-eye view under the perspective of the virtual camera at time t based on the vehicle-mounted camera data, if the spatial position of the vehicle-mounted camera corresponding to the vehicle-mounted camera data in the world coordinate system is closer to the spatial position of the virtual camera at time t in the world coordinate system, then the image acquisition area of the vehicle-mounted image in the world coordinate system in the vehicle-mounted camera data is more likely to fall within the area covered by the bird's-eye view. Therefore, the pixel values of these vehicle-mounted images can more accurately represent the pixel values of the bird's-eye view at time t. Based on this, in order to improve the quality of the bird's-eye view, in some embodiments of the above step S104, it is possible to Figure 4 generate the bird's-eye view at time t through the following steps S201 to S203 shown in

[0163] Step S201: According to the first camera parameter in the vehicle-mounted camera data, obtain the first spatial position of the vehicle-mounted camera corresponding to the vehicle-mounted camera data in the world coordinate system, and obtain the second spatial position of the virtual camera at time t in the world coordinate system according to the second camera parameter.

[0164] Specifically, obtain the first spatial position according to the translation matrix in the first camera parameter, and obtain the second spatial position according to the translation matrix in the second camera parameter.

[0165] Step S202: According to the first spatial position corresponding to the vehicle-mounted camera data, obtain multiple vehicle-mounted camera data closest to the second spatial position from all vehicle-mounted camera data as multiple neighboring data. Specifically, the distance between the first and second spatial positions can be obtained, and the distances corresponding to all vehicle-mounted camera data are arranged in ascending order, and the vehicle-mounted camera data ranked from the 1st to the kth are selected as k neighboring data, where k > 1.

[0166] In some embodiments, a neighboring area of the virtual camera can be set. The neighboring area can be a circular area centered at the second spatial position of the virtual camera, and the neighboring area is larger than the image acquisition range of the virtual camera in the world coordinate system. First, according to the first spatial position of the vehicle-mounted camera, the vehicle-mounted cameras located within the neighboring area are acquired, and the vehicle-mounted camera data of these vehicle-mounted cameras are used as target camera data. According to the first spatial positions corresponding to the target camera data, multiple target camera data closest to the second spatial position are acquired from all the target camera data as neighboring data.

[0167] See the appendix Figure 5 , Figure 5 exemplarily shows the neighboring area of the virtual camera at time t0. As Figure 5 shown, the vehicle-mounted cameras located within the neighboring area include: vehicle-mounted cameras A, B, C, D at time t0, and vehicle-mounted cameras B, C, D at time t1. The vehicle-mounted camera data of these 7 vehicle-mounted cameras will be used as target camera data. Assuming the number of neighboring data is 5, the target camera data closest to the second spatial position of the virtual camera includes: the vehicle-mounted camera data of vehicle-mounted cameras A, B, C, D at time t0, and the vehicle-mounted camera data of vehicle-mounted cameras C, D at time t1.

[0168] Step S203: Generate a bird's-eye view at time t from the perspective of the virtual camera based on all the neighboring data. Specifically, the projection points of each pixel point on the imaging plane of the virtual camera on the vehicle-mounted images in the neighboring data can be obtained, and the pixel values of the projection points on the vehicle-mounted images are used as the pixel values of the corresponding pixel points on the imaging plane. Finally, an imaging image of the virtual camera is generated based on the pixel values of each pixel point on the imaging plane, and this imaging image is the bird's-eye view at time t. If a pixel point can be projected onto multiple vehicle-mounted images, all the projection points corresponding to this pixel point can be obtained, the average value of the pixel values of all the projection points is obtained, and this average value is used as the pixel value of this pixel point on the imaging plane.

[0169] Based on the method described in the above steps S201 to S203, the quality of the bird's-eye view can be further improved by using multiple vehicle-mounted camera data closest to the spatial position of the virtual camera.

[0170] In some embodiments of the above step S203, the neighboring data can be updated through the following steps S2031 to S2033, and a bird's-eye view at time t is generated from the perspective of the virtual camera based on the updated neighboring data.

[0171] Step S2031: Obtain the camera types of the vehicle-mounted cameras corresponding to the neighboring data, and obtain the pixel parameters corresponding to the camera types. The pixel parameters include C A and C b , represents a real number matrix with 3 rows and 3 columns, represents a real number matrix with 3 rows and 1 column, and the element values in the matrix are all real numbers. The camera type can be the model of the in-vehicle camera. There may be different models of in-vehicle cameras installed on the vehicle. The models (i.e., camera types) of all in-vehicle cameras on the vehicle can be obtained in advance, and a set of pixel parameters is set for each camera type respectively.

[0172] Step S2032: Obtain the first in-vehicle image in the neighboring data, and transform the pixel value I of the i-th pixel point in the first in-vehicle image according to the pixel parameters i into the pixel value I ci , to obtain the second in-vehicle image, I ci = C A I i + C b , where i = 1, …, n, and n is the total number of pixel points in the in-vehicle image.

[0173] Among the second in-vehicle images of all neighboring data, the pixel points on each second in-vehicle image that can be projected to the same spatial position in the world coordinate system have similar pixel values. That is to say, the role of the pixel parameters is to constrain the pixel values of the same spatial position on different in-vehicle images, so that the pixel values of the same spatial position on different in-vehicle images are as identical as possible.

[0174] When cameras of different models take pictures of the same spatial position, due to factors such as different exposure times, the pixel values of the same spatial position in different in-vehicle images may vary greatly. If the pixel values vary greatly, it may affect the accuracy of the bird's-eye view when obtaining the bird's-eye view using the in-vehicle images. Therefore, using the above pixel parameters to transform the pixel values in the in-vehicle images can reduce this difference, thereby ensuring the accuracy of the bird's-eye view.

[0175] In order to make the pixel parameters corresponding to each camera type more accurate, after obtaining the second in-vehicle images of all neighboring data, it is also possible to obtain the pixel value differences between the pixel points on each second in-vehicle image that can be projected to the same spatial position in the world coordinate system, and update the pixel parameters according to the pixel value differences, that is, with the goal of minimizing the pixel value differences, optimize the pixel parameters. When generating the bird's-eye view at time t + 1, the optimized pixel parameters will be used to transform the first in-vehicle image in the neighboring data at time t + 1 to obtain the second in-vehicle image at time t + 1, and then use the second in-vehicle image at time t + 1 to generate the bird's-eye view at time t + 1, and after obtaining the second in-vehicle image at time t + 1, the pixel parameters will still be optimized by the above method. Based on this, as time increases, the pixel parameters will also be continuously optimized.

[0176] For example, when updating pixel parameters based on pixel value differences, multiple pixel groups can be obtained according to all the second vehicle-mounted images. Each pixel group includes target pixel points from multiple different second vehicle-mounted images, and the spatial positions of the target pixel points projected onto the world coordinate system are the same, that is, all the target pixel points represent the same spatial position. Then, for each pixel group, the difference in pixel values between every two target pixel points within the pixel group is obtained, and the average value of all the differences is taken as the loss value of the pixel group. Further, the average loss value of all the pixel groups is obtained, and with the goal of minimizing the average loss value, the pixel parameters corresponding to each camera type are optimized to obtain optimized pixel parameters. For example, the network parameter update method in neural network training is used to calculate the gradient of the pixel parameters according to the average loss value, and the pixel parameters are updated according to the gradient.

[0177] Step S2033: Replace the first vehicle-mounted image with the second vehicle-mounted image to obtain updated neighboring data, and generate a bird's-eye view based on the updated neighboring data.

[0178] Based on the method described in the above steps S2031 to S2033, the pixel value correction of the vehicle-mounted image is realized, ensuring that the pixel points that can be projected onto the same spatial position in the world coordinate system on different vehicle-mounted images have similar pixel values, thereby further improving the accuracy of the bird's-eye view.

[0179] Next, step S104 will be further described.

[0180] In some embodiments of the above step S104, it can be achieved through Figure 6 the following steps S301 to S305 shown to generate a bird's-eye view from the perspective of the virtual camera at time t.

[0181] Step S301: Initialize the first bird's-eye view from the perspective of the virtual camera, and set the corresponding relationship between the pixel coordinate system of the first bird's-eye view and the vehicle body coordinate system. The pixel values of each pixel point in the first bird's-eye view are all preset values initialized, for example, the pixel values are all 0.

[0182] The above corresponding relationship is similar to the corresponding relationship shown in the foregoing embodiments Figure 3 that is, the above corresponding relationship includes: the u-axis and v-axis of the pixel coordinate system are respectively in the same direction as the Y-axis and X-axis of the vehicle body coordinate system, and the origin of the vehicle body coordinate system is orthogonally projected onto the midpoint of the first bird's-eye view. Among them, the imaging plane size of the virtual camera is S×S, and the image acquisition range of the virtual camera in the world coordinate system is L×L, where the unit of S is pixel and the unit of L is meter.

[0183] Step S302: According to the corresponding relationship, project each pixel point in the first bird's-eye view from the pixel coordinate system to the vehicle body coordinate system to obtain the first projection point P with a height of 0 corresponding to each pixel point in the vehicle body coordinate systemv , P v = (x v , y v , 0). The first projection point P v The corresponding pixel point in the first bird's-eye view has coordinates

[0184] Step S303: Project the first projection point P v from the vehicle body coordinate system onto the vehicle-mounted image of the vehicle-mounted camera data to obtain the second projection point p v corresponding to the first projection point P in the vehicle-mounted image c .

[0185] Taking the i-th vehicle-mounted camera data as an example, set the vehicle-mounted image and the first camera parameter in the i-th vehicle-mounted camera data to be m i and c i .

[0186] The pixel point P on the first bird's-eye view b is projected onto the vehicle-mounted image m i , and the process of obtaining the second projection point can be simplified as follows: T w2v has the same meaning as T in the aforementioned step S1031 w2v , T w2v represents the external parameter from the world coordinate system to the vehicle body coordinate system of the vehicle, and T i represents the external camera parameter in the first camera parameter c i , and K i represents the internal camera parameter in the first camera parameter c i , and π represents the process of projecting the pixel point P on the first bird's-eye view b to obtain the second projection point .

[0187] Step S304: Obtain the pixel value of the second projection point in the vehicle-mounted image.

[0188] Step S305: According to the pixel values of the second projection points corresponding to the pixel points in the first bird's-eye view, update the pixel values of each pixel point in the first bird's-eye view respectively to obtain the second bird's-eye view.

[0189] Specifically, the pixel value of the second projection point can be used as the pixel value of the pixel point corresponding to the second projection point in the first bird's-eye view in the first bird's-eye view.

[0190] For each pixel point in the first bird's-eye view, if the first projection point P corresponding to the current pixel point v can be projected onto multiple vehicle-mounted images, then obtain the first projection point P vThe corresponding second projection points in each vehicle-mounted image are obtained, and the average value of the pixel values of all the second projection points on the vehicle-mounted image is acquired. The pixel value of the current pixel point on the first bird's-eye view is updated according to this average value.

[0191] In some embodiments, the vehicle-mounted camera data may further include the semantic segmentation result of the vehicle-mounted image. The semantic segmentation result contains the semantics of each pixel point on the vehicle-mounted image, and the semantics may include road surface, building, plant, vehicle, pedestrian, etc. In this embodiment, conventional semantic segmentation methods can be used to perform semantic segmentation on the vehicle-mounted image to obtain the semantics of each pixel point. When updating the pixel value of a pixel point in the first bird's-eye view, the semantics of the second projection point in the vehicle-mounted image can be obtained; if the semantics is the road surface, then according to the pixel value of the second projection point, the pixel value of the pixel point corresponding to the second projection point is updated; otherwise, the pixel value of the pixel point corresponding to the second projection point is not updated.

[0192] For example, the pixel value of each pixel point in the first bird's-eye view is 0, and the second projection point The pixel coordinates on the vehicle-mounted image are The second projection point The semantics is If Is the road surface, then the second projection point The corresponding pixel point P b (The pixel point on the first bird's-eye view) The pixel value on the first bird's-eye view is updated from 0 to the pixel value of the second projection point Otherwise, the pixel value of the pixel point P b On the first bird's-eye view remains 0.

[0193] Through the above embodiments, a bird's-eye view representing the road surface in the driving scenario can be obtained.

[0194] Based on the method described in the above steps S301 to S305, the vehicle body coordinate system can be used to realize the projection of image pixels between the virtual camera and the vehicle-mounted camera, and the pixel points on the bird's-eye view can be conveniently and accurately projected onto the vehicle-mounted image to obtain the projection points on the vehicle-mounted image (i.e., the above-mentioned second projection points), so as to accurately obtain the pixel values of the pixel points on the bird's-eye view by using the pixel values of the projection points.

[0195] The following continues to describe the above step S104.

[0196] In the method described in the above steps S301 to S305, the first projection point P with a height of 0 is used v , to realize the projection of image pixels between the virtual camera and the vehicle-mounted camera. In some other embodiments of the above step S104, the road surface height corresponding to the pixel points on the bird's-eye view in the vehicle body coordinate system can be predicted, and the third projection point P with the road surface height is re-obtained v', using the third projection point P v ' to achieve the projection of image pixels between the virtual camera and the vehicle-mounted camera. Specifically, in this embodiment, through the following steps S306 to S309, using the second bird's-eye view obtained in the foregoing embodiment, a third bird's-eye view at the virtual camera's perspective at time t can be generated again.

[0197] Step S306: Use a neural network to predict the road surface height z corresponding to each pixel point in the second bird's-eye view in the vehicle body coordinate system v . The road surface height z v can be understood as the road surface height of the projection point on the road surface when the pixel point is projected from the pixel coordinate system to the vehicle body coordinate system. The structure of the neural network in this embodiment is not specifically limited. For example, in some embodiments, a U-Net network can be used.

[0198] Step S307: According to the first projection point P corresponding to each pixel point v and the road surface height z v , obtain the corresponding third projection point P of each pixel point in the vehicle body coordinate system v ', P v '=(x v , y v , z v ).

[0199] The first projection point P v is obtained in the foregoing step S302. For the third projection point P v '=(x v , y v , z v ), x v and y v are respectively equal to x v and y v in the first projection point P v .

[0200] Step S308: Project the third projection point P v ' from the vehicle body coordinate system to the vehicle-mounted image of the vehicle-mounted camera data, and obtain the corresponding fourth projection point of the third projection point P v ' in the vehicle-mounted image, and obtain the pixel value of the fourth projection point in the vehicle-mounted image. The implementation method of this step is the same as that of the foregoing steps S303 and S304, and will not be elaborated here.

[0201] Step S309: According to the pixel values of the fourth projection points corresponding to each pixel point in the second bird's-eye view, update the pixel values of each pixel point in the second bird's-eye view respectively to obtain the third bird's-eye view. The implementation method of this step is the same as that of the foregoing step S305, and will not be elaborated here.

[0202] In some embodiments, when updating the pixel value of a pixel point in the second bird's-eye view, the semantics of the fourth projection point in the vehicle-mounted image may also be obtained; if the semantics is a road surface, then according to the pixel value of the fourth projection point, the pixel value of the pixel point corresponding to the fourth projection point is updated; otherwise, the pixel value of the pixel point corresponding to the fourth projection point is not updated. Through the above embodiments, a bird's-eye view representing the road surface in the driving scene can also be obtained.

[0203] See the attached Figure 7 , Figure 7 Each of the bird's-eye views, vehicle-mounted images, and road surface height maps in

[0204] In Figure 7 , the road surface height map shows the road surface height z corresponding to each pixel point in the second bird's-eye view when projected from the pixel coordinate system to the vehicle body coordinate system v , the height z v includes a total of 5 values from h0 to h5. In each bird's-eye view, if the pattern of the grid is blank, it means the pixel value of the grid is 0; if the pattern of the grid is the same as the pattern of the grid in the vehicle-mounted image, it means the pixel value of the grid in the bird's-eye view is the same as the pixel value of this grid in the vehicle-mounted image. The pixel value of the second bird's-eye view is obtained using the first projection point P v , P v =(x v , y v , 0), and the pixel value of the third bird's-eye view is obtained using the third projection point P v ′, P v ′=(x v , y v , z v ). Since the third projection point P v ′ has height information added compared to the first projection point P v , it will cause the fourth projection point projected by the third projection point P v ′ to be different from the second projection point projected by the first projection point P v , which in turn causes the pixel values of the third bird's-eye view and the second bird's-eye view to be different.

[0205] In some embodiments, if the semantics of the fourth projection point in the vehicle-mounted image is a road surface, the prior road surface height corresponding to the fourth projection point in the vehicle body coordinate system is obtained, and according to the prior road surface height and the road surface height z vUpdate the network parameters of the neural network based on the height difference between them. The prior road surface height is the road surface height of the road surface where the fourth projection point is located. The on-vehicle camera data will include the road surface heights corresponding to each road surface point (pixel points with the semantics of the road surface) on the on-vehicle image, and these road surface heights are described as prior road surface heights. When updating the network parameters of the neural network, aim to minimize the height difference and update the network parameters of the neural network. For example, a conventional parameter update method in neural network training can be adopted, and based on the prior road surface height and the road surface height z v Calculate the network loss value, calculate the gradient of the network parameters based on the network loss value, and update the network parameters by backpropagation according to the gradient.

[0206] In some embodiments, if the semantics of the fourth projection point in the on-vehicle image is the road surface, the target point cloud corresponding to the fourth projection point in the point cloud map can also be obtained, and the height of the target point cloud is also used as the prior road surface height, and the network parameters of the neural network are updated according to this prior road surface height. Further, in some embodiments, the prior road surface height obtained from the point cloud map (hereinafter described as the first prior road surface height) and the prior road surface height in the on-vehicle camera data (hereinafter described as the second prior road surface height) can be used simultaneously to update the network parameters of the neural network. For example, according to the first prior road surface height and the road surface height z v Calculate the first network loss value, and according to the second prior road surface height and the road surface height z v Calculate the second network loss value, calculate the weighted sum of the first and second network loss values to obtain the third network loss value, calculate the gradient of the network parameters according to the third network loss value, and update the network parameters by backpropagation according to the gradient.

[0207] Based on the method described in the above steps S306 to S309, the third projection point P v ' with a predicted height can be used to realize the projection of image pixels between the virtual camera and the on-vehicle camera, which can further improve the accuracy of the projection, and thus further improve the accuracy of obtaining the bird's-eye view representing the road surface using the projection point (i.e., the fourth projection point).

[0208] Next, in combination with the attached Figure 8 , an embodiment of the three-dimensional reconstruction method for the vehicle driving scenario provided by this application will be described. In Figure 8 , the training data is the on-vehicle camera data, and the Gaussian sputtering standard reconstruction training has the same meaning as the first training mode in the foregoing embodiment, and the Gaussian sputtering road surface reconstruction training has the same meaning as the second training mode in the foregoing embodiment.

[0209] During the training of Gaussian sputtering standard reconstruction, data is extracted from the training data. The three-dimensional Gaussian sputtering model is rendered to obtain the first rendered image according to the first camera parameters in the extracted data. Then, based on the pixel value difference between the in-vehicle image and the first rendered image in the extracted data, the model loss value is obtained, and the model parameters of the three-dimensional Gaussian sputtering model are updated (i.e., the position, direction, size, color, etc. of the Gaussian distribution are updated).

[0210] During the training of Gaussian sputtering road surface reconstruction, the method described in the foregoing steps S201 to S203 can be adopted. N neighboring data are extracted from the training data, and then the color embedding module is used to process the first in-vehicle image in the neighboring data to obtain the second in-vehicle image. The first in-vehicle image is replaced with the second in-vehicle image to obtain the updated neighboring data. The implementation method of the color embedding module is the same as the method described in the foregoing steps S2031 to S2033. Then, the pixel recombination module is used to process the updated neighboring data to generate the road surface recombination image. The implementation method of the pixel recombination module is the same as the method described in the foregoing steps S301 to S309. The road surface recombination image has the same meaning as the third bird's-eye view in the foregoing embodiment. The road surface recombination image is a bird's-eye view representing the road surface in the driving scene.

[0211] After obtaining the road surface recombination image, the three-dimensional Gaussian sputtering model is rendered to obtain the second rendered image according to the second camera parameters in the neighboring data. Then, based on the pixel value difference between the road surface recombination image and the second rendered image, the model loss value is obtained, and the model parameters of the three-dimensional Gaussian sputtering model are updated according to the model loss value.

[0212] It should be noted that although the above embodiments describe the various steps in a specific order, those skilled in the art can understand that in order to achieve the effects of the present application, it is not necessary for different steps to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders. These adjusted solutions are equivalent technical solutions to the technical solutions described in the present application, and therefore will also fall within the protection scope of the present application.

[0213] Those skilled in the art can understand that all or part of the processes in the methods of the above-mentioned embodiments of the present application can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium, etc., that can carry the computer program code.

[0214] Another aspect of the present application also provides a computer-readable storage medium.

[0215] In an embodiment of a computer-readable storage medium according to the present application, the computer-readable storage medium can be configured to store a program for executing the three-dimensional reconstruction method of the vehicle driving scenario in the above-mentioned method embodiment. This program can be loaded and run by a processor to implement the three-dimensional reconstruction method of the vehicle driving scenario. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The computer-readable storage medium can be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present application is a non-transitory computer-readable storage medium.

[0216] Another aspect of the present application also provides an electronic device.

[0217] In an embodiment of an electronic device according to the present application, the electronic device can include at least one processor; and a memory communicatively connected to at least one processor; wherein, a computer program is stored in the memory, and when the computer program is executed by at least one processor, the method described in any of the above embodiments is realized. Refer to the attached Figure 9 , Figure 9 In which, it is exemplarily shown that the memory 11 and the processor 12 are communicatively connected through a bus. The electronic device described in the present application can be, but is not limited to, a desktop type, a laptop type, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, etc. The embodiments of the present application do not make any limitations in this regard.

[0218] So far, the technical solution of the present application has been described in conjunction with an embodiment shown in the accompanying drawings. However, those skilled in the art can easily understand that the protection scope of the present application is obviously not limited to these specific embodiments. Without departing from the principle of the present application, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present application.

Claims

1. A three-dimensional reconstruction method for a vehicle driving scene, characterized in that: The method comprises: Using a three-dimensional Gaussian sputtering model of a vehicle driving scene, performing image rendering under multiple perspectives on the driving scene to obtain rendered images under the multiple perspectives, wherein the perspectives include a bird's-eye view; Acquire a reconstructed driving scene according to the rendered images under the multi-viewing angles; The three-dimensional Gaussian sputtering model is trained in the following way: Setting an initial model of the three-dimensional Gaussian sputtering model; Acquire vehicle-mounted camera data of each vehicle-mounted camera on the vehicle at a plurality of consecutive moments, wherein the vehicle-mounted camera data includes a vehicle-mounted image collected for the driving scene and a first camera parameter; Setting a virtual camera of the vehicle, wherein the perspective of the virtual camera is different from that of the vehicle-mounted camera, and the perspective of the virtual camera is a bird's-eye view, and setting a second camera parameter of the virtual camera at time t, wherein time t is any one of the multiple times; Generate a bird's-eye view at time t from the perspective of the virtual camera according to the vehicle-mounted camera data, and obtain virtual camera data at time t according to the bird's-eye view and the second camera parameters; The initial model is trained according to all vehicle-mounted camera data and virtual camera data of the virtual camera at multiple moments to obtain a final model of the three-dimensional Gaussian sputtering model.

2. The method according to claim 1, characterized in that The step of generating a bird's-eye view from the perspective of the virtual camera at time t according to the vehicle-mounted camera data includes: According to the first camera parameter in the vehicle camera data, obtain a first spatial position of the vehicle camera corresponding to the vehicle camera data in the world coordinate system, and according to the second camera parameter, obtain a second spatial position of the virtual camera in the world coordinate system at time t; According to the first spatial position corresponding to the vehicle camera data, acquiring a plurality of vehicle camera data closest to the second spatial position from all vehicle camera data as a plurality of neighboring data; Based on all neighbor data, a bird's-eye view at time t from the perspective of the virtual camera is generated.

3. The method according to claim 1 or 2, characterized in that: The step of generating a bird's-eye view from the perspective of the virtual camera at time t according to the vehicle-mounted camera data includes: Initialize a first bird's-eye view from the perspective of the virtual camera, and set a correspondence between a pixel coordinate system of the first bird's-eye view and a body coordinate system of the vehicle, wherein the correspondence includes that the u-axis and the v-axis of the pixel coordinate system are respectively in the same direction as the Y-axis and the X-axis of the body coordinate system, and the origin of the body coordinate system is orthographically projected at the midpoint of the first bird's-eye view; wherein the imaging plane size of the virtual camera is S×S, and the image acquisition range of the virtual camera in the world coordinate system is L×L, where the unit of S is pixel and the unit of L is meter; According to the corresponding relationship, each pixel point in the first bird's-eye view is projected from the pixel coordinate system to the vehicle body coordinate system to obtain a first projection point P with a height of 0 corresponding to each pixel point in the vehicle body coordinate system. v , P v =(x v ,y v ,0); The first projection point P v The first projection point P is obtained by projecting the vehicle body coordinate system onto the vehicle image of the vehicle camera data. v a second projection point corresponding to the vehicle-borne image, and obtaining a pixel value of the second projection point in the vehicle-borne image; According to the pixel values ​​of the second projection points corresponding to the pixel points in the first bird's-eye view, the pixel values ​​of the pixel points in the first bird's-eye view are updated respectively to obtain the second bird's-eye view.

4. The method according to claim 3, characterized in that The respectively updating the pixel value of each pixel point in the first bird's-eye view includes: If the semantics of the second projection point in the vehicle-borne image is a road surface, the pixel value of the pixel point corresponding to the second projection point is updated according to the pixel value of the second projection point.

5. The method according to claim 3, characterized in that: The step of generating a bird's-eye view from the perspective of the virtual camera at time t according to the vehicle-mounted camera data further includes: A neural network is used to predict the road surface height z corresponding to each pixel point in the second bird's-eye view in the vehicle body coordinate system. v ; According to the first projection point P corresponding to each pixel point v and the road surface height z v , obtain the third projection point P corresponding to each pixel point in the vehicle body coordinate system v ′,P v ′=(x v ,y v ,z v ); The third projection point P v ' Project the vehicle body coordinate system onto the vehicle image of the vehicle camera data to obtain the third projection point P v 'a fourth projection point corresponding to the vehicle-borne image, and obtains a pixel value of the fourth projection point in the vehicle-borne image; According to the pixel value of the fourth projection point corresponding to each pixel point in the second bird's-eye view, the pixel value of each pixel point in the second bird's-eye view is updated respectively to obtain a third bird's-eye view.

6. The method according to claim 5, characterized in that The respectively updating the pixel value of each pixel point in the second bird's-eye view includes: If the semantics of the fourth projection point in the vehicle-borne image is a road surface, the pixel value of the pixel point corresponding to the fourth projection point is updated according to the pixel value of the fourth projection point.

7. The method according to claim 5, characterized in that The method further comprises: If the semantics of the fourth projection point in the vehicle-borne image is a road surface, the a priori road surface height corresponding to the fourth projection point in the vehicle body coordinate system is obtained, and the a priori road surface height and the road surface height z are compared. v The height difference between them is used to update the network parameters of the neural network.

8. The method according to claim 2, characterized in that: The step of generating a bird's-eye view at time t from the perspective of the virtual camera according to all neighboring data includes: Obtain the camera type of the vehicle-mounted camera corresponding to the neighboring data, and obtain the pixel parameters corresponding to the camera type, wherein the pixel parameters include C A and C b , Obtain a first vehicle image in the neighboring data, and transform the pixel value I of the i-th pixel point in the first vehicle image according to the pixel parameter i Transformed into pixel value I ci , get the second vehicle image, I ci =C A I i +C b , i=1,…,n, n is the total number of pixels in the vehicle-mounted image; Replacing the first vehicle-mounted image with the second vehicle-mounted image to obtain updated neighbor data, and generating the bird's-eye view according to the updated neighbor data; in, In the second vehicle-mounted images of all neighboring data, pixel points on each second vehicle-mounted image that can be projected to the same spatial position in the world coordinate system have similar pixel values.

9. The method according to claim 8, characterized in that The method further comprises: Obtain pixel value differences between pixel points on the second vehicle-mounted images that can be projected to the same spatial position in the world coordinate system, and update the pixel parameters according to the pixel value differences.

10. The method according to claim 1, characterized in that The step of setting the second camera parameter of the virtual camera at time t includes: Randomly select the camera data of a vehicle camera at time t, and obtain the external parameter T at time t according to the first camera parameter in the vehicle camera data. w2v , T w2v An external parameter representing the world coordinate system to the body coordinate system of the vehicle; According to the external parameter T w2v The preset external parameters of the virtual camera Get the camera external parameters of the virtual camera at time t The preset external parameter An external parameter representing the distance from the vehicle body coordinate system to the camera coordinate system of the virtual camera; According to the camera external parameters and the preset camera intrinsic parameter K of the virtual camera b , obtain the second camera parameters of the virtual camera at time t.

Citation Information

Cited By

  • Parking visualization method, device and equipment based on Gaussian occupation field and medium

    CN122244330A

  • Parking visualization method and device based on gaussian occupancy field, equipment and medium

    CN122244330B