Single-map vehicle reconstruction method and system based on 3DGS

Through a single-image vehicle reconstruction method based on 3DGS, the three-dimensional model of the vehicle is reconstructed using a single RGB image, and the dependence on multiple input images or a large number of computing resources in the prior art is solved, and high-quality reconstruction effect is achieved under limited resources.

CN120147537AActive Publication Date: 2025-06-13UNIV OF SCI & TECH BEIJING
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510243879.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Existing three-dimensional reconstruction technologies usually require multiple input images or a large number of computing resources, making it difficult to adapt to application scenarios with limited resources or sparse image information.

Method used

A single-graphic vehicle reconstruction method based on 3DGS is adopted to pre-process the autonomous driving data set, a vehicle reconstruction model is constructed, and a single RGB image is used to realize the three-dimensional reconstruction of the vehicle.

Benefits of technology

It realizes high-quality three-dimensional vehicle model reconstruction with limited resources or scarce image information, reducing the cost of sensor and data acquisition, and is suitable for dynamic driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147537A_ABST
    Figure CN120147537A_ABST
Patent Text Reader

Abstract

The invention discloses a single-map vehicle reconstruction method and system based on 3DGS, and belongs to the technical field of machine vision, and the method comprises the steps: obtaining an automatic driving data set, and carrying out the preprocessing of the data in the automatic driving data set; wherein the data in the automatic driving data set is multi-view RGB image data obtained by a plurality of RGB cameras installed on the automatic driving vehicle to shoot the surrounding environment at different view angles; a vehicle reconstruction model based on 3DGS is constructed, the input of the vehicle reconstruction model is a single RGB image, and the output of the vehicle reconstruction model is a vehicle image obtained after reconstruction; training a vehicle reconstruction model by using the preprocessed automatic driving data set; and utilizing the trained vehicle reconstruction model to realize vehicle reconstruction according to a single RGB image. According to the technical scheme provided by the invention, the cost of sensors and data acquisition is remarkably reduced, and the method is particularly suitable for dynamic driving scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and particularly to a single-image vehicle reconstruction method and system based on 3DGS. Background Art

[0002] Autonomous driving scenario simulation is one of the key links in realizing autonomous driving technology, aiming to reconstruct the three-dimensional scene around the vehicle through sensors, provide a comprehensive understanding of the surrounding environment for autonomous driving vehicles, and thus support their safe and effective decision-making and navigation.

[0003] With the rapid development of computer vision and artificial intelligence technologies, three-dimensional reconstruction technology has become an important part of many application fields. Especially in the fields of autonomous driving, intelligent transportation, virtual reality, and augmented reality, accurate three-dimensional vehicle model reconstruction technology is the basis for realizing complex scene understanding and perception. Currently, common three-dimensional reconstruction methods are mainly based on technologies such as multi-view stereo, structured light, and laser scanning. These technologies can reconstruct the three-dimensional structure in the scene through the coordinated work of multiple angles and multiple sensors.

[0004] Traditional three-dimensional reconstruction methods usually rely on multi-view data or multi-modal priors (such as lidar point clouds), but these methods have limitations, such as high sensor costs, poor real-time performance, and difficulty in multi-view alignment. In addition, dynamic objects (such as other vehicles and pedestrians) in complex urban environments also increase the difficulty of scene reconstruction. Depth information is limited in monocular reconstruction, and monocular depth information cannot reflect the true depth of the invisible part of the vehicle. For example, given a view of the rear of the vehicle, the depth of the front of the vehicle is occluded, and during reconstruction, the network still needs to imagine to obtain this part of the point cloud.

[0005] Currently, the existing sparse-view three-dimensional reconstruction work based on deep learning can be divided into three technical routes. One is to first perform artificial or inference modeling on the target object in a single image to directly obtain its three-dimensional shape, and then use a texture generation model to color and process the lighting. The entire process requires staged training, resulting in an extended synchronization of training and testing times. This method requires a large amount and various types of input data, and the output three-dimensional shapes are very similar, lacking uniqueness.

[0006] Second, directly use the given prior template for body posture learning, fine-tune the template geometry, and simultaneously perform sampling to learn color features to achieve grid-based 3D model reconstruction. Specifically, use a differentiable renderer for rendering and then perform post-processing to generate the projection of a certain 3D shape onto a 2D plane; generate a segmentation map, key point coordinates, RGB pixels, etc. based on the projection, and then optimize the segmentation and color. The disadvantage of this approach is that there is a prior assumption that the training object is a symmetric object. The training needs to initialize the template based on the shape prior of the synthetic object. For objects with complex and asymmetric structures, the learning difficulty is significantly increased; there is still some work using generative adversarial networks to improve the geometric texture quality. However, due to the use of self-supervised learning, there is no clear definition of the ground truth, which is prone to converging to a sub-optimal state or failing to converge, and is limited by the volume and complexity of the object, resulting in poor learning effects for complex objects.

[0007] Third, use Neural Radiance Fields (NeRF) and volume rendering for 3D structure learning, which can directly learn the 3D shape and color of an object. Currently, using this method for 3D reconstruction is the mainstream in the academic community, and most of the relevant papers in recent years are centered around this technology. Compared with other methods, the NeRF-based method can generate virtual images with higher clarity. However, since this type of algorithm is based on an implicit 3D representation method, it will reduce the explicit geometric quality and multi-view 3D consistency, and requires a long training time and large computational overhead.

[0008] Recovering 3D information from a single 2D image requires overcoming various limitations such as perspective, lighting, and occlusion. Existing single-image reconstruction work mainly relies on deep learning methods, using CNN or Transformer to extract geometric information such as depth and normal in the image, thereby generating corresponding 3D representations, such as point clouds, meshes, voxels, or implicit functions. 3D reconstruction for autonomous driving scenarios in natural environments can reduce the difficulty of data acquisition, does not require a complex calibration process, and does not rely on multi-view or multi-modal priors. Compared with multi-view or multi-sensor data fusion, single-image reconstruction reduces the dependence on multi-source data and helps improve the real-time performance of the system.

[0009] However, using only a single 2D RGB image as input means that some parts of the target vehicle must be invisible when extracting features. Without using timestamps and obj_id, it is impossible to track the target vehicle and match multiple perspectives. The sparse shooting perspectives of the autonomous driving dataset also lead to uneven distribution of vehicle poses. For the ego vehicle, the visible perspective distribution of the vehicle is uneven, mainly at the rear and front of the vehicle. If there are more front cameras than rear cameras, it will result in more rear perspectives than front perspectives. At the same time, the autonomous driving dataset is limited by the shooting method of the ego vehicle, and the observation perspective of the driving scene is very limited because the ego vehicle can only drive along the motorway and needs to comply with traffic rules and cannot turn randomly. Therefore, although the autonomous driving dataset often has RGB cameras at multiple angles, such as the ego vehicles in pandaset and KITTI360 are equipped with 6 cameras, this shooting angle centered on the ego vehicle is actually 6 different single perspectives and is not helpful for constructing the target object under multiple perspectives. When the ego vehicle is driving straight along the road, for a partially occluded object, such as a car parked by the roadside, none of the ego vehicle's cameras can see the side of the car away from the road. Moreover, the autonomous driving dataset contains a large number of dynamic foregrounds. The relative speed of some foregrounds with the ego vehicle is too fast, resulting in these foreground objects only appearing in very few video frames, which brings great difficulties to the reconstruction of this part of the foreground.

[0010] In summary, the existing 3D reconstruction technologies usually require multiple input images or a large amount of computing resources and are difficult to adapt to application scenarios with limited resources or scarce image information. Summary of the Invention

[0011] The present invention provides a single-image vehicle reconstruction method and system based on 3DGS to solve the technical problem that the existing 3D reconstruction technologies usually require multiple input images or a large amount of computing resources and are difficult to adapt to application scenarios with limited resources or scarce image information.

[0012] To solve the above technical problems, the present invention provides the following technical solutions:

[0013] On the one hand, the present invention provides a single-image vehicle reconstruction method based on 3DGS, including:

[0014] Obtain an autonomous driving dataset and preprocess the data in the autonomous driving dataset; wherein, the data in the autonomous driving dataset is multi-perspective RGB image data obtained by cameras installed on an autonomous driving vehicle shooting the surrounding environment from different perspectives;

[0015] Constructing a vehicle reconstruction model based on 3DGS; wherein the input of the vehicle reconstruction model is a single RGB image, and the output is a reconstructed vehicle image;

[0016] Use the preprocessed autonomous driving dataset to train the vehicle reconstruction model;

[0017] Using the trained vehicle reconstruction model, vehicle reconstruction is achieved based on a single RGB image.

[0018] Furthermore, preprocessing includes: label generation, data enhancement, data storage, and data quality control.

[0019] Furthermore, the tag generation includes:

[0020] Use the pre-trained segmentation model and target detection model to generate vehicle mask data from the original RGB image and add category labels, location information, and posture angle data to the vehicle.

[0021] Furthermore, the data enhancement includes any one or more combinations of random rotation, cropping, scaling, and adding noise.

[0022] Furthermore, the data storage includes:

[0023] According to the purpose and scenario of the data, the data is divided into training set, validation set and test set in preset proportions, and the data is stored in structured folders; when storing, RGB images are stored in PNG format, the generated point cloud data is stored in PLY format, and the label data is stored in json format; the data in the training set contains a variety of vehicle types, angles and environments, which are used for model training; the data in the validation set contains some scenes and angles that have not been seen by the model, which are used to evaluate the model performance; the test set is used for the final performance test of the model.

[0024] Furthermore, the data quality control includes:

[0025] Perform quality checks on the raw data to remove data that is ambiguous, incomplete, does not meet the required size (too small), or has incorrect labels.

[0026] Furthermore, the process of processing the input RGB image by the vehicle reconstruction model includes:

[0027] Generate a mask for the vehicle using a pre-trained mask generator; and complete the mask and texture of the vehicle by fine-tuning the lama model on the pre-processed autonomous driving dataset;

[0028] Using a preset pose feature extractor, learn the pose features of the vehicle, fuse the features with the completed vehicle texture and mask, and then send them to a pre-trained encoder to generate latent encoded features; among them, the input RGB image and mask are loaded into the GPU memory and stored in tensor format; the latent encoded features and the fused feature tensors are stored in a dedicated buffer in the GPU memory in the form of continuous floating-point arrays;

[0029] Using a sphere template and the Fibonacci sampling method, generate a uniformly distributed point cloud on the surface of the sphere;

[0030] Convert the three-dimensional point cloud coordinates in the sphere template from Cartesian coordinates to polar coordinates, convert the updated polar coordinates back to Cartesian coordinates, and return the deformed 3DGS point cloud attributes to obtain the vehicle point cloud;

[0031] Use the vehicle point cloud for rendering and optimization to obtain the RGB image of the vehicle.

[0032] Further, when converting the three-dimensional point cloud coordinates in the sphere template from Cartesian coordinates to polar coordinates, the extracted features are used by a bi-plane decoder to predict the deformation parameters of each point, and the deformation parameters are used to calculate the collapse factor and the updated polar coordinates.

[0033] Further, the using the vehicle point cloud for rendering and optimization includes:

[0034] Use the object-camera rotation matrix, camera position, image height, width, and internal parameters provided by the dataset to create an original camera object, and use this original camera object to render the picture of the original perspective;

[0035] Create a camera object symmetric to the original camera object, and create camera pairs at certain intervals, and render the normalized vehicle onto a 2D plane to generate a rotation map of the vehicle;

[0036] By calculating the loss between the rendered image and the original image, perform backpropagation to optimize the model parameters, so that the reconstruction result is closer to the input image, thereby gradually approaching the three-dimensional structure of the real vehicle.

[0037] On the other hand, the present invention also provides a single-image vehicle reconstruction system based on 3DGS, including:

[0038] A data processing module for obtaining an autonomous driving dataset and preprocessing the data in the autonomous driving dataset; among them, the data in the autonomous driving dataset is multi-view RGB image data obtained by cameras installed on an autonomous driving vehicle shooting the surrounding environment from different perspectives;

[0039] A model construction module for constructing a vehicle reconstruction model based on 3DGS; wherein, the input of the vehicle reconstruction model is a single RGB image, and the output is the reconstructed vehicle image;

[0040] A model training module for training the vehicle reconstruction model using the preprocessed autonomous driving data set;

[0041] An application module for realizing vehicle reconstruction according to a single RGB image by using the trained vehicle reconstruction model.

[0042] On the other hand, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.

[0043] On another hand, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.

[0044] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0045] The present invention proposes a method for reconstructing a three-dimensional model of a vehicle based on three-dimensional Gaussian splash (3DGS) that only relies on single two-dimensional image data. It only uses a single 2D RGB picture and does not rely on camera parameter embedding networks, and can learn the 3D point cloud of the vehicle with a normalized distribution. By using a simple network structure, by aligning the mask shape of the vehicle with the vehicle pose and utilizing the spherical template and the geometric symmetry of the vehicle, high-quality point cloud reconstruction of the vehicle is achieved. Different from existing methods such as AutoRF, the present invention explicitly learns the three-dimensional point cloud of the vehicle, which combines the data representation form of 3DGS well, can achieve real-time rendering, and solves the discontinuity problem of traditional point-based rendering. It significantly reduces the cost of sensors and data acquisition, and is especially suitable for dynamic driving scenarios. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0047] Figure 1 It is a schematic flowchart of a single-image vehicle reconstruction method based on 3DGS provided by an embodiment of the present invention;

[0048] Figure 2 It is an architecture diagram of a vehicle reconstruction model provided by an embodiment of the present invention;

[0049] Figure 3 is a schematic diagram of the generated effect provided by an embodiment of the present invention;

[0050] Figure 4 is a system block diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0052] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Specifically, the use of the word "exemplarily" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0053] The first embodiment

[0054] This embodiment provides a single-image vehicle reconstruction method based on 3DGS. This method can be implemented by an electronic device, and the execution process of this method is as Figure 1 shown and includes the following steps:

[0055] S1. Obtain an autonomous driving dataset and preprocess the data in the autonomous driving dataset; wherein, the data in the autonomous driving dataset is multi-view RGB image data obtained by cameras installed on an autonomous driving vehicle taking pictures of the surrounding environment from different perspectives;

[0056] Among them, it should be noted that the original data required by the model of the present invention mainly includes RGB images and related label data. Data preprocessing includes label generation, data augmentation, data storage, data quality control, etc.

[0057] RGB images: Use multi-view on-vehicle camera data in autonomous driving datasets (such as Waymo, KITTI, etc.). These cameras are installed on autonomous driving vehicles and take pictures of the surrounding environment from different perspectives. The images cover various scenarios such as urban roads, highways, and rural roads.

[0058] Label generation: Use pre-trained segmentation models (such as Segment-Anything) and object detection models (such as GroundingDINO) to generate mask data of vehicles from the original RGB images, and add category labels, location information (such as bounding boxes), and pose angles and other data to the vehicles.

[0059] Data Augmentation: During the preprocessing process, data augmentation operations such as random rotation, cropping, scaling, adding noise, etc. are performed to improve the robustness and generalization ability of the model.

[0060] Data Storage:

[0061] Storage Format: The preprocessed RGB and images are stored in PNG format, and the generated point cloud data is stored in PLY format. All labels are stored in json format.

[0062] Hierarchical Storage: According to the data usage and scenarios, the data is divided into training set, validation set, and test set, and stored in structured folders according to the ratio (8:1:1) for easy subsequent access and use.

[0063] Training Set: It mainly contains diverse vehicle types, angles, and environments.

[0064] Validation Set: It contains some unseen scenarios and angles for evaluating the model performance.

[0065] Test Set: Independent of the training and validation sets, it is specifically used for the final performance test.

[0066] In addition, this embodiment also controls the data quality, automatically checks the quality of the collected raw data, and eliminates the data with blurriness, incompleteness, or label errors.

[0067] S2. Construct a vehicle reconstruction model based on 3DGS; wherein, the input of the vehicle reconstruction model is a single RGB image, and the output is the reconstructed vehicle image.

[0068] Specifically, the architecture of the model of the present invention is as Figure 2 shown, and its data processing process is as follows:

[0069] 1. Mask Generation and Completion

[0070] Before training, first, a mask generator centered on the pre-trained language segment-anything model generates a mask for the vehicle by taking "car" as the prompt word. Language Segment-Anything is based on the segment-anything segmentation model and the Grounding DINO detection model, and can generate masks for specific objects in images. The RGB image of the vehicle is cropped, deformed, and normalized to a resolution of 256×256 so that the network has a unified input. Since the Waymo dataset lacks occlusion rate and truncation rate labels for objects, it is necessary to repair the data with low quality. By fine-tuning the lama (Large Mask Inpainting) model on the self-made vehicle dataset, the mask and texture of the vehicle are complemented.

[0071] 2. Feature Extraction and Fusion

[0072] Use the EgoNet pose feature extractor to learn pose features such as the pitch angle and yaw angle of the vehicle, perform feature fusion with the complemented vehicle texture and mask features, and then send them into the pre-trained FeatUp ResNet50 encoder to generate latent codes. The input RGB image and mask are loaded into the GPU memory and stored in tensor format, using the CUDA version of torch.Tensor. On the GPU side, the tensors are allocated to a shared memory area for direct access by subsequent deep learning networks. The latent code features and the fused feature tensors are stored in a dedicated buffer in the GPU memory in continuous floating-point array format (float32) to reduce memory access overhead.

[0073] 3. Gaussian Point Cloud Deformation

[0074] A spherical template is proposed to generate an initial point cloud. The initial spherical point cloud uses the Fibonacci sampling method to generate a uniformly distributed point cloud on the sphere surface.

[0075]

[0076] where \(i = 1, 2, \ldots, N\) i represents the index of the point, \(\varphi\) is the golden ratio, and \(N\) is the total number of points.

[0077] When deforming the template point cloud, the three-dimensional point cloud coordinates are converted from Cartesian coordinates to polar coordinates.

[0078]

[0079] where the polar coordinate parameter \(r\) i , \(\theta\) i , \(\varphi\) irepresent the radial distance, polar angle, and azimuth angle respectively, and x i , y i , z i are the x, y, and z coordinate values of the Cartesian coordinate system respectively.

[0080] The extracted features are used by a bi - plane decoder to predict the deformation parameters for each point, and the deformation parameters are used to calculate the collapse factor and the updated polar coordinates.

[0081]

[0082] Among them, δr i , δθ i , are the deformation parameters, F feat is the extracted feature, D deform is the decoder, and p i represents a three - dimensional coordinate point.

[0083] The features are decoded into two planes, θ - Φ and Φ - θ, to increase the number of parameters. After deformation, the updated polar coordinates are converted back to Cartesian coordinates, and the deformed 3DGS point cloud attributes (coordinates, scaling factor, rotation factor, opacity, and color) are returned.

[0084] 4. Rendering and Optimization

[0085] First, use the object - camera rotation matrix, camera position, image height, width, and internal parameters provided by the dataset to create an original camera object C 0 , and use this camera to render the image of the original perspective. Then, create symmetric camera objects, and create camera pairs at regular intervals, and render the normalized vehicle onto a 2D plane to generate the rotation map of the vehicle. By calculating the loss with the original image, backpropagation is used to adjust the template deformation amount and Gaussian point cloud parameters, so that the generated 2D image is closer to the input image, thereby gradually approaching the three - dimensional structure of the real vehicle.

[0086] S3. Use the pre - processed autonomous driving dataset to train the vehicle reconstruction model;

[0087] S4. Use the trained vehicle reconstruction model to realize vehicle reconstruction according to a single RGB image.

[0088] The vehicle image generated by the model of the present invention is as Figure 3 shown.

[0089] In summary, this embodiment proposes a method for reconstructing a 3D model of a vehicle based on 3D Gaussian representation that relies only on single 2D image data. Using only a single 2D RGB picture and without embedding camera parameters into the network, it can learn the 3D point cloud of the vehicle with a normalized distribution. By using a simple network structure, aligning the mask shape of the vehicle with the vehicle pose, and leveraging the spherical template and the geometric symmetry of the vehicle, high-quality point cloud reconstruction of the vehicle is achieved. Different from existing methods such as AutoRF, the present invention explicitly learns the 3D point cloud of the vehicle, which combines the data representation form of 3DGS well, enables real-time rendering, and solves the discontinuity problem of traditional point-based rendering. It significantly reduces the costs of sensors and data acquisition and is especially suitable for dynamic driving scenarios.

[0090] Second Embodiment

[0091] This embodiment provides a single-image vehicle reconstruction system based on 3DGS, which includes the following modules:

[0092] A data processing module, configured to obtain an autonomous driving dataset and preprocess the data in the autonomous driving dataset; wherein, the data in the autonomous driving dataset is multi-view RGB image data obtained by cameras installed on an autonomous driving vehicle capturing the surrounding environment from different perspectives;

[0093] A model construction module, configured to construct a vehicle reconstruction model based on 3DGS; wherein, the input of the vehicle reconstruction model is a single RGB image, and the output is the reconstructed vehicle image;

[0094] A model training module, configured to train the vehicle reconstruction model using the preprocessed autonomous driving dataset;

[0095] An application module, configured to use the trained vehicle reconstruction model to implement vehicle reconstruction based on a single RGB image.

[0096] It should be noted that the single-image vehicle reconstruction system based on 3DGS in this embodiment corresponds to the single-image vehicle reconstruction method based on 3DGS in the above first embodiment; wherein, the functions implemented by each functional module in the single-image vehicle reconstruction system based on 3DGS in this embodiment correspond one by one to each process step in the single-image vehicle reconstruction method based on 3DGS in the above first embodiment; therefore, it will not be elaborated herein.

[0097] Third Embodiment

[0098] This embodiment provides an electronic device, such as Figure 4As shown, the electronic device includes: a central processing unit, a graphics processing unit, and a memory; wherein, the processor and the memory can be connected through a communication bus; at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the first embodiment above. In addition, the electronic device may further include a transceiver, and the processor and the transceiver can be connected through a communication bus, and the transceiver is used to communicate with other devices.

[0099] Next, specific introductions will be made to the various components of the electronic device in conjunction with Figure 4 :

[0100] Among them, the processor is the control center of the electronic device. The electronic device may include multiple central processing units (CPUs) and multiple graphics processing units (GPUs). Each of these central processing units can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.

[0101] In a specific implementation, as an embodiment, the processor may include one or more CPUs, such as Figure 4 CPU0 and CPU1 shown in

[0102] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation manner can refer to the above method embodiment and will not be elaborated here.

[0103] Optionally, the memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 4 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.

[0104] The transceiver may include a receiver and a transmitter ( Figure 4 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver may be integrated with the processor or exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 4 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.

[0105] In addition, it should be noted that Figure 4 the structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above can refer to the technical effects described in the first embodiment above, so they will not be elaborated here.

[0106] Fourth Embodiment

[0107] This embodiment provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the method of the first embodiment above. Among them, the computer-readable storage medium may be ROM, random access memory, CD-ROM, magnetic tape, floppy disk, and optical data storage devices, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to execute the above method.

[0108] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present invention can take the form of all or part of a hardware embodiment, all or part of a software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented using software, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0109] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0110] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes. Figure 1 One process or more processes and / or boxes Figure 1 The steps of the functions specified in one box or more boxes.

[0111] It should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element. In addition, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be understood specifically with reference to the context. "At least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural (items). For example, at least one of a, b or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.

[0112] In addition, it can be understood that in various embodiments of the present invention, the magnitude of the sequence numbers of the above processes does not mean the sequence of execution order. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0113] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0114] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0115] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0116] Finally, it should be noted that the above description is only the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those of ordinary skill in the art, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle described in the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A single-image vehicle reconstruction method based on 3DGS, characterized in that: include: Acquire an autonomous driving data set, and preprocess the data in the autonomous driving data set; wherein the data in the autonomous driving data set is multi-view RGB image data obtained by capturing the surrounding environment from different perspectives by a camera installed on the autonomous driving vehicle; Constructing a vehicle reconstruction model based on 3DGS; wherein the input of the vehicle reconstruction model is a single RGB image, and the output is a reconstructed vehicle image; Use the preprocessed autonomous driving dataset to train the vehicle reconstruction model; Using the trained vehicle reconstruction model, vehicle reconstruction is achieved based on a single RGB image.

2. The single-image vehicle reconstruction method based on 3DGS according to claim 1, characterized in that: The preprocessing includes: label generation, data enhancement, data storage and data quality control.

3. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 2, characterized in that: The label generation includes: Use the pre-trained segmentation model and target detection model to generate vehicle mask data from the original RGB image and add category labels, location information, and posture angle data to the vehicle.

4. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 2, characterized in that: The data enhancement includes any one or more combinations of random rotation, cropping, scaling, and adding noise.

5. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 2, characterized in that: The data storage includes: According to the purpose and scenario of the data, the data is divided into training set, validation set and test set in preset proportions, and the data is stored in structured folders; when storing, RGB images are stored in PNG format, the generated point cloud data is stored in PLY format, and the label data is stored in json format; the data in the training set contains a variety of vehicle types, angles and environments, which are used for model training; the data in the validation set contains some scenes and angles that have not been seen by the model, which are used to evaluate the model performance; the test set is used for the final performance test of the model.

6. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 2, characterized in that: The data quality control includes: Perform quality checks on raw data to remove data that is ambiguous, incomplete, does not meet size requirements, or has incorrect labels.

7. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 1, characterized in that: The process of the vehicle reconstruction model processing the input RGB image includes: Generate a mask for the vehicle using a pre-trained mask generator; and complete the mask and texture of the vehicle by fine-tuning the lama model on the pre-processed autonomous driving dataset; Use the preset posture feature extractor to learn the posture features of the vehicle, perform feature fusion with the completed vehicle texture and mask, and then send them to the pre-trained encoder to generate latent coding features; the input RGB image and mask are loaded into the GPU memory and stored in tensor format; the latent coding features and the fused feature tensor are stored in a dedicated buffer in the GPU memory in a continuous floating-point array format; Using the ball template and the Fibonacci sampling method, a uniformly distributed point cloud is generated on the surface of the ball. Convert the 3D point cloud coordinates in the spherical template from Cartesian coordinates to polar coordinates, convert the updated polar coordinates back to rectangular coordinates, and return the deformed 3DGS point cloud attributes to obtain the vehicle point cloud; The vehicle point cloud is used for rendering and optimization to obtain the RGB image of the vehicle.

8. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 7, characterized in that: When converting the 3D point cloud coordinates in the spherical template from Cartesian coordinates to polar coordinates, the extracted features are used to predict the deformation parameters of each point through a biplane decoder, which are used to calculate the collapse factor and the updated polar coordinates.

9. The single-image vehicle reconstruction method based on 3DGS as claimed in claim 7, characterized in that: The rendering and optimization using the vehicle point cloud includes: Use the object-camera rotation matrix, camera position, image height, width, and intrinsic parameters provided by the dataset to create a primitive camera object, and use this primitive camera object to render the image of the original perspective; Create a camera object symmetrical to the original camera object, create camera pairs at certain angles, render the normalized vehicle to a 2D plane, and generate a rotation map of the vehicle; By calculating the loss between the rendered image and the original image, back propagation is performed, and the model parameters are optimized so that the reconstruction result is closer to the input image, thereby gradually approaching the three-dimensional structure of the real vehicle.

10. A single-image vehicle reconstruction system based on 3DGS, characterized in that: include: A data processing module, used to obtain an autonomous driving data set and pre-process the data in the autonomous driving data set; wherein the data in the autonomous driving data set is multi-view RGB image data obtained by capturing the surrounding environment at different viewing angles by a camera installed on the autonomous driving vehicle; A model building module, used to build a vehicle reconstruction model based on 3DGS; wherein the input of the vehicle reconstruction model is a single RGB image, and the output is a reconstructed vehicle image; A model training module, used to train the vehicle reconstruction model using the preprocessed autonomous driving dataset; The application module is used to use the trained vehicle reconstruction model to achieve vehicle reconstruction based on a single RGB image.

Citation Information

Patent Citations

  • Single-view vehicle reconstruction method and device based on implicit template mapping

    CN113160382A

  • Self-supervised diffusion method and system for generating three-dimensional Gaussian with consistent views

    CN119295647A

  • Scene reconstruction method for automatic driving simulation test, electronic equipment and product

    CN119339014A

  • Self-adaptive single object three-dimensional reconstruction and image point cloud synthesis method for automatic driving scene

    CN119516098A