Three-dimensional reconstruction method, model training method and device, vehicle, chip and medium
By introducing a generation module into the 3D reconstruction model, multi-view image information is automatically expanded, solving the problem of poor 3D reconstruction quality, achieving more efficient and accurate 3D reconstruction results, and optimizing the intelligent driving experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAOMI EV TECH CO LTD
- Filing Date
- 2026-05-26
- Publication Date
- 2026-07-31
AI Technical Summary
In existing technologies, the reconstruction quality of 3D reconstruction results is poor, making it difficult to convert 2D sensor data into a 3D world model that vehicles can understand and use, thus affecting the intelligent driving experience.
By introducing a generation module into the 3D reconstruction model, image information from the second perspective is automatically expanded based on image information from the first perspective. The generation module and reconstruction module are used to reconstruct the 3D scene from multiple perspectives, generating the 3D reconstruction result of the target scene.
It improves the geometric consistency and texture detail accuracy of 3D reconstruction results, enhances the robustness of the model to the complexity of the real world, simplifies the operation process, and improves the efficiency and quality of 3D reconstruction.
Smart Images

Figure CN122492938A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the application of image processing technology in the vehicle field, and more particularly to a three-dimensional reconstruction method, a model training method, an apparatus, a vehicle, a chip, a medium, and a program product. Background Technology
[0002] How to convert two-dimensional sensor data into a three-dimensional world model that vehicles can understand and use to improve the intelligent driving experience is an urgent problem that the industry needs to solve. Summary of the Invention
[0003] This disclosure provides a 3D reconstruction method, a model training method, an apparatus, an electronic device, a vehicle, a chip, a storage medium, and a program product to at least solve the problem of poor reconstruction quality in 3D reconstruction results in related technologies. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, a three-dimensional reconstruction method is provided, comprising: responding to a three-dimensional reconstruction instruction, determining first image information for reconstructing a target scene based on the three-dimensional reconstruction instruction; wherein the first image information is image information of the target scene from a first viewpoint; generating second image information of the target scene from at least one second viewpoint based on a generation module in a three-dimensional reconstruction model and the first image information; wherein the first viewpoint and the second viewpoint are different viewpoints; and performing three-dimensional scene reconstruction based on the reconstruction module in the three-dimensional reconstruction model, the first image information, and each of the second image information to generate a three-dimensional reconstruction result of the target scene.
[0004] The 3D reconstruction method provided in the embodiments of this disclosure, in response to a 3D reconstruction command, determines first image information for reconstructing a target scene based on the 3D reconstruction command. The first image information is image information of the target scene from a first perspective. Based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene from at least one second perspective is generated. The first and second perspectives are different perspectives. 3D scene reconstruction is performed based on the reconstruction module in the 3D reconstruction model, the first image information, and each second image information to generate a 3D reconstruction result of the target scene. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. Compared to 3D scene reconstruction based solely on image information from a single perspective, the 3D reconstruction model in this disclosure, by adding a generation module, can automatically expand the second image information from the second perspective based on the generation module and the first image information. The second image information includes scene information from the second perspective. Reconstructing the 3D scene based on image information from more perspectives helps improve the reconstruction quality of the 3D reconstruction result. For example, it can improve the consistency of the geometric structure and the accuracy of texture details in the 3D reconstruction result, making the model more robust to understanding the complexity of the real world.
[0005] In addition, this disclosure can automatically expand the second image information from the second perspective based on the generation module and the first image information, without the need for additional acquisition, input and analysis of images from other different perspectives, which reduces the acquisition requirements, simplifies the operation and improves the efficiency and quality of 3D reconstruction.
[0006] In some possible implementations, generating second image information of the target scene from at least one second perspective based on the generation module in the 3D reconstruction model and the first image information includes: extracting features from the first image information through the generation module to determine the symmetry features of the target scene; and performing a geometric transformation on the first image information based on the symmetry features of the target scene through the generation module to generate the second image information.
[0007] Therefore, if the target scene has symmetrical features, the generation module can perform geometric transformations on the first image information based on the symmetrical features of the target scene, which can accurately and efficiently expand the second image information from the second perspective.
[0008] In some possible implementations, generating second image information of the target scene from at least one second perspective based on the generation module in the 3D reconstruction model and the first image information includes: using the generation module to determine the image information from any second perspective corresponding to the first image information based on the first image information and a pre-established second mapping relationship between the image information from the first perspective and the image information from any second perspective; and using the image information from any second perspective as the second image information from any second perspective.
[0009] Therefore, the generation module can achieve accurate and efficient conversion between the first image information and the second image information based on the second mapping relationship.
[0010] In some possible implementations, generating second image information of the target scene from at least one second perspective based on the generation module in the 3D reconstruction model and the first image information includes: performing data augmentation on the first image information through the generation unit corresponding to any second perspective in the generation module to generate the second image information from any second perspective.
[0011] Therefore, considering the differences in multi-view image augmentation, setting up multiple generation units can help improve the accuracy of the second image information.
[0012] In some possible implementations, the reconstruction module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and 3D reconstruction results. The step of reconstructing a 3D scene based on the reconstruction module in the 3D reconstruction model, the first image information, and each of the second image information to generate a 3D reconstruction result of the target scene includes: generating target voxels of the target scene in 3D space using the reconstruction module, based on the first image information and each of the second image information; wherein the target voxels are used to indicate voxel information with geometric information associated with the target scene; performing 3D scene reconstruction based on the first mapping relationship and the target voxels to generate 3D reconstruction results for each target voxel; and determining the 3D reconstruction result of the target scene based on the 3D reconstruction results of each target voxel.
[0013] Therefore, by using the reconstruction module to generate target voxels in three-dimensional space based on image information from more perspectives, the problem of incomplete target voxels and large errors caused by the observation blind spots of a single perspective can be overcome. This helps to improve the integrity and accuracy of target voxels, thereby improving the reconstruction quality of the three-dimensional reconstruction results.
[0014] In addition, only the target voxels need to be reconstructed in 3D, without the need to reconstruct the background area in 3D. This can significantly reduce the computational complexity required for 3D reconstruction, thereby improving the efficiency of 3D reconstruction and saving the computational resources required for 3D reconstruction.
[0015] In some possible implementations, the step of reconstructing a 3D scene based on the first mapping relationship and the target voxels to generate a 3D reconstruction result for each target voxel includes: reconstructing a 3D scene based on the mapping relationship between real image information and the 3D reconstruction result under a viewpoint matched with the first viewpoint, the first image information, and the target voxels to determine the scene features of each target voxel; updating the scene features of each target voxel based on the mapping relationship between real image information and the 3D reconstruction result under a viewpoint matched with the second viewpoint, the second image information, and the target voxels; and generating a 3D reconstruction result for each target voxel based on the updated scene features of each target voxel.
[0016] Therefore, considering the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the first viewpoint, the first image information and the target voxels can be used to reconstruct the 3D scene to determine the scene features of each target voxel. Considering the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the second viewpoint, the second image information and the target voxels can be used to update the scene features of each target voxel. The scene features of each target voxel can be iteratively optimized using image information from multiple views to generate the 3D reconstruction result of each target voxel, which helps to improve the reconstruction quality of the 3D reconstruction result.
[0017] In some possible implementations, generating target voxels of the target scene in three-dimensional space based on the first image information and each of the second image information by the reconstruction module includes: performing a reversible transformation on the first image information and each of the second image information through a first stream network in the reconstruction module to determine a first latent feature; wherein the first latent feature is used to indicate whether each voxel in three-dimensional space has geometric information associated with the target scene; and decoding the first latent feature through a first decoder in the reconstruction module to generate the target voxels.
[0018] Therefore, flow networks have the characteristic of lossless information transmission, which can more accurately preserve the detailed information in the image (such as lighting and texture), so that the first latent features contain rich detailed information, which helps to improve the integrity and accuracy of the target voxels, and thus improve the reconstruction quality of the 3D reconstruction results.
[0019] In some possible implementations, the step of reconstructing a 3D scene based on the first mapping relationship and the target voxels to generate 3D reconstruction results for each target voxel includes: performing a reversible transformation based on the first mapping relationship and the target voxels through a second stream network in the reconstruction module to determine a second latent feature; wherein the second latent feature is configured as a scene feature for each target voxel; and decoding the second latent feature through a second decoder in the reconstruction module to generate 3D reconstruction results for each target voxel.
[0020] Therefore, flow networks have the characteristic of lossless information transmission, which can more accurately preserve the detailed information in the image (such as lighting and texture), so that the second latent features contain rich detailed information, which helps to improve the reconstruction quality of the 3D reconstruction results.
[0021] In some possible implementations, the method further includes at least one of the following: In response to the target scenario including a driving scenario, the vehicle's display screen visualizes the three-dimensional reconstruction results of the driving scenario; In response to the target scenario including a driving scenario, a path is planned for the vehicle based on the 3D reconstruction results of the driving scenario.
[0022] Therefore, in driving scenarios, the vehicle's display screen can visualize the 3D reconstruction results of the driving scenario, which can more intuitively show the driver the driving scenario (such as road conditions, obstacles, and traffic facilities), thus optimizing the driving experience.
[0023] And / or, in driving scenarios, the 3D reconstruction results of the driving scenario can be taken into account to plan the vehicle's path. The 3D reconstruction results in this solution have high reconstruction quality, which helps to improve the accuracy of path planning and ensure driving safety and user experience.
[0024] According to a second aspect of the present disclosure, a model training method is provided, comprising: determining a training dataset; wherein the training dataset includes sample real image information of a sample scene from a first viewpoint; generating predicted image information of the sample scene from at least one second viewpoint based on the sample real image information from the first viewpoint using a first initial module in a 3D reconstruction model; wherein the first viewpoint and the second viewpoint are different viewpoints; performing 3D scene reconstruction based on the sample real image information from the first viewpoint and the predicted image information from each of the second viewpoints using a second initial module in the 3D reconstruction model to generate a predicted 3D reconstruction result of the sample scene; and training the first initial module and the second initial module based on the predicted 3D reconstruction result to obtain a generation module trained by the first initial module and a reconstruction module trained by the second initial module.
[0025] The model training method provided in this disclosure determines a training dataset. The training dataset includes sample real image information of a sample scene from a first perspective. A first initial module in the 3D reconstruction model generates predicted image information of the sample scene from at least one second perspective based on the sample real image information from the first perspective. The first and second perspectives are different. A second initial module in the 3D reconstruction model performs 3D scene reconstruction based on the sample real image information from the first perspective and the predicted image information from each second perspective to generate a predicted 3D reconstruction result of the sample scene. Based on the predicted 3D reconstruction result, the first and second initial modules are trained to obtain a generation module trained by the first initial module and a reconstruction module trained by the second initial module. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. By adding a generation module, the 3D reconstruction model in this disclosure allows the first initial module to learn the ability to generate predicted image information of the sample scene from at least one second perspective based on the sample real image information from the first perspective during training, and the second initial module to learn the ability to perform 3D scene reconstruction based on the sample real image information from the first perspective and the predicted image information from each second perspective during training.
[0026] In some possible implementations, training the first initial module and the second initial module based on the predicted 3D reconstruction result to obtain a generation module trained from the first initial module and a reconstruction module trained from the second initial module includes: performing pose estimation on sample real image information under the first viewpoint to determine the pose corresponding to the sample real image information under the first viewpoint; generating predicted image information under the first viewpoint based on the predicted 3D reconstruction result and the pose corresponding to the sample real image information under the first viewpoint; and training the first initial module and the second initial module based on the difference information between the predicted image information under the first viewpoint and the sample real image information under the first viewpoint to obtain the trained generation module and the trained reconstruction module.
[0027] Therefore, considering the differences between the predicted image information and the real image information of the sample from the first perspective, the first initial module and the second initial module can be trained. That is, the difference information between the image information can be used to train the 3D reconstruction model. Compared with the 3D reconstruction result, the image information contains more detailed information (such as lighting and texture), and can speed up the model convergence. Moreover, the computational complexity required for the difference information between the image information is lower, which helps to improve the training accuracy and training efficiency of the 3D reconstruction model.
[0028] In some possible implementations, the training dataset further includes the sample 3D reconstruction results of the sample scene; the step of training the first initial module and the second initial module based on the predicted 3D reconstruction results to obtain the generation module trained by the first initial module and the reconstruction module trained by the second initial module includes: training the first initial module and the second initial module based on the difference information between the predicted 3D reconstruction results and the sample 3D reconstruction results to obtain the trained generation module and the trained reconstruction module.
[0029] Therefore, taking into account the difference between the predicted 3D reconstruction result and the sample 3D reconstruction result, the first initial module and the second initial module are trained to obtain the trained generation module and the trained reconstruction module. Thus, during the training process, the first initial module and the second initial module can jointly learn the correlation between the sample real image information and the sample 3D reconstruction result from the first perspective.
[0030] According to a third aspect of the present disclosure, a three-dimensional reconstruction apparatus is provided, comprising: a determining module configured to, in response to a three-dimensional reconstruction instruction, determine first image information for reconstructing a target scene based on the three-dimensional reconstruction instruction; wherein the first image information is image information of the target scene from a first viewpoint; a first processing module configured to generate second image information of the target scene from at least one second viewpoint based on a generating module in a three-dimensional reconstruction model and the first image information; wherein the first viewpoint and the second viewpoint are different viewpoints; and a second processing module configured to perform three-dimensional scene reconstruction based on the reconstruction module in the three-dimensional reconstruction model, the first image information, and each of the second image information, to generate a three-dimensional reconstruction result of the target scene.
[0031] In some possible implementations, the first processing module is further configured to: extract features from the first image information through the generation module to determine the symmetry features of the target scene; and perform geometric transformation on the first image information based on the symmetry features of the target scene through the generation module to generate the second image information.
[0032] In some possible implementations, the first processing module is further configured to: determine, through the generation module, the image information under any second viewpoint corresponding to the first image information based on the first image information and a pre-established second mapping relationship between the image information under the first viewpoint and the image information under any second viewpoint; and use the image information under any second viewpoint as the second image information under any second viewpoint.
[0033] In some possible implementations, the first processing module is further configured to: perform data enhancement on the first image information through the generation unit corresponding to any second viewpoint in the generation module, so as to generate second image information under any second viewpoint.
[0034] In some possible implementations, the reconstruction module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and 3D reconstruction results; the second processing module is further configured to: generate target voxels of the target scene in 3D space based on the first image information and each of the second image information, using the reconstruction module; wherein the target voxels are used to indicate voxel information with geometric information associated with the target scene; perform 3D scene reconstruction based on the first mapping relationship and the target voxels using the reconstruction module to generate 3D reconstruction results for each of the target voxels; and determine the 3D reconstruction result of the target scene based on the 3D reconstruction results of each of the target voxels using the reconstruction module.
[0035] In some possible implementations, the second processing module is further configured to: perform 3D scene reconstruction based on the mapping relationship between real image information and 3D reconstruction results under the viewpoint matched with the first viewpoint, the first image information and the target voxels, to determine the scene features of each target voxel; update the scene features of each target voxel based on the mapping relationship between real image information and 3D reconstruction results under the viewpoint matched with the second viewpoint, the second image information and the target voxels; and generate 3D reconstruction results for each target voxel based on the updated scene features of each target voxel.
[0036] In some possible implementations, the second processing module is further configured to: perform a reversible transformation on the first image information and each of the second image information through a first stream network in the reconstruction module to determine a first latent feature; wherein the first latent feature is used to indicate whether each voxel in three-dimensional space has geometric information associated with the target scene; and decode the first latent feature through a first decoder in the reconstruction module to generate the target voxel.
[0037] In some possible implementations, the second processing module is further configured to: perform a reversible transformation based on the first mapping relationship and the target voxels through a second stream network in the reconstruction module to determine a second latent feature; wherein the second latent feature is configured as scene features of each of the target voxels; and decode the second latent feature through a second decoder in the reconstruction module to generate a three-dimensional reconstruction result of each of the target voxels.
[0038] In some possible implementations, the second processing module is further configured to perform at least one of the following: In response to the target scenario including a driving scenario, the vehicle's display screen visualizes the three-dimensional reconstruction results of the driving scenario; In response to the target scenario including a driving scenario, a path is planned for the vehicle based on the 3D reconstruction results of the driving scenario.
[0039] According to a fourth aspect of the present disclosure, a model training apparatus is provided, comprising: a determining module configured to determine a training dataset; wherein the training dataset includes sample real image information of a sample scene from a first viewpoint; a first processing module configured to generate predicted image information of the sample scene from at least one second viewpoint based on the sample real image information from the first viewpoint using a first initial module in a 3D reconstruction model; wherein the first viewpoint and the second viewpoint are different viewpoints; a second processing module configured to perform 3D scene reconstruction based on the sample real image information from the first viewpoint and the predicted image information from each of the second viewpoints using a second initial module in the 3D reconstruction model, to generate a predicted 3D reconstruction result of the sample scene; and a training module configured to train the first initial module and the second initial module based on the predicted 3D reconstruction result to obtain a generation module trained by the first initial module and a reconstruction module trained by the second initial module.
[0040] In some possible implementations, the training module is further configured to: perform pose estimation on the sample real image information under the first viewpoint to determine the pose corresponding to the sample real image information under the first viewpoint; generate predicted image information under the first viewpoint based on the predicted 3D reconstruction result and the pose corresponding to the sample real image information under the first viewpoint; and train the first initial module and the second initial module based on the difference information between the predicted image information under the first viewpoint and the sample real image information under the first viewpoint to obtain the trained generation module and the trained reconstruction module.
[0041] In some possible implementations, the training dataset further includes sample 3D reconstruction results of the sample scene; the training module is further configured to: train the first initial module and the second initial module based on the difference information between the predicted 3D reconstruction results and the sample 3D reconstruction results to obtain the trained generation module and the trained reconstruction module.
[0042] According to a fifth aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the three-dimensional reconstruction method described in the first aspect of the present disclosure and / or the steps of the model training method described in the second aspect of the present disclosure.
[0043] According to a sixth aspect of the present disclosure, a vehicle is provided, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the three-dimensional reconstruction method of the first aspect of the present disclosure, and / or implement the steps of the model training method of the second aspect of the present disclosure.
[0044] According to a seventh aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which computer program instructions are stored, which, when executed by a processor, implement the steps of the three-dimensional reconstruction method described in the first aspect of the present disclosure, and / or implement the steps of the model training method described in the second aspect of the present disclosure.
[0045] According to an eighth aspect of the present disclosure, a chip is provided, the chip including an interface circuit and a processing circuit coupled to each other, the interface circuit being used to input or output signals, and the processing circuit being configured to implement the steps of the three-dimensional reconstruction method of the first aspect of the present disclosure, and / or implement the steps of the model training method of the second aspect of the present disclosure.
[0046] According to a ninth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the three-dimensional reconstruction method described in the first aspect of the present disclosure, and / or implements the steps of the model training method described in the second aspect of the present disclosure.
[0047] The technical solution provided by the embodiments of this disclosure brings at least the following beneficial effects: In response to a 3D reconstruction command, first image information for reconstructing a target scene is determined based on the 3D reconstruction command; wherein the first image information is image information of the target scene from a first perspective; based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene from at least one second perspective is generated; wherein the first perspective and the second perspective are different perspectives; 3D scene reconstruction is performed based on the reconstruction module in the 3D reconstruction model, the first image information, and each second image information to generate a 3D reconstruction result of the target scene. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. Compared to 3D scene reconstruction based solely on image information from a single perspective, the 3D reconstruction model in this disclosure, by adding a generation module, can automatically expand the second image information from the second perspective based on the generation module and the first image information. The second image information includes scene information from the second perspective. Reconstructing the 3D scene based on image information from more perspectives helps improve the reconstruction quality of the 3D reconstruction result, such as improving the consistency of the geometric structure and the accuracy of texture details in the 3D reconstruction result, making the model more robust to understanding the complexity of the real world.
[0048] In addition, this disclosure can automatically expand the second image information from the second perspective based on the generation module and the first image information, without the need for additional acquisition, input and analysis of images from other different perspectives, which reduces the acquisition requirements, simplifies the operation and improves the efficiency and quality of 3D reconstruction.
[0049] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0050] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0051] Figure 1 This is a flowchart illustrating a three-dimensional reconstruction method according to an exemplary embodiment.
[0052] Figure 2 This is a flowchart illustrating a three-dimensional reconstruction method according to another exemplary embodiment.
[0053] Figure 3 This is a flowchart illustrating a three-dimensional reconstruction method according to another exemplary embodiment.
[0054] Figure 4 This is a flowchart illustrating a three-dimensional reconstruction method according to another exemplary embodiment.
[0055] Figure 5 This is a flowchart illustrating a model training method according to an exemplary embodiment.
[0056] Figure 6 This is a schematic diagram of a model according to an exemplary embodiment.
[0057] Figure 7 This is a schematic diagram of the structure of a three-dimensional reconstruction device according to an exemplary embodiment.
[0058] Figure 8 This is a schematic diagram of the structure of a model training device according to an exemplary embodiment.
[0059] Figure 9 This is a schematic diagram of the structure of a vehicle according to an exemplary embodiment.
[0060] Figure 10 This is a schematic diagram of the structure of a chip according to an exemplary embodiment. Detailed Implementation
[0061] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0062] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0063] It should be noted that the acquisition, storage, use, and processing of data in this disclosed technical solution comply with relevant laws and regulations and do not violate public order and good morals. It should also be noted that all data processed in this disclosure is explicitly authorized by users or relevant parties, and has undergone de-identification or anonymization processing before collection and use, and does not contain any personally identifiable information or user privacy content; all data is used solely for 3D reconstruction purposes, ensuring that data security and user privacy rights are fully protected while achieving the technical effects.
[0064] The following description, with reference to the accompanying drawings, describes a three-dimensional reconstruction method, a model training method, an apparatus, an electronic device, a vehicle, a chip, and a storage medium according to embodiments of the present disclosure.
[0065] Figure 1 This is a flowchart illustrating a three-dimensional reconstruction method according to an exemplary embodiment. The three-dimensional reconstruction method of this disclosure includes the following steps.
[0066] S101, in response to the 3D reconstruction command, determines the first image information for reconstructing the target scene based on the 3D reconstruction command; wherein, the first image information is the image information of the target scene from a first perspective.
[0067] It should be noted that the execution subject of the 3D reconstruction method in this disclosure is an electronic device, such as a vehicle, terminal device, server, chip, etc. Vehicles include fuel-powered vehicles, electric vehicles, hybrid vehicles, etc. Terminal devices include mobile phones, wearable devices (such as smart glasses, smartwatches), tablet computers, etc. Servers include cloud servers, distributed servers, etc. The 3D reconstruction method in this disclosure can be executed by the 3D reconstruction device in this disclosure. The 3D reconstruction device in this disclosure can be configured in any electronic device to execute the 3D reconstruction method in this disclosure.
[0068] The target scenario is not overly restricted; it can include driving scenarios, home scenarios, etc. For example, in a driving scenario, the scene elements include vehicles, pedestrians, roads, etc., while in a home scenario, the scene elements include floors, furniture, home appliances, etc.
[0069] There are no strict restrictions on how 3D reconstruction commands are generated; for example, they can be generated automatically or based on user actions. Similarly, there are no strict restrictions on the user actions used to trigger 3D reconstruction commands; these actions can include mechanical operations, touch operations, and voice interaction.
[0070] In some possible implementations, the method further includes generating 3D reconstruction instructions in response to the target scene being a preset scene. Thus, if the target scene is a preset scene, 3D reconstruction instructions are automatically generated. It should be noted that the preset scene is not overly limited; for example, it may include highly complex intelligent driving scenarios, parking lot scenarios, etc.
[0071] In some possible implementations, the method further includes generating a 3D reconstruction command in response to the time required for the vehicle to enter the target scene corresponding area being less than a set time, and / or in response to the travel distance required for the vehicle to enter the target scene corresponding area being less than a set distance. Thus, if the vehicle is about to enter the target scene corresponding area, and / or if the vehicle is close to the target scene corresponding area, a 3D reconstruction command is automatically generated and can be displayed to the user on the vehicle's infotainment screen.
[0072] In some possible implementations, the method further includes at least one of the following: Responding to touch input, determine the 3D reconstruction instructions; Voice information is collected and voice recognition is performed to determine the 3D reconstruction instructions.
[0073] Thus, users can interact with electronic devices via touch and / or voice, and the electronic devices can respond to the above-mentioned interactive operations to determine the three-dimensional reconstruction instructions.
[0074] For example, in a driving scenario, you can click the "3D Display" button on the vehicle's interactive interface, and the vehicle can respond to the touch operation and confirm the 3D reconstruction command.
[0075] Alternatively, if the user says "Generate 3D road conditions for me," the vehicle can collect voice information and perform voice recognition to generate 3D reconstruction instructions.
[0076] It should be noted that there are no strict limitations on the first-person perspective, such as including forward view, left front, left rear, right front, right rear, and rear view. There is at least one first-person perspective, and there is at least one first image information from any first-person perspective.
[0077] In some possible implementations, the 3D reconstruction instruction carries index information of the first image information. Based on the 3D reconstruction instruction, the first image information used to reconstruct the target scene is determined, including extracting the index information of the first image information from the 3D reconstruction instruction and determining the first image information based on the index information. It should be noted that the index information is not subject to excessive limitations; for example, it may include the frame number of the image frame, the memory address it occupies, etc.
[0078] In some possible implementations, determining first image information for reconstructing the target scene based on 3D reconstruction instructions includes determining the target scene to be reconstructed based on 3D reconstruction instructions and determining the first image information based on image information associated with the target scene.
[0079] S102, based on the generation module in the 3D reconstruction model and the first image information, generate second image information of the target scene in at least one second viewpoint; wherein, the first viewpoint and the second viewpoint are different viewpoints.
[0080] It should be noted that there are no strict limitations on the 3D reconstruction model, such as including mechanistic models and data-driven models (e.g., deep learning models).
[0081] In this disclosure, such as Figure 6 As shown, the 3D reconstruction model includes a generation module and a reconstruction module.
[0082] It should be noted that the second viewpoint is not subject to many restrictions, and may include forward view, left front view, left rear view, right front view, right rear view, and rear view. There is at least one second image information from any second viewpoint.
[0083] In some possible implementations, based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene in at least one second perspective is generated. This includes using the generation module to perform geometric transformations on the first image information to generate the second image information. It should be noted that the specific methods of geometric transformation are not limited in many ways, and may include affine transformations (such as rotation, mirror flip, translation, scaling), elastic deformation, perspective transformation, etc.
[0084] In some possible implementations, based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene in at least one second perspective is generated. This includes determining generation units from the generation module that match the first and second perspectives, and processing the first image information through the generation units to generate the second image information in the second perspective. It should be noted that in this embodiment, the generation module may include multiple generation units, with different generation units employing different generation methods for multi-view augmentation. Therefore, considering the differences in multi-view image augmentation, setting multiple generation units helps improve the accuracy of the second image information.
[0085] S103, based on the reconstruction module in the 3D reconstruction model, the first image information and each second image information, perform 3D scene reconstruction to generate the 3D reconstruction result of the target scene.
[0086] 3D reconstruction refers to creating a mathematical model of a 3D object suitable for computer representation and processing, such as constructing a digital 3D model of the geometric structure, surface properties, and spatial relationships of an object or environment in the real world. However, existing 3D reconstruction methods suffer from poor reconstruction quality.
[0087] For example, how to convert two-dimensional sensor data into a three-dimensional world model that vehicles can understand and use to improve the intelligent driving experience is an urgent problem that the industry needs to solve.
[0088] To address the aforementioned issues, this disclosure proposes a three-dimensional reconstruction mechanism with expanded perspectives. Compared to reconstructing a three-dimensional scene based solely on image information from a single perspective, the three-dimensional reconstruction model in this disclosure, by adding a generation module, can automatically expand the second image information from a second perspective based on the generation module and the first image information. The second image information includes scene information from the second perspective. Reconstructing a three-dimensional scene based on image information from more perspectives helps improve the reconstruction quality of the three-dimensional reconstruction results. For example, it can improve the consistency of the geometric structure and the accuracy of texture details in the three-dimensional reconstruction results, making the model more robust to understanding the complexity of the real world.
[0089] In addition, this disclosure can automatically expand the second image information from the second perspective based on the generation module and the first image information, without the need for additional acquisition, input and analysis of images from other different perspectives, which reduces the acquisition requirements, simplifies the operation and improves the efficiency and quality of 3D reconstruction.
[0090] In some possible embodiments, this disclosure can automatically expand the second image information from a second perspective and generate a three-dimensional space based on an image taken by the user, without requiring the user to manually enter multi-view images, thus simplifying user operation and optimizing user experience.
[0091] In some possible implementations, the reconstruction module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and the 3D reconstruction results. Thus, the real image information contains complex geometric structures, lighting conditions, and material properties, which helps improve the visual fidelity of the 3D reconstruction results.
[0092] In some possible implementations, a three-dimensional scene is reconstructed based on the reconstruction module in the three-dimensional reconstruction model, the first image information, and each of the second image information to generate a three-dimensional reconstruction result of the target scene. This includes extracting features from the first image information and each of the second image information through the reconstruction module to determine the scene features of the target scene in multiple perspectives, and reconstructing the three-dimensional scene based on the scene features of the target scene in multiple perspectives and the first mapping relationship to generate a three-dimensional reconstruction result of the target scene.
[0093] In some possible implementations, the method further includes at least one of the following: In response to the target scenario, including the driving scenario, the vehicle's display screen visualizes the 3D reconstruction results of the driving scenario; In response to the target scenario, including the driving scenario, the vehicle's path is planned based on the 3D reconstruction results of the driving scenario.
[0094] Therefore, in driving scenarios, the vehicle's display screen can visualize the 3D reconstruction results of the driving scenario, which can more intuitively show the driver the driving scenario (such as road conditions, obstacles, and traffic facilities), thus optimizing the driving experience.
[0095] And / or, in driving scenarios, the 3D reconstruction results of the driving scenario can be taken into account to plan the vehicle's path. The 3D reconstruction results in this solution have high reconstruction quality, which helps to improve the accuracy of path planning and ensure driving safety and user experience.
[0096] In some possible implementations, the method further includes, in response to the target scenario, creating a simulation environment for the driving scenario, and based on the 3D reconstruction results of the driving scenario, creating a 3D simulation environment to train the intelligent driving model. Thus, in the driving scenario, the 3D reconstruction results of the driving scenario can be taken into account to create a 3D simulation environment for training the intelligent driving model. The high reconstruction quality of the 3D reconstruction results in this solution helps improve the training accuracy of the intelligent driving model, thereby improving the accuracy of the intelligent driving model and ensuring driving safety.
[0097] It should be noted that this disclosure does not impose any restrictions on the execution sequence of steps S101-S103. Figure 1 The example only demonstrates the execution of steps S101-S103 in sequence.
[0098] The 3D reconstruction method provided in the embodiments of this disclosure, in response to a 3D reconstruction command, determines first image information for reconstructing a target scene based on the 3D reconstruction command. The first image information is image information of the target scene from a first perspective. Based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene from at least one second perspective is generated. The first and second perspectives are different perspectives. 3D scene reconstruction is performed based on the reconstruction module in the 3D reconstruction model, the first image information, and each second image information to generate a 3D reconstruction result of the target scene. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. Compared to 3D scene reconstruction based solely on image information from a single perspective, the 3D reconstruction model in this disclosure, by adding a generation module, can automatically expand the second image information from the second perspective based on the generation module and the first image information. The second image information includes scene information from the second perspective. Reconstructing the 3D scene based on image information from more perspectives helps improve the reconstruction quality of the 3D reconstruction result. For example, it can improve the consistency of the geometric structure and the accuracy of texture details in the 3D reconstruction result, making the model more robust to understanding the complexity of the real world.
[0099] Figure 2 This is a flowchart illustrating a three-dimensional reconstruction method according to another exemplary embodiment, such as... Figure 2 As shown, the three-dimensional reconstruction method of this disclosure includes the following steps.
[0100] S201, in response to the 3D reconstruction command, determine the first image information for reconstructing the target scene based on the 3D reconstruction command; wherein, the first image information is the image information of the target scene from a first perspective.
[0101] The details of step S201 can be found in the above embodiments and will not be repeated here.
[0102] S202, the generation module extracts features from the first image information to determine the symmetrical features of the target scene.
[0103] S203, the generation module performs geometric transformation on the first image information based on the symmetry features of the target scene to generate the second image information.
[0104] In this embodiment, the target scene has symmetrical features, and the symmetrical features are not limited in many ways, such as including axial symmetry, central symmetry, rotational symmetry, etc.
[0105] In some possible implementations, the generation module performs a geometric transformation on the first image information based on the symmetry features of the target scene to generate the second image information. This includes performing a geometric transformation on the first image information that matches the symmetry features of the target scene to generate the second image information. Therefore, a matching relationship exists between the symmetry features of the scene and the geometric transformation process, which helps improve the accuracy of the second image information.
[0106] In some possible implementations, the generation module performs a geometric transformation on the first image information to match the symmetry features of the target scene in order to generate the second image information. This includes mirroring the first image information in response to the symmetry features of the target scene being axisymmetric, thereby generating the second image information.
[0107] For example, taking a driving scenario, if the first image information is the image information of the vehicle from the left perspective, and the symmetrical feature of the vehicle is axisymmetric, the first image information is horizontally mirrored and flipped by the generation module to generate the second image information from the right perspective.
[0108] In some possible implementations, the generation module performs a geometric transformation on the first image information to match the symmetry features of the target scene in order to generate the second image information. This includes rotating the first image information by 180 degrees in response to the symmetry features of the target scene as the center of symmetry, in order to generate the second image information.
[0109] In some possible implementations, the generation module performs a geometric transformation on the first image information to match the symmetry features of the target scene, thereby generating the second image information. This includes rotating the first image information by a target angle in response to the rotational symmetry feature of the target scene, to generate the second image information. The target angle is the angle of rotational symmetry of the target scene.
[0110] S204, through the generation module, based on the first image information and the second mapping relationship between the pre-established image information under the first viewpoint and the image information under any second viewpoint, determine the image information under any second viewpoint corresponding to the first image information.
[0111] S205, take the image information from any second viewpoint as the second image information from any second viewpoint.
[0112] In this embodiment, a second mapping relationship can be pre-established between image information from a first perspective and image information from any second perspective.
[0113] In some examples, the generation module determines the image information in the right view corresponding to the first image information based on the first image information in the left view and the second mapping relationship between the pre-established image information in the left view and the image information in the right view, and uses the image information in the right view as the second image information in the right view.
[0114] In some examples, the generation module determines the image information in the rear view corresponding to the first image information based on the first image information in the front view and the second mapping relationship between the image information in the front view and the image information in the rear view, and uses the image information in the rear view as the second image information in the rear view.
[0115] S206, the first image information is augmented by the generation unit corresponding to any second viewpoint in the generation module to generate second image information under any second viewpoint.
[0116] In this embodiment, the generation module may include multiple generation units, and there is a correspondence between the generation units and the second viewpoints. Different generation units use different generation methods to expand the multiple viewpoints. Different second viewpoints may correspond to different generation units.
[0117] There are no strict limitations on the specific methods of data augmentation, such as geometric transformations and color transformations.
[0118] In some examples, the generation module includes generation units corresponding to the forward view, the left front view, the left rear view, the right front view, the right rear view, and the rear view.
[0119] If the second perspective includes a forward-looking perspective, the first image information is augmented by the generation unit corresponding to the forward-looking perspective in the generation module to generate the second image information under the forward-looking perspective.
[0120] If the second perspective includes a left front perspective, the first image information is augmented by the generation unit corresponding to the left front perspective in the generation module to generate the second image information under the left front perspective.
[0121] It should be noted that steps S202-S203, S204-S205, and S206 are respectively one implementation method for expanding the second image information under the second perspective. These three implementation methods can be used in one way or in combination of at least two implementation methods. No further restrictions are imposed here.
[0122] In some examples, in response to a 3D reconstruction instruction, first image information for reconstructing the target scene is determined based on the 3D reconstruction instruction; wherein, the first image information is the image information of the target scene from a first perspective.
[0123] The generation module extracts features from the first image information to determine the symmetry features of the target scene. Based on the symmetry features of the target scene, the generation module rotates the first image information to generate the second image information from the second perspective 1.
[0124] The generation module determines the image information under the second perspective corresponding to the first image information based on the first image information and the second mapping relationship between the image information under the first perspective and the image information under the second perspective 2, and uses the image information under the second perspective 2 as the second image information under the second perspective 2.
[0125] The first image information is augmented by the generation unit corresponding to the second perspective 3 in the generation module to generate the second image information under the second perspective 3.
[0126] In this embodiment, the first perspective and the second perspectives 1 to 3 are different perspectives.
[0127] S207, Based on the reconstruction module in the 3D reconstruction model, the first image information and each second image information, perform 3D scene reconstruction to generate the 3D reconstruction result of the target scene.
[0128] The details of step S207 can be found in the above embodiments and will not be repeated here.
[0129] It should be noted that this disclosure does not limit the execution sequence of steps S201-S207. For example, steps S201-S203 and S207 can be implemented as independent embodiments, steps S201, S204-S205 and S207 can be implemented as independent embodiments, and steps S201 and S206-S207 can be implemented as independent embodiments.
[0130] The 3D reconstruction method provided in this disclosure extracts features from first image information using a generation module to determine the symmetry features of the target scene. Based on the symmetry features of the target scene, the generation module performs a geometric transformation on the first image information to generate second image information. Therefore, if the target scene has symmetry features, the generation module can perform a geometric transformation on the first image information based on the symmetry features of the target scene, which can accurately and efficiently expand the second image information from a second perspective.
[0131] The generation module, based on the first image information and a pre-established second mapping relationship between image information from a first viewpoint and image information from any second viewpoint, determines the image information from any second viewpoint corresponding to the first image information, and uses this image information as the second image information from that second viewpoint. Thus, the generation module can achieve accurate and efficient conversion from first image information to second image information based on the second mapping relationship.
[0132] By using a generation unit corresponding to any second viewpoint in the generation module, data augmentation is performed on the first image information to generate second image information from any second viewpoint. Therefore, considering the differences in multi-view image augmentation, setting up multiple generation units helps improve the accuracy of the second image information.
[0133] Figure 3 This is a flowchart illustrating a three-dimensional reconstruction method according to another exemplary embodiment, such as... Figure 3 As shown, the three-dimensional reconstruction method of this disclosure includes the following steps.
[0134] S301, in response to the 3D reconstruction command, determines the first image information for reconstructing the target scene based on the 3D reconstruction command; wherein, the first image information is the image information of the target scene from a first perspective.
[0135] S302, based on the generation module in the 3D reconstruction model and the first image information, generate second image information of the target scene in at least one second perspective; wherein the first perspective and the second perspective are different perspectives.
[0136] The details of steps S301-S302 can be found in the above embodiments and will not be repeated here.
[0137] S303, through the reconstruction module, based on the first image information and each of the second image information, a target voxel of the target scene in three-dimensional space is generated; wherein, the target voxel is used to indicate voxel information that has geometric information associated with the target scene.
[0138] It should be noted that a voxel, also called a volumetric pixel, refers to a basic unit in three-dimensional space. A target voxel contains geometric information associated with the target scene; that is, the target voxel is the voxel containing the foreground object.
[0139] There are no strict limitations on voxel information, such as including position information, occupancy information (e.g., the probability of the presence of foreground objects), surface and shape information (e.g., normal vectors, curvature), topological and structural information (e.g., connectivity between voxels, skeleton properties), etc.
[0140] In some possible implementations, the target voxel is characterized using a voxel skeleton.
[0141] In some possible implementations, the reconstruction module generates target voxels of the target scene in three-dimensional space based on the first image information and each of the second image information. This includes determining the geometric information of the target scene based on the first image information and each of the second image information, determining voxel information in three-dimensional space that contains geometric information associated with the target scene, and determining the target voxels based on the voxel information of the geometric information associated with the target scene.
[0142] In some possible implementations, the geometric information of the target scene is determined by the reconstruction module based on the first image information and each of the second image information. This includes extracting features from the first image information and each of the second image information by the reconstruction module to determine the geometric information of the target scene.
[0143] S304, Reconstruct the three-dimensional scene based on the first mapping relationship and the target voxels to generate the three-dimensional reconstruction results of each target voxel.
[0144] In this embodiment, the reconstruction module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and 3D reconstruction results.
[0145] In some possible implementations, a three-dimensional scene is reconstructed based on a first mapping relationship and target voxels to generate a three-dimensional reconstruction result for each target voxel. This includes reconstructing a three-dimensional scene based on the first mapping relationship, target voxels, first image information, and each second image information to generate a three-dimensional reconstruction result for each target voxel.
[0146] In some possible implementations, a three-dimensional scene reconstruction is performed based on a first mapping relationship and target voxels to generate a three-dimensional reconstruction result for each target voxel. This includes performing a three-dimensional scene reconstruction based on the first mapping relationship, target voxels, first image information, and each second image information to determine the scene features of each target voxel, and generating a three-dimensional reconstruction result for each target voxel based on the scene features of each target voxel.
[0147] In some possible implementations, a 3D scene reconstruction is performed based on a first mapping relationship and target voxels to generate a 3D reconstruction result for each target voxel. This includes performing a 3D scene reconstruction based on the mapping relationship between real image information and the 3D reconstruction result under a viewpoint matched with a first viewpoint, the first image information, and the target voxels to determine the scene features of each target voxel. Based on the mapping relationship between real image information and the 3D reconstruction result under a viewpoint matched with a second viewpoint, the second image information, and the target voxels, the scene features of each target voxel are updated. Based on the updated scene features of each target voxel, a 3D reconstruction result for each target voxel is generated. Therefore, considering the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the first viewpoint, the first image information and the target voxels can be used to reconstruct the 3D scene to determine the scene features of each target voxel. Considering the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the second viewpoint, the second image information and the target voxels can be used to update the scene features of each target voxel. The scene features of each target voxel can be iteratively optimized using image information from multiple views to generate the 3D reconstruction result of each target voxel, which helps to improve the reconstruction quality of the 3D reconstruction result.
[0148] In some cases, 3D scene reconstruction is performed based on the mapping relationship between real image information and 3D reconstruction results from the viewpoint matched with the first viewpoint, the first image information and the target voxels, in order to determine the scene features of each target voxel.
[0149] Based on the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the second viewpoint 1, the second image information under the second viewpoint 1 and the target voxels, the scene features of each target voxel are updated.
[0150] Based on the updated scene features of each target voxel, the 3D reconstruction results of each target voxel are generated.
[0151] In some cases, 3D scene reconstruction is performed based on the mapping relationship between real image information and 3D reconstruction results from the viewpoint matched with the first viewpoint, the first image information and the target voxels, in order to determine the scene features of each target voxel.
[0152] Based on the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the second viewpoint 1, the second image information under the second viewpoint 1 and the target voxels, the scene features of each target voxel are updated.
[0153] Based on the mapping relationship between the real image information and the 3D reconstruction results under the viewpoint matched with the second viewpoint 2, the second image information under the second viewpoint 2 and the target voxels, the scene features of each target voxel are updated.
[0154] Based on the updated scene features of each target voxel, the 3D reconstruction results of each target voxel are generated.
[0155] S305, Based on the 3D reconstruction results of each target voxel, determine the 3D reconstruction results of the target scene.
[0156] In some possible implementations, the three-dimensional reconstruction result of the target scene is determined based on the three-dimensional reconstruction results of each target voxel, including fusing the three-dimensional reconstruction results of each target voxel to determine the three-dimensional reconstruction result of the target scene.
[0157] It should be noted that this disclosure does not impose any restrictions on the execution sequence of steps S301-S305. Figure 3 The example only demonstrates the execution of steps S301-S305 in sequence.
[0158] The 3D reconstruction method provided in this disclosure, through a reconstruction module, generates target voxels of a target scene in 3D space based on first image information and various second image information. The target voxels indicate voxel information with geometric information associated with the target scene. 3D scene reconstruction is performed based on a first mapping relationship and the target voxels to generate 3D reconstruction results for each target voxel. Based on the 3D reconstruction results of each target voxel, the 3D reconstruction result of the target scene is determined. Therefore, by generating target voxels of the target scene in 3D space based on image information from multiple perspectives through the reconstruction module, the problem of incomplete target voxels and large errors caused by the observation blind spots of a single perspective can be overcome. This helps to improve the completeness and accuracy of target voxels, thereby improving the reconstruction quality of the 3D reconstruction result.
[0159] In addition, only the target voxels need to be reconstructed in 3D, without the need to reconstruct the background area in 3D. This can significantly reduce the computational complexity required for 3D reconstruction, thereby improving the efficiency of 3D reconstruction and saving the computational resources required for 3D reconstruction.
[0160] Figure 4 This is a flowchart illustrating a three-dimensional reconstruction method according to another exemplary embodiment, such as... Figure 4 As shown, the three-dimensional reconstruction method of this disclosure includes the following steps.
[0161] S401, in response to the 3D reconstruction command, determines the first image information for reconstructing the target scene based on the 3D reconstruction command; wherein, the first image information is the image information of the target scene from a first perspective.
[0162] S402, based on the generation module in the 3D reconstruction model and the first image information, generate second image information of the target scene in at least one second viewpoint; wherein the first viewpoint and the second viewpoint are different viewpoints.
[0163] The details of steps S401-S402 can be found in the above embodiments and will not be repeated here.
[0164] S403, through the first-stream network in the reconstruction module, reversible transformation is performed on the first image information and each second image information to determine the first latent feature; wherein, the first latent feature is used to indicate whether each voxel in the three-dimensional space has geometric information associated with the target scene.
[0165] For example, such as Figure 6 As shown, the reconstruction module includes a first-stream network and a first-stream decoder. A stream network is a deep neural network model based on normalized streams, and is essentially a reversible neural network.
[0166] The first latent feature refers to the representation of the first image information and each second image information in the latent space. The latent space is a low-dimensional, continuous, and structured vector space used to extract and retain key information from high-dimensional raw data.
[0167] A reversible transformation refers to a mathematical transformation that can be reversed. There are no strict limitations on the specific methods of reversible transformations, such as affine coupling transformations, reversible convolution, activation normalization, and feature rearrangement.
[0168] S404, the first latent feature is decoded by the first decoder in the reconstruction module to generate the target voxel.
[0169] It should be noted that there are no excessive restrictions on the first decoder, such as including VAE (Variational Autoencoder) encoders.
[0170] The first latent feature is decoded by the first decoder in the reconstruction module to generate the target voxel. Any latent feature decoding method in the relevant technology can be used to achieve this, and no further restrictions are imposed here.
[0171] S405, through the second-stream network in the reconstruction module, a reversible transformation is performed based on the first mapping relationship and the target voxel to determine the second latent feature; wherein, the second latent feature is configured as the scene feature of each target voxel.
[0172] For example, such as Figure 6 As shown, the reconstruction module includes a second-stream network and a second decoder. The second latent feature refers to the representation of each target voxel in the latent space.
[0173] In some possible implementations, the second latent feature is determined by performing a reversible transformation based on the first mapping relationship and the target voxel through the second stream network in the reconstruction module. This includes determining the second latent feature by performing a reversible transformation based on the first mapping relationship, the target voxel, the first image information, and each of the second image information through the second stream network.
[0174] In some possible implementations, a second latent feature is determined by performing a reversible transformation based on a first mapping relationship, a target voxel, first image information, and each second image information through a second-stream network. This includes performing a reversible transformation on the first image information, each second image information, and the target voxel based on a first mapping relationship through a second-stream network to determine the second latent feature.
[0175] S406, through the second decoder in the reconstruction module, decodes the second latent feature to generate the three-dimensional reconstruction results of each target voxel.
[0176] It should be noted that there are no excessive restrictions on the first decoder, such as including VAE (Variational Autoencoder) encoders.
[0177] The second latent feature is decoded by the second decoder in the reconstruction module to generate the three-dimensional reconstruction results of each target voxel. Any latent feature decoding method in related technologies can be used to achieve this, and no further restrictions are imposed here.
[0178] S407, Based on the 3D reconstruction results of each target voxel, determine the 3D reconstruction results of the target scene.
[0179] The details of step S407 can be found in the above embodiments and will not be repeated here.
[0180] It should be noted that this disclosure does not impose any restrictions on the execution sequence of steps S401-S407. Figure 4 The example only demonstrates the execution of steps S401-S407 in sequence.
[0181] The 3D reconstruction method provided in this disclosure uses a first stream network in the reconstruction module to perform reversible transformations on first image information and each second image information to determine a first latent feature. The first latent feature indicates whether each voxel in 3D space has geometric information associated with the target scene. The first latent feature is decoded by a first decoder in the reconstruction module to generate the target voxel. Therefore, the stream network has the characteristic of lossless information transmission, which can more accurately preserve detailed information in the image (such as lighting and texture), making the first latent feature contain rich detailed information. This helps improve the integrity and accuracy of the target voxel, thereby improving the reconstruction quality of the 3D reconstruction result.
[0182] The second latent features are determined by a second stream network in the reconstruction module, based on the first mapping relationship and the target voxels, through a reversible transformation. These second latent features are configured as scene features for each target voxel. The second decoder in the reconstruction module decodes these latent features to generate the 3D reconstruction results for each target voxel. Thus, the stream network, with its lossless information transmission characteristics, can more accurately preserve detailed information in the image (such as lighting and texture), resulting in rich detailed information in the second latent features, which helps improve the reconstruction quality of the 3D reconstruction results.
[0183] The training content of the 3D reconstruction model will be described below.
[0184] Figure 5 This is a flowchart illustrating a model training method according to an exemplary embodiment, such as... Figure 5 As shown, the model training method of this disclosure includes the following steps.
[0185] S501, Determine the training dataset; wherein, the training dataset includes sample real image information of the sample scene from a first-person perspective.
[0186] S502, through the first initial module in the 3D reconstruction model, based on the sample real image information from the first perspective, generates predicted image information of the sample scene from at least one second perspective; wherein, the first perspective and the second perspective are different perspectives.
[0187] It should be noted that the execution subject of the model training method in this embodiment is an electronic device, such as a vehicle, terminal device, server, or chip. The model training method in this embodiment can be executed by the model training device, which can be configured in any electronic device to execute the model training method.
[0188] The relevant content of step S502 can be referred to the relevant content of generating second image information of the target scene in at least one second perspective in the above embodiment, and will not be repeated here.
[0189] In some possible implementations, a first initial module in the 3D reconstruction model generates predicted image information of the sample scene from at least one second perspective, based on the sample's real image information from a first viewpoint, including at least one of the following: Method 1: The first initial module extracts features from the real image information of the sample from the first perspective to determine the symmetry features of the sample scene. Based on the symmetry features of the sample scene, the first initial module rotates the real image information of the sample from the first perspective to generate predicted image information.
[0190] Method 2: Through the first initial module, based on the sample real image information under the first viewpoint and the pre-established second mapping relationship between the image information under the first viewpoint and the image information under any second viewpoint, determine the image information under any second viewpoint corresponding to the sample real image information under the first viewpoint, and use the image information under any second viewpoint as the predicted image information under any second viewpoint.
[0191] Method 3: Through the generation unit corresponding to any second viewpoint in the first initial module, data augmentation is performed on the real image information of the sample under the first viewpoint to generate the predicted image information under any second viewpoint.
[0192] S503 uses the second initial module in the 3D reconstruction model to reconstruct the 3D scene based on the real image information of the sample from the first perspective and the predicted image information from each second perspective, so as to generate the predicted 3D reconstruction result of the sample scene.
[0193] It should be noted that the relevant content of step S503 can be referred to the relevant content of generating the three-dimensional reconstruction result of the target scene in the above embodiment, and will not be repeated here.
[0194] In some possible implementations, the second initial module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and the 3D reconstruction results. It should be noted that the establishment and updating of the first mapping relationship can be found in the following embodiments, and will not be repeated here.
[0195] In some possible implementations, a second initial module in the 3D reconstruction model performs 3D scene reconstruction based on the sample's real image information from the first viewpoint and the predicted image information from each second viewpoint to generate a predicted 3D reconstruction result of the sample scene. This includes generating predicted voxels of the sample scene in 3D space based on the sample's real image information from the first viewpoint and the predicted image information from each second viewpoint using the second initial module; wherein the predicted voxels are used to indicate voxel information with geometric information associated with the sample scene; 3D scene reconstruction is performed based on the first mapping relationship and the predicted voxels to generate 3D reconstruction results for each predicted voxel; and the predicted 3D reconstruction result of the sample scene is determined based on the 3D reconstruction results for each predicted voxel.
[0196] In some possible implementations, a second initialization module generates a predicted voxel of the sample scene in three-dimensional space based on the sample real image information from the first viewpoint and the predicted image information from each second viewpoint. This includes performing a reversible transformation on the sample real image information from the first viewpoint and the predicted image information from each second viewpoint using a first stream network in the second initialization module to determine a predicted first latent feature. The predicted first latent feature is used to indicate whether each voxel in three-dimensional space has geometric information associated with the sample scene. The predicted first latent feature is decoded by a first decoder in the second initialization module to generate the predicted voxel.
[0197] In some possible implementations, 3D scene reconstruction is performed based on the first mapping relationship and the predicted voxels to generate 3D reconstruction results for each predicted voxel. This includes performing a reversible transformation based on the first mapping relationship and the predicted voxels through a second stream network in the second initialization module to determine predicted second latent features. The predicted second latent features are configured as scene features for each predicted voxel. The predicted second latent features are decoded by a second decoder in the second initialization module to generate 3D reconstruction results for each predicted voxel.
[0198] S504, Based on the predicted 3D reconstruction results, the first initial module and the second initial module are trained to obtain the generated module after the first initial module is trained and the reconstructed module after the second initial module is trained.
[0199] It should be noted that, based on the predicted 3D reconstruction results, the first initial module and the second initial module can be trained using any model training method from the relevant technologies, without further restrictions here.
[0200] In some possible implementations, based on the predicted 3D reconstruction results, a first initial module and a second initial module are trained to obtain a generator module trained from the first initial module and a reconstruction module trained from the second initial module. This includes estimating the pose of the sample real image information from a first viewpoint to determine the pose corresponding to the sample real image information from the first viewpoint; generating predicted image information from the first viewpoint based on the predicted 3D reconstruction results and the pose corresponding to the sample real image information from the first viewpoint; and training the first initial module and the second initial module based on the difference information between the predicted image information from the first viewpoint and the sample real image information from the first viewpoint to obtain the trained generator module and the trained reconstruction module. Therefore, by taking into account the difference information between the predicted image information from the first viewpoint and the sample real image information from the first viewpoint when training the first initial module and the second initial module, the difference information between the image information can be used to train the 3D reconstruction model. Compared to the 3D reconstruction results, image information contains more detailed information (such as lighting and texture), which can accelerate model convergence. Furthermore, the computational complexity required for the difference information between the image information is lower, which helps to improve the training accuracy and efficiency of the 3D reconstruction model. In addition, the 3D reconstruction results of the sample scenes without manual annotation can significantly reduce the annotation difficulty of the training dataset and save the annotation cost of the training dataset.
[0201] In some possible implementations, a first initial module and a second initial module are trained based on the difference information between the predicted image information and the sample real image information in the first viewpoint to obtain a trained generation module and a trained reconstruction module. This includes determining a loss function based on the difference information between the predicted image information and the sample real image information in the first viewpoint, and training the first initial module and the second initial module based on the loss function to obtain a trained generation module and a trained reconstruction module.
[0202] In some possible implementations, the training dataset also includes the sample 3D reconstruction results of the sample scene; based on the predicted 3D reconstruction results, a first initial module and a second initial module are trained to obtain a generator module trained from the first initial module and a reconstruction module trained from the second initial module. This includes training the first initial module and the second initial module based on the difference information between the predicted 3D reconstruction results and the sample 3D reconstruction results to obtain the trained generator module and the trained reconstruction module. Thus, the difference information between the predicted 3D reconstruction results and the sample 3D reconstruction results can be taken into account when training the first initial module and the second initial module to obtain the trained generator module and the trained reconstruction module. Therefore, during the training process, the first initial module and the second initial module can jointly learn the correlation between the sample's real image information and the sample's 3D reconstruction results from a first-view perspective.
[0203] In some possible implementations, a first initial module and a second initial module are trained based on the difference information between the predicted 3D reconstruction result and the sample 3D reconstruction result to obtain a trained generation module and a trained reconstruction module. This includes determining a loss function based on the difference information between the predicted 3D reconstruction result and the sample 3D reconstruction result, and training the first initial module and the second initial module based on the loss function to obtain the trained generation module and the trained reconstruction module.
[0204] It should be noted that there are no strict restrictions on the loss function, such as cross-entropy loss function, mean squared error loss function, mean absolute error loss function, KL (Kullback-Leibler) divergence loss function, etc.
[0205] This disclosure does not impose any restrictions on the execution sequence of steps S501-S504. Figure 5 The example only demonstrates the execution of steps S501-S504 in sequence.
[0206] The model training method provided in this disclosure determines a training dataset. The training dataset includes sample real image information of a sample scene from a first perspective. A first initial module in the 3D reconstruction model generates predicted image information of the sample scene from at least one second perspective based on the sample real image information from the first perspective. The first and second perspectives are different. A second initial module in the 3D reconstruction model performs 3D scene reconstruction based on the sample real image information from the first perspective and the predicted image information from each second perspective to generate a predicted 3D reconstruction result of the sample scene. Based on the predicted 3D reconstruction result, the first and second initial modules are trained to obtain a generation module trained by the first initial module and a reconstruction module trained by the second initial module. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. By adding a generation module, the 3D reconstruction model in this disclosure allows the first initial module to learn the ability to generate predicted image information of the sample scene from at least one second perspective based on the sample real image information from the first perspective during training, and the second initial module to learn the ability to perform 3D scene reconstruction based on the sample real image information from the first perspective and the predicted image information from each second perspective during training.
[0207] Based on any of the above embodiments, the second initial module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and 3D reconstruction results.
[0208] The training dataset includes real image information of sample scenes from multiple perspectives and sample 3D reconstruction results of sample scenes. The method also includes establishing a first mapping relationship between real image information and 3D reconstruction results from multiple perspectives based on real image information and sample 3D reconstruction results from multiple perspectives.
[0209] In some possible implementations, a first mapping relationship is established between the real image information of the sample from multiple perspectives and the three-dimensional reconstruction result of the sample, including establishing a first mapping relationship between the real image information of the sample and the three-dimensional reconstruction result of the sample in the same sample scene from multiple perspectives.
[0210] In some examples, the training dataset includes real image information of sample scene 1 from multiple perspectives and the sample 3D reconstruction results of sample scene 1. A first mapping relationship is established between the real image information of sample scene 1 from multiple perspectives and the sample 3D reconstruction results of sample scene 1.
[0211] The training dataset includes real image information of sample scene 2 from multiple perspectives and 3D reconstruction results of sample scene 2. The first mapping relationship between the real image information of sample scene 2 from multiple perspectives and the 3D reconstruction results of sample scene 2 is established.
[0212] The training dataset includes real image information of sample scene 3 from multiple perspectives and 3D reconstruction results of sample scene 3. The first mapping relationship between the real image information of sample scene 3 from multiple perspectives and the 3D reconstruction results of sample scene 3 is established.
[0213] It should be noted that the second initial module can learn the first mapping relationship between real image information from multiple perspectives and 3D reconstruction results during the training process, and thus the first mapping relationship is updated as the second initial module is trained.
[0214] Figure 7 This is a schematic diagram illustrating the structure of a three-dimensional reconstruction apparatus according to an exemplary embodiment. (Refer to...) Figure 7 The three-dimensional reconstruction apparatus 700 of this disclosure includes: a determination module 701, a first processing module 702, and a second processing module 703.
[0215] The determining module 701 is configured to, in response to a 3D reconstruction instruction, determine first image information for reconstructing a target scene based on the 3D reconstruction instruction; wherein the first image information is image information of the target scene from a first viewpoint; The first processing module 702 is configured to generate second image information of the target scene from at least one second perspective based on the generation module in the 3D reconstruction model and the first image information; wherein the first perspective and the second perspective are different perspectives. The second processing module 703 is configured to perform three-dimensional scene reconstruction based on the reconstruction module in the three-dimensional reconstruction model, the first image information, and each of the second image information, so as to generate a three-dimensional reconstruction result of the target scene.
[0216] In some possible implementations, the first processing module 702 is further configured to: extract features from the first image information through the generation module to determine the symmetry features of the target scene; and perform geometric transformation on the first image information based on the symmetry features of the target scene through the generation module to generate the second image information.
[0217] In some possible implementations, the first processing module 702 is further configured to: determine, through the generation module, the image information under any second viewpoint corresponding to the first image information based on the first image information and a pre-established second mapping relationship between the image information under the first viewpoint and the image information under any second viewpoint; and use the image information under any second viewpoint as the second image information under any second viewpoint.
[0218] In some possible implementations, the first processing module 702 is further configured to: perform data enhancement on the first image information through the generation unit corresponding to any second viewpoint in the generation module, so as to generate second image information under any second viewpoint.
[0219] In some possible implementations, the reconstruction module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and 3D reconstruction results; the second processing module 703 is further configured to: generate target voxels of the target scene in 3D space based on the first image information and each of the second image information, using the reconstruction module; wherein the target voxels are used to indicate voxel information with geometric information associated with the target scene; perform 3D scene reconstruction based on the first mapping relationship and the target voxels to generate 3D reconstruction results for each of the target voxels; and determine the 3D reconstruction result of the target scene based on the 3D reconstruction results of each of the target voxels.
[0220] In some possible implementations, the second processing module 703 is further configured to: perform three-dimensional scene reconstruction based on the mapping relationship between real image information and three-dimensional reconstruction results under the viewpoint matched with the first viewpoint, the first image information and the target voxels, to determine the scene features of each target voxel; update the scene features of each target voxel based on the mapping relationship between real image information and three-dimensional reconstruction results under the viewpoint matched with the second viewpoint, the second image information and the target voxels; and generate a three-dimensional reconstruction result for each target voxel based on the updated scene features of each target voxel.
[0221] In some possible implementations, the second processing module 703 is further configured to: perform a reversible transformation on the first image information and each of the second image information through a first stream network in the reconstruction module to determine a first latent feature; wherein the first latent feature is used to indicate whether each voxel in three-dimensional space has geometric information associated with the target scene; and decode the first latent feature through a first decoder in the reconstruction module to generate the target voxel.
[0222] In some possible implementations, the second processing module 703 is further configured to: perform a reversible transformation based on the first mapping relationship and the target voxels through the second stream network in the reconstruction module to determine a second latent feature; wherein the second latent feature is configured as scene features of each of the target voxels; and decode the second latent feature through the second decoder in the reconstruction module to generate a three-dimensional reconstruction result of each of the target voxels.
[0223] In some possible implementations, the second processing module 703 is further configured to perform at least one of the following: In response to the target scenario including a driving scenario, the vehicle's display screen visualizes the three-dimensional reconstruction results of the driving scenario; In response to the target scenario including a driving scenario, a path is planned for the vehicle based on the 3D reconstruction results of the driving scenario.
[0224] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0225] The 3D reconstruction apparatus provided in the embodiments of this disclosure, in response to a 3D reconstruction command, determines first image information for reconstructing a target scene based on the 3D reconstruction command. The first image information is image information of the target scene from a first perspective. Based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene from at least one second perspective is generated. The first and second perspectives are different perspectives. 3D scene reconstruction is performed based on the reconstruction module in the 3D reconstruction model, the first image information, and each second image information to generate a 3D reconstruction result of the target scene. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. Compared to 3D scene reconstruction based solely on image information from a single perspective, the 3D reconstruction model in this disclosure, by adding a generation module, can automatically expand the second image information from the second perspective based on the generation module and the first image information. The second image information includes scene information from the second perspective. Reconstructing the 3D scene based on image information from more perspectives helps improve the reconstruction quality of the 3D reconstruction result. For example, it can improve the consistency of the geometric structure and the accuracy of texture details in the 3D reconstruction result, making the model more robust to understanding the complexity of the real world.
[0226] Figure 8 This is a schematic diagram illustrating the structure of a model training device according to an exemplary embodiment. (Refer to...) Figure 8 The model training apparatus 800 of this embodiment includes: a determination module 801, a first processing module 802, a second processing module 803, and a training module 804.
[0227] The determination module 801 is configured to determine the training dataset; wherein the training dataset includes real image information of the sample scene from a first-view perspective; The first processing module 802 is configured to generate predicted image information of the sample scene in at least one second viewpoint based on the sample real image information from the first viewpoint through the first initial module in the 3D reconstruction model; wherein the first viewpoint and the second viewpoint are different viewpoints. The second processing module 803 is configured to perform three-dimensional scene reconstruction based on the sample real image information from the first viewpoint and the predicted image information from each of the second viewpoints through the second initial module in the three-dimensional reconstruction model, so as to generate the predicted three-dimensional reconstruction result of the sample scene. Training module 804 is configured to train the first initial module and the second initial module based on the predicted 3D reconstruction results, so as to obtain the generation module after the first initial module is trained and the reconstruction module after the second initial module is trained.
[0228] In some possible implementations, the training module 804 is further configured to: perform pose estimation on the sample real image information under the first viewpoint to determine the pose corresponding to the sample real image information under the first viewpoint; generate predicted image information under the first viewpoint based on the predicted 3D reconstruction result and the pose corresponding to the sample real image information under the first viewpoint; and train the first initial module and the second initial module based on the difference information between the predicted image information under the first viewpoint and the sample real image information under the first viewpoint to obtain the trained generation module and the trained reconstruction module.
[0229] In some possible implementations, the training dataset further includes sample 3D reconstruction results of the sample scene; the training module 804 is further configured to: train the first initial module and the second initial module based on the difference information between the predicted 3D reconstruction results and the sample 3D reconstruction results to obtain the trained generation module and the trained reconstruction module.
[0230] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0231] The model training apparatus provided in the embodiments of this disclosure determines a training dataset. The training dataset includes sample real image information of a sample scene from a first perspective. A first initial module in the 3D reconstruction model generates predicted image information of the sample scene from at least one second perspective based on the sample real image information from the first perspective. The first and second perspectives are different. A second initial module in the 3D reconstruction model performs 3D scene reconstruction based on the sample real image information from the first perspective and the predicted image information from each second perspective to generate a predicted 3D reconstruction result of the sample scene. Based on the predicted 3D reconstruction result, the first and second initial modules are trained to obtain a generation module trained by the first initial module and a reconstruction module trained by the second initial module. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. By adding a generation module, the 3D reconstruction model in this disclosure allows the first initial module to learn the ability to generate predicted image information of the sample scene from at least one second perspective based on the sample real image information from the first perspective during training, and the second initial module to learn the ability to perform 3D scene reconstruction based on the sample real image information from the first perspective and the predicted image information from each second perspective during training.
[0232] To implement the above embodiments, this disclosure also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the three-dimensional reconstruction method provided by this disclosure and / or the steps of the model training method provided by this disclosure.
[0233] Among them, electronic devices include, but are not limited to: terminals, servers (or cloud, servers, etc.). Terminals can be cars with communication functions, smart cars, mobile phones, wearable devices, tablets (Pads), computers with wireless transceiver functions, supercomputers, etc.
[0234] To implement the above embodiments, this disclosure also proposes a vehicle, including a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the steps of the three-dimensional reconstruction method provided in this disclosure, and / or the steps of the model training method provided in this disclosure.
[0235] Figure 9This is a schematic diagram illustrating the structure of a vehicle according to an exemplary embodiment. For example, vehicle 900 can be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicles. Vehicle 900 can be an intelligent driving vehicle, a semi-intelligent driving vehicle, or a non-intelligent driving vehicle. It should be noted that intelligent driving (also called assisted driving) refers to the technology that uses sensors, algorithms, and artificial intelligence technologies to perform environmental perception, decision-making, planning, and control command execution of the vehicle to assist the driver in driving more safely and efficiently.
[0236] Reference Figure 9 The vehicle 900 may include various subsystems, such as an infotainment system 910, a perception system 920, a decision control system 930, a drive system 940, and a computing platform 950. The vehicle 900 may also include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of the vehicle 900 can be interconnected via wired or wireless means.
[0237] In some embodiments, the infotainment system 910 may include a communication system, an entertainment system, and a navigation system, etc.
[0238] The perception system 920 may include several sensors for sensing information about the environment surrounding the vehicle 900. For example, the perception system 920 may include a global positioning system (which may be GPS, BeiDou, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter-wave radar, ultrasonic radar, and a camera device.
[0239] The decision control system 930 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.
[0240] The drive system 940 may include components that provide powered motion to the vehicle 900. In one embodiment, the drive system 940 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of internal combustion engines, electric motors, and compressed air engines. The engine is capable of converting energy provided by the energy source into mechanical energy.
[0241] Some or all of the functions of the vehicle 900 are controlled by a computing platform 950. The computing platform 950 may include at least one processor 951 and a memory 952, the processor 951 being able to execute instructions 953 stored in the memory 952.
[0242] Processor 951 can be any conventional processor. Processors may also include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), systems on chips (SoCs), application-specific integrated circuits (ASICs), or combinations thereof.
[0243] The memory 952 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0244] In addition to instruction 953, memory 952 can also store data, such as road maps, route information, vehicle position, direction, speed, and other data. The data stored in memory 952 can be used by computing platform 950.
[0245] In this embodiment of the disclosure, processor 951 may execute instructions 953 to implement all or part of the steps of the three-dimensional reconstruction method provided in this disclosure, and / or to implement all or part of the steps of the model training method provided in this disclosure.
[0246] The vehicle in this embodiment of the present disclosure, in response to a 3D reconstruction command, determines first image information for reconstructing a target scene based on the 3D reconstruction command. The first image information is an image of the target scene from a first perspective. Based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene from at least one second perspective is generated. The first and second perspectives are different perspectives. 3D scene reconstruction is performed based on the reconstruction module in the 3D reconstruction model, the first image information, and each second image information to generate a 3D reconstruction result of the target scene. Therefore, this disclosure proposes a perspective-expanded 3D reconstruction mechanism. Compared to 3D scene reconstruction based solely on image information from a single perspective, the 3D reconstruction model in this disclosure, by adding a generation module, can automatically expand the second image information from the second perspective based on the generation module and the first image information. The second image information includes scene information from the second perspective. Reconstructing the 3D scene based on image information from more perspectives helps improve the reconstruction quality of the 3D reconstruction result. For example, it can improve the consistency of the geometric structure and the accuracy of texture details in the 3D reconstruction result, making the model more robust to understanding the complexity of the real world.
[0247] To implement the above embodiments, this disclosure also proposes a non-transitory computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the steps of the three-dimensional reconstruction method provided in this disclosure, and / or implement the steps of the model training method provided in this disclosure.
[0248] In some possible implementations, the non-transitory computer-readable storage medium may be a ROM, CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0249] To implement the above embodiments, this disclosure also proposes a chip including an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the three-dimensional reconstruction method provided by this disclosure, and / or implement the steps of the model training method provided by this disclosure.
[0250] Figure 10 This is a schematic diagram illustrating the structure of a chip according to an exemplary embodiment. See also... Figure 10 The diagram shown is a schematic representation of the structure of chip 1000, but it is not limited to this.
[0251] Chip 1000 includes processing circuit 1001, which is configured to perform any of the above-described 3D reconstruction methods and / or perform any of the above-described model training methods.
[0252] In some embodiments, the chip 1000 further includes one or more interface circuits 1002. In some possible implementations, the interface circuit 1002 is connected to the memory 1003, and the interface circuit 1002 can be used to receive signals from the memory 1003 or other devices, and the interface circuit 1002 can be used to send signals to the memory 1003 or other devices. For example, the interface circuit 1002 can read instructions stored in the memory 1003 and send the instructions to the processing circuit 1001.
[0253] In some embodiments, the interface circuit 1002 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 1001 performs other steps.
[0254] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0255] In some embodiments, chip 1000 further includes one or more memories 1003 for storing instructions. In some possible embodiments, all or part of the memories 1003 may be located outside of chip 1000.
[0256] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program that, when executed by a processor, implements the steps of the three-dimensional reconstruction method provided by this disclosure, and / or the steps of the model training method provided by this disclosure.
[0257] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0258] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0259] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of this disclosure includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of this disclosure pertain.
[0260] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, wherein the sequenced list is a collection of elements in a document structure or data structure arranged in a specific order and the order of the elements has a clear meaning, wherein changing the order of the elements affects the meaning or function of the list information, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and compact disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0261] It should be understood that various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0262] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0263] Furthermore, the modules in the various embodiments of this disclosure can be implemented either in hardware or as software functional modules. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0264] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A three-dimensional reconstruction method, characterized in that, include: In response to a 3D reconstruction command, first image information for reconstructing a target scene is determined based on the 3D reconstruction command; wherein, the first image information is image information of the target scene from a first viewpoint; Based on the generation module in the 3D reconstruction model and the first image information, second image information of the target scene in at least one second viewpoint is generated; wherein, the first viewpoint and the second viewpoint are different viewpoints; Based on the reconstruction module in the three-dimensional reconstruction model, the first image information, and each of the second image information, a three-dimensional scene is reconstructed to generate a three-dimensional reconstruction result of the target scene.
2. The method according to claim 1, characterized in that, The generation module based on the 3D reconstruction model and the first image information generates second image information of the target scene from at least one second perspective, including: The generation module extracts features from the first image information to determine the symmetry features of the target scene. The generation module performs a geometric transformation on the first image information based on the symmetry features of the target scene to generate the second image information.
3. The method according to claim 1, characterized in that, The generation module based on the 3D reconstruction model and the first image information generates second image information of the target scene from at least one second perspective, including: The generation module determines the image information corresponding to the first image information based on the first image information and the pre-established second mapping relationship between the image information under the first viewpoint and the image information under any second viewpoint. The image information from any second perspective is used as the second image information from any second perspective.
4. The method according to claim 1, characterized in that, The generation module based on the 3D reconstruction model and the first image information generates second image information of the target scene from at least one second perspective, including: The first image information is augmented by the generation unit corresponding to any second viewpoint in the generation module to generate the second image information under any second viewpoint.
5. The method according to claim 1, characterized in that, The reconstruction module supports 3D reconstruction by establishing a first mapping relationship between real image information from multiple perspectives and the 3D reconstruction result; the 3D scene reconstruction based on the reconstruction module in the 3D reconstruction model, the first image information, and each of the second image information to generate the 3D reconstruction result of the target scene includes: The reconstruction module generates target voxels of the target scene in three-dimensional space based on the first image information and each of the second image information; wherein, the target voxels are used to indicate voxel information that has geometric information associated with the target scene; Based on the first mapping relationship and the target voxels, a three-dimensional scene is reconstructed to generate three-dimensional reconstruction results for each target voxel; Based on the three-dimensional reconstruction results of each target voxel, the three-dimensional reconstruction result of the target scene is determined.
6. The method according to claim 5, characterized in that, The step of reconstructing a 3D scene based on the first mapping relationship and the target voxels to generate 3D reconstruction results for each target voxel includes: Based on the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the first viewpoint, the first image information and the target voxels are used to reconstruct the 3D scene to determine the scene features of each target voxel; Based on the mapping relationship between the real image information and the 3D reconstruction result under the viewpoint matched with the second viewpoint, the second image information and the target voxel, the scene features of each target voxel are updated; Based on the updated scene features of each target voxel, the three-dimensional reconstruction results of each target voxel are generated.
7. The method according to claim 5, characterized in that, The step of generating target voxels of the target scene in three-dimensional space through the reconstruction module, based on the first image information and each of the second image information, includes: The first image information and each of the second image information are reversibly transformed by the first-stream network in the reconstruction module to determine the first latent feature; wherein the first latent feature is used to indicate whether each voxel in three-dimensional space has geometric information associated with the target scene; The first latent feature is decoded by the first decoder in the reconstruction module to generate the target voxel.
8. The method according to claim 5, characterized in that, The step of reconstructing a 3D scene based on the first mapping relationship and the target voxels to generate 3D reconstruction results for each target voxel includes: The second latent feature is determined by performing a reversible transformation based on the first mapping relationship and the target voxel through the second stream network in the reconstruction module; wherein the second latent feature is configured as the scene feature of each target voxel; The second latent feature is decoded by the second decoder in the reconstruction module to generate the three-dimensional reconstruction results of each target voxel.
9. The method according to any one of claims 1-8, characterized in that, The method further includes at least one of the following: In response to the target scenario including a driving scenario, the vehicle's display screen visualizes the three-dimensional reconstruction results of the driving scenario; In response to the target scenario including a driving scenario, a path is planned for the vehicle based on the 3D reconstruction results of the driving scenario.
10. A model training method, characterized in that, include: Determine the training dataset; wherein, the training dataset includes real image information of sample scenes from a first-view perspective; Using the first initial module in the 3D reconstruction model, based on the real image information of the sample from the first perspective, predictive image information of the sample scene in at least one second perspective is generated; wherein, the first perspective and the second perspective are different perspectives; The second initial module in the 3D reconstruction model performs 3D scene reconstruction based on the real image information of the sample from the first perspective and the predicted image information from each of the second perspectives, so as to generate the predicted 3D reconstruction result of the sample scene. Based on the predicted 3D reconstruction results, the first initial module and the second initial module are trained to obtain the generation module after the first initial module is trained and the reconstruction module after the second initial module is trained.
11. The method according to claim 10, characterized in that, The step of training the first initial module and the second initial module based on the predicted 3D reconstruction results to obtain the generation module after training the first initial module and the reconstruction module after training the second initial module includes: Pose estimation is performed on the sample real image information under the first viewpoint to determine the pose corresponding to the sample real image information under the first viewpoint. Based on the predicted 3D reconstruction result and the pose corresponding to the sample real image information under the first viewpoint, the predicted image information under the first viewpoint is generated. Based on the difference between the predicted image information from the first viewpoint and the sample real image information from the first viewpoint, the first initial module and the second initial module are trained to obtain the trained generation module and the trained reconstruction module.
12. The method according to claim 10, characterized in that, The training dataset also includes the sample 3D reconstruction results of the sample scene; the step of training the first initial module and the second initial module based on the predicted 3D reconstruction results to obtain the generation module after training the first initial module and the reconstruction module after training the second initial module includes: Based on the difference information between the predicted 3D reconstruction result and the sample 3D reconstruction result, the first initial module and the second initial module are trained to obtain the trained generation module and the trained reconstruction module.
13. A three-dimensional reconstruction device, characterized in that, include: The determination module is configured to, in response to a 3D reconstruction instruction, determine first image information for reconstructing a target scene based on the 3D reconstruction instruction; wherein the first image information is image information of the target scene from a first viewpoint; The first processing module is configured to generate second image information of the target scene from at least one second perspective based on the generation module in the 3D reconstruction model and the first image information; wherein the first perspective and the second perspective are different perspectives. The second processing module is configured to perform three-dimensional scene reconstruction based on the reconstruction module in the three-dimensional reconstruction model, the first image information, and each of the second image information, so as to generate a three-dimensional reconstruction result of the target scene.
14. The apparatus according to claim 13, characterized in that, The first processing module is further configured to: The generation module extracts features from the first image information to determine the symmetry features of the target scene. The generation module performs a geometric transformation on the first image information based on the symmetry features of the target scene to generate the second image information.
15. The apparatus according to claim 13, characterized in that, The first processing module is further configured to: The generation module determines the image information corresponding to the first image information based on the first image information and the pre-established second mapping relationship between the image information under the first viewpoint and the image information under any second viewpoint. The image information from any second perspective is used as the second image information from any second perspective.
16. A model training device, characterized in that, include: The determination module is configured to determine the training dataset; wherein the training dataset includes real image information of sample scenes from a first-view perspective; The first processing module is configured to generate predicted image information of the sample scene in at least one second viewpoint based on the sample real image information from the first viewpoint through the first initial module in the 3D reconstruction model; wherein the first viewpoint and the second viewpoint are different viewpoints. The second processing module is configured to perform three-dimensional scene reconstruction based on the sample real image information from the first viewpoint and the predicted image information from each of the second viewpoints through the second initial module in the three-dimensional reconstruction model, so as to generate the predicted three-dimensional reconstruction result of the sample scene. The training module is configured to train the first initial module and the second initial module based on the predicted 3D reconstruction results, so as to obtain the generation module after the first initial module is trained and the reconstruction module after the second initial module is trained.
17. The apparatus according to claim 16, characterized in that, The training module is also configured to: Pose estimation is performed on the sample real image information under the first viewpoint to determine the pose corresponding to the sample real image information under the first viewpoint. Based on the predicted 3D reconstruction result and the pose corresponding to the sample real image information under the first viewpoint, the predicted image information under the first viewpoint is generated. Based on the difference between the predicted image information from the first viewpoint and the sample real image information from the first viewpoint, the first initial module and the second initial module are trained to obtain the trained generation module and the trained reconstruction module.
18. The apparatus according to claim 16, characterized in that, The training dataset also includes the 3D reconstruction results of the sample scene; the training module is further configured to: Based on the difference information between the predicted 3D reconstruction result and the sample 3D reconstruction result, the first initial module and the second initial module are trained to obtain the trained generation module and the trained reconstruction module.
19. A vehicle, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured as follows: The steps of implementing the method according to any one of claims 1-9, and / or the steps of implementing the method according to any one of claims 10-12.
20. A non-transitory computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method according to any one of claims 1-9, and / or implement the steps of the method according to any one of claims 10-12.
21. A chip, characterized in that, The chip includes an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the method according to any one of claims 1-9, and / or implement the steps of the method according to any one of claims 10-12.
22. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-9, and / or implements the steps of the method according to any one of claims 10-12.