A bottom-up single-image panoramic reconstruction method, device, and computer device
Through the bottom-up single-image panoramic reconstruction method, the spatially-aware reverse projection module and the combination of the 2D instance center point and 3D offset is solved, and the problem of uncertainty in the arrangement of occluded areas and channels in the prior art is achieved, achieving more accurate and efficient panoramic reconstruction.
Patent Information
- Application Number
- CN202310650872.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-06-02
AI Technical Summary
In the single-image panoramic reconstruction, the top-down frame based on depth estimation leads to uncertainty in the occluded area and uncertainty in the arrangement of the channel, affecting the accuracy of the panoramic reconstruction.
Using a bottom-up single-image panoramic reconstruction method, the space-aware reverse projection module is used to predict the depth space that an object may exist, and the 2D semantic segmentation results are reverse projected to the 3D space to generate complete 3D space initialization features, and instance grouping is performed through the 2D instance center point and 3D offset.
The results of single-image panoramic reconstruction are significantly improved, the problem of uncertainty in occluded areas and channel arrangement is solved, and the accuracy and performance of panoramic reconstruction are improved.
Smart Images

Figure CN116681831B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology. Specifically, it relates to a bottom-up single-image panoramic reconstruction method, apparatus, and computer device. Background Art
[0002] Single-image panoramic reconstruction mainly studies how to use a single 2D image as input to reconstruct the entire 3D scene, and at the same time perform foreground instance individual segmentation and background semantic segmentation on the scene. To achieve this task, sufficient spatial information and semantic information need to be extracted from the single 2D image, so that the initial 3D features obtained by back-projecting the extracted 2D information are more accurate, and then the 3D model can better reconstruct and segment the scene. Currently, the solutions for single-image panoramic reconstruction are mainly divided into three stages: 1) 2D stage: Use a 2D model to predict 2D segmentation and 2D spatial information (such as depth estimation); 2) 2D-3D stage: Use 2D spatial information to back-project 2D segmentation information into 3D for initializing 3D features; 3) 3D stage: Input the initialized 3D features into a 3D model to reconstruct the scene and predict the panoramic segmentation result of the scene.
[0003] In the 2D stage, 2D segmentation models usually include two types: top-down (usually an instance segmentation model) and bottom-up (usually capable of panoramic segmentation). The general top-down method first predicts the object category and its bounding box, and then predicts the instance mask within the bounding box to obtain the instance segmentation result; while the typical bottom-up method predicts the center point of each instance and the relative offset of the pixels within the instance to the center point, and at the same time predicts 2D semantic segmentation. In the post-processing stage, the predicted center points and relative offsets are first used for instance grouping, and then the semantic segmentation result is used to classify the instances. Finally, the foreground instance individuals and the semantic background are combined to form the panoramic segmentation result; in the 2D-3D stage, the usual method uses the depth estimated by the 2D model to back-project to obtain the 3D object surface, and then uses the 2D segmentation result to fill the object surface to form the initialized 3D features. However, the 3D features obtained by this solution only exist on the object surface and cannot obtain information about occluded areas; in the 3D stage, a 3D model composed of 3D sparse convolution is used to reconstruct the scene and perform panoramic segmentation on the scene.
[0004] The prior art usually uses top-down Mask R-CNN for instance segmentation in the 2D stage to obtain the instance mask, and at the same time uses monocular depth estimation to predict the object surface depth; in the 2D-3D stage, the 2D instance masks are randomly arranged and then back-projected into 3D using the depth and camera intrinsics to form the 3D features of the initialized object surface; in the 3D stage, the 3D features are input into a 3D model to predict the corresponding arranged 3D instance masks, and at the same time predict semantic segmentation for instance classification. Finally, the predicted instance individuals and the semantic background are combined to form the panoramic reconstruction result.
[0005] However, the existing technology is only based on a top-down framework of depth estimation, which will cause uncertainties in occluded areas and uncertainties in channel arrangements when back-projecting the 2D instance mask to obtain the initial 3D features. Summary of the Invention
[0006] In order to solve the problem that the existing technology, which is only based on a top-down framework of depth estimation, will cause uncertainties in occluded areas and uncertainties in channel arrangements when back-projecting the 2D instance mask to obtain the initial 3D features, this application provides a bottom-up single-image panoramic reconstruction method, device, and computer device. With the help of a bottom-up framework of spatial perception, the present invention greatly improves the final panoramic reconstruction result.
[0007] The embodiments of this application are implemented as follows:
[0008] In a first aspect, this application provides a bottom-up single-image panoramic reconstruction method, including:
[0009] Obtain an image and input the single image into a 2D model;
[0010] Predict the depth space where an object may exist, the 2D semantic segmentation result, the object surface depth, and the 2D instance center point according to the 2D model;
[0011] Generate initial features of a complete 3D space based on the depth space, the 2D semantic segmentation result, and the object surface depth, using a spatially aware back-projection module;
[0012] Predict the initial features as a 3D reconstruction result, a 3D semantic segmentation result, and a 3D offset result based on a 3D model;
[0013] According to the 3D reconstruction result, the 3D semantic segmentation result, and the 3D offset result, perform instance grouping and synthesis based on a panoramic reconstruction module in combination with the 2D instance center point to obtain the final result of panoramic reconstruction.
[0014] In a possible implementation, the spatially aware back-projection module obtains a complete spatially aware 3D feature according to the depth space where an object may exist predicted by the 2D model and the object surface depth.
[0015] In a possible implementation, in the step of generating initial features of a complete 3D space based on the depth space, the 2D semantic segmentation result, and the object surface depth, using a spatially aware back-projection module, it further includes:
[0016] Back-project the depth space where an object may exist predicted by the 2D model to the 3D space using the camera internal parameters and the predicted depth;
[0017] Back-project the 2D semantic segmentation results and fill the entire 3D space;
[0018] Multiply the two after 3D sparse convolution to obtain spatially-aware initialized 3D features.
[0019] In a possible implementation, the 2D semantic segmentation results are used by a spatially-aware back-projection module to obtain initialized 3D features, and the 2D instance center points are used to group 3D voxels in combination with 3D offsets.
[0020] In a possible implementation, in the step of performing instance grouping and synthesis based on the panoramic reconstruction module in combination with the 2D instance center points according to the 3D reconstruction results, 3D semantic segmentation results, and 3D offset results to obtain the final result of panoramic reconstruction, it further includes:
[0021] Filter the 3D semantic segmentation results through the 3D reconstruction results to obtain refined 3D semantic reconstruction results;
[0022] Filter the 3D offset results through the 3D reconstruction results to obtain 3D offset reconstruction results;
[0023] In the 3D semantic reconstruction results, each foreground semantic category is input into an instance grouping module to generate instances in combination with 2D instance center points, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result.
[0024] In a possible implementation, in the step that in the 3D semantic reconstruction results, each foreground semantic category is input into an instance grouping module to generate instances in combination with 2D instance center points, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result, it further includes:
[0025] According to the foreground semantic category, obtain the 3D offset reconstruction result and 2D instance center point of this category;
[0026] Based on the 3D offset reconstruction result and 2D instance center point, perform grouping to obtain 3D instance segmentation results;
[0027] The 3D instance segmentation results are combined with the instance segmentation results and background semantics to obtain the panoramic reconstruction result.
[0028] In a possible implementation, in the step of obtaining the 3D offset reconstruction result and 2D instance center point of this category according to the foreground semantic category, it further includes:
[0029] The 3D offset reconstruction result and the 3D semantic reconstruction result of this category are projected and transformed into a multi-layer depth space;
[0030] Filter the 3D offset reconstruction results through the 3D semantic reconstruction results of the category to obtain the 3D offset reconstruction results of the category;
[0031] Extract the 2D instance center points of the category from all the 2D instance center points for category instance grouping.
[0032] In a possible implementation manner, in the step of grouping based on the 3D offset reconstruction results and 2D instance center points to obtain the 3D instance segmentation results, it further includes:
[0033] Add the 3D offset reconstruction result of each voxel of the category to the coordinates of the voxel to obtain the predicted 2D instance center point of the voxel;
[0034] According to the distance from the predicted 2D instance center point to all the actual 2D instance center points, assign the voxel to the actual 2D instance center point closest to its predicted 2D instance center point;
[0035] After completing the instance grouping of all voxels of all categories, the 3D instance segmentation results can be obtained.
[0036] In a second aspect, the present application provides a bottom-up single-image panoramic reconstruction device, including:
[0037] A 2D acquisition module, configured to acquire an image and input the single image into a 2D model;
[0038] A 2D prediction module, configured to predict the depth space where an object may exist, 2D semantic segmentation results, object surface depth, and 2D instance center points according to the 2D model;
[0039] A 3D conversion module, configured to generate initial features of a complete 3D space based on the depth space, 2D semantic segmentation results, object surface depth, and through a spatially aware back-projection module;
[0040] A 3D prediction module, configured to predict the 3D reconstruction results, 3D semantic segmentation results, and 3D offset results from the initial features based on a 3D model;
[0041] A 3D reconstruction module, configured to perform instance grouping and synthesis based on the 3D reconstruction results, 3D semantic segmentation results, and 3D offset results, and combine with the 2D instance center points through a panoramic reconstruction module to obtain the final result of panoramic reconstruction.
[0042] In a third aspect, the present application provides a computer device, which includes a memory and a processor. When the processor calls and executes a computer program stored in the memory, the steps of the bottom-up single-image panoramic reconstruction method shown in any item of the first aspect are implemented.
[0043] The technical solution provided by this application can at least achieve the following beneficial effects:
[0044] This application provides a bottom-up single-image panoramic reconstruction method, device, and computer device. Among them, a bottom-up panoramic reconstruction framework is proposed, which is the first bottom-up solution for single-image panoramic reconstruction. And to avoid the uncertainty of instance channel arrangement, 2D semantic segmentation results with fixed channels are designed to be used for the initialization of 3D features. At the same time, 2D instance center points are designed to combine with 3D offsets to perform 3D instance grouping on voxels.
[0045] This application proposes a space-aware backprojection module. By additionally predicting the possible depth space of objects through a 2D model, the 2D semantic segmentation results are backprojected into the entire 3D space to obtain more complete and space-aware initialized 3D features, so as to optimize the final panoramic reconstruction result.
[0046] And this application has obtained the optimal performance of panoramic reconstruction on both the synthetic dataset 3D-Front and the real-scene dataset Matterport-3D by solving the above two uncertainties. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0048] Figure 1 It is a schematic flowchart of a bottom-up single-image panoramic reconstruction method shown in an exemplary embodiment of this application;
[0049] Figure 2 It is a schematic flowchart of how to convert 2D features into initialized 3D features shown in an exemplary embodiment of this application;
[0050] Figure 3 It is a schematic flowchart of panoramic reconstruction shown in an exemplary embodiment of this application;
[0051] Figure 4 It is a schematic flowchart of obtaining the panoramic reconstruction result shown in an exemplary embodiment of this application;
[0052] Figure 5 It is a schematic flowchart of obtaining the 3D offset reconstruction result shown in an exemplary embodiment of this application;
[0053] Figure 6It is a schematic flow diagram showing how to group instances in an exemplary embodiment of the present application;
[0054] Figure 7 It is a schematic diagram of the framework structure of a bottom-up single-image panoramic reconstruction framework shown in an exemplary embodiment of the present application;
[0055] Figure 8 It is a schematic diagram of the framework structure of a space-aware backprojection module shown in an exemplary embodiment of the present application;
[0056] Figure 9 It is a schematic diagram of the framework structure of a panoramic reconstruction module shown in an exemplary embodiment of the present application;
[0057] Figure 10 It is a schematic diagram of the framework structure of an instance grouping module shown in an exemplary embodiment of the present application;
[0058] Figure 11 It is a schematic diagram of the structure of a bottom-up single-image panoramic reconstruction device shown in an exemplary embodiment of the present application;
[0059] Figure 12 It is a schematic diagram of the structure of a computer device shown in an exemplary embodiment of the present application. Detailed implementation manners
[0060] In order to make the purpose, implementation manners and advantages of the present application clearer and more understandable, the following will clearly and completely describe the exemplary implementation manners of the present application in combination with the drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. It should be understood that the specific embodiments described here are only used to explain the present application, and are not used to limit the present application.
[0061] It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the following described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood in their ordinary and common meanings.
[0062] The terms "first", "second", "third", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such terms used can be interchanged under appropriate circumstances.
[0063] The terms "include" and "have" and any of their variations are intended to cover but not be exclusive of inclusion. For example, a product or device including a series of components does not necessarily have to be limited to all the components clearly listed, but may include other components not clearly listed or inherent to these products or devices.
[0064] Before explaining the bottom-up single-image panoramic reconstruction method provided by the embodiments of the present application, the application scenarios and implementation environments of the embodiments of the present application will be introduced first.
[0065] Single-image panoramic reconstruction mainly studies how to use a single 2D image as input to reconstruct the entire 3D scene while performing foreground instance segmentation and background semantic segmentation on the scene. To achieve this task, sufficient spatial information and semantic information need to be extracted from the single 2D image, so that the initial 3D features obtained by back-projecting the extracted 2D information are more accurate, and then the 3D model can better reconstruct and segment the scene.
[0066] Currently, the solutions for single-image panoramic reconstruction are mainly divided into three stages:
[0067] 1) 2D stage: Use a 2D model to predict 2D segmentation and 2D spatial information (such as depth estimation);
[0068] 2) 2D-3D stage: Use the 2D spatial information to back-project the 2D segmentation information into 3D for initializing 3D features;
[0069] 3) 3D stage: Input the initialized 3D features into a 3D model to reconstruct the scene and predict the panoramic segmentation result of the scene.
[0070] Existing 2D segmentation models usually include two types: top-down (usually an instance segmentation model) and bottom-up (usually capable of panoramic segmentation). The general top-down method first predicts the object category and its bounding box, and then predicts the instance mask within the bounding box to obtain the instance segmentation result; while the typical bottom-up method predicts the center point of each instance and the relative offset of the pixels within the instance to the center point, and at the same time predicts 2D semantic segmentation. In the post-processing stage, the predicted center points and relative offsets are first used for instance grouping, and then the semantic segmentation result is used to classify the instances. Finally, the foreground instance individuals and the semantic background are combined to form the panoramic segmentation result.
[0071] In the 2D-3D stage, the general method obtains the 3D object surface by back-projecting the depth estimated by the 2D model, and then uses the 2D segmentation result to fill the object surface to form the initialized 3D features. However, the 3D features obtained by this solution only exist on the object surface and cannot obtain the information of the occluded area.
[0072] In the 3D stage, a 3D model composed of 3D sparse convolution is used to reconstruct the scene and perform panoramic segmentation on the scene.
[0073] The existing state-of-the-art technology uses a top-down Mask R-CNN for instance segmentation in the 2D stage to obtain instance masks, and at the same time uses monocular depth estimation to predict the surface depth of objects; in the 2D-3D stage, the 2D instance masks are randomly arranged and then back-projected into 3D using depth and camera intrinsics to form 3D features for initializing the object surface; in the 3D stage, the 3D features are input into a 3D model to predict the 3D instance masks corresponding to the arrangement, and at the same time semantic segmentation is predicted for instance classification, and finally the predicted instance individuals and semantic backgrounds are combined to form a panoramic reconstruction result.
[0074] However, the existing technology is only based on a top-down framework of depth estimation, which will cause uncertainties in occluded regions and channel arrangements when back-projecting 2D instance masks to obtain initial 3D features.
[0075] Specifically:
[0076] 1) Only extracting the spatial information of the object surface through depth estimation causes uncertainties in occluded regions, making it difficult to support the reconstruction of the entire scene.
[0077] 2) Using a top-down framework, instance masks with uncertain categories and quantities will be randomly arranged during the back-projection from 2D to 3D, resulting in uncertainties in channel arrangements, thus affecting the segmentation results of the scene. Therefore, the present invention proposes a bottom-up single-image panoramic reconstruction framework (BUOL) to solve the uncertainties in occluded regions and channel arrangements, thereby improving the performance of single-image panoramic reconstruction.
[0078] Based on this, the present application provides a bottom-up single-image panoramic reconstruction method.
[0079] 1) Aiming at the uncertainty of occluded regions caused by only obtaining the object surface information through depth estimation, the present invention additionally predicts the spatial information of occluded regions in the 2D model, so as to obtain more complete 3D features for initialization.
[0080] 2) Aiming at the uncertainty of channel arrangements caused by the random arrangement of instance masks in channels, the present invention adopts a bottom-up framework, back-projects the 2D semantic segmentation results with fixed channels into 3D to obtain initial 3D features, and then groups voxels through the predicted 2D center points and 3D offsets, which not only solves the uncertainty of channel arrangements but also realizes 3D instance segmentation.
[0081] With the help of a space-aware bottom-up framework, the present invention greatly improves the final panoramic reconstruction result.
[0082] Next, the technical solutions of the present application will be specifically described through embodiments in conjunction with the accompanying drawings, as well as how the technical solutions of the present application solve the above technical problems. The embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments.
[0083] Figure 1 It is a schematic flowchart of a bottom-up single-image panoramic reconstruction method shown in an exemplary embodiment of the present application; Figure 7 It is a schematic diagram of the framework structure of a bottom-up single-image panoramic reconstruction framework shown in an exemplary embodiment of the present application.
[0084] In an exemplary embodiment, as Figure 1 shown, a bottom-up single-image panoramic reconstruction method is provided. In this embodiment, the method includes the following steps:
[0085] Step 100: Obtain an image and input the single image into a 2D model.
[0086] Step 200: Predict the depth space where the object may exist, the 2D semantic segmentation result, the object surface depth, and the 2D instance center point according to the 2D model;
[0087] Step 300: Generate initial features of the complete 3D space based on the depth space, the 2D semantic segmentation result, and the object surface depth, based on a spatially aware back-projection module;
[0088] Step 400: Predict the initial features as a 3D reconstruction result, a 3D semantic segmentation result, and a 3D offset result based on a 3D model;
[0089] Step 500: Based on the 3D reconstruction result, the 3D semantic segmentation result, and the 3D offset result, perform instance grouping and synthesis based on the panoramic reconstruction module in combination with the 2D instance center point to obtain the final result of panoramic reconstruction.
[0090] Among them, steps 100 - 200 are the 2D stage, step 300 is the 2D-3D stage, and steps 400 - 500 are the 3D stage. All three parts belong to the bottom-up single-image panoramic reconstruction framework proposed by the present application. This framework is as Figure 7 shown, and its most important part is the final 3D panoramic reconstruction process.
[0091] It can be seen that some embodiments of the present application propose a bottom-up single-image panoramic reconstruction framework, including a bottom-up panoramic reconstruction framework and a spatially aware backprojection module. Among them, the bottom-up framework backprojects the predicted 2D semantic segmentation results into initialized 3D features in a fixed channel, avoiding channel uncertainty caused by random permutations of instance masks, and uses the predicted 2D instance center points and 3D offsets to group each voxel when synthesizing 3D instances; the spatially aware backprojection module utilizes the depth space occupied by the additionally predicted objects and combines the predicted depth to fill the 2D semantic information in the entire 3D space, thereby obtaining initialized 3D features of the complete space, and further optimizing the final panoramic reconstruction result.
[0092] In a possible implementation manner, the spatially aware backprojection module obtains complete spatially aware 3D features according to the depth space where the objects may exist predicted by the 2D model and the object surface depth.
[0093] Figure 2 It is a schematic flow diagram showing how to convert 2D features into initialized 3D features shown in an exemplary embodiment of the present application; Figure 8 It is a schematic framework diagram of the spatially aware backprojection module shown in an exemplary embodiment of the present application.
[0094] In a possible implementation manner, as Figure 2 shown, in the step of generating initialized features of the complete 3D space based on the depth space, 2D semantic segmentation results, and object surface depth through the spatially aware backprojection module, it further includes:
[0095] Step 310: Backproject the depth space where the objects may exist predicted by the 2D model to the 3D space by using the camera intrinsic parameters and the predicted depth;
[0096] Step 320: Backproject the 2D semantic segmentation results and fill the entire 3D space;
[0097] Step 330: Multiply the two after 3D sparse convolution to obtain spatially aware initialized 3D features.
[0098] Among them, the spatially aware backprojection module is as Figure 8 shown.
[0099] In a possible implementation manner, the 2D semantic segmentation results are used for the spatially aware backprojection module to obtain initialized 3D features, and the 2D instance center points are used to group 3D voxels in combination with 3D offsets.
[0100] Among them, both the 2D semantic segmentation result and the 2D instance center point are predicted in the 2D model through a bottom-up panoramic reconstruction framework for bottom-up panoramic segmentation.
[0101] Figure 3 It is a schematic flowchart of panoramic reconstruction shown in an exemplary embodiment of the present application; Figure 9 It is a schematic diagram of the framework structure of the panoramic reconstruction module shown in an exemplary embodiment of the present application.
[0102] In a possible implementation manner, as Figure 3 shown, in the step of performing instance grouping and synthesis based on the panoramic reconstruction module in combination with the 2D instance center point according to the 3D reconstruction result, 3D semantic segmentation result, and 3D offset result to obtain the final result of panoramic reconstruction, it further includes:
[0103] Step 510: Filter the 3D semantic segmentation result through the 3D reconstruction result to obtain a refined 3D semantic reconstruction result;
[0104] Step 520: Filter the 3D offset result through the 3D reconstruction result to obtain a 3D offset reconstruction result;
[0105] In the 3D semantic reconstruction result, each foreground semantic category is input into the instance grouping module to generate an instance in combination with the 2D instance center point, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result.
[0106] Among them, the panoramic reconstruction module is as Figure 9 shown.
[0107] Figure 4 It is a schematic flowchart of obtaining the panoramic reconstruction result shown in an exemplary embodiment of the present application.
[0108] In a possible implementation manner, as Figure 4 shown, in the step that in the 3D semantic reconstruction result, each foreground semantic category is input into the instance grouping module to generate an instance in combination with the 2D instance center point, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result, it further includes:
[0109] Step 531: Obtain the 3D offset reconstruction result and the 2D instance center point of this category according to the foreground semantic category;
[0110] Step 532: Perform grouping based on the 3D offset reconstruction result and the 2D instance center point to obtain a 3D instance segmentation result;
[0111] The 3D instance segmentation result combines the instance segmentation result and the background semantics to obtain the panoramic reconstruction result.
[0112] Figure 5 It is a schematic flow chart of obtaining a 3D offset reconstruction result shown in an exemplary embodiment of the present application.
[0113] In a possible implementation, as Figure 5 shown, in the step of obtaining the 3D offset reconstruction result and the 2D instance center point of this category according to the foreground semantic category, it further includes:
[0114] Step 5311: The 3D offset reconstruction result and the 3D semantic reconstruction result of this category are projected and transformed into a multi-layer depth space;
[0115] Step 5312: Filter the 3D offset reconstruction result through the 3D semantic reconstruction result of this category to obtain the 3D offset reconstruction result of this category;
[0116] Step 5313: Extract the 2D instance center point of this category from all the 2D instance center points for category instance grouping.
[0117] Figure 6 It is a schematic flow chart of how to perform instance grouping shown in an exemplary embodiment of the present application; Figure 10 It is a schematic framework structure diagram of an instance grouping module shown in an exemplary embodiment of the present application.
[0118] In a possible implementation, as Figure 6 shown, in the step of performing grouping based on the 3D offset reconstruction result and the 2D instance center point to obtain a 3D instance segmentation result, it further includes:
[0119] Step 5321: Add the 3D offset reconstruction result of each voxel of this category to the coordinate of this voxel to obtain the predicted 2D instance center point of this voxel;
[0120] Step 5322: According to the distance from the predicted 2D instance center point to all the actual 2D instance center points, assign this voxel to the actual 2D instance center point closest to its predicted 2D instance center point;
[0121] Step 5323: After completing the instance grouping of all voxels of all categories, a 3D instance segmentation result can be obtained.
[0122] Among them, the instance grouping module is as Figure 10 shown.
[0123] It can be seen that some embodiments of the present application use the semantic segmentation results of fixed channels through a bottom-up panoramic reconstruction framework, solve the uncertainty of instance channel arrangement caused by the top-down framework, and then complete instance synthesis by using 2D center points and 3D offsets; a spatial perception module is proposed to solve the uncertainty of occluded areas caused by only predicting depth. The overall solution proposed has achieved better performance in both scene reconstruction and panoramic segmentation.
[0124] It should be understood that although the steps in the flowcharts involved in the above embodiments are displayed in sequence as indicated, these steps are not necessarily executed in the order indicated. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0125] Moreover, the present invention has been experimentally verified, and its effectiveness has been verified on the synthetic dataset 3D-Front and the real-scene dataset Matterport-3D, both of which have achieved the current optimal performance. In the experimental results on 3D-Front and Matterport-3D, the panoramic reconstruction quality (PRQ) is 11.81% and 7.46% higher than the existing optimal solutions respectively, significantly improving the performance indicators of single-image panoramic reconstruction.
[0126] Corresponding to the embodiments of the bottom-up single-image panoramic reconstruction method described above, with the same technical concept, the present application also provides embodiments of a bottom-up single-image panoramic reconstruction device.
[0127] Figure 11 It is a schematic structural diagram of a bottom-up single-image panoramic reconstruction device shown in an exemplary embodiment of the present application.
[0128] In an exemplary embodiment, as Figure 11 shown, the bottom-up single-image panoramic reconstruction device includes:
[0129] A 2D acquisition module 1, configured to acquire an image and input the single image into a 2D model;
[0130] A 2D prediction module 2, configured to predict the depth space where an object may exist, the 2D semantic segmentation result, the surface depth of the object, and the 2D instance center point according to the 2D model;
[0131] A 3D transformation module 3, configured to generate initial features of a complete 3D space based on the depth space, 2D semantic segmentation result, and object surface depth, based on a spatially aware back-projection module;
[0132] A 3D prediction module 4, configured to predict the initial features into a 3D reconstruction result, a 3D semantic segmentation result, and a 3D offset result based on a 3D model;
[0133] A 3D reconstruction module 5, configured to perform instance grouping and synthesis based on the panoramic reconstruction module and the 2D instance center points according to the 3D reconstruction result, 3D semantic segmentation result, and 3D offset result, to obtain the final result of panoramic reconstruction.
[0134] For the specific limitations of the bottom-up single-image panoramic reconstruction device, reference may be made to the limitations of the bottom-up single-image panoramic reconstruction method in the foregoing text, which will not be elaborated herein. Each module in the above bottom-up single-image panoramic reconstruction device may be implemented in whole or in part by software, hardware, and their combination. The above modules may be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0135] In an exemplary embodiment, the above bottom-up single-image panoramic reconstruction method may be applied to Figure 11 the computer device 10 shown in the figure. At this time, the present application can use the computer device to construct a neural network model combining convolution and an attention mechanism. The convolutional layer is used for extracting sequence local features, and the self-attention layer is used for learning the relationships between positions in the global range, which can effectively extract the cleavage site features, and combining the public database and real data improves the prediction ability of the model, realizing accurate prediction of RNA cleavage sites.
[0136] Figure 12 It is a schematic structural diagram of a computer device shown in an exemplary embodiment of the present application.
[0137] In a possible implementation manner, the structure of the computer device is as Figure 12 shown. The computer device 10 includes at least a processor 11, a memory 12, a communication bus 13, and a communication interface 14.
[0138] Among them, the processor 11 may be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or may be one or more integrated circuits for implementing the solution of this application. For example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0139] Optionally, the processor 11 may include one or more CPUs. The computer device 10 may include multiple processors 11. Each of these processors 11 may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU).
[0140] It should be noted that the processor 11 here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0141] The memory 12 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, or may be a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions. It may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0142] Optionally, the memory 12 can exist independently and be connected to the processor 11 through a communication bus 13; the memory 12 can also be integrated with the processor 11.
[0143] The communication bus 13 is used to transfer information between components (such as between the processor and the memory). The communication bus 12 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, Figure 7 only one communication bus is used for illustration here, but it does not mean that there is only one bus or one type of bus.
[0144] The communication interface 14 is used for the computer device 10 to communicate with other devices or communication networks. The communication interface 14 includes a wired communication interface or a wireless communication interface. Among them, the wired communication interface can be, for example, an Ethernet interface. The Ethernet interface can be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface can be a Wireless Local Area Networks (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.
[0145] In some embodiments, the computer device 10 may further include an output device 15 and an input device 16 ( Figure 12 not shown in the figure). The output device 15 communicates with the processor 11 and can display information in various ways. For example, the output device 15 can be a Liquid Crystal Display (LCD), a Light Emitting Diode (LED) display device, a Cathode Ray Tube (CRT) display device, or a projector, etc. The input device 16 communicates with the processor 11 and can receive user input in various ways. For example, the input device 16 can be a mouse, a keyboard, a touch screen device, or a sensing device, etc.
[0146] In some embodiments, the memory 12 is used to store a computer program for executing the solution of this application, and the processor 11 can execute the computer program stored in the memory 12. For example, the computer device 10 can call and execute the computer program stored in the memory 12 through the processor 11 to implement the steps of the bottom-up single-image panoramic reconstruction method provided by the embodiments of this application.
[0147] It should be understood that the bottom-up single-image panoramic reconstruction method provided by this application can be applied to a bottom-up single-image panoramic reconstruction device. The device can be implemented as part or all of the processor 11 through software, hardware, or a combination of software and hardware, and integrated in the computer device 10
[0148] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0149] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A bottom-up single-image panoramic reconstruction method, characterized in that Including: Obtain an image and input a single said image into a 2D model; Predict the possible depth space, 2D semantic segmentation result, object surface depth, and 2D instance center point of the object according to the 2D model; Generate initial features of a complete 3D space based on the depth space, 2D semantic segmentation result, and object surface depth, based on a spatially aware inverse projection module; Predict the initial features into a 3D reconstruction result, 3D semantic segmentation result, and 3D offset result based on a 3D model; Filter the 3D semantic segmentation result through the 3D reconstruction result to obtain a refined 3D semantic reconstruction result; Filter the 3D offset result through the 3D reconstruction result to obtain a 3D offset reconstruction result; In the 3D semantic reconstruction result, each foreground semantic category is input into an instance grouping module to generate instances in combination with the 2D instance center point, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result.
2. The bottom-up single-image panoramic reconstruction method according to claim 1, wherein The spatially aware inverse projection module obtains a complete spatially aware 3D feature according to the depth space and object surface depth.
3. The bottom-up single-image panoramic reconstruction method according to claim 2, wherein, In the step of generating initial features of a complete 3D space based on the depth space, 2D semantic segmentation result, and object surface depth, based on a spatially aware inverse projection module, it further includes: Inverse project the possible depth space of the object predicted by the 2D model to the 3D space using the camera internal parameters and the predicted depth; Inverse project and fill the entire 3D space with the 2D semantic segmentation result; Multiply the two after 3D sparse convolution to obtain spatially aware initial 3D features.
4. The bottom-up single-image panoramic reconstruction method according to claim 1, wherein The 2D semantic segmentation result is used for the spatially aware inverse projection module to obtain initial 3D features, and the 2D instance center point is used for instance grouping of 3D voxels in combination with 3D offsets.
5. The bottom-up single-image panoramic reconstruction method according to claim 1, characterized in that In the step that in the 3D semantic reconstruction result, each foreground semantic category is input into an instance grouping module to generate instances in combination with the 2D instance center point, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result, it further includes: Obtain the 3D offset reconstruction result and 2D instance center point of this category according to the foreground semantic category; Group based on the 3D offset reconstruction result and 2D instance center point to obtain a 3D instance segmentation result; The 3D instance segmentation result combines the instance segmentation result and the background semantics to obtain the panoramic reconstruction result.
6. The bottom-up single-image panoramic reconstruction method according to claim 5, characterized in that In the step of obtaining the 3D offset reconstruction result and 2D instance center point of this category according to the foreground semantic category, it further includes: The 3D offset reconstruction result and the 3D semantic reconstruction result of this category are projected and transformed into a multi-layer depth space; Filter the 3D offset reconstruction result through the 3D semantic reconstruction result of this category to obtain the 3D offset reconstruction result of this category; Extract the 2D instance center point of this category from all the 2D instance center points for category instance grouping.
7. The bottom-up single-image panoramic reconstruction method according to claim 5, characterized in that In the step of grouping based on the 3D offset reconstruction result and 2D instance center point to obtain a 3D instance segmentation result, it further includes: Add the 3D offset reconstruction result of each voxel in the category to the coordinates of the voxel to obtain the predicted 2D instance center point of the voxel; According to the distance from the predicted 2D instance center point to all actual 2D instance center points, assign the voxel to the actual 2D instance center point closest to its predicted 2D instance center point; After completing the instance grouping of all voxels of all categories, the 3D instance segmentation result can be obtained.
8. A bottom-up single-image panoramic reconstruction device, characterized in that Comprising: A 2D acquisition module, configured to acquire an image and input the single image into a 2D model; A 2D prediction module, configured to predict the depth space where an object may exist, the 2D semantic segmentation result, the surface depth of the object, and the 2D instance center point according to the 2D model; A 3D conversion module, configured to generate initial features of a complete 3D space based on the depth space, the 2D semantic segmentation result, and the surface depth of the object, based on a spatially aware back-projection module; A 3D prediction module, configured to predict the initial features into a 3D reconstruction result, a 3D semantic segmentation result, and a 3D offset result based on a 3D model; A 3D reconstruction module, configured to filter the 3D semantic segmentation result through the 3D reconstruction result to obtain a refined 3D semantic reconstruction result; filter the 3D offset result through the 3D reconstruction result to obtain a 3D offset reconstruction result; in the 3D semantic reconstruction result, each foreground semantic category is input into an instance grouping module to generate an instance in combination with the 2D instance center point, and the background semantic category is used for final stitching to obtain the entire panoramic reconstruction result.
9. A computer device, characterized in that, Comprising a memory and a processor, the memory stores a computer program, wherein when the processor calls and executes the computer program from the memory, the steps of the method according to any one of claims 1 to 7 above are implemented.
Citation Information
Patent Citations
Instant positioning and map construction system and method with semantic perception
CN111968129A
Methods and devices for three-dimensional image reconstruction using single-view projection image
US20230097133A1