Method, system, electronic device and storage medium for generating three-dimensional reconstruction model
By using a virtual camera for simulation imaging and error adjustment during the three-dimensional reconstruction model generation process, the problem of low quality of the three-dimensional reconstruction model of consumer-grade sensor equipment is solved, and a higher-precision three-dimensional reconstruction model generation is achieved, which is suitable for VR/AR and other applications.
Patent Information
- Application Number
- CN202111255424.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-10-27
AI Technical Summary
The quality of real-time three-dimensional reconstruction models based on consumer-grade sensor devices is generally not ideal enough and it is difficult to meet the needs of VR/AR applications, mainly due to the errors introduced in the process of sensor measurement accuracy, data matching and model fusion.
By acquiring the target image acquired by the camera, calculating the camera pose, determining the three-dimensional point cloud data, combining the point cloud data to obtain the intermediate reconstruction model, using a virtual camera to simulate and image the intermediate reconstruction model, comparing the target image and simulation imaging results to determine the model error, and adjusting the intermediate reconstruction model according to the error to generate a high-precision three-dimensional reconstruction model.
The generation accuracy of the three-dimensional reconstruction model is improved, and the error in the intermediate reconstruction model is eliminated, making the three-dimensional reconstruction model more suitable for VR/AR applications such as VR.
Smart Images

Figure CN114022639B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of three-dimensional vision technology, and particularly to a method, a system, an electronic device and a storage medium for generating a three-dimensional reconstruction model. Background Art
[0002] Three-dimensional vision is a technology that intersects multiple disciplines such as computer vision and computer graphics, and is one of the core technologies in many application fields such as robotics, autonomous driving, VR / AR, etc. Three-dimensional reconstruction is a technology that uses various sensor devices to scan and calculate a real physical scene and generate a corresponding digital model, and is one of the key issues in three-dimensional vision research. With the progress of sensor and computing technologies, three-dimensional reconstruction has shown some new changes and characteristics, and real-time, dense, and stable three-dimensional reconstruction methods have become the goals pursued by industry insiders.
[0003] Currently, due to the influence of factors such as the measurement error of the sensor itself, the quality of image data, the motion state of the sensor, and the richness of scene features, the quality of the three-dimensional models reconstructed in real time based on consumer-grade sensor devices is generally not ideal, and it is difficult to be applied in VR / AR. The rapid three-dimensional reconstruction process based on consumer-grade devices mainly includes two parts: camera tracking and model fusion. Camera tracking is to estimate the camera pose (position and orientation, six degrees of freedom) data corresponding to each frame of data by finding matching relationships in a series of data collected by the sensor, and then transform the data of different frames into the same world coordinate system for data fusion. Since errors will be introduced in the processes of sensor measurement accuracy, data matching, and model fusion, and the entire three-dimensional reconstruction process can only obtain an approximate solution through iterative optimization methods, the accuracy of the reconstruction results is thus low.
[0004] Therefore, how to improve the generation accuracy of the three-dimensional reconstruction model is a technical problem that needs to be solved by those skilled in the art at present. Summary of the Invention
[0005] The purpose of the present application is to provide a method, a system, an electronic device and a storage medium for generating a three-dimensional reconstruction model, which can improve the generation accuracy of the three-dimensional reconstruction model.
[0006] To solve the above technical problem, the present application provides a method for generating a three-dimensional reconstruction model, and the method for generating the three-dimensional reconstruction model includes:
[0007] Obtain a target image collected by a camera, and calculate the camera pose corresponding to each target image;
[0008] Determine the three-dimensional point cloud data corresponding to the target image according to the camera pose, and merge the three-dimensional point cloud data corresponding to all the target images to obtain an intermediate reconstruction model;
[0009] Performing simulated imaging on the intermediate reconstruction model using a virtual camera;
[0010] Determining the model error of the intermediate reconstruction model by comparing the target image and the simulated imaging result, and adjusting the intermediate reconstruction model using the model error to obtain a three-dimensional reconstruction model.
[0011] Optionally, performing simulated imaging on the intermediate reconstruction model using a virtual camera includes:
[0012] Adjusting the pose of the virtual camera to the camera pose;
[0013] Controlling the virtual camera to project light rays onto the intermediate reconstruction model, and generating the simulated imaging result according to the intersection result of the light rays and the intermediate reconstruction model.
[0014] Optionally, generating the simulated imaging result according to the intersection result of the light rays and the intermediate reconstruction model includes:
[0015] Determining whether the light rays intersect with the intermediate reconstruction model;
[0016] If so, setting the image pixel value at the corresponding position of the light rays in the simulated imaging result to 0;
[0017] If not, determining the target pixel value according to the color information or distance information at the light ray intersection position, and setting the image pixel value at the corresponding position of the light rays in the simulated imaging result to the target pixel value.
[0018] Optionally, determining the model error of the intermediate reconstruction model by comparing the target image and the simulated imaging result includes:
[0019] Calculating the pixel difference of the pixel points at the same position in the target image and the simulated imaging result;
[0020] Determining the model error of the intermediate reconstruction model according to the pixel difference.
[0021] Optionally, after merging the three-dimensional point cloud data corresponding to all the images to obtain an intermediate reconstruction model, it further includes:
[0022] Dividing the intermediate reconstruction model into multiple spatial units according to a preset spatial resolution;
[0023] Setting the confidence level of the spatial unit according to the number of point cloud data fused in the spatial unit;
[0024] Correspondingly, adjusting the intermediate reconstruction model using the model error to obtain a three-dimensional reconstruction model includes:
[0025] Set the pixel points in the simulated imaging result with pixel differences greater than the first threshold as abnormal pixel points according to the model error, and reduce the confidence of the spatial unit where the abnormal pixel points are located;
[0026] Determine whether the confidence of the spatial unit is less than the second threshold;
[0027] If so, set the spatial unit with a confidence less than the second threshold as the spatial unit to be adjusted;
[0028] Adjust the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted, and generate the three-dimensional reconstruction model using the adjusted pixel points.
[0029] Optionally, adjusting the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted includes:
[0030] Obtain the pixel points corresponding to the spatial unit to be adjusted in all the target images and perform weighting to obtain the standard value corresponding to the spatial unit to be adjusted;
[0031] Adjust the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted with the standard value as the optimization target.
[0032] Optionally, after determining whether the confidence of the spatial unit is less than the second threshold, it further includes:
[0033] Restore the confidence of the spatial units with a confidence greater than or equal to the second threshold of the spatial unit to the initial value.
[0034] This application also provides a system for generating a three-dimensional reconstruction model, and the system includes:
[0035] A pose calculation module, configured to obtain the target images collected by the camera and calculate the camera pose corresponding to each of the target images;
[0036] A reconstruction module, configured to determine the three-dimensional point cloud data corresponding to the target images according to the camera pose, and merge the three-dimensional point cloud data corresponding to all the target images to obtain an intermediate reconstruction model;
[0037] A simulated imaging module, configured to perform simulated imaging on the intermediate reconstruction model using a virtual camera;
[0038] A model adjustment module, configured to determine the model error of the intermediate reconstruction model by comparing the target images and the simulated imaging results, and adjust the intermediate reconstruction model using the model error to obtain a three-dimensional reconstruction model.
[0039] The present application also provides a storage medium, on which a computer program is stored, and the steps executed by the above method for generating a three-dimensional reconstruction model are implemented when the computer program runs.
[0040] The present application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the steps executed by the above method for generating a three-dimensional reconstruction model are implemented when the processor calls the computer program in the memory.
[0041] The present application provides a method for generating a three-dimensional reconstruction model, including: acquiring a target image collected by a camera, and calculating the camera pose corresponding to each target image; determining three-dimensional point cloud data corresponding to the target image according to the camera pose, and merging the three-dimensional point cloud data corresponding to all the target images to obtain an intermediate reconstruction model; simulating imaging of the intermediate reconstruction model by using a virtual camera; determining a model error of the intermediate reconstruction model by comparing the target image with the simulated imaging result, and adjusting the intermediate reconstruction model by using the model error to obtain a three-dimensional reconstruction model.
[0042] The present application determines the three-dimensional point cloud data of a target image according to the camera pose when collecting the target image, and obtains three-dimensional point cloud data by merging the three-dimensional point cloud data. An intermediate reconstruction model can be obtained by fusing the three-dimensional point cloud data. The present application uses a virtual camera to simulate imaging of the intermediate reconstruction model, and compares the simulated imaging result with the target image to determine the difference between the intermediate reconstruction model and the actual imaging, that is, to obtain a model error. The three-dimensional reconstruction model obtained by adjusting the intermediate reconstruction model according to the model error can eliminate the error in the intermediate reconstruction model. Therefore, the present application can improve the generation accuracy of the three-dimensional reconstruction model. The present application also provides a system for generating a three-dimensional reconstruction model, a storage medium, and an electronic device, which have the above beneficial effects and will not be elaborated here. Description of the Drawings
[0043] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0044] Figure 1 It is a flowchart of a method for generating a three-dimensional reconstruction model provided by an embodiment of the present application;
[0045] Figure 2 It is a flowchart of a method for three-dimensional reconstruction with visual consistency provided by an embodiment of the present application;
[0046] Figure 3Schematic diagram of the structure of a three-dimensional reconstruction model generation system provided by an embodiment of the present application. Detailed implementation manners
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0048] Please refer to the following Figure 1 , Figure 1 Flowchart of a three-dimensional reconstruction model generation method provided by an embodiment of the present application.
[0049] The specific steps may include:
[0050] S101: Obtain target images collected by a camera, and calculate the camera pose corresponding to each of the target images;
[0051] Among them, this embodiment can be applied to a three-dimensional model reconstruction device, and the camera can capture target images of a specific object or scene at multiple poses. Based on the obtained target images collected by the camera, this embodiment can perform a camera pose estimation operation on the target images; specifically, after the camera captures the target images, the camera pose corresponding to each target image can be calculated. The above camera can be a consumer-grade sensor device, such as an ordinary camera, a depth camera, etc.
[0052] For each captured target image, the camera pose corresponding to the target image can be estimated. The estimation of the camera pose can be achieved by using optimization based on geometric constraints or by means of a method for constructing an optimization problem based on learning. The objective function of the optimization can be constructed using various geometric and photometric constraints, including but not limited to reprojection error, photometric consistency, etc. According to the different types of sensors, the construction of the optimization problem can be based on color images, depth images alone, or both color images and depth images.
[0053] S102: Determine the three-dimensional point cloud data corresponding to the target images according to the camera poses, and merge the three-dimensional point cloud data corresponding to all the target images to obtain an intermediate reconstruction model;
[0054] Among them, based on the obtained camera poses, the three-dimensional point cloud data can be obtained by restoring the target images according to the camera poses, and the three-dimensional point cloud data corresponding to multiple target images can be merged to obtain an intermediate reconstruction model.
[0055] S103: Simulate imaging of the intermediate reconstruction model using a virtual camera;
[0056] Among them, based on obtaining the intermediate reconstruction model, simulation imaging can be performed on the intermediate reconstruction model to evaluate the model quality of the intermediate reconstruction model using the simulation imaging results. Specifically, in this embodiment, the pose of the virtual camera can be determined according to the camera pose corresponding to each target image, so that the virtual camera performs simulation imaging on the intermediate reconstruction model at the corresponding camera pose, and the simulation imaging result is the image obtained by the virtual camera taking a picture of the intermediate reconstruction model at the camera pose.
[0057] S104: Determine the model error of the intermediate reconstruction model by comparing the target image and the simulation imaging result, and use the model error to adjust the intermediate reconstruction model to obtain a three-dimensional reconstruction model.
[0058] Among them, in this embodiment, the target image actually captured by the camera is compared with the virtual imaging result of the virtual camera to obtain the difference between the intermediate reconstruction model and the actual object, that is, the model error, so as to be used as a measure of the model reconstruction quality. Based on obtaining the model error, in this embodiment, the intermediate reconstruction model is adjusted using the model error to obtain a three-dimensional reconstruction model.
[0059] In this embodiment, the three-dimensional point cloud data of the target image is determined according to the camera pose when collecting the target image, and the three-dimensional point cloud data is obtained by merging the three-dimensional point cloud data. The intermediate reconstruction model can be obtained by fusing the three-dimensional point cloud data. In this embodiment, a virtual camera is used to perform simulation imaging on the intermediate reconstruction model, and the simulation imaging result is compared with the target image to determine the difference between the intermediate reconstruction model and the actual imaging, that is, the model error. The three-dimensional reconstruction model obtained by adjusting the intermediate reconstruction model according to the model error can eliminate the error in the intermediate reconstruction model. Therefore, this embodiment can improve the generation accuracy of the three-dimensional reconstruction model.
[0060] As a further introduction to the Figure 1 corresponding embodiment, in this embodiment, the intermediate reconstruction model can be simulated and imaged in the following manner: adjust the pose of the virtual camera to the camera pose; control the virtual camera to project light onto the intermediate reconstruction model, and generate the simulation imaging result according to the intersection result of the light and the intermediate reconstruction model.
[0061] Each ray projected by the virtual camera corresponds to a pixel point in the simulated imaging result. Therefore, after projecting rays onto the intermediate reconstruction model, in this embodiment, it can be determined whether the ray intersects the intermediate reconstruction model; if so, the image pixel value at the corresponding position of the ray in the simulated imaging result is set to 0; if not, the target pixel value is determined according to the color information or distance information at the ray intersection position, and the image pixel value at the corresponding position of the ray in the simulated imaging result is set to the target pixel value.
[0062] Specifically, after the point cloud data of a frame of target image is fused into the intermediate reconstruction model, according to the estimated camera pose of this frame of target image, the fused model is simulated for imaging, which is used for subsequent evaluation of the estimated camera pose and the accuracy of the reconstruction model. It mainly includes the following operations:
[0063] The camera pose corresponding to the estimated frame of target image is used as the pose of the virtual camera, and the virtual camera is used to simulate the imaging of the reconstruction model. The internal and external camera parameters such as the imaging resolution are determined according to the parameters of the actual camera, and the two are kept consistent.
[0064] The simulated imaging can be carried out by methods similar to ray projection. Specifically, starting from the camera pose, rays can be projected towards each pixel of the target image of the simulated imaging respectively. The rays may or may not intersect the reconstruction model. The image pixel values corresponding to the rays that do not intersect the model are set to 0 or other default values, and the image pixel values corresponding to the rays that intersect the model are determined by the information at the intersection position of the model and the rays. Each pixel can correspond to one projected ray, and in this case, the value of the pixel is determined by this ray. Each pixel can also project multiple rays, and in this case, the value of the pixel is jointly determined by these multiple rays in a weighted manner, and the position distribution of each ray within the pixel size range can be random sampling, the four corner positions, or other methods.
[0065] The simulated imaging result can be a color map, a depth map, or can also include both a color map and a depth map. When generating a color map, the value of each pixel is determined by the model color value at the intersection position of the model and the ray. When generating a depth map, the value of each pixel is the distance between the camera and the ray intersection of the model. This distance can be the direct distance between two points, or the projection distance of the direct distance between two points on the camera optical axis direction. This distance can be a dimensional value with a certain measurement unit, or a dimensionless value after a certain normalization.
[0066] As for Figure 1Further introduction to the corresponding embodiment, the model error of the intermediate reconstructed model is determined by comparing the target image and the simulated imaging result, including: calculating the pixel difference between the pixel points at the same position in the target image and the simulated imaging result; and determining the model error of the intermediate reconstructed model based on the pixel difference.
[0067] Furthermore, after merging the three-dimensional point cloud data corresponding to all the images to obtain an intermediate reconstructed model, the intermediate reconstructed model can be divided into multiple spatial units according to a preset spatial resolution; and the confidence of the spatial unit can be set according to the number of point cloud data fused in the spatial unit.
[0068] Specifically, the target image can be restored to three-dimensional point cloud data according to the calculated camera pose, and the three-dimensional point cloud data corresponding to different target images can be fused (transformed into the same coordinate system and the points corresponding to the same position are merged) to obtain an intermediate reconstructed model. The intermediate reconstructed model adopts a spatial occupancy representation method, and the space where the intermediate reconstructed model is located is divided into spatial units of the same size according to a certain resolution. The efficiency of spatial representation can be improved with the help of algorithms such as voxel hashing. Each spatial unit stores the surface points of the reconstructed object located in the unit, or stores the signed closest distance between the unit and the reconstructed surface. Specifically, it includes the following operations:
[0069] Operation 1: Transform the image into 3D point cloud data according to the camera pose.
[0070] Operation 2: Divide the space occupied by the newly reconstructed intermediate reconstruction model according to certain spatial resolution requirements.
[0071] If the 3D point cloud reconstructed from the new image exceeds the representation range of all the cells that have been segmented in the space, the newly added spatial position will be segmented according to the predetermined resolution based on the original segmentation. The segmentation resolution is determined based on the accuracy of the reconstructed model representation and the available storage resources. Each spatial unit is given an initial confidence value when segmenting. The confidence value represents the accuracy of the reconstructed point located in the spatial unit and is used to evaluate the quality of the subsequent reconstructed model.
[0072] Operation 3: Fuse the newly reconstructed 3D point cloud into the corresponding spatial unit.
[0073] Each point in the reconstructed point cloud is mapped to a certain spatial unit according to its spatial position, and each spatial unit stores the coordinates of the reconstructed points within that unit. The coordinates of the newly added point (or the signed closest distance of the unit from the reconstructed surface, determined according to the data stored in the unit) and the coordinates (or distances) of the existing points within the unit are weighted and averaged. The weight size can be determined according to the number of points that have been fused. For example, if n points have been fused within the unit, the weight of the unit is n, the weight of the newly added point is 1, and the weight of the unit becomes n + 1 after fusion. The above process of point fusion can also be implemented using other strategies. During the model fusion process, the confidence of each unit can remain unchanged or can change according to certain rules, and the change in confidence should be able to reflect the change in the model reconstruction quality.
[0074] As for Figure 1 For a further introduction to the corresponding embodiment, the three-dimensional reconstruction model can be obtained by adjusting the intermediate reconstruction model in the following way: According to the model error, the pixel points in the simulated imaging result with pixel differences greater than the first threshold are set as abnormal pixel points, and the confidence of the spatial unit where the abnormal pixel points are located is reduced; it is judged whether the confidence of the spatial unit is less than the second threshold; if so, the spatial unit with a confidence less than the second threshold is set as the spatial unit to be adjusted; the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted are adjusted, and the three-dimensional reconstruction model is generated using the adjusted pixel points.
[0075] Specifically, in this embodiment, all the pixel points in the target images corresponding to the spatial unit to be adjusted can be weighted to obtain the standard value corresponding to the spatial unit to be adjusted; the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted are adjusted with the standard value as the optimization target. After judging whether the confidence of the spatial unit is less than the second threshold, the confidence of the spatial unit with a confidence greater than or equal to the second threshold can also be restored to the initial value.
[0076] The above process can compare the image obtained by simulated imaging with the image actually captured by the camera, and the difference between the reconstruction model and the actual object can be obtained, which is used as a measure of the model reconstruction quality. It mainly includes the following operations:
[0077] According to the difference between the corresponding pixel values of the actually captured image and the simulated imaging image, the confidence of the spatial unit is calibrated. If the pixel values differ by more than a predetermined threshold, it is considered that the confidence of the spatial unit corresponding to the pixel is poor, and the confidence of the corresponding spatial unit is reduced according to certain rules. If the difference between the pixel values is less than the predetermined threshold, the confidence of the spatial unit corresponding to the pixel remains unchanged.
[0078] According to the set calibration times, the spatial units are evaluated and adjusted and optimized in segments. During the reconstruction process of each frame of the image, all spatial units visible in the image are calibrated once. Each spatial unit maintains a counter of the calibration times. When the calibration reaches a certain number of times, the spatial unit is evaluated once. If the confidence of the spatial unit is less than a predetermined threshold at this time, it is considered that the spatial unit needs to be adjusted and optimized. If the confidence of the spatial unit is greater than the above-mentioned predetermined threshold at this time, it is considered that the spatial unit does not need to be adjusted and optimized, and the confidence of the spatial unit is restored to the initialized value. The purpose of restoring the initialization is to minimize the influence of accumulated errors in the camera pose estimation and reconstruction process.
[0079] The process of adjusting the intermediate reconstruction model to generate the three-dimensional reconstruction model is as follows: some spatial units in the reconstruction model are adjusted and optimized, with the purpose of making the corresponding values of the spatial units in the simulated imaging image and the actual acquired image consistent. The adjustment process uses the actual value in the acquired image as the accurate value, and the specific adjustment scheme is determined according to the meaning of the data represented by the spatial unit or the content of the stored data.
[0080] The main operations include:
[0081] Operation 1: Calculate and obtain the spatial unit adjustment optimization target based on the collected image.
[0082] For a spatial unit that needs to be adjusted, first obtain the pixel values of the spatial unit corresponding to all the acquired images involved in the calibration of the unit in step 4, then restore the pixels to three-dimensional space points according to the camera posture and other data corresponding to each acquired image, and weight each point according to certain rules to obtain the accurate value corresponding to the spatial unit as the target of adjustment and optimization.
[0083] Operation 2: If the accurate value of the spatial unit position obtained in Operation 1 does not exceed the range of the current spatial unit, the data in the spatial unit to be adjusted in the current reconstructed model is updated according to the accurate value.
[0084] Operation 3: If the accurate value of the spatial unit position obtained in operation 1 exceeds the range of the current spatial unit, the data in the spatial unit to be adjusted in the current reconstructed model and the data in the spatial unit corresponding to the accurate value are weighted fused. The confidence in the spatial unit corresponding to the accurate value can remain unchanged or be adjusted according to the preset strategy.
[0085] The process described in the above embodiment is explained below through an embodiment in actual application.
[0086] The types of sensors used in different application scenarios are different, and the data formats collected are also different. Therefore, 3D reconstruction presents different technical architectures and implementation methods in different applications. In VR / AR, it is crucial to construct a virtual environment / object that is highly consistent with the real world, which will directly affect the VR / AR application experience. Real-time, convenient, and automated 3D reconstruction based on real scenarios / objects is an important way to achieve this goal and is also one of the research hotspots in this field. High-precision 3D reconstruction technologies represented by volumetric photography can achieve very high-quality 3D reconstruction with the support of complex sensor devices and sufficient computing power, and have been well applied in many professional fields. However, due to limitations such as the site, cost, and ease of use of complex sensor devices, this method cannot be widely promoted, and fast and high-quality 3D reconstruction based on consumer-grade sensor devices is an urgent need for VR / AR applications.
[0087] Existing 3D model reconstruction schemes mainly use depth maps and color maps to construct joint optimization objectives, construct multi-dimensional constraints by means of information such as shadows / contours, and use a layer-by-layer optimization strategy from coarse to fine to improve the accuracy of camera pose estimation, thereby indirectly improving the model reconstruction accuracy. Although the quality of the reconstructed models obtained by these methods has been greatly improved, it is still difficult to meet the actual application requirements. The current reconstruction process cannot effectively measure the reconstruction result error in real time during the reconstruction process. Generally, after the reconstruction is completed, the quality of the reconstructed model is indirectly measured by estimating the camera trajectory offset. At the same time, for the problem of error accumulation during the reconstruction process, the current methods generally correct the camera pose by performing global optimization during the reconstruction process to indirectly improve the quality of the reconstructed model.
[0088] Considering that 3D reconstruction is a process of gradual accumulation and fusion, this embodiment starts from the reconstruction result in reverse. According to the requirement that the reconstructed model and the actual target object should be visually consistent, by simulating the imaging of the intermediate results obtained during the reconstruction process, the difference between the simulation imaging and the actual camera imaging is found, and based on this difference, the reconstruction result is optimized specifically, so that the geometric details of the model during the reconstruction process can be better retained and restored, thereby directly improving the quality of 3D reconstruction.
[0089] In view of the problem that the quality of real-time 3D reconstruction models based on consumer-grade devices in the field of 3D vision is not high, this embodiment provides a 3D reconstruction method that meets visual consistency, which is used to achieve fast and convenient 3D reconstruction with high geometric quality. In view of the problem that the quality of the model is difficult to measure during the reconstruction process, a method for simulating imaging of the reconstruction results is provided, which can obtain a simulated image of the reconstruction results, so as to measure the quality of the reconstructed model from the perspective of visual consistency; in view of the problems such as the existence of noise in the images used for reconstruction and the cumulative errors caused by the reconstruction algorithm, a model credibility measurement method based on the segmented accumulation of image sequences is provided, so as to achieve accurate evaluation of the difference between the reconstructed model and the real object; in view of the problem of error correction during the reconstruction process, a method for optimizing and adjusting the reconstruction model based on the simulated imaging results is provided, so that the reconstruction results and the real object meet visual consistency.
[0090] This embodiment is based on consumer-grade sensor devices, such as ordinary cameras, depth cameras, etc., and uses the collected image data to perform three-dimensional reconstruction. By simulating imaging of intermediate reconstruction results, the geometric quality of the reconstructed model is evaluated from the perspective of visual consistency. An image sequence composed of multiple images is used to provide a credibility model for measuring the reconstruction results. The high-error parts of the reconstructed model are obtained by comparing with the images directly measured by the camera, and these high-error positions are optimized and adjusted in a targeted manner to improve the model accuracy.
[0091] See also Figure 2 , Figure 2 A flowchart of a visually consistent 3D reconstruction method provided in an embodiment of the present application may include the following steps:
[0092] Step 1: The camera captures images.
[0093] Step 2: Camera pose estimation.
[0094] Step 3: Point cloud recovery and model fusion.
[0095] According to the camera pose estimated in step 2, the image can be restored to a 3D point cloud, and the point cloud data corresponding to different images can be fused (transformed into the same coordinate system and the points corresponding to the same position are merged). The reconstructed model uses a spatial occupancy representation method to divide the space where the reconstructed model is located according to a certain resolution and divide it into spatial units of the same size. The efficiency of spatial representation can be improved with the help of algorithms such as voxel hashing. Each unit stores the surface points of the reconstructed object located within the unit, or stores the signed closest distance between the unit and the reconstructed surface.
[0096] Step 4: Simulate imaging of the reconstruction results.
[0097] After the point cloud of a certain frame is fused into the model, according to the camera pose of this frame estimated in Step 2, simulate imaging on the fused model, which is used to evaluate the estimated camera pose and the accuracy of the reconstructed model subsequently.
[0098] Step 5: Error metric of the reconstructed model based on the simulated imaging result.
[0099] Compare the image obtained from the simulated imaging in Step 4 with the image actually captured by the camera, and the difference between the reconstructed model and the actual object can be obtained, which is used as a measure of the model reconstruction quality.
[0100] Step 6: Adjustment and optimization of the reconstruction result based on the error metric.
[0101] According to the evaluation in Step 5, adjust and optimize some spatial units in the reconstructed model, aiming to make the corresponding values of the spatial unit in the simulated imaging image and the actually captured image consistent. The actual value in the captured image is used as the accurate value during the adjustment process, and the specific adjustment scheme is determined according to the data meaning or the stored data content represented by the spatial unit.
[0102] Step 7: Repeat the operations of Step 1 to Step 6 for each frame of data collected until the camera scanning is completed.
[0103] In this embodiment, simulated imaging can be performed on the intermediate reconstruction result to measure the quality of the reconstructed model and guide the optimization adjustment of the reconstruction result, so that the reconstruction result and the actual object are visually consistent. Each of the above spatial units is given an initial confidence value during meshing for subsequent evaluation of the quality of the reconstructed model.
[0104] In this embodiment, by comparing the image obtained from the simulated imaging of the reconstruction result with the image actually captured by the camera, the difference between the reconstructed model and the actual object can be obtained, which is used as a measure of the model reconstruction quality.
[0105] In this embodiment, the confidence of the spatial unit can be calibrated according to the corresponding pixel positions of the actually captured image and the simulated imaging image. If the pixel values differ by more than a predetermined threshold, it is considered that the confidence of the spatial unit corresponding to the pixel is poor, and the confidence of the corresponding spatial unit is reduced according to certain rules. If the difference between the pixel values is less than the predetermined threshold, the confidence of the spatial unit corresponding to the pixel remains unchanged.
[0106] The reconstruction process of each frame of the image will calibrate all the spatial units visible in the image. Each spatial unit maintains a counter of the number of calibrations. When the number of calibrations reaches a certain value, the spatial unit is evaluated. If the confidence of the spatial unit is less than a predetermined threshold at this time, it is considered that the spatial unit needs to be adjusted and optimized. If the confidence of the spatial unit is greater than a predetermined threshold at this time, it is considered that the control unit does not need to be adjusted and optimized, and the confidence of the spatial unit is restored to the initialized value. The purpose of restoring the initialization is to minimize the influence of accumulated errors in the camera pose estimation and reconstruction process.
[0107] This embodiment can determine the adjustment method for the spatial unit that needs to be adjusted and optimized in the reconstructed model according to the data meaning represented by the spatial unit or the stored data content. The overall purpose is to make the corresponding value of the spatial unit in the simulated imaging and the actual collected image consistent. During the adjustment process, the value in the collected image is used as the accurate value for adjustment.
[0108] For a spatial unit that needs to be adjusted, first obtain the pixel values corresponding to the spatial unit in all the acquired images for calibrating the unit, then restore the pixels to points in three-dimensional space according to their corresponding camera postures and other data, and then weight each point according to certain rules to obtain an accurate value corresponding to the spatial unit as the target of adjustment and optimization.
[0109] If the accurate value of the adjusted spatial unit position does not exceed the range of the current spatial unit, the data in the current spatial unit will be updated according to the accurate value; if the accurate value of the adjusted spatial unit position exceeds the range of the current spatial unit, the data in the current spatial unit and the data in the spatial unit corresponding to the accurate value will be weightedly fused, and the confidence in the spatial unit corresponding to the accurate value will remain unchanged.
[0110] This embodiment can address the problem of low quality of real-time 3D reconstruction models based on consumer-grade devices in the field of 3D vision, and provide a 3D reconstruction method that meets visual consistency, which can achieve high-quality and fast 3D reconstruction of physical scenes. A proposed method for simulating imaging of reconstruction results can obtain simulated images of reconstruction results, measure the geometric quality of the reconstructed model from the perspective of visual consistency, and solve the problem that the quality of the model is difficult to measure effectively during the reconstruction process; a proposed method for reconstructed model confidence based on segmented accumulation of image sequences can achieve accurate evaluation of the difference between the reconstructed model and the real object, and solve the problems of inaccurate evaluation results caused by image noise and algorithm cumulative errors; a proposed method for optimizing and adjusting the reconstruction model based on simulated imaging results can make the reconstruction result and the real object meet visual consistency, and solve the problem of error correction during the reconstruction process.
[0111] See alsoFigure 3 , Figure 3 It is a schematic structural diagram of a system for generating a three-dimensional reconstruction model provided by an embodiment of the present application. The system may include:
[0112] A pose calculation module 301, configured to obtain a target image collected by a camera and calculate the camera pose corresponding to each target image;
[0113] A reconstruction module 302, configured to determine three-dimensional point cloud data corresponding to the target image according to the camera pose, and merge all the three-dimensional point cloud data corresponding to the target images to obtain an intermediate reconstruction model;
[0114] A simulated imaging module 303, configured to perform simulated imaging on the intermediate reconstruction model by using a virtual camera;
[0115] A model adjustment module 304, configured to determine a model error of the intermediate reconstruction model by comparing the target image and the simulated imaging result, and use the model error to adjust the intermediate reconstruction model to obtain a three-dimensional reconstruction model.
[0116] In this embodiment, the three-dimensional point cloud data of the target image is determined according to the camera pose when collecting the target image, and the three-dimensional point cloud data is obtained by merging the three-dimensional point cloud data. An intermediate reconstruction model can be obtained by fusing the three-dimensional point cloud data. In this embodiment, a virtual camera is used to perform simulated imaging on the intermediate reconstruction model, and the simulated imaging result is compared with the target image to determine the difference between the intermediate reconstruction model and the actual imaging, that is, the model error. The three-dimensional reconstruction model obtained by adjusting the intermediate reconstruction model according to the model error can eliminate the error in the intermediate reconstruction model. Therefore, this embodiment can improve the generation accuracy of the three-dimensional reconstruction model.
[0117] Further, the simulated imaging module 303 includes:
[0118] A pose adjustment unit, configured to adjust the pose of the virtual camera to the camera pose;
[0119] An imaging unit, configured to control the virtual camera to project light rays onto the intermediate reconstruction model, and generate the simulated imaging result according to the intersection result of the light rays and the intermediate reconstruction model.
[0120] Further, the imaging unit is configured to determine whether the light ray intersects with the intermediate reconstruction model; if so, set the image pixel value at the corresponding position of the light ray in the simulated imaging result to 0; if not, determine the target pixel value according to the color information or distance information at the light ray intersection position, and set the image pixel value at the corresponding position of the light ray in the simulated imaging result to the target pixel value.
[0121] Further, the model adjustment module 304 is configured to calculate the pixel differences between the pixel points at the same positions in the target image and the simulated imaging result; and is further configured to determine the model error of the intermediate reconstruction model according to the pixel differences.
[0122] Further, it further includes:
[0123] The model division module is configured to divide the intermediate reconstruction model into multiple spatial units according to a preset spatial resolution after merging the three-dimensional point cloud data corresponding to all the images; and is further configured to set the confidence level of the spatial unit according to the number of point cloud data fused in the spatial unit.
[0124] Correspondingly, the model adjustment module 304 is configured to set the pixel points in the simulated imaging result with pixel differences greater than a first threshold as abnormal pixel points according to the model error, and reduce the confidence level of the spatial unit where the abnormal pixel points are located; and is further configured to determine whether the confidence level of the spatial unit is less than a second threshold; if so, set the spatial unit with a confidence level less than the second threshold as a spatial unit to be adjusted; and is further configured to adjust the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted, and generate the three-dimensional reconstruction model by using the adjusted pixel points.
[0125] Further, the adjustment of the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted by the model adjustment module 304 includes: obtaining the pixel points corresponding to the spatial unit to be adjusted in all the target images and performing weighting to obtain the standard value corresponding to the spatial unit to be adjusted; and adjusting the pixel points in the intermediate reconstruction model corresponding to the spatial unit to be adjusted with the standard value as the optimization target.
[0126] Further, it further includes:
[0127] The confidence restoration module is configured to restore the confidence levels of the spatial units with confidence levels greater than or equal to the second threshold to the initial values after determining whether the confidence level of the spatial unit is less than the second threshold.
[0128] This embodiment provides a three-dimensional reconstruction solution that conforms to visual consistency. Through the provided quality evaluation and adjustment optimization methods for the reconstruction results, fast three-dimensional reconstruction with high geometric quality based on consumer-grade devices is achieved. A simulated imaging solution for the reconstruction results is provided, which can obtain the simulated images of the reconstruction results and measure the quality of the reconstruction model from the perspective of visual consistency. A model confidence measurement solution based on the accumulation of image segmentation sequences is provided, which can solve the influence of image noise and algorithm cumulative errors on quality evaluation and accurately evaluate the difference between the reconstruction model and the real object. A reconstruction model optimization and adjustment solution based on confidence is provided to solve the problem of error correction during the reconstruction process, so that the reconstruction results and the real object meet visual consistency.
[0129] Since the embodiments in the system part correspond to the embodiments in the method part, please refer to the description of the embodiments in the method part for the embodiments in the system part, and will not be elaborated here temporarily.
[0130] This application also provides a storage medium on which a computer program is stored. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0131] This application also provides an electronic device, which may include a memory and a processor. When the processor calls the computer program stored in the memory, the steps provided in the above embodiments can be implemented. Of course, the electronic device may also include various network interfaces, power supplies and other components.
[0132] The embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
[0133] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
Claims
1. A method for generating a three-dimensional reconstruction model, characterized in that, comprising: Obtaining target images collected by a camera and calculating the camera pose corresponding to each of the target images; Determining the three-dimensional point cloud data corresponding to the target images according to the camera poses, and merging the three-dimensional point cloud data corresponding to all the target images to obtain an intermediate reconstruction model; Simulating imaging of the intermediate reconstruction model using a virtual camera; Determining the model error of the intermediate reconstruction model by comparing the target images with the simulated imaging results, and adjusting the intermediate reconstruction model using the model error to obtain a three-dimensional reconstruction model; After merging the three-dimensional point cloud data corresponding to all the images to obtain an intermediate reconstruction model, it further includes: Dividing the intermediate reconstruction model into multiple spatial units according to a preset spatial resolution; Setting the confidence level of the spatial unit according to the number of point cloud data fused in the spatial unit; Correspondingly, adjusting the intermediate reconstruction model using the model error to obtain a three-dimensional reconstruction model, including: Setting the pixel points in the simulated imaging result with pixel differences greater than a first threshold as abnormal pixel points according to the model error, and reducing the confidence level of the spatial unit where the abnormal pixel points are located; Judging whether the confidence level of the spatial unit is less than a second threshold; If so, setting the spatial unit with a confidence level less than the second threshold as a spatial unit to be adjusted; Adjusting the pixel points corresponding to the spatial unit to be adjusted in the intermediate reconstruction model, and generating the three-dimensional reconstruction model using the adjusted pixel points; Adjusting the pixel points corresponding to the spatial unit to be adjusted in the intermediate reconstruction model, including: Obtaining and weighting the pixel points corresponding to the spatial unit to be adjusted in all the target images to obtain a standard value corresponding to the spatial unit to be adjusted; Adjusting the pixel points corresponding to the spatial unit to be adjusted in the intermediate reconstruction model with the standard value as an optimization target.
2. The method for generating a three-dimensional reconstruction model according to claim 1, characterized in that, The simulating imaging of the intermediate reconstruction model using a virtual camera includes: Adjusting the pose of the virtual camera to the camera pose; Controlling the virtual camera to project light rays onto the intermediate reconstruction model, and generating the simulated imaging result according to the intersection result of the light rays and the intermediate reconstruction model.
3. The method for generating a three-dimensional reconstruction model according to claim 2, characterized in that, Generating the simulated imaging result according to the intersection result of the light rays and the intermediate reconstruction model, including: Judging whether the light rays intersect with the intermediate reconstruction model; If not, setting the image pixel value at the corresponding position of the light rays in the simulated imaging result to 0; If so, determining a target pixel value according to the color information or distance information at the light ray intersection position, and setting the image pixel value at the corresponding position of the light rays in the simulated imaging result to the target pixel value.
4. The method for generating a three-dimensional reconstruction model according to claim 1, characterized in that, Determining the model error of the intermediate reconstruction model by comparing the target image and the simulated imaging result includes: Calculating the pixel difference of the pixel points at the same position in the target image and the simulated imaging result; Determining the model error of the intermediate reconstruction model according to the pixel difference.
5. The method for generating a three-dimensional reconstruction model according to claim 1, wherein, after determining whether the confidence level of the spatial unit is less than the second threshold, it further includes: Restoring the confidence levels of the spatial units whose confidence levels are greater than or equal to the second threshold to the initial values.
6. A system for generating a three-dimensional reconstruction model, wherein, configured to implement the steps of the method for generating a three-dimensional reconstruction model according to claim 1, including: A pose calculation module, configured to obtain a target image collected by a camera and calculate the camera pose corresponding to each target image; A reconstruction module, configured to determine the three-dimensional point cloud data corresponding to the target image according to the camera pose, and merge the three-dimensional point cloud data corresponding to all the target images to obtain an intermediate reconstruction model; A simulated imaging module, configured to perform simulated imaging on the intermediate reconstruction model by using a virtual camera; A model adjustment module, configured to determine the model error of the intermediate reconstruction model by comparing the target image and the simulated imaging result, and adjust the intermediate reconstruction model by using the model error to obtain a three-dimensional reconstruction model.
7. An electronic device, wherein, it includes a memory and a processor, a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the method for generating a three-dimensional reconstruction model according to any one of claims 1 to 5 are implemented.
8. A storage medium, wherein, computer-executable instructions are stored in the storage medium, and when the computer-executable instructions are loaded and executed by a processor, the steps of the method for generating a three-dimensional reconstruction model according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Computer implemented method and device for performing three-dimensional blood vessel reconstruction by using contrast image
CN109448072A
RGB-D image-based indoor scene three-dimensional reconstruction method
CN109658449A