Camera pose correction method, three-dimensional reconstruction method and related device
By acquiring common-view images and loopback images, using pose map optimization targets and constraints to correct camera poses, the problem of low camera pose accuracy in 3D reconstruction is solved, and the accuracy of 3D reconstruction is improved.
Patent Information
- Application Number
- CN202510387625.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-15
AI Technical Summary
In the existing three-dimensional reconstruction technology, due to factors such as noise interference, feature extraction errors and sensor accuracy limitations, the camera position accuracy is low, which in turn affects the accuracy of the three-dimensional reconstruction results.
By acquiring the common-view image and loopback images from the image dataset, the camera position is determined based on the common-view relationship and similarity conditions, and the camera position is corrected using the optimization target and constraint terms of the pose map to eliminate cumulative errors and improve the accuracy of the camera position.
While maintaining the consistency of relative poses between images, cumulative errors are eliminated and the accuracy and stability of the three-dimensional reconstruction results are improved.
Smart Images

Figure CN120495092A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a camera posture correction method, a three-dimensional reconstruction method, and related devices. Background Art
[0002] 3D reconstruction is the process of recovering the 3D structure of a scene or object from 2D images or other measurement data. Through 3D reconstruction, point cloud models of the corresponding scene or object can be generated, which are very useful in practical applications such as visual relocalization.
[0003] 3D reconstruction is a time-series process. However, when using existing technical solutions for 3D reconstruction, cumulative errors are often introduced due to various factors, such as noise interference, feature extraction errors, insufficient optimization convergence, and sensor accuracy limitations. This error significantly reduces the accuracy of the point cloud model obtained through 3D reconstruction, adversely affecting the practical application of the 3D reconstruction results. Therefore, for those skilled in the art, improving the accuracy of 3D reconstruction results remains a key technical issue in the field of 3D reconstruction. Summary of the Invention
[0004] The present application provides a camera posture correction method, a three-dimensional reconstruction method and related devices to solve the problem of low accuracy of camera posture in the prior art and achieve the purpose of improving the accuracy of camera posture.
[0005] In a first aspect, the present application provides a method for correcting a camera pose, comprising:
[0006] Acquire, from an image dataset, an image that satisfies a preset co-viewing condition with the current image as a co-viewing image, wherein the image dataset includes one or more two-dimensional images captured at the same target location as the current image;
[0007] Acquire a relative pose between a camera that captures the current image and a camera that captures the co-viewing image as a co-viewing relative pose;
[0008] Acquire an image from the image dataset that satisfies a preset similarity condition with the current image, and if the image is not the same as the co-viewed image, use the image as a loop closure image;
[0009] Obtaining a relative pose between a camera that captures the current image and a camera that captures the loop-closed image as a loop-closed relative pose;
[0010] With the optimization goal of narrowing the gap between the loop relative pose and the common view relative pose, and with the constraint of minimizing the relative pose change of the edges of the pose graph, the camera pose corresponding to each node in the pose graph is corrected to eliminate the cumulative error of the camera pose, wherein the pose graph includes: nodes and edges, a node corresponds to an image in the image dataset and the node records the camera pose of the camera that took the image, an edge connects two nodes, there is a common view relationship between the images corresponding to the two nodes, and the edge records the relative pose of the cameras recorded by the two nodes.
[0011] In a second aspect, the present application provides a three-dimensional reconstruction method, comprising:
[0012] Optimizing the camera pose of the image captured at the target location based on the method described in the first aspect above;
[0013] The camera pose after image optimization is used as one of the inputs of three-dimensional reconstruction to perform three-dimensional reconstruction of the target location.
[0014] In a third aspect, the present application provides a camera posture correction device, comprising:
[0015] an acquisition module, configured to acquire, from an image dataset, an image that satisfies a preset co-viewing condition with the current image as a co-viewing image, wherein the image dataset includes one or more two-dimensional images captured at the same target location as the current image;
[0016] The acquisition module is further configured to acquire a relative posture between a camera that captures the current image and a camera that captures the co-viewing image as a co-viewing relative posture;
[0017] The acquisition module is further configured to acquire an image from the image dataset that satisfies a preset similarity condition with the current image, and if the image is not the same as the co-viewed image, use the image as a loop closure image;
[0018] The acquisition module is further configured to acquire a relative posture between a camera that captures the current image and a camera that captures the loop-closed image as a loop-closed relative posture;
[0019] A correction module is used to correct the camera pose corresponding to each node in the pose graph with the optimization goal of narrowing the gap between the loop relative pose and the common view relative pose, and with the minimum relative pose change of the edge of the pose graph as the constraint item, so as to eliminate the cumulative error of the camera pose, wherein the pose graph includes: nodes and edges, a node corresponds to an image in the image dataset and the node records the camera pose of the camera that took the image, an edge connects two nodes, there is a common view relationship between the images corresponding to the two nodes, and the edge records the relative pose of the cameras recorded by the two nodes.
[0020] In a fourth aspect, the present application provides a three-dimensional reconstruction device, comprising:
[0021] An optimization module, configured to optimize the camera pose of the image captured at the target location based on the method described in the first aspect above;
[0022] The reconstruction module is used to use the camera pose after the image optimization as one of the inputs of three-dimensional reconstruction to perform three-dimensional reconstruction of the target place.
[0023] In a fifth aspect, the present application provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to perform the method provided in the first or second aspect of the present application.
[0024] In a sixth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method provided in the first aspect or the second aspect of the present application is implemented.
[0025] In a seventh aspect, the present application provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method provided in the first or second aspect of the present application.
[0026] The present application provides a method for correcting camera pose, a three-dimensional reconstruction method, and related devices. The method for correcting camera pose is to obtain a common view image based on a preset common view condition from an image data set, and obtain a loop image from the image data set. In this way, the accumulated error in the camera pose of the image can be determined based on the common view relationship between the images, and the determination process does not rely on the temporal correlation between the images. Furthermore, when correcting the camera pose corresponding to each node in the pose graph to eliminate the accumulated error, the optimization goal is to reduce the gap between the relative pose of the loop and the relative pose of the common view, and the constraint term is to minimize the relative pose change of the edges of the pose graph. In this way, the estimated camera pose of the image can be corrected while keeping each edge slightly changed or unchanged, so that the relative pose between the images can be kept as consistent as possible with that before correction. Through this optimization goal and this constraint term, the accumulated error can be avoided from being transferred to the corrected camera poses in the process of eliminating the accumulated error of the camera pose, and the camera pose of the image can be adjusted in a direction close to the true value, thereby improving the accuracy of the camera pose after correction. Since the optimization objective and the constraint item are used for correction during the correction process, the change amplitude of the three-dimensional points recursively obtained during the three-dimensional reconstruction process can be made smaller, thereby ensuring that the overall deformation of the three-dimensional reconstruction result is small and the accuracy of the three-dimensional reconstruction result is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0028] Figure 1 is a schematic diagram of a cyclic factor graph in the related art;
[0029] Figure 2 A flowchart of a method for correcting camera posture provided in an embodiment of the present application;
[0030] Figure 3 A schematic diagram of a pose graph provided in an embodiment of the present application;
[0031] Figure 4 A schematic diagram of a flow chart of a three-dimensional reconstruction method provided in an embodiment of the present application;
[0032] Figure 5 A schematic diagram of a trajectory curve in three-dimensional reconstruction provided in an embodiment of the present application;
[0033] Figure 6 A schematic diagram of the flow of the SFM loop detection module provided in an embodiment of the present application;
[0034] Figure 7 A schematic diagram of the structure of a camera posture correction device provided in an embodiment of the present application;
[0035] Figure 8 A schematic diagram of the structure of a three-dimensional reconstruction device provided in an embodiment of the present application;
[0036] Figure 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0037] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0038] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0039] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0040] 3D reconstruction is an important branch of computer vision technology. There are many different types of 3D reconstruction, including monocular 3D reconstruction, binocular 3D reconstruction, and 3D reconstruction based on red, green, blue, and depth (RGBD). Compared to binocular and RGBD 3D reconstruction, monocular 3D reconstruction requires less image acquisition and offers a more cost-effective reconstruction. Structure from Motion (SFM) is one type of monocular 3D reconstruction.
[0041] Existing SFM 3D reconstruction methods typically perform steps such as feature point detection, feature point matching, pose estimation, and 3D point cloud generation on multiple 2D images of a 3D scene, resulting in a 3D reconstruction result such as 3D point cloud data or a 3D model of the 3D scene. SFM 3D reconstruction tasks are typically performed offline on multiple 2D images that are not temporally related. In related applications, this can be achieved using tools such as Constrained Local Models for Automatic Pose Estimation (COLMAP) or Open Multi-View Stereo (OpenMVS).
[0042] Although the two-dimensional images collected by crowd-sources can be sorted according to the shooting time of the two-dimensional images, crowd-source means that the acquisition device that collects the images is not a single device. That is, the two-dimensional images collected by crowd-sources can be understood as being collected by multiple different acquisition devices. Therefore, even if the two-dimensional images collected by crowd-sources are sorted by time, because the acquisition devices that collect these two-dimensional images are not the same, although the two-dimensional images after sorting have a time sequence, the changes in the camera posture when the acquisition device takes these two-dimensional images are not related to the time sequence.
[0043] For example, in a shopping mall, different collectors use their mobile phone cameras to capture images of the mall's interior. These collectors upload the captured images and obtain crowdsourced 2D images of the mall. Since these 2D images are captured by different collectors, even if two collectors use their mobile phones to capture the same scene at similar times, the camera poses of the two phones capturing the two images will differ due to the different ways the collectors hold their phones or their different positions or orientations. Furthermore, since the two phones are independent devices, the difference in their camera poses is unrelated to the time at which they captured the images. Therefore, even if crowdsourced 2D images can be sorted by time, the camera poses used to capture these 2D images have no correlation with the time sequence, and the camera pose of a temporally adjacent 2D image cannot be used to infer the camera pose of another 2D image.
[0044] Different from multiple 2D images without a temporal connection are multiple 2D images with a temporal connection. For example, in a shopping mall, a subject records a video of some scenes there. From this video, multiple 2D images with a temporal connection can be captured. Because these images were captured by the subject using the same handheld acquisition device over a continuous period of time, the multiple camera poses corresponding to each image change with the time series. Therefore, there is a strong coupling between the camera poses of the multiple images in the preceding and following frames. When determining the camera pose of a single image, the camera poses of the preceding and following frames can be combined to infer it. Therefore, for multiple 2D images with a temporal connection, it is easier to accurately predict the camera pose of each image.
[0045] In SFM 3D reconstruction, it is necessary to estimate the camera pose when capturing each 2D image. The camera pose includes the camera's position and orientation at the time of capture. The camera pose can indicate the position and orientation of the camera when capturing the corresponding 2D image. Since the camera pose is estimated by an estimation algorithm, it can be understood as the estimated pose of the 2D image. After obtaining the camera pose of each 2D image, the distance between each pixel in the corresponding 2D image and the camera in 3D space can be calculated based on each camera pose through triangulation and other methods. In other words, the position data of the 3D (3-Dimensional) point corresponding to each pixel can be obtained, thereby achieving the purpose of obtaining 3D point cloud data for each 2D point in the 2D image based on the planar 2D image, and thus realizing 3D reconstruction.
[0046] When estimating the camera pose of each two-dimensional image, it is usually done recursively. For example, after determining the camera pose of the initial two-dimensional image, the camera pose of other two-dimensional images related to the initial two-dimensional image is estimated based on the estimated camera pose. Due to factors such as sensor noise of the shooting device, feature point extraction errors, or data association errors, the subsequent estimated camera poses will produce cumulative errors in the process of continuous recursive estimation, causing the subsequent camera poses to gradually deviate from the true pose. In particular, when reconstructing large scenes based on a large number of two-dimensional images, estimating the camera pose image by image will cause the errors generated by each estimation to be continuously transmitted backward and accumulated, forming cumulative errors, resulting in the accuracy of the subsequent estimated camera poses becoming lower and lower, gradually moving away from the true value. After the cumulative error is generated, it is reflected in the pose trajectory formed by each camera pose, that is, the shape of the pose trajectory gradually diverges and deviates from the true trajectory, resulting in the inability to close the trajectory points that should be closed in the pose trajectory; it is reflected in the 3D points corresponding to the pixel points, that is, the 3D point position data reconstructed based on each camera pose gradually deviates from the true position of the 3D point, resulting in the divergence of the 3D points. It can be understood that both images include the same 3D point in the three-dimensional space, but due to the divergence caused by the deviation in camera pose estimation, the 3D point in the two images cannot be closed together in the three-dimensional reconstruction results and cannot be aligned.
[0047] The divergence of the camera pose trajectory caused by the accumulated error will cause significant drift in the 3D point cloud data calculated based on the camera pose, greatly affecting the accuracy of the 3D reconstruction results. Therefore, if the impact of the accumulated error cannot be effectively eliminated or reduced, the accuracy of the 3D reconstruction results will be low.
[0048] To solve the above problems, we can consider drawing on the ideas of the Simultaneous Localization and Mapping (SLAM) task. The SLAM task continuously collects multiple two-dimensional images with a time-series relationship in real time and at short intervals based on the image capture device, and establishes a corresponding ring factor graph based on each two-dimensional image. By optimizing the ring factor graph, the cumulative error of the camera pose corresponding to each two-dimensional image can be eliminated or reduced. Since the two-dimensional images of the SLAM system are continuously collected in real time and at short intervals, many common feature points can be observed between consecutive two-dimensional images. Therefore, the ring factor graph it constructs can naturally form effective constraints between the camera poses, thereby achieving effective ring factor graph optimization.
[0049] Figure 1 is a schematic diagram of a cyclic factor graph in the related art, such as Figure 1As shown, the cyclic factor graph is a graph structure data composed of nodes and edges, where the boxes represent nodes and the dotted lines represent edges between nodes. Among them, each node represents the camera pose corresponding to the two-dimensional images collected at different time points within a continuous time period, and each edge represents the relative relationship between the camera poses. When performing a SLAM task, usually in the process of collecting images, the shooting device will continuously collect multiple two-dimensional images in the current scene (for example, shooting a video), and return to the starting position when the collection of the current scene is completed. In the process of constructing the cyclic factor graph, the node 101 corresponding to the first frame image at the starting position can be added first, and then the nodes corresponding to each image can be added in sequence according to the temporal relationship of the time points of shooting each image. The node corresponding to the nth image is fixedly connected to the node corresponding to its previous image (i.e., the n-1th image), and the camera pose of the current node is estimated based on the camera pose of the previous node. When the last frame of image is captured, the camera returns to the starting position. Therefore, the last frame of image and the first frame of image are in the same posture when they are captured. Therefore, the node 102 corresponding to the last frame of image and the node 101 corresponding to the first frame of image have a loop relationship. Then, the previously estimated camera postures can be optimized based on the loop relationship.
[0050] In summary, in SLAM tasks, pose graph optimization when loop closure is triggered needs to rely on the time sequence between images to construct a ring factor graph. If SLAM tasks are applied to SFM systems, since the images collected by the SFM system are offline, multi-source two-dimensional images without time sequence association, these two-dimensional images do not have a time sequence relationship. Therefore, it is impossible to construct a loop factor graph. Figure 1 The ring factor graph shown cannot eliminate the cumulative error when estimating the poses of each camera according to the ring factor graph, and thus cannot improve the accuracy of the camera poses. Therefore, it is difficult to improve the accuracy of the 3D reconstruction results.
[0051] Under the premise of overcoming the above difficulties and in order to effectively reduce the impact of cumulative errors on the accuracy of camera pose during the 3D reconstruction process and improve the accuracy of the 3D reconstruction results, the present application provides a method for correcting camera pose. The method obtains an image from the image data set that meets a preset common view condition with the current image as a common view image; and obtains an image from the image data set that meets a preset similarity condition with the current image. If the image is not the same as the common view image, the image is used as a loop closure image. In this way, based on the common view relationship between images and the similarity between images, it can be determined that there is a cumulative error in the camera pose of the image. This determination process does not rely on the temporal relationship between images, and thus can overcome the defect of not being able to trigger loop closure when performing 3D reconstruction based on multiple images without temporal connection.
[0052] This method also obtains the relative pose between the camera capturing the current image and the camera capturing the coview image as the commonview relative pose; and obtains the relative pose between the camera capturing the current image and the camera capturing the loop-closure image as the loop-closure relative pose. Furthermore, with the optimization objective of narrowing the gap between the loop-closure relative pose and the commonview relative pose, and the constraint of minimizing the relative pose change of the pose graph edges as the constraint, the camera pose corresponding to each node in the pose graph is corrected to eliminate the cumulative error in the camera pose. This correction, which takes into account the constraint of minimizing the relative pose change of the pose graph edges, avoids transferring the accumulated error to the corrected camera poses during the process of eliminating the cumulative error in the camera poses. The joint optimization based on this optimization objective and the constraint can adjust each camera pose closer to the true value, thereby improving the accuracy of the corrected camera poses.
[0053] In addition, since the optimization objective and the constraint item are used to correct the camera pose during the correction process, the change amplitude of each three-dimensional point recursively obtained during the three-dimensional reconstruction process can be smaller, and the overall deformation of the three-dimensional reconstruction result can be smaller. Under this premise, the camera pose is corrected again, so the accuracy of the three-dimensional reconstruction result can be improved.
[0054] Figure 2 This is a flow chart of a method for correcting the camera posture provided by an embodiment of the present application. This method can be executed by an electronic device with corresponding data processing capabilities, which can be in the form of a server or terminal device. Figure 2 As shown, the camera posture correction method includes the following steps:
[0055] Step S201 : obtaining an image that satisfies a preset co-viewing condition with the current image from an image dataset as a co-viewing image, wherein the image dataset includes one or more two-dimensional images captured at the same target location as the current image.
[0056] In this step, the target location can be a real three-dimensional space at any location or structure. For example, the target location can be an outdoor space or an indoor space, such as a shopping mall, a square, or an underground parking lot. The current image can be a two-dimensional image captured at any time and / or by any image acquisition device of the target location. The image dataset can be a dataset including multiple two-dimensional images, one or more of which are images captured of the target location. Exemplarily, the common view condition can be any condition used to determine whether a common view relationship exists between two or more images. A common view relationship between images can be understood as a visual association relationship exhibited when the viewpoints of two or more image capture devices can simultaneously observe specific points or objects in the same scene. It can be understood that if two or more images share certain specific points or objects, then the images have a common view relationship; if two or more images do not share certain specific points or objects, then the images do not have a common view relationship.
[0057] Taking two images as an example, if object C exists in both image A and image B, it can be understood that there is a common view relationship between image A and image B; if all objects in image A do not appear in image B, it can be understood that there is no common view relationship between image A and image B.
[0058] Detecting or determining whether images have a common view relationship typically involves computer vision and image processing technologies. The common view conditions for determining a common view relationship can include feature point matching, perspective comparison, and / or geometric consistency. For example, the common view condition may include whether the number of successfully matched feature points exceeds a preset threshold. If so, the common view condition is satisfied; otherwise, the common view condition is not satisfied.
[0059] Exemplarily, when acquiring a co-viewing image, co-viewing detection can be performed on the current image and any image in the image data set using preset co-viewing conditions. If it is detected that at least one image satisfies the preset co-viewing condition with the current image, one of the at least one images can be acquired as a co-viewing image. For example, the image with the strongest co-viewing relationship with the current image among the at least one image is determined as the co-viewing image.
[0060] For example, for any image in the image dataset and the current image, multiple feature points of each are extracted from the two images respectively. For example, feature point detection algorithms such as Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), or Oriented FAST and Rotated BRIEF (ORB) are used to extract multiple feature points from the image. Each feature point is described to obtain a feature descriptor for each feature point. The feature descriptor can be understood as an image information descriptor used to describe the corresponding feature point in the adjacent pixel area. The feature descriptor can be a multi-dimensional feature vector.
[0061] By matching each feature descriptor, the number and / or ratio of the same feature points between the two images are calculated. If the calculated number and / or ratio of the same feature points is less than or equal to a preset number and / or preset ratio, it can be determined that the image does not meet the preset co-viewing condition with the current image and cannot be obtained as a co-viewing image of the current image. If the calculated number and / or ratio of the same feature points is greater than the preset number and / or preset ratio, it can be determined that the image meets the preset co-viewing condition with the current image and can be obtained as a co-viewing image of the current image. The ratio can be the ratio of the number of the same feature points to the total number of all feature points.
[0062] For example, in order to determine the co-viewing relationship with higher accuracy, the 3D points corresponding to the same feature points in the two images may be matched. If the 3D point matching meets the matching conditions, it can be determined that there is a co-viewing relationship between the two images.
[0063] For example, by matching the descriptors of each feature point, if 95% of the feature points in the current image and any image in the image dataset are identical, and the calculated proportion of identical features (95%) is greater than the preset proportion (90%), then it can be determined that the current image and the image in question are highly likely to have a co-viewing relationship. The 3D point position data for these identical feature points is calculated based on the camera pose of the current image, and the 3D point position data for these identical feature points is calculated based on the camera pose of the image in question. Matching is performed based on the calculated 3D point position data. If the matching conditions are met, it can be determined that the two images have a co-viewing relationship.
[0064] The matching condition is used to determine whether the 3D point position data of each common feature point in the two images is the same or close. For example, the matching condition can be: for all common feature points, whether the difference between the two 3D point position data of the common feature points in the two images is less than a preset difference. If so, the matching condition is met; otherwise, the matching condition is not met.
[0065] It should be understood that since the camera poses are estimated values, the 3D point position data calculated for the same feature point under two different camera poses may deviate significantly. By setting a preset difference, a more reliable match of the 3D point position data can be achieved. Using matching conditions similar to those in the examples above, it is possible to determine whether the viewpoints of the two images observe a sufficient number of identical 3D points, and thus whether the two images have a co-viewing relationship. Based on this, the co-viewing image of the current image can be determined with relatively high accuracy.
[0066] Step S202 : Acquire the relative posture between the camera that captures the current image and the camera that captures the co-viewing image as the co-viewing relative posture.
[0067] For example, the camera pose of a 2D image can be solved based on a matrix describing the geometric relationship between the two images, as well as known 3D points and their 2D projections in the 2D image. For example, the camera pose of the current image or any image in the image dataset can be calculated using the Perspective-n-Point (PnP) problem for estimating the camera pose. Alternatively, the camera pose can be solved using algorithms such as the Efficient Perspective-n-Point (EPnP) and Robust Perspective-n-Point (RPnP) algorithms.
[0068] For example, in order to describe more concisely, the camera pose of a camera that takes an image can also be called the camera pose of the image, and the camera pose of the image can be understood as the camera pose of the camera that takes the image; further, the relative pose between cameras that take two images can also be called the relative pose between the two images, and the relative pose between images can be understood as the relative pose between the cameras that take these images.
[0069] After obtaining the camera pose of the current image and the camera pose of the coview image, the relative pose between the camera that took the current image and the camera that took the coview image can be obtained by calculating the difference between the two camera poses, and then the relative pose can be used as the coview relative pose.
[0070] For example, the camera pose includes the position and direction of the camera when taking the image. The position difference in two camera poses is obtained by subtracting the position, and the direction difference in two camera poses is obtained by subtracting the direction. Therefore, the relative pose can include the position difference and the direction difference.
[0071] Step S203: Obtain an image from the image dataset that meets a preset similarity condition with the current image. If the image is not the same as the co-viewed image, use the image as a loop image.
[0072] Illustratively, the similarity condition may be any condition for determining whether two or more images are similar. For example, the similarity condition may include comparing the similarity between the two images with a preset similarity threshold; if the similarity is greater than the preset similarity threshold, the preset similarity condition is satisfied; if the similarity is not greater than the preset similarity threshold, the preset similarity condition is not satisfied.
[0073] The preset similarity threshold can be any preset value for judging the similarity between two two-dimensional images, and whether the two two-dimensional images are similar images can be determined based on the preset similarity threshold. When determining the similarity between two two-dimensional images, it can be determined by a model-based image similarity detection algorithm, for example, it can be based on a pre-trained model such as the Visual Geometry Group (VGG) and the Residual Network (ResNet). Compared with the Bag of Words (BoW) algorithm, the model-based image similarity detection algorithm has the advantages of strong robustness and high accuracy, and can improve the detection rate of loop relationships.
[0074] If at least one image that meets a preset similarity condition with the current image is obtained from the image data set, and if the at least one image is not the same image as the co-viewed image, one of the at least one images can be used as a loop image. For example, the image with the highest similarity to the current image among the at least one image is determined as the loop image.
[0075] A loop closure image can be understood as an image that is sufficiently similar to the current image but is not a co-viewing image. Since the loop closure image is sufficiently similar to the current image, there is a high probability that the loop closure image and the current image were captured at the same or similar camera position and using the same or similar camera orientation. A co-viewing image can be understood as an image that has a sufficient co-viewing relationship with the current image, and there is a high probability that the co-viewing image and the current image are sufficiently similar. However, if the loop closure image and the co-viewing image are not the same image, it can be indicated that when determining the co-viewing image of the current image, the camera pose of the co-viewing image has a significant deviation. This can be understood as the cumulative error is already large enough when recursively estimating the camera pose of the co-viewing image. Assuming that the cumulative error is not large enough, the loop closure image that is sufficiently similar to the current image should be a co-viewing image that has a sufficient co-viewing relationship with the current image.
[0076] Therefore, in the process of 3D reconstruction, when the 3D reconstruction comes to the current image, both the co-view image and the loop closure image are obtained in the image data set, which indicates that the cumulative error of the recursive estimation of each camera pose is already large. At this time, the loop closure optimization can be triggered to correct the estimated camera poses of each image to reduce or eliminate the cumulative error of the camera pose and improve the accuracy of the camera pose of each image.
[0077] Step S204: Obtain the relative pose between the camera that captures the current image and the camera that captures the loop-closed image as the loop-closed relative pose.
[0078] For example, after obtaining the camera pose of the current image and the camera pose of the loop closure image, the relative pose between the camera that captured the current image and the camera that captured the loop closure image can be obtained by calculating the difference between the two camera poses, and the relative pose can then be used as the loop closure relative pose. For example, the position difference value can be obtained by subtracting the camera position in the camera pose of the current image from the camera pose of the loop closure image, and the direction difference value can be obtained by subtracting the camera direction in the camera pose of the current image from the camera pose of the loop closure image.
[0079] Step S205, with the optimization goal of narrowing the gap between the loop relative pose and the common view relative pose, and with the constraint of minimizing the relative pose change of the edges of the pose graph, the camera pose corresponding to each node in the pose graph is corrected to eliminate the cumulative error of the camera pose, wherein the pose graph includes: nodes and edges, a node corresponds to an image in an image dataset and the node records the camera pose of the camera that took the image, an edge connects two nodes, there is a common view relationship between the images corresponding to the two nodes, and the edge records the relative pose of the cameras recorded by the two nodes.
[0080] Exemplarily, a pose graph includes nodes and edges. A node corresponds to an image in an image dataset and records the camera pose of the camera that captured the image. This can be understood as a node including an image and the camera pose corresponding to the image. This camera pose can be an estimated camera pose. An edge connects two nodes. There is a co-viewing relationship between the images corresponding to these two nodes. This edge records the relative poses of the cameras recorded by these two nodes. The existence of a co-viewing relationship can be understood as having a co-viewing relationship, but it does not necessarily mean that the preset co-viewing conditions are met. The edge records the relative poses of the cameras of the two nodes. This can be understood as the edge representing the difference between the camera poses of the two nodes.
[0081] The "graph" in the pose graph refers to a graph, which is a data structure used to record the co-viewing relationship between images and the camera pose. A pose graph can be understood as a tool or means for adjusting the camera pose during correction. Therefore, when a loop closure image is detected, triggering loop closure optimization to correct the camera pose, a pose graph can be generated and corrections can be made based on the pose graph. Alternatively, instead of generating a pose graph, the camera pose can be corrected using a graph-structured data structured by data such as the images, their camera poses, the relative poses between the camera poses, and the co-viewing relationship between the images as an invisible pose graph.
[0082] For example, when performing 3D reconstruction, the corresponding pose graph can be generated when processing the first image. As more images are processed, nodes and edges in the pose graph are added, allowing the pose graph to be updated as image processing progresses. When loop closure optimization is detected and loop closure optimization is required, the pose graph can be used to modify the camera pose of each node in the pose graph by combining the optimization objectives and constraints.
[0083] Alternatively, when performing a 3D reconstruction task, a pose graph is not generated first. As the number of processed images increases, when loop closure images are detected and loop closure optimization is required, a pose graph is generated based on the previously processed images and the camera pose of each image. Based on the pose graph, the camera pose of each node in the pose graph is corrected in combination with the optimization objectives and constraints.
[0084] Alternatively, when performing 3D reconstruction tasks, a pose graph is not generated. As the number of processed images increases, when loop closure images are detected and loop closure optimization is required, the poses of each camera are corrected based on the previously processed images, the camera poses of each image, the relative poses between the camera poses, and the common view relationship between the images as graph-structured data. Although a pose graph is not constructed in this process, each image and its camera pose can be invisibly used as a node. Edges are established between the nodes corresponding to two images with a common view relationship, and the relative poses between the nodes are used as the edge information. An optimization method is used to combine the optimization objectives and constraints to correct the poses of each camera.
[0085] For example, the loop closure relative pose is the relative pose between the current image and the loop closure image, which can be understood as the difference between the camera poses of these two images. The common view relative pose is the relative pose between the current image and the common view image, which can be understood as the difference between the camera poses of these two images. Detecting the presence of loop closure images in the image dataset can trigger loop closure optimization to correct the estimated camera poses of each image.
[0086] The cumulative error of camera pose is the main factor causing the difference between loop closure relative pose and common view relative pose. Therefore, the optimization goal is to reduce the gap between loop closure relative pose and common view relative pose, which is to reduce or eliminate the cumulative error of camera pose.
[0087] It should be understood that relative pose change and relative pose are not the same concept. The relative pose change of the edge of the pose graph can be understood as the difference between the relative pose represented by the edge before and after correction.
[0088] With the optimization goal of narrowing the gap between the relative pose of the loop closure and the relative pose of the common view, and the constraint of minimizing the relative pose change of the edges of the pose graph, the camera pose corresponding to each node in the pose graph can be corrected using optimization methods from related technologies. For example, global optimization algorithms or local optimization algorithms can be used.
[0089] Taking the global optimization algorithm as an example, if the loop closure relative pose is a and the common view relative pose is a', then the difference between the loop closure relative pose and the common view relative pose is a-a'. During optimization, with the goal of reducing a-a', a new camera pose is generated for each node in the pose graph. After obtaining the new camera pose for each node, the difference in the camera pose of the two nodes with an edge is calculated based on the new camera pose. This can be understood as calculating a new relative pose for each edge. Then, the difference between the relative pose of any edge before the update and the relative pose after the update is determined. This difference can be regarded as the relative pose change of the edge. The relative pose changes of all edges can be added together to obtain the total change of all edges corresponding to the new camera pose generated in the first round.
[0090] The above process is repeated again to generate a new camera pose for each node with the goal of reducing a-a', calculate a new relative pose for each edge, and determine the total change of all edges corresponding to the new camera pose generated in the second round.
[0091] Compare the total changes from the first round with the total changes from the second round, and determine the smaller total change. The smaller total change corresponds to the camera pose of each node generated in that round, and can be used as the corrected camera pose for each node. As can be seen from the above example, this process achieves the goal of narrowing the gap between the loop closure relative pose and the common view relative pose, and constrains the relative pose change of the edges of the pose graph to minimize the relative pose change of each node in the pose graph.
[0092] Since the camera poses of all nodes in the pose graph are corrected in the above example, it can be understood as a correction based on a global optimization algorithm executed globally. If the camera poses of some nodes in the pose graph are corrected, it can be understood as a correction based on a local optimization algorithm. In addition, the above example can continue to generate and compare multiple rounds, with the smallest total change in multiple rounds as the optimal solution, and the camera poses generated corresponding to the round with the smallest total transformation as the corrected camera pose, thereby achieving a camera pose correction with higher accuracy. Alternatively, when comparing any two rounds, the two dimensions of the degree of reduction of a-a' and the total change can be jointly compared, and the optimization objectives and constraints can be comprehensively considered to determine the corrected camera pose, which can further improve the accuracy of the corrected camera pose.
[0093] For example, the optimization goal is to reduce the gap between the relative pose of the loop and the relative pose of the common view, which is an optimization goal set for the purpose of eliminating cumulative errors. With the constraint of minimizing the relative pose change of the edges of the pose graph, the camera pose corresponding to each node in the pose graph is corrected. The estimated camera pose of each image can be adjusted while keeping all edges as unchanged as possible or changing them slightly. In this way, the relative pose between each image can be kept as close as possible to that before correction, which allows for slight changes to each 3D point recursively obtained during the 3D reconstruction process, minimizing the overall deformation of the 3D reconstruction result and preventing excessive deviation between the 3D reconstruction result before and after the correction.
[0094] If the optimization goal is only to reduce the gap between the relative pose of the loop closure and the relative pose of the common view, and no constraint is set to minimize the relative pose change of the edges of the pose graph, although the cumulative error is eliminated between the loop closure image and the common view image, when correcting the camera pose, the cumulative error may be transferred to the corrected camera pose of each image during correction due to the lack of this constraint. This may cause a significant change in the 3D points recursively obtained during the 3D reconstruction process, resulting in excessive deformation of the 3D reconstruction result and reduced accuracy of the 3D reconstruction result.
[0095] The camera pose correction method provided in the embodiment of the present application obtains a common view image from the image data set based on a preset common view condition, and obtains a loop image from the image data set. In this way, the cumulative error in the camera pose of the image can be determined based on the common view relationship between the images, and the determination process does not rely on the temporal correlation between the images. Furthermore, when correcting the camera pose corresponding to each node in the pose graph to eliminate the cumulative error, the optimization goal is to reduce the gap between the relative pose of the loop and the relative pose of the common view, and the constraint term is to minimize the relative pose change of the edges of the pose graph. In this way, the estimated camera pose of the image can be corrected while keeping each edge slightly changed or unchanged, so that the relative pose between the images can be kept as consistent as possible with that before correction. Through this optimization goal and this constraint term, the cumulative error can be avoided from being transferred to the corrected camera poses in the process of eliminating the cumulative error of the camera pose, and the camera pose of the image can be adjusted in a direction close to the true value, thereby improving the accuracy of the camera pose after correction. Since the optimization objective and the constraint item are used for correction during the correction process, the change amplitude of the three-dimensional points recursively obtained during the three-dimensional reconstruction process can be made smaller, thereby ensuring that the overall deformation of the three-dimensional reconstruction result is small and the accuracy of the three-dimensional reconstruction result is improved.
[0096] Exemplarily, when correcting the camera pose, the correction can be performed based on the pose graph.
[0097] In a possible implementation, the method further includes: generating a pose graph based on the co-viewing relationship between images in the image dataset and the camera pose of a camera that captures each image.
[0098] In 3D reconstruction, since multiple images from multiple sources are used without temporal correlation, it is difficult to establish camera pose constraints between images through temporal correlation. The present invention overcomes this shortcoming by using the common view relationship between images.
[0099] The co-view relationship between images can more accurately reflect the pose correlation of viewpoints between images. Therefore, based on the co-view relationship between images in an image dataset, a pose graph can be generated that effectively constrains the camera pose of each image with a co-view relationship. When generating a pose graph, any image and the camera pose of the camera that captured it can be used as a node. This allows multiple nodes to be created based on multiple images of the target location in the image dataset. After creating the nodes, by determining whether the images corresponding to any two nodes have a co-view relationship, edges can be established between the nodes corresponding to the images with a co-view relationship, thereby generating a pose graph.
[0100] To determine whether there is a co-viewing relationship between any two images, any method for determining the co-viewing relationship between images can be used. For example, similar to the method for determining co-viewing images in the above embodiment, the preset number in the above co-viewing condition can be adjusted to a smaller value as the judgment value for determining the co-viewing relationship, for example, it can be adjusted to 0 or other smaller values. When the number of identical feature points calculated between the two images is greater than the judgment value, the two images can be determined to be images with a co-viewing relationship; when the number of identical feature points calculated between the two images is less than or equal to the judgment value, the two images can be determined to be images without a co-viewing relationship. Alternatively, other methods can be used to determine the co-viewing relationship between images, which is not limited in this embodiment of the present application. In this embodiment of the present application, each node in the pose graph can be constructed based on the camera pose of the camera that captured each image, and based on determining whether there is a co-viewing relationship between the images in the image data set, edges between the nodes can be constructed, thereby generating a pose graph more accurately.
[0101] In 3D reconstruction tasks such as SFM, the camera pose estimation of the image and the generation of 3D point cloud data are both carried out recursively. Therefore, in order to better fit the overall process of 3D reconstruction, a recursive method can be used to generate the pose graph to improve the accuracy and timeliness of 3D reconstruction.
[0102] In one possible implementation, when generating a pose graph based on the co-viewing relationship between images in an image dataset and the camera pose of a camera that captured each image, it can be specifically implemented as follows: select any two images with a co-viewing relationship from the image dataset, construct two nodes of the pose graph based on the camera pose of the camera that captured the any two images, and construct edges between the two nodes in the pose graph with the relative poses of the cameras recorded by the two nodes; use the two nodes as seed nodes respectively, search for other images that have a co-viewing relationship with the image of the seed node in the remaining images in the image dataset, and construct new nodes and edges, and continue to search and construct new nodes and edges in the pose graph with the newly added nodes as new seed nodes, until nodes and edges corresponding to all images with a co-viewing relationship in the image dataset are constructed in the pose graph.
[0103] Specifically, for each image of the target place in the image data set, any two images with a co-viewing relationship can be selected from each image in a random or sequential manner. For example, after arbitrarily selecting an image, another image with a co-viewing relationship with the selected image can be determined by determining the co-viewing relationship between images. Then, these two images can be selected as two images with a co-viewing relationship.
[0104] After identifying the two images, the two images and their corresponding camera poses are used as nodes in the pose graph, and an edge is constructed between these two nodes to generate an initial pose graph. This initial pose graph contains two nodes corresponding to the two images, along with an edge between them that records the relative poses of the cameras that captured the two images. These two nodes can serve as seed nodes, and subsequent nodes can be expanded using these two seed nodes. This can be understood as a tree-like pose graph, with the seed node as the root node and leaf nodes continuously expanding.
[0105] For example, for any seed node, the remaining images in the image dataset are searched for other images that share a co-viewing relationship with the seed node's image. When an image sharing a co-viewing relationship with the seed node's image is found, the image and the camera pose of the camera that captured it are added as new nodes to the initial pose graph. When constructing new nodes, edges can be constructed. For example, an edge can be added between two nodes corresponding to two images sharing a co-viewing relationship.
[0106] The process of searching and building new nodes can be repeated for multiple rounds. Using the newly built nodes as new seed nodes, the above process is continued to search and build new nodes and edges in the pose graph until all nodes and edges corresponding to the images in the image dataset that have a common view relationship are built in the pose graph.
[0107] Figure 3 A schematic diagram of a pose graph provided in an embodiment of the present application is shown in FIG. Figure 3 As shown in the figure, based on two images with a co-viewing relationship and the camera poses of the images, seed nodes 301 and 302 are constructed. For seed nodes 301 and 302 respectively, other images with a co-viewing relationship with the seed node images are searched in the remaining images, and new nodes and edges are constructed along the poses of the found images. Figure 3 The middle nodes 303 to 310 are all new nodes corresponding to the images found to have a common view relationship with the seed nodes 301 and 302. Corresponding edges can be constructed between the nodes according to the relative postures.
[0108] The newly constructed nodes are used as new seed nodes to continue searching and construct new nodes and edges. The above process of searching and constructing new nodes and edges can be repeated multiple times. For example, using nodes 303 to 310 as new seed nodes, new nodes 311 to 322 and their corresponding edges are searched and constructed. Figure 3The figure only shows two rounds of constructing new nodes and edges. In practice, the number of rounds for finding and constructing new nodes and edges can be unlimited. This process continues in this way until all nodes corresponding to images with a common view relationship under the target location in the image dataset are added to the pose graph.
[0109] In this embodiment, by creating a seed node and then constructing other new nodes based on the seed node, the generation of the pose graph can be completed quickly in a diffusion manner, thereby improving the generation efficiency. In addition, this method can also be understood as a recursive pose graph generation method, which is consistent with the method of recursively processing three-dimensional point clouds for each image during three-dimensional reconstruction, and can cooperate with the advancement of three-dimensional reconstruction. For example, a pose graph can be generated while three-dimensional reconstruction is being performed, or three-dimensional reconstruction can be advanced while the pose graph is generated. In this way, during the three-dimensional reconstruction process, the camera pose can be corrected based on the pose graph at any time, the three-dimensional reconstruction results can be corrected in time, the accumulation of errors in the three-dimensional reconstruction process can be reduced, and the accuracy of the three-dimensional reconstruction results can be improved. Moreover, since the two processes of three-dimensional reconstruction and pose graph generation can be progressively expanded with each other, three-dimensional reconstruction can be performed in a timely manner, which can improve the timeliness of three-dimensional reconstruction.
[0110] In one possible implementation, for a seed node, if at least two other images having a common-view relationship with the image of the seed node are found in the remaining images in the image data set, the method further includes: obtaining an image with the highest common-view relationship strength from the at least two other images having a common-view relationship with the image of the seed node; wherein the common-view relationship strength is used to characterize the degree of matching of three-dimensional points corresponding to feature points of the two images; and constructing new nodes and edges, specifically: constructing a new node in the pose graph based on the camera pose of the camera that took the image with the highest common-view relationship strength, and constructing a new node and an edge of the seed node in the pose graph using the relative pose of the camera recorded by the new node and the seed node.
[0111] Specifically, when constructing new nodes corresponding to images one by one, based on the new seed node, all images without added nodes are traversed, and two or more images having a co-viewing relationship with the image of the new seed node may be found.
[0112] For example, performing PnP calculations on the new seed node image and all remaining images may yield multiple 2D images with varying co-viewing strengths. Among these images, the image with the highest co-viewing strength (Frame & 1st PnP Inlier) is determined as the new node image, i.e., the new seed node's strong co-viewing neighbor (Neighbor). The camera pose of this image can be added as a new node to the pose graph, and an edge can be established between the two nodes.
[0113] For another example, for other images with the highest non-co-viewing relationship strength, nodes may not be added temporarily. When other new nodes are expanded, these images can still participate in the step of constructing new nodes.
[0114] Although the multiple images in the SFM system lack temporal correlation, the aforementioned method can still be used to construct a pose graph that represents the co-viewing situation between the images. The nodes in the graph reflect the camera pose of the image, and the edges reflect the difference in the camera pose between two images, that is, the relative pose between the two images. Based on this, a pose graph generation method that is independent of time sequence is designed. The relative pose between each camera and the camera with the strongest co-viewing relationship during the PnP solution is used as the residual, replacing the relative pose constraints at adjacent moments. This also serves to constrain the shape of the local camera pose trajectory.
[0115] For example, the strength of the common view relationship can be understood as the degree of matching between the three-dimensional points corresponding to the feature points of the two images. The three-dimensional points corresponding to the feature points can be the positions of the feature points in three-dimensional space obtained by 3D projection of the feature points based on the camera pose of the images.
[0116] For example, when matching three-dimensional points, the higher the consistency of the 3D point position data for the same feature point in two images, the stronger the co-viewing relationship between the two images, and vice versa. Alternatively, the more feature points that successfully match the 3D point positions, the stronger the co-viewing relationship between the two images, and vice versa. For example, if the 3D point position data corresponding to the largest number of feature points in two images are successfully matched, the co-viewing relationship between the two images is the strongest.
[0117] In this embodiment, new nodes and edges are constructed in the pose graph based on the image with the strongest common-view relationship. This allows for optimal selection of new nodes for various subnodes based on the strength of the common-view relationship. This allows for the establishment of more reliable edge constraints, i.e., more reliable relative pose constraints. This ensures that when the pose graph is generated, two nodes with an edge have stronger relative pose constraints. This improves the reliability of the corrections made based on the pose graph, thereby increasing the accuracy of the corrected camera pose and reducing the cumulative error generated during recursive camera pose derivation.
[0118] The embodiment of the present application also provides a three-dimensional reconstruction method. Figure 4 This is a flow chart of a three-dimensional reconstruction method provided in an embodiment of the present application. This method can be executed by an electronic device with corresponding data processing capabilities, which can be in the form of a server or terminal device. Figure 4 As shown, the three-dimensional reconstruction method includes the following steps:
[0119] Step S401 : Optimizing the camera pose of an image captured at a target location based on the camera pose correction method of any of the above embodiments.
[0120] In this step, the camera pose correction method described in any of the above embodiments can be used to correct the camera pose of the image being reconstructed. For example, based on the generated pose graph, the camera pose of each node can be corrected using the pose graph. This can eliminate the cumulative error in the estimated camera poses of each image during the 3D reconstruction process, thereby improving the accuracy of the camera pose. The specific implementation process can be referred to the description of the above embodiments and will not be repeated here.
[0121] Step S402: Using the camera pose after image optimization as one of the inputs for 3D reconstruction, the target location is reconstructed in 3D.
[0122] For example, the process of generating a pose graph and the process of generating three-dimensional point cloud data can be regarded as an overall process that has a mutual dependence and can jointly promote the three-dimensional reconstruction process.
[0123] For example, after generating a seed node in the pose graph, a triangulation operation is performed based on the camera pose of the seed node and the 2D coordinates of each feature point in the seed node's image. This generates the corresponding 3D point cloud data for each feature point in 3D space, which then forms part of the 3D reconstruction result. By continuously adding new nodes to the pose graph and generating point cloud data for each new node, the 3D reconstruction process can be gradually advanced, ultimately resulting in a 3D reconstruction of the target location.
[0124] In the process of generating a pose graph, a node newly added to the pose graph can be used as a target node to establish a corresponding edge to the target node. When the current image is added to the pose graph, the target node is the node corresponding to the current image. The processing process for any target node includes: connecting the target node with a common view node, which is a node in the pose graph whose corresponding image has a common view relationship with the image of the target node; if the existing nodes also contain nodes with a loop relationship with the target node, then loop optimization is performed on the camera pose of each node in the current pose graph based on the loop relationship to correct each camera pose. Among them, the loop relationship can be understood as a node-to-node relationship that can trigger loop optimization. For example, if the current image and the loop image of the current image can trigger loop optimization, then the two have a loop relationship.
[0125] Optionally, in an embodiment of the present application, a pose graph can be generated based on a common view relationship. When a pose graph is generated using a common view relationship, a loop relationship can be understood as: the images corresponding to two nodes are taken at the same or similar camera poses, but there is no edge between the two nodes. Based on this loop relationship, loop optimization can be triggered. The principle is: for any two nodes, if the image similarity of the two nodes is high, it can be considered that the two images are taken at the same or similar camera poses, and the captured image content is relatively similar, and they should have a common view relationship. Therefore, in the pose graph, there should be an edge between the two nodes. In addition, the camera pose of the node added to the pose graph later can be inferred based on the camera pose of the node added to the pose graph earlier. If the image similarity of the two nodes is high, but there is no edge between them, it means that the cumulative error of the camera pose is relatively obvious at this time. As a result, the images that should have a common view relationship have a low matching degree of the solved three-dimensional points due to the large camera pose error, thus not having a common view relationship, and thus no edge is established between the two nodes.
[0126] For example, the image of the target node and the image of any existing node are subjected to image similarity detection or feature point matching to determine whether the image content of the two images is similar. If the two images are determined to be the same image or highly similar images, it indicates that there is a high probability that the two images were taken at the same or similar location in the target location using the same or similar camera position and orientation. At this time, based on whether there is a common view relationship between the two images, it can be further determined whether the camera pose corresponding to the two images has drifted due to cumulative error. If drift is determined to have occurred, it can be determined that the existing nodes contain nodes that have a loop relationship with the target node.
[0127] Since the 3D reconstruction task is not a real-time reconstruction task, there is no temporal correlation between the images, so it is difficult to trigger loop optimization based on the time series between the images. Based on this, the 3D reconstruction method provided in the embodiment of the present application, based on the above-mentioned camera pose correction method, can correct the estimated camera pose through loop optimization. The corrected camera pose is then used as one of the inputs for 3D reconstruction, and the target location is reconstructed in 3D. This can improve the accuracy of the 3D reconstruction result while improving the accuracy of the camera pose.
[0128] Figure 5 This is a schematic diagram of a trajectory curve in a three-dimensional reconstruction provided in an embodiment of the present application, such as Figure 5 As shown in the figure, the red curve is the camera pose trajectory curve formed by each camera pose corresponding to each image involved in 3D reconstruction. Figure 5In Figure (a), the starting point on the left side of the lower right corner of the trajectory curve is the starting point of the trajectory curve. By gradually recursively deducing the position of each camera and constructing a three-dimensional model, the end point on the right side of the trajectory curve can be gradually approached. Figure 5 The middle figure (b) shows the divergence of the trajectory curve caused by not eliminating the cumulative error of the camera pose. It can be seen that due to the drift of the camera pose and the 3D point cloud data, the trajectory start and end points corresponding to the same position in the 3D scene cannot be closed in the trajectory curve. In other words, the 3D point cloud data cannot be aligned, resulting in a decrease in the accuracy of the 3D reconstruction result. Figure 5 The figure (c) on the right side is the trajectory curve corresponding to each optimized camera posture after the camera posture is corrected based on the camera posture correction method of the above embodiment. It can be seen that by correcting each camera posture, the trajectory starting point and trajectory end point that should be closed can be loop closed (Loop Closure), and the camera postures corresponding to other nodes can be optimized to reduce the error of each camera posture, and then the three-dimensional point cloud data formed before optimization can be recalculated to obtain more accurate three-dimensional point cloud data, thereby improving the accuracy of the three-dimensional reconstruction results. This method can realize loop detection and correction of three-dimensional reconstruction, and after loop detection and correction, the effect of reducing or eliminating cumulative errors can be achieved, greatly improving the accuracy of the camera posture trajectory, such as Figure 5 As shown in Figure (c), the divergent trajectory curve is closed after correction. In addition, this method can reduce the data acquisition cost for indoor and outdoor 3D reconstruction by eliminating the need for sensors other than vision.
[0129] In addition, the three-dimensional reconstruction method provided in the embodiment of the present application has better stability than reconstruction algorithms based on neural networks, such as Neural Radiance Fields (NerF) and Dense and Unconstrained Stereo 3D Reconstruction (DUSt3R), because it involves less uncertainty brought about by the training and prediction process of neural networks. It can provide more reliable three-dimensional reconstruction results. Moreover, since this method has relatively low requirements for computing resources such as video memory, it can handle the three-dimensional reconstruction of larger target places.
[0130] In some embodiments, connecting the target node to the common view node can be achieved by:
[0131] If at least two of the existing nodes have a co-view relationship with the target node, then an edge is established between the target node and each node with a co-view relationship.
[0132] If the existing nodes include nodes that have a loop relationship with the target node, when optimizing the current pose graph based on the loop relationship, it can be: among the nodes that already exist in the pose graph and do not have a common view relationship with the target node, if there is a node whose corresponding image has a similarity with the image of the target node greater than a preset similarity threshold, then the node is regarded as a node with a loop relationship with the target node, and the current pose graph is optimized.
[0133] by Figure 3 For example, in the process of generating the pose graph, node 322 is newly added at the current moment. Node 322 is the target node. Target node 322 has a co-view relationship with both node 306 and node 305. Node 306 and node 305 are both co-view nodes of target node 322. An edge can be connected between target node 322 and node 306, and an edge can be connected between target node 322 and node 305. Based on this, all nodes with a co-view relationship can be constrained through connected edges. In addition, for each node in the pose graph that already exists and is not connected to the target node 322, it can be determined whether a node with a loop relationship with the target node 322 is included.
[0134] Exemplarily, the purpose of determining whether nodes with loop relationships are included can be understood as determining, for the current target node, whether the camera poses in the pose graph at this time have accumulated a large cumulative error, so as to determine whether loop optimization should be triggered to correct the camera poses of the current nodes.
[0135] If two images are taken at the same or similar locations in a target location using the same or similar shooting directions, the true camera poses of the two images should be the same or similar, and the strength of the common view relationship between the two images should also be high. If, through determination, it is determined that the common view relationship between the two is weak or non-existent, this indicates that the determination of the common view relationship is incorrect. This error is caused by a large error in the camera pose of at least one of the two images, which causes the 3D point position data of the feature points determined based on the erroneous camera pose to drift, resulting in an error in the determined common view relationship. Based on the above analysis, it can be seen that in the pose graph, when two identical or highly similar images do not have a common view relationship, it indicates that there are certain errors in the camera poses of the pose graph. In this case, corrections should be made to improve the accuracy of the camera poses.
[0136] Based on this, through the similarity and co-viewing relationship between the images corresponding to the nodes, nodes with loop relationships can be effectively detected, and it can be timely and accurately judged whether the current pose graph should be optimized based on the loop relationship to avoid further accumulation of errors, thereby improving the accuracy of three-dimensional reconstruction.
[0137] In some embodiments, when the camera pose of each node in the current pose graph is corrected based on the loop relationship, a loop relationship edge can be established between two nodes with a loop relationship, and the camera pose of each node in the pose graph is corrected with the constraint of minimizing the relative pose represented by the loop relationship edge and the residual of each edge in the pose graph to obtain an optimized pose graph. The residual of the edge is used to represent the change in the edge before and after optimization.
[0138] Specifically, when adding a target node to the pose graph, an edge is connected between the target node and the common-view node based on the common-view relationship. This edge can constrain the relative pose of the two connected nodes. Since there is no edge between two nodes with a loop relationship, in order to establish a constraint between the two nodes with a loop relationship, a loop relationship edge can be established between the two nodes with a loop relationship. This loop relationship edge represents the difference in the camera pose of the two nodes with a loop relationship, that is, the relative pose of the two nodes.
[0139] by Figure 3 For example, if the current target node 322 has a loop relationship with the existing node 311, then the loop relationship edge (such as Figure 3 The two nodes are connected by a dotted line (the dashed line between them) to form a constraint on the relative pose of the two nodes. The camera pose of each node in the pose graph is corrected by minimizing the relative pose represented by the loop relationship edge and the residual of each edge in the pose graph. The optimized pose graph can be expressed as follows:
[0140]
[0141] Among them, ||(T′ Neighbor ) -1 T Neighbor || 2 represents the residual of the edge, T′ Neighbor represents the value of the edge before optimization, T Neighbor represents the value of the edge after optimization, ||T′ PnP T loop || 2 Indicates the change in the loop relationship edge before and after optimization, T′ PnP Represents the value of the loop relationship edge before optimization, T loop Indicates the value of the loop relationship edge after optimization.
[0142] It can be understood that the relative pose between two nodes with a common view relationship is a relatively accurate estimate. Therefore, by taking minimizing the residual of the edge as one of the constraints, when optimizing the camera pose, the relative pose of the two nodes with a common view relationship can be kept unchanged or changed slightly. Then, when optimizing each camera pose, the local shape of the formed camera pose trajectory curve can be maintained as much as possible, so as to achieve the purpose of fine-tuning the camera pose while slightly changing the shape of the formed trajectory curve, so as to eliminate the error of the camera pose.
[0143] The above expression can be used as the objective function to minimize the sum of the residuals of the loop relationship edge table plus all edges. This allows the global camera pose to be optimized while pulling or closing the camera pose of two nodes with a loop relationship and maintaining the shape of the local trajectory curve formed by the relative poses of each common view. When optimizing the pose of each camera, existing optimization algorithms can be combined, such as nonlinear least squares methods such as General Graph Optimization (g2o), Ceres Solver, and Georgia Tech Smoothing and Mapping (GTSAM).
[0144] In practical applications, since the images of two nodes with a loop relationship are not necessarily the same image, the loop relationship edge established can be a smaller value to represent the relative pose between two images with similar real camera poses. PnP It can be a value that tends to zero, and the degree of its tendency to zero may depend on the closeness of the real camera poses of the two images. When it is determined that the two images are images captured under the same real camera pose, T′ PnP can be set to zero. Similarly, T loop It is expected to be a smaller value, or a value close to zero, or zero, and the size of this value is also related to the difference in the actual camera pose between the images.
[0145] In this embodiment, the camera pose of each node in the pose graph is optimized with the constraint of minimizing the residual of the loop relationship edge and each edge in the pose graph. The global camera pose can be optimized and adjusted through loop optimization to improve the accuracy of each camera pose.
[0146] For example, when performing three-dimensional reconstruction of a target location based on a pose graph, it can be achieved through an SFM loop detection module. Figure 6 A flow chart of the SFM loop detection module provided in the embodiment of the present application is shown as follows: Figure 6As shown in the figure, incremental registration means traversing each image and adding corresponding nodes to the pose graph. During the incremental registration process, the loop relationship between the target node and the existing nodes is determined, that is, loop detection is performed. If the loop detection result does not contain any nodes with a loop relationship, that is, there is no loop image for the current image, loop optimization is not triggered and incremental registration continues. If the loop detection result contains nodes with a loop relationship, that is, there is a loop image for the current image, loop optimization is triggered and operations such as merging 3D points and optimizing the pose graph are performed. Since camera pose recursion and 3D point cloud data generation can be regarded as two interrelated branch tasks, the 3D point cloud data can be corrected synchronously when the loop is detected and pose optimization is performed.
[0147] Among them, merging 3D points can be understood as correcting the 3D point position data of each 3D point of the corresponding image of two nodes with a loop relationship when determining to perform loop optimization, that is, merging the 3D point position data of the same feature points of the corresponding images of two nodes with a loop relationship. The merging method can be set according to the actual situation. For example, the 3D point position data of the same feature points of the image of the existing node can be replaced with the 3D point position data of the corresponding feature points in the image of the target node. Pose graph optimization can be to optimize the pose graph based on the loop relationship, that is, to correct the camera pose corresponding to each node. After performing pose graph optimization, each camera pose can be globally optimized, that is, global bundling adjustment (GlobalBA). After optimizing each camera pose, the incremental registration step can be continued.
[0148] Alternatively, camera pose estimation and 3D point cloud data generation can be performed in two sequential stages. For example, after optimizing the camera poses for each image, 3D point cloud data generation can be performed to obtain the 3D reconstruction results. This avoids multiple revisions to the 3D point cloud data by performing point cloud data generation when no further camera pose optimization is required, thus saving computing resources to a certain extent.
[0149] In practical applications, methods that suppress divergence through loop closure optimization rely on the presence of some loop images in the image dataset. When the image dataset lacks loops or lacks sufficient loops, divergence is difficult to effectively eliminate. In this case, eliminating divergence through purely visual observation is difficult and requires the introduction of other observation sources. Therefore, an SFM backend that supports multi-source joint optimization can be designed to achieve multi-source joint optimization. In 3D reconstruction tasks, integrating and utilizing multiple different types of data sources or information sources for multi-source joint optimization can optimize overall task performance.
[0150] In some embodiments, the three-dimensional reconstruction method also includes: obtaining measurement data corresponding to the image, the measurement data being used to characterize the posture of the shooting device collected when the image is taken; in the process of generating a pose graph, based on the measurement data of two nodes connected by any edge, determining the relative relationship between the postures of the shooting devices of the two nodes; accordingly, optimizing the current pose graph based on the loop relationship, including: optimizing the current pose graph based on the loop relationship and the relative relationship corresponding to each edge.
[0151] Measurement data can be data collected by sensors on a camera and used to characterize the camera's pose. The data type of the measurement data depends on the sensor type and can include one or more types of data. To effectively constrain the camera pose of each node across different dimensions, during pose graph generation, a relative relationship reflecting the camera's pose can be established between the two nodes based on the measurement data of any two nodes connected by an edge, forming corresponding constraints. This relative relationship can be understood as the difference between the measurement data corresponding to the two nodes.
[0152] When optimizing the current pose graph based on loop relationships, the current pose graph can be optimized based on the loop relationships and the relative relationships corresponding to each edge. For example, when determining nodes with loop relationships, the loop relationship edges are set to a reasonably small value, and the optimization solution is performed with the constraint of minimizing the change in each relative relationship before and after optimization. This can achieve optimization of the current pose graph based on the loop relationships and the relative relationships corresponding to each edge.
[0153] According to the method provided by this application, when there is no common view between two highly similar images, it is considered necessary to perform loop optimization. At this time, the pose graph optimization construction method without time-series dependency proposed by this method can be used to optimize the pose graph, which greatly eliminates the cumulative error between the two frames of image. Moreover, the pose of other related frames is optimized through the relative pose constraints between frames, which plays the role of reducing overall divergence and improving reconstruction accuracy. In addition, in response to the limitations of single-source vision, a multi-source joint optimization backend is introduced. For example, sparse position constraints such as GNSS (such as GPS, Beidou positioning, etc.) measurement data and manual control points are used to control the overall scale and heading divergence. Relative constraints such as inertial measurement units (IMUs) are used to improve overall accuracy.
[0154] In this embodiment, multi-source joint optimization can be achieved by integrating measurement data measured by non-visual sensors that are different from visual data. The measurement data can complement the visual information, making up for the defect of easy divergence caused by the poor robustness of a single visual signal source, and effectively improving the accuracy of the three-dimensional reconstruction results.
[0155] This embodiment can be applied to situations where only a few loop closure images exist among multiple images. For situations where there are no loop closure images among multiple images, optimization can be performed based solely on the relative relationships or on the relative relationships and the residuals of each edge, which can also eliminate or reduce the cumulative error in the camera pose to a certain extent.
[0156] Taking into account the differences in the ability of measurement data collected by different sensors to accurately represent the true posture, processing multiple types of measurement data in a weighted manner can improve the relative relationship of the measurement data to form more reliable and effective constraints. Optimization under effective constraints can improve optimization efficiency.
[0157] In some embodiments, the measurement data includes at least one type of measurement data of inertial measurement data, positioning measurement data, and augmented reality engine measurement data;
[0158] When determining the relative relationship between the camera postures of the two nodes based on the measurement data of the two nodes connected on any edge, specifically: when there are at least two types of measurement data, determine the relative relationship between the camera postures of the two nodes under the corresponding type for each type of measurement data; based on the weight values corresponding to each type of measurement data, weight the relative relationships corresponding to each type of measurement data to obtain the weighted relative relationships.
[0159] Specifically, the shooting device can usually be a camera, camera, mobile phone, tablet computer or other electronic device with integrated visual sensors. When these devices integrate sensors (such as gyroscopes, inertial measurement units IMU, positioning receivers, etc.) in addition to visual sensors, they can also collect measurement data of these non-visual sensors when shooting images, such as gyroscope data, GNSS positioning data, etc.
[0160] Inertial measurement data can be understood as measurement data collected by sensors that measure inertial-related data, such as attitude data collected by gyroscopes or inertial measurement units (IMUs). Positioning measurement data can be understood as data related to GNSS positioning signals and position measurement results measured by positioning receivers.
[0161] In this embodiment, when measurement data from non-visual sensors is present, this data can be added as constraints to the optimization problem and jointly optimized with the feature reprojection residuals to mitigate divergence and improve reconstruction accuracy. The addition of multiple types of measurement data allows for the establishment of multi-dimensional constraints. By weighting the relative relationships between each type of measurement data, the relative relationship between measurement data of varying accuracy can be adjusted to determine their relative strength within the overall constraint, improving the reliability of the multi-source constraints and enabling more accurate correction of each camera pose.
[0162] In some embodiments, when optimizing the current pose graph based on the loop relationship and the relative relationship corresponding to each edge, specifically: establishing a loop relationship edge between two nodes with a loop relationship, and optimizing the camera pose of each node in the pose graph with the constraint of minimizing the relative pose represented by the loop relationship edge, the residual of each edge in the pose graph, and the residual of each relative relationship, to obtain an optimized pose graph and optimized camera poses, the relative pose relationship is used to characterize the pose difference between the two nodes connected to the edge, the residual of the edge is used to characterize the change of the edge before and after optimization, and the residual of the relative relationship is used to characterize the change of the relative relationship before and after optimization.
[0163] The relative pose relationship represented by minimizing the loop relationship edge can unify the camera poses of two nodes with a loop relationship as much as possible, and with the minimization of the residuals of each edge in the pose graph and the residuals of each relative relationship as constraints, joint optimization can be performed under multiple effective constraints to obtain better optimization results.
[0164] In this embodiment, the camera pose of nodes with loop relationships can be corrected with the constraints of minimizing the relative pose relationships represented by the loop relationship edges, the residuals of each edge in the pose graph, and the residuals of each relative relationship. In addition, the camera pose of other nodes can be effectively globally optimized while maintaining the local trajectory shape as much as possible, thereby improving the effectiveness of the optimization.
[0165] When there is no loop closure image that triggers loop closure optimization in the multiple images corresponding to the target location, it is difficult to optimize the camera pose based on the loop closure relationship. In this case, the relative relationship formed by the measurement data can be used to perform periodic relative relationship optimization to avoid the divergence of the trajectory curve caused by not performing camera pose optimization for a long time when loop closure optimization is not possible.
[0166] In some embodiments, the method further includes: if the existing nodes do not include a node having a loop relationship with the target node, and the count value of the newly added node is less than a preset count threshold, the pose graph is not optimized and the count value is increased by one; if the existing nodes do not include a node having a loop relationship with the target node, and the count value of the newly added node is equal to the preset count threshold, the current pose graph is optimized based on the relative relationship of the shooting device poses to eliminate the accumulated error of each camera pose, and the count value is cleared after optimization.
[0167] A count value can be set for the number of added nodes to count the number of nodes added during the current counting cycle. If the count value of newly added nodes is less than the preset count threshold and loop optimization has not been triggered, the pose graph will not be optimized, that is, global camera pose optimization will not be performed. In this case, the count value will be continuously accumulated based on the number of newly added nodes. It should be noted that the preset count threshold is a judgment value for the count value and can be set according to the actual application scenario.
[0168] When loop optimization is not triggered and the accumulated count reaches a preset threshold, the current pose graph can be optimized based on the relative relationships of the camera poses. The relative relationships of the camera poses can be established based on various measurement data, and these relationships can serve as auxiliary constraints for the relative pose relationships between nodes. After optimization, the count can be reset to zero and the count can be restarted based on the number of newly added nodes after optimization.
[0169] For example, if the count threshold is set to 50, and loop closure optimization was not triggered when adding the first 50 nodes, global optimization can be performed based on the relative relationships established by the measurement data of each node to adjust the camera poses to reduce or eliminate the cumulative error. If new nodes are added after optimization, the count value will be accumulated starting from 1 based on the number of new nodes added. If the count value reaches 50 again and loop closure optimization is not triggered, global optimization can be performed again based on the relative relationships established by the measurement data of each node, and so on.
[0170] This approach solves the problem of being unable to globally optimize the poses of each camera due to the absence of loops in the nodes corresponding to the 2D image, thus enabling timely optimization of the pose graph. By setting a count value, each time the real-time accumulated count reaches a preset threshold, the global poses of each camera are forcibly optimized based on their relative relationships, thus avoiding the problem of continued error accumulation caused by long-term non-optimization.
[0171] Whether loop closure optimization can be triggered depends on whether there are loop closure images that can trigger loop closure optimization. At the same time, the number of loop closures performed also depends on the number of loop closure images that can trigger loop closure optimization in multiple images. Therefore, to perform loop closure optimization a reasonable number of times, the number of loop closure images that can trigger loop closure optimization can be increased in multiple images.
[0172] In some embodiments, the multiple images are captured by a photographing device at multiple locations in the target location, and the multiple images include two-dimensional images captured by the photographing device at least twice at the same location, and at least two shots correspond to the same or similar shooting directions.
[0173] It is understandable that when multiple images are taken at the same location in the target location using the same or similar shooting directions, multiple images with high similarity can be obtained. When generating the pose graph, these images may be added to corresponding nodes one after another. During the addition process, these nodes will have a high probability of triggering loop closure optimization to perform global camera pose optimization. Similarly, by taking multiple images at multiple locations using the same or similar shooting directions, more loop closure images that can trigger loop closure optimization can be obtained, allowing pose graph optimization based on loop closure relationships to be performed at different stages of updating the pose graph.
[0174] For example, when performing 3D reconstruction of a larger target location, multiple similar images that can trigger loop closure optimization can be captured in advance during the image preparation phase. For example, a predetermined shooting route is defined throughout the target location, and multiple loop points, such as loop points A, B, C, and D, are determined at equal or approximately equal intervals along the route. Images can be captured along the route or randomly, but multiple images must be captured at loop point A using the same or similar shooting orientation to obtain at least two images with the same or similar camera poses. Multiple images must also be captured in the same manner at loop points B, C, and D. Based on this, when generating the pose graph for 3D reconstruction of the target location, the images are shuffled. Loop closure optimization may be triggered at least once for each of the four images captured at loop points A, B, C, and D, enabling multiple rounds of camera pose correction.
[0175] Based on this, by adding loop points in long-distance acquisition, the optimization frequency of the pose graph can be increased to perform multiple loop optimizations. As a result, the pose of each camera can be optimized multiple times in the pose graph through each loop optimization, and the accumulated errors can be eliminated in multiple rounds, thereby obtaining highly accurate three-dimensional reconstruction results.
[0176] In practical applications, the accuracy and stability of GPS positioning are easily affected by GPS signal quality. Therefore, high-quality GPS positioning is difficult to achieve in urban canyon scenes or other environments with significant building obstruction. Visual relocalization is an important technology in computer vision and robotics. It involves re-determining a device's current position in a known environment using visual information. The core concept is to re-locate the device by identifying known features or landmarks in the image. Visual relocalization technology uses a camera to capture an image of the environment, matches it to a visual base map, and then calculates the pose. This technology provides decimeter-level positioning results and is immune to signal interference, addressing the pain point of low positioning accuracy in these scenarios. Visual relocalization requires pre-constructing a base map, namely, performing a 3D reconstruction of the target scene to obtain a point cloud model of the corresponding scene. However, due to cumulative errors, large-scale 3D reconstructions often exhibit divergence in track scale and heading, seriously affecting the accuracy of the 3D reconstruction results. Therefore, when visual relocalization is performed based on a less accurate 3D model, the positioning results are also less accurate.
[0177] To address the above issues, this application proposes a visual relocalization method that improves the accuracy of visual relocalization results by improving the accuracy of the three-dimensional model. The method comprises: obtaining a query image of a target location for which visual relocalization is requested; and, based on the query image, performing a calculation in a three-dimensional model corresponding to the target location to determine a positioning result corresponding to the query image. The three-dimensional model is a three-dimensional model obtained using the three-dimensional reconstruction method provided in any of the above embodiments.
[0178] Since the three-dimensional reconstruction method provided by each of the above embodiments includes a pose graph generated independently of the time sequence between images, and can correct the camera pose of each image in the presence of a loop image to eliminate the cumulative error of the camera pose and improve the accuracy of the camera pose. The problem that the trajectory curve of the camera pose cannot be closed due to divergence is solved, and the accuracy of the three-dimensional reconstruction result can be effectively improved. In addition, in order to solve the limitations of a single visual data source, multiple data sources are introduced into the three-dimensional reconstruction, and a multi-source joint optimization algorithm is proposed to improve the robustness of the algorithm and the accuracy of the three-dimensional reconstruction result. Therefore, since the method in each of the above embodiments can effectively improve the accuracy of the three-dimensional model of the target location, based on a more accurate three-dimensional model, a more accurate visual relocation result can be obtained, thereby improving the accuracy of the positioning result.
[0179] It should be understood that the execution order between the steps provided in the aforementioned method embodiments of the present application is only an example. Under the premise of being logical, some parallel steps can be executed serially, or in order to improve efficiency, some serial steps can be executed in parallel, or the execution order of some steps can be changed. This application does not limit this.
[0180] The present application also provides a device for correcting camera posture. Figure 7 A schematic diagram of the structure of the camera posture correction device provided in the embodiment of the present application is shown as follows: Figure 7 As shown, the camera posture correction device 700 includes:
[0181] An acquisition module 701 is configured to acquire, from an image dataset, an image that satisfies a preset co-viewing condition with the current image as a co-viewing image, wherein the image dataset includes one or more two-dimensional images captured at the same target location as the current image;
[0182] The acquisition module 701 is further configured to acquire a relative pose between a camera that captures a current image and a camera that captures a co-viewing image as a co-viewing relative pose;
[0183] The acquisition module 701 is further configured to acquire an image from the image dataset that satisfies a preset similarity condition with the current image, and if the image is not the same as the co-viewed image, use the image as a loopback image;
[0184] The acquisition module 701 is further configured to acquire a relative pose between a camera that captures a current image and a camera that captures a loop-closed image as a loop-closed relative pose;
[0185] Correction module 702 is used to optimize the gap between the loop relative pose and the common view relative pose, and to minimize the relative pose change of the edges of the pose graph as a constraint, to correct the camera pose corresponding to each node in the pose graph to eliminate the cumulative error of the camera pose. The pose graph includes: nodes and edges, where a node corresponds to an image in an image dataset and the node records the camera pose of the camera that took the image. An edge connects two nodes, and there is a common view relationship between the images corresponding to the two nodes. The edge records the relative pose of the cameras recorded by the two nodes.
[0186] Optionally, the camera pose correction device 700 further includes a generation module, which is used to generate a pose graph based on the co-viewing relationship between images in the image data set and the camera pose of the camera that takes each image.
[0187] Optionally, the generation module is specifically used to: select any two images with a co-viewing relationship from the image dataset, construct two nodes of the pose graph based on the camera pose of the camera that took the any two images, and construct an edge between the two nodes in the pose graph with the relative pose of the camera recorded by the two nodes; use the two nodes as seed nodes respectively, search for other images that have a co-viewing relationship with the image of the seed node in the remaining images in the image dataset, and construct new nodes and edges, use the new nodes as new seed nodes to continue searching and constructing new nodes and edges in the pose graph, until the nodes and edges corresponding to all images with a co-viewing relationship in the image dataset are constructed in the pose graph.
[0188] Optionally, for a seed node, if at least two other images having a common-view relationship with the image of the seed node are found in the remaining images in the image data set, the generation module is further used to: obtain an image with the highest common-view relationship strength from at least two other images having a common-view relationship with the image of the seed node; wherein the common-view relationship strength is used to characterize the degree of matching of the three-dimensional points corresponding to the feature points of the two images; the generation module is specifically used to: construct a new node in the pose graph based on the camera pose of the camera that took the image with the highest common-view relationship strength, and construct the new node and the edge of the seed node in the pose graph with the relative pose of the camera recorded by the new node and the seed node.
[0189] The camera posture correction device provided in the embodiment of the present application can be used to implement the technical solution of the camera posture correction method provided in any of the above embodiments of the present application. Its implementation principle and technical effect are similar, and this embodiment will not be repeated here. The embodiment of the present application also provides a three-dimensional reconstruction device, Figure 8 A schematic diagram of the structure of a three-dimensional reconstruction device provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the 3D reconstruction device 800 includes:
[0190] An optimization module 801 is configured to optimize the camera pose of an image captured at a target location based on the camera pose correction method of any of the above embodiments;
[0191] The reconstruction module 802 is used to use the camera pose after image optimization as one of the inputs of three-dimensional reconstruction to perform three-dimensional reconstruction of the target location.
[0192] The three-dimensional reconstruction device provided in the embodiment of the present application can be used to execute the technical solution of the three-dimensional reconstruction method provided in any of the above embodiments of the present application. Its implementation principle and technical effects are similar, and will not be repeated here in this embodiment.
[0193] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 9As shown, the electronic device of this embodiment may include: at least one processor 901; and a memory 902 communicatively connected to the at least one processor; wherein the memory 902 stores instructions that can be executed by the at least one processor 901, and the instructions are executed by the at least one processor 901 to enable the electronic device to execute the method described in any of the above embodiments.
[0194] Optionally, the memory 902 may be independent or integrated with the processor 901 .
[0195] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the aforementioned embodiments and will not be described in detail here.
[0196] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in any of the above embodiments is implemented.
[0197] An embodiment of the present application further provides a computer program product, including a computer program, which implements the method described in any of the aforementioned embodiments when executed by a processor.
[0198] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, other division methods may be used. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not implemented.
[0199] The above-mentioned integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform some steps of the method described in each embodiment of the present application.
[0200] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include RAM (Random Access Memory), and may also include NVM (Non-Volatile Memory), such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0201] The storage medium may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0202] An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an application-specific integrated circuit (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0203] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0204] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0205] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0206] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A method for correcting camera posture, characterized in that: include: Acquire, from an image dataset, an image that satisfies a preset co-viewing condition with the current image as a co-viewing image, wherein the image dataset includes one or more two-dimensional images captured at the same target location as the current image; Acquire a relative pose between a camera that captures the current image and a camera that captures the co-viewing image as a co-viewing relative pose; Acquire an image from the image dataset that satisfies a preset similarity condition with the current image, and if the image is not the same as the co-viewed image, use the image as a loop closure image; Obtaining a relative pose between a camera that captures the current image and a camera that captures the loop-closed image as a loop-closed relative pose; With the optimization goal of narrowing the gap between the loop relative pose and the common view relative pose, and with the constraint of minimizing the relative pose change of the edges of the pose graph, the camera pose corresponding to each node in the pose graph is corrected to eliminate the cumulative error of the camera pose, wherein the pose graph includes: nodes and edges, a node corresponds to an image in the image dataset and the node records the camera pose of the camera that took the image, an edge connects two nodes, there is a common view relationship between the images corresponding to the two nodes, and the edge records the relative pose of the cameras recorded by the two nodes.
2. The method according to claim 1, characterized in that Also includes: A pose graph is generated based on the co-viewing relationship between images in the image dataset and the camera pose of a camera that takes each image.
3. The method according to claim 2, characterized in that Generating a pose graph based on the co-viewing relationship between images in the image dataset and the camera pose of a camera that captures each image, including: Selecting any two images with a co-viewing relationship from the image dataset, constructing two nodes of a pose graph based on the camera poses of the cameras that captured the two images, and constructing an edge between the two nodes in the pose graph based on the relative poses of the cameras recorded by the two nodes; Taking the two nodes as seed nodes respectively, searching for other images that have a co-viewing relationship with the image of the seed node in the remaining images in the image dataset, and constructing new nodes and edges, and continuing to search and construct new nodes and edges in the pose graph with the new nodes as new seed nodes, until the nodes and edges corresponding to all images with a co-viewing relationship in the image dataset are constructed in the pose graph.
4. The method according to claim 3, characterized in that For a seed node, if at least two other images having a co-viewing relationship with the image of the seed node are found among the remaining images in the image dataset, the method further includes: Obtaining an image with the highest common-view relationship strength from at least two other images that have a common-view relationship with the image of the seed node; wherein the common-view relationship strength is used to represent the degree of matching between the three-dimensional points corresponding to the feature points of the two images; The construction of new nodes and edges is specifically as follows: A new node in the pose graph is constructed based on the camera pose of the camera that captured the image with the highest co-viewing relationship strength, and an edge between the new node and the seed node in the pose graph is constructed using the relative pose of the camera recorded by the new node and the seed node.
5. A three-dimensional reconstruction method, characterized in that: include: Optimizing the camera pose of the image captured at the target location based on the method described in any of claims 1 to 4; The camera pose after image optimization is used as one of the inputs of three-dimensional reconstruction to perform three-dimensional reconstruction of the target location.
6. A camera posture correction device, characterized in that: include: an acquisition module, configured to acquire, from an image dataset, an image that satisfies a preset co-viewing condition with the current image as a co-viewing image, wherein the image dataset includes one or more two-dimensional images captured at the same target location as the current image; The acquisition module is further configured to acquire a relative posture between a camera that captures the current image and a camera that captures the co-viewing image as a co-viewing relative posture; The acquisition module is further configured to acquire an image from the image dataset that satisfies a preset similarity condition with the current image, and if the image is not the same as the co-viewed image, use the image as a loop closure image; The acquisition module is further configured to acquire a relative posture between a camera that captures the current image and a camera that captures the loop-closed image as a loop-closed relative posture; A correction module is used to correct the camera pose corresponding to each node in the pose graph with the optimization goal of narrowing the gap between the loop relative pose and the common view relative pose, and with the minimum relative pose change of the edge of the pose graph as the constraint item, so as to eliminate the cumulative error of the camera pose, wherein the pose graph includes: nodes and edges, a node corresponds to an image in the image dataset and the node records the camera pose of the camera that took the image, an edge connects two nodes, there is a common view relationship between the images corresponding to the two nodes, and the edge records the relative pose of the cameras recorded by the two nodes.
7. A three-dimensional reconstruction device, characterized in that: include: An optimization module, configured to optimize the camera pose of the image captured at the target location based on the method described in any one of claims 1 to 4; The reconstruction module is used to use the camera pose after the image optimization as one of the inputs of three-dimensional reconstruction to perform three-dimensional reconstruction of the target place.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the electronic device to execute the method according to any one of claims 1 to 4 or the method according to claim 5.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method according to any one of claims 1 to 4 or the method according to claim 5 is implemented.
10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 4 or the method according to claim 5.