Camera pose drift detection and correction
By employing geometric-based data analysis and correction techniques, the method addresses the issue of pose data drift in AR systems, ensuring accurate 3-D modeling and measurement in environments with geometric constraints.
Patent Information
- Application Number
- PCT/US2025/026468
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-26
- Filing Date
- 2025-04-25
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional visual inertial odometry systems in mobile devices, such as ARKit or ARCore, suffer from rapid and accumulating pose data drift, leading to inaccurate 3-D modeling, particularly in environments with geometric constraints, which affects measurement and design accuracy.
A method that utilizes geometric-based data from detected linear features, compares it to predetermined environmental geometry, detects drift values, and applies corrections to pose data using a set of processors to create a corrected 3-D model by grouping images based on drift error, employing geometric transforms to align model elements with image elements.
The method effectively corrects camera pose drift, ensuring accurate 3-D modeling and measurement by aligning model elements with image elements, thereby improving the reliability of 3-D models created using mobile devices.
Smart Images

Figure US2025026468_30102025_PF_FP_ABST
Abstract
Description
Camera Pose Drift Detection and CorrectionTECHNICAL FIELD
[0001] The disclosed implementations relate generally to three-dimensional (3-D) reconstruction and more specifically to systems and methods for 3-D modeling, drift detection, and drift correction for visual inertial odometry (VIO).BACKGROUND
[0002] Three-dimensional (3-D) models and visualization tools can produce significant cost savings when applied to building contexts. For example, a series of images may be captured by moving in an environment, such as the exterior or interior of a building, to provide image data of the environment used to generate a 3-D model. Accurate 3-D models can be used, for instance, to estimate and plan during a project. For example, with near real-time feedback, contractors may capture images of a home and have a 3-D model created from the images. The 3-D model may be used to provide related data, such as measurements, in order to provide a customer with an instant quote for a remodeling project. Interactive design tools using the 3-D model allows users to view it under various conditions, such as different lighting, different weather conditions, and the like. A contractor may select regions of the 3-D model data in an interactive design tool, e.g., a house roof, to get accurate measurements used to automatically estimate cost and amount of material required.
[0003] Thus, it is useful to be able to build an accurate 3-D model to supply users with the ability to design projects, produce estimates, etc. There are many approaches to creating 3-D models from image data. One approach used in numerous industries is visual inertial odometry (VIO) in which image data is supplemented with sensor data (such as from an inertial measurement unit, IMU), that indicates the camera position and orientation for the image, thus aiding determination of the location of image features in 3-D space used for modeling. Another approach assumes that some features of the environment have geometric constraints, such as lines or right-angles (Manhattan constraints), which define image features with a presumed geometry that can be used to locate the images features in the 3-D space.SUMMARY
[0004] A 3-D model may be built using relatively inexpensive equipment, for example a mobile device with one or more cameras (“mobile device”). For example, a mobile device in the form of a smartphone, runs an operating system such as iOS or Android, which provide augmented reality (AR) modeling tools, for example ARKit or ARCore respectively. The AR tools use VIO systems to provide camera pose data, but the VIO systems have inherent inaccuracies which drift over time, making any 3-D model built using this data inaccurate, and so its use, such as measurement, cost estimation, animating scenes for illustrating a new design, etc. Thus, while such devices greatly increase the convenience of the image capture procedure, the relative pose data of mobile devices may be unreliable in some situations. For example, conventionally these VIO systems generate pose data for indicating camera orientation that is initially quite accurate, but it tends to drift quite rapidly, over seconds. Further, the pose data from these conventional systems is relative, increasing the effective drift as the sequence of images progresses. Thus, any modeling of an environment will suffer, particularly as the drift accumulates over time.
[0005] Separately, while it is known to use geometric constraints of image features to determine an accurate 3-D location for the features, such as the angles between groups of lines, this provides relative data reliant on detecting the features. Thus, even if the environment is constrained, such as an indoor environment with many regular features including parallel and perpendicular lines (formed from walls, edges of doors, etc.), a process that uses fixed geometric relationships to create a 3-D model will be reliant on the underlying geometric assumptions and any noise in the captured image. Further, detecting and measuring geometry is process-intensive.
[0006] Accordingly, an embodiment provides a method, comprising: obtaining, using a set of one or more processors, pose data derived from an augmented reality (AR) system of a mobile device indicative of a position and orientation of a camera of the mobile device in three- dimensional (3-D) space; the pose data being associated with the capture of one or more images of an environment using the camera of the mobile device; determining, using the set of one or more processors, geometric-based data for respective ones of the images based on detected linear features and the pose data; comparing, using the set of one or more processors, the geometric-based data to a predetermined geometry applicable to the environment; employing,using the set of one or more processors, the comparison to detect a drift value, wherein the drift value is based on an increasing error in the pose data; grouping, using the set of one or more processors, the images into first and second subsets based on a drift error value associated with respective images; applying, using the set of one or more processors, different corrections to the pose data for images of the first and second subsets; and creating, using the set of one or processors, a 3-D model of the environment based on image features and corrected pose data for respective images.
[0007] In an embodiment, the grouping comprises sorting the first and second subsets based on determining that a change in the geometric-based data is statistically significant when compared to a monotonic increasing function.
[0008] In an embodiment, a first one of the different corrections comprises rotating one of three axes of respective image data by a predetermined or computed factor.
[0009] In an embodiment, a first one of the different corrections comprises indicating to a user a need for intervention.
[0010] In an embodiment, the environment is a building interior.
[0011] In an embodiment, the model comprises one or more measurements of a real-world object.
[0012] In an embodiment, the method comprises providing the model of the building interior to a design tool.
[0013] In an embodiment, the geometric-based data is line-based data, and the determining the geometric-based data comprises determining a Manhattan basis for the respective ones of the images.
[0014] In an embodiment, determining the Manhattan basis comprises, for each respective image, identifying a minimum set of lines of the respective image useable for determining the Manhattan basis.
[0015] In an embodiment, the minimum set of lines comprises a first line in a first direction, and one or more lines in a second direction, different from the first direction.
[0016] In an embodiment, the minimum set of lines of the respective image are identified using a line detection process.
[0017] In an embodiment, the minimum set of lines of the respective image are identified using a line markup process, and the line markup process accepts user input via a graphical user interface (GUI) to mark one or more lines within the image.
[0018] In another aspect, a computer system includes one or more processors, memory, and one or more programs stored in the memory. The programs are configured for execution by the one or more processors. The programs include instructions for performing any of the methods described herein.
[0019] In another aspect, a non-transitory computer readable storage medium stores one or more programs configured for execution by one or more processors of a computer system. The programs include instructions for performing any of the methods described herein.
[0020] The foregoing is a summary and is not intended to be in any way limiting. For a better understanding of the example embodiments, reference can be made to the detailed description and the drawings. The scope of the invention is defined by the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic diagram of a computing system for drift detection and correction, in accordance with some implementations.
[0022] Figure 2 is a block diagram of a device capable of capturing images, in accordance with some implementations.
[0023] Figure 3 is a schematic diagram of example drifts for camera poses, according to some implementations.
[0024] Figures 4A-4I illustrate examples of camera poses with drift errors, modeling, and corrections, according to some implementations.
[0025] Figure 5 is a flowchart of an example method for drift detection and correction, in accordance with some implementations.
[0026] Figures 6 shows an example user interface for determining Manhattan axis, according to some implementations.
[0027] FIG. 7 illustrates an example method according to some implementations.DESCRIPTION OF IMPLEMENTATIONS
[0028] Reference will now be made to various implementations, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments and the described implementations. However, the claims may be practiced without these specific details or in alternate sequences or combinations. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to unnecessarily obscure aspects of the implementations.
[0029] Figure l is a block diagram of a computer system 100 for detecting and correcting drift in camera poses, in accordance with some implementations. In some implementations, the computer system 100 includes mobile device(s) such as image capture device(s) 104, and a computing device 108.
[0030] An image capture device 104 communicates with the computing device 108 through one or more networks 110. The image capture device 104 provides image capture functionality (e.g., take photos) and communications with the computing device 108. In some implementations, the image capture device 104 is connected to an image preprocessing server system (not shown) that provides server-side functionality (e.g., preprocessing images, such as creating textures, storing environment maps (or world maps) and images and handling requests to transfer images) for any number of image capture devices 104.
[0031] In some implementations, the image capture device 104 is a computing device 108, such as desktops, laptops, smartphones, and other mobile devices, from which users 106 can capture images (e.g., take photos), discover, view, edit, or transfer images. In some implementations, the users 106 are robots or automation systems that are pre-programmed to capture images of the building structure 102 at various angles (e.g., by activating the image capture device 104). In some implementations, the image capture device 104 is a device capable of (or configured to) capture images and generate (or dump) world map data for scenes. In some implementations, the image capture device 104 is an augmented reality (AR) camera or a smartphone capable of performing the image capture and world map generation functions. In some implementations, the world map data includes (camera) pose data, tracking states, or environment data (e.g., illumination data, such as ambient lighting).
[0032] In some implementations, a user 106 walks inside a building structure 102 (e.g., a house), and takes pictures of rooms of the building structure 102 using the image capture device 104 (e.g., an IPHONE) at different poses (e.g., poses 112-2, 112-4, 112-6, and 112-8). Each pose corresponds to a different perspective or view of a room of the building structure 102 and its surrounding environment, including one or more objects (e.g., a wall, furniture within the room) within the building structure 102. Each pose alone may be insufficient to reconstruct a complete 3-D model of the rooms of the building structure 102, but the data from the different poses can be collectively used to generate the 3-D model or portions thereof, according to some implementations. In some instances, the user 106 completes a loop inside and / or around the building structure 102. In some implementations, the loop provides validation of data collected around and / or within the building structure 102. In some implementations, data collected at a pose is used to validate data collected at an earlier pose. For example, data collected at the pose 112-8 is used to validate data collected at the pose 112-2.
[0033] At each pose, the image capture device 104 obtains (118) images of the building structure 102, and / or data for objects (sometimes called anchors) visible to the image capture device 104 at the respective pose. For example, the image capture device 104 captures data 118-1 at the pose 112-2, the image capture device 104 captures data 118-2 at the pose 112-4, and so on. As indicated by the dashed lines around the data 118, in some instances, the image capture device 104 fails to capture images or cameras have a drift (described below). For example, the user 106 switches the image capture device 104 from a landscape to a portrait mode or receives a call. In such circumstances of system interruption, the image capture device 104 fails to capture valid data or fails to correlate data to a preceding or subsequent pose.
[0034] Although the description above refers to a single device 104 used to obtain (or generate) the data 118, any number of devices 104 may be used to generate the data 118. Similarly, any number of users 106 may operate the image capture device 104 to produce the data 118.
[0035] In some implementations, the data 118 is collectively a wide baseline image set, which is collected at sparse positions (or poses 112) inside the building structure 102. In other words, the data collected may not be a continuous video of the building structure 102 or its environment, but rather still images or related data with substantial rotation or translation between successive positions. In some implementations, the data 118 is a dense capture set, wherein the successive frames and poses 112 are taken at frequent intervals. Notably, in sparsedata collection such as wide baseline differences, there are fewer features common among the images and deriving a reference pose is more difficult or not possible. Additionally, sparse collection also produces fewer corresponding real-world poses and filtering these, as described further below, to candidate poses may reject too many real -world poses such that scaling is not possible.
[0036] In some implementations, the computing device 108 obtains drift-related data 224 via the network 110. Based on the data received, the computing device 108 detects and / or corrects drifts in camera poses for the building structure 102.
[0037] The computer system 100 shown in Figure 1 includes both a client-side portion (e.g., the image capture devices 104) and a server-side portion (e.g., a module in the computing device 108). In some implementations, data preprocessing is implemented as a standalone application installed on the computing device 108 or the image capture device 104. In addition, the division of functionality between the client and server portions can vary in different implementations. For example, in some implementations, the image capture device 104 uses a thin-client module that provides only image search requests and output processing functions, and delegates all other data processing functionality to a backend server (e.g., the server system 108). In some implementations, the computing device 108 delegates image processing functions to the image capture device 104, or vice-versa.
[0038] The communication network(s) 110 can be any wired or wireless local area network (LAN) or wide area network (WAN), such as an intranet, an extranet, or the Internet. It is sufficient that the communication network 110 provides communication capability between the image capture devices 104, the computing device 108, or external servers (e.g., servers for image processing, not shown). Examples of one or more networks 110 include local area networks (LAN) and wide area networks (WAN) such as the Internet. One or more networks 110 are, optionally, implemented using any known network protocol, including various wired or wireless protocols, such as Ethernet, Universal Serial Bus (USB), FIREWIRE, Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wi-Fi, voice over Internet Protocol (VoIP), Wi-MAX, or any other suitable communication protocol.
[0039] The computing device 108 or the image capture devices 104 are implemented on one or more standalone data processing apparatuses or a distributed network of computers. In someimplementations, the computing device 108 and the image capture device 104 are a single device. In some implementations, the computing device 108 and the image capture device 104 support real-time drift detection and / or correction (e.g., during capture). In some implementations, the computing device 108 and the image capture device 104 support off-line drift detection and / or correction (e.g., post capture). In some implementations, the computing device 108 or the image capture devices 104 also employ various virtual devices or services of third party service providers (e.g., third-party cloud service providers) to provide the underlying computing resources or infrastructure resources.
[0040] Figure 2 is a block diagram illustrating a representative image capture and visual inertial odometry device 104 that is capable of capturing images (or taking photos) of building structures 102 (e.g., a house) and performing visual inertial odometry from which image related data 258 is extracted, in accordance with some implementations. The image capture device 104, typically, includes one or more processing units (e.g., CPUs or GPUs) 122, one or more network interfaces 240, memory 244, optionally display 242, optionally one or more cameras and / or sensors 238 (e.g., IMUs), and one or more communication buses 236 for interconnecting these components (sometimes called a chipset).
[0041] Memory 244 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid state memory devices; and, optionally, includes nonvolatile memory, such as one or more magnetic disk storage devices, one or more optical disk storage devices, one or more flash memory devices, or one or more other non-volatile solid state storage devices. Memory 244, optionally, includes one or more storage devices remotely located from one or more processing units 122. Memory 244, or alternatively the non-volatile memory within memory 244, includes a non-transitory computer readable storage medium. In some implementations, memory 244, or the non-transitory computer readable storage medium of memory 244, stores the following programs, modules, and data structures, or a subset or superset thereof: an operating system 246; a network communication module 248 for connecting the image capture and visual inertial odometry device 104 to other computing devices (e.g., the computing device 108 or image-related data sources) connected to one or more networks 110 via one or more network interfaces 240 (wired or wireless); an image capture module 250 for capturing (or obtaining) images captured by the image capture device 104, including, but not limited to: a transmitting module 252 to transmit image-relatedinformation (similar to the transmitting module 218); and an image processing module 254 to post-process images captured by the image capture and visual inertial odometry device 104. In some implementations, the image capture module 250 controls a user interface on the display 242 to confirm (to the user 106) whether the captured images by the user satisfy threshold parameters for drift detection, drift correction, and / or generating 3-D representations. For example, the user interface displays a message for the user to move to a different location so as to capture two sides (or two rooms) of a building, or so that all sides (or all rooms) of a building are captured; a visual inertial odometry module that generates visual inertial odometry data (e.g., a state, pose and / or velocity of the image capture device 104 by using the image related data 258 and one or more Inertial Measurement Units (IMUs)); and / or a database of image-related data storing data for drift detection and / or correction.
[0042] Examples of the image capture and visual inertial odometry device 104 include, but are not limited to, a handheld computer, a wearable computing device, a personal digital assistant (PDA), a tablet computer, a laptop computer, a cellular telephone, a smartphone, an enhanced general packet radio service (EGPRS) mobile phone, a media player, a navigation device, a portable gaming device console, a tablet computer, a laptop computer, a desktop computer, or a combination of any two or more of these data processing devices or other data processing devices. In some implementations, the image capture and visual inertial odometry device 104 is an augmented-reality (AR)-enabled device that captures augmented reality maps (AR maps, sometimes called world maps). Examples include ANDROID devices with ARCore, or IPHONES with ARKit modules.
[0043] In some implementations, the image capture device 104 includes (e.g., is coupled to) a display 242 and one or more input devices (e.g., camera(s) or sensors 238). In some implementations, the image capture device 104 receives inputs (e.g., images) from the one or more input devices and outputs data corresponding to the inputs to the display for display to the user 106. The user 106 uses the image capture device 104 to transmit information (e.g., images) to the computing device 108. In some implementations, the computing device 108 receives the information, processes the information, and sends processed information to the display 116 or the display of the image capture device 104 for display to the user 106.
[0044] Figure 3 is a schematic diagram of example drifts for camera poses 300, according to some implementations. Camera poses from a VIO system, such as ARKit or ARCore, mayhave good initial relative poses. The camera poses may drift over time, and features and geometry of a scene observed from a drifted camera may be incorrectly placed relative to other cameras. This is illustrated in Figure 3 where the dots indicate image locations. The rectangles 312, 314, 316, 318, and 320 may each represent rooms in a house. The capture path (shown in dashed circles) corresponds to a path a user walks and a corresponding device takes. The data captured by the device is shown in full circles. Initially, the data that is captured by the device overlaps (e.g., at image location 302) with the path the device takes. Subsequently, the data that is captured by the device drifts relative to the path the device takes. In some embodiments, the drift may be due to an accumulation of errors from sensors, such as an IMU, that are not rectified by observations in images from cameras. In Figure 3, raw camera poses drift towards the end of the user trajectory as compared to the modeling poses. For example, camera pose aligns with modeling pose at image location 304, but drifts at image location 306. Similarly, initially camera pose and modeling pose align at image location 308, but there is a drift at image location 310.
[0045] Grouping
[0046] Given a set of camera poses, some implementations create a subset of camera poses including at least one camera pose for which drift is detected. Each group of camera poses may include camera poses with low relative drift enabling camera poses in the group to retain their relative camera poses and change their global position in a collective manner such that all the groups are consistent with one another. Camera poses within a same subset have a unique group offset transform applied to them, where the transform defines where the subset of camera poses should be placed in world space.
[0047] Example grouping logics are described herein. The number of groups can vary from one to the number of camera poses. If there is just one group, visual alignment can be difficult for a large number of camera poses in case there is significant drift towards the end of capture. If there are as many groups as camera poses, this may result in discarding any relative transformation between neighboring camera poses and trying to independently solve each of their extrinsic parameters.
[0048] During the drift correction process, some implementations use image markups to estimate the offset parameters for each subset. Some implementations use as few subsets as possible to reduce the number of image markups needed, and possibly reduce the computation,storage, and time needed. Some implementations look for acceptable alignment between model elements (e.g., geometry, points, lines) and image elements (e.g., points, lines). In some implementations, the drift correction process is independent of the modeling process.
[0049] Iterative Modeling and Drift Detection
[0050] In some implementations, modeling and camera pose correction are performed concurrently. For example, a user may first solve different parts of the geometry depending on an initial alignment of camera poses (e.g., AR camera poses). After this, the user may continue to augment the geometry, for example by growing the existing geometry or adding new geometry. If at any point a number of images associated with camera poses are not aligning well with the geometry, then that may be an indicator of camera pose drift. The geometry may be generated according to one or more camera poses of one or more groups of camera poses.
[0051] When a drift is detected, some implementations start creating a new camera pose group from that point onwards. In other words, some implementations detect a first drift in at least one camera pose based on an observed misalignment and create a new group of camera poses including the at least one drifted camera pose associated with the first drift and one or more other camera poses. In some implementations, the at least one drifted camera pose associated with the first drift is corrected, for example using a transformation. To correct for camera pose drift, users and / or the system draw markups to indicate correspondence between the model elements, such as geometry, points, and lines and their image space representation. The correspondence between the model elements and their image space representations may be used to generate the transform.
[0052] In some implementations, the at least one drifted camera pose associated with the first drift is corrected using the transformation, and the same transformation is applied to the rest of the camera poses in the group of camera poses including the at least one drifted camera pose associated with the first drift. In some implementations, once drift has been corrected for a camera pose, the drift correction is propagated to other camera poses that are in the same group, for example subsequent camera poses in the same group. Some implementations continue modeling, detect a second drift in at least one additional camera pose, create a new group of camera poses comprising the at least one additional drifted camera pose associated with the second drift and one or more other camera poses, correct the at least one additional drifted camera pose associated with the second drift, for example using a transformation, and applythe same transformation to the rest of the camera poses in the group of camera poses including the at least one drifted camera pose associated with the second drift. Because drift in later cameras is worse than earlier cameras, a drift correction to earlier cameras is unlikely to solve for drift in later cameras. In some implementations, this process is repeated until all parts of the environment subject to the capture have been modeled and there is acceptable alignment between the model elements and the image elements. The propagation step applies the same correction to all camera poses within the same group. This is because relative camera poses between nearby frames are quite accurate. A transform may be applied to a new group of camera poses that aligns a camera pose in the new group of camera poses to geometry observed by that camera pose. Suppose there is a group with three camera poses represented by rotation and translation matrices pairs, and suppose the group offset is estimated for the group. Propagation here refers to applying the same offset to all camera poses within the group. In some embodiments, the offset includes a rotation and a translation, an example of which is described above. In some embodiments, the offset includes only translation. In some embodiments, the offset includes only rotation. If a drift has been corrected using only a few camera poses and if the offset propagation looks good (e.g., by visual alignment check) on other images, then a user who is using a modeling software, and working with the images to create a model, can use the corrected and solved camera poses for modeling.
[0053] Example Grouping Methods
[0054] Figure 4A is a schematic diagram of example camera poses with drastic drift errors 400, according to some implementations. Camera poses for Visual Inertial Odometry (VIO) methods, such as ARKit or ARCore camera poses are prone to drift errors. This may happen as a slow drift in camera poses over time and length of the user trajectory. It may be a sudden large jump induced if a VIO system resets. Figure 4A shows a drastic drift 400 in camera poses, according to some implementations. A building may have several rooms or enclosures. In this example, the building has two rooms 402 and 406 , separated by a wall 404. A user (e.g., the user 106) or users may walk inside the building and capture, or take, images, or pictures (or record a video). Actual path is shown as a red dashed line 410 and camera position is shown as blue triangles 408. Camera poses 412 do not show a drift, whereas camera poses 414 have a drastic drift.
[0055] Figure 4B is a schematic diagram of example camera poses with gradual drift errors 450, according to some implementations. In contrast to Figure 4A, the camera poses 416 show a gradual drift (from the line 410).
[0056] Figure 4C is a schematic diagram of example camera poses after drift correction 452, according to some implementations. In some implementations, the process for correcting drift includes modeling, drift detection, and drift correction in an iterative way until all images are aligned with the model. Some implementations group camera poses (e.g., poses 418, poses 420, and poses 422) as a way of keeping camera poses with low relative drift together. This allows camera poses within the same group retain their relative poses and only change their global position (e.g., in relation to the line 410) and rotation and / or translation in a collective manner. The example shows the actual path 410 versus camera poses corrected in discrete steps for each group.
[0057] Figure 4D is a schematic diagram of example initial camera poses 454 for drift detection and correction, according to some implementations. The initial camera poses (e.g., poses 408) and later camera poses may be in different coordinate systems. For example, ARKit or ARCore world coordinate system (e.g., line 464) may be chosen randomly depending on a first image associated with a first camera pose of the poses 408. Hence, some implementations use the initialization process to orient the camera poses to the Manhattan modeling system. After this step, the same Manhattan offset can also be propagated to other camera poses and the system can model a portion of the building (e.g., a room) using this information.
[0058] Figure 4E is a schematic diagram of example initialization and modeling 456 for drift detection and correction, according to some implementations. After initialization (as described above in reference to Figure 4D), some implementations start modeling the geometry. The images associated with the poses 408 can be used for the modeling.
[0059] Figure 4F is a schematic diagram of example drift detection and correction 458, according to some implementations. For drift detection, beyond a certain point, the drift in camera poses may be significant such that the model elements may not be aligned with the images elements. For example, in Figure 4F, the drift after poses 426 may be significant (for poses 430) such that the model elements are not aligned with the image elements of images associated with poses 430. If model generation has used multiple prior images associated with multiple prior camera poses with sufficient baseline between them, then there is a highconfidence that the modeling is correct and it is the camera poses which have error. Confidence in the model can increase with more modeling and / or other quantitative metrics, such as covariance of model parameters. Once drift has been detected, some implementations align the model elements and the image elements of the image associated with the drifted camera pose, for example by reprojecting the model into the image associated with the drifted camera pose and adjusting one or more camera parameters of the drifted camera pose until the model elements of the model and the image elements of the image associated with the drifted pose align (e.g., have no or low reprojection error). This can be done by adding markup correspondences such as 3D to 2D point markups or 3D to 2D line markups. Some implementations optimize offset of an entire group of camera poses so as to minimize the alignment error between model elements and image elements of images associated with the group of camera poses.
[0060] Figure 4G is a schematic diagram of example modeling 460 parts of a model using corrected camera poses, according to some implementations. Once drift has been corrected, some implementations propagate the correction to one or more other camera poses, for example subsequent camera poses, and the corrected camera poses can be used for modeling new parts of the environment. Some implementations continue generating other parts of the model using corrected camera poses. For example in Figure 4G, the corrections in Figure 4F (sometimes referred to as corrected camera poses) are used when processing the poses 430, for generating the model elements of the environment that correspond to the wall 404 and the room 406.
[0061] Figure 4H is a schematic diagram of example iterative drift detection and correction 462, according to some implementations. In some implementations, modeling, drift identification and correction, are repeated iteratively. Some implementations continue generating model elements until alignment between the model elements and image elements of images associated with poses is bad (e.g., reproduction error is greater than 30 pixels). For example, a second part of the poses 430 in Figure 4H start to drift.
[0062] Figure 41 is a schematic diagram of example iterative drift detection and correction 464, according to some implementations. In some implementations, when alignment is bad, e.g., as defined by a predetermined factor, the system corrects for camera poses (e.g., by creating a new group of camera poses) and continues the modeling process iteratively. For example, in Figure 41, because the second part of the poses 430 of Figure 4H showed misalignment, thesystem splits the poses 430 of Figure 4H into two sets 436 and 438 and corrects for camera poses in the group 438 and continues the modeling.
[0063] Example Methods for Drift Detection and Correction
[0064] Figure 5 is a flowchart of an example method 500 for drift detection and drift correction, in accordance with some implementations. The method 500 is performed in a computing device. The method can be used to detect and / or correct drifts in camera pose data, for example based on visual inertial odometry.
[0065] The method includes obtaining (502) a set of images and a geometry for a building structure. Examples of geometry include a model, such as a 3D building model.
[0066] The method also includes detecting (504) a first misalignment between lines in the model and lines based on a first image. The method also includes, in response to detecting the first misalignment, correcting (506) a drift so that a reprojection error as determined from a first misaligned camera pose associated with the first image is minimized, including applying a geometric transform to the first misaligned camera pose and one or more other camera poses. In some examples, the one or more other cameras poses may include subsequent, for example temporally subsequent, camera poses with respect to the first misaligned camera pose. In some examples, the one or more camera poses may be in a same group as the first misaligned camera pose. In some implementations, detecting the first misalignment, correcting the drift, and / or propagating the drift, are performed (508) while building the geometry. In some implementations, the method further includes continuing modeling (510). The method includes detecting a second misalignment between lines in the model and lines based on a second image. The method also includes, in response to detecting the second misalignment, correcting a drift so that a reprojection error as determined from a second misaligned camera pose associated with the second image is minimized, including applying a geometric transform to the second misaligned camera pose and one or more other camera poses. In some examples, the one or more other camera poses may include subsequent, for example temporally subsequent, camera poses with respect to the second misaligned camera pose. In some examples, the one or more camera poses may be in a same group as the second misaligned camera pose.
[0067] Figure 6 shows an example user interface 600. User interface 600 may be used as part of a line markup process, for example for determining Manhattan axis according to some implementations. Some implementations determine vanishing points from a calibrated image.Figure 6 shows an example of a calibrated image 602 of a building structure. A calibrated image is an image where the focal length and principal point are known and lens distortion is assumed to have been corrected according to a pinhole model. Suppose a line, e.g., line 604, is provided. In some examples, a line may be detected using one or more line detection techniques. In some examples, the line may be provided as a markup, for example by a user. Images may include line, edges, corners of a room, intersections of planes, and the like. In Figure 6, the ceiling and the walls intersect on sides of the room and indicate lines. The line 604 helps define a first plane in 3D. Three points define a plane. The two ends of the line 604 (ends 1 and 2) define two of the points. The third point is a camera’s eye (e.g., optical center). Rays coming from the camera and intersecting the two ends of the line 604 can help define the plane. A Manhattan direction can be determined, for example assuming a rectangular room. By using two or more lines, e.g., lines 604, 605, and 606, their normals, and direction, a Manhattan axis for a line can be determined.
[0068] Some implementations automatically determine the Manhattan axis that corresponds to a line (or lines). An axis guesser 610 may be used to illustrate identification of an axis of a line, e.g., responsive to a user request or automated highlighting, selecting or identification thereof. Some implementations use an original AR pose to place the lines into a world space, find a plane normal in the world space, and then compute a value. In some implementations, the value is the XY magnitude, or lack thereof, that informs whether a line is close to the gravity vector. Otherwise, a line is in the horizontal direction, the first of which may be labelled and afterwards the system may use 45-degree thresholds to distinguish line directions.
[0069] Some implementations use a line segment detector (LSD) to detect lines in a calibrated image. After detecting the lines, the system can perform the computations to determine Manhattan basis for a 3D model of a building structure. An example of this is shown in Figure 6, according to some implementations. Some implementations provide an affordance (e.g., button) for a user to select automatic line segment detection (e.g., instead of, or in addition to, a user drawing the lines on the interface 600) to trigger the Manhattan basis calculations, one example of which is shown in the panel to the right of the image in the UI 600.
[0070] An example method can be used to detect and / or correct drifts in images obtained from visual inertial odometry, for example represented in pose data derived from an augmented reality (AR) system of a mobile device indicative of a position and orientation of a camera ofthe mobile device in three-dimensional (3-D) space. The method includes obtaining a plurality of images, a set of camera poses associated with the plurality of images, and a model for a building structure. In some implementations, the image corresponds to a rectangular room of the building structure. Manhattan world scenes are scenes that consist of piece-wise planar surfaces with dominant directions. In some implementations, the set of camera poses comprises augmented reality (AR) camera poses.
[0071] The method also includes detecting drift in at least one camera pose of the set of camera poses based on visual data of at least one associated image and the model observed from the at least one camera pose. The method also includes obtaining three lines for an image of the building structure, e.g., lines 604, 605, and 606 of FIG. 6. At least one line, e.g., line 605, corresponds to a different direction than other lines 604, 606, of the three lines. At least two lines 604, 606 correspond to a same direction. In some implementations, the three lines include lines that are (i) separated by at least a predetermined distance and (ii) having at least a predetermined length.
[0072] The method also includes identifying an axis for each line of the three lines based on a camera pose associated with the image. The method also includes calculating Manhattan basis for the image based on the three lines and the axis of each. The method also includes correcting the drift, e.g., in the AR system data, by adjusting one or more camera parameters of the at least one camera pose based on the Manhattan basis for the image. In some implementations, the one or more camera parameters includes rotation. In some implementations, the one or more camera parameters includes translation.
[0073] In some implementations, the method further includes associating the image with a reference axis based on the identified axis for each line of the three lines, and adjusting one or more camera parameters of one or more images subsequent to the image, based on at least one of: the reference axis and the adjusted one or more camera parameters.
[0074] In some implementations, the method further includes displaying, on a graphical user interface, e.g., 600 of FIG. 6, the image of the building structure, e.g., image 602, and detecting, on the graphical user interface, a user drawing the three lines, e.g., lines 604, 605, and 606. In some implementations, the method further includes continuing to detect the three lines until (i) at least one line corresponds to a different direction than other lines of the three lines and (ii) at least two lines that corresponds to a same direction, based on the axis for each line. In someimplementations, the image includes at least one horizontal line and at least one vertical line that the user can use as a guide to draw the three lines.
[0075] In some implementations, the method further includes identifying the three lines using line segment detector (LSD) to extract line features of the image, wherein the three lines includes (i) at least one line that corresponds to a different direction than other lines of the three lines and (ii) at least two lines that corresponds to a same direction.
[0076] In some implementations, the method further includes selecting images from the plurality of images for drift correction based on confidence scores for each of the images. In some implementations, the method further includes correcting the drift based on the selected images.
[0077] In some implementations, the method further includes correcting the drift based on an expected rotation from a camera pose and a previous image’s Manhattan basis. The previous image precedes the current image in the plurality of images, e.g., in a group or subset of images.
[0078] In some implementations, the method includes using a difference between rotations based on camera poses and rotations based on the Manhattan basis to detect drifts for the plurality of images. In an embodiment, a method includes using a difference between rotations based on camera poses and rotations based on the Manhattan basis, to correct drifts for the plurality of images. For example, an embodiment may use a rotational transform based on camera poses, e.g., from an AR system, for short time scale grouping of images into subsets, and rely on rotations based on the Manhattan basis for long time scale grouping of images, to detect drifts for the plurality of images. In other words, an embodiment may use the relative stability of the rotation transform based on the Manhattan basis to detect drift in the camera pose data of an AR system. Accordingly, in some implementations, an embodiment uses rotations based on camera poses for short time scale (propagating corrections within a small time scale subset) and uses rotations based on the Manhattan basis for longer time scales (to determine when camera pose drift from an AR system has changed significantly) to correct drifts for the plurality of images.
[0079] In an embodiment, the difference between a rotation transform based on camera poses and a rotation transform based on the Manhattan basis may be plotted, for example in a 2D graph with degree of difference on the y-axis and image number in the sequence on the x-axis, for human review. In such a display, a human reviewer can understand that one or morechanges in the difference represented in the plot signify transitions where image subsets should be made or groups defined, indicating changes in AR subsystem drift and also indicating that different corrections may be needed for the respective image subsets.
[0080] In an embodiment, the use of the difference between rotation transforms based on camera poses and rotation transforms based on the Manhattan basis may be utilized in a programmatic form, automated or semiautomated. For example, a method may include computing a Fourier transform of rotations based on camera poses (e.g., blur these rotations) and a Fourier transform of rotations based on the Manhattan basis; and using a low frequency of the Fourier transform of rotations based on camera poses and a high frequency of the Fourier transform of rotations based on the Manhattan basis (e.g., high-pass filter), to detect drifts for the plurality of images.
[0081] In an embodiment, the method further includes computing a Fourier transform of rotations based on camera poses (e.g., blur these rotations) and a Fourier transform of rotations based on the Manhattan basis; and using a low frequency of the Fourier transform of rotations based on camera poses and a high frequency of the Fourier transform of rotations based on the Manhattan basis (e.g., high-pass filter), to correct drifts for the plurality of images.
[0082] In this way, the techniques provided herein detect and / or correct drifts in images obtained from visual inertial odometry.
[0083] Referring to FIG. 7, an embodiment provides a method including obtaining, using a set of one or more processors, pose data derived from an augmented reality (AR) system of a mobile device indicative of a position and orientation of a camera of the mobile device in three- dimensional (3-D) space, as indicated at 710. The pose data may be associated with the capture of one or more images of an environment using the camera of the mobile device, e.g., from an iOS or Android tool such as ARkit or ARC ore.
[0084] In an embodiment, the method includes determining, using the set of one or more processors, geometric-based data for respective ones of the images based on detected linear features and the pose data, as indicated at 720. For example, a Manhattan basis may be determined using edges detected in an image.
[0085] In an embodiment, the method includes comparing, using the set of one or more processors, the geometric-based data to a predetermined geometry applicable to theenvironment, as indicated at 730. For example, a model based on Manhattan constraints may be used to compare the Manhattan basis of the image with a view of the model.
[0086] In an embodiment, the method includes employing, using the set of one or more processors, the comparison to detect a drift value, wherein the drift value is based on an increasing error in the pose data, as indicated at 740. For example, the comparison of the Manhattan basis of the images to the model may give a rotation transform for camera pose for the image which can be used to compute a drift value representing a difference between a rotation transform based on camera poses of the AR system and a rotation transform determined using the Manhattan basis. Thus, an embodiment may assign drift error values for each image in a series of images. For example, a drift error value may be computed as the absolute difference value of a rotation transform component from a Manhattan or geometrybased rotation and the rotation value provided by an AR system for the image.
[0087] In an embodiment, the method includes grouping, using the set of one or more processors, the images into first and second subsets based on a drift error value associated with respective images, as indicated at 750. For example, a first subset of images having a similar drift error value (e.g., within a threshold difference, such as a numerical percentage value) may be grouped together, e.g., the first hundred images in a sequence several hundred images. Another subset of images having a second similar drift error value (but different as compared to the first subset of images), maybe assigned or classified in a different group.
[0088] In an embodiment, the method may include applying, using the set of one or more processors, different corrections to the pose data for images of the first and second subsets, as indicated at 760. For example, a first subset of the images may have a first rotation correction applied to each image in the first subset, whereas the second subset of images may have a second, different rotation correction applied to each image in the second subset. The different corrections may be determined using a variety of mechanisms, for example choosing one of the AR system rotations, the Manhattan basis derived rotations, a combination thereof, or the correction may be determined manually by a user, for example employing interface 600.
[0089] In an embodiment, the method includes creating, using the set of one or processors, a 3-D model of the environment based on image features and corrected pose data for respective images, as indicated at 770. For example, once corrections have been applied to respectiveimages, the features of the images may be used to build or refine 3-D geometry and thus a 3-D model that can be used within a design tool.
[0090] In an embodiment, the grouping comprises sorting the first and second subsets based on determining that a change in the geometric-based data is statistically significant when compared to a monotonic increasing function. For example, a function of a plot of the differences between rotation transform components based on camera poses and on the Manhattan basis may be used to determine a statistically significant change or transition point within a series of images having drift error values associated therewith.
[0091] In an embodiment, a first one of the different corrections comprises rotating one of three axes of respective image data by a predetermined or computed factor. For example, the predetermined or computed factor may be selected from an AR camera pose rotation value, a geometrically determined rotation value, a combination thereof, of some combination of rotation values, such as a mean or average value of a subset or group of images.
[0092] In an embodiment, a first one of the different corrections comprises indicating to a user a need for intervention. For example, a user may be notified that an image or image feature detected is to be corrected, for example a detected line, a rotation applied to an image, etc.
[0093] In an embodiment, the environment is a building interior, for example an interior of a home that is to be remodeled.
[0094] In an embodiment, the model comprises one or more measurements of a real-world object. For example, the real -world objects found in the image(s), such as a kitchen wall and its determined size or area.
[0095] In an embodiment, the method comprises providing the model of the building interior to a design tool. For example, the dimensions of the real -world object may be imported into a design tool to allow a contractor to do cost estimates.
[0096] In an embodiment, the geometric-based data is line-based data, and the determining the geometric-based data comprises determining a Manhattan basis for the respective ones of the images. In an embodiment, determining the Manhattan basis comprises, for each respective image, identifying a minimum set of lines of the respective image useable for determining the Manhattan basis, as described in connection with FIG. 6. In an embodiment, the minimum set of lines comprises a first line in a first direction, and one or more lines in a second direction,different from the first direction, as illustrated in FIG. 6. In an embodiment, the minimum set of lines of the respective image are identified using a line detection process, e.g., using a button or element in GUI 600 to request automated edge detection. In an embodiment, the minimum set of lines of the respective image are identified using a line markup process, and the line markup process accepts user input via interface 600 to mark one or more lines within the image.
[0097] The foregoing description, for purpose of explanation, has been described with reference to specific implementations. However, the illustrative discussions above are not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The implementations were chosen and described in order to best explain the principles of the invention and its practical applications, to thereby enable others skilled in the art to best utilize the invention and various implementations with various modifications as are suited to the particular use contemplated.
Claims
What is claimed is:
1. A method, comprising: obtaining, using a set of one or more processors, pose data derived from an augmented reality (AR) system of a mobile device indicative of a position and orientation of a camera of the mobile device in three-dimensional (3-D) space; the pose data being associated with the capture of one or more images of an environment using the camera of the mobile device; determining, using the set of one or more processors, geometric-based data for respective ones of the images based on detected linear features and the pose data; comparing, using the set of one or more processors, the geometric-based data to a predetermined geometry applicable to the environment; employing, using the set of one or more processors, the comparison to detect a drift value, wherein the drift value is based on an increasing error in the pose data; grouping, using the set of one or more processors, the images into first and second subsets based on a drift error value associated with respective images; applying, using the set of one or more processors, different corrections to the pose data for images of the first and second subsets; and creating, using the set of one or processors, a 3-D model of the environment based on image features and corrected pose data for respective images.
2. The method of claim 1, wherein grouping comprises sorting the first and second subsets based on determining that a change in the geometric-based data is statistically significant when compared to a monotonic increasing function.
3. The method of claim 1, wherein a first one of the different corrections comprises rotating one of three axes of respective image data by a predetermined or computed factor.
4. The method of claim 1, wherein a first one of the different corrections comprises indicating to a user a need for intervention.
5. The method of claim 1, wherein the environment is a building interior.
6. The method of claim 1, wherein the model comprises one or more measurements of a real -world object.
7. The method of claim 5, comprising providing the model of the building interior to a design tool.
8. The method of claim 1, wherein the determining comprises determining a Manhattan basis for the respective ones of the images.
9. The method of claim 8, wherein determining the Manhattan basis comprises, for each respective image, identifying a minimum set of lines of the respective image useable for determining the Manhattan basis.
10. The method of claim 9, wherein the minimum set of lines comprises a first line in a first direction, and one or more lines in a second direction, different from the first direction.
11. The method of claim 10, wherein the minimum set of lines of the respective image are identified using an edge detection process.
12. The method of claim 11, wherein: the minimum set of lines of the respective image are identified using a line markup process; and the line markup process accepts user input via a graphical user interface (GUI) to mark one or more lines within the image.
13. A device, comprising: a memory; and a set of one or more processors configured to execute code configurable to perform any of the methods of any of the claims 1-12.
14. A computer program product comprising a non-transitory storage medium having executable code configurable to perform any of the methods of any of the claims 1-12.
Citation Information
Patent Citations
Stable video super-resolution by edge strength optimization
US20160210727A1
Structure annotation
US20210342600A1
System and method of hybrid scene representation for visual simultaneous localization and mapping
US20240104771A1