Direct scale level selection for multi-level feature tracking

By estimating motion blur levels and predicting scale changes in AR/VR devices, the optimal scale level of the image pyramid is identified for feature matching, thus solving the problem of tracking performance degradation caused by image blur and achieving savings in computational resources and improved robustness.

CN117321635BActive Publication Date: 2026-02-13SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280036061.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-11-08
Filing Date
2022-05-16
Publication Date
2026-02-13
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

When AR/VR devices move quickly, image blurring in visual tracking systems leads to a decrease in tracking performance, and existing image pyramid algorithms are computationally intensive, time-consuming, and energy-intensive.

Method used

By estimating the motion blur level and predicting scale changes, the optimal scale level for image pyramid processing is identified, and feature matching is performed only at this level, avoiding repeated matching at each level and reducing computational cost.

Benefits of technology

It improves the robustness of the visual inertial tracking system, reduces computational resource requirements, including processor cycles, network traffic, memory usage, data storage capacity, power consumption, and network bandwidth, and reduces cooling requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117321635B_ABST
    Figure CN117321635B_ABST
Patent Text Reader

Abstract

Methods for mitigating motion blur in a visual inertial tracking system are described. In an aspect, the method includes accessing a first image generated by an optical sensor of a visual tracking system, accessing a second image generated by the optical sensor of the visual tracking system, the second image being subsequent to the first image, determining a first motion blur level of the first image, determining a second motion blur level of the second image, identifying a scale change between the first image and the second image, determining a first optimal scale level for the first image based on the first motion blur level and the scale change, and determining a second optimal scale level for the second image based on the second motion blur level and the scale change.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims the benefit of priority to U.S. Application Serial No. 17 / 521,109, filed November 8, 2021, which claims priority to U.S. Provisional Patent Application Serial No. 63 / 190,101, filed May 18, 2021, each of which is incorporated herein by reference in its entirety. Technical Field

[0003] The subject matter disclosed herein generally relates to visual tracking systems. Specifically, this disclosure relates to systems and methods for mitigating motion blur in visual-inertial tracking systems. Background Technology

[0004] Augmented reality (AR) devices allow users to observe a scene while seeing related virtual content that can be aligned with items, images, objects, or the environment within the device's field of view. Virtual reality (VR) devices offer a more immersive experience than AR devices. VR devices use virtual content displayed based on the VR device's positioning and orientation to obscure the user's field of view.

[0005] Both AR and VR devices rely on motion tracking systems to track the device's pose (e.g., orientation, orientation, position). Motion tracking systems (also known as visual tracking systems) use images captured by the AR / VR device's optical sensors to track its pose. However, images become blurred when the AR / VR device moves rapidly. Therefore, high motion blur leads to degraded tracking performance. Alternatively, high motion blur results in higher computational operations to maintain sufficient tracking accuracy and image quality under high dynamic conditions. Attached Figure Description

[0006] To facilitate identification of any particular element or behavior being discussed, one or more of the highest-order digits in the reference numerals indicate the figure number in which the element was first introduced.

[0007] Figure 1 This is a block diagram illustrating an environment for operating an AR / VR display device according to an example implementation.

[0008] Figure 2 This is a block diagram illustrating an AR / VR display device according to an example implementation.

[0009] Figure 3 This is a block diagram illustrating a visual tracking system according to an example implementation.

[0010] Figure 4 This is a block diagram illustrating a motion blur mitigation module according to an example implementation.

[0011] Figure 5 This is a block diagram illustrating a process according to an example implementation.

[0012] Figure 6 This is a flowchart illustrating a method for mitigating motion blur according to an example implementation.

[0013] Figure 7 This is a flowchart illustrating a method for mitigating motion blur according to an example implementation.

[0014] Figure 8 An example of a first scenario of the subject matter according to one implementation is shown.

[0015] Figure 9 An example of a second scenario of the subject matter according to one implementation is shown.

[0016] Figure 10 An example of a third scenario of the subject matter according to one implementation is shown.

[0017] Figure 11 An example of a fourth scenario of the subject matter according to one implementation is shown.

[0018] Figure 12 An example of a fifth scenario of the subject matter according to one implementation is shown.

[0019] Figure 13 An example of pseudocode for motion blur mitigation according to one implementation is shown.

[0020] Figure 14 An example of an algorithm for motion blur reduction according to one embodiment is shown.

[0021] Figure 15 A network environment in which a head-mounted device can be implemented is shown according to an example implementation.

[0022] Figure 16 This is a block diagram illustrating a software architecture in which the present disclosure can be implemented according to an example embodiment.

[0023] Figure 17 It is a schematic representation of a machine in the form of a computer system according to an example implementation, in which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein. Detailed Implementation

[0024] The following description describes systems, methods, techniques, instruction sequences, and computer program products that illustrate example implementations of the subject matter. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide an understanding of various implementations of the present subject matter. It will be apparent, however, to one skilled in the art that implementations of the present subject matter can be practiced without some or all of these specific details. Examples are presented solely for purposes of illustration. Structurally (e.g., structural components such as modules) is optional and can be combined or subdivided, and operationally (e.g., in processes, algorithms, or other functions) can be changed in sequence or combined or subdivided unless explicitly otherwise stated.

[0025] The term “augmented reality” (AR) is used herein to refer to an interactive experience of a real-world environment where physical objects existing in the real world are “augmented” or enhanced by computer-generated digital content (also referred to as virtual content or synthetic content). AR can also refer to a system that enables a combination of real and virtual worlds, real-time interaction, and 3D registration of virtual objects and real objects. A user of an AR system perceives virtual content that appears to be attached to or interacting with physical objects of the real world.

[0026] The term “virtual reality” (VR) is used herein to refer to a simulated experience of a virtual world environment that is completely different from the real-world environment. Computer-generated digital content is displayed in the virtual world environment. VR also refers to a system that enables a user of the VR system to be fully immersed in the virtual world environment and to interact with virtual objects presented in the virtual world environment.

[0027] The term “AR application” is used herein to refer to a computer-operated application that enables an AR experience. The term “VR application” is used herein to refer to a computer-operated application that enables a VR experience. The term “AR / VR application” refers to a computer-operated application that is capable of enabling a combination of an AR experience or a VR experience.

[0028] The term “visual tracking system” is used herein to refer to a computer-operated application or system that enables a system to track visual features identified in images captured by one or more cameras of the visual tracking system. A visual tracking system establishes a model of a real-world environment based on tracked visual features. Non-limiting examples of visual tracking systems include: visual simultaneous localization and mapping systems (VSLAM) and visual odometry inertial (VIO) systems. VSLAM can be used to establish a target from an environment or scene based on one or more cameras of the visual tracking system. VIO (also referred to as visual inertial tracking systems and visual inertial odometry systems) determines a latest pose (e.g., position and orientation) of a device based on data acquired from multiple sensors (e.g., optical sensors, inertial sensors) of the device.

[0029] The term“inertial measurement unit” (IMU) is used herein to refer to a device capable of reporting the inertial state of a moving object including the acceleration, velocity, orientation, and position of the moving object. An IMU is capable of tracking the motion of an object by integrating the acceleration and angular velocity measured by the IMU. An IMU can also refer to a combination of an accelerometer and a gyroscope that determine and quantify linear acceleration and angular velocity, respectively. Values obtained from the IMU gyroscope can be processed to obtain the pitch, roll, and heading of the IMU, and thereby the pitch, roll, and heading of the object associated with the IMU. Signals from the accelerometer of the IMU can also be processed to obtain the velocity and displacement of the IMU.

[0030] Both AR and VR applications allow users to access information, for example, in the form of virtual content presented in a display of an AR / VR display device (also referred to as a display device). The presentation of virtual content can be based on the positioning of the display device relative to a physical object or relative to a frame of reference (external to the display device) such that the virtual content appears correctly in the display. For AR, the virtual content appears aligned with physical objects as perceived by the user and the camera of the AR display device. The virtual content appears to be attached to the physical world (e.g., the physical object of interest). To do this, the AR display device detects the physical object and tracks the pose of the positioning of the AR display device relative to the physical object. The pose identifies the positioning and orientation of the display device relative to a frame of reference or relative to another object. For VR, the virtual object appears at a location based on the pose of the VR display device. Thus, the virtual content is refreshed based on the latest pose of the device. A vision tracking system at the display device determines the pose of the display device. Examples of vision tracking systems include visual-inertial tracking systems (e.g., VIO systems) that rely on data acquired from multiple sensors (e.g., optical sensors, inertial sensors).

[0031] In cases where the camera moves quickly (e.g., spins quickly), the images captured by the vision tracking system can be blurred. Motion blur in the images can result in degraded tracking performance (of the vision tracking system). Alternatively, motion blur can also result in higher computational operations of the vision tracking system to maintain sufficient tracking accuracy and image quality under high dynamics.

[0032] In particular, vision tracking systems are typically based on image feature matching components. In an incoming video stream, the algorithm detects different 3D points in the images (features) and tries to find (match) these points in subsequent images. The first image in this matching process is referred to herein as the“source image”. A second image (e.g., a subsequent image in which features are to be matched) is referred to herein as the“target image”.

[0033] Reliable feature points are typically detected in high-contrast regions of an image, such as corners or edges. However, for head-mounted devices with a built-in camera, the camera can move rapidly as the user shakes his / her head, resulting in severe motion blur in images captured with the built-in camera. Such rapid motion results in blurred high-contrast regions. As a result, the feature detection and matching phase of a visual tracking system is negatively affected, and the overall tracking accuracy of the system suffers.

[0034] A common strategy to mitigate motion blur is to perform feature detection and matching on down-sampled versions of the source and target images if matching at the original image resolution fails due to motion blur. With visual information lost in the down-sampled image versions, motion blur is reduced. As a result, feature matching becomes more reliable. Typically, an image is down-sampled multiple times to obtain different resolutions for different severities of motion blur, and the collection of all different versions is referred to as an image pyramid. The downscaling process is also referred to as "image pyramid processing" or "image pyramid algorithm." However, image pyramid processing can be time-consuming, and the process is computationally intensive.

[0035] A typical image pyramid algorithm performs an iterative downscaling process on multiple levels of the source and target images until features from a down-scaled level of the source image match features from a down-scaled level of the target image. For example, in a fine-to-coarse process, the image pyramid algorithm starts at the finest level (highest image resolution) and continues until a successful match. In a coarse-to-fine process, the image pyramid algorithm starts at the coarsest level (lowest image resolution) and stops when a match fails. In either case, the image pyramid algorithm performs matching over many multiple levels.

[0036] This application describes a method that identifies the optimal scale level for feature matching. Instead of attempting to match features at every scale level of the image pyramid algorithm until a successful match is detected, the presently described method predicts the optimal scale level for feature matching prior to the matching process based on multiple inputs, such as motion blur estimates and predicted scale changes. As a result, only one matching attempt per image is needed for each feature, resulting in shorter processing times.

[0037] In one example implementation, the present application describes a method for mitigating motion blur in a visual inertial tracking system. The method includes accessing a first image generated by an optical sensor of the visual tracking system, accessing a second image generated by the optical sensor of the visual tracking system, the second image being subsequent to the first image, determining a first motion blur level of the first image, determining a second motion blur level of the second image, identifying a scale change between the first image and the second image, determining a first optimal scale level for the first image based on the first motion blur level and the scale change, and determining a second optimal scale level for the second image based on the second motion blur level and the scale change.

[0038] Accordingly, one or more of the methods described herein help to address the technical problem of power consumption savings by identifying an optimal scale level for a current image through image pyramid processing. The presently described methods provide improvements to computer function operations by providing for reduced power consumption while still maintaining robustness of visual inertial tracking to motion blur. Accordingly, one or more of the methods described herein can avoid the need for certain efforts or computational resources. Examples of such computational resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.

[0039] Figure 1 is a network diagram illustrating an environment 100 suitable for operating an AR / VR display device 106 in accordance with some example implementations. The environment 100 includes a user 102, an AR / VR display device 106, and a physical object 104. The user 102 operates the AR / VR display device 106. The user 102 can be a human user (e.g., a human), a machine user (e.g., a computer configured by a software program to interact with the AR / VR display device 106), or any suitable combination thereof (e.g., a human assisted by a machine or a machine supervised by a human). The user 102 is associated with the AR / VR display device 106.

[0040] The AR / VR display device 106 can be a computing device with a display, such as a smartphone, a tablet, or a wearable computing device (e.g., a watch or glasses). The computing device can be handheld or can be removably mounted to the head of the user 102. In one example, the display includes a screen that displays images captured with a camera of the AR / VR display device 106. In another example, the display of the device can be transparent, such as in the lenses of wearable computing glasses. In other examples, the display can be opaque, partially transparent, partially opaque. In other examples, the display can be worn by the user 102 to cover the field of view of the user 102.

[0041] The AR / VR display device 106 includes an AR application that generates virtual content based on images detected with a camera of the AR / VR display device 106. For example, the user 102 can direct the camera of the AR / VR display device 106 to capture an image of the physical object 104. The AR application generates virtual content that corresponds to an identified object in the image (e.g., the physical object 104) and presents the virtual content in a display of the AR / VR display device 106.

[0042] The AR / VR display device 106 includes a visual tracking system 108. The visual tracking system 108 tracks a pose (e.g., a position and an orientation) of the AR / VR display device 106 relative to the real-world environment 110 using, for example, optical sensors (e.g., 3D cameras with depth functionality, image cameras), inertial sensors (e.g., gyroscopes, accelerometers), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors. In one example, the AR / VR display device 106 displays virtual content based on the pose of the AR / VR display device 106 relative to the real-world environment 110 and / or the physical object 104.

[0043] Figure 1 Any of the illustrated machines, databases, or devices can be implemented in a general-purpose computer modified (e.g., configured or programmed) by software to be a special-purpose computer to perform one or more of the functions described herein for that machine, database, or device. For example, the following discussion refers to machines that can be used to implement any one or more of the methods described herein. Figures 6-7 A computer system capable of implementing any one or more of the methods described herein is discussed below. As used herein, a “database” is a data storage resource and can store data structured as a text file, a table, a spreadsheet, a relational database (e.g., an object-relational database), a triple store, a hierarchical data store, or any suitable combination thereof. Moreover, Figure 1 Any two or more of the illustrated machines, databases, or devices can be combined into a single machine, and the functions described herein for any single machine, database, or device can be subdivided among multiple machines, databases, or devices.

[0044] The AR / VR display device 106 can operate over a computer network. The computer network can be any network that enables communication among or between machines, databases, and devices. Thus, the computer network can be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. The computer network can include one or more portions that constitute a private network, a public network (e.g., the Internet), or any suitable combination thereof.

[0045] Figure 2is a block diagram illustrating modules (e.g., components) of an AR / VR display device 106, in accordance with some example embodiments. The AR / VR display device 106 includes sensors 202, a display 204, a processor 206, and a storage device 208. Examples of the AR / VR display device 106 include a wearable computing device, a mobile computing device, a navigation device, a portable media device, or a smartphone.

[0046] The sensors 202 include, for example, optical sensors 212 (e.g., cameras such as color cameras, thermal cameras, depth sensors, and one or more grayscale, global / rolling shutter tracking cameras) and inertial sensors 210 (e.g., gyroscopes, accelerometers, magnetometers). Other examples of the sensors 202 include proximity or location sensors (e.g., near field communication, GPS, Bluetooth, Wi-Fi), audio sensors (e.g., microphones), thermal sensors, pressure sensors (e.g., barometers), or any suitable combination thereof. Note that the sensors 202 described herein are for purposes of illustration and, thus, the sensors 202 are not limited to the sensors described above.

[0047] The display 204 includes a screen or monitor configured to display images generated by the processor 206. In one example embodiment, the display 204 can be transparent or semi-opaque such that the user 102 can view through the display 204 (in AR use cases). In another example embodiment, the display 204 covers the eyes of the user 102 and occludes the entire field of view of the user 102 (in VR use cases). In another example, the display 204 includes a touch screen display configured to receive user input via contact on the touch screen display.

[0048] The processor 206 includes an AR / VR application 214 and a visual tracking system 108. The AR / VR application 214 uses computer vision to detect and recognize physical objects 104 or a physical environment. The AR / VR application 214 retrieves virtual content (e.g., 3D object models) based on the recognized physical objects 104 or physical environment. The AR / VR application 214 renders the virtual objects in the display 204. In one example implementation, the AR / VR application 214 includes a local rendering engine that generates a visualization of virtual content overlaid (e.g., superimposed or otherwise displayed in coordination with) on images of the physical objects 104 captured by the optical sensor 212. The visualization of the virtual content can be manipulated by adjusting the positioning (e.g., physical location, orientation, or both) of the physical objects 104 relative to the AR / VR display device 106. Similarly, the visualization of the virtual content can be manipulated by adjusting the pose of the AR / VR display device 106 relative to the physical objects 104. For a VR application, the AR / VR application 214 displays the virtual content in the display 204 at a location (in the display 204) determined based on the pose of the AR / VR display device 106.

[0049] The visual tracking system 108 estimates the pose of the AR / VR display device 106. For example, the visual tracking system 108 uses image data from the optical sensor 212 and corresponding inertial data from the inertial sensor 210 to track the position and pose of the AR / VR display device 106 relative to a frame of reference (e.g., the real-world environment 110). See, e.g., FIG. 2B, which illustrates the AR / VR display device 106 in a first pose relative to the real-world environment 110. The visual tracking system 108 can track the pose of the AR / VR display device 106 relative to the real-world environment 110 using a variety of techniques, including, e.g., visual odometry, visual-inertial odometry, and / or visual-inertial SLAM. Figure 3 The visual tracking system 108 is described in more detail below.

[0050] The storage device 208 stores virtual content 216. The virtual content 216 includes, e.g., a database of visual references (e.g., images of physical objects) and corresponding experiences (e.g., three-dimensional virtual object models).

[0051] Any one or more of the modules described herein can be implemented using hardware (e.g., a processor of a machine) or a combination of hardware and software. For example, any of the modules described herein can configure a processor to perform the operations described herein for that module. Moreover, any two or more of these modules can be combined into a single module, and the functions described herein for a single module can be subdivided among multiple modules. Furthermore, according to various example implementations, modules described as being within a single machine, database, or device can be distributed across multiple machines, databases, or devices.

[0052] Figure 3A visual tracking system 108 according to one example embodiment is shown. The visual tracking system 108 includes an inertial sensor module 302, an optical sensor module 304, a blur mitigation module 306, and a pose estimation module 308. The inertial sensor module 302 accesses inertial sensor data from the inertial sensor 210. The optical sensor module 304 accesses optical sensor data (e.g., images, camera settings / operation parameters) from the optical sensor 212. Examples of camera operation parameters include, but are not limited to, an exposure time of the optical sensor 212, a field of view of the optical sensor 212, an ISO value of the optical sensor, and an image resolution of the optical sensor 212.

[0053] In one example embodiment, the blur mitigation module 306 determines an angular velocity of the optical sensor 212 based on the IMU sensor data from the inertial sensor 210. The blur mitigation module 306 estimates a motion blur level based on the angular velocity and the camera operation parameters without performing any analysis of the pixels in the image.

[0054] In another example embodiment, the blur mitigation module 306 considers both the angular velocity and the linear velocity of the optical sensor 212 based on a current velocity estimate from the visual tracking system 108, in combination with the 3D positions of the currently tracked points in the current image. For example, the blur mitigation module 306 determines a linear velocity of the optical sensor 212 based on the distance of the objects (from the optical sensor 212) in the current image (e.g., determined by the 3D positions of the feature points) and the effect of the linear velocity on different regions of the current image. Thus, objects closer to the optical sensor 212 appear more blurred than objects farther away from the optical sensor 212 (in the case that the optical sensor 212 is moving).

[0055] The blur mitigation module 306 down-scales the image captured by the optical sensor 212 based on the motion blur level of the image. For example, the blur mitigation module 306 determines that the current image is blurred and applies an image pyramid algorithm to the current image to increase the contrast. In one example embodiment, the blur mitigation module 306 identifies an optimal scale level for feature matching. Instead of attempting to match features at each scale level of the image pyramid algorithm until a successful match is detected, the blur mitigation module 306 predicts the optimal scale level for feature matching prior to the matching process based on the motion blur estimate and a predicted scale change. The higher the estimated motion blur, the lower the optimal resolution for feature matching. The higher the scale change between the source image and the target image, the more adjustment to the optimal scale level of the image pyramid algorithm. By predicting the optimal scale level for the source image and for the target image, the blur mitigation module 306 performs only one feature matching attempt for each image, resulting in a shorter processing time. Reference is made to the following example embodiments for further details. Figure 4Example components of the blur mitigation module 306 are described in more detail.

[0056] The pose estimation module 308 determines a pose (e.g., position, location, orientation) of the AR / VR display device 106 relative to a frame of reference (e.g., the real-world environment 110). In one example implementation, the pose estimation module 308 includes a VIO system that estimates the pose of the AR / VR display device 106 based on a 3D map of feature points from a current image captured with the optical sensors 212 and inertial sensor data captured with the inertial sensors 210.

[0057] In one example implementation, the pose estimation module 308 computes a location and an orientation of the AR / VR display device 106. The AR / VR display device 106 includes one or more optical sensors 212 mounted with one or more inertial sensors 210 on a rigid platform (a frame of the AR / VR display device 106). The optical sensors 212 can be mounted with non-overlapping (distributed aperture) or overlapping (stereo or more) fields of view.

[0058] In some example implementations, the pose estimation module 308 includes an algorithm that combines inertial information from the inertial sensors 210 and image information from the pose estimation module 308, where the inertial sensors 210 and the pose estimation module 308 are coupled to a rigid platform (e.g., the AR / VR display device 106) or a photography rig. In one implementation, the photography rig can consist of multiple cameras mounted with inertial navigation units (e.g., the inertial sensors 210) on a rigid platform. Thus, the photography rig can have at least one inertial navigation unit and at least one camera.

[0059] Figure 4 is a block diagram illustrating the blur mitigation module 306, according to one example implementation. The blur mitigation module 306 includes a motion blur estimation module 402, a scale variation estimation module 404, an optimal scale computation module 406, a pyramid computation engine 408, and a feature matching module 410.

[0060] The motion blur estimation module 402 estimates the level of motion blur for an image from the optical sensor 212. In one example implementation, the motion blur estimation module 402 estimates motion blur based on camera operation parameters (obtained from the optical sensor module 304) and angular velocity of the inertial sensor 210 (obtained from the inertial sensor module 302). The motion blur estimation module 402 retrieves camera operation parameters of the optical sensor 212 from the optical sensor module 304. For example, the camera operation parameters include settings of the optical sensor module 304 during the capture / exposure time of the current image. The motion blur estimation module 402 also retrieves inertial sensor data from the inertial sensor 210 (where the inertial sensor data is generated during the capture / exposure time of the current image). The motion blur estimation module 402 retrieves angular velocity from the IMU of the inertial sensor module 302. In one example, the motion blur estimation module 402 samples the angular velocity of the inertial sensor 210 based on the inertial sensor data sampled during the exposure time of the current image. In another example, the motion blur estimation module 402 identifies the maximum angular velocity of the inertial sensor 210 based on the inertial sensor data captured during the exposure time of the current image.

[0061] In another example implementation, the motion blur estimation module 402 estimates motion blur based on camera operation parameters, angular velocity, and linear velocity of the visual tracking system 108. The motion blur estimation module 402 retrieves angular velocity from the VIO data (from the pose estimation module 308). The motion blur estimation module 402 retrieves linear velocity (from the VIO data) and estimates its effect on motion blur in the current image based on the 3D positions of the feature points in the current image. As previously described above, objects depicted closer to the optical sensor 212 are shown as more blurred, while objects depicted farther from the optical sensor 212 are shown as less blurred. The pose estimation module 308 tracks the 3D positions of the feature points and computes the effect of the computed linear velocity on the various portions of the current image.

[0062] The scale change estimation module 404 estimates the scale change between the source image and the target image by tracking the 3D positions of the feature points provided by the pose estimation module 308. For example, a change in the position of the matched feature points (in the source image and the target image) can indicate that the optical sensor 212 is moving closer or farther from the scene.

[0063] The optimal scale computation module 406 determines the optimal scale level for the pyramid computation engine 408 based on the estimated motion blur and scale change. Figures 8-12 Examples of different scenarios of the operation of the optimal scale computation module 406 are shown.

[0064] In Figure 8In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution.

[0065] In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution. Figure 9 In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution.

[0066] Figure 10 In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution.

[0067] In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution. Figure 11 In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution.

[0068] Figure 12 In the case where the motion blur estimation module 402 detects no motion blur in both the source and target images, the scale change estimation module 404 estimates that the scale between the source and target images has changed because the feature points in the source image are further away from the feature points in the target image than the full resolution. Thus, the optimal scale computation module 406 determines that the source optimal scale level for the source image is increased to the first scale level, while the target optimal scale level for the target image remains at the full resolution.

[0069] ​​The pyramid computation engine 408 applies the image pyramid algorithm to the source image at the source optimal scale level to generate a down-scaled version of the source image. The pyramid computation engine 408 applies the image pyramid algorithm to the target image at the target optimal scale level to generate a down-scaled version of the target image. In other examples where the optimal scale level corresponds to the full resolution of the image, the pyramid computation engine 408 does not apply the image pyramid algorithm to the image.

[0070] The feature matching module 410 matches features between the down-scaled version of the source image and the down-scaled version of the target image based on the corresponding optimal scale levels determined by the optimal scale computation module 406. In one example, the feature matching module 410 matches features between the full resolution version of the source image and the down-scaled version of the target image. In another example, the feature matching module 410 matches features between the down-scaled version of the source image and the full resolution version of the target image.

[0071] Figure 5 is a block diagram illustrating an example process according to one example implementation. The visual tracking system 108 receives sensor data from the sensors 202 to determine a pose of the visual tracking system 108. The blur mitigation module 306 estimates motion blur of the source image and the target image based on the sensor data (e.g., angular velocity from an IMU or VIO, linear velocity from VIO data of the pose estimation module 308) and camera operation parameters (e.g., exposure time, field of view, resolution) associated with the source image and the target image. The blur mitigation module 306 also estimates scale changes between the source image and the target image by using VIO data (e.g., 3D points, pose) provided by the pose estimation module 308. The blur mitigation module 306 identifies a source optimal scale level for the pyramid computation engine 408 of the source image based on the motion blur of the source image and the scale changes between the source image and the target image. The blur mitigation module 306 identifies a target optimal scale level for the pyramid computation engine 408 of the target image based on the motion blur of the target image and the scale changes between the source image and the target image.

[0072] The pyramid computation engine 408 applies the image pyramid algorithm to the source image to down-scale the source image at the source optimal scale level. The pyramid computation engine 408 applies the image pyramid algorithm to the target image to down-scale the target image at the target optimal scale level. The pyramid computation engine 408 provides the down-scaled / full versions of the source image and the target image to the pose estimation module 308.

[0073] The pose estimation module 308 identifies a pose of the visual tracking system 108 based on the full resolution or down-scaled images provided by the pyramid computation engine 408. The pose estimation module 308 provides pose data to the AR / VR application 214.

[0074] The AR / VR application 214 retrieves the virtual content 216 from the storage device 208 and causes the virtual content 216 to be displayed at a certain location (in the display 204) based on the pose of the AR / VR display device 106. Note that the pose of the AR / VR display device 106 is also referred to as the pose of the visual tracking system 108 or the optical sensor 212.

[0075] Figure 6 is a flowchart illustrating a method 600 for mitigating motion blur according to one example implementation. The operations in the method 600 can be performed by the visual tracking system 108 using the components (e.g., modules, engines) described above with respect to Figure 4 Thus, the method 600 is described by way of example with reference to the blur mitigation module 306. However, it should be appreciated that at least some of the operations of the method 600 can be deployed on various other hardware configurations or performed by similar components residing elsewhere.

[0076] In block 602, the motion blur estimation module 402 estimates a source motion blur level in a source image. In block 604, the motion blur estimation module 402 estimates a target motion blur level in a target image. In block 606, the scale variation estimation module 404 identifies a scale variation between the source image and the target image. In block 608, the optimal scale calculation module 406 determines a source optimal scale level for the source image based on the source motion blur level and the scale variation. In block 610, the optimal scale calculation module 406 determines a target optimal scale level for the target image based on the target motion blur level and the scale variation. In block 612, the optimal scale calculation module 406 determines a selected scale level based on a maximum of the source optimal scale level and the target optimal scale level. In block 614, the pyramid calculation engine 408 updates the source optimal scale level and the target optimal scale level based on the selected scale level. The method 600 proceeds to block A 616.

[0077] It is noted that other implementations can use different ordering, additional or fewer operations, and different nomenclature or terminology to accomplish similar functions. In some implementations, various operations can be performed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein are selected to illustrate some principles of operation in a simplified form.

[0078] Figure 7 is a flowchart illustrating a method 700 for mitigating motion blur according to one example implementation. The operations in the method 600 can be performed by the visual tracking system 108 using the components (e.g., modules, engines) described above with respect to Figure 4The described components (e.g., modules, engines) are, may be, for example, stored in the memory 306, and / or executed by the processing unit 304. As such, the method 600 is described, by way of example, with reference to the blurring mitigation module 306. However, it is to be understood that at least some of the operations of the method 600 can be deployed on or performed by similar components residing elsewhere.

[0079] The method 700 continues from the method 600 at block A616. In block 702, the pyramid computation engine 408 down-scales the source image at a source optimal scale level. In block 704, the pyramid computation engine 408 down-scales the target image at a target optimal scale level. In block 706, the feature matching module 410 identifies source features in the down-scaled source image. In block 708, the feature matching module 410 identifies target features in the down-scaled target image. In block 710, the feature matching module 410 matches the source features with the target features. In block 712, the pose estimation module 308 determines a pose based on the matched features.

[0080] Figure 8 An example of a first scenario of the subject matter in accordance with one embodiment is shown.

[0081] Figure 9 An example of a second scenario of the subject matter in accordance with one embodiment is shown.

[0082] Figure 10 An example of a third scenario of the subject matter in accordance with one embodiment is shown.

[0083] Figure 11 An example of a fourth scenario of the subject matter in accordance with one embodiment is shown.

[0084] Figure 12 An example of a fifth scenario of the subject matter in accordance with one embodiment is shown.

[0085] Figure 13 An example of a pseudo code for motion blur mitigation in accordance with one embodiment is shown.

[0086] Figure 14 An example of an algorithm for motion blur mitigation in accordance with one embodiment is shown.

[0087] System with head-mounted device

[0088] Figure 15 A network environment 1500 in which a head-mounted device 1502 can be implemented is shown in accordance with one example embodiment. Figure 15 is a high-level functional block diagram of an example head-mounted device 1502 communicatively coupled to a mobile client device 1538 and a server system 1532 via various networks 1540.

[0089] The head-mounted device 1502 includes a camera, such as at least one of a visible light camera 1512, an infrared emitter 1514, and an infrared camera 1516. The client device 1538 can be able to connect with the head-mounted device 1502 using both the communication 1534 and the communication 1536. The client device 1538 is connected to the server system 1532 and the network 1540. The network 1540 can include any combination of wired and wireless connections.

[0090] The head-mounted device 1502 also includes two image displays of the optical assembly's image display 1504. The two image displays include one image display associated with the left side of the head-mounted device 1502 and one image display associated with the right side. The head-mounted device 1502 also includes an image display driver 1508, an image processor 1510, low-power circuitry 1526, and high-speed circuitry 1518. The optical assembly's image display 1504 is used to present images and video to a user of the head-mounted device 1102, including images that can include a graphical user interface.

[0091] The image display driver 1508 commands and controls the image display of the optical assembly's image display 1504. The image display driver 1508 can either transfer image data directly to an image display in the optical assembly's image display 1504 for presentation or must convert the image data into a signal or data format suitable for transfer to the image display device. For example, the image data can be video data formatted according to a compression format such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, or the like, and still image data can be formatted according to a compression format such as Portable Network Group (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable image file format (Exif), or the like.

[0092] As described above, the head-mounted device 1502 includes a frame and a stem (or temple) extending from a side of the frame. The head-mounted device 1502 also includes a user input device 1506 (e.g., a touch sensor or button) that includes an input surface on the head-mounted device 1502. The user input device 1506 (e.g., a touch sensor or button) is to receive input selections from a user to manipulate a graphical user interface of a presented image.

[0093] Figure 15The components shown for the head-mounted device 1502 are located on one or more circuit boards (e.g., PCBs or flexible PCBs) in the frame or temples. Alternatively or additionally, the depicted components may be located in blocks, frames, hinges, or beams of the head-mounted device 1502. The left and right sides may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, camera lenses, or any other corresponding visible light or light-capturing elements that can be used to capture data including images of a scene with unknown objects.

[0094] The head-mounted device 1502 includes a memory 1522 that stores instructions to perform a subset or all of the functions described herein. The memory 1522 may also include a storage device.

[0095] like Figure 15 As shown, the high-speed circuit system 1518 includes a high-speed processor 1520, a memory 1522, and a high-speed wireless circuit system 1524. In this example, an image display driver 1508 is coupled to the high-speed circuit system 1518 and operated by the high-speed processor 1520 to drive the left and right image displays of the image display 1504 of the optical components. The high-speed processor 1520 can be any processor capable of managing the operation of any general computing system required for high-speed communication and the head-mounted device 1502. The high-speed processor 1520 includes the processing resources required to manage the high-speed data transmission over communication 1536 to a wireless local area network (WLAN) using the high-speed wireless circuit system 1524. In some examples, the high-speed processor 1520 executes an operating system such as the LINUX operating system or another such operating system for the head-mounted device 1502, and this operating system is stored in the memory 1522 for execution. Among other duties, the high-speed processor 1520, which executes the software architecture of the head-mounted device 1502, is also used to manage data transmission with the high-speed wireless circuit system 1524. In some examples, the high-speed wireless circuit system 1524 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as Wi-Fi. In other examples, other high-speed communication standards can be implemented using the high-speed wireless circuit system 1524.

[0096] The low-power wireless circuit system 1530 and high-speed wireless circuit system 1524 of the headset 1502 may include a short-range transceiver (Bluetooth). TM This includes wireless wide area network, local area network, or wide area network transceivers (e.g., cellular or WiFi). Client device 1538, which includes transceivers communicating via communications 1534 and communications 1536, can be implemented using details of the architecture of headset 1502, as can other elements of network 1540.

[0097] Memory 1522 includes any storage device capable of storing various data and applications, including, in particular, camera data generated by the left and right infrared cameras 1516 and image processor 1510, and images generated by image display driver 1508 on image display 1504 of optical components for display. While memory 1522 is shown integrated with high-speed circuitry 1518, in other examples, memory 1522 may be a separate component of head-mounted device 1502. In some such examples, circuitry can be provided by lines connecting the image processor 1510 or low-power processor 1528 to memory 1522 via a chip including high-speed processor 1520. In other examples, high-speed processor 1520 may manage addressing of memory 1522 such that low-power processor 1528 will initiate high-speed processor 1522 whenever a read or write operation involving memory 1522 is required.

[0098] like Figure 15 As shown, the low-power processor 1528 or high-speed processor 1520 of the head-mounted device 1502 may be coupled to a camera device (visible light camera device 1512; infrared emitter 1514 or infrared camera device 1516), an image display driver 1508, a user input device 1506 (e.g., a touch sensor or button) and a memory 1522.

[0099] The head-mounted device 1502 is connected to a host computer. For example, the head-mounted device 1502 is paired with a client device 1538 via communication 1536, or connected to a server system 1532 via a network 1540. The server system 1532 may be one or more computing devices as part of a service or network computing system, and for example, it includes a processor, memory, and network communication interfaces to communicate with the client device 1538 and the head-mounted device 1502 via the network 1540.

[0100] Client device 1538 includes a processor and a network communication interface coupled to the processor. The network communication interface enables communication via network 1540, communication 1534, or communication 1536. Client device 1538 may also store at least a portion of the instructions for generating dual-channel audio content in the memory of client device 1538 to implement the functions described herein.

[0101] The output components of the head-worn device 1502 include visual components, e.g., a display such as a liquid crystal display (LCD), a plasma display panel (PDP), a light emitting diode (LED) display, a projector, or a waveguide, etc. The image display of the optical assembly is driven by an image display driver 1508. The output components of the head-worn device 1502 also include acoustic components (e.g., a speaker), haptic components (e.g., a vibratory motor), other signal generators, etc. The input components of the head-worn device 1502, client device 1538, and server system 1532 (e.g., user input devices 1506) can include alphanumeric input components (e.g., a keyboard, a touchscreen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., a physical button, a touch screen that provides location and force of touches or touch gestures, or other tactile input components), audio input components (e.g., a microphone), etc.

[0102] The head-worn device 1502 can optionally include additional peripheral device elements. Such peripheral device elements can include biometric sensors, additional sensors, or display elements integrated with the head-worn device 1502. For example, the peripheral device elements can include any I / O components, including output components, motion components, positioning components, or any other such elements described herein.

[0103] For example, biometric components include components that detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components include acceleration sensor components (e.g., accelerometer), gravitation sensor components, rotation sensor components (e.g., gyroscope), and so forth. The position components include location sensor components (e.g., a Global Position System (GPS) receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like. Such positioning system coordinates can also be received from the client device 1538 through the communication 1536 via the low power wireless circuitry 1530 or the high speed wireless circuitry 1524. TM For example, biometric components include components that detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components include acceleration sensor components (e.g., accelerometer), gravitation sensor components, rotation sensor components (e.g., gyroscope), and so forth. The position components include location sensor components (e.g., a Global Position System (GPS) receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like. Such positioning system coordinates can also be received from the client device 1538 through the communication 1536 via the low power wireless circuitry 1530 or the high speed wireless circuitry 1524.

[0104] In the case of using phrases such as "at least one of A, B, or C," "at least one of A, B, and C," "one or more of A, B, or C," or "one or more of A, B, and C," it is intended that the phrase be interpreted to mean that A alone can be present in an embodiment, B alone can be present in an embodiment, C alone can be present in an embodiment, or any combination of A, B, and C can be present in a single embodiment.

[0105] Changes and modifications can be made to the disclosed embodiments without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure, as expressed in the appended claims.

[0106] Figure 16 is a block diagram 1600 illustrating software architecture 1604, which can be installed on any one or more of the devices described herein. The software architecture 1604 is supported by hardware such as machine 1602 that includes processors 1620, memory 1626, and I / O components 1638. In this example, the software architecture 1604 can be conceptualized as a stack of layers, where each layer provides a particular functionality. The software architecture 1604 includes layers such as an operating system 1612, libraries 1610, frameworks 1608, and applications 1606. Operationally, the applications 1606 invoke API calls 1650 through the software stack and receive messages 1652 in response to the API calls 1650.

[0107] The operating system 1612 manages hardware resources and provides common services. The operating system 1612 includes, for example, a kernel 1614, services 1616, and drivers 1622. The kernel 1614 acts as an abstraction layer between the hardware and the other software layers. For example, the kernel 1614 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functionality. The services 1616 can provide other common services that the applications 1606 and other software layers use. The drivers 1622 or Low energy drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Audio drivers, power management drivers, and so forth.

[0108] The libraries 1610 provide a low-level common infrastructure used by the applications 1606. The libraries 1610 can include system libraries 1618 (e.g., C standard library) providing functionality such as memory allocation functions, string manipulation strings, mathematics functions, and the like. Further, the libraries 1610 can include API libraries 1624 such as media libraries (e.g., libraries to support presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., an OpenGL framework used to render two and three dimensional graphics on a display), database libraries (e.g., SQLite to provide various relational database functions), web libraries (e.g., WebKit to provide web browsing functionality), and the like. The libraries 1610 can also include a wide variety of other libraries 1628 to provide many other APIs to the applications 1606.

[0109] The frameworks 1608 provide a high-level common infrastructure used by the applications 1606. For example, the frameworks 1608 provide various graphical user interface (GUI) functions, high-level resource management, and high-level positioning

[0110] In an example implementation, the applications 1606 can include a home application 1636, a contacts application 1630, a browser application 1632, a book reader application 1634, a location application 1642, a media application 1644, a messaging application 1646, a game application 1648, and a broad assortment of other applications such as a third party application 1640. The applications 1606 are programs that execute functions defined in the programs. Programs may TM be written in various programming languages, such as an object oriented programming language such as Objective-C, Java, or C++, or a procedural programming language such as C or assembly language. In a specific example, the third party application 1640 (e.g., an application developed by an entity other than the vendor of the particular platform) can be an Android TM application written using the Android software development kit (SDK) provided by Google TM . TM ​Mobile software running on the mobile operating system of the Phone or another mobile operating system. In this example, the third-party application 1640 can invoke API calls 1650 provided by the operating system 1612 to facilitate the functionality described herein.

[0111] Figure 17 is a diagrammatic representation of the machine 1700 within which instructions 1708 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 1700 to perform any one or more of the methodologies discussed herein can be executed. For example, the instructions 1708 can cause the machine 1700 to execute any one or more of the methods described herein. The instructions 1708 transform the general, non-programmed machine 1700 into a particular machine 1700 programmed to carry out the described and illustrated functions in the manner described. The machine 1700 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1700 can operate in the capacity of a server machine or a client machine in server-client network environments, or as a peer machine in peer-to-peer (or distributed) network environments. The machine 1700 can comprise, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 1708, sequentially or otherwise, that specify actions to be taken by machine 1700. Further, while only a single machine 1700 is illustrated, the term “machine” shall also be taken to include a collection of machines 1700 that individually or jointly execute the instructions 1708 to perform any one or more of the methodologies discussed herein.

[0112] The machine 1700 can include processors 1702, memory 1704, and I / O components 1742, which can be configured to communicate with each other via a bus 1744. In example embodiments, the processors 1702 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio-frequency integrated circuit (RFIC), other processors, or any suitable combination thereof) can include, for example, a processor 1706 and a processor 1710 that execute the instructions 1708. The term “processor” is intended to include a multi-core processor that can include two or more independent processors (sometimes referred to as “cores”) that can execute instructions contemporaneously. Although FIG. 17 shows the machine 1700 as having only one bus 1744, the machine 1700 can have Figure 17Multiple processors 1702 are shown, but machine 1700 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0113] Memory 1704 includes main memory 1712, static memory 1714, and storage cells 1716, all of which are accessible by processor 1702 via bus 1744. Main memory 1712, static memory 1714, and storage cells 1716 store instructions 1708 embodying any one or more of the methods or functions described herein. Instructions 1708 may also reside wholly or partially in main memory 1712, in static memory 1714, in machine-readable medium 1718 within storage cell 1716, within at least one processor in processor 1702 (e.g., within the processor's cache memory), or in any suitable combination thereof during execution by machine 1700.

[0114] I / O component 1742 may include a wide variety of components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurement results, etc. The specific I / O component 1742 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is unlikely to include such a touch input device. It should be understood that I / O component 1742 may be included in... Figure 17 Many other components are not shown. In various example embodiments, I / O component 1742 may include output component 1728 and input component 1730. Output component 1728 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. Input component 1730 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens or other haptic input components that provide the position and / or force of a touch or touch gesture), audio input components (e.g., microphones), etc.

[0115] In other example implementations, the I / O components 1742 can include biometric components 1732, motion components 1734, environmental components 1736, or positioning components 1738, among a variety of other components. Biometric components 1732, for example, include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogram-based identification), and the like. The motion components 1734 include acceleration sensor components (e.g., accelerometer), gravitation sensor components, rotation sensor components (e.g., gyroscope), and so forth. The environmental components 1736 include, for example, illumination sensor components (e.g., photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometer), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to a physical environment. The positioning components 1738 include location sensor components (e.g., a GPS receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and the like.

[0116] Communication can be implemented using a wide variety of technologies. The I / O components 1742 further include communication components 1740 that can be components (e.g., low energy), components, and other communication components to provide communication via other modalities. The devices 1722 can be another machine or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a USB).

[0117] Moreover, the communication components 1740 can detect identifiers or include components operable to detect identifiers. For example, the communication components 1740 can include radio frequency identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes), or acoustic detection components (e.g., microphones to identify tagged audio signals). In addition, a variety of information can be derived via the communication components 1740, such as, for example, location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, location via cellular signal triangulation, location via detecting NFC beacon signals that can indicate a particular location, and so forth.

[0118] The various memories (e.g., memory 1704, main memory 1712, static memory 1714, and / or memory shared by processor 1702) and / or the storage unit 1716 can store one or more sets of instructions and data structures (e.g., software) embodying or utilized by any one or more of the methodologies or functions described herein. These instructions (e.g., instructions 1708) can be those

[0119] The instructions 1708 can be transmitted or received over the network 1720 via the network interface device (e.g., a network interface component included in the communication components 1740) using a transmission medium and any one of a number of well-known transfer protocols (e.g., hypertext transfer protocol (HTTP)). Similarly, the instructions 1708 can be transmitted or received using a transmission medium via the coupling 1726 (e.g., a peer-to-peer coupling) to the devices 1722.

[0120] As used herein, the terms "machine-storage medium," "device-storage medium," and "computer-storage medium" mean the same thing and can be used interchangeably in this disclosure. The terms refer to a single or multiple storage devices and / or media (e.g., a centralized or distributed database, and / or associated caches and servers) that store executable instructions and / or data structures. Thus, the term should be taken to include a singular entity or multiple entities and be taken to include both memory on processor(s) and storage devices external to processor(s). Specific examples of machine-storage media, computer-storage media, and / or device-storage media include nonvolatile memory, including by way of example semiconductor memory devices, e.g., erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), field-programmable gate arrays (FPGAs), and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine-storage media," "computer-storage media," and "device-storage media" specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are covered under the term "signal medium" as discussed below.

[0121] The terms "transmission medium" and "signal medium" mean the same thing and can be used interchangeably in this disclosure. The terms "transmission medium" and "signal medium" shall be taken to include any intangible medium that is capable of storing, encoding, or carrying the instructions 1416 for execution by the machine 1400, and includes digital or analog communications signals or other intangible media to facilitate communication of such software. Hence, the terms "transmission medium" and "signal medium" shall be taken to include any form of a modulated data signal, carrier wave, and so on. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.

[0122] The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure. The terms are defined to include both machine-storage media and transmission media. Thus, the terms include both storage devices / media and carrier waves / modulated data signals.

[0123] While implementations have been described with reference to particular examples, it is apparent to those skilled in the art that various modifications and changes can be made thereto without departing from the broader spirit and scope of the disclosure. The present specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. The accompanying drawings, which are incorporated in and form a part of the specification, illustrate one or more implementations of the present disclosure and together with the description serve to explain the principles of the disclosure. The implementations shown are full, complete, and disability detailed to enable one of ordinary skill in the art to make and use the teachings herein. Other implementations can be employed, and thus the general principles defined herein can be applied to other implementations without departing from the scope of the present disclosure. Accordingly, the present disclosure is not intended to be limited to the implementations shown but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0124] The implementations of the subject matter of the present disclosure can be referred to herein, individually and / or collectively, by the term "application" merely for convenience and without intending to voluntarily limit the scope of this application to any single implementation or inventive concept if there is more than one. Thus, although specific implementations have been illustrated and described herein, it should be appreciated that any arrangement can be substituted for the specific implementations shown. This disclosure is intended to cover any and all adaptations or variations of various implementations. Combinations of the above-described implementations, and other implementations not specifically described herein, will be apparent to those of reasonable skill in the art upon reviewing the above description.

[0125] The Abstract of the Disclosure is provided to allow a reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that the Abstract is not intended to be used to interpret or limit the scope or the meaning of the claims. Additionally, in the foregoing Detailed Description, it can be seen that various features are grouped together in a single embodiment for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments require more features than are explicitly recited in each claim. Rather, as the following claims reflect, inventive subject matter can lie in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, where each claim can stand as a separate embodiment.

[0126] Examples

[0127] Example 1 is a method for selective motion blur mitigation in a visual tracking system, comprising: accessing a first image generated by an optical sensor of the visual tracking system; accessing a second image generated by the optical sensor of the visual tracking system, the second image subsequent to the first image; determining a first motion blur level of the first image; determining a second motion blur level of the second image; identifying a scale change between the first image and the second image; determining a first optimal scale level for the first image based on the first motion blur level and the scale change; and determining a second optimal scale level for the second image based on the second motion blur level and the scale change.

[0128] Example 2 includes the example 1, further comprising: downscaling the first image using a multi-level downscaling algorithm at the first optimal scale level to generate a first down-scaled image; and downscaling the second image using the multi-level downscaling algorithm at the first optimal scale level to generate a second down-scaled image.

[0129] Example 3 includes the example 2, further comprising: identifying a first feature in the first down-scaled image; identifying a second feature in the second down-scaled image; and matching the first feature with the second feature.

[0130] Example 4 includes the example 1, wherein determining a first optimal scale level for the first image comprises: calculating a first matching level based on the first motion blur level; applying the scale change to the first matching level to generate a scaled matching level for the first image; identifying a selected scale level based on a maximum level between the scaled matching level for the first image and a second matching level based on the second motion blur level; and applying the selected scale level to the first optimal scale level for the first image.

[0131] Example 5 includes the example 1, wherein determining a second optimal scale level for the second image comprises: calculating a second matching level based on the second motion blur level; identifying a selected scale level based on a maximum level between the scaled matching level for the first image and the second matching level based on the second motion blur level; and applying the selected scale level to the second optimal scale level for the second image.

[0132] Example 6 includes the example 1, further comprising: calculating a first matching level based on the first motion blur level; calculating a second matching level based on the second motion blur level; determining a base matching level based on a maximum of the first matching level and the second matching level; and adjusting the base matching level based on the scale change.

[0133] Example 7 includes the example 1, wherein determining the first motion blur level comprises: identifying, for the first image, first camera operation parameters of the optical sensor; and determining, for the first image, a first motion of the optical sensor, wherein determining the second motion blur level comprises: identifying, for the second image, second camera operation parameters of the optical sensor; and determining, for the second image, a second motion of the optical sensor.

[0134] Example 8 includes the example 7, wherein determining the first motion of the optical sensor for the first image comprises: retrieving, for the first image, first inertial sensor data from an inertial sensor of the visual tracking system; and determining, based on the first inertial sensor data, a first angular velocity of the visual tracking system, wherein the first motion blur level is based on the first camera operation parameters of the visual tracking system and the first angular velocity without analyzing content of the first image, wherein determining the second motion of the optical sensor for the second image comprises: retrieving, for the second image, second inertial sensor data from the inertial sensor of the visual tracking system; and determining, based on the second inertial sensor data, a second angular velocity of the visual tracking system, wherein the second motion blur level is based on the second camera operation parameters of the visual tracking system and the second angular velocity without analyzing content of the second image.

[0135] Example 9 includes the example 7, wherein determining the first motion of the optical sensor for the first image includes: accessing first VIO data from a VIO system of the visual tracking system, the first VIO data including a first estimated angular velocity of the optical sensor, a first estimated linear velocity of the optical sensor, and locations of feature points in the first image, wherein the first motion blur level is based on the first camera operation parameters and the first VIO data without analyzing content of the first image, wherein the first motion blur in different regions of the first image is based on the first estimated angular velocity of the optical sensor, the first estimated linear velocity of the optical sensor, and 3D locations of the feature points in the corresponding different regions of the first image relative to the optical sensor, wherein determining the first motion of the optical sensor for the first image includes: accessing second VIO data from the VIO system of the visual tracking system, the second VIO data including a second estimated angular velocity of the optical sensor, a second estimated linear velocity of the optical sensor, and locations of feature points in the second image, wherein the second motion blur level is based on the second camera operation parameters and the second VIO data without analyzing content of the second image, wherein the second motion blur in different regions of the second image is based on the second estimated angular velocity of the optical sensor, the second estimated linear velocity of the optical sensor, and 3D locations of the feature points in the corresponding different regions of the second image relative to the optical sensor.

[0136] Example 10 includes the example 7, wherein the first source camera operation parameter or the second source camera operation parameter includes a combination of an exposure time of the optical sensor, a field of view of the optical sensor, an ISO value of the optical sensor, and an image resolution, wherein the first image includes a source image, wherein the second image includes a target image.

[0137] Example 11 is a computing device comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access a first image generated by an optical sensor of a visual tracking system; access a second image generated by the optical sensor of the visual tracking system, the second image being subsequent to the first image; determine a first motion blur level of the first image; determine a second motion blur level of the second image; identify a scale change between the first image and the second image; determine a first optimal scale level for the first image based on the first motion blur level and the scale change; and determine a second optimal scale level for the second image based on the second motion blur level and the scale change.

[0138] Example 12 includes the example 11, wherein the instructions further configure the device to: downscale the first image using a multi-level downscaling algorithm to generate a first down-scaled image at the first optimal scale level; and downscale the second image using the multi-level downscaling algorithm to generate a second down-scaled image at the first optimal scale level.

[0139] Example 13 includes the example 12, wherein the instructions further configure the device to: identify a first feature in the first down-scaled image; identify a second feature in the second down-scaled image; and match the first feature to the second feature.

[0140] Example 14 includes the example 11, wherein determining a first optimal scale level for the first image includes: calculating a first matching level based on the first motion blur level; applying the scale change to the first matching level to generate a scaled matching level for the first image; identifying a selected scale level based on a maximum level between the scaled matching level for the first image and a second matching level based on the second motion blur level; and applying the selected scale level to the first optimal scale level for the first image.

[0141] Example 15 includes the example 11, wherein determining a second optimal scale level for the second image includes: calculating a second matching level based on the second motion blur level; identifying a selected scale level based on a maximum level between a scaled matching level for the first image and the second matching level based on the second motion blur level; and applying the selected scale level to the second optimal scale level for the second image.

[0142] Example 16 includes the example 11, wherein the instructions further configure the device to: calculate a first matching level based on the first motion blur level; calculate a second matching level based on the second motion blur level; determine a base matching level based on a maximum of the first matching level and the second matching level; and adjust the base matching level based on the scale change.

[0143] Example 17 includes the example 11, wherein determining the first motion blur level includes: identifying first camera operation parameters of the optical sensor for the first image; and determining a first motion of the optical sensor for the first image, wherein determining the second motion blur level includes: identifying second camera operation parameters of the optical sensor for the second image; and determining a second motion of the optical sensor for the second image.

[0144] Example 18 includes the example 17, wherein determining the first motion of the optical sensor for the first image includes: retrieving first inertial sensor data from an inertial sensor of the visual tracking system for the first image; and determining a first angular velocity of the visual tracking system based on the first inertial sensor data, wherein the first motion blur level is based on the first camera operating parameters and the first angular velocity of the visual tracking system without analyzing content of the first image, wherein determining the second motion of the optical sensor for the second image includes: retrieving second inertial sensor data from the inertial sensor of the visual tracking system for the second image; and determining a second angular velocity of the visual tracking system based on the second inertial sensor data, wherein the second motion blur level is based on the second camera operating parameters and the second angular velocity of the visual tracking system without analyzing content of the second image.

[0145] Example 19 includes the example 17, wherein determining the first motion of the optical sensor for the first image includes: accessing first VIO data from a VIO system of the visual tracking system, the first VIO data including a first estimated angular velocity of the optical sensor, a first estimated linear velocity of the optical sensor, and locations of feature points in the first image, wherein the first motion blur level is based on the first camera operating parameters and the first VIO data without analyzing content of the first image, wherein the first motion blur in different regions of the first image is based on the first estimated angular velocity of the optical sensor, the first estimated linear velocity of the optical sensor, and 3D locations of the feature points relative to the optical sensor in corresponding different regions of the first image, wherein determining the first motion of the optical sensor for the first image includes: accessing second VIO data from the VIO system of the visual tracking system, the second VIO data including a second estimated angular velocity of the optical sensor, a second estimated linear velocity of the optical sensor, and locations of feature points in the second image, wherein the second motion blur level is based on the second camera operating parameters and the second VIO data without analyzing content of the second image, wherein the second motion blur in different regions of the second image is based on the second estimated angular velocity of the optical sensor, the second estimated linear velocity of the optical sensor, and 3D locations of the feature points relative to the optical sensor in corresponding different regions of the second image.

[0146] Example 20 is a non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to: access a first image generated by an optical sensor of a visual tracking system; access a second image generated by the optical sensor of the visual tracking system, the second image being subsequent to the first image; determine a first motion blur level of the first image; determine a second motion blur level of the second image; identify a scale change between the first image and the second image; determine a first optimal scale level for the first image based on the first motion blur level and the scale change; and determine a second optimal scale level for the second image based on the second motion blur level and the scale change.

Claims

1. A method for selective motion blur mitigation in a visual tracking system, comprising: accessing a first image generated by an optical sensor of the visual tracking system; accessing a second image generated by the optical sensor of the visual tracking system, the second image subsequent to the first image; determining a first motion blur level for the first image; determining a second motion blur level for the second image; identifying a scale change between the first image and the second image; determining a first optimal scale level for the first image based on the first motion blur level and the scale change by: calculating a first matching level based on the first motion blur level; applying the scale change to the first matching level to generate a scaled matching level for the first image; identifying a selected scale level based on a maximum level between the scaled matching level for the first image and a second matching level based on the second motion blur level; and applying the selected scale level to the first optimal scale level for the first image; and determining a second optimal scale level for the second image based on the second motion blur level and the scale change.

2. The method of claim 1, further comprising: downscaling the first image at the first optimal scale level to generate a first down-scaled image; and downscaling the second image at the second optimal scale level to generate a second down-scaled image.

3. The method of claim 2, further comprising: identifying a first feature in the first down-scaled image; identifying a second feature in the second down-scaled image; and matching the first feature to the second feature.

4. The method of claim 1, wherein, Determining a second optimal scale level for the second image comprises: calculating a second matching level based on the second motion blur level; and applying the selected scale level to the second optimal scale level for the second image.

5. The method of claim 1, further comprising: calculating a second matching level based on the second motion blur level; determining a base matching level based on a maximum of the first matching level and the second matching level; and adjusting the base matching level based on the scale change.

6. The method of claim 1, wherein, Determining the first motion blur level comprises: identifying a first camera operation parameter of the optical sensor for the first image; and determining a first motion of the optical sensor for the first image, wherein determining the second motion blur level comprises: identifying a second camera operation parameter of the optical sensor for the second image; and determining a second motion of the optical sensor for the second image.

7. The method of claim 6, wherein, Determining a first motion of the optical sensor for the first image comprises: retrieving first inertial sensor data from an inertial sensor of the visual tracking system for the first image; and determining a first angular velocity of the visual tracking system based on the first inertial sensor data, wherein the first motion blur level is based on the first camera operation parameters of the visual tracking system and the first angular velocity without analyzing content of the first image, wherein determining the second motion of the optical sensor for the second image comprises: retrieving second inertial sensor data from the inertial sensor of the visual tracking system for the second image; and determining a second angular velocity of the visual tracking system based on the second inertial sensor data, wherein the second motion blur level is based on the second camera operation parameters of the visual tracking system and the second angular velocity without analyzing content of the second image.

8. The method of claim 6, wherein, wherein determining the first motion of the optical sensor for the first image comprises: accessing first VIO data from a VIO system of the visual tracking system, the first VIO data comprising a first estimated angular velocity of the optical sensor, a first estimated linear velocity of the optical sensor, and positions of feature points in the first image, wherein the first motion blur level is based on the first camera operation parameters and the first VIO data without analyzing content of the first image, wherein first motion blur in different regions of the first image is based on the first estimated angular velocity of the optical sensor, the first estimated linear velocity of the optical sensor, and 3D positions of feature points in the corresponding different regions of the first image relative to the optical sensor, wherein determining the second motion of the optical sensor for the second image comprises: accessing second VIO data from the VIO system of the visual tracking system, the second VIO data comprising a second estimated angular velocity of the optical sensor, a second estimated linear velocity of the optical sensor, and positions of feature points in the second image, wherein the second motion blur level is based on the second camera operation parameters and the second VIO data without analyzing content of the second image, wherein second motion blur in different regions of the second image is based on the second estimated angular velocity of the optical sensor, the second estimated linear velocity of the optical sensor, and 3D positions of feature points in the corresponding different regions of the second image relative to the optical sensor.

9. The method of claim 6, wherein, the first camera operation parameters or the second camera operation parameters comprise a combination of an exposure time of the optical sensor, a field of view of the optical sensor, an ISO value of the optical sensor, and an image resolution, wherein the first image comprises a source image, and wherein the second image comprises a target image.

10. A computing device comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access a first image generated by an optical sensor of a visual tracking system; access a second image generated by the optical sensor of the visual tracking system, the second image being subsequent to the first image; determine a first motion blur level for the first image; determine a second motion blur level for the second image; identifying a scale change between the first image and the second image; determining a first optimal scale level for the first image based on the first motion blur level and the scale change by: computing a first matching level based on the first motion blur level; applying the scale change to the first matching level to generate a scaled matching level for the first image; identifying a selected scale level based on a maximum level between the scaled matching level for the first image and a second matching level based on the second motion blur level; and applying the selected scale level to a first optimal scale level for the first image; and determining a second optimal scale level for the second image based on the second motion blur level and the scale change.

11. The computing device of claim 10, wherein, the instructions further configure the device to: downscale the first image at the first optimal scale level to generate a first down-scaled image; and downscale the second image at the second optimal scale level to generate a second down-scaled image.

12. The computing device of claim 11, wherein, the instructions further configure the device to: identify a first feature in the first down-scaled image; identify a second feature in the second down-scaled image; and match the first feature to the second feature.

13. The computing device of claim 10, wherein, determining a second optimal scale level for the second image includes: computing a second matching level based on the second motion blur level; and applying the selected scale level to a second optimal scale level for the second image.

14. The computing device of claim 10, wherein, the instructions further configure the device to: compute a second matching level based on the second motion blur level; determine a base matching level based on a maximum of the first matching level and the second matching level; and adjust the base matching level based on the scale change.

15. The computing device of claim 10, wherein, determining the first motion blur level includes: identifying, for the first image, a first camera operation parameter of the optical sensor; and determining, for the first image, a first motion of the optical sensor, wherein determining the second motion blur level includes: identifying, for the second image, a second camera operation parameter of the optical sensor; and determining, for the second image, a second motion of the optical sensor.

16. The computing device of claim 15, wherein, determining the first motion of the optical sensor for the first image includes: retrieving, for the first image, first inertial sensor data from an inertial sensor of the visual tracking system; and determining, based on the first inertial sensor data, a first angular velocity of the visual tracking system, wherein the first motion blur level is based on the first camera operation parameter and the first angular velocity of the visual tracking system without analyzing content of the first image, wherein determining the second motion of the optical sensor for the second image includes: retrieving, for the second image, second inertial sensor data from the inertial sensor of the visual tracking system; and determining, based on the second inertial sensor data, a second angular velocity of the visual tracking system, wherein the second motion blur level is based on the second camera operating parameter and the second angular velocity of the visual tracking system without analyzing content of the second image.

17. The computing device of claim 15, wherein, determining the first motion of the optical sensor for the first image comprises: accessing first VIO data from a VIO system of the visual tracking system, the first VIO data comprising a first estimated angular velocity of the optical sensor, a first estimated linear velocity of the optical sensor, and positions of feature points in the first image, wherein the first motion blur level is based on the first camera operating parameter and the first VIO data without analyzing content of the first image, wherein the first motion blur in different regions of the first image is based on the first estimated angular velocity of the optical sensor, the first estimated linear velocity of the optical sensor, and 3D positions of feature points in the corresponding different regions of the first image relative to the optical sensor, wherein determining the second motion of the optical sensor for the second image comprises: accessing second VIO data from the VIO system of the visual tracking system, the second VIO data comprising a second estimated angular velocity of the optical sensor, a second estimated linear velocity of the optical sensor, and positions of feature points in the second image, wherein the second motion blur level is based on the second camera operating parameter and the second VIO data without analyzing content of the second image, wherein the second motion blur in different regions of the second image is based on the second estimated angular velocity of the optical sensor, the second estimated linear velocity of the optical sensor, and 3D positions of feature points in the corresponding different regions of the second image relative to the optical sensor.

18. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform the following operations: accessing a first image generated by an optical sensor of a visual tracking system; accessing a second image generated by the optical sensor of the visual tracking system, the second image being subsequent to the first image; determining a first motion blur level for the first image; determining a second motion blur level for the second image; identifying a scale change between the first image and the second image; determining a first optimal scale level for the first image based on the first motion blur level and the scale change by: calculating a first matching level based on the first motion blur level; applying the scale change to the first matching level to generate a scaled matching level for the first image; identifying a selected scale level based on a maximum level between the scaled matching level of the first image and a second matching level based on the second motion blur level; and applying the selected scale level to the first optimal scale level for the first image; and determining a second optimal scale level for the second image based on the second motion blur level and the scale change.