Systems and methods for simultaneous positioning and mapping
By initializing the SLAM process using image frames and IMU data, combining partial and complete loops, the problem of local minimum value trapped in the SLAM process is solved, and efficient positioning and mapping on finite resource devices is achieved.
Patent Information
- Application Number
- CN202310044797.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-08-30
- Filing Date
- 2017-08-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2037-08-30
AI Technical Summary
Existing SLAM processes tend to fall into local minimums when matching image features, resulting in inaccurate positioning and mapping, especially on devices that use limited computing resources, which are difficult to achieve real-time and efficient positioning and mapping.
By using image frames from two different postures in the physical environment, combining inertial measurement unit (IMU) data, the SLAM process is initialized, and by ignoring feature errors to ensure convergence to a global minimum, a combination of partial SLAM process loops and full SLAM process loops is used to reduce computing resource requirements.
It realizes efficient and real-time positioning and mapping on devices with limited computing resources, improves positioning accuracy, and reduces the consumption of computing resources.
Smart Images

Figure CN116051640B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with application number 201780052981.6 and invention title "Systems and Methods for Simultaneous Localization and Mapping" filed on August 30, 2017.
[0002] Cross - reference to related applications
[0003] This application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 381,036, filed on August 30, 2016, which is incorporated herein by reference. Technical field
[0004] The embodiments described herein relate to the localization and mapping of sensors within a physical environment, and more particularly, to systems, methods, devices, and instructions for performing Simultaneous Localization and Mapping (SLAM). Background art
[0005] A SLAM (Simultaneous Localization and Mapping) process (e.g., algorithm) can be used by a mobile computing device (e.g., a mobile phone, a tablet, a wearable augmented reality (AR) device, a wearable autonomous aerial or ground vehicle, or a robot) to map the structure of the physical environment surrounding the mobile computing device and to localize the relative position of the mobile computing device within the mapped environment. When the mobile computing device moves within its physical environment, the SLAM process can typically map and localize in real - time.
[0006] Although not specifically image - based, some SLAM processes achieve mapping and localization by using images of the physical environment provided by an image sensor associated with the mobile computing device (e.g., the built - in camera of a mobile phone). From the captured images, such SLAM processes can recover the position of the mobile computing device and construct a map of the physical environment surrounding the mobile computing device by recovering the pose of the image sensor and the structure of the map (initially without knowing either).
[0007] SLAM processes that use captured images typically require several images of corresponding physical features (hereinafter referred to as features) in the physical environment, which are captured by an image sensor (e.g., of the mobile computing device) in different poses. Images captured from different camera positions allow such SLAM processes to converge and start their localization and mapping processes. Unfortunately, the localization problem in image - based SLAM processes is often difficult to solve because of errors in matching corresponding features between the captured images - these errors tend to move the local result of the minimization problem of the SLAM process to a local minimum rather than providing the global minimum for a specific location. Brief description of the drawings
[0008] The various figures illustrate only some embodiments of the present disclosure and should not be considered as limiting its scope. The figures are not necessarily drawn to scale. To easily identify the discussion of any particular element or action, the most significant digit in the reference number refers to the figure number in which the element was first introduced, and like reference numerals may describe similar components in different views.
[0009] Figure 1 is a block diagram illustrating an exemplary client - server based advanced network architecture including a Simultaneous Localization and Mapping (SLAM) system according to some embodiments;
[0010] Figure 2 is a block diagram illustrating an example computing device including a SLAM system according to some embodiments;
[0011] Figures 3 - 7 is a flowchart illustrating an example method for a SLAM process according to various embodiments;
[0012] Figure 8 is a block diagram illustrating a representative software architecture that can be used in conjunction with the various hardware architectures described herein to implement embodiments; and
[0013] Figure 9 is a block diagram illustrating components of a machine capable of reading instructions from a machine - readable medium (e.g., a machine - readable storage medium) and performing any one or more of the methods discussed herein. Detailed Description
[0014] Various embodiments provide systems, methods, devices, and instructions for performing Simultaneous Localization and Mapping (SLAM), which involve using images (hereinafter referred to as image frames) from as few as two different poses (e.g., physical locations) of a camera in a physical environment to initialize the SLAM process. Some embodiments can achieve this by ignoring errors (hereinafter referred to as feature errors) in matching corresponding features depicted in image frames of the (physical environment) captured by an image sensor of a mobile computing device, and by updating the SLAM process in a manner that causes the minimization process to converge to a global minimum rather than fall into a local minimum. The global minimum can provide the physical location of the image sensor.
[0015] According to some embodiments, the SLAM process is initialized by detecting the movement of a mobile computing device between two physical locations within a physical environment, where the movement is bounded by two image frames (hereinafter referred to as images) that can be distinguished and captured by an image sensor of the mobile computing device. The mobile computing device can identify two distinguishable images by correlating the image blur detected via the image sensor with the motion shocks or impulses detected via a motion sensor (such as an inertial measurement unit (IMU) or an accelerometer) of the mobile computing device. The motion detected by the mobile computing device can include the shocks or impulses detected when the mobile computing device initially starts moving and can also include the shocks or impulses detected when the mobile computing device finally stops moving. In this way, various embodiments can bind data from one or more sensors of the mobile computing device to specific images captured by the image sensor of the mobile computing device, which in turn can initialize the operation of the SLAM process. Additionally, various embodiments can allow the SLAM process of the embodiments to initialize each key image frame based on a previous image frame and use the IMU to determine an initial distance.
[0016] For some embodiments, the motion includes a side step performed by a human individual holding the mobile computing device, which can provide sufficient parallax for good SLAM initialization. In particular, embodiments can use the impulses generated at the start and end of the side step (such as based on a typical human side step) to analyze the motion of the human individual and extract relevant portions of the motion. Subsequently, embodiments can use those relevant portions to identify the first and second key image frames and initialize the SLAM process based on the first and second key image frames.
[0017] Some embodiments enable the SLAM process to determine the localization of the mobile computing device and map the physical environment of the mobile computing device (e.g., with an available or acceptable accuracy), while using a motion sensor (such as a noisy IMU) that provides poor accuracy. Some embodiments enable the SLAM process to determine the localization and map the physical environment while using a limited amount of image data. Some embodiments enable the SLAM process to determine the localization and map the physical environment in real time while using limited computing resources (such as a low-power processor). Additionally, some embodiments enable the SLAM process to determine the localization and map the physical environment without using depth data.
[0018] The SLAM techniques of some embodiments can be used to: track key points (tracking points) in two-dimensional (2D) image frames (such as of a video stream); and identify three-dimensional (3D) features (such as physical objects in a physical environment) in the 2D image frame and the relative physical pose (such as position) of the camera to the 3D features.
[0019] For example, the SLAM technology of the embodiments can be used in conjunction with augmented reality (AR) image processing and image frame tracking. Specifically, the SLAM technology can be used to track the image frames captured for an AR system, and then virtual objects can be placed within the captured image frames or relative to the captured image frames as part of an AR display of a device (such as smart glasses, a smartphone, a tablet, or another mobile computing device). As used herein, augmented reality (AR) refers to a system, method, device, and instructions that can capture image frames, enhance those image frames with additional information, and then present the enhanced information on a display. For example, this can enable a user to hold up a mobile computing device (such as a smartphone or a tablet) to capture a video stream of a scene, and the output display of the mobile computing device to present the scene as visible to the user along with additional information. The additional information can include placing virtual objects in the scene such that the virtual objects are presented as if they exist in the scene. Aspects of these virtual objects are processed to occlude the virtual objects if another real or virtual object passes in front of the virtual object from the perspective of an image sensor that captures the physical environment. Such virtual objects are also processed to maintain their relationship with real objects as both the real objects and the virtual objects move over time and as the perspective of the image sensor that captures the environment changes.
[0020] Some embodiments provide a method that includes: performing a loop of a full SLAM process, performing a loop of a partial SLAM process, and performing the loop of the partial SLAM process and the loop of the full SLAM process such that the loop of the partial SLAM process is performed more frequently than the loop of the full SLAM process. According to various embodiments, the loop of the full SLAM process performs the full SLAM process, while the loop of the partial SLAM process performs a partial SLAM process, and performing the partial SLAM process requires fewer computing resources (such as processing, memory resources, or both) than performing the full SLAM process. Additionally, the partial SLAM process can be performed faster than the loop of the full SLAM process.
[0021] For some embodiments, a partial SLAM process only performs the localization part of the SLAM process. In alternative embodiments, a partial SLAM process only performs the mapping part of the SLAM process. By only performing a part of the SLAM process, the partial SLAM process can be executed using fewer computational resources than the full SLAM process and can be executed faster than the full SLAM process. Additionally, by performing full SLAM process loops less frequently than partial SLAM process loops, various embodiments achieve SLAM results (such as useful and accurate SLAM results) while limiting the computer resources required to achieve those results. Thus, various embodiments are suitable for performing the SLAM process on a device that otherwise has limited computational resources for performing traditional SLAM techniques, such as a smartphone or smart glasses with limited processing power.
[0022] For some embodiments, image frames are captured (e.g., continuously captured at a specific sampling rate) by an image sensor of a device such as a camera of a mobile phone. Some embodiments perform full SLAM process loops on those captured image frames that are identified (e.g., generated) as new key image frames and perform partial SLAM process loops on those captured image frames that are not identified as key image frames. In some embodiments, a captured image frame is identified (e.g., generated) as a new key image frame when one or more key image frame conditions are met. Various embodiments use the key image frame conditions to ensure that the new key image frames identified from the captured image frames are unique enough to ensure that each full SLAM process loop is executed as desired or expected.
[0023] For example, a new key image frame can be generated when the captured image frame has at least a predetermined quality (e.g., fair image quality), if not better. In this way, embodiments can avoid designating those image frames captured during the movement of the image sensor as new image frames, which may capture image blurring caused by the movement of the image sensor. The image quality of the captured image frame can be determined by a gradient histogram method, which can determine the quality of the current image frame based on the quality of a predetermined number of previously captured image frames. In another example, a new key image frame can be generated only after a certain amount of time or a certain number of cycles (e.g., partial SLAM process cycles) have elapsed between the recognition (e.g., generation) of the last new key image frame. In this way, embodiments can avoid each captured image frame being considered a key image frame and being processed by a full SLAM cycle, as described herein, the execution of which can be processor-intensive or memory-intensive and not suitable for continuous execution on a device with limited computing resources. In another example, a new key image frame can be generated only after a certain amount of transformation (e.g., caused by a change in the position of the image sensor relative to the X, Y, or Z coordinates in the physical environment) is detected between the current captured image frame and the previous image frame. In this way, embodiments can avoid too many image frames that capture the same point in the physical environment from being designated as new key image frames, which are not helpful for three-dimensional (3D) mapping purposes.
[0024] For some embodiments, the full SLAM process cycle and the partial SLAM process cycle can be executed in parallel, whereby the full SLAM process cycle is executed only on those captured image frames identified as new key image frames, and the partial SLAM process cycle is executed on all other image frames captured between non-key image frames. Additionally, for some embodiments, after SLAM initialization is performed as described herein, the full SLAM process cycle and the partial SLAM process cycle begin execution. For example, the SLAM initialization process of an embodiment can generate the first two key image frames (e.g., based on a human individual's side step), provide initial positioning data (e.g., including six degrees of freedom (6DOF) for the second key image frame), and provide initial mapping data (e.g., including the 3D positions of features matched between the first two key image frames). Subsequently, based on the initial positioning and mapping data provided by the SLAM initialization process, the full SLAM process cycle and the partial SLAM process cycle can begin execution.
[0025] Although various embodiments regarding the use of an IMU are described herein, it should be understood that some embodiments may use one or more other sensors in addition to or instead of an IMU, such as an accelerometer or a gyroscope. As used herein, degrees of freedom (DOF) (e.g., measured by an IMU, accelerometer, or gyroscope) may include displacement (e.g., measured according to X, Y, and Z coordinates) and orientation (e.g., measured according to psi, theta, and phi). Thus, six degrees of freedom (6DOF) parameters may include values representing distances along the x-axis, y-axis, and z-axis, and values representing rotations according to the Euler angles psi, theta, and phi. Four degrees of freedom (4DOF) parameters may include values representing distances along the x-axis, y-axis, and z-axis, and values representing rotations according to an Euler angle (e.g., phi).
[0026] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody illustrative embodiments of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of the various embodiments of the subject matter of the present invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the present invention may be practiced without these specific details. In general, well-known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.
[0027] Reference will now be made in detail to embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein.
[0028] Figure 1 is a block diagram showing an exemplary client-server based advanced network architecture 100 including a Simultaneous Localization and Mapping (SLAM) system 126 according to some embodiments. As shown, the network architecture 100 includes: client devices 102A and client device 102B (collectively referred to hereinafter as client devices 102), the SLAM system 126 is included in the client device 102B; a messaging server system 108; and a network 106 (e.g., the Internet or a Wide Area Network (WAN)) that facilitates data communication between the client devices 102 and the messaging server system 108. In the network architecture 100, the messaging server system 108 may provide server-side functionality to the client devices 102 via the network 106. In some embodiments, a user (not shown) interacts with one of the client devices 102 or the messaging server system 108 using one of the client devices 102.
[0029] The client device 102 may include a computing device that includes at least display and communication capabilities to provide communication with the messaging server system 108 via the network 106. Each client device 102 may include, but is not limited to, a remote device, a workstation, a computer, a general-purpose computer, an Internet device, a handheld device, a wireless device, a portable device, a wearable computer, a cellular or mobile phone, a personal digital assistant (PDA), a smartphone, a tablet, an ultrabook, a netbook, a laptop, a desktop, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, a network personal computer (PC), a minicomputer, etc. According to an embodiment, at least one of the client devices 102 may include one or more of a touch screen, an inertial measurement unit (IMU), an accelerometer, a gyroscope, a biometric sensor, a camera, a microphone, a global positioning system (GPS) device, etc.
[0030] For some embodiments, the client device 102B represents a mobile computing device that includes an image sensor, such as a mobile phone, a tablet, or a wearable device (e.g., smart glasses, smart visor, or smart watch). As shown, the client device 102B includes a sensor 128, which may include an image sensor (e.g., a camera) of the client device 102B and other sensors (such as an inertial measurement unit (IMU), an accelerometer, or a gyroscope). For various embodiments, the sensor 128 facilitates the operation of the SLAM system 126 on the client device 102B.
[0031] The SLAM system 126 performs the SLAM technology of the embodiment on the client device 102B, which may allow the client device 102B to map its physical environment while determining its position within the physical environment. Additionally, for some embodiments, although the client device 102B has limited computing resources (e.g., processing or memory resources), the SLAM system 126 allows the SLAM technology to be executed on the client device 102B, which may prevent traditional SLAM technology from performing as expected on the client device 102B. The SLAM technology performed by the SLAM system 126 may support image frame tracking of the augmented reality system 124 of the client device 102B.
[0032] As shown in the figure, the client device 102B includes an augmented reality system 124, which can represent an augmented reality application operating on the client device 102B. The augmented reality system 124 can provide the function of generating augmented reality images for display on a display (such as an AR display) of the client device 102B. The network architecture 100 can be used to transmit information about virtual objects to be displayed on the client device 102B through the augmented reality system 124 included in the client device 102B, or to provide data (such as street view data) for creating models used by the augmented reality system 124. The SLAM system 126 can be used to track the image frames captured for the augmented reality system 124, and then virtual objects can be placed within or relative to the captured image frames as part of the AR display of the client device 102B.
[0033] Each client device 102 can host multiple applications, including a messaging client application 104, such as an ephemeral messaging application. Each messaging client application 104 can be communicatively coupled via a network 106 (such as the Internet) to other instances of the messaging client application 104 and the messaging server system 108. Thus, each messaging client application 104 may be able to communicate and exchange data with another messaging client application 104 and the messaging server system 108 via the network 106. The data exchanged between the messaging client applications 104 and between the messaging client application 104 and the messaging server system 108 can include functions (such as commands to invoke functions) and payload data (such as text, audio, video, or other multimedia data).
[0034] The messaging server system 108 provides server-side functions to a specific messaging client application 104 via the network 106. Although some functions of the network architecture 100 are described herein as being performed by the messaging client application 104 or by the messaging server system 108, it should be understood that the location of certain functions within the messaging client application 104 or the messaging server system 108 is a design choice. For example, it is technically preferred to initially deploy certain technologies and functions within the messaging server system 108, but later migrate the technologies and functions to the messaging client application 104 where the client device 102 has sufficient processing power.
[0035] The messaging server system 108 supports various services and operations provided to the messaging client application 104 or the augmented reality system 124. These operations include transmitting data to the messaging client application 104 or the augmented reality system 124, receiving data from the messaging client application 104 or the augmented reality system 124, and processing data generated by the messaging client application 104 or the augmented reality system 124. As an example, the data can include message content, client device information, geolocation information, media annotations and overlays, message content persistence conditions, social network information, augmented reality (AR) content, and live event information. Data exchange within the network architecture 100 is invoked and controlled through functions available via the user interface (UI) of the messaging client application 104 or the augmented reality system 124.
[0036] Turning now specifically to the messaging server system 108, the application programming interface (API) server 110 is coupled to the application server 112 and provides a programming interface thereto. The application server 112 is communicatively coupled to the database server 118, which facilitates access to the database 120 in which data associated with messages or augmented reality-related data processed by the application server 112 is stored.
[0037] Specifically processed by the application programming interface (API) server 110, which receives and sends message data (such as commands and message payloads) between the client device 102 and the application server 112. Specifically, the API server 110 provides a set of interfaces (such as routines and protocols) that the messaging client application 104 can call or query in order to invoke the functions of the application server 112. The API server 110 exposes various functions supported by the application server 112, including: account registration; login functionality; sending messages from a specific messaging client application 104 to another messaging client application 104 via the application server 112; sending media files (such as images or videos) from the messaging client application 104 to the messaging server application 114 and setting up a collection of media data (such as stories) for possible access by another messaging client application 104; retrieving a friend list of the user of the client device 102; such collection retrieval; message and content retrieval; adding and deleting friends in the social graph; locating friends in the social graph; opening application events (such as related to the messaging client application 104).
[0038] The application server 112 hosts multiple applications and subsystems, including a messaging server application 114, an image processing system 116, and a social networking system 122. The messaging server application 114 implements a variety of message processing techniques and functions, particularly related to the aggregation and other processing of content (such as text and multimedia content) included in messages received from multiple instances of the messaging client application 104. Text and media content from multiple sources can be aggregated into collections of content (such as those referred to as stories or galleries). The messaging server application 114 then makes these collections available to the messaging client application 104. Given the hardware requirements for such processing, other processor- and memory-intensive data processing can also be performed by the messaging server application 114 on the server side.
[0039] The application server 112 also includes an image processing system 116 that is dedicated to performing various image processing operations typically on images or videos received within the payloads of messages at the messaging server application 114.
[0040] The social networking system 122 supports various social networking function services and makes these functions and services available to the messaging server application 114. To this end, the social networking system 122 maintains and accesses an entity graph within the database 120. Examples of functions and services supported by the social networking system 122 include the identification of other users of the messaging system 108 with whom a particular user has a relationship or whom a particular user "follows" and the identification of other entities and interests of a particular user.
[0041] The application server 112 is communicatively coupled to a database server 118 that facilitates access to a database 120 in which data associated with messages or augmented reality content processed by the application server 112 is stored.
[0042] Figure 2 FIG. is a block diagram showing an example computing device 200 including a SLAM system 240 according to some embodiments. The computing device 200 can represent a mobile computing device that a human can easily carry and move around in a physical environment, such as a mobile phone, a tablet computer, a laptop computer, a wearable device, etc. As shown, the computing device 200 includes a processor 210, an image sensor 220, an inertial measurement unit (IMU) 230, and a SLAM system 240. The SLAM system 240 includes an image frame capture module 241, an IMU data capture module 242, a key image frame module 243, a complete SLAM loop module 244, and a partial SLAM loop module 245. Depending on the embodiment, the SLAM system 240 may or may not include a SLAM initialization module 246.
[0043] Any one or more functional components (e.g., modules) of the SLAM system 240 can be implemented using hardware (e.g., the processor 210 of the computing device 200) or a combination of hardware and software. For example, any one of the components described herein can configure the processor 210 to perform the operations described herein for that component. Additionally, any two or more of these components can be combined into a single component, and the functionality described herein for a single component can be subdivided among multiple components. Further, according to various example embodiments, Figure 2 Any of the illustrated functional components can be implemented together or separately within a single machine, database, or device, or can be distributed across multiple machines, databases, or devices.
[0044] The processor 210 can include a central processing unit (CPU), the image sensor 220 can include a camera built into or externally coupled to the computing device 200, and the IMU 230 can include sensors that are capable of measuring degrees of freedom (e.g., 6DOF) if not relative to the computing device 200 then at least relative to the image sensor 220. Although not shown, the computing device 200 can include other sensors to facilitate the operation of the SLAM system 240, such as an accelerometer or a gyroscope.
[0045] The image frame capture module 241 can invoke, facilitate, or perform the continuous capture of new image frames of the physical environment of the computing device 200 via the image sensor 220. The continuous capture can be performed according to a predetermined sampling rate, such as 25 or 30 frames per second. The image frame capture module 241 can add the new image frames continuously captured by the image sensor 220 to a set of captured image frames, which can be further processed by the SLAM system 240.
[0046] The IMU data capture module 242 can invoke, facilitate, or perform the continuous capture of IMU data from the IMU 230 corresponding to the image frames captured by the image frame capture module 241. For example, the IMU data capture module 242 can capture IMU data for each captured image frame. For a given captured image frame, the captured IMU data can include the degrees of freedom (DOF) parameters of the image sensor at the time the image sensor captured the image frame. The DOF parameters can include, for example, four degrees of freedom (4DOF) or six degrees of freedom (6DOF) measured relative to the image sensor 220. Where the IMU 230, the image sensor 220, and the computing device 200 are physically integrated as a single unit, the IMU data can reflect the DOF parameters of the image sensor 220 and the computing device 200.
[0047] For each particular new image frame added to the set of captured image frames (e.g., via the image frame capture module 241), the key image frame module 243 can determine whether a set of key image frame conditions is satisfied for that particular new image frame. In response to the set of key image frame conditions being satisfied for a particular new image frame, the key image frame module 243 can identify that particular new image frame as a new key image frame. Thus, the key image frame module 243 can generate new key images frames based on the set of key image frame conditions. As described herein, the set of key image frame conditions can ensure that the new key image frames are unique enough to be processed through a full SLAM process loop. Example key image frame conditions can relate to whether the new image frame meets or exceeds a particular image quality, whether a minimum amount of time has elapsed since the last execution of the full SLAM process loop, or whether the transformation between a previous image frame and the new image frame meets or exceeds a minimum transformation threshold.
[0048] The full SLAM loop module 244 can perform a full SLAM process loop on each particular new key image frame identified by the key image frame module 243. Performing a full SLAM process loop on a particular new key image frame can include determining the 6DOF of the image sensor of the computing device associated with that particular new key image frame. Additionally, performing a full SLAM process loop on a particular new key image frame can include determining a set of 3D positions of new 3D features that match in that particular new key image frame. More on partial SLAM process loops is referenced herein Figure 6 described.
[0049] The partial SLAM loop module 245 can perform a partial SLAM process loop on each particular new image frame not identified by the key image frame module 243. For some embodiments, the partial SLAM process loop only performs the localization portion of the SLAM process. Performing a partial SLAM process loop on a particular new image frame can include determining the 6DOF of the image sensor 220 of the computing device 200 associated with the particular new image frame. Additionally, performing a partial SLAM process loop on a particular new image frame can include projecting a set of tracking points onto the particular new image frame based on the 6DOF of the image sensor 220. Alternatively, for some embodiments, the partial SLAM process loop only performs the mapping portion of the SLAM process. More on partial SLAM process loops is described herein Figure 7 described.
[0050] The SLAM initialization module 246 can detect the movement of the image sensor 220 from a first pose (e.g., the orientation or position of the image sensor 220) in a physical environment to a second pose in the physical environment based on the captured IMU data from the IMU data capture module 242. The SLAM initialization module 246 can identify a first key image frame and a second key image frame based on the movement. In particular, the first and second key image frames can be identified such that the first key image frame corresponds to the start impulse of the movement and the second key image frame corresponds to the end impulse of the movement. For example, in the case where a human individual holds the computing device 200, the start impulse of the movement can be the start of a side step performed by the human individual, and the end impulse of the movement can be the end of the side step. The impact function of the computing device 200 can be used to detect the start or end impulse.
[0051] Figures 3 - 7 is a flowchart showing an example method for a SLAM process according to various embodiments. It should be understood that, according to some embodiments, the example methods described herein can be performed by a device such as a computing device (e.g., the computing device 200). Additionally, the example methods described herein can be implemented in the form of executable instructions stored on a computer-readable medium or in the form of an electronic circuit. For example, Figure 3 one or more operations of the method 300 can be represented by executable instructions that, when executed by a processor of a computing device, cause the computing device to perform the method 300. Depending on the embodiment, the operations of the example methods described herein can be repeated in different ways or involve intermediate operations not shown. Although the operations of the example methods can be depicted and described in a specific order, the order of performing the operations can vary between embodiments, including performing some operations in parallel.
[0052] Figure 3 is a flowchart showing an example method 300 for a SLAM process according to some embodiments. In particular, the method 300 shows how an embodiment can perform a full SLAM process loop and a partial SLAM process loop. As shown, the method 300 begins at operation 302 by invoking, facilitating, or executing the continuous capture of new image frames of the physical environment of the computing device by the image sensor of the computing device. Operation 302 adds the new image frames to a set of captured image frames that can be further processed by the method 300. The method 300 continues to operation 304, which corresponds to the image frames captured by operation 302 and invokes, facilitates, or executes the continuous capture of IMU data from the inertial measurement unit of the computing device (IMU). As described herein, the IMU data for a specific image frame can include the degrees of freedom (DOF) of the image sensor measured by the IMU when the image frame was captured in operation 302.
[0053] Method 300 proceeds to operation 306, and operations 320 through 326 are performed for each particular new image frame captured by operation 302 and added to the set of captured image frames. Operation 306 begins with operation 320 determining whether a set of key image frame conditions is satisfied for the particular new image frame. Operation 306 proceeds to operation 322, in response to operation 320 determining that the set of key image frame conditions is satisfied for the particular new image frame, identifying the particular new image frame as a new key image frame.
[0054] Operation 306 proceeds to operation 324, performing a full SLAM process loop on the new key image frame. For future processing purposes, some embodiments track those image frames identified as key image frames. Performing a full SLAM process loop on a particular new key image frame can include determining the 6DOF of the image sensor of the computing device associated with the particular new key image frame. Additionally, performing a full SLAM process loop on the particular new key image frame can include determining a set of 3D positions of new 3D features matched in the particular new key image frame. More on the partial SLAM process loop is referenced herein Figure 6 described.
[0055] Operation 306 proceeds to operation 326, in response to operation 320 determining that the set of key image frame conditions is not satisfied for a particular new image frame (i.e., a non-key new image frame), performing a partial SLAM process loop on the particular new image frame. Performing a partial SLAM process loop on a particular new image frame can include determining the 6DOF of the image sensor of the computing device associated with the particular new image frame. Additionally, performing a partial SLAM process loop on the particular new image frame can include projecting a set of tracking points onto the particular new image frame based on the 6DOF of the image sensor. More on the Figure 7 partial SLAM process loop is described herein.
[0056] Figure 4 is a flowchart showing an example method 400 for a SLAM process according to some embodiments. In particular, method 400 shows how embodiments initialize a SLAM process, perform a full SLAM process loop, and perform a partial SLAM process loop. As shown, method 400 begins with operations 402 and 404, which, according to some embodiments, are respectively similar to operations 302 and 304 of method 300 described above with respect to Figure 3 described.
[0057] Method 400 proceeds to operation 406, which, based on the captured IMU data, detects the movement of the image sensor from a first pose (e.g., the orientation or position of the image sensor) in the physical environment to a second pose in the physical environment. Method 400 proceeds to operation 408, which, based on the movement detected in operation 406, identifies a first key image frame and a second key image frame. For some embodiments, the first key image frame corresponds to the start impulse of the movement, and the second key image frame corresponds to the end impulse of the movement. By operations 406 and 408, some embodiments can initialize Method 400 to perform a full and partial SLAM process loop at operation 410. As described herein, the movement can be caused by a human individual performing a side step while holding a computing device that executes Method 400 and includes an image sensor.
[0058] Method 400 proceeds to operation 410, which performs operations 420 through 426 for each specific new image frame captured by operation 402 and added to the set of captured image frames. According to some embodiments, operations 420 through 426 are respectively similar to operations 320 through 326 of Method 300 described above with reference to Figure 3 the description.
[0059] Figure 5 is a flowchart showing an example method 500 for a SLAM process according to some embodiments. In particular, Method 500 shows how an embodiment initializes a SLAM process. As shown, Method 500 begins with operations 502 and 504, which, according to some embodiments, are respectively similar to operations 302 and 304 of Method 300 described above with reference to Figure 3 the description.
[0060] Method 500 proceeds to operation 506, which identifies a first key image frame from the set of captured image frames. The identified first key image frame can include a specific image quality (e.g., fair quality) and can be an image frame captured by the image sensor (e.g., image sensor 220) when the IMU (e.g., IMU 230) indicates that the image sensor is stable. Thus, operation 506 may not identify the first key image frame until an image frame is captured when the image sensor is stable and the captured image frame meets the specific image quality.
[0061] Method 500 proceeds to operation 508, which identifies first IMU data associated with the first key image frame from the IMU data captured by operation 504. For some embodiments, the first IMU data includes 4DOF parameters (e.g., x, y, z, and phi). The first IMU data can represent the IMU data captured when the image sensor captured the first key image frame.
[0062] Method 500 proceeds to operation 510, where an IMU detects movement of an image sensor from a first pose (e.g., orientation or position) in a physical environment to a second pose in the physical environment. As described herein, the movement can be caused by a human individual who performs a side step while holding a computing device that executes Method 500 and includes an image sensor and an IMU.
[0063] Method 500 proceeds to operation 512, which executes operations 520 through 528 in response to detecting movement via operation 510. Operation 512 begins with operation 520 identifying a second key image frame from the set of captured image frames. While the first key image frame can be identified via operation 506 such that the first key image frame corresponds to the start of the movement detected via operation 510, the second key image frame can be identified via operation 520 such that the second key image frame corresponds to the end of the movement detected via operation 510.
[0064] Operation 512 proceeds to operation 522, which identifies second IMU data associated with the second key image frame from the IMU data captured via operation 504. For some embodiments, the second IMU data includes 4DOF parameters (e.g., x, y, z, and phi). The second IMU data can represent the IMU data captured when the image sensor captured the second key image frame.
[0065] Operation 512 proceeds to operation 524, which performs feature matching on at least the first and second key image frames to identify a set of matching 3D features in the physical environment. For some embodiments, operation 524 uses a KAZE- or A-KAZE-based feature matcher that extracts 3D features from the set of image frames by matching features across the image frames. Operation 512 proceeds to operation 526, which generates a filtered set of matching 3D features by filtering out at least one error feature from the set of matching 3D features generated via operation 524 based on a set of error criteria. For example, the set of error criteria can include error criteria related to the polar axis, projection error, or spatial error. If a feature error is found, Method 500 can return to operation 524 to perform feature matching again. Operation 512 proceeds to operation 528, which determines a set of 6DOF parameters for the image sensor for the second key image frame and a set of 3D positions for the set of matching 3D features. To facilitate this determination, operation 512 performs a (complete) SLAM process on the second key image frame based on the first IMU data identified via operation 506, the second IMU data identified via operation 522, and the filtered set of matching 3D features extracted via operation 526..
[0066] Figure 6It is a flowchart showing an example method 600 for the SLAM process according to some embodiments. In particular, method 600 shows how an embodiment performs a complete SLAM process loop. For some embodiments, method 600 is not executed until at least two key image frames are generated by the SLAM initialization process (e.g., method 500). As shown, method 600 begins at operation 602 by identifying specific IMU data associated with a new key image frame from captured IMU data. The IMU data may represent the IMU data captured when the image sensor captures the new key image frame.
[0067] Method 600 continues to operation 604, which performs feature matching on the new key image frame and at least one previous image frame (e.g., the last two captured image frames) to identify a set of matching 3D features in the physical environment. For some embodiments, operation 604 uses a KAZE- or A-KAZE-based feature matcher that extracts 3D features from a set of image frames by matching features on the image frames. Method 600 continues to operation 606, which performs a (complete) SLAM process on the new key image frame based on the set of matching 3D features extracted by operation 604 and the specific IMU data identified by operation 602 to determine a first set of 6DOF parameters of the image sensor for the new key image frame.
[0068] Method 600 continues to operation 608, which generates a filtered set of matching 3D features by filtering out at least one error feature from the set of matching 3D features extracted by operation 604 based on a set of error criteria and the first set of 6DOF parameters determined by operation 606. As described herein, the set of error criteria may include error criteria related to epipolar axis, projection error, or spatial error. For example, the error criteria may specify filtering out those features representing the top 3% of the worst projection error.
[0069] Method 600 continues to operation 610 by performing a (complete) SLAM process on all key image frames based on the second set of filtered matching 3D features generated by operation 608 and the specific IMU data identified by operation 602 to determine a second set of 6DOF parameters of the image sensor for the new key image frame and a set of 3D positions of new 3D features in the physical environment.
[0070] Figure 7is a flowchart showing an example method 700 for a SLAM process according to some embodiments. In particular, method 700 shows how an embodiment performs a partial SLAM process loop. As shown, method 700 begins at operation 702, where feature tracking is performed on non-critical image frames based on the 3D positions of new 3D features provided (e.g., extracted) by the last execution of a complete SLAM process loop (e.g., method 600) and a set of new image frames most recently processed by the complete SLAM process loop (e.g., via method 600). For some embodiments, operation 702 uses a 2D tracker based on the Kanade-Lucas-Tomasi (KLT) method, which extracts 2D features from new key image frames processed by the last execution of the complete SLAM process loop. Method 700 continues to operation 704, where a set of 6DOF parameters for the image sensor for the non-critical image frames is determined by performing only the localization portion of the SLAM process based on the set of 2D features from operation 702. Method 700 continues to operation 706, where a filtered set of 2D features is generated by filtering out at least one error feature from the set of 2D features identified in operation 702 based on a set of error criteria and the set of 6DOF parameters determined in operation 704. The set of error criteria can include, for example, error criteria related to epipolar axes, projection errors, or spatial errors. Method 700 continues to operation 708, where a set of tracked points is projected onto the non-critical image frames based on the filtered set of 2D features generated in operation 706 and the set of 6DOF parameters determined in operation 704. For some embodiments, the set of tracked points allows for 2D virtual tracking on the non-critical image frames, which can be useful in applications such as augmented reality.
[0071] Software architecture
[0072] Figure 8 is a block diagram showing an example software architecture 806 that can be used in conjunction with various hardware architectures described herein to implement embodiments. Figure 8 is a non-limiting example of a software architecture, and it should be understood that many other architectures can be implemented to facilitate the functions described herein. Software architecture 806 can be executed on hardware such as Figure 9 a machine 900 (which particularly includes a processor 904, a memory 914, and I / O components 918, etc.). A representative hardware layer 852 is shown and can represent, for example, Figure 9 a machine 900. The representative hardware layer 852 includes a processor unit 854 with associated executable instructions 804. The executable instructions 804 represent the executable instructions of software architecture 806, including the implementation of the methods, components, etc. described herein. The hardware layer 852 also includes a memory and / or storage module memory / memory 856 that also has executable instructions 804. The hardware layer 852 can also include other hardware 858.
[0073] In Figure 8 the example architecture, the software architecture 806 can be conceptualized as a stack of layers, where each layer provides a specific function. For example, the software architecture 806 can include layers such as an operating system 802, libraries 820, framework / middleware 818, applications 816, and a presentation layer 814. In operation, an application 816 and / or other components in a layer can make application programming interface (API) calls 808 through the software stack and receive responses as messages 812. The layers shown are representative in nature, and not all software architectures have all layers. For example, some mobile or specialized operating systems may not provide framework / middleware 818, while others may provide such layers. Other software architectures can include additional or different layers.
[0074] The operating system 802 can manage hardware resources and provide common services. The operating system 802 can include, for example, a kernel 822, services 824, and drivers 826. The kernel 822 can act as an abstraction layer between the hardware and other software layers. For example, the kernel 822 can be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. The services 824 can provide other common services for other software layers. The drivers 826 are responsible for controlling or interfacing with the underlying hardware. For example, depending on the hardware configuration, the drivers 826 include a display driver, a camera driver, a driver, a flash drive, a serial communication driver (e.g., a universal serial bus (USB) driver), a driver, an audio driver, a power management driver, etc.
[0075] The library 820 provides a common infrastructure used by the application 816 and / or other components and / or layers. The library 820 provides functions that allow other software components to perform tasks in a way that is easier than directly interfacing with the underlying operating system 802 functions (such as the kernel 822, services 824, and / or drivers 826). The library 820 may include system libraries 844 (such as the C standard library), which may provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the library 820 may include API libraries 846, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG), graphics libraries (e.g., the OpenGL framework that can be used to render 2D and 3D in graphical content on a display), database libraries (e.g., SQLite that can provide various relational database functions), web libraries (e.g., WebKit that can provide web browsing functions), etc. The library 820 may also include various other libraries 848 to provide many other APIs to the application 816 and other software components / modules.
[0076] The framework / middleware 818 (sometimes also referred to as middleware) provides a higher-level common infrastructure that can be used by the application 816 and / or other software components / modules. For example, the framework / middleware 818 may provide various graphical user interface (GUI) functions, advanced resource management, advanced location services, etc. The framework / middleware 818 may provide a wide range of other APIs that can be used by the application 816 and / or other software components / modules, some of which may be specific to a particular operating system 802 or platform.
[0077] The application 816 includes built-in applications 838 and / or third-party applications 840. Examples of representative built-in applications 838 may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, and / or game applications. The third-party applications 840 may include applications developed using the ANDROID TM or IOS TM software development kit (SDK), and may be mobile software running on a mobile operating system such as IOS TM 、ANDROID TM 、 、Phone or other mobile operating systems. The third-party applications 840 may call API calls 808 provided by the mobile operating system (such as the operating system 802) to facilitate the functions described herein.
[0078] Application 816 can use built-in operating system functions (such as kernel 822, services 824, and / or drivers 826), libraries 820, and frameworks / middleware 818 to create a user interface for interacting with the users of the system. Alternatively or additionally, in some systems, the interaction with the users can occur through a presentation layer (such as presentation layer 814). In these systems, the application / component “logic” can be separated from the aspects of the application / component that interact with the users.
[0079] Figure 9 is a block diagram showing components of a machine 900 that can read instructions 804 from a machine-readable medium (such as a machine-readable storage medium) and perform any one or more of the methods discussed herein. Specifically, Figure 9 shows a graphical representation of a machine 900 in an example form of a computer system, in which instructions 910 (such as software, programs, applications, applets, apps, or other executable code) can be executed to cause the machine 900 to perform any one or more of the methods discussed herein. Thus, the instructions 910 can be used to implement the modules or components described herein. The instructions 910 transform the general, unprogrammed machine 900 into a particular machine 900 programmed to perform the described and shown functions in the described manner. In alternative embodiments, the machine 900 operates as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 900 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 900 can include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smart phones, mobile computing devices, wearable devices (such as smart watches), smart home devices (such as smart appliances), other smart devices, web devices, network routers, network switches, bridges, or any machine capable of sequentially or otherwise executing the instructions 910 specifying the actions to be taken by the machine 900. Further, although only a single machine 900 is shown, the term “machine” shall also be taken to include a collection of machines that individually or jointly execute the instructions 910 to perform any one or more of the methods discussed herein.
[0080] Machine 900 may include a processor 904, a memory / storage device 906, and I / O components 918 that may be configured to communicate with each other, for example, via a bus 902. The memory / storage device 906 may include a memory 914, such as a main memory or other storage device, and a storage unit 916, both of which may be accessed by the processor 904, for example, via the bus 902. The storage unit 916 and the memory 914 store instructions 910 that embody any one or more of the methods or functions described herein. The instructions 910 may also reside, in whole or in part, within the memory 914, within the storage unit 916, within at least one memory of the processor 904 (such as within a cache memory of the processor), or in any suitable combination thereof during execution by the machine 900. Accordingly, the memory 914, the storage unit 916, and the memory of the processor 904 are examples of machine-readable media.
[0081] The I / O components 918 may include a variety of components to receive input, provide output, generate output, send information, exchange information, capture measurements, and so on. The specific I / O components 918 included in a particular machine 900 will depend on the type of the machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that the I / O components 918 may include Figure 9 many other components not shown. The I / O components 918 are grouped according to functionality merely to simplify the following discussion, and the grouping is in no way restrictive. In various example embodiments, the I / O components 918 may include an output component 926 and an input component 928. The output component 926 may include a visual component (such as a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), an acoustic component (such as a speaker), a haptic component (such as a vibration motor, a resistive mechanism), other signal generators, and so on. The input component 928 may include an alphanumeric input component (such as a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), a point-based input component (such as a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), a haptic input component (such as a physical button, a touch screen that provides the location and / or force of a touch or touch gesture, or other haptic input components), an audio input component (such as a microphone), and so on.
[0082] In other example embodiments, the I / O component 918 may include a biometric component 930, a motion component 934, an environmental component 936, or a location component 938 among a wide range of other components. For example, the biometric component 930 may include components for detecting expressions (such as hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (such as blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (such as voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion component 934 may include an acceleration sensor component (such as an accelerometer), a gravity sensor component, a rotational sensor component (such as a gyroscope), etc. The environmental component 936 may include, for example, a lighting sensor component (such as a photometer), a temperature sensor component (such as one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (such as a barometer), an acoustic sensor component (such as one or more microphones for detecting background noise), a proximity sensor component (such as an infrared sensor for detecting nearby objects), a gas sensor (such as a gas detection sensor for detecting the concentration of hazardous gases to ensure safety or measuring air pollutants), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. The location component 938 may include a positioning sensor component (such as a Global Positioning System (GPS) receiver component), an altitude sensor component (such as an altimeter or barometer for detecting air pressure from which altitude can be derived), an azimuth sensor component (such as a magnetometer), etc.
[0083] Various techniques may be used to implement communication. The I / O component 918 may include a communication component 940, which is operable to couple the machine 900 to the network 932 or the device 920 via the couplings 924 and 922, respectively. For example, the communication component 940 may include a network interface component or other suitable devices interfacing with the network 932. In other examples, the communication component 940 may include a wired communication component, a wireless communication component, a cellular communication component, a Near Field Communication (NFC) component, components (such as low energy ), components, and other communication components that provide communication via other means. The device 920 may be another machine or any of a variety of peripheral devices (such as a peripheral device coupled via a Universal Serial Bus (USB)).
[0084] In addition, communication component 940 may detect an identifier or include components operable to detect an identifier. For example, communication component 940 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as universal product code (UPC) barcodes, multi-dimensional barcodes such as quick response (QR) codes, Aztec codes, data matrix, Dataglyph (data format), MaxiCode, PDF417, hypercode, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various information may be derived via communication component 940, such as a location derived via Internet protocol (IP) geolocation, a location derived via signal triangulation, a location derived via detecting an NFC beacon signal that may indicate a particular location, and so on.
[0085] As used herein, a "transient message" may refer to a message that may be accessible for a limited duration of time (e.g., up to 10 seconds). Transient messages may include text content, image content, audio content, video content, etc. The access time of a transient message may be set by the message sender, or alternatively, the access time may be a default setting or a setting specified by the recipient. Regardless of the setting technique, transient messages are temporary. The message duration parameter associated with a transient message may provide a value that determines the amount of time the receiving user of the transient message may display or access the transient message. A messaging client software application (e.g., a transient messaging application) capable of receiving and displaying the content of a transient message may be used to access or display a transient message.
[0086] As also used herein, a "transient message story" may refer to a collection of transient message content that may be accessible for a limited duration of time, similar to a transient message. A transient message story may be sent from one user to another and may be accessed or displayed using a messaging client software application (e.g., a transient messaging application) capable of receiving and displaying a collection of transient content.
[0087] Throughout the specification, multiple instances may implement components, operations, or structures described as a single instance. Although the various operations of one or more methods are shown and described as separate operations, one or more of the various operations may be performed simultaneously and the operations need not be performed in the order shown. Structures and functions that are presented as separate components in an example configuration may be implemented as a combined structure or component. Similarly, structures and functions presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0088] Although the subject matter of the present invention has been described in an overview with reference to specific example embodiments, various modifications and changes can be made to these embodiments without departing from the broader scope of the embodiments of the present disclosure.
[0089] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments can be used and other embodiments can be derived therefrom, such that structural and logical substitutions and changes can be made without departing from the scope of the present disclosure. Accordingly, the detailed description should not be construed in a limiting sense, and the scope of the various embodiments is defined only by the appended claims and the full scope of equivalents given by those claims.
[0090] As used herein, a module can comprise a software module (e.g., code stored or otherwise embodied in a machine-readable medium or a transmission medium), a hardware module, or any suitable combination thereof. A “hardware module” is a tangible (e.g., non-transitory) physical component (e.g., a set of one or more processors) that is capable of performing certain operations and can be configured or arranged in a particular physical manner. In various embodiments, one or more computer systems or one or more of their hardware modules can be configured by software (e.g., an application or a portion thereof) as a hardware module for performing the operations described herein for that module.
[0091] In some embodiments, a hardware module can be implemented electronically. For example, a hardware module can include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware module can be or can include a dedicated processor, such as a field programmable gate array (FPGA) or an ASIC. A hardware module can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. As an example, a hardware module can include software contained within a CPU or other programmable processor.
[0092] Considering embodiments where a hardware module is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate every hardware module at any one time. For example, in the case where a hardware module includes a CPU that is configured by software to be a dedicated processor, the CPU can be configured to be different dedicated processors (e.g., each included in a different hardware module) at different times. Software (e.g., a software module) can accordingly configure one or more processors, e.g., to become or otherwise constitute a particular hardware module at one time and to become or otherwise constitute a different hardware module at a different time.
[0093] Hardware modules can provide information to other hardware modules and receive information from them. Thus, the described hardware modules can be considered communicatively coupled. In the case where multiple hardware modules are present simultaneously, communication can be achieved through signal transmission between or among two or more hardware modules (e.g., via suitable circuits and buses). In embodiments where multiple hardware modules are configured or instantiated at different times, communication between these hardware modules can be achieved, for example, by storing and retrieving information in a memory structure accessible to the multiple hardware modules. For example, one hardware module can perform an operation and store the output of the operation in a memory (e.g., a memory device) communicatively coupled to it. Then, another hardware module can later access the memory to retrieve and process the stored output. Hardware modules can also initiate communication with input or output devices and can operate on resources (e.g., a collection of information from computing resources).
[0094] The various operations of the example methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily configured or permanently configured, such processors can constitute processor-implemented modules for performing one or more of the operations or functions described herein. As used herein, "processor-implemented module" refers to a hardware module in which the hardware includes one or more processors. Thus, the operations described herein can be performed, at least in part, by processor implementation, hardware implementation, or a combination of both, since a processor is an example of hardware and at least some of the operations within any one or more of the methods discussed herein can be performed by a module implemented by one or more processors, a hardware-implemented module, or any suitable combination thereof.
[0095] As used herein, the term "or" can be interpreted in either an inclusive or exclusive sense. The term "a" or "an" shall be understood to mean "at least one", "one or more", etc. The use of words and phrases such as "one or more", "at least", "but not limited to" or other similar phrases should not be understood to imply a narrower scope is intended or required in the absence of such expansive phrases.
[0096] The boundaries between various resources, operations, modules, engines, and data stores are to some extent arbitrary and particular operations are shown in the context of a particular illustrative configuration. Other function assignments are envisioned and can fall within the scope of the various embodiments of the present disclosure. Generally, structures and functions presented as separate resources in an example configuration can be implemented as a combined structure or resource. Similarly, structures and functions presented as a single resource can be implemented as separate resources. These and other variations, modifications, additions, and improvements fall within the scope of the embodiments of the present disclosure as represented by the appended claims. Thus, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0097] The foregoing description includes systems, methods, devices, instructions, and computer media (such as computer program products) that embody illustrative embodiments of the present disclosure. In the specification, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of the various embodiments of the subject matter of the present invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the present invention may be practiced without these specific details. In general, well-known instruction examples, protocols, structures, and techniques are not necessarily shown in detail.
Claims
1. A method for simultaneous localization and mapping, comprising: identifying, by one or more hardware processors, a first key image frame from a set of image frames captured by an image sensor; and after movement of the image sensor from a first pose in a physical environment to a second pose in the physical environment: identifying, by the one or more hardware processors, a second key image frame from the set of captured image frames; performing, by the one or more hardware processors, feature matching on at least the first key image frame and the second key image frame to identify a set of matching three-dimensional (3D) features in the physical environment; generating, by the one or more hardware processors, a filtered set of matching 3D features by filtering out at least one error feature from the set of matching 3D features based on a set of error criteria; and determining, by the one or more hardware processors, a first set of six degrees of freedom (6DOF) parameters of the image sensor for the second key image frame and a set of 3D positions for the set of matching 3D features, the determining including performing a simultaneous localization and mapping (SLAM) process based on first inertial measurement unit (IMU) data captured by the IMU, second IMU data captured by the IMU, and the filtered set of matching 3D features, the first IMU data being associated with the first key image frame and the second IMU data being associated with the second key image frame.
2. The method according to claim 1, wherein The first IMU data includes a set of four degrees of freedom (4DOF) parameters of the image sensor, and the second IMU data includes a set of 4DOF parameters of the image sensor.
3. The method according to claim 1, wherein, The image sensor and the IMU are included in a device.
4. The method according to claim 3, wherein The movement of the image sensor is caused by a human individual holding the device performing a side step.
5. The method according to claim 4, wherein, The first key image frame is identified by detecting the start impulse of the side step, and the first key image frame is a specific image frame in the set of captured image frames corresponding to the detected start impulse.
6. The method according to claim 4, wherein The second key image frame is identified by detecting the end impulse of the side step, and the second key image frame is a specific image frame in the set of captured image frames corresponding to the detected end impulse.
7. The method according to claim 1, further comprising: for each specific new image frame added to the set of captured image frames: determining, by the one or more hardware processors, whether a set of key image frame conditions for the specific new image frame is satisfied; identifying, by the one or more hardware processors, the specific new image frame as a new key image frame in response to the set of key image frame conditions for the specific new image frame being satisfied, and performing, by the one or more hardware processors, a complete SLAM process loop on the new key image frame; and The one or more hardware processors execute a partial SLAM process loop on the specific new image frame that is a non-critical image frame in response to the set of key image frame conditions not being satisfied for the specific new image frame, the partial SLAM process loop including only the localization portion of the SLAM process.
8. The method according to claim 7, wherein The set of key image frame conditions includes at least one of the following: the new image frame meets or exceeds a specific image quality, a minimum time has elapsed since the last execution of the full SLAM process loop, and the transformation between the previous image frame and the new image frame meets or exceeds a minimum transformation threshold.
9. The method according to claim 7, wherein, Executing the full SLAM process loop on the new key image frame includes: The one or more hardware processors identify third IMU data associated with the new key image frame; The one or more hardware processors perform feature matching on the new key image frame and at least one previous image frame to identify a second set of matching 3D features in the physical environment; The one or more hardware processors determine a second set of 6DOF parameters for the image sensor for the new key image frame by performing a SLAM process on the new key image frame based on the second set of matching 3D features and the third IMU data; The one or more hardware processors generate a second filtered set of matching 3D features by filtering out at least one error feature from the second set of matching 3D features based on a second set of error criteria and the second set of 6DOF parameters; and The one or more hardware processors determine a third set of 6DOF parameters for the image sensor for the new key image frame and a set of 3D positions of new 3D features in the physical environment by performing the SLAM process on all key image frames based on the second filtered set of matching 3D features and the third IMU data.
10. The method according to claim 7, wherein, Executing the partial SLAM process loop on the non-critical image frame includes: The one or more hardware processors perform two-dimensional (2D) feature tracking on the non-critical image frame based on the set of 3D positions of the new 3D features from the execution of the full SLAM process loop and the most recently identified new key image frame to identify a set of 2D features; The one or more hardware processors determine a fourth set of 6DOF parameters for the image sensor for the non-critical image frame by performing only the localization portion of the SLAM process based on the set of 2D features; The one or more hardware processors generate a filtered set of 2D features by filtering out at least one error feature from the set of 2D features based on a third set of error criteria and the fourth set of 6DOF parameters; and The one or more hardware processors project a set of tracking points on the non-critical image frame based on the filtered set of 2D features and the fourth set of 6DOF parameters.
11. A system for simultaneous localization and mapping, comprising: An image sensor; A memory storing instructions; And A hardware processor communicatively coupled to the memory and configured by the instructions to perform operations including the following: For each specific new image frame in a set of captured image frames of a physical environment captured by the image sensor: Determine, by one or more hardware processors, whether a set of key image frame conditions is satisfied for the specific new image frame; In response to the set of key image frame conditions being satisfied for the specific new image frame, identify the specific new image frame as a new key image frame and perform a full Simultaneous Localization and Mapping (SLAM) process loop on the new key image frame; and In response to the set of key image frame conditions not being satisfied for the specific new image frame, perform a partial SLAM process loop on the specific new image frame as a non-key image frame.
12. The system according to claim 11, further comprising: Detect, based on IMU data captured from an Inertial Measurement Unit (IMU) of the system, movement of the image sensor from a first pose in the physical environment to a second pose in the physical environment; And Identify a first key image frame and a second key image frame based on the movement, the first key image frame corresponding to the start impulse of the movement and the second key image frame corresponding to the end impulse of the movement.
13. The system according to claim 11, wherein, The set of key image frame conditions includes at least one of the following: the new image frame meets or exceeds a specific image quality, a minimum time has elapsed since the last execution of the full SLAM process loop, and the transformation between a previous image frame and the new image frame meets or exceeds a minimum transformation threshold.
14. The system according to claim 11, wherein: Performing the full SLAM process loop on the new key image frame includes: Identifying second IMU data associated with the new key image frame from the captured IMU data, the captured IMU data being captured from the IMU; Performing feature matching on the new key image frame and at least one previous image frame to identify a second set of matching 3D features in the physical environment; Determining a first set of 6DOF parameters of the image sensor for the new key image frame by performing a SLAM process on the new key image frame based on the second set of matching 3D features and the second IMU data; Generating a second filtered set of matching 3D features by filtering out at least one error feature from the second set of matching 3D features based on a second set of error criteria and the first set of 6DOF parameters; and Determining a second set of 6DOF parameters of the image sensor for the new key image frame and a set of 3D positions of new 3D features in the physical environment by performing the SLAM process on all key image frames based on the second filtered set of matching 3D features and the second IMU data.
15. The system according to claim 11, wherein Performing the partial SLAM process loop on the non-key image frame includes: Performing two-dimensional (2D) feature tracking on the non-key image frame based on the set of 3D positions of new 3D features from the execution of the full SLAM process loop and the most recently identified new key image frame to identify a set of 2D features; Determine a third set of 6DOF parameters for the image sensor for the non-critical image frame by performing only the localization portion of the SLAM process based on the 2D feature set; Generate a filtered 2D feature set by filtering at least one error feature from the 2D feature set based on a third set of error criteria and the third set of 6DOF parameters; and Project a set of tracking points on the non-critical image frame based on the filtered 2D feature set and the third set of 6DOF parameters.
16. A system for simultaneous localization and mapping, comprising: An inertial measurement unit IMU; An image sensor; A memory storing instructions; And A hardware processor communicatively coupled to the memory and configured by the instructions to perform operations including: Identify a first key image frame from a set of image frames captured by the image sensor; And After the image sensor moves from a first pose in a physical environment to a second pose in the physical environment: Identify a second key image frame from the set of captured image frames; Perform feature matching on at least the first key image frame and the second key image frame to identify a set of matching three-dimensional 3D features in the physical environment; Generate a filtered set of matching 3D features by filtering at least one error feature from the set of matching 3D features based on a set of error criteria; And Determine a first set of six degrees of freedom 6DOF parameters for the image sensor for the second key image frame and a set of 3D positions for the set of matching 3D features, the determination including performing a simultaneous localization and mapping SLAM process based on first IMU data captured by the IMU, second IMU data captured by the IMU, and the filtered set of matching 3D features, the first IMU data being associated with the first key image frame and the second IMU data being associated with the second key image frame.
17. The system according to claim 16, wherein: The first IMU data includes a set of four degrees of freedom 4DOF parameters of the image sensor, and the second IMU data includes a set of 4DOF parameters of the image sensor.
18. The system according to claim 16, wherein, The movement of the image sensor is caused by a human individual holding the system performing a side step.
19. The system according to claim 18, wherein The first key image frame is identified by detecting the start impulse of the side step, and the first key image frame is a specific image frame in the set of captured image frames corresponding to the detected start impulse.
20. The system according to claim 18, wherein The second key image frame is identified by detecting the end impulse of the side step, and the second key image frame is a specific image frame in the set of captured image frames corresponding to the detected end impulse.
Citation Information
Patent Citations
Diminished and mediated reality effects from reconstruction
CN105164728A
Method and apparatus for device orientation tracking using a visual gyroscope
US20150103183A1