Method and system for implementing adaptive feature detection in a vSLAM system

By adaptively adjusting the detector threshold and pyramid level, the robustness of the vSLAM system under the influence of initialization, motion and tracking errors is solved, and the accuracy and stability of feature detection and tracking are improved.

CN115380308BActive Publication Date: 2025-07-25GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180018400.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-09
Filing Date
2021-02-08
Publication Date
2025-07-25
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

Existing augmented reality (vSLAM) systems are susceptible to initialization, motion, and tracking errors during feature detection and tracking, resulting in insufficient robustness.

Method used

Adaptive feature detection method is adopted to adjust the detector threshold and pyramid level, and dynamically adjust the feature detection parameters according to different conditions of the initialization state, motion level and tracking level to improve robustness.

Benefits of technology

Improves the robustness of the vSLAM system in the face of initialization, motion and tracking errors, and optimizes the accuracy and stability of the output pose.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115380308B_ABST
    Figure CN115380308B_ABST
Patent Text Reader

Abstract

A method includes receiving a first image, receiving a motion data set, determining a motion level, determining an initialization state, and determining a tracking level. Under a first condition, the method includes generating a first image pyramid, detecting a plurality of features in the first image pyramid using a first detector threshold, and generating a first set of detected key points from the plurality of features. Under a second condition, the method includes generating a second image pyramid, detecting a plurality of features in the second image pyramid using a second detector threshold that is less restrictive than the first detector threshold, and generating a second set of detected key points. Under a third condition, the method includes detecting a plurality of features in the first image according to the first detector threshold and generating a third set of detected key points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of augmented reality technology, and more particularly, to a method and system for implementing adaptive feature detection in a vSLAM system. Background Art

[0002] Augmented Reality (AR) superimposes virtual content on the user's real-world view. With the development of AR Software Development Kits (SDKs), the mobile industry has brought mobile device AR platforms into the mainstream. AR SDKs typically provide six degrees-of-freedom (6DoF) tracking capabilities. A user can use a camera in an electronic device (such as a smartphone or an AR system) to scan the environment, and the electronic device performs visual simultaneous localization and mapping (vSLAM) in real time. A vSLAM unit can be used in a mobile device to detect features of real-world objects and track these features as the mobile device moves in a three-dimensional environment to implement vSLAM.

[0003] Despite the progress in the AR field, there is still a need in the art to improve methods and systems related to AR. Summary of the Invention

[0004] This application generally relates to methods and systems related to augmented reality applications. More specifically, embodiments of this application provide a method and system for adaptive feature detection using variable pyramid levels and detector thresholds. This application is applicable to various applications involving vSLAM operations, including but not limited to computer vision-based online 3D modeling, AR visualization, face recognition, robotics, and autonomous vehicles.

[0005] As described in this application, embodiments of this application respond to operating conditions to adjust feature detection parameters during the vSLAM process. The robustness of the vSLAM process in response to one or more conditions (such as the movement of the vSLAM unit, tracking error, and / or initialization error) can be improved by adjusting the feature detection parameters. For example, vSLAM feature detection can implement an image pyramid and / or can adjust the detection threshold to improve the robustness of feature detection of the vSLAM unit.

[0006] By installing software, firmware, hardware, or a combination thereof on a system, a system of one or more computers can be configured to perform specific operations or actions, and the software, firmware, hardware, or combination thereof causes the system to perform actions during operation. One or more computer programs can be configured to perform specific operations or actions through included instructions that cause a data processing device to perform actions when executed. In one general aspect, an adaptive feature detection method in visual simultaneous localization and mapping (vSLAM) processing is provided. In these methods, a computer system receives a first image, receives a motion data set, determines a motion level, determines an initialization state, and determines a tracking level. The method also includes determining one of at least three conditions. Under a first condition, the method includes generating a first image pyramid, detecting a plurality of features in the first image pyramid using a first detector threshold, and generating a first set of detected key points from the plurality of features at least in part through key point fusion and selection. Under a second condition, the method includes generating a second image pyramid, detecting a plurality of features in the second image pyramid using a second detector threshold that is less restrictive than the first detector threshold, and generating a second set of detected key points at least in part through key point fusion and selection. Under a third condition, the method includes detecting a plurality of features in the first image according to the first detector threshold and generating a third set of detected key points.

[0007] Implementations of the above methods may include one or more of the following features. Optionally, the first condition is determining that the initialization state is true and the motion level is true or determining that the initialization state is false. Optionally, the second condition is determining that the initialization state is true, the motion level is false, and the tracking level is false. Optionally, the third condition is determining that the initialization state is true, the motion level is false, and the tracking level is true. The method may also include receiving a second image, performing feature tracking on the second image at least in part according to the first set of detected key points, the second set of detected key points, or the third set of detected key points, determining a tracking quality, and generating updated key points from the second image according to determining that the tracking quality is false.

[0008] Optionally, determining the initialization state includes receiving one or more initialization parameters from an initializer communicating with the computer system. Optionally, determining the initialization state further includes determining an initialization quality value at least in part based on the one or more initialization parameters, comparing the initialization quality value with a threshold criterion, and determining that the initialization state is true according to the initialization quality value meeting the threshold criterion. Optionally, according to the initialization quality value not meeting the threshold criterion, determining the initialization state further includes determining that the initialization state is false.

[0009] Optionally, determining a motion level includes receiving a motion data set from an inertial measurement unit communicating with the computer system. Optionally, determining the motion level further includes determining a displacement value by a motion monitor communicating with the computer system at least partially based on the motion data set, comparing the displacement value with a threshold criterion, and determining the motion level as true according to the displacement value satisfying the threshold criterion. Optionally, according to the displacement value not satisfying the threshold criterion, determining the motion level further includes determining the motion level as false.

[0010] Optionally, determining a tracking level includes receiving a set of key points, tracking the set of key points in a first image, selecting a set of inlier points from the set of key points tracked in the first image, determining an error value from the set of inlier points, comparing the error value with an error threshold, and determining the tracking level as true according to the error value satisfying the error threshold. Optionally, according to the error value not satisfying the error threshold, determining the tracking level further includes determining the tracking level as false.

[0011] Optionally, generating a first image pyramid includes generating N downscaled images from a first image, where the average pixel resolution of each subsequent image after the first image is lower than the image before it in the first image pyramid, and where N is a pyramid level value corresponding to a non-zero integer.

[0012] Optionally, determining a first detector threshold is at least partially based on a detector threshold used to initialize the vSLAM unit.

[0013] Optionally, receiving a first image from a camera communicating with the vSLAM unit.

[0014] In another general aspect, a computer system is provided that includes one or more processors and one or more memories storing computer-readable instructions. When the one or more processors execute the computer-readable instructions, the computer-readable instructions configure the computer system to receive a first image, receive a motion data set, determine a motion level, determine an initialization state, and determine a tracking level. The computer-readable instructions further configure the computer system to determine one of at least three conditions. Under a first condition, the computer system is further configured to generate a first image pyramid, detect a plurality of features in the first image pyramid using a first detector threshold, and generate a first set of detected key points from the plurality of features at least partially by key point fusion and selection. Under a second condition, the computer system is further configured to generate a second image pyramid, detect a plurality of features in the second image pyramid using a second detector threshold, where the second detector threshold is less restrictive than the first detector threshold, and generate a second set of detected key points by key point fusion and selection. Under a third condition, the computer system is further configured to detect a plurality of features in the first image according to the first detector threshold and generate a third set of detected key points.

[0015] The implementation of the above system may include one or more of the following features. Optionally, the first condition is to determine that the initialization state is true and the motion level is true, or to determine that the initialization state is false. Optionally, the second condition is to determine that the initialization state is true, the motion level is false, and the tracking level is false. Optionally, the third condition is to determine that the initialization state is true, the motion level is false, and the tracking level is true.

[0016] Optionally, the computer-readable instructions further configure the computer system to receive a second image, perform feature tracking on the second image at least in part based on a first set of detected key points, a second set of detected key points, or a third set of detected key points to determine tracking quality, and generate updated key points from the second image based on determining that the tracking quality is false.

[0017] In another general aspect, there is provided one or more non-transitory computer storage media storing instructions that, when executed on a computer system, cause the computer system to perform operations including receiving a first image, receiving a motion data set, determining a motion level, determining an initialization state, and determining a tracking level. The operations further include determining one of at least three conditions. Under the first condition, the operations further include generating a first image pyramid, detecting a plurality of features in the first image pyramid using a first detector threshold, and generating a first set of detected key points from the plurality of features at least in part through key point fusion and selection. Under the second condition, the operations further include generating a second image pyramid, detecting a plurality of features in the second image pyramid using a second detector threshold that is less restrictive than the first detector threshold, and generating a second set of detected key points at least in part through key point fusion and selection. Under the third condition, the operations further include detecting a plurality of features in the first image according to the first detector threshold and generating a third set of detected key points.

[0018] Compared with the prior art, many benefits can be achieved through the present application. For example, embodiments of the present application provide methods and systems for improving the robustness of vSLAM feature detection operations by adapting the operating parameters of the vSLAM unit to the operating conditions of the system communicating with the vSLAM unit. These and other embodiments of the present application and their many advantages and features are described in more detail below in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 An example of a computer system including an inertial measurement unit and an RGB optical sensor for feature detection and tracking applications according to an embodiment of the present application is shown.

[0020] Figure 2 A simplified schematic diagram of a vSLAM system according to an embodiment of the present application is shown.

[0021] Figure 3 A simplified schematic diagram of an adaptive feature detection technique according to an embodiment of the present application is shown.

[0022] Figure 4 A simplified schematic diagram of a technique for generating a set of detected key points according to an embodiment of the present application is shown.

[0023] Figure 5A A simplified schematic diagram of a technique for generating a set of detected key points according to an embodiment of the present application is shown.

[0024] Figure 5B A simplified schematic diagram of a technique for generating a set of detected key points according to an embodiment of the present application is shown.

[0025] Figure 5C A simplified schematic diagram of a technique for generating a set of detected key points according to an embodiment of the present application is shown.

[0026] Figure 6 A simplified flowchart of a method for performing adaptive feature detection according to an embodiment of the present application is shown.

[0027] Figure 7 An example computer system according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0028] In the following description, various embodiments will be described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will be apparent to those skilled in the art that the embodiments may be practiced without specific details. In addition, well-known features may be omitted or simplified in order not to obscure the described embodiments.

[0029] Among other things, embodiments of the present application relate to a vSLAM unit that includes a detection strategy processor, a motion monitor, and a tracking performance monitor. By introducing a detection strategy processor, a tracking performance monitor, and a motion monitor that communicate with the vSLAM unit, the robustness of the vSLAM unit's operation can be improved, particularly for feature detection and tracking. The detection strategy processor can implement a pyramid level detection technique to improve the robustness of feature detection in the images received by the vSLAM unit. The detection strategy processor can employ variable detection thresholds and variable pyramid level values during feature detection operations as a function of the initialization state, motion level, and / or tracking level. In this way, the detection strategy processor can reduce the impact of initialization errors and motion on the feature detection and tracking operations performed by the vSLAM unit.

[0030] In some embodiments, the detection policy processor may determine an initialization state, which is used to describe whether the vSLAM unit is initialized. The detection policy processor may also receive a motion level determined by a motion monitor at least in part based on motion data received from an inertial measurement unit (IMU). The detection policy processor may also receive a tracking level at least in part based on an error in tracking features determined by a tracking performance monitor. The detection policy processor may implement feature detection (also known as keypoint detection) using an image pyramid at least in part based on the initialization state, the motion level, and / or the tracking level, and apply a detector threshold to the feature detection operation, where the image pyramid includes multiple levels described by pyramid level values. The detection policy processor may modify the pyramid level values and / or the detector threshold as a function of the initialization state, the motion level, and / or the tracking level. The detection policy processor may generate a set of detected keypoints for the vSLAM unit for feature tracking operations on subsequent images received by the vSLAM unit.

[0031] In an example, a smartphone application may include an AR function for superimposing animated elements onto objects in the real world. For example, the animated elements may be symbols, floral patterns, cartoon animals, etc. For example, the smartphone application may detect and track a specific object such that a specific animated element appears on the phone screen only when the specific object is within the camera's field of view. To correctly place the animated elements in the display area in the appropriate size, perspective, and position so that they appear to interact with the real-world object, the smartphone application requires information about the object surfaces in the environment around the phone as well as the position and orientation (also known as pose) of the phone. In some cases, this information includes images captured by the camera and motion information of the phone in the environment. To determine the pose of the camera, the vSLAM unit may perform an initialization operation, thereby calculating an initial map of three-dimensional features into a multi-dimensional coordinate system and further providing an initial pose of the camera relative to the coordinate system.

[0032] The vSLAM unit can then initiate feature detection and tracking operations using the images received from the camera to track objects in the camera's field of view. When receiving an image, the vSLAM unit can perform feature tracking on the image using a set of detected key points determined during initialization or in the previous cycle of feature detection. The results of feature tracking can be used to determine the tracking level. At least partially based on the tracking level, the image can be used for feature detection to update the set of detected key points. In some cases, when the error in feature tracking exceeds an allowed threshold, the feature detection process can be carried out. Then, the results of feature tracking and feature detection can be used to optimize the output of the vSLAM unit, such as through bundle adjustment. This may include motion data from an inertial motion unit (IMU). In some cases, the vSLAM unit can adjust the feature detection procedure at least partially by updating the set of detected key points to correct for biases in the conditions at the time of vSLAM unit initialization.

[0033] In this example, the vSLAM unit can include other units to improve the robustness of feature detection and tracking operations. For example, the vSLAM unit can include a detection strategy processor to modify the process for updating the set of detected key points. The detection strategy processor can receive multiple inputs, including the motion level, initialization state, and / or tracking level. Each input can be determined by a unit in the vSLAM unit and can be used by the detection strategy processor to determine the pyramid level value and detector threshold applied to feature detection. A set of updated detected key points generated by the detection strategy processor can be applied to feature tracking as a technique to reduce the error in feature tracking and improve the output of the vSLAM unit.

[0034] Generally, vSLAM allows AR systems and other types of systems that use computer vision (CV) to detect features and objects in the real world to detect and track objects as they move relative to the objects. Since initialization, motion, and tracking errors can have an adverse impact on the accuracy and robustness of the vSLAM unit, the embodiments of this application provide systems for improving feature detection and tracking, reducing such errors, and improving the output pose generated by vSLAM operations.

[0035] Figure 1FIG. 0 shows an example of a computer system 110 according to an embodiment of the present application. The computer system 110 includes an inertial measurement unit (IMU) 112 and an RGB optical sensor 114 for feature detection and tracking applications. Feature detection and tracking can be implemented by the vSLAM unit 116 of the computer system 110. Generally, the RGB optical sensor 114 generates RGB images of the real-world environment. For example, the real-world environment includes real-world object 130. In some embodiments, the IMU 112 generates motion data regarding the motion of the computer system 110 in a three-dimensional environment, where the data includes, for example, the rotation and translation of the IMU 112 with respect to six degrees of freedom (e.g., translation and rotation according to three Cartesian axes). After the initialization of the AR session (where the initialization can include calibration and tracking), the vSLAM unit 116 provides an optimized output pose 120 of the real-world environment in the AR session, where the optimized output pose 120 describes the pose of the RGB optical sensor 114 at least partially with respect to the illustrated features 124 detected in the real-world object 130. The optimized output pose 120 describes the coordinate system and illustration for placing a two-dimensional AR object onto the real-world object representation 122 of the real-world object 130.

[0036] In one example, the computer system 110 represents a suitable user device that, in addition to including the IMU 112 and the RGB optical sensor 114, further includes one or more graphical processing units (GPUs), one or more general purpose processors (GPPs), and one or more memories that store computer-readable instructions executable by at least one processor to perform various functions of the embodiments of the present application. For example, the computer system 110 can be any one of a smartphone, a tablet computer, an AR headset, or a wearable AR device, etc.

[0037] The IMU 112 can have a known sampling rate (e.g., the time frequency at which data points are generated), and this value can be locally stored and / or accessible to the vSLAM unit 116. The RGB optical sensor 114 can be a color camera. The RGB optical sensor 114 and the IMU 112 can have different sampling rates. Generally, the sampling rate of the RGB optical sensor 114 is lower than the sampling rate of the IMU 112. For example, the RGB optical sensor 114 can have a sampling rate of 30 Hz, while the IMU 112 can have a sampling rate of 100 Hz.

[0038] In addition, the IMU 112 and the RGB optical sensor 114 installed in the computer system 110 can be separated by a transformation (e.g., distance offset, field of view angle difference, etc.). This transformation can be known, and its value can be stored locally and / or accessible to the vSLAM unit 116. During the movement of the computer system 110, the RGB optical sensor 114 and the IMU 112 can experience different motions relative to the center of face, the center of mass, or another rotation point of the computer system 110. In some cases, the transformation may cause errors or mismatches in the output pose optimized by vSLAM. For this reason, the computer system may include calibration data. In some cases, the calibration data can be set only based on the transformation. The calibration data can include data that is at least partially associated with the resolution of the RGB optical sensor 114.

[0039] The vSLAM unit 116 can be implemented as dedicated hardware and / or a combination of hardware and software (e.g., a general-purpose processor and computer-readable instructions stored in a memory and executable by the general-purpose processor). In addition to initializing the AR session, the computer system 110 can also execute adaptive feature detection techniques as part of the vSLAM process, such as Figures 2 to 7 shown.

[0040] In Figure 1 the example, a smartphone is used to display an AR session of the real-world environment. Specifically, the AR session includes providing an AR scene that includes a representation of a real-world table, and a vase (or some other real-world object) is placed on top of the real-world table. Virtual objects will be displayed in the AR scene. Specifically, the virtual objects will be displayed on top of the table. As part of detecting how the smartphone is positioned relative to the table and the vase in the real-world environment, the smartphone can use images from the RGB optical sensor 114 or other cameras to initialize the vSLAM unit. The vSLAM unit will define a coordinate system according to which it will detect features in the table and the vase. After initialization, the vSLAM unit will detect and track features as part of the entire AR system. When detecting and tracking features, the phone can monitor the accuracy of the tracking operation of the vSLAM unit 116, the motion level of the computer system 110, and / or the initialization quality, and can adjust the feature detection process used by the vSLAM unit 116 to improve the robustness of feature detection and tracking.

[0041] Figure 2 FIG. shows a simplified schematic diagram of a vSLAM system 200 according to an embodiment of the present application. In some cases, the vSLAM system 200 performs feature detection and tracking operations after initialization. In some cases, the vSLAM unit 116 receives data from an RGB optical sensor (e.g., Figure 1The RGB optical sensor 114) receives the first image 202. The first image 202 may form part of a set of images received by the vSLAM unit 116 such that the vSLAM unit 116 generates a set of detected key points from a previously received image of the set of images during initialization by the initializer unit 220 or during a previous feature detection operation. The vSLAM unit 116 may also receive IMU data 204.

[0042] The feature tracking unit 240 may track features detected in a previously received image of the set of images in the first image 202. At least partially based on whether the change in feature position is suitable for the model prediction of the coordinated feature offset, and at least partially based on the motion of the computer system (e.g., Figure 1 the computer system 110) of the computer system 110) relative to its environment, the output of the feature tracking unit 240 may include information about features described as inliers or outliers. For example, the initialization of the vSLAM unit 116 may determine the coordinate system and the output pose to predict that the feature will be translated by a given displacement in the first image 202 relative to a previously received image of the set of images received by the vSLAM unit 116. The displacement may be plotted in the coordinate system generated by the initializer 220 during the initialization of the vSLAM unit 116. At least partially based on the determination of the error between the model prediction and the measured displacement, the feature may be designated as an inlier or an outlier.

[0043] The tracking performance monitor 250 may analyze the feature tracking information generated by the feature tracking unit 240 to determine a tracking level, which may be a value along a range, e.g., a value between 0 and 1 along a range from 0 to 1. In some cases, the tracking performance monitor 250 may perform one or more operations using internal data from the feature tracking unit 240 to determine whether the tracked features in the first image 202 meet a predetermined criterion of the vSLAM system 200. For example, the tracking performance monitor 250 may integrate the error of the inliers tracked in the first image 202 and compare the integrated error with a threshold λ. In some cases, the tracking performance monitor 250 may determine the tracking level based on whether the error exceeds λ, such that the tracking level is false when the error exceeds λ and true when the error does not exceed λ.

[0044] The feature tracking level output by the tracking performance monitor 250 can be received as an input to the detection policy processor 260, which can also receive inputs from the initializer 220 and the motion monitor 230. In some cases, the initializer 220 can determine the initialization state at least in part based on measurements of initialization accuracy and / or quality. The initialization state can be represented as a true or false value received by the detection policy processor 260. In some cases, the initialization state can be determined by calculating the error of the currently tracked features in the image relative to the initial output pose and the coordinate system generated during initialization. For example, a computer system including the vSLAM unit 116 (e.g., Figure 1 computer system 110) can move from one environment to another (e.g., from an indoor environment to an outdoor environment), such that the initial coordinate system no longer accurately describes the environment around the computer system. In some cases, when the initialization accuracy exceeds a threshold, the vSLAM unit 116 can determine the initialization state as false.

[0045] In some cases, the motion monitor 230 can receive IMU data 204, including translational and rotational data in six degrees of freedom, described in more detail as Figure 1 shown. The motion monitor 230 can determine a motion level, which can be represented as a true or false value and can be based on accelerometer output and / or gyroscope output. In other embodiments, the motion level can be a value along a range, e.g., a value between 0 and 1 along a range from 0 to 1. In some cases, the motion level can be determined based on one or more operations that reduce the IMU data 204 to a single displacement value, which is then compared to a threshold, described in more detail as Figure 3 shown. The motion level received by the detection policy processor 260 can be used together with the initialization state and / or the tracking level to modify the operation of the feature detection unit 270, described in more detail as Figure 3 shown. The optimization unit 280 can receive the IMU data 204 as well as the output of the feature detection unit 270 to optimize the output pose of the optical sensor (e.g., Figure 1 RGB sensor 114). In some cases, the optimization may include a bundle adjustment (BA) operation. The BA operation can adjust the output pose 290 generated by the vSLAM unit 116 to minimize a cost function that quantifies the error when fitting a model to parameters, including but not limited to the camera pose and coordinates in a coordinate map associated with features detected in a three-dimensional environment (e.g., Figure 1 features 124).

[0046] Figure 3A simplified schematic diagram of a technique 300 for adaptive feature detection according to an embodiment of the present application is shown. In some cases, motion data 342 is generated by integrating 340 from IMU data 204, where the integration 340 converts acceleration data in six degrees of freedom into displacement values. For example, the translational and rotational accelerations measured by an accelerometer in an IMU (e.g., Figure 1 the IMU 112) can be integrated in three spatial dimensions to generate displacement values in units of length (e.g., meters). In some cases, the motion data 342 is received by a detection strategy processor 260. As Figure 2 shown, the detection strategy processor 260 can receive an image t-1 that forms part of a set of images 302. The detection strategy processor 260 can generate a set of detected key points 312 from the image t-1 at least in part based on the motion data, the initialization state, and / or the tracking level. The set of detected key points 312 can be applied to feature tracking 320 of features in a subsequent image t from the set of images 302. The multiple tracked feature points 322 generated during the feature tracking 320 are used to determine the tracking quality 330, which is described in more detail as Figure 2 shown.

[0047] In some cases, the tracking quality fails to meet a predetermined threshold, prompting the detection strategy processor 260 to repeat the detection operation and generate another set of detected key points 312. For example, if the tracking quality is poor, e.g., because the image contains few elements that can be tracked, the detection threshold can be reduced and / or the pyramid level can be increased as described herein. In some cases, the tracking quality 330 meets the predetermined threshold, after which the vSLAM unit can implement data alignment 350 to compensate for the motion of the computer system measured by the IMU, and / or can determine an updated initialization state 360. In some cases, the initialization state 360 is false, such that the vSLAM unit may not update the output pose 362. In some cases, the initialization state 360 is true, such that the vSLAM unit can implement optimization 370 of the output pose, which is described in more detail as Figure 2 shown, thereby generating an optimized pose 372.

[0048] In some cases, the adaptive feature detection technique 300 includes multiple iterations of the process such that each image in the set of images 302 is processed as image t-1 in the detection strategy processor 260 and subsequently as image t in the feature tracking 320. In some cases, the feature tracking quality meets a predetermined threshold such that the same set of detected key points 312 is used in the feature tracking 320 to process multiple consecutive images in the set of images 302, and for example, when the tracking quality remains true over multiple tracking cycles 330, the set of detected key points 312 is not updated. In some cases, the motion data 342 or the tracking quality 330 may require redefining the set of detected key points 312 so that the detection strategy processor 260 receives image t-1 in the set of images 302 and performs key point detection operations, which are described in more detail below Figure 4 as shown

[0049] Figure 4 FIG. shows a simplified schematic diagram of a technique 400 for generating a set of detected key points according to an embodiment of the present application. In some cases, as Figures 2 to 3 shown, the detection strategy processor 260 receives an image 202 from a set of images (e.g., Figure 2 a set of images 302), such as image 202 (e.g., Figure 2 image 202 and Figure 3 image t-1). In some cases, the detection strategy processor 260 also receives an initialization state 360, a motion level 422, and a tracking level 424, which are described in more detail as Figure 2 shown. In some cases, the motion level is determined based on comparing the motion data 342 with a threshold displacement value to generate a true or false value. In Figure 4 , I represents that the initialization state is true, T represents that the tracking level is true, and M represents that the motion level is true. Figure 4 The strategy in is named which of the determined parameters is true, while ignoring the false parameters. Operations inside the detection strategy processor may include, but are not limited to, implementing one or more key point detection strategies based on a combination of the initialization state 360, the motion level 422, and / or the tracking level 424 values. The detection strategy processor 260 may determine a pyramid level value and / or a detector threshold at least partially based on the combination of values, so as to generate the set of detected key points 312 from the image 202. In some cases, the set of detected key points 312 may include the output of a single detection strategy for each cycle. In some cases, the detection strategy processor 260 may implement a single strategy for each cycle at least partially based on the combination of the received values.

[0050] Generally speaking, it is described in more detail as Figures 5A to 5CAs shown, the pyramid level value describes a number of downscaling steps by which additional images are generated from image 202 for subsequent processing in keypoint detection according to a detector threshold. The detector threshold in keypoint detection refers to the criterion for recording the measured features in image 202 in the set of detected keypoints 312 or discarding them. In some cases, the detection strategy processor uses the threshold as the first detector threshold applied during initialization.

[0051] In some cases, when the initialization state 360 and the tracking level 424 are true while only the motion level 422 is false, it corresponds to satisfactory initialization, tracking, and motion. Based on this combination of values, the detection strategy processor 260 can use a pyramid level value of zero to implement the keypoint detection strategy IT 430 without the need to modify the detector threshold from the default value or the current value. The keypoint detection strategy IT 430 can be referred to as the default keypoint detection strategy, and this default keypoint detection strategy can be implemented when the vSLAM unit is initialized and the tracking error and motion level are small.

[0052] In some cases, only the initialization state 360 is true while the tracking level 424 and the motion level 422 are false. Based on this combination of values, the detection strategy processor 260 can use a pyramid level value N to implement the keypoint detection strategy I 432, where N is an integer greater than zero. The keypoint detection strategy I 432 can be referred to as the tracking error keypoint detection strategy, which can be implemented when the vSLAM unit is initialized and the motion level is small, but the vSLAM unit measures a tracking error outside a predetermined threshold. The pyramid level value can be determined at least in part based on the parameters of the hardware that makes up the computer system (e.g., Figure 1 computer system 110). In some cases, in the detection strategy I 432, the detection strategy processor 260 can reduce the detector threshold to a reduced threshold, which is less restrictive than the default threshold or the current threshold. In some cases, performing keypoint detection on N images at least in part according to the reduced detector threshold allows the detection strategy I 432 to detect more keypoints relative to the keypoint detection strategy IT 430, such that the set of detected keypoints 312 allows for improved tracking based on higher-quality detection results, as described in more detail in FIG. 5.

[0053] In some cases, the initialization state 360 and the motion level 422 are true, while the tracking level 424 is false. Based on this combination of values, the detection strategy processor 260 can implement the key point detection strategy IM 434a using a non-zero integer pyramid level value N and the default or current value of the detector threshold. The key point detection strategy IM 434a can be referred to as a high-motion key point detection strategy, which can be implemented when the vSLAM unit is initialized and feature tracking is very small but the vSLAM unit determines that the motion exceeds a predetermined threshold. In some cases, the motion level being true indicates that the displacement and thus the motion of the computer system measured by the IMU have exceeded the threshold (e.g., the computer system may be "moving fast" and / or may have experienced non-optimal acceleration during the period when the IMU measurements were recently generated). The detection strategy IM 434a can include a non-zero pyramid level value to improve the robustness of feature detection by selecting features that appear between pyramid levels, as described in more detail in FIG. 5, but keep the detector threshold unchanged, at least in part because the tracking level indicates that the tracking quality meets a predetermined threshold. As an example, the detection strategy processor can apply the strategy IM 434 in response to the motion level crossing from false to true due to the environment around the computer system moving simultaneously with the computer system (e.g., a smartphone held in the passenger cabin of a turning or accelerating vehicle), so that the vSLAM unit (e.g., Figures 1 to 2 the vSLAM unit 116) can track key points in the environment while recording displacements that exceed a predetermined threshold.

[0054] In some cases, the initialization state 360 may be false. Based on this combination of values, the detection strategy processor 260 can implement the key point detection zero strategy 434b using a non-zero integer pyramid level value N and the default or current value of the detector threshold. The key point detection zero strategy 434b can be referred to as an initialization key point detection strategy, which can be implemented when the detection strategy processor determines that the vSLAM unit is not initialized. The term zero (null) means that none of the parameters are true, in which case the most robust detection method can be applied to compensate for the lack of initialization. The key point detection zero strategy 434b can correspond to the same parameters as the strategy IM 434a, at least in part for correcting the initialization of the vSLAM unit, and no longer provides an accurate initial coordinate map or initial pose to produce accurate vSLAM operations, including but not limited to the optimized output pose. As described in more detail in Figure 3As shown, the initialization state can be an important parameter for several points in the vSLAM technology 300, such that when the initialization state 360 is false, the vSLAM unit can maintain the output pose unchanged. Therefore, in some cases, the detection strategy processor adopts the key point detection zero strategy 434b until the vSLAM unit completes re-initialization. In some cases, the detection strategy processor can use the detector threshold for initialization as the detector threshold used in the key point detection zero strategy 434b.

[0055] Figures 5A to 5C are all simplified schematic diagrams illustrating the technique for generating a set of detected key points (e.g., Figure 3 a set of detected key points 312) according to an embodiment of the present application. More details are described as Figure 4 shown, the detection strategy processor (e.g., Figure 4 the detection strategy processor 260) can implement a detection strategy based on a combination of the initialization state, the motion level, and / or the tracking level value. As Figure 4 shown, the four strategies can include different combinations of the detector threshold and the pyramid level value, which are described in more detail below.

[0056] Figure 5A shows a simplified schematic diagram of the technique for generating a set of detected key points according to an embodiment of the present application. In some cases, as Figure 4 shown, the detector strategy processor uses the detection strategy IT 430 according to the very small operation of the vSLAM unit (e.g., Figure 1 the vSLAM unit 116). Therefore, the detection strategy IT 430 can include a detector strategy processor that receives the original image 502 (e.g., Figure 2 the image 202 and Figure 3 the image t-1). In some cases, the detection strategy processor (e.g., Figure 2 the detection strategy processor 260) generates a set of detected key points 312 by detecting features in the original image 502 using the default or current threshold according to the original T detection 510. In some cases, the original T detection 510 does not include the pyramid level and includes it as an option for the detection strategy processor to illustrate the cumulative error in feature tracking during multiple feature tracking cycles performed by the vSLAM system.

[0057] Figure 5B is a simplified schematic diagram of the technique for generating a set of detected key points according to an embodiment of the present application. In some cases, as Figure 4As shown, the detector policy processor adopts detection policy I 432 based on the operation of the vSLAM unit under conditions where the movement is satisfactory but the tracking is not satisfactory. Thus, detection policy I 432 can generate one or more images through pyramid construction 540 according to the pyramid level value N, where N is a non-zero integer. The detector policy processor can process the original image 502 and a set of N downscaled images 542a - N, where a and N are integers greater than zero, and N is the pyramid level value. In some cases, N is equal to a. In detection policy I432, the pixel resolution of each downscaled image 542a to downscaled image 542n may be lower than the pixel resolution of the original image 502, and each pixel resolution may be progressively lower than the previous downscaled image 542a - N in the image pyramid. In some cases, the degree of downscaling can be at least partially based on a downscaling factor (e.g., binomial filter downscaling), or can be based on spatially weighted downscaling to emphasize one or more regions in the original image 502. Pyramid construction 540 can include, but is not limited to, Gaussian, Laplacian, and steerable pyramid construction techniques. For example, the Gaussian method can employ a context smoothing function based on a Gaussian filter. In contrast, the steerable pyramid method can use multi-scale, multi-directional bandpass filters to modify the scaling operation at each level of the image pyramid. In some cases, according to policy I432, the detector policy processor generates a set of detected key points 312 by detecting features in the original image 502 using a reduced T detection 544. Thus, the detector policy processor can use a reduced T detection 544a - n for each downscaled image 542a - n. In some cases, after key point detection 544 - 544n, detection policy I432 can include key point fusion and selection 546. Key point fusion and selection 546 can include selecting the set of detected key points 312 by combining the results of the reduced T detection 544 with the results of each downscaled image 542a - n, fusing key points that may be associated with the same feature in the images in the image pyramid, and selecting key points at least partially based on the scores of the fused key points. In some cases, the fusion of key points is at least partially based on the spatial positioning of the key points relative to each other in the coordinate system generated during initialization. In some cases, detection policy I 432 adopts other techniques, such as key point descriptor fusion, which includes comparing the context information of each key point in an effort to identify two or more key points with each other. After key point fusion and selection, detection policy I 432 can produce a set of detected key points 312 for use by technique 300.

[0058] Figure 5C FIG. shows a simplified schematic diagram of a technique for generating a set of detected key points according to an embodiment of the present application. In some cases, the detector policy processor implements detection policy IM or a zero policy according to detection operation 534 (e.g.,Figure 4 Strategy 434a - b). In some cases, both detection strategies employ similar methods, using N levels of pyramids and either default detector thresholds or current detector thresholds. As described in more detail for detection strategy I 432, the original image 502 can be reduced to m downscaled images 552a - m by a pyramid construction 550, where m is a non - zero integer equivalent to the pyramid level value associated with operation 534. In some cases, operation 534 includes detecting key points of the original image 502 according to the original T - detection 554, and detecting each downscaled image according to the original T - detections 554a - m, and then combining the detected key points through key - point fusion and selection 556, as described for detection strategy I 432 previously. In some cases, operation 534 generates the set of detected key points 312.

[0059] Figure 6 FIG. shows a simplified flowchart of an adaptive feature detection method using a vSLAM unit according to at least one embodiment of the present application. This process is described in conjunction with a computer system as an example of the computer system above. Some or all of the operations of the process can be implemented by specific hardware on the computer system, and / or can be implemented as computer - readable instructions stored on a non - transitory computer - readable medium of the computer system. When stored, the computer - readable instructions represent a programmable module including code executable by a processor of the computer system. Execution of these computer - readable instructions configures the computer system to perform the corresponding operations. Each programmable module combined with a processor represents a way of performing its respective operation. Although the above operations are shown in a specific order, it should be understood that a specific order may not be required, and one or more operations may be omitted, skipped, and / or reordered.

[0060] The method includes receiving a first image (602). As described in more detail Figure 1 shown, the first image can form part of a set of images received by a vSLAM unit (e.g., Figure 1 the vSLAM unit 116) from an optical sensor (e.g., Figure 1 the RGB sensor 114). Optionally, the first image can be received from a camera in communication with the vSLAM unit. In some cases, the camera may generate images at the original or native pixel resolution.

[0061] The method further includes receiving a motion data set (604). As described in more detail Figure 2 shown, a computer system (e.g., Figure 1 the computer system 110) can include an IMU. The IMU can measure motion in six degrees of freedom and provide motion data to the vSLAM unit. In some cases, the vSLAM unit processes the motion data to determine displacement values, equivalent to translational motion over a period of time.

[0062] The method further includes determining an initialization state (606). Optionally, determining the initialization state includes receiving one or more initialization parameters from an initializer communicating with the computer system, determining an initialization quality value at least in part based on the one or more initialization parameters, and comparing the displacement value with a threshold criterion. Based on the initialization quality value that meets the threshold criterion, the method may include determining that the initialization state is true. Optionally, the method may include determining that the initialization state is false based on an initialization quality value that does not meet the threshold criterion. More specifically described as Figure 1 As shown, the initialization may provide an initial output pose and an initial coordinate map for the vSLAM unit, and the vSLAM unit detects and tracks features in each image of the set of images therewith.

[0063] The method further includes determining a motion level (608). In one embodiment, determining the motion level includes receiving a motion data set from an inertial measurement unit communicating with the computer system, and determining a displacement value by a motion monitor communicating with the computer system at least in part based on the motion data set. In this embodiment, the method further includes comparing the displacement value with a threshold criterion, and determining that the motion level is true based on the displacement value meeting the threshold criterion. Optionally, the method may include determining that the motion level is false based on a displacement value that does not meet the threshold criterion. In some cases, more specifically described as Figure 3 As shown, the motion level may reflect the acceleration of the computer system based on displacement and / or the motion level.

[0064] The method further includes determining a tracking level (610). In a particular embodiment, determining the tracking level includes receiving a set of key points and tracking the set of key points in a first image. In this specific embodiment, the method further includes selecting a set of inliers from the set of key points tracked in the first image, determining an error value from the set of inliers, and comparing the error value with an error threshold. If the error value meets the error threshold, it is determined that the tracking level is true. If the error value does not meet the error threshold, it is determined that the tracking level is false. More specifically described as Figure 3 As shown, the tracking level may reflect the integrated error at least in part based on the inliers tracked in a set of detected key points.

[0065] The method further includes generating a first image pyramid based on determining that the initialization state is true and the motion level is true or determining that the initialization state is false (i.e., the first condition), detecting a plurality of features in the first image pyramid using a first detector threshold, and generating a set of detected key points (612) at least in part by key point fusion and selection. Optionally, generating the first image pyramid includes generating N downsampled images from the first image, where the average pixel resolution of each subsequent image after the first image is lower than the image in front of it in the image pyramid, and where N is a pyramid level value corresponding to a non-zero integer. Optionally, the first detector threshold is determined at least in part based on the detector threshold used to initialize the vSLAM unit.

[0066] The method further includes generating a second image pyramid based on determining that the initialization state is true, the motion level is false, and the tracking level is false (i.e., the second condition), detecting a plurality of features in the second image pyramid using a second detector threshold, where the second detector threshold is less restrictive than the first detector threshold, and generating a second set of detected key points (614) at least in part by key point fusion and selection.

[0067] The method further includes detecting a plurality of features in the first image according to the first detector threshold based on determining that the initialization state is true, the motion level is false, and the tracking level is true (i.e., the third condition); and generating a third set of detected key points (616).

[0068] In a particular embodiment, the method further includes receiving a second image, performing feature tracking on the second image at least in part based on the set of detected key points, determining a tracking quality at least in part based on a plurality of tracked feature points in the second image, and when determining that the tracking quality is false, generating updated key points from the second image; and replacing the set of detected key points with the updated key points.

[0069] It should be understood that Figure 6 The specific steps shown provide a particular method for detecting features in an image according to an embodiment of the present application. Other sequences of steps may also be performed according to alternative embodiments. For example, alternative embodiments of the present application may perform the above steps in a different order. Additionally, Figure 6 Each of the steps shown may include multiple sub-steps, which may be performed in various orders suitable for each step. Additionally, additional steps may be added or removed according to a particular application. Those of ordinary skill in the art will recognize many variations, modifications, and alternatives.

[0070] Figure 7 An example of the components of a computer system 700 according to certain embodiments is shown. Computer system 700 is an example of the above computer system. Although these components are shown as belonging to the same computer system 700, computer system 700 may also be distributed.

[0071] Computer system 700 includes at least a processor 702, a memory 704, a storage device 706, input / output peripherals (I / O) 708, a communication peripheral 710, and an interface bus 712. The interface bus 712 is configured to communicate, transfer, and transmit data, control, and commands among the various components of the computer system 700. The memory 704 and the storage device 706 include computer-readable storage media, such as RAM, ROM, electrically erasable programmable read-only memory (EEPROM), hard disk drives, CD-ROMs, optical storage devices, magnetic storage devices, electronic non-volatile computer storage, such as memory and other tangible storage media. Any such computer-readable storage media can be configured to store instructions or program code embodying various aspects of the present application. The memory 704 and the storage device 706 also include computer-readable signal media. A computer-readable signal media includes a propagated data signal that contains computer-readable program code. Such a propagated signal can be in any of a variety of forms, including but not limited to electromagnetic, optical, or any combination thereof. A computer-readable signal media includes any computer-readable media that is not a computer-readable storage media and that can communicate, propagate, or transmit a program for use in conjunction with the computer system 700.

[0072] In addition, the memory 704 includes an operating system, programs, and applications. The processor 702 is configured to execute stored instructions and includes, for example, a logic processing unit, a microprocessor, a digital signal processor, and other processors. The memory 704 and / or the processor 702 can be virtualized and can be hosted within another computer system, such as a cloud network or a data center. The input / output peripherals 708 include a user interface, such as a keyboard, a screen (e.g., a touch screen), a microphone, a speaker, other input / output devices, and computing components, such as a graphics processing unit, a serial port, a parallel port, a universal serial bus, and other input / output peripherals. The input / output peripherals 708 are connected to the processor 702 through any port coupled to the interface bus 712. The communication peripheral 710 is configured to facilitate communication between the computer system 700 and other computing devices through a communication network, and the communication peripheral 710 includes, for example, a network interface controller, a modem, wireless and wired interface cards, antennas, and other communication peripherals.

[0073] Although the subject matter has been described in detail with respect to specific embodiments, it should be understood that those skilled in the art can readily make modifications, variations, and equivalents to these embodiments after understanding the above. Therefore, it should be understood that the present application is presented for purposes of illustration rather than limitation, and does not exclude obvious modifications, changes, and / or additions by those of ordinary skill in the art to the subject matter. In fact, the methods and systems described herein can be embodied in various other forms. Additionally, various omissions, substitutions, and changes can be made to the methods and systems described herein without departing from the spirit of the present application. The appended claims and their equivalents are intended to cover forms or modifications that will fall within the scope and spirit of the present application.

[0074] Unless otherwise explicitly stated, it should be understood that throughout the discussion of this specification, the use of terms such as "processing", "computing", "calculating", "determining", and "identifying" refers to actions or processes of a computing device, such as one or more computers or similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.

[0075] One or more of the systems discussed in this application are not limited to any particular hardware architecture or configuration. A computing device can include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on a general-purpose microprocessor that access software stored thereon, which programs or configures the computer system from a general-purpose computing device into a special-purpose computing device that implements one or more embodiments of the subject matter. Any suitable programming, scripting, or other type of language or combination of languages can be used to implement the teachings contained in this application, which are included in the software that will be used to program or configure the computing device.

[0076] Embodiments of the methods disclosed in this application can be executed in the operation of such computing devices. The order of the blocks presented in the above examples can be changed, e.g., the blocks can be reordered, combined, and / or decomposed into sub-blocks. Certain blocks or processes can be executed in parallel.

[0077] The conditional language used in this application, e.g., among other things, "able to", "can", "may", "could", "for example", etc., unless otherwise explicitly stated or otherwise understood in the context in which it is used, is generally intended to convey that certain examples include, while other examples do not include, certain features, elements, and / or steps. Thus, such conditional language generally does not mean that one or more examples in any way require the features, elements, and / or steps, or that one or more examples necessarily include logic for deciding whether these features, elements, and / or steps are included in any particular example or will be performed in any particular example, with or without author input or prompting.

[0078] The terms "comprising", "having", etc. are synonyms and are used to indicate inclusion in an open-ended manner without excluding other elements, features, acts, operations, etc. In addition, the term "or" is used herein in its inclusive sense (rather than in its exclusive sense), so when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list. In this application, "adapted to" or "configured to" refers to open and inclusive language and does not exclude a device adapted to or configured to perform other tasks or steps. In addition, the use of "based on" means open and inclusive because a process, step, calculation, or other action "based on" one or more of the stated conditions or values may in practice be based on other conditions or values in addition to the stated conditions or values. Similarly, the use of "at least partially based on" means open and inclusive because a process, step, calculation, or other action "at least partially based on" one or more of the stated conditions or values may actually be based on other conditions or values in addition to the stated conditions or values. The headings, lists, and numbers included in this application are for ease of explanation only and are not restrictive.

[0079] The various features and processes described above can be used independently of each other or combined in various ways. All possible combinations and sub-combinations are within the scope of this application. In addition, in some embodiments, certain method or process blocks may be omitted. The methods and processes described in this application are also not limited to any particular sequence, and the associated blocks or states can be executed in other suitable sequences. For example, the described blocks or states can be executed in an order not specifically disclosed, or multiple blocks or states can be combined into a single block or state. Example blocks or states can be executed serially, in parallel, or in some other manner. Blocks or states can be added to or removed from the disclosed examples. Similarly, the configurations of the example systems and components described in this application may be different from those described. For example, elements can be added, removed, or rearranged compared to the disclosed examples.

Claims

1. A method implemented by a computer system, comprising: Receiving a first image via a Visual Simultaneous Localization and Mapping (vSLAM) unit, the first image being generated by an optical sensor communicating with the computer system; Receiving a motion data set, the motion data set being generated by an Inertial Measurement Unit (IMU) communicating with the vSLAM unit; The vSLAM unit determining a motion level using a motion monitor; The vSLAM unit determining an initialization state using an initializer; The vSLAM unit determining a tracking level using a tracking performance monitor; and Under a first condition, using a detection strategy processor of the vSLAM unit, the first condition being that the initialization state is determined to be true and the motion level is true, or the initialization state is determined to be false: Generating a first image pyramid based on the first image; the pyramid level value of the first image pyramid being a non - zero integer; Detecting a plurality of features in the first image pyramid using a first detector threshold; And Generating a first set of detected key points from the plurality of features at least in part by key point fusion and selection; Under a second condition, using a detection strategy processor of the vSLAM unit, the second condition being that the initialization state is determined to be true, the motion level is false, and the tracking level is false: Generating a second image pyramid based on the first image; the pyramid level value of the second image pyramid being a non - zero integer; Detecting a plurality of features in the second image pyramid using a second detector threshold, the second detector threshold being less restrictive than the first detector threshold; And Generating a second set of detected key points at least in part by key point fusion and selection; And Under a third condition, using a detection strategy processor of the vSLAM unit, the third condition being that the initialization state is determined to be true, the motion level is false, and the tracking level is true: Detecting the plurality of features in the first image according to the first detector threshold; And Generating a third set of detected key points.

2. The method according to claim 1, wherein The method further comprises: Receiving a second image; Performing feature tracking on the second image at least in part according to the first set of detected key points, the second set of detected key points, or the third set of detected key points; Determining tracking quality; and Generating updated key points from the second image according to determining that the tracking quality is false.

3. The method according to claim 1, characterized in that, The determining the initialization state includes: Receiving one or more initialization parameters from an initializer communicating with the computer system; Determining an initialization quality value at least in part based on the one or more initialization parameters; Comparing the initialization quality value with a threshold criterion; and Determining that the initialization state is true according to the initialization quality value meeting the threshold criterion; or Determining that the initialization state is false according to the initialization quality value not meeting the threshold criterion.

4. The method according to claim 1, wherein The determining the motion level includes: Receiving the motion data set from an Inertial Measurement Unit (IMU) communicating with the computer system; Determining a displacement value at least in part based on the motion data set by a motion monitor communicating with the computer system; Compare the displacement value with a threshold criterion; and Determine that the motion level is true based on the displacement value satisfying the threshold criterion; or Determine that the motion level is false based on the displacement value not satisfying the threshold criterion.

5. The method according to claim 1, wherein The determining the tracking level includes: Receiving a set of key points; Tracking the set of key points in the first image; Selecting a set of inlier points from the set of key points tracked in the first image; Determining an error value from the set of inlier points; Comparing the error value with an error threshold; and Determining that the tracking level is true based on the error value satisfying the error threshold; or Determining that the tracking level is false based on the error value not satisfying the error threshold.

6. The method according to claim 1, wherein Generating the first image pyramid based on the first image includes generating N downscaled images from the first image, with the average pixel resolution of each subsequent image after the first image being lower than the image before it in the first image pyramid, where N is a pyramid level value corresponding to a non-zero integer.

7. The method according to claim 1, wherein Determine the first detector threshold at least in part based on a detector threshold for initializing the vSLAM unit.

8. The method according to claim 1, wherein The first image is received from a camera communicating with the vSLAM unit.

9. A computer system, comprising: One or more processors; And One or more memories configured to store computer-readable instructions that, when executed by the one or more processors, configure the computer system to: Receive a first image through a vSLAM unit, the first image being generated by an optical sensor communicating with the computer system; Receive a motion data set generated by an inertial measurement unit communicating with the vSLAM unit; The vSLAM unit determines a motion level using a motion monitor; The vSLAM unit determines an initialization state using an initializer; The vSLAM unit determines a tracking level using a tracking performance monitor; and Under a first condition, using a detection strategy processor of the vSLAM unit, the first condition being determining that the initialization state is true and the motion level is true, or determining that the initialization state is false: Generate a first image pyramid based on the first image; the pyramid level value of the first image pyramid is a non-zero integer; Detect a plurality of features in the first image pyramid using a first detector threshold; And Generate a first set of detected key points from the plurality of features at least in part through key point fusion and selection; Under a second condition, using a detection strategy processor of the vSLAM unit, the second condition being determining that the initialization state is true, the motion level is false, and the tracking level is false: Generate a second image pyramid based on the first image; the pyramid level value of the second image pyramid is a non-zero integer; Detect a plurality of features in the second image pyramid using a second detector threshold, the second detector threshold being less restrictive than the first detector threshold; And Generate a second set of detected key points at least in part through key point fusion and selection; And Under a third condition, using a detection strategy processor of the vSLAM unit, the third condition being that the initialization state is determined to be true, the motion level is false, and the tracking level is true: Detect the plurality of features in the first image according to the first detector threshold; and Generate a third set of detected key points.

10. The computer system according to claim 9, wherein The computer-readable instructions further configure the computer system to perform: Receive a second image; Perform feature tracking on the second image at least in part according to the first set of detected key points, the second set of detected key points, or the third set of detected key points; Determine the tracking quality; and Generate updated key points from the second image according to determining that the tracking quality is false.

11. The computer system according to claim 9, wherein, The determining the initialization state includes: Receiving one or more initialization parameters from an initializer communicating with the computer system; Determining an initialization quality value at least in part based on the one or more initialization parameters; Comparing the initialization quality value with a threshold criterion; and Determining that the initialization state is true according to the initialization quality value meeting the threshold criterion; or Determining that the initialization state is false according to the initialization quality value not meeting the threshold criterion.

12. The computer system according to claim 9, wherein Determining the motion level includes: Receiving the motion data set from an inertial measurement unit communicating with the computer system; Determining a displacement value at least in part based on the motion data set by a motion monitor communicating with the computer system; Comparing the displacement value with a threshold criterion; and Determining that the motion level is true according to the displacement value meeting the threshold criterion; or Determining that the motion level is false according to the displacement value not meeting the threshold criterion.

13. The computer system according to claim 9, wherein The determining the tracking level includes: Receiving a set of key points; Tracking the set of key points in the first image; Selecting a set of inliers from the set of key points tracked in the first image; Determining an error value from the set of inliers; Comparing the error value with an error threshold; and Determining that the tracking level is true according to the error value meeting the error threshold; or Determining that the tracking level is false according to the error value not meeting the error threshold.

14. The computer system according to claim 9, wherein Generating the first image pyramid based on the first image includes generating N downsampled images from the first image, the average pixel resolution of each subsequent image after the first image being lower than that of the image before it in the first image pyramid, where N is a pyramid level value corresponding to a non-zero integer.

15. A non-transitory computer storage medium storing instructions that, when executed on a computer system, cause the computer system to perform the following operations: Receive a first image through a vSLAM unit, the first image being generated by an optical sensor communicating with the computer system; Receive a motion data set, the motion data set being generated by an inertial measurement unit communicating with the vSLAM unit; The vSLAM unit determines the motion level using a motion monitor; The vSLAM unit determines the initialization state using an initializer; The vSLAM unit determines the tracking level using a tracking performance monitor; and Under a first condition, use the detection strategy processor of the vSLAM unit: Generate a first image pyramid based on the first image; the pyramid level value of the first image pyramid is a non-zero integer; Detect a plurality of features in the first image pyramid using a first detector threshold; and Generate a first set of detected key points from the plurality of features at least in part by key point fusion and selection; Under a second condition, use the detection strategy processor of the vSLAM unit: Generate a second image pyramid based on the first image; the pyramid level value of the second image pyramid is a non-zero integer; Detect a plurality of features in the second image pyramid using a second detector threshold, the second detector threshold being less restrictive than the first detector threshold; and Generate a second set of detected key points at least in part by key point fusion and selection; and Under a third condition, use the detection strategy processor of the vSLAM unit: Detect the plurality of features in the first image according to the first detector threshold; and Generate a third set of detected key points, wherein: The first condition is to determine that the initialization state is true and the motion level is true, or to determine that the initialization state is false; The second condition is to determine that the initialization state is true, the motion level is false, and the tracking level is false; and The third condition is to determine that the initialization state is true, the motion level is false and the tracking level is true.

16. The non-transitory computer storage medium according to claim 15, wherein The determining of the initialization state includes: Receiving one or more initialization parameters from an initializer communicating with the computer system; Determining an initialization quality value at least in part based on the one or more initialization parameters; Comparing the initialization quality value with a threshold criterion; and Determining that the initialization state is true according to the initialization quality value meeting the threshold criterion; or Determining that the initialization state is false according to the initialization quality value not meeting the threshold criterion.

17. The non-transitory computer storage medium according to claim 15, wherein The determining of the motion level includes: Receiving the motion data set from an inertial measurement unit communicating with the computer system; Determining a displacement value at least in part based on the motion data set by a motion monitor communicating with the computer system; Comparing the displacement value with a threshold criterion; and Determining that the motion level is true according to the displacement value meeting the threshold criterion; or Determining that the motion level is false according to the displacement value not meeting the threshold criterion.

18. The non-transitory computer storage medium according to claim 15, wherein The determining of the tracking level includes: Receiving a set of key points; Tracking the set of key points in the first image; Selecting a set of inliers from the set of key points tracked in the first image; Determining an error value from the set of inliers; Comparing the error value with an error threshold; and Determining that the tracking level is true according to the error value meeting the error threshold; or Determining that the tracking level is false according to the error value not meeting the error threshold.

Citation Information

Patent Citations

  • Global positioning method based on depth information

    CN110132284A

  • Posture data processing method and device, terminal and computer readable storage medium

    CN110310326A