Pose estimation method and device, electronic equipment and storage medium
By combining standard cameras and event cameras, dynamically selecting the camera type suitable for the current scene, solving the problem of insufficient accuracy of visual information in different scenarios of visual inertial odometers, achieving more accurate pose estimation.
Patent Information
- Application Number
- CN202411966422.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
The existing visual inertial odometers provide insufficient accuracy of visual information in different scenarios, which affects the estimation accuracy of position.
Using a combination of standard cameras and event cameras, through the fusion of feature point tracking and inertial measurement data, dynamically selecting the camera type suitable for the current scene to improve the accuracy of visual information.
Accurate feature point tracking results can be obtained in different scenarios, which improves the accuracy of posture estimation of electronic devices.
Smart Images

Figure CN119935126A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual positioning technology, and in particular to a posture estimation method, device, electronic device and storage medium based on visual inertial odometer. Background Art
[0002] Visual-inertial odometry (VIO), sometimes also called visual-inertial system, is used to fuse camera and inertial measurement unit (IMU) data to achieve simultaneous localization and mapping (SLAM). Currently, the accuracy of visual information provided by visual-inertial odometry in different scenarios cannot be guaranteed, which in turn affects the accuracy of pose estimation. Summary of the invention
[0003] The embodiments of the present invention provide a method, device, electronic device and storage medium for posture estimation based on a visual inertial odometry, aiming to improve the accuracy of visual information provided by the visual inertial odometry in different scenarios, so as to improve the accuracy of posture estimation.
[0004] In a first aspect, an embodiment of the present invention provides a posture estimation method based on a visual inertial odometer, wherein the visual inertial odometer includes a standard camera, an event camera, and an inertial measurement unit, and the visual inertial odometer is loaded on an electronic device, and the method includes:
[0005] Acquire a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment, and perform a feature point tracking operation according to the first image and the second image to obtain a first tracking result;
[0006] Generate a target event image according to an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment;
[0007] In response to the first tracking result being inaccurate, performing a feature point tracking operation according to the target event image and the reference event image to obtain a second tracking result;
[0008] Determine the position and posture of the electronic device according to the second tracking result and the inertial measurement data output by the inertial measurement unit between the first moment and the second moment;
[0009] In response to the first tracking result being accurate, updating the reference event image to the target event image; and determining the position and posture of the electronic device according to the first tracking result and the inertial measurement data.
[0010] In a second aspect, an embodiment of the present invention further provides a posture estimation device based on a visual inertial odometer, wherein the visual inertial odometer includes a standard camera, an event camera, and an inertial measurement unit, and the visual inertial odometer is loaded on an electronic device, and the device includes:
[0011] An image acquisition module, used for acquiring a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment;
[0012] A tracking module, configured to perform a feature point tracking operation according to the first image and the second image to obtain a first tracking result;
[0013] An image generation module, configured to generate a target event image according to an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment;
[0014] The tracking module is further configured to, in response to the first tracking result being inaccurate, perform a feature point tracking operation based on the target event image and the reference event image to obtain a second tracking result;
[0015] A posture determination module, used to determine the posture of the electronic device according to the second tracking result and the inertial measurement data output by the inertial measurement unit between the first moment and the second moment;
[0016] an image updating module, configured to update the reference event image to the target event image in response to the first tracking result being accurate;
[0017] The posture determination module is further used to determine the posture of the electronic device according to the first tracking result and the inertial measurement data.
[0018] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a visual inertial odometer, a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the visual inertial odometer, the processor and the memory, wherein the visual inertial odometer comprises a standard camera, an event camera and an inertial measurement unit, wherein when the computer program is executed by the processor, the pose estimation method as described in the first aspect is implemented.
[0019] In a fourth aspect, an embodiment of the present invention further provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the pose estimation method as described in the first aspect.
[0020] The embodiment of the present invention provides a method, device, electronic device and storage medium for posture estimation based on visual inertial odometer. Since the visual inertial odometer in the embodiment of the present invention includes two different types of cameras, namely a standard camera and an event camera, and the standard camera can provide comprehensive and accurate visual information in a static or low-speed motion scene, and the event camera can provide comprehensive and accurate visual information in a high-speed motion or high-dynamic scene, after using two adjacent image frames provided by the standard camera to track feature points, if the feature point tracking result is inaccurate, it can be determined that the accuracy of the visual information provided by the standard camera does not meet the requirements of the current scene, and thus the visual information provided by the event camera whose accuracy meets the requirements of the current scene is used to track the feature points to ensure the accuracy of the feature point tracking result. If the feature point tracking result is accurate, it can be determined that the accuracy of the visual information provided by the standard camera meets the requirements of the current scene. In this way, whether in a static or low-speed motion scene or in a high-speed motion or high-dynamic scene, an accurate feature point tracking result can be obtained, and based on the accurate feature point tracking result and inertial measurement data, the posture of the electronic device can be accurately determined, thereby ensuring the accuracy of the posture of the electronic device in different scenes and improving the posture accuracy of the electronic device. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 It is a flowchart of a method for posture estimation based on visual inertial odometer provided by an embodiment of the present invention;
[0023] Figure 2 for Figure 1 Schematic diagram of the sub-step flow chart of the pose estimation method in ;
[0024] Figure 3 It is a flowchart of another method for posture estimation based on visual inertial odometer provided by an embodiment of the present invention;
[0025] Figure 4 It is a flowchart of another method for posture estimation based on visual inertial odometer provided by an embodiment of the present invention;
[0026] Figure 5 It is a structural schematic block diagram of a posture estimation device based on visual inertial odometer provided by an embodiment of the present invention;
[0027] Figure 6 It is a schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0030] It should be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0031] It should be noted that a standard camera is a camera that captures image frames at a fixed frame rate, and its output is a complete image frame. An event camera is an event-based camera, and its output is a sequence of events at a variable rate. Each event represents a change in light brightness. When the light intensity changes from the previous moment by more than a certain threshold, an event is generated, and the event's location, positive or negative polarity (light becomes stronger or weaker) and timestamp are recorded. An event camera has a microsecond delay and a dynamic range of up to 140dB, which is much higher than the 60dB of a traditional camera, enabling it to provide reliable visual information in high-speed motion or high dynamic range scenes. Standard cameras produce blur when shooting high-speed moving objects, but event cameras rarely have this problem because they only record events with brightness changes and are not affected by motion blur.
[0032] Visual-inertial odometry (VIO), sometimes also called visual-inertial system, is used to fuse camera and inertial measurement unit (IMU) data to achieve simultaneous localization and mapping (SLAM). Currently, visual-inertial odometry is mainly implemented using a single type of camera, which has limitations. The comprehensiveness and accuracy of visual information in different scenes cannot be guaranteed, affecting the accuracy of pose estimation.
[0033] To solve the above problems, an embodiment of the present invention provides a method, device, electronic device and storage medium for posture estimation based on visual inertial odometer. Since the visual inertial odometer includes two different types of cameras, a standard camera and an event camera, and the standard camera can provide comprehensive and accurate visual information in static or low-speed motion scenes, and the event camera can provide comprehensive and accurate visual information in high-speed motion or high-dynamic scenes, after using two adjacent image frames provided by the standard camera for feature point tracking, if the feature point tracking result is inaccurate, it can be determined that the accuracy of the visual information provided by the standard camera does not meet the requirements of the current scene, and thus the visual information provided by the event camera whose accuracy meets the requirements of the current scene is used for feature point tracking to ensure the accuracy of the feature point tracking result. If the feature point tracking result is accurate, it can be determined that the accuracy of the visual information provided by the standard camera meets the requirements of the current scene. In this way, whether in a static or low-speed motion scene or in a high-speed motion or high-dynamic scene, an accurate feature point tracking result can be obtained, and based on the accurate feature point tracking result and inertial measurement data, the posture of the electronic device can be accurately determined, thereby ensuring the accuracy of the posture of the electronic device in different scenes and improving the posture accuracy of the electronic device.
[0034] Among them, the pose estimation method based on visual inertial odometer can be applied to electronic devices, which may include mobile phones, tablet computers, laptops, desktop computers, personal digital assistants, head-mounted display devices, etc. Head-mounted display devices may include augmented reality (AR) glasses, AR helmets, mixed reality (MR) glasses, MR helmets, etc.
[0035] Some embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0036] See also Figure 1 , Figure 1It is a flowchart of a method for posture estimation based on visual inertial odometer provided in an embodiment of the present invention.
[0037] like Figure 1 As shown, the posture estimation method includes steps S101 to S106.
[0038] Step S101: obtaining a first image output by a standard camera at a first moment and a second image output at a second moment after the first moment, and performing a feature point tracking operation according to the first image and the second image to obtain a first tracking result.
[0039] In this embodiment, the second moment may be the current moment, and the first moment may be the previous moment. For example, if the second moment is t, then the first moment is t-1. The time difference between the first moment and the second moment may be determined according to the frame rate of a standard camera, or may be set by the user, and this is not specifically limited in the embodiment of the present invention.
[0040] In some embodiments, performing a feature point tracking operation based on the first image and the second image to obtain a first tracking result may include: extracting multiple first feature points in the first image, and determining an optical flow field between the first image and the second image based on a preset optical flow algorithm; and tracking the extracted multiple first feature points in the second image based on the optical flow field to obtain a first tracking result, wherein the first tracking result includes some or all of the first feature points tracked in the second image. The multiple first feature points in the first image are points with obvious gradient changes in the first image, such as corner points or edge points, and the preset optical flow algorithm may be a KLT (Kanade-Lucas-Tomasi) algorithm, a Lucas-Kanade algorithm, or a Horn-Schunck algorithm, etc.
[0041] In some embodiments, the method further includes: determining the number of first feature points tracked in the second image included in the first tracking result; dividing the number of first feature points tracked in the second image by the number of extracted first feature points to obtain a first tracking accuracy; if the first tracking accuracy is greater than a preset accuracy, determining that the first tracking result is accurate; if the first tracking accuracy value is less than or equal to the preset accuracy, determining that the first tracking result is inaccurate. The preset accuracy can be set based on actual conditions, and the embodiments of the present invention do not specifically limit this.
[0042] In some embodiments, performing a feature point tracking operation based on the first image and the second image to obtain a first tracking result may include: extracting multiple first feature points in the first image, and extracting multiple second feature points in the second image; matching the multiple first feature points with the multiple second feature points based on a preset feature point matching algorithm to obtain a first tracking result, wherein the first tracking result includes multiple first matching point pairs, and the first matching point pair includes a first feature point and a matched second feature point. The multiple first feature points in the first image are points with obvious gradient changes in the first image, such as corner points or edge points, and the multiple second feature points in the second image are points with obvious gradient changes in the second image, such as corner points or edge points, and the preset feature point matching algorithm may be a Brute-Force Matcher algorithm, a KNN (K-Nearest Neighbors) matching algorithm, or a FLANN (Fast Library for Approximate Nearest Neighbors) matching algorithm.
[0043] In some embodiments, the method also includes: determining the number of first matching point pairs included in the first tracking result; dividing the number of first matching point pairs by the larger of the number of extracted first feature points and the number of extracted second feature points to obtain a first tracking accuracy; if the first tracking accuracy is greater than a preset accuracy, determining that the first tracking result is accurate; if the first tracking accuracy value is less than or equal to the preset accuracy, determining that the first tracking result is inaccurate.
[0044] Step S102: Generate a target event image according to an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment.
[0045] In this embodiment, the event sequence includes multiple events, the inertial measurement data includes acceleration measurement data, angular velocity measurement data and magnetic measurement data, and the target event image is a motion compensated event image, which can reduce motion blur to improve the clarity of the event image and facilitate subsequent feature point extraction and tracking.
[0046] In some embodiments, Figure 2 As shown, step S102 includes: sub-steps S1021 to S1023.
[0047] Sub-step S1021, generating an event sub-sequence according to the event sequence.
[0048] In this embodiment, the event subsequence includes all or part of the events in the event sequence.
[0049] In some embodiments, according to the event sequence, generating an event subsequence may include: obtaining events from the event sequence in order of timestamps until N events are obtained, and the time window corresponding to the N events is greater than the first time window and less than the second time window, N is an integer greater than 1, and the second time window is related to the frame rate of the standard camera; sorting the N events in order of timestamps to obtain an event subsequence. For example, N and the first time window can be set based on actual conditions, and the embodiment of the present invention does not specifically limit this, and the second time window is related to the frame rate of the standard camera, including that the second time window is less than or equal to the inverse of the frame rate of the standard camera. The content and time corresponding to the event subsequence in this embodiment can be adjusted according to the dynamic changes in the scene, thereby ensuring the accuracy of the target event image subsequently generated based on the event subsequence.
[0050] In some embodiments, according to the event sequence, generating an event subsequence may include: acquiring events from the event sequence in order of timestamps until the time window corresponding to the acquired multiple events is equal to the second time window, and the second time window is related to the frame rate of the standard camera; sorting the acquired multiple events in order of timestamps to obtain an event subsequence. Among them, the second time window is related to the frame rate of the standard camera, including that the second time window is less than or equal to the inverse of the frame rate of the standard camera. The content and time corresponding to the event subsequence in this embodiment can be adjusted according to the dynamic changes in the scene, thereby ensuring the accuracy of the target event image subsequently generated based on the event subsequence.
[0051] In some embodiments, acquiring events from the event sequence in order of timestamps may include: acquiring events from the event sequence in order of timestamps from earliest to latest, or acquiring events from the event sequence in reverse order of timestamps from earliest to latest.
[0052] Sub-step S1022: determining a reference event from the event subsequence according to the second moment.
[0053] In this embodiment, an event closer to the second moment may be selected from the event subsequence as a reference event, thereby ensuring the correlation between the motion compensated target event image and the second image output by the standard camera at the second moment.
[0054] In some embodiments, according to the second moment, determining the reference event from the event subsequence may include: determining the absolute value of the difference between the second moment and the timestamp of each event in the event subsequence, obtaining the deviation value between the second moment and the moment of recording each event in the event subsequence; determining any event in the event subsequence whose deviation value is within a preset range as the reference event or determining the event corresponding to the smallest deviation value in the event subsequence as the reference event. The preset range may be set based on actual conditions, and the embodiments of the present invention do not specifically limit this.
[0055] For example, the event subsequence includes event 1, event 2, event 3, event 4, event 5, event 6, event 7, event 8, event 9, and event 10. If the deviation value between the second moment and the timestamp of recording event 8, event 9, and event 10 is within a preset range, event 8, event 9, or event 10 can be determined as a reference event. If the deviation value between the second moment and the timestamp of recording event 9 is the smallest, event 9 can be determined as a reference event.
[0056] Sub-step S1023: projecting all events in the event subsequence except the reference event onto an image plane corresponding to the reference event according to the inertial measurement data, to obtain a motion-compensated target event image.
[0057] In this embodiment, since the reference event is determined based on the second moment, the target event image obtained by projecting all events except the reference event onto the image plane corresponding to the reference event has a higher correlation with the second image output by the standard camera at the second moment, which facilitates subsequent feature extraction and tracking.
[0058] In some embodiments, based on inertial measurement data, all events in an event subsequence except a reference event are projected onto an image plane corresponding to the reference event to obtain a motion compensated target event image, which includes: determining a homogeneous transformation matrix from each event in the event subsequence except the reference event to the image plane corresponding to the reference event based on the inertial measurement data; determining a reprojected position of each event in the event subsequence except the reference event on the image plane corresponding to the reference event based on the homogeneous transformation matrix from each event in the event subsequence except the reference event to the image plane corresponding to the reference event and the position of each event in the event subsequence except the reference event; and drawing each event in the event subsequence except the reference event in the image plane corresponding to the reference event based on the reprojected position of each event in the event subsequence except the reference event on the image plane corresponding to the reference event to obtain a motion compensated target event image.
[0059] In some embodiments, determining the homogeneous transformation matrix from each event in an event subsequence except a reference event to an image plane corresponding to the reference event based on inertial measurement data includes: determining each event in the event subsequence except the reference event as a reprojection event; integrating the acceleration data and angular velocity data corresponding to the duration from the reprojection event to the reference event in the inertial measurement data to obtain the homogeneous transformation matrix from the reprojection event to the image plane corresponding to the reference event.
[0060] Step S103 : In response to the first tracking result being inaccurate, a feature point tracking operation is performed according to the target event image and the reference event image to obtain a second tracking result.
[0061] In this embodiment, the reference event image is generated earlier than the target event image, and the reference event image is a motion compensated event image.
[0062] In some embodiments, performing a feature point tracking operation based on the target event image and the reference event image to obtain a second tracking result may include: extracting multiple third feature points in the reference event image, determining an optical flow field between the target event image and the reference event image based on a preset optical flow algorithm; and tracking the extracted multiple third feature points in the target event image based on the optical flow field to obtain a second tracking result, wherein the second tracking result includes some or all of the third feature points tracked in the target event image. The preset optical flow algorithm may be a KLT (Kanade-Lucas-Tomasi) algorithm, a Lucas-Kanade algorithm, or a Horn-Schunck algorithm, etc.
[0063] In some embodiments, the method further includes: determining the number of third feature points tracked in the target event image included in the second tracking result; dividing the number of third feature points tracked in the target event image by the number of extracted third feature points to obtain a second tracking accuracy; if the second tracking accuracy is greater than a preset accuracy, determining that the second tracking result is accurate; if the second tracking accuracy value is less than or equal to the preset accuracy, determining that the second tracking result is inaccurate. The preset accuracy can be set based on actual conditions, and the embodiments of the present invention do not specifically limit this.
[0064] In some embodiments, according to the target event image and the reference event image, performing a feature point tracking operation to obtain a second tracking result may include: extracting a plurality of third feature points in the reference event image, and extracting a plurality of fourth feature points in the target event image; based on a preset feature point matching algorithm, matching the plurality of third feature points with the plurality of fourth features to obtain a second tracking result, the second tracking result including a plurality of second matching point pairs, the second matching point pair including a third feature point and a matched fourth feature point. The plurality of third feature points are points with obvious gradient changes in the reference event image, such as corner points or edge points, etc., the plurality of fourth feature points are points with obvious gradient changes in the target event image, such as corner points or edge points, etc., and the preset feature point matching algorithm may be a Brute-Force Matcher algorithm, a KNN (K-Nearest Neighbors) matching algorithm, or a FLANN (Fast Library for Approximate Nearest Neighbors) matching algorithm.
[0065] In some embodiments, the method also includes: determining the number of second matching point pairs included in the second tracking result; dividing the number of second matching point pairs by the larger of the number of extracted third feature points and the number of extracted fourth feature points to obtain a second tracking accuracy; if the second tracking accuracy is greater than a preset accuracy, determining that the second tracking result is accurate; if the second tracking accuracy value is less than or equal to the preset accuracy, determining that the first tracking result is inaccurate.
[0066] Step S104: Determine the position and posture of the electronic device according to the second tracking result and the inertial measurement data.
[0067] In this embodiment, the second tracking result and the inertial measurement data can be fused and processed based on a filter algorithm to obtain the position and posture of the electronic device, or the second tracking result and the inertial measurement data can be fused and processed based on a nonlinear optimization algorithm to obtain the position and posture of the electronic device.
[0068] In response to the first tracking result being accurate, step S105 is executed to determine the position and posture of the electronic device according to the first tracking result and the inertial measurement data.
[0069] In this embodiment, the first tracking result and the inertial measurement data can be fused and processed based on a filter algorithm to obtain the position and posture of the electronic device, or the first tracking result and the inertial measurement data can be fused and processed based on a nonlinear optimization algorithm to obtain the position and posture of the electronic device.
[0070] In some embodiments, the method further includes: acquiring a first image output by a standard camera at a first moment and a second image output at a second moment after the first moment, performing a feature point tracking operation based on the first image and the second image to obtain a first tracking result; generating a target event image based on an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment; performing a feature point tracking operation based on the target event image and the reference event image to obtain a second tracking result; determining the position and posture of the electronic device based on the tracking result with a higher tracking accuracy in the first tracking result and the second tracking result and the inertial measurement data. This embodiment uses the tracking result with a higher tracking accuracy and the inertial measurement data to determine the position and posture of the electronic device, thereby improving the estimation accuracy of the position and posture.
[0071] In response to the first tracking result being accurate, step S106 is executed to update the reference event image to the target event image.
[0072] This embodiment updates the baseline event image to the target event image when the first tracking result is accurate, so as to realize continuous updating of the baseline event image, thereby ensuring that the baseline event image is up to date, and facilitating the use of the latest baseline event image for subsequent feature extraction and tracking when the tracking result of the standard camera is inaccurate, so as to ensure the accuracy of pose estimation.
[0073] In some embodiments, updating the reference event image to the target event image may include: determining the number of events and the time window corresponding to the target event image; when the number of events is greater than or equal to a preset threshold, and the time window is greater than or equal to a second time window, updating the reference event image to the target event image (i.e., taking the target event image as a new reference event image, and replacing the old reference event image with the new reference event image). The preset threshold may be set based on actual conditions, and the embodiment of the present invention does not specifically limit this. This embodiment can ensure that the content of the reference event image is rich, which is convenient for subsequent feature extraction and tracking.
[0074] In some embodiments, Figure 3 As shown, after step S103, the method further includes: updating the reference event image to the target event image. In this embodiment, after performing the feature point tracking operation according to the target event image and the reference event image, the reference event image is updated to the target event image, so as to continuously update the reference event image, so as to ensure that the reference event image is the latest, so as to facilitate the subsequent feature extraction and tracking using the latest reference event image when the tracking result of the standard camera is inaccurate, so as to ensure the accuracy of the pose estimation.
[0075] For example, in the first time period, the reference event image is IT0 , according to the first image and the second image in the first time period, perform a feature point tracking operation to obtain a first tracking result G in the first time period T1 At the same time, according to the event sequence output by the event camera in the first time period and the inertial measurement data output by the inertial measurement unit in the first time period, the target event image I is determined T1 , if the first tracking result G in the first time period T1 If it is not accurate, then according to the benchmark event image I T0 and target event image I T1 , perform feature point tracking operation, obtain the second tracking result, and convert the benchmark event image I T0 Update to target event image I T1 Therefore, after the second time period, the benchmark event image is I T2 ; If the first tracking result G in the first time period T1 If accurate, the benchmark event image I T0 Update to target event image I T1 Therefore, in the second time period after the first time period, the reference event image is I T1 .
[0076] In the second time period, a feature point tracking operation is performed according to the first image and the second image in the second time period to obtain a first tracking result G in the second time period. T2 At the same time, according to the event sequence output by the event camera in the second time period and the inertial measurement data output by the inertial measurement unit in the second time period, the target event image I is determined T2 , if the first tracking result G in the second time period T2 If it is not accurate, then according to the benchmark event image I T1 and target event image I T2 , perform feature point tracking operation, obtain the second tracking result, and convert the benchmark event image I T1 Update to target event image I T2 Therefore, after the second time period, the benchmark event image is I T2 .
[0077] In some embodiments, Figure 4 As shown, before step S101, it also includes:
[0078] Step S107: Initialize the visual inertial odometer.
[0079] In this embodiment, initializing the visual inertial odometer may include static initialization or dynamic initialization of the visual inertial odometer to determine scale information, gravity vector, initial velocity, initial position, initial attitude, accelerometer and gyroscope deviations, etc. Wherein, after the visual inertial odometer is successfully initialized, step S101 is executed. In other embodiments, step S102 may be executed while the visual inertial odometer is initialized, or after the visual inertial odometer is successfully initialized, which is not specifically limited in the embodiment of the present invention.
[0080] See also Figure 5 , Figure 5 It is a schematic block diagram of the structure of a posture estimation device based on a visual inertial odometer provided in an embodiment of the present invention.
[0081] like Figure 5 As shown, the pose estimation device 110 based on visual inertial odometer includes:
[0082] An image acquisition module 111 is used to acquire a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment;
[0083] A tracking module 112, configured to perform a feature point tracking operation according to the first image and the second image to obtain a first tracking result;
[0084] An image generation module 113, configured to generate a target event image according to the event sequence output by the event camera between the first moment and the second moment and the inertial measurement data output by the inertial measurement unit between the first moment and the second moment;
[0085] The tracking module 112 is further configured to, in response to the first tracking result being inaccurate, perform a feature point tracking operation based on the target event image and the reference event image to obtain a second tracking result;
[0086] A posture determination module 114, configured to determine the posture of the electronic device according to the second tracking result and the inertial measurement data;
[0087] An image updating module 115 is configured to update the reference event image to the target event image in response to the first tracking result being accurate;
[0088] The posture determination module 114 is further configured to determine the posture of the electronic device according to the first tracking result and the inertial measurement data.
[0089] In some embodiments, the image generation module 113 includes:
[0090] An event generation submodule, used to generate an event subsequence according to the event sequence, wherein the event subsequence includes all or part of the events in the event sequence;
[0091] A reference event determination submodule, configured to determine a reference event from the event subsequence according to the second moment;
[0092] The reprojection submodule is used to project all events in the event subsequence except the reference event onto an image plane corresponding to the reference event according to the inertial measurement data, so as to obtain the target event image after motion compensation.
[0093] In some embodiments, the event generation submodule is further used to:
[0094] Acquire events from the event sequence in order of timestamps until N events are acquired, and a time window corresponding to the N events is greater than a first time window and less than a second time window, where N is an integer greater than 1, and the second time window is related to a frame rate of the standard camera;
[0095] The N events are sorted in order of timestamps to obtain the event subsequence.
[0096] In some embodiments, the event generation submodule is further used to:
[0097] Acquire events from the event sequence in order of timestamps until a time window corresponding to a plurality of acquired events is equal to a second time window, where the second time window is related to a frame rate of the standard camera;
[0098] The N events are sorted in order of timestamps to obtain the event subsequence.
[0099] In some embodiments, the reference event determination submodule is further used to:
[0100] determining a deviation value between the second moment and a timestamp of each event in the event subsequence;
[0101] Any event in the event subsequence whose deviation value is within a preset range is determined as the reference event.
[0102] In some embodiments, the image updating module is further used to update the reference event image to the target event image after obtaining the second tracking result.
[0103] In some embodiments, the visual inertial odometer-based pose estimation device 110 further includes:
[0104] An initialization module is used to initialize the visual inertial odometer.
[0105] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described posture estimation device based on visual inertial odometer can refer to the corresponding process in the aforementioned posture estimation method embodiment, and will not be repeated here.
[0106] See also Figure 6 , Figure 6 It is a schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0107] like Figure 6 As shown, the electronic device 100 includes a processor 101, a memory 102 and a visual inertial odometer 103, and the processor 101, the memory 102 and the visual inertial odometer 103 are connected via a bus 104, such as an I2C (Inter-integrated Circuit) bus.
[0108] Specifically, the processor 101 is used to provide computing and control capabilities to support the operation of the entire head-mounted display device. The processor 301 can be a central processing unit (CPU), and the processor 101 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0109] Specifically, the memory 102 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk, etc. The visual inertial odometer 103 includes a standard camera, an event camera, and an inertial measurement unit.
[0110] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a partial structure related to the embodiment of the present invention, and does not constitute a limitation on the head-mounted display device to which the embodiment of the present invention is applied. The specific head-mounted display device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0111] The processor 101 is used to run a computer program stored in the memory 102, and implement any one of the pose estimation methods provided by the embodiments of the present invention when executing the computer program.
[0112] In one embodiment, the processor 101 is used to run a computer program stored in a memory, and implement the following steps when executing the computer program:
[0113] Acquire a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment, and perform a feature point tracking operation according to the first image and the second image to obtain a first tracking result;
[0114] Generate a target event image according to an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment;
[0115] In response to the first tracking result being inaccurate, performing a feature point tracking operation according to the target event image and the reference event image to obtain a second tracking result;
[0116] Determining a position and posture of the electronic device according to the second tracking result and the inertial measurement data;
[0117] In response to the first tracking result being accurate, updating the reference event image to the target event image; and
[0118] The position and posture of the electronic device are determined according to the first tracking result and the inertial measurement data.
[0119] In some embodiments, when the processor 101 generates a target event image according to the event sequence output by the event camera between the first moment and the second moment and the inertial measurement data output by the inertial measurement unit between the first moment and the second moment, it is configured to implement:
[0120] Generate an event subsequence according to the event sequence, wherein the event subsequence includes all or part of the events in the event sequence;
[0121] determining a reference event from the event subsequence according to the second moment;
[0122] According to the inertial measurement data, all events in the event subsequence except the reference event are projected onto an image plane corresponding to the reference event to obtain the target event image after motion compensation.
[0123] In some embodiments, when generating an event subsequence according to the event sequence, the processor 101 is configured to implement:
[0124] Acquire events from the event sequence in order of timestamps until N events are acquired, and a time window corresponding to the N events is greater than a first time window and less than a second time window, where N is an integer greater than 1, and the second time window is related to a frame rate of the standard camera;
[0125] The N events are sorted in order of timestamps to obtain the event subsequence.
[0126] In some embodiments, when generating an event subsequence according to the event sequence, the processor 101 is configured to implement:
[0127] Acquire events from the event sequence in order of timestamps until a time window corresponding to a plurality of acquired events is equal to a second time window, where the second time window is related to a frame rate of the standard camera;
[0128] The N events are sorted in order of timestamps to obtain the event subsequence.
[0129] In some embodiments, when determining the reference event from the event subsequence according to the second moment, the processor 101 is configured to implement:
[0130] determining a deviation value between the second moment and a timestamp of each event in the event subsequence;
[0131] Any event in the event subsequence whose deviation value is within a preset range is determined as the reference event.
[0132] In some embodiments, after performing a feature point tracking operation according to the target event image and the reference event image to obtain a second tracking result, the processor 101 is further configured to implement:
[0133] The reference event image is updated to the target event image.
[0134] In some embodiments, before acquiring a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment, the processor 101 is further configured to implement:
[0135] Initialize the visual inertial odometry.
[0136] It should be noted that those skilled in the art can clearly understand that, for the convenience and conciseness of description, the specific working process of the electronic device described above can refer to the corresponding process in the aforementioned posture estimation method embodiment, and will not be repeated here.
[0137] An embodiment of the present invention also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement any pose estimation method provided in the description of the embodiment of the present invention.
[0138] The storage medium may be an internal storage unit of the electronic device described in the above embodiment, such as a hard disk or memory of the electronic device. The storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device.
[0139] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0140] It should be understood that the term "and / or" used in the present specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system including the element.
[0141] The serial numbers of the embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments. The above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A pose estimation method based on visual inertial odometry, characterized in that: The visual inertial odometer includes a standard camera, an event camera, and an inertial measurement unit. The visual inertial odometer is loaded on an electronic device. The method includes: Acquire a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment, and perform a feature point tracking operation according to the first image and the second image to obtain a first tracking result; Generate a target event image according to an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment; In response to the first tracking result being inaccurate, performing a feature point tracking operation according to the target event image and the reference event image to obtain a second tracking result; Determining a position and posture of the electronic device according to the second tracking result and the inertial measurement data; In response to the first tracking result being accurate, updating the reference event image to the target event image; and The position and posture of the electronic device are determined according to the first tracking result and the inertial measurement data.
2. The method for posture estimation according to claim 1, characterized in that: The generating a target event image according to the event sequence output by the event camera between the first moment and the second moment and the inertial measurement data output by the inertial measurement unit between the first moment and the second moment comprises: Generate an event subsequence according to the event sequence, wherein the event subsequence includes all or part of the events in the event sequence; determining a reference event from the event subsequence according to the second moment; According to the inertial measurement data, all events in the event subsequence except the reference event are projected onto an image plane corresponding to the reference event to obtain the target event image after motion compensation.
3. The method for posture estimation according to claim 2, characterized in that: Generating an event subsequence according to the event sequence includes: Acquire events from the event sequence in order of timestamps until N events are acquired, and a time window corresponding to the N events is greater than a first time window and less than a second time window, where N is an integer greater than 1, and the second time window is related to a frame rate of the standard camera; The N events are sorted in order of timestamps to obtain the event subsequence.
4. The method for posture estimation according to claim 2, characterized in that: Generating an event image according to the event sequence includes: Acquire events from the event sequence in order of timestamps until a time window corresponding to a plurality of acquired events is equal to a second time window, where the second time window is related to a frame rate of the standard camera; The multiple events are sorted in order of timestamps to obtain the event subsequence.
5. The method for posture estimation according to claim 2, characterized in that: The determining a reference event from the event subsequence according to the second moment includes: determining a deviation value between the second moment and a timestamp of each event in the event subsequence; Any event in the event subsequence whose deviation value is within a preset range is determined as the reference event.
6. The method for posture estimation according to any one of claims 1 to 5, characterized in that: After generating the target event image according to the event sequence output by the event camera between the first moment and the second moment and the inertial measurement data output by the inertial measurement unit between the first moment and the second moment, the method further includes: The reference event image is updated to the target event image.
7. The method for posture estimation according to any one of claims 1 to 5, characterized in that: Before acquiring a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment, the method further includes: Initialize the visual inertial odometry.
8. A posture estimation device based on visual inertial odometer, characterized in that: The visual inertial odometer includes a standard camera, an event camera and an inertial measurement unit. The visual inertial odometer is loaded on an electronic device. The device includes: An image acquisition module, used for acquiring a first image output by the standard camera at a first moment and a second image output at a second moment after the first moment; A tracking module, configured to perform a feature point tracking operation according to the first image and the second image to obtain a first tracking result; An image generation module, configured to generate a target event image according to an event sequence output by the event camera between the first moment and the second moment and inertial measurement data output by the inertial measurement unit between the first moment and the second moment; The tracking module is further configured to, in response to the first tracking result being inaccurate, perform a feature point tracking operation based on the target event image and the reference event image to obtain a second tracking result; A posture determination module, used to determine the posture of the electronic device according to the second tracking result and the inertial measurement data; an image updating module, configured to update the reference event image to the target event image in response to the first tracking result being accurate; The posture determination module is further used to determine the posture of the electronic device according to the first tracking result and the inertial measurement data.
9. An electronic device, characterized in that: The electronic device includes a visual inertial odometer, a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the visual inertial odometer, the processor and the memory. The visual inertial odometer includes a standard camera, an event camera and an inertial measurement unit. When the computer program is executed by the processor, the pose estimation method as described in any one of claims 1 to 7 is implemented.
10. A storage medium for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the pose estimation method described in any one of claims 1 to 7.