A visual-inertial odometry method fusing events and distances

By integrating an event camera and a distance sensor into a visual inertial odometry method, the problem of inaccurate positioning of micro UAVs in high dynamic scenarios is solved, achieving accurate UAV positioning, improving positioning accuracy and stability, and making it suitable for complex environments.

CN115479602BActive Publication Date: 2026-01-02BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211258632.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2026-01-02
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

In highly dynamic scenarios, micro drones face significant positioning challenges. Traditional visual inertial odometry (VIO) systems are inaccurate in high-speed motion or environments with drastic brightness changes, and motion drift occurs when the IMU cannot be excited, resulting in large positioning errors.

Method used

A visual-inertial odometry method integrating event cameras and distance sensors generates clear event images through IMU state prediction, distance observation, and back-end optimization. It constructs a nonlinear sliding window to optimize the solution of system state and improves positioning accuracy and stability by combining visual geometry, IMU pre-integration, and range coplanar constraints.

Benefits of technology

In highly dynamic scenarios, it achieves precise positioning of UAVs, reduces drift of traditional technologies, improves positioning accuracy and stability, is suitable for various environments, and performs particularly well in conditions without satellite positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115479602B_ABST
    Figure CN115479602B_ABST
Patent Text Reader

Abstract

The application discloses a kind of fusion event and distance visual inertial odometry method, belong to multi-sensor fusion navigation positioning field.Aiming at the deficiency of traditional visual inertial odometry calculation method under the scene of high-speed motion or brightness changes sharply, the application proposes a pose estimation method fusing event, IMU and distance based on event camera, uses IMU state prediction, distance observation and back-end estimation map point to motion-compensate the event stream output by event camera in front end, and synthesizes clear event image;Then back end jointly makes nonlinear sliding window optimization with visual geometric constraint, IMU pre-integration constraint and the coplanar constraint of distance observation, finally constructs cost function optimization to solve and estimate system state.The method fuses distance information, makes up the influence on positioning caused by IMU failure, simultaneously realizes accurate positioning under high-speed motion, reduces drift, and has good positioning accuracy, stability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of multi-sensor fusion navigation and positioning, and particularly relates to a visual-inertial odometer method fusing events and distances. BACKGROUND

[0002] The micro unmanned aerial vehicle has the characteristics of small size, light weight, strong operability, single-person carrying, good concealment, and convenient operation, and has great military and civil values. Unlike traditional aircraft, the positioning of the micro unmanned aerial vehicle in a high dynamic environment is a challenging research topic. With the increasing complexity of the scenarios that the unmanned aerial vehicle responds to, in indoor, building, jungle and other global satellite positioning and navigation denial environments, there is no external positioning source for small and light devices such as micro unmanned aerial vehicles, and it is also impossible to carry large sensors such as laser radars in terms of size and weight, and the positioning difficulty is greatly increased. In the face of the above situation, monocular visual-inertial odometer (VIO) of light and small sensors such as cameras and inertial measurement units (IMU) can play an important role.

[0003] The target will appear motion blur in a high-speed motion or a scene with a sharp change in brightness, and thus the image quality will be reduced, and it is difficult to achieve accurate positioning by using the traditional vision-based VIO. The event camera is a new type of neural stimulation vision sensor and has potential advantages in a high dynamic environment. Unlike the traditional frame camera that outputs images at a fixed frame rate, the event camera outputs asynchronous and high-frequency 'event stream', that is, an event is output when the brightness of each pixel changes by more than a certain threshold. Compared with the traditional frame camera, the event camera has the characteristics of low latency (microsecond level) and wide dynamic range (140 dB). These characteristics of the event camera have great application value in a high dynamic scene.

[0004] Meanwhile, the unmanned aerial vehicle may not be effectively excited by the IMU during high-speed motion, and the VIO may produce serious motion drift and large positioning error due to the inability to obtain effective observation scale information. The distance sensor can measure the distance within tens of meters with an error of centimeters, and its light weight and small size can well complement the visual (inertial) module without significantly increasing the load of the system. The distance sensor can provide the visual scale information required by the VIO when the IMU cannot be excited, and improve the positioning accuracy and stability of the VIO.

[0005] Based on the deficiencies in the prior art and the characteristics of the event camera and the distance sensor, the application provides a visual-inertial odometer fusing events and distances, which fuses the event camera and the distance sensor on the basis of the existing monocular VIO to provide more accurate positioning. SUMMARY

[0006] The present application is directed to the problem of difficult navigation and positioning of unmanned aerial vehicles in high dynamic scenes, and proposes a navigation and positioning method that fuses events, distances and inertial measurement units (IMUs) to improve the positioning accuracy and stability of unmanned aerial vehicles in fast-moving environments. The main content of the method is as follows: first, the event stream output by the event camera is motion compensated using IMU state prediction, distance observation and back-end estimated map points in the front end to synthesize clear event images; then, the visual geometric constraint, IMU pre-integration constraint and coplanar constraint of distance observation are combined together for nonlinear sliding window optimization, and finally a cost function is constructed to optimize and solve the estimated system state. The specific technical solutions are as follows:

[0007] A visual inertial odometer method that fuses events and distances, comprising the following steps:

[0008] S1: Build a visual inertial odometer system: define the world coordinate system as W, the camera coordinate system as C, the IMU coordinate system as B, and the distance sensor coordinate system as R, and determine the IMU dynamics model;

[0009] S2: Generate event motion compensation images: fix the number of events and generate event frame images, use IMU and distance sensor to obtain pose and image depth information, and obtain event motion compensation images through motion compensation;

[0010] S3: Feature extraction, detection and tracking: extract Harris corner points from event motion compensation images, track them through optical flow method, obtain feature matching between image frames, and remove false matches based on RANSAC algorithm to obtain accurate feature point inter-frame association;

[0011] S4: Construct nonlinear constraint conditions: fix the size of the sliding window, and construct the marginalization prior constraint, IMU pre-integration constraint, visual re-projection constraint and coplanar constraint of the distance sensor and the feature points.

[0012] S5: Construct a cost function and solve the pose: construct a cost function based on the constraint error term, and optimize and solve the pose.

[0013] Preferably, the IMU dynamics model in step S1 is:

[0014]

[0015]

[0016]

[0017] where g w is the gravity vector in the world coordinate system, t iis the time of the ith frame, Δt is the time interval between two frames, and is the translation, velocity and rotation matrix of the IMU coordinate system relative to the world coordinate system at the ith frame; is the quaternion of , denotes quaternion multiplication; a t and ω t are the measured values of acceleration and angular velocity, respectively; and are the biases of the gravity accelerometer and the gyroscope.

[0018] Preferably, the step S2 of generating the event motion compensation image is as follows:

[0019] S21: Let the event e = [u t p] output by the event camera, where u = [u x u y ] represents the pixel coordinates, t represents the time, and p represents the polarity;

[0020] S22: Divide the event stream into multiple windows in chronological order, each window containing an event image composed of the same number of events;

[0021] S23: Motion compensate the event stream using the IMU and the distance sensor. For an event stream within a period of time, select one time instant as the reference time t ref , and then project all events at other time instants within the period of time onto the image plane corresponding to the reference time t ref . The new coordinates after projection are:

[0022]

[0023] where K is the intrinsic matrix of the event camera, T k and T ref are the poses of the event camera relative to the reference system at t k and t ref , respectively, obtained by integrating the angular velocity and acceleration of the IMU, s k and s ref are the depth information in the corresponding camera coordinate system before and after the event projection, respectively, obtained by distance measurement and plane constraint;

[0024] S24: Based on the new coordinates p′ k after motion compensation, accumulate the events to generate a clear event motion compensation image.

[0025] Preferably, the step S3 specifically includes:

[0026] S31: Extract Harris corner points from the event motion compensation image obtained from S2, divide the event motion compensation image into MxN regions, maintain a limited number of feature points in each region, and evenly distribute the feature points in each region;

[0027] S32: For the current frame image, first perform forward optical flow from the previous frame to the current frame, and provide a prediction value of the coordinates on the current frame image for each feature point of the previous frame image;

[0028] S33: After obtaining the tracking value of each feature point in the current frame image, perform reverse optical flow from the current frame to the previous frame to obtain feature matching between image frames;

[0029] S34: Solve the basis matrix from the previous frame to the current frame based on the RANSAC algorithm, and remove the wrong matching.

[0030] Preferably, the step S4 adopts a Schur complement method to construct the marginalized prior constraint, specifically:

[0031] Suppose the residual is r, and the Jacobian of the residual with respect to the optimization variable is J. Solve the following equation:

[0032] Hδx=b

[0033] where H=J T J, b=J T r, δx is the increment of variable x in this iteration;

[0034] The Schur complement method is used to divide the variable x into a part δx1 that needs to be marginalized and a part δx2 that is not marginalized:

[0035]

[0036] Correspondingly, H and b are also divided into:

[0037]

[0038] The δx2 is marginalized by using Gaussian elimination method, and the constraint unrelated to δx1 is obtained as follows:

[0039]

[0040] Preferably, the IMU pre-integration constraint in step S4 is constructed as follows:

[0041] According to the IMU dynamics model in step S1:

[0042]

[0043]

[0044]

[0045] wherein is the pre-integration quantity, respectively

[0046]

[0047]

[0048]

[0049] The pre-integration provides position, velocity and attitude constraints between consecutive frames, and the pre-integration error is constructed for two adjacent frames i and i+1 as follows:

[0050]

[0051] wherein is the error function, χ is the optimization variable, is the variable observation value;

[0052] Preferably, the visual re-projection constraint in the step S4 is constructed as follows:

[0053] For each feature point, it is projected onto other key frames using its inverse depth and the pose of the key frame, and the difference between the corresponding observation coordinates of other key frames is calculated to calculate the re-projection error.

[0054] Preferably, the coplanar constraint in the step S4 is constructed as follows:

[0055] Assuming that the points in the field of view of the event camera are all in the same plane, the distance between the IMU and the horizontal plane is obtained by the distance observation, and the distance between the IMU and the horizontal plane is obtained by the feature point, and the coplanar constraint of the distance observation is constructed.

[0056] Preferably, the step S5 specifically comprises:

[0057] S51: Construct the cost function according to the nonlinear constraint condition in S4 as follows:

[0058]

[0059] wherein, r p represents the residual error, H p is the information matrix. is the IMU pre-integration error function, is the re-projection error function, is the coplanar constraint error function. represents the observation value of the optimization variable. ρ is the kernel function, represents the error covariance of the pre-integration between the i-th frame and the i+1-th frame in the IMU coordinate system, represents the reprojection error covariance between the l-th frame and the j-th frame in the camera coordinate system, represents the coplanarity constraint error covariance between the l-th frame and the j-th frame based on the distance measurement point r and the feature point k in the camera coordinate system;

[0060] S52: obtaining the pose estimation by solving through the Ceres optimizer.

[0061] Compared with the prior art, the visual inertial odometer fusing events and distances has the following advantages:

[0062] The present application fuses distance, vision and inertial information to realize multi-source common positioning, overcomes the deficiency of high dynamic scene positioning technology of the unmanned aerial vehicle, solves the problems of inaccurate positioning and difficult positioning of the unmanned aerial vehicle under high-speed motion, and reduces the drift of the traditional technology. The method can be used in various scenes, is less affected by the environment, has good positioning accuracy, stability and robustness, and plays an important role in the satellite positioning environment. BRIEF DESCRIPTION OF DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows, and the features and advantages of the present application will be more clearly understood by referring to the drawings. The drawings are schematic and should not be understood as any limitation on the present application. For those skilled in the art, other drawings can be obtained without creative labor on the basis of these drawings. Among them:

[0064] Figure 1 is a flowchart of the visual inertial odometer fusing events and distances of the present application;

[0065] Figure 2 is an event motion compensation schematic diagram of the present application;

[0066] Figure 3 is a multi-point coplanar constraint schematic diagram of the present application;

[0067] Figure 4 is a pose estimation schematic diagram of the present application; DETAILED DESCRIPTION

[0068] In order to more clearly understand the above-mentioned purposes, features and advantages of the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0069] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without the specific details set forth in this description.

[0070] A visual-inertial odometry method fusing events and distance, embodiments include the following steps:

[0071] S1: Build a visual-inertial odometry system: define the world coordinate system as W, the camera coordinate system as C, the IMU coordinate system as B, and the distance sensor coordinate system as R, and determine the IMU dynamics model.

[0072] The IMU dynamics model is:

[0073]

[0074]

[0075]

[0076] wherein the point P in the A coordinate system is defined as p A The conversion matrix of the point coordinate transformed from the A coordinate system to the B coordinate system is The corresponding rotation matrix is The translation matrix is g W is the gravity vector in the world coordinate system, t i is the time of the i-th frame, and Δt is the interval between two frames, and are the translation, velocity and rotation matrix of the i-th frame IMU coordinate system relative to the world coordinate system. is the corresponding quaternion, denotes quaternion multiplication, a t and ω t are the measured values of acceleration and angular velocity, and are the biases of the accelerometer and the gyroscope.

[0077] S2: Generate event motion compensation images: generate event frame images with a fixed number of events, and obtain clearer event motion compensation images through motion compensation by using the pose and image depth information obtained by the IMU and the distance sensor.

[0078] According to the timestamp order, the event stream is divided into windows containing the same number of events, and each window W i An event image is synthesized, wherein the intensity of each pixel on the picture is positively correlated with the number of events in the window at the pixel coordinate. The motion-compensated image is represented as:

[0079]

[0080]

[0081] where x is the original event position, x i is the event position after motion compensation, f(x) is the intensity of x pixel position in the event frame image.

[0082] Since each event corresponds to a different timestamp, when the relative motion is fast, direct accumulation will produce serious motion blur, which is not conducive to subsequent feature extraction and tracking. Therefore, the events are motion compensated and then accumulated into event images. Select a time window event stream at a certain time as the reference time t ref For events occurring at time t k , project them to t ref , and the new coordinates after projection are

[0083]

[0084] where p k is the coordinate before projection, K is the intrinsic matrix of the camera, T k and T ref are the poses of the camera relative to the reference system at t k and t ref , which are obtained by integrating the angular velocity and acceleration of the IMU, s k and s ref are the depth information of the event before and after projection in the corresponding camera coordinate system, which is obtained by distance measurement and plane constraint.

[0085] Finally, based on the new normalized coordinates p′ k after motion compensation, the events are accumulated to generate clearer motion compensated event images.

[0086] S3: Feature extraction, detection and tracking: extract Harris corner points, and then track through the optical flow method to obtain feature matching between image frames, and remove a small amount of false matches based on RANSAC to obtain more accurate feature point inter-frame association.

[0087] First, divide the motion compensated image into MxN regions to extract Harris corner points, then perform forward optical flow from the previous frame to the current frame for the current frame image, and provide a coordinate prediction value for each feature point of the previous frame in the current frame. For a triangulated feature point k, the i normalized coordinate of the previous frame is Use the IMU to predict the pose T i+1 of the current frame i+1

[0088]

[0089] For the feature points which are not triangulated, the average optical flow is calculated to get the predicted coordinates of the feature points in the current frame, thus the forward optical flow is completed.

[0090] After getting the tracking value of the feature points in the current frame, to ensure the quality of tracking, the backward optical flow from the current frame to the previous frame is performed again, and the coordinates of the feature points in the previous frame are calculated reversely. Only when the coordinate error of the two times is less than the set threshold, it is considered that the tracking is correct.

[0091] Finally, the fundamental matrix from the previous frame to the current frame is solved based on the RANSAC algorithm to remove a small amount of incorrect matches, and a more accurate inter-frame association is obtained.

[0092] S4: Construct nonlinear constraints: fix the size of the sliding window, and construct the marginalization prior constraint, IMU pre-integration constraint, visual re-projection constraint, and coplanar constraint of the distance sensor and the feature points.

[0093] S4-1: Use a sliding window to solve the back-end optimization, and the window size is N+1. The optimization variables are constructed as follows

[0094] χ=[x0,x1,...,x N ,ρ0,ρ1,...,ρ m ]

[0095]

[0096] where ρ k is the inverse depth of feature point k in the starting frame. is the external parameter between the IMU and the camera, is the external parameter between the camera and the distance sensor, which can be calibrated online as a variable during optimization.

[0097] According to the IMU dynamics model in step S1, we get

[0098]

[0099]

[0100]

[0101] where is the pre-integration quantity, respectively

[0102]

[0103]

[0104]

[0105] The pre-integration provides position, velocity and attitude constraints between consecutive frames, and the pre-integration error is constructed for two adjacent frames i and i+1 as:

[0106]

[0107] where is the error function, χ is the optimization variable, is the variable observation.

[0108] S4-2: For each feature point, project it onto other keyframes using its inverse depth and the poses of the keyframes, and calculate the re-projection error by differencing its corresponding observations in other keyframes. For a feature point k, the coordinates projected from the ith frame onto the jth frame are:

[0109]

[0110] where are the normalized coordinates in the ith and jth frames, λ k is the inverse depth of the feature point in the starting frame.

[0111] The re-projection error is denoted as

[0112]

[0113] S4-3: When a new keyframe arrives, in order to maintain the size of the sliding window unchanged, the oldest keyframe in the window needs to be marginalized out. In order to guarantee the observation or constraint information carried by the old keyframe not to be lost while controlling the scale, the past state and observation are converted into prior constraints of the state still in the window by using the Schur complement.

[0114] Suppose the residual is r, and the Jacobian of the residual with respect to the optimization variable is J. Taking the Gauss-Newton method as an example, finally the following equation will be solved

[0115] Hδx=b

[0116] where H=J T J, b=J T r, δx is the increment of the variable x in this iteration.

[0117] The variable x is divided into the part that needs to be marginalized and the part that is not marginalized by using the Schur complement:

[0118]

[0119] Correspondingly, the H matrix and b are also divided into:

[0120]

[0121] By using Gaussian elimination method, the constraint which is irrelevant to δx1 can be obtained as follows

[0122]

[0123] S4-4: Assuming that the points in the camera view are on the same plane, the distance between the IMU and the horizontal plane can be obtained by distance observation, and the distance between the IMU and the horizontal plane at this moment can also be obtained by using the feature points. By using this point, a coplanar constraint of distance observation is constructed.

[0124] The coordinates of the distance measurement point in the IMU coordinate system can be expressed as:

[0125]

[0126] where r j is the distance observation of the jth frame, and are the rotation and translation parameters of the camera and the IMU, and are the rotation and translation parameters of the camera and the distance sensor.

[0127] The distance from the jth frame IMU to the plane can be expressed as:

[0128]

[0129] where R j is the rotation matrix of the jth frame IMU relative to the world coordinate system, and n is the unit normal vector of the plane in the world coordinate system.

[0130] For the feature point k on the jth frame, the coordinates of the point in the world coordinate system can be expressed as

[0131]

[0132] where λ k is the inverse depth of the starting frame i, R i and t i are the rotation and translation of the i-th frame IMU relative to the world coordinate system.

[0133] Since the feature point is also on the plane, the distance from the jth frame IMU to the plane can also be expressed as the inner product of the line connecting the jth frame IMU position and the feature point and the normal of the plane:

[0134]

[0135] and d j should be equal, that is, the line connecting the feature point and the distance observation point belongs to the plane and is perpendicular to the normal.

[0136]

[0137] Since each frame of camera has corresponding distance sensor data, for each feature point, constraints can be established between it and all the observed frames, which is consistent with the form of visual reprojection error.

[0138] S4-5: The camera view field is mostly a plane occupying most of the area, but there may be some different planes, i.e. the feature points in the view field may not be in the same plane, and it is necessary to determine whether the feature points are in the same plane as the distance measurement points. Assuming that the feature points are on a plane, the depth of the feature point on the current frame j can be obtained based on the coplanar constraint

[0139]

[0140] wherein is the normalized coordinate of the feature point k in the jth frame.

[0141] The estimated depth of the feature point in the starting frame is Determining whether the feature point is in the same plane as the distance measurement point can be converted into whether the depth calculated based on this assumption is reasonable. The reprojection error under the two depths is calculated respectively, and by comparing the two reprojection errors, it is determined whether the feature point is in the same plane.

[0142] First, based on the estimated depth The reprojection error of the feature point from the ith frame to the jth frame is calculated, and the projected coordinates of the feature point in the jth frame camera coordinate system are obtained:

[0143]

[0144] wherein, represents the coordinates of the feature point k in the ith frame;

[0145] After normalization of the coordinates, the reprojection error is

[0146]

[0147] Further, based on the coplanar constraint depth The reprojection error is calculated

[0148]

[0149] The reprojection error is

[0150]

[0151] Compare the two reprojection errors, if |e2|≤|e1|, it means The depth of the more current pose constraint, i.e. coplanar constraint, is reasonable, the feature point belongs to the same plane with the distance measurement point, therefore the plane constraint is added to the optimization solution, otherwise not.

[0152] S5: Constructing cost function, solving pose: constraint error term constructing cost function, optimizing solving pose.

[0153] According to the nonlinear constraint condition in S4, the cost function is constructed as follows:

[0154]

[0155] Among them, from left to right is the edge prior constraint, IMU pre-integration constraint, visual re-projection, distance coplanar constraint, and ρ is the kernel function. Finally, the pose estimation can be obtained by solving the Ceres optimizer.

[0156] The above only describes the embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.​

Claims

1. A visual-inertial odometry method fusing events and distance, characterized in that, The method comprises the following steps: S1: building a visual-inertial odometry system: defining a world coordinate system as W, a camera coordinate system as C, an IMU coordinate system as B, and a distance sensor coordinate system as R, and determining an IMU dynamics model; S2: generating an event motion compensation image: fixing the number of events and generating an event frame image, obtaining a pose and image depth information by using an IMU and a distance sensor, and obtaining an event motion compensation image by motion compensation; S3: feature extraction, detection and tracking: extracting Harris corner points from the event motion compensation image, tracking by an optical flow method, obtaining feature matching between image frames, and removing false matching based on an RANSAC algorithm to obtain accurate feature point inter-frame association; S4: constructing a nonlinear constraint condition: fixing a sliding window size, constructing an edge prior constraint, an IMU pre-integration constraint, a visual re-projection constraint, and a co-planar constraint of a distance sensor and a feature point; S5: constructing a cost function and solving a pose: constructing a cost function by a constraint error term, and optimizing and solving a pose.

2. The fusion event and distance visual-inertial odometry method according to claim 1, wherein, The IMU dynamics model in the step S1 is as follows: where g w is the gravity vector in the world coordinate frame, t i is the time of the i-th frame, and Δt is the time interval between two frames, and are the translation, velocity and rotation matrix of the IMU coordinate frame relative to the world coordinate frame at the i-th frame; is the quaternion of , denotes the quaternion multiplication; a t and ω t are the measured values of the acceleration and angular velocity, respectively; and are the bias of the accelerometer and gyroscope.

3. The fusion event and distance visual-inertial odometry method of claim 1, wherein, The event motion compensation image in the step S2 is generated in the following manner: S21: let the event camera output an event e = [u t p], where u = [u x u y ] represents a pixel coordinate, t represents a time, and p represents a polarity; S22: dividing an event stream into multiple windows according to a time stamp sequence, each window containing an event image composed of the same number of events; S23: Compensate the motion of the event stream using IMU and distance sensor, select one time as the reference time t for the event stream in a period of time ref Then project all events at other times in this period of time to the reference time t ref The new coordinates after projection on the corresponding image plane are: where p k is the pre-projection coordinate, K is the intrinsic matrix of the event camera, T k and T ref are the poses of the event camera relative to the reference frame at t k and t ref , which are obtained by integrating the angular velocity and acceleration of the IMU, s k and s ref are the depth information of the event camera before and after projection in the corresponding camera coordinate system, which are obtained by distance measurement and plane constraint; S24: Based on the new coordinates p' after motion compensation k , the cumulative generation of events generates a clear event motion compensation image.

4. The fusion event and distance visual-inertial odometry method of claim 1, wherein, The step S3 specifically comprises: S31: extracting Harris corner points from the event motion compensation image obtained in S2, dividing the event motion compensation image into MxN regions, maintaining a limited number of feature points in each region, and uniformly distributing the feature points in each region; S32: for a current frame image, first performing a forward optical flow from a previous frame to the current frame, and simultaneously providing a predicted value of a coordinate on the current frame image for each feature point of the previous frame image; S33: after obtaining the tracking values of the feature points in the current frame image, performing a reverse optical flow from the current frame to the previous frame to obtain feature matching between the image frames; S34: solving a fundamental matrix from the previous frame to the current frame based on an RANSAC algorithm, and removing false matching.

5. The fusion event and distance visual-inertial odometry method of claim 1, wherein, The edge prior constraint in the step S4 is constructed in a Schur complement manner, and specifically: Assuming that a residual error is r, a Jacobian of the residual error with respect to an optimization variable is J, and the following equation is solved: Hδx=b where H = J T J, b = J T r, δx is the increment of variable x in this iteration; The variable x is divided into a part δx1 that does not need to be marginalized and a part δx2 that needs to be marginalized in a Schur complement manner: Correspondingly, H and b are also divided into: The δx2 is marginalized by a Gaussian elimination method to obtain a constraint that is irrelevant to δx1 as follows:

6. The fusion event and distance visual-inertial odometry method of claim 2, wherein, The IMU pre-integration constraint in the step S4 is constructed in the following manner: According to the IMU dynamics model in the step S1, the following is obtained: wherein are pre-integration quantities, respectively The pre-integration provides position, velocity and attitude constraints between consecutive frames, and a pre-integration error is constructed for two adjacent frames i and i+1 as follows: wherein is the error function, χ is the optimization variable, is the variable observation.

7. The fusion event and distance visual-inertial odometry method of claim 1, wherein, The visual re-projection constraint in the step S4 is constructed in the following manner: For each feature point, the inverse depth thereof and the pose of a key frame are used to project the feature point onto other key frames, and a difference between the projection and a corresponding observation coordinate of the other key frames is calculated to calculate a re-projection error.

8. The fusion event and distance visual-inertial odometry method of claim 1, wherein, The co-planar constraint in the step S4 is constructed in the following manner: Assuming that the points in the field of view of the event camera are in the same plane, the distance between the IMU and the horizontal plane is obtained by distance observation, and the distance between the IMU and the horizontal plane is obtained by using the feature points, and a coplanar constraint of distance observation is constructed.

9. The fusion event and distance visual-inertial odometry method of claim 6, wherein, The step S5 specifically includes: S51: Construct a cost function according to the nonlinear constraint condition in S4 as follows: where r p represents the residual, H p is the information matrix, is the IMU pre-integration error function, is the re-projection error function, is the coplanarity constraint error function, represents the observation of the optimization variable, and p is the kernel function, represents the error covariance of the pre-integration between the i-th frame and the i+1-th frame in the IMU coordinate system, represents the error covariance of the re-projection between the l-th frame and the j-th frame in the camera coordinate system, represents the error covariance of the coplanarity constraint between the l-th frame and the j-th frame in the camera coordinate system based on the distance measurement point r and the feature point k. S52: Solve the pose estimation by using a Ceres optimizer.

Citation Information

Patent Citations

  • A method for a high precision solution of pose parameters based on minimization of reconstruction errors

    CN109102567A

  • Robot multi-camera visual-inertial real-time positioning method and device

    CN109506642A