Binocular visual odometry method based on event contrast maximization
By employing a binocular visual odometry method based on maximizing event contrast, and combining data from event cameras and standard cameras, the stability problem of combining a monocular camera with an IMU in visual-inertial odometry is solved, achieving more efficient and accurate camera trajectory tracking.
Patent Information
- Application Number
- CN202211082743.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-09-06
AI Technical Summary
Visual-inertial odometry (VIO) systems combining monocular cameras and IMUs exhibit poor accuracy and robustness under constant acceleration conditions, and the IMUs are susceptible to noise interference, leading to instability in the VIO system.
A binocular visual odometry method based on event contrast maximization is adopted. By fusing data from event cameras and standard cameras and combining IMU data, feature point tracking and depth estimation are performed using the contrast maximization algorithm and Beta-Gaussian filter. The camera motion trajectory is estimated using the ICP algorithm.
It improves the computation speed of event streams and the accuracy of depth estimation, enables lower latency camera trajectory tracking, and enhances the robustness and accuracy of the system.
Smart Images

Figure CN115375767B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robot vision, and particularly relates to a binocular vision odometer method based on event contrast maximization. BACKGROUND
[0002] VO (Visual Odometry) is the most core part of the VSLAM (Visual SLAM) system, which is to solve the camera pose and motion trajectory through a series of continuous images captured by the camera in the movement process. However, there are still many problems in VO which only relies on the standard camera. For example: in the area with less texture, the number of feature points is insufficient, and VO cannot work; because for objects with high dynamic range or high-speed movement, the frame-based camera cannot obtain a clear image, and it is difficult for the visual front-end to extract feature points; in dynamic scenes, moving objects and pedestrians will cause feature point matching errors; therefore, in this case, it is difficult for the frame-based camera to realize feature tracking and camera pose estimation.
[0003] In addition, the frame-based camera captures redundant information in the static scene, which not only wastes storage resources, but also consumes a large amount of additional computing resources in the processing process. Bionic event cameras, such as dynamic vision sensors (DVS), overcome the above limitations of frame-based cameras. As a new type of bio-inspired visual sensor, the event camera has a completely different mode from the traditional camera, and the event camera only outputs the change depending on the brightness of the pixel. When the intensity of each pixel changes to a certain threshold, an event will be triggered, and the event carries information such as pixel coordinates, timestamp and polarity. The IMU (Inertial Measurement Unit) can well capture the acceleration and rotation information of the camera in three coordinate axes during high-speed movement, and assist the positioning of the camera; at the same time, the pose information of the camera is also helpful to reduce the influence of IMU drift; in addition, the IMU has the advantages of high integration, lightness, durability and low price. Therefore, by fusing the camera information and the IMU information to form the VIO (Visual-Inertial Odometry), the advantages and disadvantages of the two can be complementary, and the accuracy and reliability of the system are improved. At the same time, the VIO which fuses the camera and IMU information also greatly promotes the application of VISLAM (Visual-Inertial SLAM) in rescue robots, AR / VR rapid 3D reconstruction, automatic driving and other aspects.
[0004] Currently, the mainstream VIO framework is based on the fusion of monocular camera and IMU, but limited by the scale uncertainty of monocular camera, the algorithm will fail under constant acceleration motion conditions, resulting in poor accuracy and robustness of visual-inertial odometer. In addition, although the absolute scale of monocular camera pose can be obtained by aligning with IMU information, IMU is usually disturbed by noise, and the IMU used by ground robots is usually cheap, and the alignment result is poor and unstable. SUMMARY
[0005] The embodiment of the application aims to provide a binocular visual odometer method based on event contrast maximization, which solves the problem of unstable operation of monocular frame-based camera combined with IMU.
[0006] The technical scheme of the application is as follows:
[0007] In order to achieve the above-mentioned purpose, the application provides a binocular visual odometer method based on event contrast maximization, comprising the following steps:
[0008] The binocular visual odometer method based on event contrast maximization comprises:
[0009] S1: Preprocessing the pictures from the standard camera, and establishing a binocular visual odometer model based on event contrast maximization;
[0010] S2: Synchronizing the event stream and the pictures from the event camera on the time stamp, and selecting a space-time window according to the corresponding period of the event;
[0011] S3: Maximizing the contrast of the events in each space-time window, and calculating the optical flow between the corresponding events of the template;
[0012] S4: Using the IMU data and the calculated optical flow to correct and update the position of the template edge;
[0013] S5: Estimating the depth of the template edge using the Beta-Gaussian filter, and obtaining the position of the template edge in the 3D space;
[0014] S6: Estimating the motion trajectory of the event camera by the ICP pose solving algorithm.
[0015] Further, the above step S2 is specifically as follows:
[0016] Suppose that a feature point x is detected from the image frame at time t0, then the motion of the feature point can be described as:
[0017]
[0018]
[0019] In formula (1), To represent the velocity of feature point x at time s; select a set of events from the event stream, and select the time of feature point x as the spatiotemporal window corresponding to the initial time t0;
[0020] In formula (2), W is the set of events within the spatiotemporal window; e i This represents the i-th event within the spatiotemporal window; The value of t1 represents the time when the event occurred; n represents the number of events; since [t0, t1] is the first sub-time interval, the value of t1 can be calculated by setting the size of the time-space window;
[0021] To achieve the asynchronous nature of the feature point tracking method, the size of the sub-time interval is determined by the method during real-time execution, and the specific calculation process is as follows:
[0022] After obtaining the optical flow of all feature points in this iteration, the size of the next sub-time interval is calculated using the optical flow; define x. i Let i = {1, ..., m} represent the i-th feature point, and let m be the number of feature points. It is in the nth sub-time interval [t] n-1 , t n Interior feature point x i Optical flow; given sub-time interval [t] n-1 , t n Optical flow of all feature points within ] The next sub-time interval [t] can be calculated. n , t n+1 ];
[0023]
[0024] In formula (3), the unit of the number 3 is pixels; Represents the sub-time interval [t] n-1 , t n The average optical flow of all feature points within the area; t is calculated using formula (3). n+1 The time required for the feature points in the previous time interval to move an average of 3 pixels is used as the estimate for the current interval.
[0025] Furthermore, step S3 above is specifically as follows:
[0026] After determining the spatio-temporal window by corresponding feature points, a contrast maximization algorithm is used to match the event set W around the feature point x with the template point set; it is assumed that all template points have the same optical flow as the feature point x, the optical flow θ of all pixels in the region is the same, and the optical flow of the feature point is constant in the sub-time interval [t0, t1]; the optical flow of the feature point x in the time interval [t0, t1] is defined as v; for the event c in W i ; the position X' k of the event at time t0 is calculated using the warped event image (IWE), and the formula is as follows:
[0027]
[0028]
[0029] In formula (5), X' kj is the position of the kth event after being warped along the jth group of optical flows; N e represents the number of events; δ represents the Dirac function; P kj represents the probability that the kth event belongs to the jth group of optical flows; I j (x) represents the IWE corresponding to the optical flow;
[0030] The events are aligned by the contrast of the image, and the contrast of the image is defined by the sharpness ratio dispersion metric, such as variance
[0031] Var(I j )=∫ Ω (I j (x)-μ j ) 2 dx (6)
[0032] In formula (6), Ω is the image plane; μ j is the average of the warped event image;
[0033]
[0034] In formula (7), μ represents the step size; N l represents the number of clusters; θ is the optical flow in the corresponding spatio-temporal window.
[0035] Further, the above step S4 is specifically as follows:
[0036] After obtaining the optical flow of the feature point, the feature point and the template edge are updated; however, when the camera rotates, the template edge and the template point move at obviously different speeds, and the points far from the rotation center move faster; the process of updating the template edge is divided into two steps; the position of the template edge is updated using the optical flow, and then the position of the template edge is corrected using the IMU data;
[0037] Optical flow θ is used to update the position of feature point x and template edge x. j The corresponding position; assuming that the optical flow of feature point x is constant within the sub-time interval [t0, t1], the position of feature point is updated using optical flow, as shown in the following formula:
[0038] x(t1)=x(t0)+v.(t1-t0) (8)
[0039] In formula (8), x(t0) is the position of the initially extracted feature point; x(t1) is the position of the updated feature point.
[0040] To eliminate the influence of rotation on template edge updates, IMU data is introduced to correct the template edge position; the corrected position is closer to the true value of the template edge position; the relative position of the template point with respect to the feature point is calculated at time t0, and then the relative position is corrected using IMU data; for the template point, x i We define its relative position:
[0041]
[0042] Define symbol X j X and X represent the template point x in the camera coordinate system, respectively. j The 3D coordinates corresponding to the feature point x;
[0043]
[0044]
[0045] In formulas (10) and (11), R and t are the camera's rotation matrix and translation vector, and K represents the camera's intrinsic parameter matrix; Substitute the above formulas into (9) and normalize the 3D coordinates;
[0046]
[0047] x j The position at time t1 is determined by the position of the feature point and its relative position at that time. The result of adding them together:
[0048]
[0049] Finally, all template points can be updated using the formula.
[0050] Furthermore, step S5 above is detailed as follows:
[0051] After obtaining the corresponding position of each frame's edge map, the 3D position coordinates of the edge in space can be recovered by matching the pixel points on the edge between different edge maps. Using triangulation, the coordinates of the 3D point in space are recovered by utilizing the pixel positions of the 3D point observed from different viewpoints. The same 3D point can be observed in multiple frames, and the corresponding spatial 3D point coordinates can be calculated for any two frames. In the final estimation of the same 3D point, this strategy is based on a uniform Gaussian mixture distribution.
[0052] First, the depth of pixels in the edge image is recovered using triangulation. Assuming that a pixel p0 and a pixel p1 in the edge image are a pair of matching points, and the two pixels correspond to the same 3D point in space, formula (14) holds as follows:
[0053] y0K -1 p0=y1RK -1 p1+t (14)
[0054] Where R and t are the rotation matrix and translation vector of the event camera from edge map to edge map, respectively, and y0 and y1 represent the 3D points in the event camera coordinate system at time t. i and time t j The depth; Formula (14) is further simplified to:
[0055]
[0056] Then a depth value y can be calculated from corresponding points in the two images. i The distribution of y can be represented by a combination of Gaussian and uniform distributions.
[0057]
[0058] in the formula Based on the true value A Gaussian distribution centered on the center, It is its variance, π is the internal probability, and the π of a well-tracked feature is close to 1, y min and y max These are the characteristics of the lower and upper limits of a uniform distribution.
[0059] Furthermore, step S6 above is specifically as follows:
[0060] Given a sparse map of the scene, the current camera pose (R, t) is obtained using the ICP algorithm; the corresponding 3D point set P = {p1, p2, ..., pt} can be calculated through steps S4 and S5. n} and Q = {q1, q2, ..., q n} where n is the number of corresponding point pairs; then the optimal coordinate transformation, i.e. rotation matrix R and translation vector t, is calculated by the ICP algorithm, which can be described as follows:
[0061]
[0062] w in equation (17) i is the weight of each point; R and t are the rotation matrix and translation vector we need; in order to be robust to outliers, the bell-shaped Tukey weight function is used as follows:
[0063]
[0064] in equation (18) and b = 5; in addition, features that have been tracked well for a long time are given the highest weight, while features that have lost tracking are usually deleted due to large re-projection errors.
[0065] The technical scheme provided by the embodiments of the present application can include the following beneficial effects:
[0066] 1) Compared with the traditional event-by-event tracking method, the present application proposes a contrast maximization algorithm to solve the data association of events and images, which greatly improves the calculation speed of event stream.
[0067] 2) Since the contrast maximization algorithm highly depends on the depth of the scene, the present application proposes a robust Beta-Gaussian distribution depth filter to obtain more accurate line segment template depth than depth estimation using only triangulation.
[0068] 3) The evaluation experiment of the present application applied to public event camera dataset can achieve better performance and obtain lower delay camera trajectory compared with the visual odometry algorithm of ORB. BRIEF DESCRIPTION OF DRAWINGS
[0069] Figure 1 is the flowchart of the present application;
[0070] Figure 2(a) is the corner detection of the present application on the image frame;
[0071] Figure 2(b) is the edge detection of the present application on the image frame;
[0072] Figure 3 is the event accumulation graph of the first time window selected by the present application;
[0073] Figure 4 is the update schematic diagram of the line segment template of the present application;
[0074] Figure 5 is the variance convergence curve diagram of the present application;
[0075] Figure 6 is a schematic diagram of the triangulation depth estimation of the present application;
[0076] Figure 7 is a comparison diagram of the trajectory after pose estimation of the present application and different methods and the real trajectory;
[0077] Figure 8 is a comparison diagram of the absolute pose error of the present application and different methods. DETAILED DESCRIPTION
[0078] The present application will be described in detail below in conjunction with the drawings and specific embodiments.
[0079] As shown in the figure, the binocular visual odometry method based on event contrast maximization is as follows: Figure 1
[0080] Step 1, pre-process the pictures from the standard camera, and establish the event contrast maximization binocular visual odometry model; in the present application, we use DAVIS event camera, DAVIS is composed of event-based dynamic vision sensor (DVS) and frame-based active pixel sensor (APS) in the same pixel array, considering that the events of the edge area of the scene are triggered more frequently than the events of the low-texture area, we design the feature tracking combined with image and event to realize the event contrast maximization binocular visual odometry, as shown in the figure; first, the feature points and edge maps are extracted from the images from the standard camera through Harris corner detection and Canny edge detection, and then the edges in the specified area around each feature point are selected as the template edges of the feature points, as shown in Figs. 2(a) and (b). Figure 1
[0081] Step 2, synchronize the event stream and pictures from the event camera on the time stamp, and select the space-time window according to the corresponding period of the event, as shown in the figure; assuming that the feature point x is detected from the image frame at time t0, the motion of the feature point can be described as: Figure 3
[0082]
[0083]
[0084] In order to realize the asynchrony of the feature point tracking method, the size of the sub-time interval is determined by the method when it is running in real time, and the specific calculation process is as follows:
[0085] After obtaining the optical flow of all feature points in this iteration, the size of the next sub-time interval is calculated through the optical flow; define x i Let i = {1, ..., m} represent the i-th feature point, and let m be the number of feature points. It is in the nth sub-time interval [t] n-1 , t n Interior feature point x i Optical flow; given sub-time interval [t] n-1 , t n Optical flow of all feature points within ] The next sub-time interval [t] can be calculated. n , t n+1 ];
[0086]
[0087] Step 3: Maximize the contrast of events within each spatiotemporal window and calculate the optical flow between corresponding events in the template, such as... Figure 4 As shown; after determining the spatiotemporal window through corresponding feature points, the contrast maximization algorithm is used to match the event set W around feature point x with the template point set; for event e in W... i ;Calculate its position X' at time t0 using the image of the warped event (IWE). k The formula is as follows:
[0088]
[0089]
[0090] The event is aligned via image contrast, which is defined by a ratio of sharpness to dispersion, such as variance.
[0091] Var(I j )=∫ Ω (I j (x)-μ j ) 2 dx (6)
[0092]
[0093] θ represents the optical flow within the corresponding spatiotemporal window.
[0094] Step 4: Use IMU data and calculated optical flow to correct and update the position of the template edges, such as... Figure 5 As shown; after obtaining the optical flow of the feature points, the feature points and template edges are updated; the process of updating the template edges consists of two steps: updating the position of the template edges using optical flow, and then correcting the position of the template edges using IMU data;
[0095] Optical flow θ is used to update the position of feature point x and template edge x. jThe position of the feature point x is updated by using the optical flow, and the formula is as follows:
[0096] x(t1)=x(t0)+v.(t1-t0) (8)
[0097] The IMU data is introduced to correct the position of the template edge, and the corrected position is closer to the real value of the position of the template edge; the relative position of the template point with respect to the feature point is calculated at the time t0, and then the relative position is corrected by using the IMU data; for the template point, x j We define its relative position as follows:
[0098]
[0099] Define the symbol X j and X respectively represent the 3D coordinates of the template point x j and the feature point x corresponding to the camera coordinate system;
[0100]
[0101]
[0102] In the formulas (10) and (11), R and t are the rotation matrix and translation vector of the camera, and K represents the intrinsic matrix of the camera; the above formula is substituted into (9), and the 3D coordinates are normalized;
[0103]
[0104] x j The position at the time t1 is obtained by adding the position of the feature point and the relative position at that time
[0105]
[0106] Finally, the update of all template points can be completed by the formula.
[0107] Step 5, the depth of the template edge is estimated by using the filter of Beta-Gaussian, and the position of the template edge in the 3D space is obtained, as shown in Figure 6 After obtaining the corresponding position of each edge map, the 3D position coordinates of the edge in the space can be recovered by the matching relationship of the pixel points on the edge between different edge maps; the triangular measurement method is adopted to recover the coordinates of the 3D point in the space by using the pixel positions of the 3D point observed under different angles; in the final estimation of the same 3D point, the strategy is based on the uniform Gaussian mixture distribution;
[0108] First, the depth of pixels in the edge map is recovered using the triangulation method; assume that the pixel point p0 of the edge map and the pixel point p1 of the edge map are a pair of matching points, and the two pixel points correspond to the same spatial 3D point in space, and formula (14) is as follows:
[0109] y0K -1 p0=y1RK -1 p1+t (14)
[0110] Formula (14) is further simplified as:
[0111]
[0112] Then a depth value y can be calculated from the corresponding points in the two images i ; the distribution of y can be represented by a combination of Gaussian distribution and uniform distribution;
[0113]
[0114] Step 6, the motion trajectory of the event camera is estimated by the ICP pose solving algorithm, and the absolute pose error between the estimated camera trajectory and the 0RB estimated camera trajectory is drawn, as shown in Figure 7 and Figure 8 . Given the sparse map of the scene, the current camera pose (R, t) is obtained by the ICP algorithm; the corresponding 3D point sets P = {p1, p2, …, p n} and Q = {q1, q2, …, q n} can be calculated by step S4 and step S5, where the number of corresponding point pairs is n; then the optimal coordinate transformation, i.e. the rotation matrix R and the translation vector t, is calculated by the ICP algorithm, and the problem can be described as follows:
[0115]
[0116] In formula (17), w i represents the weight of each point; R and t are the rotation matrix and the translation vector we need; in order to be robust to outliers, the bell-shaped Tukey weight function is used, as follows:
[0117]
[0118] In formula (18), a and b = 5; in addition, features that have been tracked well for a long time are given the highest weight, and features that have lost tracking are usually deleted due to large re-projection errors.
[0119] The motion trajectory of the event camera is finally obtained by the ICP algorithm.
[0120] In the description of the application, reference can be made to terms such as "one embodiment", "some embodiments", "an example", "a specific example" or "some examples" etc. It is to be understood that such terms are merely used to describe a particular feature, structure, material or characteristic under discussion. Therefore, in no way these terms restrict the scope of the application. Further, when used in the description, the terms "a" or "an" are used in the sense that they mean "one or more" unless the context clearly dictates otherwise. In addition, the term "based on" is used to describe one or more feature, structure, material or characteristic that is / are used for an action or a step which is / are performed one or more times. Also, the term "based on" is used to describe one or more feature, structure, material or characteristic that is / are used for performing an action or a step which is / are performed one or more times.
[0121] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the application being indicated by the following claims.
Claims
1. A binocular vision odometry method based on event contrast maximization, characterized in that, Comprise: S1: pre-processing the pictures from standard camera, establishing binocular vision odometry model of event contrast maximization; S2: synchronization on time stamp of event stream and pictures from event camera, selecting space-time window according to corresponding period of event; S3: contrast maximization of event in each space-time window, and calculating optical flow between corresponding events of template; The step S3 is specifically as follows: After the spatiotemporal window is determined by the corresponding feature points, a contrast maximization algorithm is used to match the event set W around the feature point x with the template point set; all template points have the same optical flow as the feature point x, the optical flow θ of all pixels in the region is the same, and the optical flow of the feature point is constant in the sub-time interval [t0, t1]; the optical flow of the feature point x in the time interval [t0, t1] is defined as v; for the event e in W i ; the position X' of the event at time t0 is calculated using the warped event image IWE k , and the formula is as follows: In Equation (5), X kj is the position of the kth event after being warped by the jth group of optical flow; N e represents the number of events; δ represents the Dirac function; P kj represents the probability that the kth event belongs to the jth group of optical flow; I j (x) represents the IWE corresponding to the optical flow; Event is aligned through image contrast, and the contrast of image is defined by sharpness and color dispersion, variance In equation (6), Ω is the image plane; μ j is the average of the warped event image; In formula (7), μ represents a step size; N l denotes the number of clusters; θ is the optical flow in the corresponding spatio-temporal window; S4: using IMU data and calculated optical flow to correct and update the position of template edge; S5: using Beta-Gaussian filter to estimate the depth of template edge, and obtaining the position of template edge in 3D space; S6: estimating the motion trajectory of event camera through ICP pose solving algorithm.
2. The event-contrast-maximization-based binocular visual odometry method according to claim 1, characterized in that, The step S2 is specifically as follows: The motion of feature point x detected from image frame at time t0 can be described as: In formula (1), is the velocity of feature point x at time s; a set of events is selected from the event stream, and the time of feature point x is selected as the initial time t0 corresponding to the space-time window; In formula (2), W is a set of events within a space-time window; e i represents the i-th event in the space-time window; represents the time at which the event occurred; n represents the number of events; since [t0, t1] is the first sub-time interval, the value of t1 can be calculated by setting the size of the space-time window; In order to realize the asynchrony of feature point tracking method, the size of sub-time interval is determined by the method in real-time running, and the specific calculation process is as follows: After obtaining the optical flow of all feature points in this iteration, the size of the next sub-time interval is calculated using the optical flow; define x. i Let i = {1, ..., m} represent the i-th feature point, and let m be the number of feature points. It is in the nth sub-time interval [t] n-1 , t n Interior feature point x i Optical flow; given sub-time interval [t] n-1 ,t n Optical flow of all feature points within ] The next sub-time interval [t] can be calculated. n ,t n+1 ]; In equation (3), the unit of the number 3 is pixels; the average optical flow of all feature points in the sub-interval [t n-1 , t n ]; calculated by equation (3) t n+1 , the time required to move the feature points in the previous interval by an average of 3 pixels as the estimate of the current interval.
3. The event-contrast-maximization-based binocular visual odometry method according to claim 2, characterized in that, The step S4 is specifically as follows: After obtaining the optical flow of feature point, the feature point and template edge are updated; however, when the camera rotates, the template edge and template point move at obviously different speeds, and the points far away from the rotation center move faster; the process of updating the template edge is divided into two steps; the position of the template edge is updated using the optical flow, and then the position of the template edge is corrected using the IMU data; The optical flow θ is used to update the position of the feature point x and the corresponding position of the template edge x j In the sub-time interval [t0, t1], the optical flow of the feature point x is constant, so the position of the feature point is updated by using the optical flow, and the formula is as follows: x(t1)=x(t0)+v.(t1-t0) (8) In formula (8), x(t0) is the position of the feature point extracted initially; x(t1) is the position of the updated feature point, and v represents the motion speed; In order to eliminate the influence of rotation on the template edge update, the IMU data is introduced to correct the template edge position; the corrected position is closer to the true value of the template edge position; the relative position of the template point with respect to the feature point is calculated at t0, and then the relative position is corrected by using the IMU data; for the template point, x j We define its relative position: Definition of the symbol X j and X respectively represent the 3D coordinates corresponding to the template point x j and the feature point x in the camera coordinate system; In formulas (10) and (11), R and t are the rotation matrix and translation vector of the camera, and K represents the intrinsic matrix of the camera; the above formula is substituted into (9) and normalized to 3D coordinates; x j The position at time t1 is obtained by the position of the feature point and the relative position at that time The sum is: Finally, the update of all template points can be completed through formula.
4. The event-contrast-maximization-based binocular visual odometry method according to claim 3, characterized in that, The step S5 is specifically as follows: After obtaining the corresponding position of each frame edge map, the 3D position coordinates of the edge in space can be recovered through the matching relationship of the pixel points on the edge between different edge maps; the triangular measurement method is used to recover the coordinates of the 3D point in space by using the pixel positions of the 3D point observed under different viewing angles; the same 3D point can be observed in multiple frames, and the corresponding space 3D point coordinates can be calculated in any two frames; in the final estimation of the same 3D point, the strategy is based on uniform Gaussian mixture distribution; Firstly, the triangular measurement method is used to recover the depth of the pixels in the edge map; a pixel point p0 of the edge map and a pixel point p1 of the edge map are a pair of matching points, and the two pixel points correspond to the same space 3D point in space, and formula (14) is established as follows: y0K -1 p0 = y1RK -1 p1 + t (14) where R and t are the rotation matrix and translation vector of the event camera from edge map to edge map, y0and y1denote the depth of the 3D point in the event camera coordinate system at time t i and time t j respectively; equation (14) is further simplified as: A depth value y can then be calculated from the corresponding points in the two images i The distribution of y can be represented jointly by a Gaussian distribution and a uniform distribution in the formula Based on the true value A Gaussian distribution centered on the center, It is its variance, π is the internal probability, and the π of a well-tracked feature is close to 1, y min and y max These are the characteristics of the lower and upper limits of a uniform distribution.
5. The event contrast maximization based binocular visual odometry method according to claim 4, characterized in that, The step S6 is specifically as follows: Given the sparse map of the scene, the current camera pose (R, t) is obtained by the ICP algorithm; the corresponding 3D point set P = {p1, p2, …, pn} and Q = {q1, q2, …, qn} can be calculated by steps S4 and S5, where the number of corresponding point pairs is n; then the optimal coordinate transformation, i.e. the rotation matrix R and the translation vector t, is calculated by the ICP algorithm, and the problem can be described as follows: n} and Q = (q1, q2, …, q n} where the number of corresponding point pairs is n; then the optimal coordinate transformation, i.e. the rotation matrix R and the translation vector t, is calculated by the ICP algorithm, and the problem can be described as follows: w in equation (17) i denotes the weight of each point; R and t are the rotation matrix and translation vector we need; for robustness to outliers, a bell-shaped Tukey weight function is used, as follows: In equation (18) and b = 5; In addition, the features well tracked for a long time are given the highest weight, and the features losing tracking are usually deleted due to large re-projection error.
Citation Information
Patent Citations
Image defogging method with optimal contrast ratio and minimal information loss
CN104200445A
Monocular vision odometer positioning method and positioning system based on semi-direct method
CN108986037A