A High-Frequency Visual Odometry Method Based on Roller Shutter Camera

By employing line-by-line exposure and epipolar curve matching technology with a rolling camera array, the problems of low tracking accuracy and frequency of rolling cameras under high-speed motion are solved, realizing high-frequency, low-latency visual odometry, which is suitable for robot navigation, autonomous driving and augmented reality/virtual reality fields.

CN117274385BActive Publication Date: 2025-10-31ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311306232.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-10
Publication Date
2025-10-31
Estimated Expiration
2043-10-10

AI Technical Summary

Technical Problem

Existing visual odometry methods suffer from decreased tracking accuracy, low frequency, and high system latency when using rolling cameras, especially during high-speed movement or rapid rotation, leading to inaccurate positioning.

Method used

A high-frequency visual odometry method based on rolling shutter cameras is adopted. The distorted images from different perspectives are acquired through a group of rolling shutter cameras. By using epipolar constraints and depth filtering techniques, feature points are matched and depth is estimated row by row to construct a high-frequency tracking and optimization camera pose.

Benefits of technology

The tracking frequency of the visual odometry was increased, the system latency was reduced, the accuracy and robustness of positioning under rapid movement were ensured, the computational load was reduced, and real-time performance was achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274385B_ABST
    Figure CN117274385B_ABST
Patent Text Reader

Abstract

This invention discloses a high-frequency visual odometry method based on a rolling shutter camera. The method includes: acquiring rolling shutter distortion images from different perspectives using a rolling shutter camera group; obtaining tracking points on rows of the rolling shutter distortion images and finding the optimal matching point on the previous frame image; obtaining the matched pixels on each row of camera images and estimating the current row pose of the camera group using epipolar constraints; performing depth filtering on the feature points of each row on a keyframe, inputting a new frame of the rolling shutter image, calculating the epipolar curve of the feature points in the new frame, and searching for the optimal matching position on the current frame's epipolar curve; recovering the depth and uncertainty of a single measurement of the feature points using triangulation based on geometric relationships; fusing the previous state depth and the current measured depth of the camera feature points; adding the feature points to the map after the depth of the feature points on the reconstructed frame converges, and optimizing the camera pose. This invention features a low-cost camera, fast computation, and can be applied to scenarios with high real-time requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a high-frequency visual odometry method based on a rolling shutter camera. Background Technology

[0002] Simultaneous Localization and Mapping (SLAM) refers to a robot equipped with specific sensors building an environmental map while moving in an unknown environment, and simultaneously estimating its own motion. Visual SLAM (VSLAM), which uses cameras as sensors, plays an important role in augmented reality / virtual reality (AR / VR), 3D reconstruction, autonomous driving, path planning, intelligent robots, and drones due to its lower cost, higher image resolution, and ability to build denser maps. Most cameras are divided into global cameras and rolling cameras. Rolling cameras are widely used in consumer electronics due to their lower cost. Unlike global cameras that expose all pixels simultaneously, rolling cameras expose pixels row by row, with each row having a different pose. Currently, mainstream VSLAM algorithms and visual odometry (VO) methods are based on global exposure. However, the rolling effect is ignored when using rolling cameras, which leads to a decrease in system tracking accuracy, especially during rapid movement where tracking is easily lost.

[0003] In situations involving rapid robot movement, such as when a user wearing an AR / VR headset walks and rotates their head quickly, the tracking frequency required to accurately match the user's head movements is far higher than the camera's frame rate. Conventional visual odometry often incorporates an inertial measurement unit (IMU) to estimate the camera's position and rotation, but this introduces calibration issues between different sensors and increases costs. Devices using global cameras for tracking cannot achieve extremely fast tracking due to the mechanical limitations of their minimum exposure time. Roll-up cameras, on the other hand, have shorter exposure times, can achieve higher frame rates, and are less expensive. For high-frequency, high-quality tracking in situations involving high-speed camera rotation, maintaining low-latency targets and preventing them from being lost, roll-up cameras are the optimal and most cost-effective choice in static and well-lit environments. Therefore, researching high-frequency visual odometry based on roll-up cameras in static scenes to improve the accuracy of visual odometry during high-speed movement and reduce end-to-end system latency is of significant research and practical value. Summary of the Invention

[0004] The purpose of this invention is to address the problems of decreased or even lost camera tracking accuracy, low tracking frequency, and high system latency in existing technologies under high-speed motion conditions, and to provide a high-frequency visual odometry method based on a rolling shutter camera. This invention is applicable to achieving accurate, low-latency tracking, especially under high-speed rotation, significantly improving the system's tracking frequency and positioning accuracy.

[0005] The objective of this invention is achieved through the following technical solution: a high-frequency visual odometry method based on a rolling shutter camera, comprising the following steps:

[0006] S1. Obtain the roller shutter distortion images from different perspectives in the current environment through the roller shutter camera group, and simultaneously obtain the image row index of the roller shutter camera group exposed at the same time.

[0007] S2. Obtain the tracking point on the current roller shutter distortion image row, remove the distortion of the roller shutter distortion image, reconstruct the previous frame image row block through rotation compensation, and find the best matching point on the reconstructed image row.

[0008] S3. Based on step S2, obtain the best matching point on the row image of each rolling shutter camera exposed at the same time, and estimate the row pose of the initial rolling shutter camera group image by using epipolar constraints.

[0009] S4. Perform depth filtering on each row of feature points in the key frame. The new frame of the roller image is input. Calculate the parameter matrix of the epipolar equation of the current frame using the pose of the inter-frame row image. Construct the roller epipolar curve using the least squares optimization method. Search to obtain the best matching point position of the feature point in the current frame.

[0010] S5. Based on the epipolar geometric constraints satisfied by the feature points and their best matching points, use triangulation to recover the depth and uncertainty of a single measurement of the feature points;

[0011] S6. The optimal depth of the previous state of the camera feature points and the current measured depth are fused. After the feature point depths on the reconstructed row converge, the feature points are added to the map, and the reprojection error is minimized to obtain the optimized row pose of the rolling shutter camera group image.

[0012] Furthermore, the roller shutter camera group includes multiple roller shutter cameras, which are placed in different directions, and each roller shutter camera has an overlapping field of view with at least one other roller shutter camera. The roller shutter cameras rotate and translate in the x, y, and z directions, and the even-numbered roller shutter cameras rotate 90° on the image plane.

[0013] Furthermore, the step of obtaining the image row index of the rolling shutter camera group being exposed at the same time specifically includes: installing an LED light that flashes at a fixed frequency for each rolling shutter camera, with the light occupying the entire field of view of the rolling shutter camera; comparing the difference in the average grayscale value of the row images of the rolling shutter camera in two consecutive frames; if the difference exceeds a preset light intensity threshold, it indicates that the current row of the current frame is being exposed when the LED light flashes; and returning the image row index of the current row, which is the image row index of the rolling shutter camera group being exposed at the same time.

[0014] Further, step S2 specifically includes:

[0015] Calculate the gray-level gradient of each row in the distorted image of the roller blind, and take the first few points with the largest gradient changes as the tracking points of the row image;

[0016] In the camera coordinate system, the first two parameters k1 and k2 of the Brownian radial distortion model are used to model the radial distortion of the row image, and the distortion point p on the normalized plane is then used. c distorted =[x d ,y d ,1] T Perform distortion correction to obtain the original point p. c_undistorted =[x u ,y u ,1] T ;

[0017] Each new image row is passed in. t Distortion correction is performed on the tracking points on the row image. The same row of the roller image has the same y-value. d But different x d Different y in corresponding space u The position, that is, the line of the distorted image of the roller blind, after being deformed and restored to the original image, is a curve;

[0018] Extract the row image from the previous frame. i Its rotation is R i First, the camera image is rotated and compensated using the homography matrix, then the distortion is removed and the row image is reconstructed. For the input image Row... t Specify reference row j , with its rotation R j To replace the rotation of the current row image, when Row t with Row j When the rotation between rows exceeds the set rotation threshold, the current row will be updated to a new reference row;

[0019] Given the camera intrinsic parameter matrix K, determine the row image Row. j Row image iThe homography relationship between them is expressed as H j,i =KR j R i T K -1 Based on this expression, homography transformation is performed on the pixel rows of the previous frame image. Block matching is then performed in the previous frame image using the tracking points of the current row image: image blocks A of length L are taken on both sides of the tracking point, and image blocks B of the same length are also taken on both sides of the N points with the largest gray-level gradients in the reconstructed rows of the previous frame image. i Let i = 1, 2, ..., N, and calculate the similarity between two image patches. Select the two image patches with the highest similarity as the matching image patches to obtain the best matching point.

[0020] Furthermore, the expression for the Brownian radial distortion model is:

[0021] x d =x u (1+k1r 2 +k2r 4 )

[0022] y d =y u (1+k1r 2 +k2r 4 )

[0023]

[0024] Where, x d Represents the coordinates of the distortion point in the x-direction, y-direction... d The x-coordinate represents the coordinate of the distorted point in the y-direction. u Represents the coordinates of the original point in the x-direction, y-direction... u This represents the coordinate of the original point in the y-direction, and r represents the distance from the original point to the center of the image.

[0025] The formula for calculating the similarity between the two image patches is:

[0026]

[0027] Where S(A,B) NCC Let A(i,j) represent the similarity between image patch A and image patch B, where A(i,j) represents the gray value of the pixel in the i-th row and j-th column of the gray-level matrix of image patch A. Let B(i,j) represent the mean gray value of image block A, and let B(i,j) represent the gray value of the pixel in the i-th row and j-th column of the gray matrix of image block B. This represents the average gray level of image block B.

[0028] Furthermore, step S3 specifically includes:

[0029] make n ξ cam ∈SE(3) represents the relative pose of the rolling shutter camera n with respect to the coordinate system of the center of the rolling shutter camera group, that is, the relative pose of the center of the rolling shutter camera group to the rolling shutter camera n, which is fixed and obtained through external parameter calibration;

[0030] Let P be a 3D world point. W =[x W ,y W ,z W ,1] T The coordinates of the rolling shutter camera in the n-coordinate system at times t1 and t2 are P(n,t1)=[x1,y1,z1,1], respectively. T P(n,t2)=[x2,y2,z2,1] T ,have:

[0031] P(n,t1)= n ξ cam cam ξ W (t1)P W

[0032]

[0033] in, cam ξ W This indicates the relative pose of the center of the rolling shutter camera group with respect to the world coordinate system. replace cam ξ W (t2) cam ξ W -1 (t1) represents the pose change of the roller shutter camera group from time t1 to time t2. cam ξ W (t1) represents the pose of the rolling shutter camera group relative to the world coordinate system at time t1. cam ξ W (t2) represents the pose of the rolling shutter camera group relative to the world coordinate system at time t2. cam ξ n This represents the relative pose of the rolling shutter camera group with respect to the rolling shutter camera n. n ξ cam This represents the relative pose of the rolling shutter camera n with respect to the rolling shutter camera group.

[0034] The normalized coordinates of 3D points P(n,t1) and P(n,t2) are respectively P c (n,t1)=[x1 / z1,y1 / z1,1] T P c (n,t2)=[x2 / z2,y2 / z2,1]T Through the camera intrinsic parameter matrix K n Projected onto pixel p uv (n,t1) and p uv The pixel positions of (n, t2) are as follows:

[0035] p uv (n,t1)=K n P c (n,t1)

[0036] p uv (n,t2)=K n P c (n,t2)

[0037] Substituting the pixel coordinates of the tracking points and their best matching points of the current row of images from each rolling shutter camera at the same time, normalized coordinates are obtained. The essential matrix E is obtained according to the epipolar constraint. The rotation matrix R and translation vector s are decomposed according to E = s^R to obtain the relative pose of the rolling shutter camera group at times t1 and t2. The relative pose and the initial pose are iteratively multiplied to obtain the row pose of the rolling shutter camera group image.

[0038] Furthermore, step S4 specifically includes:

[0039] A roller shutter image with radial distortion is used to generate a single line of curves in the distortion-free image by removing distortion from the scan lines. The epipolar curve of the roller shutter is formed by the intersection points of each line of curves in the distortion-free image with the epipolar line of the image. A distortion-free pixel is defined as u. undistorted =[u u ,v u ,1] T and the distorted pixel is u distorted =[u d ,v d ,1] T , where u u v u These represent the u-axis and v-axis coordinates of a distortion-free pixel on the pixel plane, respectively. d v d Let u and v represent the coordinates of the distorted pixel on the pixel plane, respectively. The camera intrinsic parameter matrix is: The relationships between distortion-free pixels, distorted pixels, and the camera intrinsic parameter matrix are as follows:

[0040] u undistorted =KP c_undistorted

[0041] u distorted =KP c_distorted

[0042] Among them, Pc_undistorted =[x u ,y u ,1] T and P c_distorted =[x d ,y d ,1] T x represents the distortion-free points and distortion points on the normalized plane in the camera coordinate system. u y u Let x and y represent the x-axis and y-axis coordinates of the distortion-free point on the normalized plane in the camera coordinate system, respectively. d y d f represents the x-axis and y-axis coordinates of the distortion point on the normalized plane in the camera coordinate system, respectively. x f y c represents the focal length along the x-axis and y-axis of the image plane, respectively. x c y These represent the coordinates of the principal point on the image plane;

[0043] Based on the above steps for modeling radially distorted images, the intersection of the distorted image curve and the epipolar line is found using the following least squares optimization method: If the epipolar line equation is known to be au + bv + c = 0, where a, b, and c are the epipolar line equation parameters;

[0044] When a≠0, for the distorted image row y d Solve for the objective function:

[0045]

[0046] in, x u =[(-b / a)v u -(c / a)-c x ] / f x y u =(v u -c y ) / f y The intersection point of the curve row and the epipolar line in the distorted image is obtained as u. * =[u u * ,v u * ,1] T , where u u * =-(b / a)v u * -(c / a), where r is the distance from the distortion-free point to the center of the image, and k1 and k2 are distortion parameters;

[0047] When a = 0, i.e., when the polar line is horizontal, solve for the objective function:

[0048]

[0049] Where, x u =(u u -c x ) / f x y u =(-c / bc) y ) / f y The intersection point u can be obtained. * =[u u * ,v u * ,1] T , where v u * =-c / b, the two intersection points are obtained in the current situation, the reprojection error between the feature point and the obtained intersection point is calculated based on the prior depth, and non-target intersection points are eliminated;

[0050] The tracking points in each row of the keyframe are taken as feature points for depth filtering. For a feature point p, the row of the image is called Row. p Based on the known inter-frame relative poses obtained in step S3, the parameter matrix [a,b,c] of the epipolar equation for the current frame is obtained. T Thus, a roller shutter epipolar curve is obtained in the current frame. The projection points of the feature points are searched on the roller shutter epipolar curve, and the optimal matching point position q is obtained through the comparison algorithm with minimum photometric error.

[0051] Further, step S5 specifically includes:

[0052] Let the inverse depth of a pixel in the image be ρ = z. -1 Satisfies the Gaussian distribution model:

[0053] p(ρ)=N(m,τ 2 )

[0054] Where z represents depth, m is the inverse depth mean, and τ 2 This represents inverse deep uncertainty;

[0055] Given a feature point p on a keyframe projected onto a normal frame as a projection point q, what is the relative pose of the lines containing these two points? q ξ p The depth estimate of a single measurement of a feature point is recovered by using triangulation.

[0056] The uncertainty of depth estimation is calculated by assuming that the feature point has a one-pixel deviation on the epipolar curve of the curtain.

[0057] Further, step S6 specifically includes:

[0058] The inverse depth mean and inverse depth uncertainty of camera feature points after the previous fusion are m and τ, respectively. 2 ;

[0059] When a new image arrives, a new triangulation is performed on the feature points to obtain a new inverse depth observation value corresponding to the current state, which also follows a Gaussian distribution:

[0060]

[0061] The new inverse depth observation corresponding to the current state is depth-fused with the inverse depth fused from the previous state. The resulting inverse depth distribution is as follows: have:

[0062]

[0063]

[0064] When inverse depth uncertainty If the depth is less than the preset inverse depth uncertainty threshold, it indicates that the depth has converged and the fusion calculation stops; otherwise, if the depth has not converged, the depth fusion continues to update the depth.

[0065] By incorporating depth-converged feature points into the map and minimizing reprojection error, the optimized pose of the rolling shutter camera group image is obtained.

[0066] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention utilizes the line-by-line exposure characteristic of a rolling shutter camera, sets up a rolling shutter camera group, and performs high-frequency tracking of the feature points of the camera group line by line; it uses the epipolar curve of the rolling shutter to achieve pixel search and matching, and performs depth filtering on the feature points through triangulation and depth fusion strategies; this invention effectively improves the tracking frequency of visual odometry, reduces system latency, and enables the robot to achieve accurate positioning under rapid movement, especially rapid rotation; this invention greatly reduces the amount of computation while ensuring similar accuracy, thereby achieving real-time performance. Attached Figure Description

[0067] Figure 1 This is a flowchart of the high-frequency visual odometry method based on a rolling shutter camera according to the present invention;

[0068] Figure 2 This is a schematic diagram of the polarity curve of the roller shutter in this invention. Detailed Implementation

[0069] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0070] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0071] It should be understood that although the terms first, second, third, etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first information may also be referred to as second information without departing from the scope of this invention, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0072] The present invention will now be described in detail with reference to the accompanying drawings. Unless otherwise specified, the features of the following embodiments and implementations can be combined with each other.

[0073] See Figure 1 The high-frequency visual odometry method based on a rolling shutter camera of the present invention specifically includes the following steps:

[0074] S1. Obtain the roller shutter distortion images from different perspectives in the current environment through the roller shutter camera group, and simultaneously obtain the image row index of the roller shutter camera group exposed at the same time.

[0075] It should be understood that the image row index indicates which row of each camera in the rolling camera group is being exposed at the same time. When estimating the pose of the camera group, a system of linear equations needs to be constructed for the row images exposed at the same time. The row index can be used to know the v coordinate of the pixel (u,v) in the current row of a rolling image.

[0076] Furthermore, the rolling shutter camera group includes multiple rolling shutter cameras, for example, four rolling shutter cameras. The four rolling shutter cameras cam1, cam2, cam3, and cam4 with radial distortion are placed in different directions, and each rolling shutter camera has an overlapping field of view with at least one other rolling shutter camera. The rolling shutter cameras rotate and translate in the x, y, and z directions, and the even-numbered rolling shutter cameras rotate 90° on the image plane. By setting the positions of multiple rolling shutter cameras in the rolling shutter camera group in this way, the rolling shutter distortion images from different perspectives of the current environment can be obtained through the rolling shutter camera group.

[0077] In this embodiment, obtaining the image row index of the roller shutter camera group exposed at the same time specifically includes: installing an LED light that flashes at a fixed frequency for each roller shutter camera, with the light occupying the entire field of view of the roller shutter camera; comparing the difference in the average grayscale value of the roller shutter camera row images of two consecutive frames; if the difference exceeds a preset light intensity threshold, it indicates that the current row of the current frame is being exposed when the LED light flashes; and returning the image row index of the current row, which is the image row index of the roller shutter camera group exposed at the same time.

[0078] It should be understood that the preset photometric threshold can be adjusted according to the brightness of the LED light.

[0079] S2. Obtain the tracking point on the current roller shutter distortion image row, remove the distortion from the roller shutter distortion image, reconstruct the previous frame image row block through rotation compensation, and find the best matching point on the reconstructed image row.

[0080] Specifically, the grayscale gradient of each row in the distorted image is calculated, and the points with the largest gradient changes are selected as tracking points for that row. It should be understood that the number of tracking points can be adjusted according to the degree of image distortion. If the distortion is severe, more points can be tracked per row, i.e., multiple tracking points are selected. For example, in this embodiment, the first three points are selected as tracking points. In the camera coordinate system, the first two parameters k1 and k2 of the Brownian radial distortion model are used to model the radially distorted row image, and the distortion points p on the normalized plane are... c_distorted =[x d ,y d ,1] T Perform distortion correction to obtain the original point p. c_undistorted =[x u ,y u ,1] T Each new image row is passed in. t Distortion correction is performed on the tracking points on the row image. The same row of the roller image has the same y-value. d But different x d Different y in corresponding space uThe position, i.e., the row of the distorted image of the roller shutter, after distortion correction, is restored to a curve on the original image. The row image is extracted from the previous frame. i Its rotation is R i First, the camera image is rotated and compensated using the homography matrix, then the distortion is removed and the row image is reconstructed. For the input image Row... t Specify reference row j , with its rotation R j To replace the rotation of the current row image, when Row t with Row j When the rotation between rows exceeds a set rotation threshold, the current row is updated to a new reference row. Based on the given camera intrinsic parameter matrix K, the row image is determined. j Row image i The homography relationship between them is expressed as H j,i =KR j R i T K -1 Based on this expression, homography transformation is performed on the pixel rows of the previous frame image, and block matching is performed on the previous frame image using the tracking points of the current row image to obtain the best matching point.

[0081] When performing block matching to obtain the best matching point, specifically, image blocks A of length L are taken on both sides of the tracking point, and then image blocks B of the same length are also taken on both sides of the N points with the largest gray-level gradients in the reconstruction row of the previous frame image. i Let i = 1, 2, ..., N, and calculate the similarity between two image patches. Select the two image patches with the highest similarity as the matching image patches to obtain the optimal matching point. For example, image patches A with a length of 8 can be selected on both sides of the tracking point. This length only needs to be sufficient for the window image to be distinguishable and the computational cost not to be too high. Usually, the top 3 points with the largest gray-level gradients are selected. Of course, the top N points with the largest gray-level gradients can also be selected according to actual needs. For example, when considering the number of comparisons, the top 6 points with the largest gray-level gradients can be selected, that is, image patches B of the same length are selected on both sides of the top 6 points. i ,i=1,2,...,6.

[0082] It should be understood that, let A be the window image of a tracking point in the current frame, and let B1, B2, B3, etc. be the window images of the points with the largest gradient changes in the previous frame. The similarity between this A and each B is compared, and the B with the highest calculated similarity is selected as the matching block of A. The center gradient point of this window image is the matching point of the tracking point. Essentially, for one A and multiple Bs, multiple similarities are calculated, and the image blocks A and B corresponding to the highest similarity are selected as the matching image blocks.

[0083] It should be noted that the reference row index j is not necessarily equal to the image row index t-1. The pixel offset between two consecutive rows in the same frame is very small and can be regarded as noise, which cannot be accurately measured.

[0084] It should be understood that when the camera only translates, the changes in image blocks are negligible, so it can be assumed that the image blocks remain unchanged when the camera only translates; however, when the camera rotates, the changes in image blocks are significant, so rotation compensation of the camera image is required before performing block matching.

[0085] Furthermore, the expression for the Brownian radial distortion model is:

[0086] x d =x u (1+k1r 2 +k2r 4 )

[0087] y d =y u (1+k1r 2 +k2r 4 )

[0088]

[0089] Where, x d Represents the coordinates of the distortion point in the x-direction, y-direction... d The x-coordinate represents the coordinate of the distorted point in the y-direction. u Represents the coordinates of the original point in the x-direction, y-direction... u This represents the coordinates of the original point in the y-direction, and r represents the distance from the original point to the center of the image.

[0090] Furthermore, the mean-reduced NCC is used to compare the similarity between two image patches. The formula for calculating the similarity between two image patches is:

[0091]

[0092] Where S(A,B) NCC Let A(i,j) represent the similarity between image patch A and image patch B, where A(i,j) represents the gray value of the pixel in the i-th row and j-th column of the gray-level matrix of image patch A. Let B(i,j) represent the mean gray value of image block A, and let B(i,j) represent the gray value of the pixel in the i-th row and j-th column of the gray matrix of image block B. This represents the average grayscale value of image patch B. It should be understood that a similarity close to 0 indicates that the two image patches are dissimilar, while a similarity close to 1 indicates that they are similar. Therefore, the image patch with the highest similarity is selected as the matching image patch, and the best matching pixel is obtained, which is the optimal matching point.

[0093] S3. Based on step S2, obtain the best matching point on the row image of each rolling shutter camera exposed at the same time, and estimate the row pose of the initial rolling shutter camera group image through epipolar constraints.

[0094] Specifically, let n ξ cam ∈SE(3) represents the relative pose of the rolling shutter camera n (n=1,2,3,4) with respect to the coordinate system of the center of the rolling shutter camera group, that is, the relative pose from the center of the rolling shutter camera group to the rolling shutter camera n, which is fixed and obtained through extrinsic parameter calibration. Let 3D world point P W =[x W ,y W ,z W ,1] T The coordinates of the rolling shutter camera in the n-coordinate system at times t1 and t2 are P(n,t1)=[x1,y1,z1,1], respectively. T P(n,t2)=[x2,y2,z2,1] T ,have:

[0095] P(n,t1)= n ξ cam cam ξ W (t1)P W

[0096]

[0097] in, cam ξ W This indicates the relative pose of the center of the rolling shutter camera group with respect to the world coordinate system. replace cam ξ W (t2) cam ξ W -1 (t1) represents the pose change of the roller shutter camera group from time t1 to time t2. cam ξ W (t1) represents the pose of the rolling shutter camera group relative to the world coordinate system at time t1. cam ξ W (t2) represents the pose of the rolling shutter camera group relative to the world coordinate system at time t2. cam ξ n This represents the relative pose of the rolling shutter camera group with respect to the rolling shutter camera n. n ξ cam This represents the relative pose of the rolling shutter camera n with respect to the rolling shutter camera group.

[0098] The normalized coordinates of 3D points P(n,t1) and P(n,t2) are respectively P c(n,t1)=[x1 / z1,y1 / z1,1] T P c (n,t2)=[x2 / z2,y2 / z2,1] T Through the camera intrinsic parameter matrix K n Projected onto pixel p uv (n,t1) and p uv The pixel positions of (n, t2) are as follows:

[0099] p uv (n,t1)=K n P c (n,t1)

[0100] p uv (n,t2)=K n P c (n,t2)

[0101] Substituting the pixel coordinates of the tracking points and their best matching points of the current row of images from each rolling shutter camera at the same time, normalized coordinates are obtained. The essential matrix E is obtained according to the epipolar constraint. The rotation matrix R and translation vector s are decomposed according to E = s^R to obtain the relative pose of the rolling shutter camera group at times t1 and t2. The relative pose and the initial pose are iteratively multiplied to obtain the row pose of the rolling shutter camera group image.

[0102] S4. Perform depth filtering on each row of feature points in the keyframe. When the new frame of the roller shutter image is input, calculate the parameter matrix of the epipolar equation of the current frame using the pose of the inter-frame row image. Construct the roller shutter epipolar curve using the least squares optimization method and search to obtain the best matching point position of the feature point in the current frame.

[0103] It should be understood that the first frame is taken as the initial keyframe, and a new keyframe is created when no feature points are observed on the keyframe in the new frame.

[0104] like Figure 2 As shown, the epipolar curve of the shutter is obtained by using the intersection of the epipolar line and the line scan line. A search is then performed on the epipolar curve to obtain the coordinates of the matching point, and the optimal matching point position is further determined. It should be understood that a matching point refers to a pixel that matches a feature point on a keyframe.

[0105] Specifically, for a roller shutter image with radial distortion, the line scan lines are distorted to generate a single line of curves in the distortion-free image. The epipolar curve of the roller shutter is formed by the intersection points of each line of curves in the distorted image and the epipolar line of the image. The distortion-free pixel is defined as u. undistorted =[u u ,v u ,1] T and the distorted pixel is udistorted =[u d ,v d ,1] T , where u u v u These represent the u-axis and v-axis coordinates of a distortion-free pixel on the pixel plane, respectively. d v d Let u and v represent the coordinates of the distorted pixel on the pixel plane, respectively. The camera intrinsic parameter matrix is: The relationships between distortion-free pixels, distorted pixels, and the camera intrinsic parameter matrix are as follows:

[0106] u undistorted =KP c_undistorted

[0107] u distorted =KP c_distorted

[0108] Among them, P c_undistorted =[x u ,y u ,1] T and P c_distorted =[x d ,y d ,1] T x represents the distortion-free points and distortion points on the normalized plane in the camera coordinate system. u y u Let x and y represent the x-axis and y-axis coordinates of the distortion-free point on the normalized plane in the camera coordinate system, respectively. d y d f represents the x-axis and y-axis coordinates of the distortion point on the normalized plane in the camera coordinate system, respectively. x f y c represents the focal length along the x-axis and y-axis of the image plane, respectively. x c y These represent the coordinates of the principal point on the image plane.

[0109] Based on the above steps for modeling radially distorted images, the intersection points of the distorted image curve rows and epipolar lines are found using the following least-squares optimization method. If the epipolar equation is known to be au + bv + c = 0, where a, b, and c are epipolar equation parameters; when a ≠ 0, for the distorted image row y... d Solve for the objective function:

[0110]

[0111] in, x u =[(-b / a)v u -(c / a)-c x ] / f xy u =(v u -c y ) / f y The intersection point of the curve row and the epipolar line in the distorted image can be obtained as u. * =[u u * ,v u * ,1] T , where u u * =-(b / a)v u * -(c / a), where r is the distance from the distortion-free point to the image center, and k1 and k2 are distortion parameters. When a = 0, i.e., the epipolar line is horizontal, the objective function is solved as follows:

[0112]

[0113] Where, x u =(u u -c x ) / f x y u =(-c / bc) y ) / f y The intersection point u can be obtained. * =[u u * ,v u * ,1] T , where v u * =-c / b, the current situation can obtain two intersection points. Calculate the reprojection error between the feature point and the obtained intersection point based on the prior depth, and eliminate non-target intersection points.

[0114] The tracking points in each row of the keyframe are taken as feature points for depth filtering. For a feature point p, the row of the image is called Row. p Based on the known inter-frame relative poses obtained in step S3, the parameter matrix [a,b,c] of the epipolar equation for the current frame can be obtained. T Thus, a roller shutter epipolar curve can be obtained on the current frame. The projection points of the feature points are searched on the roller shutter epipolar curve, and the optimal matching point position q is obtained through the comparison algorithm with minimum photometric error.

[0115] It should be understood that the minimum photometric error comparison algorithm can calculate the grayscale difference between two pixels. The pixel with the smallest difference can be considered to be the same pixel. It is based on the theory that the grayscale of the same pixel is consistent and the calculation method is simple.

[0116] S5. Based on the epipolar geometric constraints satisfied by the feature points and their best matching points, use triangulation to recover the depth and uncertainty of a single measurement of the feature points.

[0117] Specifically, assume that the inverse depth ρ = z of a pixel in the image -1 Satisfies the Gaussian distribution model:

[0118] p(ρ)=N(m,τ 2 )

[0119] Where z represents depth, m is the inverse depth mean, and τ 2 This represents inverse deep uncertainty.

[0120] Given a feature point p on a keyframe projected onto a normal frame as a projection point q, what is the relative pose of the lines containing these two points? q ξ p The depth estimate of a single measurement of a feature point is recovered using triangulation. The uncertainty of the depth estimate is calculated by assuming a one-pixel deviation of the feature point on the epipolar curve of the roller blind.

[0121] S6. The optimal depth of the previous state of the camera feature points and the current measured depth are fused. After the feature point depths on the reconstructed row converge, the feature points are added to the map, and the reprojection error is minimized to obtain the optimized row pose of the rolling shutter camera group image.

[0122] It should be understood that deep fusion is essentially a process of continuous depth optimization. The depth obtained at the previous moment before the current fusion is the optimized optimal depth. It's easy to understand that if this is the first deep fusion, then the optimal depth of the previous state is the initial depth obtained in step S5.

[0123] The inverse depth mean and inverse depth uncertainty of camera feature points after the previous fusion are m and τ, respectively. 2 When a new image arrives, a new triangulation is performed on the feature points to obtain a new inverse depth observation value corresponding to the current state, which also follows a Gaussian distribution:

[0124]

[0125] The new inverse depth observation corresponding to the current state is depth-fused with the inverse depth fused from the previous state. The resulting inverse depth distribution is as follows: have:

[0126]

[0127]

[0128] When inverse depth uncertainty If the depth uncertainty is less than the preset inverse depth uncertainty threshold, it indicates depth convergence, and the fusion calculation stops; otherwise, if the depth has not converged, depth fusion continues to update the depth. The feature points that have converged in depth are added to the map to minimize the reprojection error and obtain the optimized row pose of the rolling camera group image.

[0129] In summary, the high-frequency visual odometry method based on a rolling shutter camera described in this invention aims to improve the tracking frequency of visual odometry. This invention utilizes the line-by-line exposure characteristic of a rolling shutter camera, setting up a rolling shutter camera group and performing high-frequency tracking of feature points of the camera group line by line. Furthermore, this invention uses the epipolar curve of the rolling shutter to achieve pixel search and matching, and performs depth filtering on feature points through triangulation and depth fusion strategies to improve the robustness of the system. Ultimately, this invention enables visual odometry to achieve accurate positioning under rapid motion, especially rapid rotation. In experiments, the tracking frequency of this method is significantly higher than that of traditional visual odometry methods, indicating that the method of this invention can significantly reduce system latency under high-speed motion and, while ensuring similar accuracy, greatly reduce the computational load compared to existing methods. The innovation of this invention lies in the use of a geometric algorithm based on radial distortion of the rolling shutter line image, and the method of using the epipolar curve of the rolling shutter to achieve pixel search and matching and perform depth filtering on feature points. These innovations greatly improve the robustness and real-time performance of the system, enabling the rolling shutter camera visual odometry method to achieve accurate positioning under high-speed motion. Therefore, the method of the present invention has significant practical application value and can be widely applied in fields such as robot navigation, autonomous driving, and augmented reality / virtual reality, providing strong support for technological innovation and progress in related industries.

[0130] The above description is merely a preferred embodiment of the present invention. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall still fall within the protection scope of the technical solutions of the present invention.

Claims

1. A high-frequency visual odometry method based on a rolling shutter camera, characterized in that, Includes the following steps: S1. Obtain the roller shutter distortion images from different perspectives in the current environment through the roller shutter camera group, and simultaneously obtain the image row index of the roller shutter camera group exposed at the same time. S2. Obtain the tracking point on the current roller shutter distortion image row, remove the distortion of the roller shutter distortion image, reconstruct the previous frame image row block through rotation compensation, and find the best matching point on the reconstructed image row. S3. Based on step S2, obtain the best matching point on the row image of each rolling shutter camera exposed at the same time, and estimate the row pose of the initial rolling shutter camera group image by using epipolar constraints. S4. Perform depth filtering on each row of feature points in the keyframe. With the new frame's roller shutter image input, calculate the parameter matrix of the epipolar equation for the current frame using the inter-frame row image pose. Construct the roller shutter epipolar curve using the least squares optimization method and search for the optimal matching point position of the feature points in the current frame. Step S4 specifically includes: A roller shutter image with radial distortion is used to generate a single line of curves in the distortion-free image by removing distortion from the scan lines. The epipolar curve of the roller shutter is formed by the intersection points of each line of curves in the distortion-free image with the epipolar line of the image. A distortion-free pixel is defined as u. undistorted =[u u ,v u ,1] T and the distorted pixel is u distorted =[u d ,v d ,1] T , where u u v u These represent the u-axis and v-axis coordinates of a distortion-free pixel on the pixel plane, respectively. d v d Let u and v represent the coordinates of the distorted pixel on the pixel plane, respectively. The camera intrinsic parameter matrix is: The relationships between distortion-free pixels, distorted pixels, and the camera intrinsic parameter matrix are as follows: u undistorted =KP c_undistorted u distorted =KP c_distorted Among them, P c_undistorted =[x u ,y u ,1] T and P c_distorted =[x d ,y d ,1] T x represents the distortion-free points and distortion points on the normalized plane in the camera coordinate system. u y u Let x and y represent the x-axis and y-axis coordinates of the distortion-free point on the normalized plane in the camera coordinate system, respectively. d y d f represents the x-axis and y-axis coordinates of the distortion point on the normalized plane in the camera coordinate system, respectively. x f y c represents the focal length along the x-axis and y-axis of the image plane, respectively. x c y These represent the coordinates of the principal point on the image plane; Based on the above modeling steps for the roller blind image with radial distortion, the intersection point of the distortion-free image curve and the epipolar line is found by the following least squares optimization method: If the epipolar line equation is known to be au + bv + c = 0, where a, b, and c are the epipolar line equation parameters; When a≠0, for the distorted image row y d Solve for the objective function: in, x u =[(-b / a)v u -(c / a)-c x ] / f x y u =(v u -c y ) / f y The intersection point of the curve row and the epipolar line in the distorted image is obtained as u. * =[u u * ,v u * ,1] T ,in r is the distance from the distortion-free point to the center of the image, and k1 and k2 are the distortion parameters; When a = 0, i.e., when the polar line is horizontal, solve for the objective function: Where, x u =(u u -c x ) / f x y u =(-c / bc) y ) / f y The intersection point u can be obtained. * =[u u * ,v u * ,1] T , where v u * =-c / b, the two intersection points are obtained in the current situation, the reprojection error between the feature point and the obtained intersection point is calculated based on the prior depth, and non-target intersection points are eliminated; The tracking points in each row of the keyframe are taken as feature points for depth filtering. For a feature point p, the row of the image is called Row. p Based on the known inter-frame relative poses obtained in step S3, the parameter matrix [a,b,c] of the epipolar equation for the current frame is obtained. T Thus, a roller shutter epipolar curve is obtained in the current frame. The projection points of the feature points are searched on the roller shutter epipolar curve, and the optimal matching point position q is obtained through the comparison algorithm with minimum photometric error. S5. Based on the epipolar geometric constraints satisfied by the feature points and their best matching points, use triangulation to recover the depth and uncertainty of a single measurement of the feature points; S6. The optimal depth of the previous state of the camera feature points and the current measured depth are fused. After the feature point depths on the reconstructed row converge, the feature points are added to the map, and the reprojection error is minimized to obtain the optimized row pose of the rolling shutter camera group image.

2. The high-frequency visual odometry method based on a rolling shutter camera according to claim 1, characterized in that, The roller shutter camera group includes multiple roller shutter cameras, which are placed in different directions, and each roller shutter camera has an overlapping field of view with at least one other roller shutter camera. The roller shutter cameras rotate and translate in the x, y, and z directions, and the even-numbered roller shutter cameras rotate 90° on the image plane.

3. The high-frequency visual odometry method based on a rolling shutter camera according to claim 1, characterized in that, The process of obtaining the image row index of the rolling shutter camera group exposed at the same time specifically includes: installing an LED light that flashes at a fixed frequency for each rolling shutter camera, with the light occupying the entire field of view of the rolling shutter camera; comparing the difference in the average grayscale value of the row images of the rolling shutter camera in two consecutive frames; if the difference exceeds a preset light intensity threshold, it indicates that the current row of the current frame is being exposed when the LED light flashes; and returning the image row index of the current row, which is the image row index of the rolling shutter camera group exposed at the same time.

4. The high-frequency visual odometry method based on a rolling shutter camera according to claim 1, characterized in that, Step S2 specifically includes: Calculate the gray-level gradient of each row in the distorted image of the roller blind, and take the first few points with the largest gradient changes as the tracking points of the row image; In the camera coordinate system, the first two parameters k1 and k2 of the Brownian radial distortion model are used to model the radial distortion of the row image, and the distortion point p on the normalized plane is then used. c_distorted =[x d ,y d ,1] T Perform distortion correction to obtain the original point p. c_undistorted =[x u ,y u ,1] T ; Each new image row is passed in. t Distortion correction is performed on the tracking points on the row image. The same row of the roller image has the same y-value. d But different x d Different y in corresponding space u The position, that is, the line of the distorted image of the roller blind, after being deformed and restored to the original image, is a curve; Extract the row image from the previous frame. i Its rotation is R i First, the camera image is rotated and compensated using the homography matrix, then the distortion is removed and the row image is reconstructed. For the input image Row... t Specify reference row j , with its rotation R j To replace the rotation of the current row image, when Row t with Row j When the rotation between rows exceeds the set rotation threshold, the current row will be updated to a new reference row; Given the camera intrinsic parameter matrix K, determine the row image Row. j Row image i The homography relationship between them is expressed as follows: Based on this expression, homography transformation is performed on the pixel rows of the previous frame image. Block matching is then performed in the previous frame image using the tracking points of the current row image: image blocks A of length L are taken on both sides of the tracking point, and image blocks B of the same length are also taken on both sides of the N points with the largest gray-level gradients in the reconstructed rows of the previous frame image. i Let i = 1, 2, ..., N, and calculate the similarity between two image patches. Select the two image patches with the highest similarity as the matching image patches to obtain the best matching point.

5. The high-frequency visual odometry method based on a rolling shutter camera according to claim 4, characterized in that, The expression for the Brownian radial distortion model is: x d =x u (1+k1r 2 +k2r 4 ) and d / and u (1+k1r 2 +k2r 4 ) Where, x d Represents the coordinates of the distortion point in the x-direction, y-direction... d The x-coordinate represents the coordinate of the distorted point in the y-direction. u Represents the coordinates of the original point in the x-direction, y-direction... u This represents the coordinate of the original point in the y-direction, and r represents the distance from the original point to the center of the image. The formula for calculating the similarity between the two image patches is: Where S(A,B) NCC Let A(i,j) represent the similarity between image patch A and image patch B, where A(i,j) represents the gray value of the pixel in the i-th row and j-th column of the gray-level matrix of image patch A. Let B(i,j) represent the mean gray value of image block A, and let B(i,j) represent the gray value of the pixel in the i-th row and j-th column of the gray matrix of image block B. This represents the average gray level of image block B.

6. The high-frequency visual odometry method based on a rolling shutter camera according to claim 1, characterized in that, Step S3 specifically includes: make n ξ cam ∈SE(3) represents the relative pose of the rolling shutter camera n with respect to the coordinate system of the center of the rolling shutter camera group, that is, the relative pose of the center of the rolling shutter camera group to the rolling shutter camera n, which is fixed and obtained through external parameter calibration; Let P be a 3D world point. W =[x W ,y W ,z W ,1] T The coordinates of the rolling shutter camera in the n-coordinate system at times t1 and t2 are P(n,t1)=[x1,y1,z1,1], respectively. T P(n,t2)=[x2,y2,z2,1] T ,have: P(n,t1)= n ξ cam cam ξ W (t1)P W in, cam ξ W This indicates the relative pose of the center of the rolling shutter camera group with respect to the world coordinate system. replace cam ξ W (t2) cam ξ W -1 (t1) represents the pose change of the roller shutter camera group from time t1 to time t2. cam ξ W (t1) represents the pose of the rolling shutter camera group relative to the world coordinate system at time t1. cam ξ W (t2) represents the pose of the rolling shutter camera group relative to the world coordinate system at time t2. cam ξ n This represents the relative pose of the rolling shutter camera group with respect to the rolling shutter camera n. n ξ cam This represents the relative pose of the rolling shutter camera n with respect to the rolling shutter camera group. The normalized coordinates of 3D points P(n,t1) and P(n,t2) are respectively P c (n,t1)=[x1 / z1,y1 / z1,1] T P c (n,t2)=[x2 / z2,y2 / z2,1] T Through the camera intrinsic parameter matrix K n Projected onto pixel p uv (n,t1) and p uv The pixel positions of (n, t2) are as follows: p uv (n,t1)=K n P c (n,t1) p uv (n,t2)=K n P c (n,t2) Substituting the pixel coordinates of the tracking points and their best matching points of the current row of images from each rolling shutter camera at the same time, normalized coordinates are obtained. The essential matrix E is obtained according to the epipolar constraint. The rotation matrix R and translation vector s are decomposed according to E = s^R to obtain the relative pose of the rolling shutter camera group at times t1 and t2. The relative pose and the initial pose are iteratively multiplied to obtain the row pose of the rolling shutter camera group image.

7. The high-frequency visual odometry method based on a rolling shutter camera according to claim 1, characterized in that, Step S5 specifically includes: Let the inverse depth of a pixel in the image be ρ = z. -1 Satisfies the Gaussian distribution model: p(ρ)=N(m,τ 2 ) Where z represents depth, m is the inverse depth mean, and τ 2 This represents inverse deep uncertainty; Given a feature point p on a keyframe projected onto a normal frame as a projection point q, what is the relative pose of the lines containing these two points? q ξ p The depth estimate of a single measurement of a feature point is recovered by using triangulation. The uncertainty of depth estimation is calculated by assuming that the feature point has a one-pixel deviation on the epipolar curve of the curtain.

8. The high-frequency visual odometry method based on a rolling shutter camera according to claim 1, characterized in that, Step S6 specifically includes: The inverse depth mean and inverse depth uncertainty of camera feature points after the previous fusion are m and τ, respectively. 2 ; When a new image arrives, a new triangulation is performed on the feature points to obtain a new inverse depth observation value corresponding to the current state, which also follows a Gaussian distribution: The new inverse depth observation corresponding to the current state is depth-fused with the inverse depth fused from the previous state. The resulting inverse depth distribution is as follows: have: When inverse depth uncertainty If the depth is less than the preset inverse depth uncertainty threshold, it indicates that the depth has converged and the fusion calculation stops; otherwise, if the depth has not converged, the depth fusion continues to update the depth. By incorporating depth-converged feature points into the map and minimizing reprojection error, the optimized pose of the rolling shutter camera group image is obtained.

Citation Information

Patent Citations

  • Lightweight monocular vision positioning method based on dot-line features and depth filter

    CN110807809A

  • Automatic precision solving method for rotating linear array scanning image pose

    CN113962853A