A visual SLAM method and system for binocular fisheye cameras

Through the visual SLAM method of binocular fisheye camera, visual images are acquired in real time, feature points are extracted, the moving speed and parallax of visual feature points are calculated, and the visual image is corrected using the IMU external parameter matrix. This solves the problems of narrow field of view of binocular cameras and stitching errors of multiple monocular cameras, and realizes high-precision visual SLAM.

CN115731235BActive Publication Date: 2025-09-19BEIJING GREEN VALLEY TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211579115.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-08
Publication Date
2025-09-19
Estimated Expiration
2042-12-08

AI Technical Summary

Technical Problem

The existing binocular camera has a narrow field of view, and the stitching of images from multiple monocular cameras results in insufficient accuracy, resulting in high uncertainty in the calculated carrier pose.

Method used

A binocular fisheye camera is used to obtain visual images in real time, extract visual feature points, calculate the moving speed and parallax of the visual feature points, and use the IMU external parameter matrix to transform the residual information, correct the visual image, eliminate the residual, and improve the accuracy.

Benefits of technology

It realizes visual SLAM with a 360-degree field of view, reduces the amount of calculation, reduces image stitching errors, and improves the accuracy and stability of the carrier's posture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731235B_ABST
    Figure CN115731235B_ABST
Patent Text Reader

Abstract

The present invention discloses a visual SLAM method and system for a binocular fisheye camera, wherein the visual SLAM method for the binocular fisheye camera includes: acquiring a visual image of the binocular fisheye camera in real time and extracting visual feature points of the visual image; calculating the moving speed of the visual feature points based on the pixel positions of the visual feature points in the current frame and the previous frame visual image; calculating the disparity of the visual feature points based on the moving speed of the visual feature points; extracting residual information of the visual image of the current frame based on the magnitude relationship between the disparity of the visual feature points and a preset disparity threshold; transforming the residual information of the visual image of the current frame based on the extrinsic parameter transformation matrix of the binocular fisheye camera and the IMU to obtain the residual information of the visual image in the IMU coordinate system; and using the residual information of the visual image to correct the visual image of the binocular fisheye camera. The technical solution of the present invention can solve the problem in the prior art that images obtained by multiple monocular cameras through SLAM technology are insufficient in accuracy and the calculated carrier posture has high uncertainty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of binocular vision technology, and in particular to a visual SLAM method and system for a binocular fisheye camera. Background Art

[0002] SLAM (Simultaneous Localization and Mapping) technology allows a sensor-equipped object to model its surroundings while in motion, even without knowing the surroundings, while simultaneously estimating the device's own motion. This makes it widely used in fields such as radar scanning, mapping, and sensor monitoring. SLAM technology is generally divided into two categories based on the different SLAM sensors: laser SLAM when the sensor is a lidar, and visual SLAM when the sensor is a camera.

[0003] In visual SLAM, commonly used cameras include RGBD cameras, monocular cameras, and binocular cameras, each with its own advantages and disadvantages. Monocular cameras have a simple sensor structure and low cost, allowing the SLAM process to be completed using only a few cameras. However, monocular cameras suffer from a narrow field of view and scale uncertainty. RGBD cameras, using infrared structured light or the time-of-flight principle, can directly measure the distance between each pixel in an image, thus addressing the narrow field of view and scale uncertainty of monocular cameras. However, they consume high power, are costly, and are susceptible to interference. Therefore, binocular cameras are commonly used in the prior art for visual SLAM. Binocular cameras can estimate depth both in motion and at rest, eliminating the scale uncertainty inherent in monocular cameras.

[0004] Although binocular cameras have a larger field of view than monocular cameras and can observe richer information about the surrounding environment, the field of view of most binocular cameras currently on the market is only a little over 100 degrees, installed on the same side. Examples include the well-known ZED2 (110°(H)x70°(V)x120°(D)) and T265 (163°Fov(±5°)). Due to this phenomenon, many researchers use multiple monocular cameras to stitch images together to achieve a 360° panoramic view in order to expand the field of view. For example, autonomous parking vision solutions use multiple monocular cameras. This undoubtedly increases the computational workload, and errors are inevitable when stitching images from multiple cameras. This, in turn, results in images obtained by SLAM technology using multiple monocular cameras being less accurate and exhibiting significant positional offsets, increasing the uncertainty of the calculated vehicle pose. Summary of the Invention

[0005] The present invention provides a visual SLAM solution for a binocular fisheye camera, aiming to solve the problems of narrow field of view of binocular cameras provided by the prior art, insufficient accuracy of images obtained by multiple monocular cameras through SLAM technology, and high uncertainty in the calculated carrier pose.

[0006] To achieve the above object, according to a first aspect of the present invention, the present invention proposes a visual SLAM method for a binocular fisheye camera, comprising:

[0007] Acquire the visual image of the binocular fisheye camera in real time and extract the visual feature points of the visual image;

[0008] Calculate the moving speed of the visual feature point according to the pixel position of the visual feature point in the current frame visual image and the previous frame visual image;

[0009] Calculate the parallax of the visual feature points according to the moving speed of the visual feature points;

[0010] Extracting residual information of the current frame visual image based on the relationship between the disparity of the visual feature points and the preset disparity threshold;

[0011] According to the external parameter transformation matrix of the binocular fisheye camera and the IMU, the residual information of the visual image of the current frame is transformed to obtain the residual information of the visual image in the IMU coordinate system;

[0012] The visual image of the binocular fisheye camera is corrected using the residual information of the visual image.

[0013] Preferably, in the above-mentioned visual SLAM method, the step of extracting visual feature points of the visual image includes:

[0014] Extract multiple visual feature points of the current frame visual image;

[0015] Using the visual feature points of the current frame visual image, tracking the matching visual feature points in the previous frame visual image;

[0016] Use the F matrix to verify the visual feature points of the previous visual image frame and filter out the visual feature points with incorrect matching in the previous visual image frame;

[0017] The visual feature points that are successfully matched in the previous frame of visual image are dedistorted to obtain the visual feature points with corrected coordinates.

[0018] Preferably, the above-mentioned visual SLAM method further comprises, after the step of tracking the matching visual feature points in the previous frame of visual image:

[0019] When the visual feature points of the current frame visual image fail to match, a blank image with the same size as the current frame visual image is set;

[0020] Calculate the pixel positions of the blank image corresponding to the visual feature points where matching fails;

[0021] In a blank area outside a predetermined distance range around the pixel position in the blank image, visual feature points that match the visual feature points that failed to match are tracked and obtained as visual feature points in the previous frame of visual image.

[0022] Preferably, in the above-mentioned visual SLAM method, the step of calculating the moving speed of the visual feature point according to the pixel position of the visual feature point in the current frame visual image and the previous frame visual image includes:

[0023] Obtain the interval time between the current frame visual image and the previous frame visual image;

[0024] Calculate the moving distance of the visual feature point according to the pixel position of the visual feature point of the current frame visual image and the pixel position of the previous frame visual image;

[0025] The moving distance and interval time of the visual feature points are used to calculate the moving speed of the visual feature points.

[0026] Preferably, in the above-mentioned visual SLAM method, the step of calculating the parallax of the visual feature points according to the moving speed of the visual feature points includes:

[0027] According to the moving speed of the visual feature points, the sum of the coordinate differences of all visual feature points in two adjacent frames of visual images is calculated;

[0028] The average value of the sum of the coordinate differences of all visual feature points is calculated to obtain the disparity of the visual feature points.

[0029] Preferably, in the above-mentioned visual SLAM method, the step of extracting residual information of the current frame visual image according to the magnitude relationship between the disparity of the visual feature points and the preset disparity threshold comprises:

[0030] Determine whether the disparity of the visual feature point is greater than or equal to a preset disparity threshold;

[0031] If the disparity of the visual feature point is greater than or equal to the preset disparity threshold, the IMU pre-integrated velocity increment difference between the current frame visual image and the previous frame visual image is calculated;

[0032] Determine whether the second norm of the IMU pre-integrated velocity increment difference is greater than a predetermined increment difference threshold;

[0033] If the second norm of the IMU pre-integrated velocity increment difference is greater than the predetermined increment difference threshold, the current frame image is set as a key frame, and the residual information of the key frame is extracted using a sliding window.

[0034] Preferably, the above-mentioned visual SLAM method further comprises, after the step of extracting the residual information of the current frame visual image:

[0035] Initialize the camera parameters of any camera in the binocular fisheye camera;

[0036] The relative position relationship of the binocular fisheye cameras is used to update the residual information of the other camera in the binocular fisheye camera to obtain the updated residual information of the current frame visual image.

[0037] Preferably, in the above-mentioned visual SLAM method, the step of transforming the residual information of the current frame visual image according to the extrinsic parameter transformation matrix of the binocular fisheye camera and the IMU includes:

[0038] Use the IMU's external parameter transformation matrix to transform the residual information of the current frame visual image to obtain the residual information of the visual image in the IMU coordinate system;

[0039] The LM optimization algorithm is used to optimize the residual information of the visual image in the IMU coordinate system to obtain the optimized residual information of the visual image.

[0040] According to a second aspect of the present invention, the present invention further provides a visual SLAM system for a binocular fisheye camera, the visual SLAM system being used for a binocular fisheye camera, the visual SLAM system comprising:

[0041] A visual feature extraction module is used to obtain the visual image of the binocular fisheye camera in real time and extract the visual feature points of the visual image;

[0042] A moving speed calculation module is used to calculate the moving speed of the visual feature point based on the pixel position of the visual feature point in the current frame visual image and the previous frame visual image;

[0043] A disparity calculation module, used to calculate the disparity of the visual feature points according to the moving speed of the visual feature points;

[0044] The residual information extraction module is used to extract the residual information of the current frame visual image based on the size relationship between the disparity of the visual feature points and the preset disparity threshold;

[0045] The residual information transformation module is used to transform the residual information of the current frame visual image according to the external parameter transformation matrix of the binocular fisheye camera and the IMU, and obtain the residual information of the visual image in the IMU coordinate system;

[0046] The visual image correction module is used to correct the visual image of the binocular fisheye camera using the residual information of the visual image.

[0047] According to a third aspect of the present invention, the present invention further provides a visual SLAM system of a binocular fisheye camera, comprising:

[0048] A memory, a processor, and a visual SLAM program for a binocular fisheye camera stored in the memory and running on the processor. When the visual SLAM program is executed by the processor, the steps of the visual SLAM method for a binocular fisheye camera described in any of the above technical solutions are implemented.

[0049] In summary, the visual SLAM scheme of the binocular fisheye camera mentioned above in the present invention obtains the visual image of the binocular fisheye camera in real time, extracts the visual feature points of the visual image, and then calculates the moving speed of the visual feature point based on the pixel position of the visual feature point in the current frame and the previous frame visual image. In this way, the disparity of the visual feature point can be calculated based on the moving speed of the visual feature point, and the size of the disparity of the visual feature point and the preset disparity threshold can be used to determine whether there is a residual in the current frame visual image. If there is a residual, the residual information of the current frame visual image is extracted, and then the residual information of the current frame visual image is transformed according to the external parameter transformation matrix of the binocular fisheye camera and the IMU, so that the residual information of the visual image in the IMU coordinate system can be obtained, thereby unifying the current frame visual image and the previous frame visual image to the IMU coordinate system, and then using the residual information of the visual image to correct the visual image of the binocular fisheye camera, thereby eliminating the residual in the visual image. Because a binocular fisheye camera is used, the field of view of the binocular fisheye camera can reach 360 degrees, and there is no need to stitch multiple monocular cameras, thereby reducing the amount of calculation and the error caused by stitching multiple camera images. In addition, the above scheme can quickly and accurately calculate the residual information of the current frame visual image of the binocular fisheye camera, and use this residual information to verify the visual image of the binocular fisheye camera, thereby eliminating feature points with position offset, adjusting the object posture in the visual image, and improving the accuracy of the binocular fisheye camera. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying any creative work.

[0051] Figure 1 1 is a schematic structural diagram of a binocular fisheye camera provided by an embodiment of the present invention;

[0052] Figure 2 Schematic diagram of the principle of a visual SLAM method provided by an embodiment of the present invention;

[0053] Figure 3 1 is a flow chart of a first visual SLAM method using a binocular fisheye camera provided by an embodiment of the present invention;

[0054] Figure 4 yes Figure 3 A schematic flow chart of a method for extracting visual feature points provided by the illustrated embodiment;

[0055] Figure 5 yes Figure 4 A schematic flow chart of a method for tracking visual feature points provided by the illustrated embodiment;

[0056] Figure 6 yes Figure 3 A schematic flow chart of a method for calculating the moving speed of a visual feature point provided by the illustrated embodiment;

[0057] Figure 7 yes Figure 3 A schematic flow chart of a method for calculating the disparity of visual feature points provided by the illustrated embodiment;

[0058] Figure 8 yes Figure 3 A schematic flow chart of a method for calculating residual information provided by the illustrated embodiment;

[0059] Figure 9 1 is a flow chart of a second visual SLAM method using a binocular fisheye camera provided by an embodiment of the present invention;

[0060] Figure 10 yes Figure 3 A schematic flow chart of a method for transforming residual information provided by the illustrated embodiment;

[0061] Figure 11 1 is a schematic structural diagram of a visual SLAM system using a first binocular fisheye camera according to an embodiment of the present invention;

[0062] Figure 12 1 is a schematic structural diagram of a visual SLAM system using a second binocular fisheye camera according to an embodiment of the present invention.

[0063] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0064] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0065] The main technical problems solved by the embodiments of the present invention are:

[0066] In current research, although binocular cameras have a larger field of view than monocular cameras and can observe richer information about the surrounding environment, the field of view of most binocular cameras currently on the market is only a little over 100 degrees, installed on the same side. Examples include the well-known ZED2 (110°(H)x70°(V)x120°(D)) and T265 (163°Fov(±5°)). Due to this phenomenon, many researchers have adopted image stitching from multiple monocular cameras to achieve a 360° panoramic view in order to expand the field of view. For example, autonomous parking vision solutions use multiple monocular cameras. This undoubtedly increases the computational complexity, and errors are inevitable when stitching images from multiple cameras. This, in turn, results in images obtained using SLAM technology from multiple monocular cameras being less accurate and exhibiting significant positional offsets, increasing the uncertainty of the calculated vehicle pose.

[0067] To solve the above problems, see Figure 1 , Figure 1 The embodiment of the present invention provides a visual SLAM solution for binocular fisheye cameras. The binocular fisheye cameras include a left eye 1 and a right eye 2. The two fisheye cameras can provide a 360-degree wide field of view and can be applied to outdoor scenes with a large field of view. Compared with the images obtained by multiple monocular cameras through SLAM technology, it uses the rich features of the two camera images and the software algorithm to improve the carrier's field of view. In addition, combined with Figure 2 As can be seen from the flowchart shown, the visual images captured by the fisheye camera 1 and the fisheye camera 2 of the present application undergo the steps of feature extraction 1, residual information extraction 2, initialization 3, optimization 4 and output pose 5, thereby completing the entire visual SLAM process.

[0068] In addition, the above-mentioned visual feature points can be integrated with IMU and GPS information to accurately calculate the disparity between the current frame visual image and the previous frame visual image, and then calculate the residual information of the visual image in the IMU coordinate system, thereby eliminating the residual of the current visual image, reducing the error caused by the image stitching of multiple monocular cameras, and then eliminating the visual feature points with position offset in the visual image, so as to achieve the purpose of adjusting the object posture and improving the accuracy and stability of the binocular fisheye camera.

[0069] To achieve the above purpose, please see Figure 3 , Figure 3 FIG. 1 is a flow chart of a first binocular fisheye camera visual SLAM method provided by an embodiment of the present invention. Figure 3 As shown, the visual SLAM method of the binocular fisheye camera includes:

[0070] S110: Acquire the visual image of the binocular fisheye camera in real time and extract the visual feature points of the visual image. For the visual image captured by the binocular fisheye camera, the visual feature points in the visual image can be extracted using visual feature extraction technology. Taking into account the real-time problem, the embodiment of the present application extracts a certain number (for example, 200) of feature points from the two images of the binocular fisheye camera respectively, and then uses multi-layer optical flow technology for tracking. Due to the existence of visual feature points to be tracked and identified, the number of visual feature points obtained in the current frame image will be insufficient. Therefore, after tracking, in order to avoid the visual feature points being too concentrated, it is necessary to track and extract features again outside a certain radius range of the previously extracted visual feature points.

[0071] Specifically, as a preferred embodiment, Figure 4 As shown, the above step S110: extracting visual feature points of the visual image includes:

[0072] S111: Extract multiple visual feature points from the current frame visual image. In this embodiment of the present application, the binocular fisheye camera captures images in real time at predetermined intervals. In this way, the present embodiment extracts a certain number of visual feature points (e.g., 200) from the current frame visual image. This allows the use of multiple visual feature points to verify errors in the current frame visual image.

[0073] S112: Use the visual feature points of the current frame visual image to track the matching visual feature points in the previous frame visual image. The embodiment of the present application uses multi-layer optical flow tracking technology to track the visual feature points of the previous frame visual image. Specifically, the N visual feature points mentioned in the current frame visual image are used to track and match the visual feature points in the previous frame visual image. The method of tracking matching visual feature points by multi-layer optical flow tracking technology is mainly to track by utilizing the photometric consistency of pixel points, that is, when the time interval is very small (for example, between two consecutive frames of a video), because the interval time is too short, the brightness of the same visual feature point will not change when it moves between different frames.

[0074] S113: Using the F matrix to verify the visual feature points of the previous visual image frame, and filter out the visual feature points with incorrect matching in the previous visual image frame.

[0075] The F matrix is ​​a basic matrix commonly used in visual feature extraction. The specific content of the F matrix will not be repeated here. The classification of the solution methods of the F matrix can be roughly divided into linear estimation based on algebraic error and nonlinear estimation based on geometric error. Among them, in the linear estimation based on algebraic error, no matter which F matrix is ​​used, its final form is generally Ax=b, where A is a matrix composed of matched points, and x is a vector composed of elements of the basic matrix to be solved, and b varies according to the solution method. By using the F matrix to test the tracked and matched visual feature points, a part of the mismatched visual feature points can be screened out, thereby improving the matching accuracy of the visual images of the previous and next frames.

[0076] S114: Dedistorting the successfully matched visual feature points in the previous frame of visual image to obtain visual feature points with corrected coordinates.

[0077] During the dedistortion process for visual feature points, distortion parameters are obtained through camera calibration using intrinsic parameters. The pixel coordinates of the feature points are then corrected based on the camera's distortion model. Typically, the camera's distortion model is built-in and customized for each camera type. Dedistortion processing yields visual feature points with corrected coordinates.

[0078] In addition, in the above-mentioned multi-layer optical flow tracking technology, there may be a situation where some of the visual feature points of the previous frame visual image are matched, while some are not matched. In order to solve this problem and ensure that the number of visual feature points of the previous frame visual image is sufficient, it is necessary to count the number of visual feature points that failed to match in the previous frame visual image and extract the corresponding number of visual feature points in the previous frame visual image.

[0079] Specifically, as a preferred embodiment, Figure 5 As shown, the above-mentioned visual SLAM method further includes, after step S112: tracking the matching visual feature points in the previous frame visual image:

[0080] S1121: When the visual feature points of the previous visual image fail to match, a blank image of the same size as the current visual image is set. The matching failure here refers to the visual feature points that are screened out in the previous visual image and the visual feature points that are distorted.

[0081] S1122: Calculate the pixel positions of the blank image corresponding to the visual feature points where matching fails.

[0082] S1123: In a blank area outside a predetermined distance range around the pixel position in the blank image, visual feature points that match the visual feature points that failed to match are tracked and obtained as visual feature points in the previous frame of visual image.

[0083] The technical solution provided by the embodiment of the present application is to set a blank image of the same size as the current frame visual image when the visual feature points of the previous frame visual image fail to match. In this way, a blank image is set, the size of the blank image is the same as the size of the current frame visual image, and in the blank image, N feature points (for example, 200) of the previous frame visual image are used as the center, and the surrounding moving range of the visual feature point is set to black. Then, the visual feature points that match the above-mentioned visual feature points that failed to match are tracked in the white area outside the black area, so as to be used as the visual feature points in the above-mentioned previous frame visual image, thereby ensuring that the number of visual feature points of the current frame visual image is the same as that of the previous frame visual image.

[0084] After matching the visual feature points of the previous frame visual image through the stream tracing technology, it is necessary to calculate the moving speed of the visual feature points. The moving speed of the visual feature points in the horizontal and vertical directions can be obtained by the pixel position of the same visual feature point in the current frame visual image and the previous frame visual image. Figure 3 As shown, the visual SLAM method also includes:

[0085] S120: Calculating the moving speed of the visual feature point according to the pixel position of the visual feature point in the current frame visual image and the previous frame visual image.

[0086] Specifically, as a preferred embodiment, Figure 6 As shown, the step of calculating the moving speed of the visual feature point according to the pixel position of the visual feature point in the current frame visual image and the previous frame visual image includes:

[0087] S121: Obtaining the interval time between the current frame visual image and the previous frame visual image.

[0088] S122: Calculate the moving distance of the visual feature point according to the pixel position of the visual feature point of the current frame visual image and the pixel position of the previous frame visual image.

[0089] S123: Calculate the moving speed of the visual feature point using the moving distance and interval time of the visual feature point.

[0090] In the technical solution provided in the embodiment of the present application, for the successfully matched visual feature points, the time difference between two adjacent frames of visual images is calculated. Since the capture of different visual images has a certain time stamp, the calculation formula for the interval time is as follows:

[0091] dt = cur_time - prev_time; / / Current frame timestamp - previous frame timestamp

[0092] After obtaining the pixel position of the visual feature point in the current frame visual image and the pixel position of the previous frame visual image, the movement distance of the visual feature point can be calculated. The movement distance includes the movement distance in the horizontal and vertical directions. Then, the movement distance in the horizontal and vertical directions is divided by the above time interval to calculate the speed in the horizontal and vertical directions. The calculation formula is as follows:

[0093] v_x=(cur_un_pts[i].x-prev_un_pts[i].x) / dt;

[0094] v_y=(cur_un_pts[i].y-prev_un_pts[i].y) / dt.

[0095] Because there are a large number of visual feature points in the current frame visual image and the previous frame visual image, by calculating the moving speed of the above visual features, it is possible to determine whether there is a moving object in the current frame visual image and whether there is parallax between the current frame visual image and the previous frame visual image. When it is determined that there is parallax in the current frame visual image, the parallax of the visual feature points can be calculated. Specifically, after calculating the moving speed of the visual feature points, Figure 3 The visual SLAM method provided by the illustrated embodiment further includes the following steps:

[0096] S130: Calculating the parallax of the visual feature point according to the moving speed of the visual feature point.

[0097] As a preferred embodiment, Figure 7 As shown, in the above-mentioned visual SLAM method, step S130: the step of calculating the parallax of the visual feature point according to the moving speed of the visual feature point includes:

[0098] S131: Calculate the sum of the coordinate differences of all visual feature points in two adjacent visual image frames based on the movement speed of the visual feature points. A large number of visual feature points of stationary objects in the two images are selected. By combining the movement speed of each visual feature point with the aforementioned interval time, the coordinate difference between the two adjacent visual image frames can be obtained, and then the sum of the coordinate differences can be calculated.

[0099] S132: Calculate the average value of the sum of the coordinate differences of all visual feature points to obtain the disparity of the visual feature points.

[0100] In the technical solution provided by the embodiment of the present application, after obtaining the movement speed of a large number of visual feature points, it is possible to determine whether there is parallax in the current frame visual image based on the movement speed of the large number of visual feature points. If there is parallax, the parallax of the visual feature points can be calculated by summing the coordinate differences of all visual feature points in two adjacent frames of visual images. The specific parallax calculation method is as follows:

[0101] Assume that the 3D point (normalized coordinates) P1(u_1,v_1,1) in the current visual image and the matching visual feature point P0(u_0,v_0,1) in the previous visual image are in the same world coordinate system. In this way, the coordinate difference between the two visual feature points can be calculated: d_u = u_1 - u_0, d_v = v_1 - v_0. The parallax of a single point is p = (d_u*d_u+d_v*d_v)0.5. The final average parallax of the two images is: Parallax = (pi+pi+1+...+pi+n) / n, where n is the number of matching points.

[0102] After obtaining the disparity of the visual feature points, Figure 3 The visual SLAM method provided by the illustrated embodiment also includes:

[0103] S140: Extracting residual information of the current frame visual image according to the magnitude relationship between the disparity of the visual feature points and a preset disparity threshold.

[0104] By calculating the sum of the coordinate differences of the visual feature points in two adjacent images and then averaging them as the average disparity, the system can determine whether the current visual frame is a key frame by comparing the average disparity with a preset disparity threshold. If the disparity of the visual feature points is greater than or equal to the preset disparity threshold, the current visual frame is preliminarily determined to be a key frame. The IMU pre-integrated velocity of the key frame is then used to determine whether the key frame has residual information.

[0105] Specifically, as a preferred embodiment, Figure 8 As shown, the step of extracting residual information of the current frame visual image based on the magnitude relationship between the disparity of the visual feature points and the preset disparity threshold includes:

[0106] S141: Determine whether the disparity of the visual feature point is greater than or equal to a preset disparity threshold.

[0107] S142: If the disparity of the visual feature point is greater than or equal to the preset disparity threshold, the IMU pre-integrated velocity increment difference between the current frame visual image and the previous frame visual image is calculated.

[0108] S143: Determine whether the second norm of the IMU pre-integrated velocity increment difference is greater than a predetermined increment difference threshold;

[0109] S144: If the binary norm of the IMU pre-integrated velocity increment difference is greater than a predetermined increment difference threshold, the current frame image is set as a key frame, and the residual information of the key frame is extracted using a sliding window.

[0110] In the technical solution provided by the embodiment of the present application, it is considered that when the camera is stationary, the movement of dynamic targets will cause excessive parallax. Such excessive parallax will be continuously and erroneously added to the key frame, causing a large offset in the image coordinates of the key frame. To solve the above problem, the embodiment of the present application introduces the pre-integration speed of the IMU, and uses the pre-integration speed as a reference. Specifically, when the incremental difference between the IMU pre-integration speed of the current frame visual image and the previous frame visual image is greater than the predetermined incremental difference threshold, the current frame image is finally determined as a key frame, and the residual information of the key frame is collected using a sliding window. The specific calculation formula is as follows:

[0111] Assume that the IMU pre-integrated velocity increment of the current visual image is deltaV0, and the IMU pre-integrated velocity increment of the previous visual image is deltaV1. Determine whether the two-norm of the difference between deltaV0 and deltaV1 is greater than th, that is:

[0112] ||deltaV0-deltaV1||2>th, so as to finally determine whether the visual feature of the current frame is a key frame. When it is finally determined that the visual image of the current frame is not a key frame, the sliding window is used to collect the residual information of the key frame.

[0113] After using the sliding window to collect the residual information of the key frame, the binocular camera needs to be initialized, and then the residual information of the other camera is updated using the relative pose relationship between the two cameras of the binocular fisheye camera.

[0114] Specifically, as a preferred embodiment, Figure 9 As shown, the visual SLAM method provided in the embodiment of the present application further includes, after the above-mentioned step of extracting the residual information of the visual image of the current frame:

[0115] S210: Initializing the camera parameters of any one of the binocular fisheye cameras; the camera parameters include project bias, velocity, gravity, scale factor, etc. Initializing the camera parameters resets the original values ​​of the camera parameters.

[0116] S220: Using the relative position relationship of the binocular fisheye cameras, updating the residual information of the other camera in the binocular fisheye cameras to obtain the updated residual information of the current frame visual image.

[0117] In the technical solution provided in the embodiment of the present application, the binocular fisheye camera is initialized. After one of the binocular cameras completes the calculation of the bias, velocity, gravity and scale factor through initialization, the feature point information extracted from the other camera (i.e., the above-mentioned residual information) is updated using the relative posture relationship between the two cameras.

[0118] In the technical solution provided by the embodiment of the present application, after initializing the camera parameters of the binocular fisheye camera and updating the residual information, the feature point information extracted by the two cameras is used for optimization, specifically as follows: Figure 3 As shown, the above-mentioned visual SLAM method also includes:

[0119] S150: According to the extrinsic parameter transformation matrix of the binocular fisheye camera and the IMU, the residual information of the visual image of the current frame is transformed to obtain the residual information of the visual image in the IMU coordinate system. The residual information of the visual image includes the visual residual of the visual feature points, the IMU pre-integration residual, the marginalization residual, and the GPS residual.

[0120] Specifically, as a preferred embodiment, Figure 10 As shown, the above-mentioned step of transforming the residual information of the current frame visual image according to the external parameter transformation matrix of the binocular fisheye camera and the IMU includes:

[0121] S151: Using the IMU's extrinsic parameter transformation matrix to transform the residual information of the current frame visual image, obtaining the residual information of the visual image in the IMU coordinate system. By matching the residual information of the current frame visual image extracted by the two cameras with the IMU's extrinsic parameter transformation matrix, the residual information of the current frame visual image can be converted to the IMU coordinate system, thereby optimizing the residual information of the visual image.

[0122] S152: Optimize the residual information of the visual image in the IMU coordinate system using the Levenberg-Marquardt optimization algorithm to obtain optimized residual information of the visual image. The Levenberg-Marquardt optimization algorithm is a commonly used optimization algorithm for binocular cameras and will not be described in detail in the present embodiment.

[0123] In the technical solution provided in the embodiment of the present application, by optimizing the residual information extracted by the two cameras, the above residual information can be unified into the same IMU coordinate system. According to the external parameter transformation matrix between the residual information of the two visual cameras and the IMU, the residual information such as the visual residual of the visual feature point, the IMU pre-integration residual, the marginalization residual and the GPS residual are calculated. Finally, the LM method is used to optimize the many residuals, so that the above residual information is unified into the IMU coordinate system. Then, the residual information of the visual image is used to correct the visual image to eliminate the residual of the visual image.

[0124] After obtaining the residual information of the visual image, Figure 3 The visual SLAM method shown also includes:

[0125] S160: Correcting the visual image of the binocular fisheye camera using residual information of the visual image.

[0126] In summary, the visual SLAM method of the binocular fisheye camera provided by the above embodiment of the present invention obtains the visual image of the binocular fisheye camera in real time, extracts the visual feature points of the visual image, and then calculates the moving speed of the visual feature point based on the pixel position of the visual feature point in the current frame and the previous frame visual image. In this way, the disparity of the visual feature point can be calculated based on the moving speed of the visual feature point, and whether there is a residual in the current frame visual image can be determined based on the size of the disparity of the visual feature point and the preset disparity threshold. If there is a residual, the residual information of the current frame visual image is extracted, and then the residual information of the current frame visual image is transformed according to the external parameter transformation matrix of the binocular fisheye camera and the IMU, so that the residual information of the visual image in the IMU coordinate system can be obtained, thereby unifying the current frame visual image and the previous frame visual image into the IMU coordinate system, and then using the residual information of the visual image to correct the visual image of the binocular fisheye camera, thereby eliminating the residual in the visual image. Because a binocular fisheye camera is used, the field of view of the binocular fisheye camera can reach 360 degrees, and there is no need to stitch multiple monocular cameras, thereby reducing the amount of calculation and the error caused by stitching multiple camera images. In addition, the above scheme can quickly and accurately calculate the residual information of the current frame visual image of the binocular fisheye camera, and use this residual information to verify the visual image of the binocular fisheye camera, thereby eliminating feature points with position offset, adjusting the object posture in the visual image, and improving the accuracy of the binocular fisheye camera.

[0127] In addition, based on the same concept of the above-mentioned method embodiment, the embodiment of the present invention also provides a visual SLAM system of a binocular fisheye camera, which is used to implement the above-mentioned method of the present invention. Since the principles and methods of solving the problems in this system embodiment are similar, it has at least all the beneficial effects brought by the technical solutions of the above-mentioned embodiments, and will not be repeated here one by one.

[0128] See also Figure 11 , Figure 11 The following is a schematic diagram of the structure of a visual SLAM system using a binocular fisheye camera according to an embodiment of the present invention. Figure 11 As shown, the visual SLAM system of the binocular fisheye camera is used Figure 1 The binocular fisheye camera shown in the figure, the visual SLAM system includes:

[0129] A visual feature extraction module 110 is used to acquire a visual image from a binocular fisheye camera in real time and extract visual feature points of the visual image;

[0130] A moving speed calculation module 120 is used to calculate the moving speed of the visual feature point based on the pixel position of the visual feature point in the current frame visual image and the previous frame visual image;

[0131] a disparity calculation module 130 for calculating the disparity of the visual feature points according to the moving speed of the visual feature points;

[0132] The residual information extraction module 140 is used to extract the residual information of the current frame visual image based on the relationship between the disparity of the visual feature points and a preset disparity threshold;

[0133] The residual information transformation module 150 is used to transform the residual information of the visual image of the current frame according to the external parameter transformation matrix of the binocular fisheye camera and the IMU to obtain the residual information of the visual image in the IMU coordinate system;

[0134] The visual image correction module 160 is configured to correct the visual image of the binocular fisheye camera using residual information of the visual image.

[0135] In summary, the visual SLAM system of the binocular fisheye camera provided by the above embodiment of the present invention obtains the visual image of the binocular fisheye camera in real time through the visual feature extraction module 110, extracts the visual feature points of the visual image, and then the moving speed calculation module 120 can calculate the moving speed of the visual feature point according to the pixel position of the visual feature point in the current frame and the previous frame visual image, so that the disparity calculation module 130 can calculate the disparity of the visual feature point according to the moving speed of the visual feature point, and the residual information extraction module 140 can calculate the disparity of the visual feature point according to the disparity of the visual feature point and the preset disparity. The size of the threshold can determine whether there is a residual in the current frame visual image. If there is a residual, the residual information of the current frame visual image is extracted. Then, the residual information transformation module 150 transforms the residual information of the current frame visual image according to the external parameter transformation matrix of the binocular fisheye camera and the IMU, so that the residual information of the visual image in the IMU coordinate system can be obtained, thereby unifying the current frame visual image and the previous frame visual image into the IMU coordinate system. Then, the visual image correction module 160 uses the residual information of the visual image to correct the visual image of the binocular fisheye camera, thereby achieving the purpose of eliminating the error in the visual image. Because a binocular fisheye camera is used, the field of view of the binocular fisheye camera can reach 360 degrees, and there is no need to splice multiple monocular cameras, thereby reducing the amount of calculation and the error caused by splicing multiple camera images. In addition, the above scheme can quickly and accurately calculate the residual information of the current frame visual image of the binocular fisheye camera, and use this residual information to verify the visual image of the binocular fisheye camera, thereby eliminating the feature points of position offset, adjusting the object posture in the visual image, and improving the accuracy of the binocular fisheye camera.

[0136] In addition, if Figure 12 As shown, the present invention also provides a visual SLAM system of a binocular fisheye camera, comprising:

[0137] A communication bus 1002, a communication module 1003, a memory 1004, a processor 1001, and a visual SLAM program for a binocular fisheye camera stored in the memory 1004 and running on the processor 1001. When the visual SLAM program is executed by the processor 1001, the steps of the visual SLAM method for a binocular fisheye camera provided in any of the above embodiments are implemented.

[0138] In summary, compared with the prior art, this patent has at least one of the following advantages:

[0139] 1. The binocular camera's field of view is improved, and richer environmental information is obtained. There is no need to add multiple monocular cameras to achieve a large field of view, which improves the stability of the system.

[0140] 2. Utilize multi-sensor fusion technology to closely integrate vision, IMU, and GPS information, which improves stability to a certain extent while also improving the accuracy of pose estimation.

[0141] 3. The algorithm itself is relatively lightweight and fully meets real-time requirements. It can be applied to products such as in-vehicle and handheld devices.

[0142] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0143] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0144] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0146] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claim. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, third etc. does not indicate any order. These words may be interpreted as names.

[0147] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0148] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A visual SLAM method for a binocular fisheye camera, characterized in that: include: Acquire a visual image from a binocular fisheye camera in real time, and extract visual feature points of the visual image; Calculating a moving speed of the visual feature point according to the pixel positions of the visual feature point in the current frame visual image and the previous frame visual image; Calculating the disparity of the visual feature point according to the moving speed of the visual feature point; Extracting residual information of the visual image of the current frame according to a magnitude relationship between the disparity of the visual feature point and a preset disparity threshold; According to the external parameter transformation matrix of the binocular fisheye camera and the IMU, the residual information of the visual image of the current frame is transformed to obtain the residual information of the visual image in the IMU coordinate system; The visual image of the binocular fisheye camera is corrected using the residual information of the visual image.

2. The visual SLAM method according to claim 1, wherein The steps of extracting visual feature points of a visual image include: Extracting a plurality of visual feature points of the current frame visual image; Using the visual feature points of the current frame visual image, tracking the matching visual feature points in the previous frame visual image; Verifying the visual feature points of the previous visual image using the F matrix, and filtering out the visual feature points with incorrect matching in the previous visual image; Dedistortion processing is performed on the successfully matched visual feature points in the previous frame of visual image to obtain the visual feature points with corrected coordinates.

3. Visual SLAM method according to claim 2, characterized in that, After the step of tracking the matching visual feature points in the previous frame of visual image, the method further includes: When the visual feature points of the previous frame visual image fail to match, setting a blank image with the same size as the current frame visual image; Calculating pixel positions of the blank image corresponding to visual feature points where matching fails; In a blank area outside a predetermined distance range around the pixel position in the blank image, visual feature points that match the visual feature points that failed to match are tracked and obtained as visual feature points in the previous frame of visual image.

4. Visual SLAM method according to claim 1, characterized in that, The step of calculating the moving speed of the visual feature point according to the pixel position of the visual feature point in the current frame visual image and the previous frame visual image comprises: Obtaining the interval time between the current frame visual image and the previous frame visual image; Calculating a moving distance of the visual feature point according to a pixel position of the visual feature point of the current frame visual image and a pixel position of the previous frame visual image; The moving speed of the visual feature point is calculated using the moving distance of the visual feature point and the interval time.

5. The visual SLAM method according to claim 1, wherein The step of calculating the parallax of the visual feature point according to the moving speed of the visual feature point includes: Calculating the sum of the coordinate differences of all the visual feature points in two adjacent frames of visual images according to the moving speed of the visual feature points; The average value of the sum of the coordinate differences of all visual feature points is calculated to obtain the disparity of the visual feature points.

6. The visual SLAM method according to claim 1 or 5, wherein The step of extracting residual information of the current frame visual image according to the magnitude relationship between the disparity of the visual feature point and a preset disparity threshold comprises: Determining whether the disparity of the visual feature point is greater than or equal to a preset disparity threshold; If the disparity of the visual feature point is greater than or equal to a preset disparity threshold, then calculating the IMU pre-integrated velocity increment difference between the current frame visual image and the previous frame visual image; Determining whether a second norm of the IMU pre-integrated velocity increment difference is greater than a predetermined increment difference threshold; If the second norm of the IMU pre-integrated velocity increment difference is greater than the predetermined increment difference threshold, the current frame image is set as a key frame, and the residual information of the key frame is extracted using a sliding window.

7. The visual SLAM method according to claim 1, wherein After the step of extracting residual information of the current frame visual image, the method further includes: Initializing the camera parameters of any one of the binocular fisheye cameras; The relative position relationship of the binocular fisheye cameras is used to update the residual information of the other camera in the binocular fisheye cameras to obtain the updated residual information of the current frame visual image.

8. The visual SLAM method according to claim 1, wherein The step of transforming the residual information of the current frame visual image according to the extrinsic parameter transformation matrix of the binocular fisheye camera and the IMU includes: Using the IMU's external parameter transformation matrix to transform the residual information of the visual image of the current frame, to obtain the residual information of the visual image in the IMU coordinate system; The LM optimization algorithm is used to optimize the residual information of the visual image in the IMU coordinate system to obtain the optimized residual information of the visual image.

9. A visual SLAM system using a binocular fisheye camera, characterized in that: The visual SLAM system is used for a binocular fisheye camera, and the visual SLAM system includes: A visual feature extraction module is used to acquire the visual image of the binocular fisheye camera in real time and extract visual feature points of the visual image; A moving speed calculation module, configured to calculate the moving speed of the visual feature point according to the pixel positions of the visual feature point in the current frame visual image and the previous frame visual image; a disparity calculation module, configured to calculate the disparity of the visual feature point according to the moving speed of the visual feature point; A residual information extraction module, configured to extract residual information of the visual image of the current frame based on a magnitude relationship between the disparity of the visual feature points and a preset disparity threshold; A residual information transformation module is used to transform the residual information of the visual image of the current frame according to the external parameter transformation matrix of the binocular fisheye camera and the IMU to obtain the residual information of the visual image in the IMU coordinate system; A visual image correction module is used to correct the visual image of the binocular fisheye camera using residual information of the visual image.

10. A visual SLAM system using a binocular fisheye camera, characterized in that: include: A memory, a processor, and a visual SLAM program for a binocular fisheye camera stored in the memory and running on the processor, wherein when the visual SLAM program is executed by the processor, the steps of the visual SLAM method for a binocular fisheye camera as described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Binocular vision positioning method for target grabbing of underwater robot

    CN111062990A

  • Visual inertia fusion positioning method and system facing indoor SLAM (Simultaneous Localization and Mapping)

    CN115218906A