A method for visual inertial calibration

Through synchronous data acquisition and processing of vision sensors and inertial measurement units, the visual-inertial external parameter calibration matrix is ​​calculated using visual motion feature points and optical flow tracking, which solves the problem of increasing navigation error in the prior art and achieves higher navigation accuracy and robustness.

CN119832089BActive Publication Date: 2025-05-27YULIN SHENHUA ENERGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510299781.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing visual inertial calibration methods rely on a large number of manual adjustments and meticulous operations, resulting in an increase in navigation errors and affecting navigation accuracy.

Method used

Visual motion image data in the road environment is obtained through vision sensors, and inertial data is collected simultaneously through the inertial measurement unit built into the carrier, and the visual-inertial guide external parameter calibration matrix is ​​calculated using motion feature point analysis, optical flow tracking and three-dimensional camera pose conversion.

Benefits of technology

It improves the accuracy and robustness of the positioning and navigation system, reduces navigation errors, and enhances the understanding and tracking of the moving trajectory of the target carrier.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832089B_ABST
    Figure CN119832089B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of inertial navigation calibration, and particularly to a visual inertial navigation calibration method. The method includes the following steps: obtaining visual motion image data corresponding to a target carrier in a road environment through a visual sensor and synchronously collecting corresponding inertial data of the target carrier; analyzing motion feature points of the visual motion image data to obtain a set of visual image feature points; determining the pixel displacement of feature points between adjacent frames of visual images through the set of visual image feature points, performing optical flow tracking of motion trajectories and calculating three-dimensional camera pose conversion to obtain a visual pose conversion matrix between adjacent frames of images in the camera coordinate system; performing adjacent frame pose pre-integration and visual-inertial extrinsic parameter calibration on the inertial data of the target carrier to generate a corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system. The present invention can achieve high-precision calibration of an inertial measurement unit, thereby improving its corresponding navigation accuracy and reliability in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of inertial navigation calibration, and particularly to a visual-inertial calibration method. Background Art

[0002] Visual-Inertial Calibration technology is an important technology that fuses computer vision and inertial measurement unit (IMU) data, and is widely used in fields such as autonomous driving, robot navigation, and augmented reality (AR). By fusing the data of visual sensors and inertial sensors, this technology can improve the accuracy of positioning and attitude estimation, especially in scenarios where GPS signals are missing or the environment is complex. Currently, common visual-inertial calibration methods mainly rely on means such as calibration boards, known scenes, or motion constraints, and use optimization algorithms to solve the internal parameters of the camera, external parameters, and the pose relationship between the IMU and the camera. Traditional inertial navigation calibration methods mainly rely on high-precision turntable equipment. By conducting tests at multiple positions and postures on the turntable and using the known motion information of the turntable to estimate the sensor error parameters, however, due to the large differences in the sensor characteristics of the camera and the IMU, the calibration of the two often requires a large amount of manual adjustment and meticulous operation, resulting in an increasing navigation error and thus seriously affecting the navigation accuracy. Summary of the Invention

[0003] Based on this, it is necessary for the present invention to provide a visual-inertial calibration method to solve at least one of the above technical problems.

[0004] To achieve the above object, a visual-inertial calibration method includes the following steps:

[0005] Step S1: Obtain the visual motion image data corresponding to the target carrier in the road environment through a visual sensor, and synchronously collect the corresponding inertial data of the target carrier through the inertial measurement unit built in the target carrier; analyze the motion feature points of the visual motion image data to obtain a set of visual image feature points;

[0006] Step S2: Determine the pixel displacement of feature points between adjacent frames of visual images through the set of visual image feature points, and perform optical flow tracking on the corresponding adjacent frame feature points within the set of visual image feature points based on the pixel displacement of feature points between adjacent frames of visual images to generate the motion trajectory of feature points between adjacent frames of images;

[0007] Step S3: Calculate the visual motion transformation matrix between adjacent frames of images according to the motion trajectory of feature points between adjacent frames of images, and perform three-dimensional camera pose conversion calculation according to the visual motion transformation matrix to obtain the visual pose conversion matrix between adjacent frames of images in the camera coordinate system;

[0008] Step S4: Perform adjacent-frame pose pre-integration on the inertial data of the target carrier to obtain the inertial pose change matrix between adjacent frame times in the inertial coordinate system; perform visual-inertial extrinsic parameter calibration based on the visual pose conversion matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame times in the inertial coordinate system to generate the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

[0009] Further, step S1 includes the following steps:

[0010] Step S11: Deploy a visual sensor on the target carrier to obtain the visual motion image data corresponding to the target carrier in the road environment;

[0011] Step S12: Send a synchronization pulse through a synchronization signal generator to synchronously collect the corresponding inertial data of the target carrier using the inertial measurement unit built into the target carrier;

[0012] Step S13: Perform histogram equalization processing on each frame of the visual motion image data to obtain the visual motion contrast equalized image data;

[0013] Step S14: Perform pixel blurriness analysis on each frame of the visual motion contrast equalized image data to obtain the pixel blurriness corresponding to each frame of the visual motion image; perform pixel blurring filtering on each frame of the visual motion contrast equalized image data based on the pixel blurriness corresponding to each frame of the visual motion image to filter out the corresponding pixel high-frequency noise and pixel outliers to obtain the visual motion filtered image data;

[0014] Step S15: Input the visual motion filtered image data into a pre-trained convolutional neural network model to extract the corresponding feature key points and their descriptors in each frame of the image, and perform non-maximum suppression on the corresponding feature key points and their descriptors in each frame of the image to obtain the visual image feature point set.

[0015] Further, the non-maximum suppression of the corresponding feature key points and their descriptors in each frame of the image includes:

[0016] Perform in-depth feature analysis on the corresponding feature key points and their descriptors in each frame of the image to analyze the local neighborhood structure corresponding to the feature key points in the visual motion image, and calculate the histogram of gradient directions of the pixels in the neighborhood of the feature key points to characterize the texture direction distribution characteristics of the feature key points. For the descriptors, analyze the numerical distribution law and dimensional correlation in different color spaces to obtain the in-depth feature data set of the feature key points and descriptors;

[0017] Based on the dataset of characteristic key points and descriptor depth characteristics, the corresponding characteristic key points and their descriptors in each frame of the image are initially screened for feature points. Considering the temporal continuity between frames, the characteristic key points and their descriptors in each frame of the image are associated in the time dimension. The position changes corresponding to the characteristic key points of adjacent frames and the similarity corresponding to the descriptors are analyzed, and the characteristic key points with abnormal position changes or the descriptors with descriptor similarity lower than the threshold are marked as objects to be excluded for screening and exclusion, obtaining the set of characteristic key points and descriptors after preliminary screening;

[0018] For each point in the set of characteristic key points and descriptors after preliminary screening, a multi-scale local extreme region is defined to determine the corresponding local extreme region according to the pixel texture condition in its neighborhood, obtaining the multi-scale local extreme region;

[0019] Based on the multi-scale local extreme region, non-maximum suppression is performed on the set of characteristic key points and descriptors after preliminary screening to obtain the set of visual image feature points.

[0020] Further, the multi-scale local extreme region is specifically a 1x1 scale local extreme region set in the texture complex area to capture detailed features, and a 5x5 scale local extreme region is set in the texture smooth area to consider overall features.

[0021] Further, step S2 includes the following steps:

[0022] Step S21: Perform adjacent frame feature point matching on the set of visual image feature points to obtain adjacent frame image feature point matching pairs;

[0023] Step S22: Obtain the motion direction and speed range corresponding to the target carrier, and perform adjacent frame feature displacement constraint analysis according to the motion direction and speed range corresponding to the target carrier to analyze the displacement direction and magnitude constraints corresponding to the feature points in adjacent frames, obtaining adjacent frame feature point displacement constraints;

[0024] Step S23: Based on the adjacent frame feature point displacement constraints, calculate the feature point pixel displacement between each pair of feature points in the adjacent frame image feature point matching pairs to obtain the feature point pixel displacement between adjacent frame visual images;

[0025] Step S24: Based on the feature point pixel displacement between adjacent frame visual images, perform motion trajectory optical flow tracking on the corresponding adjacent frame feature points in the set of visual image feature points to generate the feature point motion trajectory between adjacent frame images.

[0026] Further, step S24 includes the following steps:

[0027] Step S241: Use the pixel displacement of feature points between adjacent frame visual images as a spatial vector to generate a pixel displacement vector corresponding to each feature point within the adjacent frame images.

[0028] Step S242: Decompose the direction components of the pixel displacement vector corresponding to each feature point within the adjacent frame images to obtain the pixel displacement components corresponding to each feature point within the adjacent frame images in the horizontal, vertical, and diagonal directions.

[0029] Step S243: Inverse-solve the displacement angle between each feature point based on the pixel displacement components corresponding to each feature point within the adjacent frame images in the horizontal, vertical, and diagonal directions to obtain the angle between the adjacent frame displacement vectors.

[0030] Step S244: Perform optical flow tracking on the corresponding adjacent frame feature points within the visual image feature point set based on the angle between the adjacent frame displacement vectors and in combination with the pixel displacement of feature points between the adjacent frame visual images to generate the feature point motion trajectories between the adjacent frame images.

[0031] Further, Step S244 includes the following steps:

[0032] Perform displacement change analysis on the pixel displacement of feature points between the adjacent frame visual images in the time dimension to obtain the change values corresponding to the pixel displacement of adjacent frame feature points over time.

[0033] Calculate the displacement time derivative based on the change values corresponding to the pixel displacement of adjacent frame feature points over time to obtain the pixel displacement velocity and pixel displacement acceleration corresponding to the adjacent frame feature points.

[0034] Perform optical flow tracking on the corresponding adjacent frame feature points within the visual image feature point set based on the angle between the adjacent frame displacement vectors, the pixel displacement velocity, and the pixel displacement acceleration corresponding to the adjacent frame feature points, and in combination with the pyramid hierarchical optical flow method to generate the feature point motion trajectories between the adjacent frame images.

[0035] Further, Step S3 includes the following steps:

[0036] Step S31: Project the feature point motion trajectories between the adjacent frame images into a three-dimensional feature space and determine the three-dimensional coordinates corresponding to the feature points of the adjacent frame images within the three-dimensional feature space to obtain the three-dimensional coordinates of each adjacent frame image feature point.

[0037] Step S32: Reproject the three-dimensional coordinates of each adjacent frame image feature point into the image plane coordinate system to obtain the reprojection coordinates of each adjacent frame image feature point.

[0038] Step S33: Obtain the actual observed coordinates corresponding to the feature points of each adjacent frame image on the image plane, and calculate the visual error based on the actual observed coordinates corresponding to the feature points of each adjacent frame image on the image plane and the reprojection coordinates of the feature points of each adjacent frame image, so as to obtain the visual reprojection error vector corresponding to the feature points of each adjacent frame image;

[0039] Step S34: Perform visual motion transformation calculation on the three-dimensional coordinates of the feature points of each adjacent frame image based on the visual reprojection error vector corresponding to the feature points of each adjacent frame image, so as to obtain the visual motion transformation matrix between adjacent frame images;

[0040] Step S35: Perform three-dimensional camera pose conversion calculation according to the visual motion transformation matrix between adjacent frame images, so as to obtain the visual pose transformation matrix between adjacent frame images in the camera coordinate system.

[0041] Further, the three-dimensional coordinates of the feature points of each adjacent frame image described in step S31 are specifically:

[0042] Obtain the camera internal parameter matrix corresponding to the visual sensor , where and are the focal lengths corresponding to the camera in the and directions respectively, and are the coordinates of each point of the visual motion image respectively;

[0043] Let the adjacent frame images be the th frame and the th frame , and obtain the corresponding feature point set in the th frame , and at the same time obtain the matching feature point set in the th frame , where represents the total number of adjacent frame image feature points, represents the th th adjacent frame image feature point corresponding in the th frame, specifically , where represents the abscissa of the th th adjacent frame image feature point in the th frame, represents the ordinate of the th th adjacent frame image feature point in the th frame, Indicates the frame corresponding to the n-th adjacent frame image feature point, specifically , where Indicates the frame corresponding to the n-th adjacent frame image feature point's abscissa, Indicates the frame corresponding to the n-th adjacent frame image feature point's ordinate, Indicates the transpose symbol;

[0044] Based on the camera internal parameter matrix for the frame feature point set or the frame feature point set perform normalized coordinate transformation to obtain the frame normalized coordinates or the frame normalized coordinates ;

[0045] Perform three-dimensional feature space mapping on the frame normalized coordinates or the frame normalized coordinates to obtain the three-dimensional coordinates of each adjacent frame image feature point , where the three-dimensional coordinates of the n-th adjacent frame image feature point .

[0046] Furthermore, the reprojection coordinates of each adjacent frame image feature point described in step S32 are specifically:

[0047] Assume that the external parameters corresponding to the camera in the frame are the rotation matrix and the translation vector respectively, and convert it from the three-dimensional coordinate system to the camera coordinate system according to the corresponding external parameters. The coordinates corresponding to the frame in the camera coordinate system are specifically ;

[0048] Project it onto the image plane according to the projection formula corresponding to the camera. The projection formula is specifically:

[0049] ;

[0050] Among them, Indicates the The reprojection coordinates of the feature points of adjacent frame images on the image plane of the frame image, representing the scale factor;

[0051] The corresponding reprojection coordinates are obtained by solving the above projection formula , , and all the reprojection coordinates are aggregated to obtain the reprojection coordinates of the feature points of each adjacent frame image .

[0052] Furthermore, the visual reprojection error vector described in step S33 is specifically , where represents the total number of feature points of adjacent frame images, represents the transpose symbol, represents the th visual reprojection error value between adjacent frame image feature points, represents the th actual observed coordinate corresponding to the adjacent frame image feature point on the image plane , represents the th reprojection coordinate of the adjacent frame image feature point .

[0053] Furthermore, step S35 includes the following steps:

[0054] Perform singular value decomposition on the visual motion transformation matrix between adjacent frame images to obtain the rotation matrix and translation vector between adjacent frame image feature points;

[0055] Perform three-dimensional camera pose transformation calculation based on the rotation matrix and translation vector between adjacent frame image feature points to obtain the visual pose transformation matrix between adjacent frame images in the camera coordinate system.

[0056] Furthermore, the visual pose transformation matrix is specifically:

[0057] ;

[0058] where represents the visual pose transformation matrix between the th and the th frame images in the camera coordinate system, represents the rotation matrix between adjacent frame image feature points, represents the translation vector between adjacent frame image feature points.

[0059] Furthermore, step S4 includes the following steps:

[0060] Step S41: Perform adjacent-frame pose pre-integration on the inertial data of the target carrier to obtain the inertial pose change matrix between adjacent-frame times in the inertial coordinate system;

[0061] Step S42: Perform visual-inertial extrinsic parameter calibration based on the visual pose conversion matrix between adjacent-frame images in the camera coordinate system and the inertial pose change matrix between adjacent-frame times in the inertial coordinate system to generate the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

[0062] Further, the inertial pose change matrix described in Step S41 is specifically:

[0063] ;

[0064] where represents the inertial pose change matrix between the th frame and the th frame in the inertial coordinate system, represents the rotational pre-integration quantity between the th frame and the th frame, represents the position pre-integration quantity between the th frame and the th frame.

[0065] Advantages of the present invention:

[0066] The visual inertial calibration method proposed by the present invention, compared with the prior art, the beneficial effects of the present application are as follows: visual motion image data in the road environment is obtained through a visual sensor, and inertial data is synchronously collected through an inertial measurement unit (IMU) built in the vehicle. The image data provided by the visual sensor contains static and dynamic information of the environment, while the IMU can provide dynamic information such as acceleration and angular velocity. By synchronously collecting these two types of data, the motion state of the target vehicle in a complex dynamic environment can be effectively captured, enhancing the understanding of the motion trajectory of the target vehicle. The image data obtained by the visual sensor provides a necessary basis for subsequent analysis of motion feature points. By processing the image sequence, stable and representative visual feature points are extracted. These feature points can help us better capture the motion changes between adjacent frames. The acceleration and angular velocity information provided by the IMU data can help us estimate the motion state of the vehicle, filling the blind spots existing in the visual sensor in specific environments (for example, during dynamic motion or rapid rotation, the visual sensor will produce blurring or distortion). By synchronously collecting these two types of data, the accuracy and robustness of the positioning and navigation system can be effectively improved. Secondly, the pixel displacement of feature points between adjacent frames is determined through the visual image feature point set. This process involves the extraction and matching of visual image feature points. By analyzing the changes between the feature points of adjacent image frames, the pixel displacement amount is calculated, thereby understanding the motion pattern of the image. Through the optical flow tracking algorithm, the motion trajectories of feature points in adjacent frames are tracked and inferred, and the motion trajectory of the target vehicle can be further accurately understood. Through this process, the motion trajectory of the feature points, that is, the optical flow, can be constructed. Because the optical flow of feature points can not only reflect the translational motion of the target vehicle, but also indirectly reflect the rotation information of the vehicle. Compared with other traditional methods, the optical flow method has higher time resolution and real-time performance, and is suitable for scenes with high-speed motion or large dynamic changes. Through accurate feature point tracking, the interference caused by image noise or blurring can also be effectively reduced, thereby improving the accuracy and stability of motion estimation. Then, a visual motion transformation matrix is calculated through the motion trajectories of feature points between adjacent frame images. This matrix can describe the geometric transformation relationship between adjacent frames, that is, the visual motion change from the previous frame image to the current frame image. Based on this transformation matrix, further conversion calculations of the three-dimensional camera pose can be performed. Through this step, the camera motion between adjacent frames can be accurately estimated, and then the spatial motion trajectory of the target vehicle can be deduced. The visual motion transformation matrix reflects the position and direction changes in the three-dimensional space calculated from the visual data, so that the motion trajectory of the target vehicle can be continuously tracked in the global coordinate system, thereby providing accurate visual information for subsequent visual inertial fusion.Finally, through the pre-integration of the inertial data of the target carrier over adjacent frames, the pre-integration technology of inertial data can help estimate the motion trajectory of the carrier based on the data from accelerometers and gyroscopes without relying on external signals. Different from the processing of visual data, inertial data is usually not affected by problems such as environmental light changes and motion blur. Therefore, they can provide more stable motion estimation in certain special cases. By combining the inertial pose change matrix with the visual pose transformation matrix, visual-inertial extrinsic calibration can be performed. Visual-inertial extrinsic calibration is to register the coordinate systems of the visual sensor and the inertial measurement unit so that they can work in harmony with each other. Through this calibration, the advantages of both visual data and inertial data can be utilized simultaneously to improve the accuracy and reliability of motion estimation. The calibrated extrinsic matrix can perform precise coordinate transformation in practical applications, enabling more accurate positioning and navigation, and thus reducing the navigation error corresponding to the inertial measurement unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0068] Figure 1 is a schematic flow chart of the steps of the visual-inertial calibration method of the present invention;

[0069] Figure 2 is Figure 1 a detailed schematic flow chart of step S1 in

[0070] Figure 3 is Figure 1 a detailed schematic flow chart of step S2 in DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0072] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0073] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0074] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a visual-inertial calibration method, and the method includes the following steps:

[0075] Step S1: Obtain the visual motion image data corresponding to the target carrier in the road environment through a visual sensor, and synchronously collect the corresponding inertial data of the target carrier through the inertial measurement unit built in the target carrier; perform motion feature point analysis on the visual motion image data to obtain a set of visual image feature points;

[0076] Step S2: Determine the feature point pixel displacement between adjacent frame visual images through the set of visual image feature points, and perform optical flow tracking on the corresponding adjacent frame feature points within the set of visual image feature points based on the feature point pixel displacement between adjacent frame visual images to generate the feature point motion trajectory between adjacent frame images;

[0077] Step S3: Calculate the visual motion transformation matrix between adjacent frame images according to the feature point motion trajectory between adjacent frame images, and perform three-dimensional camera pose conversion calculation according to the visual motion transformation matrix to obtain the visual pose conversion matrix between adjacent frame images in the camera coordinate system;

[0078] Step S4: Perform adjacent frame pose pre-integration on the inertial data of the target carrier to obtain the inertial pose change matrix between adjacent frame times in the inertial coordinate system; perform visual-inertial extrinsic parameter calibration according to the visual pose conversion matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame times in the inertial coordinate system to generate the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

[0079] In the embodiments of the present invention, please refer to Figure 1 shown, which is a schematic flow chart of the steps of the visual-inertial calibration method of the present invention. In this example, the visual-inertial calibration method includes the following steps:

[0080] Step S1: Obtain the visual motion image data corresponding to the target carrier in the road environment through a visual sensor, and synchronously collect the corresponding inertial data of the target carrier through the inertial measurement unit built in the target carrier; perform motion feature point analysis on the visual motion image data to obtain a visual image feature point set;

[0081] In an embodiment of the present invention, by firmly installing a high-resolution and high-frame-rate visual sensor, such as an industrial-grade camera, on the target carrier, so that its lens can cover a certain range of the road environment around the target carrier. When the target carrier runs on the road, the visual sensor continuously captures images at a speed of 30 frames per second, converts the captured images into digital signals and stores them, thereby obtaining visual motion image data. At the same time, the inertial measurement unit (IMU) built in the target carrier is connected to the visual sensor through a synchronous signal line, and a synchronous pulse signal is sent by a synchronous signal generator to make the IMU and the visual sensor work synchronously, and the acceleration, angular velocity and other inertial data of the target carrier are collected and stored in real time. For the obtained visual motion image data, the Scale-Invariant Feature Transform (SIFT) algorithm is used to process each frame of the image. The SIFT algorithm first detects the extreme points in the image, then screens and locates these extreme points to determine them as feature points, and calculates the descriptors of each feature point, and finally obtains a visual image feature point set containing a large number of feature points and their descriptors.

[0082] Step S2: Determine the pixel displacement of the feature points between adjacent frames of visual images through the visual image feature point set, and perform optical flow tracking on the corresponding adjacent frame feature points in the visual image feature point set based on the pixel displacement of the feature points between adjacent frames of visual images to generate the feature point motion trajectory between adjacent frames of images;

[0083] In an embodiment of the present invention, by using a feature point matching algorithm, such as a matching method based on the Hamming distance of descriptors, to match the feature points of two adjacent frames of images in the visual image feature point set. For each feature point in the first frame of the image, calculate the Hamming distance between its descriptor and the descriptors of all feature points in the second frame of the image, and determine the feature point with the smallest distance and less than a specific threshold (such as 60) as the matching point, so as to obtain the feature point matching pairs between adjacent frames of visual images. According to the pixel coordinates of the feature points in these matching pairs in the two frames of images, calculate the pixel displacement of the feature points. For example, if the coordinate of a certain feature point in the first frame is , and the pixel coordinate of the feature point matching it in the second frame of the image is , then the pixel displacement of this pair of feature points is , the Lucas-Kanade optical flow algorithm is adopted. Using these pixel displacements as the initial conditions, optical flow tracking of feature points is performed between adjacent frame images. During the tracking process, a search window is defined centered on the position of the feature point in the current frame. The best matching position is found in the subsequent frame through iterative calculation, and the positions of the feature points in adjacent frames are connected in sequence, finally generating the motion trajectories of the feature points between adjacent frame images.

[0084] Step S3: Calculate the visual motion transformation matrix between adjacent frame images based on the motion trajectories of the feature points between adjacent frame images, and perform three-dimensional camera pose transformation calculation according to the visual motion transformation matrix to obtain the visual pose transformation matrix between adjacent frame images in the camera coordinate system;

[0085] In the embodiment of the present invention, by projecting the motion trajectories of the feature points between adjacent frame images into a three-dimensional feature space, and determining the three-dimensional coordinates corresponding to the feature points of adjacent frame images in the three-dimensional feature space , and re-projecting the three-dimensional coordinates of the feature points of each adjacent frame image into the image plane coordinate system to obtain the re-projection coordinates of the feature points of each adjacent frame image , by obtaining the actual observed coordinates corresponding to the feature points of each adjacent frame image on the image plane , and comparing them with the re-projection coordinates of the feature points of adjacent frame images to perform error metric calculation to calculate the visual re-projection error vector, which is specifically , at the same time, by constructing and combining the previously obtained th frame th visual re-projection error vector corresponding to the image feature points to construct the corresponding error function , and by using a non-linear optimization algorithm (such as the Levenberg-Marquardt algorithm) to minimize the error function , during the optimization process, continuously adjust and values to make the value of the error function gradually decrease. Through iterative update , where is the Jacobian matrix of the error function E with respect to and , is the damping factor, is the vector composed of the visual re-projection error vectors of all feature points. After multiple iterations, when the error function converges to the minimum value, the visual motion transformation matrix between adjacent frame images is obtained , then, by performing singular value decomposition on the visual motion transformation matrix between adjacent frame images, the rotation matrix between the feature points of adjacent frame images is obtained , where and represents an orthogonal matrix, represents a diagonal matrix. Since the rotation matrix is an orthogonal matrix, take , as the rotation matrix between the feature points of adjacent frame images, and the translation vector is directly obtained from the visual motion transformation matrix. The pose matrix of the i-th frame camera is calculated by conversion in the camera coordinate system as , then the visual pose transformation matrix between adjacent frame images in the camera coordinate system is specifically , where, represents the visual pose transformation matrix between the -th frame and the -th frame images in the camera coordinate system, represents the rotation matrix between the feature points of adjacent frame images, represents the translation vector between the feature points of adjacent frame images.

[0086] Step S34: Perform visual motion transformation calculation on the three-dimensional coordinates of each adjacent frame image feature point based on the corresponding visual reprojection error vector of each adjacent frame image feature point to obtain the visual motion transformation matrix between adjacent frame images;

[0087] Step S35: Perform three-dimensional camera pose transformation calculation according to the visual motion transformation matrix between adjacent frame images to obtain the visual pose transformation matrix between adjacent frame images in the camera coordinate system.

[0088] Step S4: Perform adjacent frame pose pre-integration on the inertial data of the target carrier to obtain the inertial pose change matrix between adjacent frame times in the inertial coordinate system; perform visual-inertial extrinsic parameter calibration according to the visual pose transformation matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame times in the inertial coordinate system to generate the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

[0089] In the embodiment of the present invention, by acquiring the inertial data corresponding to the target carrier, which is provided by an inertial measurement unit (IMU) and includes the angular velocity and linear acceleration between adjacent frame times. For the angular velocity , at each sampling point , its measured value is . Considering the zero bias of the gyroscope, the true angular velocity is . And for the linear acceleration , at each sampling point , its measured value is . Considering the zero bias of the accelerometer and the gravitational acceleration , the true linear acceleration in the inertial coordinate system is , where is the rotation matrix from the vehicle coordinate system to the inertial coordinate system at time . By integrating the true angular velocity, the rotation pre-integration quantity and the position pre-integration quantity between adjacent frame times can be obtained. The specific calculation uses the integration method of quaternions or rotation matrices. For example, the midpoint integration method is used: , where is 's skew-symmetric matrix, is the sampling time interval. By double-integrating the linear acceleration, the position pre-integration quantity between adjacent frame times can be obtained. The corresponding calculation formula is specifically . Combine the rotation pre-integration quantity and the position pre-integration quantity into the inertial pose change matrix between adjacent frame times in the inertial coordinate system. At the same time, find the transformation matrix from the camera coordinate system to the inertial coordinate system by combining the visual pose transformation matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame times in the inertial coordinate system calculated previously, so that the visual and inertial data can be aligned in the same coordinate system. According to the coordinate transformation relationship, there is . Expand the above equation to obtain an equation about and . For the rotation part, there is ; for the translation part, there is . In order to solve and , collect multiple sets of and data of adjacent frames, construct a non-linear least squares problem, define the error function , and use a non-linear optimization algorithm (such as the Levenberg-Marquardt algorithm) to minimize the error function . During the optimization process, continuously adjust and values so that the value of the error function gradually decreases. When the error function converges to the minimum value, the obtained and are combined into the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

[0090] Furthermore, step S1 includes the following steps:

[0091] Step S11: Deploy a visual sensor on the target vehicle to obtain the corresponding visual motion image data of the target vehicle in the road environment;

[0092] Step S12: Send a synchronization pulse through a synchronization signal generator to synchronously collect corresponding target carrier inertial data using the inertial measurement unit built in the target carrier;

[0093] Step S13: Perform histogram equalization processing on each frame of the visual motion image data to obtain visual motion contrast equalized image data;

[0094] Step S14: Analyze the pixel blurriness of each frame of the visual motion contrast equalized image data to obtain the pixel blurriness corresponding to each frame of the visual motion image; Based on the pixel blurriness corresponding to each frame of the visual motion image, perform pixel blur filtering on each frame of the visual motion contrast equalized image data to filter out the corresponding pixel high-frequency noise and pixel outliers, and obtain visual motion filtered image data;

[0095] Step S15: Input the visual motion filtered image data into a pre-trained convolutional neural network model to extract the corresponding feature key points and their descriptors in each frame of the image, and perform non-maximum suppression on the corresponding feature key points and their descriptors in each frame of the image to obtain a visual image feature point set.

[0096] As an embodiment of the present invention, refer to Figure 2 shown, for Figure 1 the detailed step flow schematic diagram of step S1 in

[0097] Step S11: Deploy a visual sensor on the target carrier to obtain corresponding visual motion image data of the target carrier in the road environment;

[0098] In the embodiment of the present invention, a visual sensor, such as a high-definition camera, is firmly installed at a suitable position on the target carrier. The visual sensor has the characteristics of high resolution and high frame rate, and can perform image acquisition at a speed of 30 frames per second. During the driving process of the target carrier in the road environment, the visual sensor continuously shoots the surrounding environment. The viewing angle of the camera lens can cover a certain range of road scenes in front of and on both sides of the target carrier, and transmits the captured image data to the data storage device in the form of a digital signal, and finally forms the corresponding visual motion image data of the target carrier in the road environment.

[0099] Step S12: Send a synchronization pulse through a synchronization signal generator to synchronously collect corresponding target carrier inertial data using the inertial measurement unit built in the target carrier;

[0100] In an embodiment of the present invention, by using a high-precision pulse generation circuit in a synchronization signal generator, a stable and accurate synchronization pulse signal can be generated. When the target carrier starts to operate, the synchronization signal generator is activated and sends synchronization pulses at a fixed time interval (such as every 10 milliseconds). The inertial measurement unit (IMU) built into the target carrier is connected to the synchronization signal generator. After receiving the synchronization pulse signal, it immediately collects inertial data such as the acceleration and angular velocity of the target carrier. The accelerometer and gyroscope inside the IMU can accurately measure the motion state of the target carrier in three-dimensional space. The collected inertial data is stored in the data storage module to ensure synchronization with the acquisition time of the visual motion image data, and finally, the corresponding inertial data of the target carrier is obtained.

[0101] Step S13: Perform histogram equalization processing on each frame of the visual motion image data to obtain visual motion contrast equalized image data;

[0102] In an embodiment of the present invention, for each frame of the visual motion image data, the histogram equalization algorithm is used for processing. First, calculate the gray-level histogram of the image and count the frequency of each gray level in the image. Then, calculate the cumulative distribution function (CDF) according to the histogram and perform normalization processing on the CDF. Next, map the gray value of each pixel in the image according to the normalized CDF, and replace the original gray value with the mapped gray value. In this way, the contrast of the image is enhanced, making the bright parts in the image brighter and the dark parts darker, and finally, the visual motion contrast equalized image data is obtained, which is convenient for subsequent analysis and processing of the image.

[0103] Step S14: Analyze the pixel blurriness of each frame of the visual motion contrast equalized image data to obtain the pixel blurriness corresponding to each frame of the visual motion image; perform pixel blur filtering on each frame of the visual motion contrast equalized image data based on the pixel blurriness corresponding to each frame of the visual motion image to filter out the corresponding pixel high-frequency noise and pixel outliers, and obtain visual motion filtered image data;

[0104] In an embodiment of the present invention, a Laplace operator is used to perform pixel blur analysis on each frame of the visual motion contrast equalization image data to quantify the Laplace value of each pixel in the image. The larger the Laplace value, the more drastic the local change of the pixel, that is, the lower the blur. By performing a statistical analysis on the Laplace value of the entire frame of the image, the pixel blur corresponding to the frame of the image is obtained. Based on the pixel blur, a Gaussian filtering algorithm is used to perform pixel blur filtering on the image. The Gaussian filtering performs a convolution operation on the image through a two-dimensional Gaussian kernel function, and adjusts the standard deviation of the Gaussian kernel according to the pixel blur, so that the filtering process can adaptively remove pixel high-frequency noise and pixel outliers, and finally obtain visual motion filtered image data.

[0105] Step S15: Input the visual motion filtered image data into the pre-trained convolutional neural network model to extract the corresponding feature key points and their descriptors in each frame image, and perform non-maximum suppression on the corresponding feature key points and their descriptors in each frame image to obtain a visual image feature point set.

[0106] In an embodiment of the present invention, a pre-trained convolutional neural network model is used, which is a model trained on a large-scale image data set, such as a model based on a ResNet architecture, and the visual motion filter image data is input into the convolutional neural network model frame by frame. The model extracts features of the image through structures such as a convolution layer and a pooling layer. In the convolution layer, different convolution kernels are used to extract local features of the image. After multiple convolution and pooling operations, a high-level feature representation of the image is obtained. Representative feature key points are found from these high-level features, and corresponding descriptors are generated. The descriptors are used to describe the feature information of the feature key points. Then, non-maximum suppression is performed on the corresponding feature key points and their descriptors in each frame of the image. For each feature key point, its response value is compared with the response values ​​of other feature key points in the surrounding neighborhood. If the response value of the feature key point is not a local maximum, it is suppressed, and finally a set of visual image feature points is obtained.

[0107] Furthermore, the non-maximum suppression of the corresponding feature key points and their descriptors in each frame of the image includes:

[0108] The deep feature analysis is performed on the corresponding key points and their descriptors in each frame of the image, so as to analyze the local neighborhood structure corresponding to the key points in the visual motion image, and the texture direction distribution characteristics of the key points are characterized by calculating the gradient direction histogram of the pixels in the neighborhood of the key points. For the descriptors, their numerical distribution laws and dimensional correlations in different color spaces are analyzed to obtain the deep feature data set of the key points and descriptors.

[0109] In an embodiment of the present invention, for each feature key point in each frame of image, a local neighborhood with a fixed size is defined centered thereon, such as a region of 16×16 pixels. Within this local neighborhood, the gradient direction and magnitude of each pixel are calculated, the gradient direction is divided into several intervals, and the sum of the gradient magnitudes within each interval is statistically calculated, so as to obtain the gradient direction histogram of the pixels within the neighborhood of the feature key point, thereby characterizing its texture direction distribution characteristics. For the descriptor, it is respectively converted into different color spaces such as RGB and HSV, and the numerical distribution of the descriptor in different color spaces is statistically calculated, such as the mean value, variance, etc. At the same time, the correlation coefficient between each dimension of the descriptor is calculated, such as the Pearson correlation coefficient, and finally these information are integrated to form a dataset of the depth characteristics of the feature key point and the descriptor.

[0110] Preferably, based on the dataset of the depth characteristics of the feature key point and the descriptor, the feature key points and their descriptors corresponding to each frame of image are preliminarily screened. Considering the temporal continuity between frames, the feature key points and their descriptors in each frame of image are associated in the time dimension, the position changes corresponding to the feature key points in adjacent frames and the similarity of the corresponding descriptors are analyzed therefrom, and the feature key points with abnormal position changes or the descriptors with descriptor similarity lower than the threshold are marked as objects to be excluded for screening and exclusion, so as to obtain the set of feature key points and descriptors after preliminary screening;

[0111] In an embodiment of the present invention, by relying on the dataset of the depth characteristics of the feature key point and the descriptor, the feature key points and their descriptors in adjacent frames of images are associated. For the feature key points in adjacent frames, the position offset amount thereof on the image plane is calculated, and a reasonable position change threshold is set, such as 20 pixels. If the position change of a certain feature key point between adjacent frames exceeds this threshold, it is marked as an object to be excluded. For the descriptor, the cosine similarity algorithm is used to calculate the similarity between the descriptors in adjacent frames, and the similarity threshold is set to 0.7. If the descriptor similarity is lower than this threshold, the corresponding descriptor is marked as an object to be excluded. Finally, all the marked objects to be excluded are removed from the original set of feature key points and descriptors, and finally the set of feature key points and descriptors after preliminary screening is obtained.

[0112] Preferably, a multi-scale local extreme value region is defined for each point in the set of feature key points and descriptors after preliminary screening, so as to determine the corresponding local extreme value region according to the pixel texture condition within its neighborhood, and a multi-scale local extreme value region is obtained;

[0113] In an embodiment of the present invention, for each point in the set of feature key points and descriptors after primary screening, neighborhoods of different scales are constructed, such as 1x1, 5x5, etc. Within the neighborhood of each scale, the texture condition of the pixels is analyzed. For example, the complexity corresponding to the pixel texture is calculated. If the complexity corresponding to a certain point is greater than the threshold, it is determined as a texture complex region, and a 1x1 scale local extreme value region is set within its region to capture detailed features, while a 5x5 scale local extreme value region is set in the texture smooth region (i.e., when it is less than or equal to the threshold) to consider overall features. Such operations are performed on each point at different scales, and finally, a multi-scale local extreme value region corresponding to each point is obtained.

[0114] Preferably, non-maximum suppression is performed on the set of feature key points and descriptors after primary screening based on the multi-scale local extreme value region to obtain a set of visual image feature points.

[0115] In an embodiment of the present invention, for each point in the set of feature key points and descriptors after primary screening, non-maximum suppression operations are performed within its multi-scale local extreme value region. Within each local extreme value region, the response value (such as feature intensity) of this point is compared with the response values of other points in the neighborhood. If the response value of this point is not the local maximum, it is suppressed, that is, removed from the set. By performing such comparisons and screenings on all points within their multi-scale local extreme value regions, the set of remaining feature key points and descriptors is finally obtained, and this set is the set of visual image feature points.

[0116] Further, the multi-scale local extreme value region is specifically a 1x1 scale local extreme value region set in the texture complex region to capture detailed features, while a 5x5 scale local extreme value region is set in the texture smooth region to consider overall features.

[0117] Further, step S2 includes the following steps:

[0118] Step S21: Perform adjacent frame feature point matching on the set of visual image feature points to obtain adjacent frame image feature point matching pairs;

[0119] Step S22: Obtain the motion direction and speed range corresponding to the target carrier, and perform adjacent frame feature displacement constraint analysis according to the motion direction and speed range corresponding to the target carrier to analyze the displacement direction and magnitude constraints corresponding to the feature points in adjacent frames, and obtain adjacent frame feature point displacement constraints;

[0120] Step S23: Perform feature point pixel displacement calculation between each pair of feature points in the adjacent frame image feature point matching pair based on the adjacent frame feature point displacement constraint to obtain the feature point pixel displacement between adjacent frame visual images;

[0121] Step S24: Perform optical flow tracking on the corresponding adjacent-frame feature points within the visual image feature point set based on the pixel displacement of feature points between adjacent-frame visual images to generate the feature point motion trajectories between adjacent-frame images.

[0122] As an embodiment of the present invention, referring to Figure 3 shown in Figure 1 is the detailed step flow diagram of step S2 in

[0123] Step S21: Perform adjacent-frame feature point matching on the visual image feature point set to obtain adjacent-frame image feature point matching pairs;

[0124] In the embodiment of the present invention, by adopting a feature descriptor matching algorithm, such as a matching algorithm based on Hamming distance, perform adjacent-frame feature point matching on the visual image feature point set. For the feature points in two adjacent frames of images, extract their descriptors respectively, calculate the Hamming distance between the descriptor of each feature point in the first frame image and the descriptors of all feature points in the second frame image. If a certain feature point has a unique matching point in the second frame image with the smallest Hamming distance and less than a set threshold (such as 50), then determine these two feature points as a matching pair. Perform such matching operations on all adjacent-frame images within the visual image feature point set, and finally obtain a set of adjacent-frame image feature point matching pairs.

[0125] Step S22: Obtain the motion direction and speed range corresponding to the target carrier, and perform adjacent-frame feature displacement constraint analysis according to the motion direction and speed range corresponding to the target carrier to analyze the displacement direction and magnitude constraints corresponding to the feature points in adjacent frames, and obtain adjacent-frame feature point displacement constraints;

[0126] In the embodiment of the present invention, obtain the motion direction and speed range of the target carrier through the positioning system and speed sensor built in the target carrier. Assume that the motion direction of the target carrier is forward and the speed range is 10 - 30 m / s. According to the installation position and parameters of the camera, as well as the motion information of the target carrier, calculate the possible displacement direction and magnitude range of the feature points in adjacent frames. For example, if the camera is installed in front of the target carrier and the target carrier moves forward, then the displacement direction of the feature points in adjacent frames is most likely forward, and the displacement magnitude is related to the speed of the target carrier. By establishing a mathematical model, combining the internal and external parameters of the camera and the motion information of the target carrier, obtain the displacement direction and magnitude constraints corresponding to the feature points in adjacent frames, and finally form adjacent-frame feature point displacement constraints.

[0127] Step S23: Perform feature point pixel displacement calculation between each pair of feature points within the adjacent-frame image feature point matching pairs based on the adjacent-frame feature point displacement constraints to obtain the feature point pixel displacement between adjacent-frame visual images;

[0128] In an embodiment of the present invention, for each pair of feature points in the adjacent frame image feature point matching pair, first check whether its displacement conforms to the adjacent frame feature point displacement constraint. If it conforms, calculate the pixel coordinate difference of this pair of feature points on the image plane. Assume that the pixel coordinate of a certain feature point in the first frame image is , and the pixel coordinate of the feature point matching it in the second frame image is , then the pixel displacement of this pair of feature points is . Perform such calculations for all adjacent frame image feature point matching pairs. If the displacement of a certain pair of feature points does not conform to the displacement constraint, exclude it. Finally, obtain the set of pixel displacements of feature points that meet the constraint conditions between adjacent frame visual images.

[0129] Step S24: Perform optical flow tracking on the corresponding adjacent frame feature points in the visual image feature point set based on the pixel displacements of feature points between adjacent frame visual images to generate the feature point motion trajectories between adjacent frame images.

[0130] In an embodiment of the present invention, perform optical flow tracking on the corresponding adjacent frame feature points in the visual image feature point set by using the Lucas-Kanade optical flow algorithm. This algorithm is based on the assumption that the image gray value remains unchanged in a short period of time, and estimates the optical flow (i.e., motion speed) of feature points by solving a least squares problem. According to the pixel displacements of feature points between adjacent frame visual images, use them as the initial estimate and input them into the Lucas-Kanade optical flow algorithm. For each feature point, in the adjacent frame images, draw a search window centered on its position in the previous frame, and find the exact position of this feature point in the current frame through iterative calculation. Connect the positions of feature points in adjacent frames in sequence, and then generate the feature point motion trajectories between adjacent frame images. Perform such optical flow tracking operations on all feature points in the visual image feature point set, and finally obtain the complete feature point motion trajectories between adjacent frame images.

[0131] Further, step S24 includes the following steps:

[0132] Step S241: Use the pixel displacements of feature points between adjacent frame visual images as spatial vectors to generate pixel displacement vectors corresponding to each feature point in the adjacent frame images;

[0133] In an embodiment of the present invention, by using the pixel displacements of feature points between adjacent frame visual images It is regarded as a two-dimensional spatial vector, that is, a pixel displacement vector corresponding to the feature point is generated. Such an operation is performed on each feature point in adjacent frame images. For example, if there are 100 feature points in adjacent frame images, the pixel displacements of these 100 feature points are calculated respectively and converted into vector form. Finally, pixel displacement vectors corresponding to each feature point in adjacent frame images are obtained.

[0134] Step S242: Decompose the direction components of the pixel displacement vector corresponding to each feature point in adjacent frame images to obtain the pixel displacement components corresponding to each feature point in the horizontal, vertical, and diagonal directions in adjacent frame images;

[0135] In the embodiment of the present invention, for the pixel displacement vector corresponding to each feature point in adjacent frame images, a vector decomposition method is adopted for processing. Taking the horizontal direction (x-axis), vertical direction (y-axis), and diagonal direction (taking 45° and 135° with the x-axis as the reference directions) as the benchmark directions for the pixel displacement vector , the pixel displacement component in the horizontal direction is , the pixel displacement component in the vertical direction is , for the diagonal direction, according to the trigonometric function relationship, the displacement component in the 45° direction is calculated as , and the displacement component in the 135° direction is . Such a direction component decomposition operation is performed on the pixel displacement vector of each feature point in adjacent frame images. Finally, the pixel displacement components corresponding to each feature point in different directions are obtained.

[0136] Step S243: Based on the pixel displacement components corresponding to each feature point in the horizontal, vertical, and diagonal directions in adjacent frame images, perform an inverse solution of the displacement angle between each feature point to obtain the angle between adjacent frame displacement vectors;

[0137] In the embodiment of the present invention, for any two feature points in adjacent frame images, according to the pixel displacement components corresponding to them in the horizontal, vertical, and diagonal directions, the angle between their displacement vectors is calculated using the vector dot product formula. Let the displacement vectors of the two feature points be and , and their components in the horizontal, vertical, and diagonal directions are and , and according to the vector dot product formula (where is the angle between the two vectors), then, by first calculating , then calculating and , and finally through Inverse solve the angle between the two vectors, and perform such calculations for every two feature points within adjacent frame images, finally obtaining the angle between adjacent frame displacement vectors.

[0138] Step S244: Based on the angle between adjacent frame displacement vectors and in combination with the pixel displacement of feature points between adjacent frame visual images, perform optical flow tracking on corresponding adjacent frame feature points within the visual image feature point set to generate the feature point motion trajectories between adjacent frame images.

[0139] In the embodiment of the present invention, by adopting the Lucas-Kanade optical flow algorithm, in combination with the angle between adjacent frame displacement vectors and the pixel displacement of feature points between adjacent frame visual images, perform optical flow tracking on corresponding adjacent frame feature points within the visual image feature point set. For each feature point, use its pixel displacement in adjacent frames as the initial estimate value, and at the same time consider the angle information between adjacent frame displacement vectors to optimize the optical flow calculation. In the Lucas-Kanade optical flow algorithm, by demarcating a search window centered on the position of the previous frame of the feature point in adjacent frame images, adjust the search direction and range using pixel displacement and angle information, and iteratively calculate to find the exact position of the feature point in the current frame. Connect the positions of the feature points in adjacent frames in sequence to generate the feature point motion trajectories between adjacent frame images. Perform such optical flow tracking operations on all feature points within the visual image feature point set, and finally obtain the complete feature point motion trajectories between adjacent frame images.

[0140] Further, step S244 includes the following steps:

[0141] Perform displacement change analysis on the pixel displacement of feature points between adjacent frame visual images in the time dimension to obtain the corresponding change value of adjacent frame feature point pixel displacement over time;

[0142] In the embodiment of the present invention, through the known pixel displacement of feature points calculated between adjacent frame visual images, and each frame image has a corresponding acquisition time. For a certain feature point in adjacent frames, if the pixel coordinates in the nth frame image are , and the pixel coordinates in the (n + 1)th frame image are , then the pixel displacement of this feature point between these two frames is , let the acquisition time of the nth frame image be , the acquisition time of the (n + 1)th frame image be , the time interval , compare the pixel displacements of this feature point in multiple groups of adjacent frame images (such as from n to n + 1, from n + 1 to n + 2, etc.) to calculate the difference in adjacent frame feature point pixel displacement, such as , so as to obtain the corresponding change value of the pixel displacement of the feature point over time. Such an analysis operation is performed on each feature point in adjacent frame images, and finally the corresponding change value of the pixel displacement of the feature points in adjacent frames over time is obtained.

[0143] Preferably, according to the corresponding change value of the pixel displacement of the feature points in adjacent frames over time, a displacement time derivative calculation is performed to obtain the pixel displacement velocity and pixel displacement acceleration corresponding to the feature points in adjacent frames;

[0144] In the embodiment of the present invention, for the corresponding change value of the pixel displacement of the feature points in adjacent frames over time, the definition of the derivative is used to calculate the pixel displacement velocity and pixel displacement acceleration. For a certain feature point in the horizontal direction, the change value of the pixel displacement in the horizontal direction between adjacent frames is known and the time interval , then the pixel displacement velocity in the horizontal direction . Similarly, the pixel displacement velocity in the vertical direction can be obtained . For the calculation of the pixel displacement acceleration, the velocity is differentiated again, that is, the pixel displacement acceleration in the horizontal direction , where is the change value of the pixel displacement velocity in the horizontal direction within the adjacent time interval. Similarly, the pixel displacement acceleration in the vertical direction can be obtained . Such a displacement time derivative calculation is performed on each feature point in adjacent frame images in both the horizontal and vertical directions, and finally the pixel displacement velocity and pixel displacement acceleration corresponding to each feature point are obtained.

[0145] Preferably, based on the angle between adjacent frame displacement vectors, the pixel displacement velocity and pixel displacement acceleration corresponding to the feature points in adjacent frames, and in combination with the pyramid hierarchical optical flow method, optical flow tracking of the motion trajectories is performed between the corresponding adjacent frame feature points in the visual image feature point set to generate the feature point motion trajectories between adjacent frame images.

[0146] In an embodiment of the present invention, by adopting the pyramid hierarchical optical flow method, an image pyramid of adjacent frame images is first constructed, and the original image is continuously downsampled to obtain image layers with different resolutions. For each feature point in the visual image feature point set, on the bottom layer image (high-resolution image), in combination with the included angle between the known displacement vectors of adjacent frames, the pixel displacement velocity, and the pixel displacement acceleration information, taking this information as the initial condition for optical flow calculation, based on the Lucas-Kanade optical flow algorithm, the size and moving direction of the search window are adjusted according to the pixel displacement velocity and acceleration, and the included angle information between the adjacent frame displacement vectors is used to optimize the angle of the search direction. Starting from the bottom layer image, the optical flow is calculated to obtain a preliminary displacement estimate of the feature point on this layer image, and then this estimated value is passed to the upper layer (slightly lower resolution) image, and the optical flow calculation and displacement estimate optimization are continued. By analogy, until the top layer image, finally, the displacement estimates of the feature points obtained on each layer image are combined to obtain the accurate motion trajectory of the feature point between the adjacent frame images. Such optical flow tracking operations are performed on all feature points in the visual image feature point set, thereby generating a complete feature point motion trajectory between the adjacent frame images.

[0147] Further, step S3 includes the following steps:

[0148] Step S31: Project the motion trajectory of the feature points between adjacent frame images into a three-dimensional feature space, and determine the three-dimensional coordinates corresponding to the feature points of the adjacent frame images in the three-dimensional feature space to obtain the three-dimensional coordinates of each adjacent frame image feature point;

[0149] In an embodiment of the present invention, by obtaining the camera internal parameter matrix corresponding to the visual sensor , where and are the focal lengths corresponding to the camera in the and directions respectively, and are the coordinates of each point of the visual motion image respectively, and by assuming that the adjacent frame images are the th frame and the th frame , and obtaining the corresponding feature point set in the th frame , and at the same time obtaining the matching feature point set in the th frame , where represents the total number of feature points of the adjacent frame images, represents the th frame Adjacent frame image feature points, specifically , where represents the frame and the abscissa of the th adjacent frame image feature point, represents the frame and the ordinate of the th adjacent frame image feature point, represents the frame and the corresponding th adjacent frame image feature point, specifically , where represents the frame and the abscissa of the th adjacent frame image feature point, represents the frame and the ordinate of the th adjacent frame image feature point, represents the transpose symbol; based on the camera intrinsic matrix perform normalization coordinate transformation on the frame feature point set or the frame feature point set to convert the pixel coordinates to normalized plane coordinates to obtain the frame normalized coordinates or the frame normalized coordinates ; project the frame normalized coordinates or the frame normalized coordinates into a three-dimensional feature space to obtain the three-dimensional coordinates of each adjacent frame image feature point , where the three-dimensional coordinates of the th adjacent frame image feature point .

[0150] Step S32: Reproject the three-dimensional coordinates of each adjacent frame image feature point onto the image plane coordinate system to obtain the reprojection coordinates of each adjacent frame image feature point;

[0151] In the embodiment of the present invention, by assuming that the external parameters corresponding to the frame of the camera are the rotation matrix and the translation vector respectively, and converting it from the three-dimensional coordinate system to the camera coordinate system according to the corresponding external parameters, the The corresponding coordinates in the frame camera coordinate system are specifically , and project it onto the image plane according to the projection formula corresponding to the camera. The specific projection formula is , where represents the reprojection coordinate of the th adjacent frame image feature point on the th frame image plane, represents the scale factor, and the corresponding reprojection coordinates are obtained by solving the above projection formula , , and collect all the reprojection coordinates to obtain the reprojection coordinates of each adjacent frame image feature point .

[0152] Step S33: Obtain the actual observed coordinates corresponding to each adjacent frame image feature point on the image plane, and calculate the visual error based on the actual observed coordinates corresponding to each adjacent frame image feature point on the image plane and the reprojection coordinates of each adjacent frame image feature point, so as to obtain the visual reprojection error vector corresponding to each adjacent frame image feature point;

[0153] In the embodiment of the present invention, by obtaining the actual observed coordinates corresponding to each adjacent frame image feature point on the image plane , and measure the error with the reprojection coordinates of the adjacent frame image feature points to calculate the visual reprojection error vector, which is specifically , where represents the total number of adjacent frame image feature points, represents the transpose symbol, represents the visual reprojection error value between the th adjacent frame image feature points, represents the actual observed coordinates corresponding to the th adjacent frame image feature point on the image plane , represents the th adjacent frame image feature point reprojection coordinate .

[0154] Step S34: Perform visual motion transformation calculation on the three-dimensional coordinates of each adjacent frame image feature point based on the visual reprojection error vector corresponding to each adjacent frame image feature point, so as to obtain the visual motion transformation matrix between adjacent frame images;

[0155] In the embodiment of the present invention, by constructing the combination of the previously obtained visual reprojection error vector corresponding to the th frame and the th image feature point construct the corresponding error function , where is the rotation matrix between adjacent frame images, is the translation vector, and the error function is minimized by using a non - linear optimization algorithm (such as the Levenberg - Marquardt algorithm). During the optimization process, and are continuously adjusted so that the value of the error function gradually decreases. Let the initial rotation matrix be and the translation vector be . Through iterative update of , where is the Jacobian matrix of the error function E with respect to and , is the damping factor, is the vector composed of the visual reprojection error vectors of all feature points. After multiple iterations, when the error function converges to the minimum value, the visual motion transformation matrix between adjacent frame images is finally obtained.

[0156] Step S35: Perform three - dimensional camera pose conversion calculation according to the visual motion transformation matrix between adjacent frame images to obtain the visual pose conversion matrix between adjacent frame images in the camera coordinate system.

[0157] In the embodiment of the present invention, by performing singular value decomposition on the visual motion transformation matrix between adjacent frame images, the rotation matrix between the feature points of adjacent frame images is obtained, where and represent orthogonal matrices, represents a diagonal matrix. Since the rotation matrix is an orthogonal matrix, take as the rotation matrix between the feature points of adjacent frame images, and the translation vector is directly obtained from the visual motion transformation matrix. The pose matrix of the i - th frame camera is calculated as in the camera coordinate system. Then the visual pose conversion matrix between adjacent frame images in the camera coordinate system is specifically , where represents the visual pose conversion matrix between the -th and -th frame images in the camera coordinate system, represents the rotation matrix between the feature points of adjacent frame images, represents the translation vector between the feature points of adjacent frame images.

[0158] Furthermore, step S4 includes the following steps:

[0159] Step S41: Perform adjacent-frame pose pre-integration on the inertial data of the target carrier to obtain the inertial pose change matrix between adjacent-frame times in the inertial coordinate system;

[0160] In an embodiment of the present invention, by obtaining the inertial data corresponding to the target carrier, which is provided by an inertial measurement unit (IMU) and includes the angular velocity and linear acceleration between adjacent-frame times. Assuming that the adjacent-frame times are and , within this time interval , the inertial data is discretized. Let the sampling frequency be , then there are sampling points within this time interval. For the angular velocity , at each sampling point , its measured value is . Considering the zero bias of the gyroscope, the true angular velocity is . And for the linear acceleration , at each sampling point , its measured value is . Considering the zero bias of the accelerometer and the gravitational acceleration , the true linear acceleration in the inertial coordinate system is , where is the rotation matrix from the carrier coordinate system to the inertial coordinate system at time . By integrating the true angular velocity, the rotational pre-integration quantity between adjacent-frame times can be obtained. The specific calculation uses the integration method of quaternions or rotation matrices, for example, using the midpoint integration method: , where is the skew-symmetric matrix of , is the sampling time interval. By double-integrating the linear acceleration, the position pre-integration quantity and the velocity pre-integration quantity between adjacent-frame times can be obtained. Taking the position pre-integration quantity as an example, the corresponding calculation formula is specifically . Combine the rotational pre-integration quantity and the position pre-integration quantity into the inertial pose change matrix between adjacent-frame times in the inertial coordinate system, and finally obtain the inertial pose change matrix between adjacent-frame times in the inertial coordinate system.

[0161] Step S42: Perform visual-inertial extrinsic parameter calibration based on the visual pose transformation matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame times in the inertial coordinate system, so as to generate the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

[0162] In the embodiment of the present invention, by using the previously calculated visual pose transformation matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame times in the inertial coordinate system to find the transformation matrix from the camera coordinate system to the inertial coordinate system , so that the visual and inertial data can be aligned in the same coordinate system. According to the coordinate transformation relationship, there is . Expand the above equation to obtain an equation about and . For the rotation part, there is ; for the translation part, there is . In order to solve and , collect multiple sets of and data of adjacent frames, construct a non-linear least squares problem, define the error function

[0163] , and use a non-linear optimization algorithm (such as the Levenberg-Marquardt algorithm) to minimize the error function . During the optimization process, continuously adjust the and values so that the value of the error function gradually decreases. When the error function converges to the minimum value, the obtained and are combined into the corresponding visual-inertial extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system .

[0164] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.

[0165] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A visual inertial navigation calibration method, characterized in that: The following steps are involved: Step S1: obtaining visual motion image data corresponding to the target carrier in the road environment through a visual sensor, and synchronously collecting corresponding target carrier inertial data through an inertial measurement unit built into the target carrier; performing motion feature point analysis on the visual motion image data to obtain a visual image feature point set; Step S2: determining the pixel displacement of feature points between adjacent frames of visual images through the visual image feature point set, and performing motion trajectory optical flow tracking between corresponding adjacent frame feature points in the visual image feature point set based on the pixel displacement of feature points between adjacent frames of visual images, so as to generate a motion trajectory of feature points between adjacent frame images; Step S3: Calculate the visual motion transformation matrix between adjacent frame images according to the motion trajectory of the feature points between adjacent frame images, and perform three-dimensional camera posture transformation calculation according to the visual motion transformation matrix to obtain the visual posture transformation matrix between adjacent frame images in the camera coordinate system; wherein step S3 includes the following steps: Step S31: projecting the motion trajectory of the feature points between adjacent frame images into the three-dimensional feature space, and determining the three-dimensional coordinates corresponding to the feature points of the adjacent frame images in the three-dimensional feature space to obtain the three-dimensional coordinates of the feature points of each adjacent frame image; Step S32: reprojecting the three-dimensional coordinates of the feature points of each adjacent frame image into the image plane coordinate system to obtain the reprojected coordinates of the feature points of each adjacent frame image; Step S33: obtaining the actual observed coordinates corresponding to the feature points of each adjacent frame image on the image plane, and performing visual error calculation according to the actual observed coordinates corresponding to the feature points of each adjacent frame image on the image plane and the reprojected coordinates of each adjacent frame image feature point, so as to obtain the visual reprojection error vector corresponding to each adjacent frame image feature point; Step S34: performing visual motion transformation calculation on the three-dimensional coordinates of the feature points of each adjacent frame image based on the visual reprojection error vector corresponding to the feature points of each adjacent frame image, so as to obtain a visual motion transformation matrix between adjacent frame images; Step S35: performing three-dimensional camera posture conversion calculation according to the visual motion transformation matrix between adjacent frame images to obtain the visual posture conversion matrix between adjacent frame images in the camera coordinate system; Step S4: Pre-integrate the adjacent frame poses of the target carrier inertial data to obtain the inertial pose change matrix between adjacent frame moments in the inertial coordinate system; perform visual-inertial navigation extrinsic parameter calibration according to the visual pose conversion matrix between adjacent frame images in the camera coordinate system and the inertial pose change matrix between adjacent frame moments in the inertial coordinate system to generate the corresponding visual-inertial navigation extrinsic parameter calibration matrix from the camera coordinate system to the inertial coordinate system.

2. The visual inertial navigation calibration method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: acquiring visual motion image data corresponding to the target carrier in the road environment by deploying a visual sensor on the target carrier; Step S12: Sending a synchronization pulse through a synchronization signal generator to synchronously collect corresponding target carrier inertial data using an inertial measurement unit built into the target carrier; Step S13: performing histogram equalization processing on each frame of the visual motion image data to obtain visual motion contrast balanced image data; Step S14: performing pixel blur analysis on each frame of the visual motion contrast balanced image data to obtain pixel blur corresponding to each frame of the visual motion image; performing pixel blur filtering on each frame of the visual motion contrast balanced image data based on the pixel blur corresponding to each frame of the visual motion image to filter and remove corresponding pixel high-frequency noise and pixel abnormal values ​​to obtain visual motion filtered image data; Step S15: Input the visual motion filtered image data into the pre-trained convolutional neural network model to extract the corresponding feature key points and their descriptors in each frame image, and perform non-maximum suppression on the corresponding feature key points and their descriptors in each frame image to obtain a visual image feature point set.

3. The visual inertial navigation calibration method according to claim 2, characterized in that: The non-maximum suppression of the corresponding feature key points and their descriptors in each frame of image comprises: The deep feature analysis is performed on the corresponding key points and their descriptors in each frame of the image, so as to analyze the local neighborhood structure corresponding to the key points in the visual motion image, and the texture direction distribution characteristics of the key points are characterized by calculating the gradient direction histogram of the pixels in the neighborhood of the key points. For the descriptors, their numerical distribution laws and dimensional correlations in different color spaces are analyzed to obtain the deep feature data set of the key points and descriptors. Based on the feature key point and descriptor deep characteristic data set, the feature key points and their descriptors corresponding to each frame of the image are preliminarily screened, and the feature key points and their descriptors in each frame of the image are associated in the time dimension considering the temporal continuity between frames, from which the position changes corresponding to the feature key points of adjacent frames and the similarities corresponding to the descriptors are analyzed, and the feature key points with abnormal position changes or the descriptors with descriptor similarity lower than the threshold are marked as objects to be excluded for screening and exclusion, so as to obtain a set of feature key points and descriptors after the preliminary screening; A multi-scale local extreme value region is defined for each point in the feature key points and descriptor set after the initial screening, so as to determine the corresponding local extreme value region according to the pixel texture condition in its neighborhood, and obtain the multi-scale local extreme value region; Based on the multi-scale local extreme value region, non-maximum suppression is performed on the feature key points and descriptor sets after the initial screening to obtain the visual image feature point set.

4. The visual inertial navigation calibration method according to claim 3, characterized in that: The multi-scale local extreme value region specifically sets a 1x1 scale local extreme value region in a complex texture region to capture detail features, and sets a 5x5 scale local extreme value region in a smooth texture region to consider overall features.

5. The visual inertial navigation calibration method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: performing adjacent frame feature point matching on the visual image feature point set to obtain adjacent frame image feature point matching pairs; Step S22: obtaining the motion direction and speed range corresponding to the target carrier, and performing adjacent frame feature displacement constraint analysis according to the motion direction and speed range corresponding to the target carrier, so as to analyze the displacement direction and size constraint corresponding to the feature point in the adjacent frames, and obtain the adjacent frame feature point displacement constraint; Step S23: calculating the pixel displacement of feature points between each pair of feature points in the feature point matching pairs of the adjacent frame images based on the displacement constraint of the feature points of the adjacent frames, so as to obtain the pixel displacement of feature points between the adjacent frame visual images; Step S24: Based on the pixel displacement of feature points between adjacent frame visual images, optical flow tracking of motion trajectories is performed between corresponding adjacent frame feature points in the visual image feature point set to generate feature point motion trajectories between adjacent frame images.

6. The visual inertial navigation calibration method according to claim 5, characterized in that: Step S24 includes the following steps: Step S241: taking the pixel displacement of feature points between adjacent frame visual images as a spatial vector to generate a pixel displacement vector corresponding to each feature point in the adjacent frame images; Step S242: performing directional component decomposition on the pixel displacement vector corresponding to each feature point in the adjacent frame images to obtain the pixel displacement components corresponding to each feature point in the adjacent frame images in the horizontal, vertical and diagonal directions; Step S243: inversely solving the displacement angle between each feature point based on the pixel displacement components corresponding to each feature point in the horizontal, vertical and diagonal directions in the adjacent frame images to obtain the angle between the displacement vectors of adjacent frames; Step S244: Based on the angle between the displacement vectors of adjacent frames and in combination with the pixel displacement of the feature points between the adjacent frame visual images, the motion trajectory optical flow tracking is performed between the corresponding adjacent frame feature points in the visual image feature point set to generate the feature point motion trajectory between the adjacent frame images.

7. The visual inertial navigation calibration method according to claim 6, characterized in that: Step S244 includes the following steps: Perform displacement change analysis on the pixel displacement of feature points between adjacent frames of visual images in the time dimension to obtain the change value of the pixel displacement of feature points of adjacent frames corresponding to time; The displacement time derivative is calculated according to the change value of the pixel displacement of the feature point of the adjacent frames corresponding to time, so as to obtain the pixel displacement speed and pixel displacement acceleration corresponding to the feature point of the adjacent frames; Based on the angle between the displacement vectors of adjacent frames, the pixel displacement speed and pixel displacement acceleration corresponding to the feature points of adjacent frames, and combined with the pyramid hierarchical optical flow method, the motion trajectory optical flow tracking is performed between the corresponding adjacent frame feature points in the visual image feature point set to generate the motion trajectory of the feature points between adjacent frame images.

8. The visual inertial navigation calibration method according to claim 1, characterized in that: The three-dimensional coordinates of the feature points of each adjacent frame image described in step S31 are specifically: Get the camera intrinsic parameter matrix corresponding to the visual sensor ,in and The camera is and The focal length corresponding to the direction, and They are the point-by-point coordinates of the visual motion image; Let the adjacent frame image be frame With frame , and in frame Get the corresponding feature point set , while in frame Get the matching feature point set ,in Represents the total number of feature points in adjacent frame images, Indicates frame The corresponding adjacent frame image feature points, specifically: ,in Indicates frame Neidi The horizontal coordinates of the feature points of the adjacent frame images are Indicates frame Neidi The vertical coordinates of the feature points of the adjacent frame images are Indicates frame The corresponding adjacent frame image feature points, specifically: ,in Indicates frame Neidi The horizontal coordinates of the feature points of the adjacent frame images are Indicates frame Neidi The vertical coordinates of the feature points of the adjacent frame images are represents the transpose symbol; Based on the camera intrinsic parameter matrix For Frame feature point set Or Frame feature point set Perform normalized coordinate transformation to obtain Frame Normalized Coordinates Or Frame Normalized Coordinates ; The first Frame Normalized Coordinates Or Frame Normalized Coordinates Perform three-dimensional feature space mapping to obtain the three-dimensional coordinates of feature points of each adjacent frame image , among which 3D coordinates of feature points in adjacent frames .

9. The visual inertial navigation calibration method according to claim 8, characterized in that: The reprojection coordinates of the feature points of each adjacent frame image described in step S32 are specifically: Assume that the camera is The external parameters corresponding to the frames are the rotation matrices and the translation vector , and transform it from the three-dimensional coordinate system to the camera coordinate system according to the corresponding external parameters. The corresponding coordinates in the frame camera coordinate system are ; It is projected onto the image plane according to the projection formula corresponding to the camera, and the projection formula is specifically: ; in, Indicates The feature points of the adjacent frame images are The horizontal coordinate of the reprojection on the frame image plane, Indicates The feature points of the adjacent frame images are The ordinate of the reprojection on the frame image plane, represents the scale factor; By solving the above projection formula, we can get the corresponding reprojection coordinates. , , and collect all the reprojection coordinates to obtain the reprojection coordinates of the feature points of each adjacent frame image .

10. The visual inertial navigation calibration method according to claim 1, characterized in that: The visual reprojection error vector described in step S33 is specifically: ,in Represents the total number of feature points in adjacent frame images, represents the transpose symbol, Indicates The visual reprojection error value between feature points of adjacent frame images, Indicates The actual observed coordinates corresponding to the feature points of the adjacent frames on the image plane , Indicates Reprojection coordinates of feature points in adjacent frames .

11. The visual inertial navigation calibration method according to claim 1, characterized in that: Step S35 includes the following steps: Perform singular value decomposition on the visual motion transformation matrix between adjacent frame images to obtain the rotation matrix and translation vector between the feature points of adjacent frame images; The three-dimensional camera pose transformation is calculated based on the rotation matrix and translation vector between the feature points of adjacent frame images to obtain the visual pose transformation matrix between adjacent frame images in the camera coordinate system.

12. The visual inertial navigation calibration method according to claim 11, characterized in that: The visual pose conversion matrix is ​​specifically: ; in, Indicates the camera coordinate system Frame and The visual pose transformation matrix between frame images, Represents the rotation matrix between feature points of adjacent frame images, Represents the translation vector between feature points of adjacent frame images.

13. The visual inertial navigation calibration method according to claim 1, characterized in that: The inertial posture change matrix described in step S4 is specifically: ; in, In the inertial coordinate system, Frame and The inertial pose change matrix between frame images, Indicates Frame and The rotation pre-integration quantity between frame images, Indicates Frame and Position pre-integration quantity between frame images.

Citation Information

Patent Citations

  • Visual navigation and inertial navigation combination navigation method based on aircraft

    CN107014380A

  • Positioning method and system based on visual inertial navigation information fusion

    CN107869989A