A vision-based inertial navigation calibration method
By installing vision sensors and inertial navigation sensors on the target carrier, the data is measured and aligned in real time, and the visual-inertial calibration method is used to solve the problem of inertial navigation sensor error accumulation, achieving high-precision navigation and motion state tracking.
Patent Information
- Application Number
- CN202510308762.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-17
AI Technical Summary
During long-term operation, the existing inertial guidance calibration methods have accumulated sensor errors, resulting in the continuous increase in navigation errors. There is a lack of systematic and efficient inertial guidance calibration methods, which cannot accurately compensate for the errors of inertial guidance sensors.
The vision-based inertial calibration method is adopted, by installing a vision sensor and an inertial navigation sensor on the target carrier, the visual image data and inertial motion data are measured in real time, and the visual-inertial calibration is performed through time stamp alignment to obtain the error calibration compensation result of the inertial navigation sensor.
Through visual-inertial calibration, the errors of the inertial navigation sensor can be accurately identified and compensated, navigation accuracy can be improved, and the motion state tracking and navigation performance of the target carrier can be significantly improved.
Smart Images

Figure CN119879989B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of calibration measurement, and particularly to a vision-based inertial navigation calibration method. Background Art
[0002] The inertial navigation system is an important autonomous navigation system. By measuring the acceleration and angular velocity of the carrier, it uses integral operations to calculate the position, velocity, and attitude information of the carrier. Traditional inertial navigation calibration methods are mainly based on turntable tests. By testing the inertial navigation system at multiple positions and attitudes on a high-precision turntable and using the known motion information of the turntable to estimate the error parameters of the inertial navigation sensors. However, during the long-term operation of the inertial navigation system, due to the accumulation of sensor errors (such as gyroscope drift, accelerometer zero bias, etc.), the navigation error will continuously increase. At the same time, the current methods of combining vision technology with inertial navigation calibration are not yet perfect, lacking systematic and efficient inertial navigation calibration, resulting in the inability to accurately compensate for the errors of inertial navigation sensors. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a vision-based inertial navigation calibration method to solve at least one of the above technical problems.
[0004] To achieve the above object, a vision-based inertial navigation calibration method includes the following steps:
[0005] Step S1: Install and deploy corresponding vision sensors and inertial navigation sensors on the target carrier, and use the vision sensors and inertial navigation sensors to measure the target vision image data and target inertial navigation motion data of the target carrier in real time; add corresponding timestamps to align the target vision image data and target inertial navigation motion data in the time dimension to obtain the corresponding target vision image sequence and target inertial navigation motion sequence in the same time dimension;
[0006] Step S2: Perform static and dynamic motion feature fusion analysis on the target vision image corresponding to each time point in the target vision image sequence to obtain the static and dynamic motion vision features of the target carrier;
[0007] Step S3: Search for the nearest neighbor feature point pairs of the static and dynamic motion vision features of the target carrier to obtain the motion matching feature point pairs of the target carrier; perform visual motion estimation on the target carrier based on the motion matching feature point pairs of the target carrier to obtain the visual motion state variables of the target carrier, including the visual motion translation vector of the target carrier and the visual motion rotation matrix of the target carrier;
[0008] Step S4: Perform visual-inertial calibration on the error measurement corresponding to the inertial navigation sensor based on the visual motion state variables of the target carrier and the target inertial navigation motion sequence to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
[0009] Further, step S1 includes the following steps:
[0010] Step S11: Install and deploy corresponding vision sensors and inertial navigation sensors on the target carrier;
[0011] Step S12: Use the built-in synchronization signal generator to perform hardware synchronization triggering to synchronously and real-time measure the target carrier using the vision sensor and the inertial navigation sensor, so as to obtain target visual image data and target inertial navigation motion data;
[0012] Step S13: Perform data cleaning and standardization on the target visual image data and the target inertial navigation motion data, so as to enhance the contrast, remove noise, and standardize pixels for the target visual image data, while removing duplicate data, outliers, missing value compensation, and standardizing motion parameters for the target inertial navigation motion data, so as to obtain target visual standard image data and target inertial navigation standard data;
[0013] Step S14: Align the target visual standard image data and the target inertial navigation standard data in the time dimension by adding corresponding timestamps, so as to obtain corresponding target visual image sequences and target inertial navigation motion sequences in the same time dimension.
[0014] Further, step S11 includes the following steps:
[0015] Perform material and dynamics analysis on the target carrier to obtain the corresponding structural material, motion characteristics, and environmental characteristics of the target carrier;
[0016] Based on the corresponding structural material, motion characteristics, and environmental characteristics of the target carrier, perform sensing adaptation planning analysis on the vision sensor and the inertial navigation sensor to adapt and plan the corresponding types, accuracy requirements, and protection levels of the vision sensor and the inertial navigation sensor, and obtain the target carrier sensor adaptation planning result;
[0017] Based on the target carrier sensor adaptation planning result, perform sensing installation topology design on the to-be-installed corresponding vision sensor and inertial navigation sensor, so as to determine the corresponding optimal installation angle and position for the vision sensor according to the changing motion field of the target carrier, and ensure that the target area is covered to the greatest extent under various motion states, while installing the inertial navigation sensor at the position with the least vibration and most easily perceiving motion according to the dynamic characteristics of the target carrier, so as to generate the target carrier sensor installation topology structure;
[0018] Based on the target carrier sensor installation topology, the corresponding visual sensors and inertial navigation sensors are installed on the target carrier and sensor calibration is performed. The corresponding internal parameters of the visual sensors, including focal length and distortion coefficient, are calibrated in combination with the optical imaging principle. Gravity and turntable calibration equipment is used to calibrate the corresponding zero bias and scale factor of the inertial navigation sensor according to the gravity field and the law of rotational motion, so as to realize the corresponding calibration work of the visual sensors and inertial navigation sensors.
[0019] Furthermore, the visual sensor includes a binocular camera or a monocular camera combined with a laser radar.
[0020] Furthermore, the inertial navigation sensor includes a gyroscope and an accelerometer.
[0021] Further, step S12 includes the following steps:
[0022] The built-in synchronization signal generator of the target carrier is used as the core node to synchronize the trigger topology with the visual sensor and the inertial navigation sensor in a star structure, and the feedback signals corresponding to each sensor are aggregated to the synchronization signal generator using a bus structure to achieve two-way communication, so as to generate the hardware synchronization trigger topology of the target carrier;
[0023] Based on the hardware synchronous trigger topology of the target carrier, a hardware synchronous trigger method is used to send a synchronous trigger signal to the corresponding visual sensor and inertial navigation sensor to generate a target carrier synchronous trigger signal pulse;
[0024] The target carrier is synchronized with the trigger signal pulse response to control the corresponding visual sensor and inertial navigation sensor, and the transmission delay of the synchronization signal pulse from the generator to the sensor is measured. At the same time, the corresponding measurement time is fine-tuned based on the transmission delay to perform synchronous real-time measurement of the target carrier to obtain the target visual image data and target inertial navigation motion data.
[0025] Further, step S2 includes the following steps:
[0026] Step S21: obtaining the target visual images corresponding to the adjacent time frames before and after the time point through the target visual image corresponding to the target visual image sequence at each time point;
[0027] Step S22: performing target carrier change analysis between target visual images corresponding to the corresponding time points based on the target visual images corresponding to the previous and next adjacent time frames at the time point, so as to compare and analyze the target edge and texture difference changes between the target visual images corresponding to the previous and next adjacent time frames at the time point, so as to obtain the target visual image difference changes between each time point and the previous and next adjacent time frames;
[0028] Step S23: Based on the change in the target visual image difference between each time point and the adjacent time frames before and after, perform static and dynamic image judgment on the corresponding target visual image at the corresponding time point. If there is a difference change between each time point and the adjacent time frames before and after, determine the corresponding target visual image at this time point as a dynamic visual image; if there is no difference change between each time point and the adjacent time frames before and after, determine the corresponding target visual image at this time point as a static visual image;
[0029] Step S24: Perform static motion feature analysis on the static visual image to obtain the static motion features of the target carrier; perform dynamic motion feature analysis on the dynamic visual image to obtain the dynamic motion features of the target carrier;
[0030] Step S25: Perform static and dynamic motion feature fusion analysis on the static motion features of the target carrier and the dynamic motion features of the target carrier to obtain the static and dynamic motion visual features of the target carrier.
[0031] Further, the static motion feature analysis of the static visual image described in step S24 includes the following steps:
[0032] Perform Gaussian blur and downsampling operations on the static visual image to obtain the corresponding target static visual sub-images at each scale;
[0033] Perform target edge structure feature analysis on the corresponding target static visual sub-images at each scale to extract the target edge contour and regional shape corresponding to the static visual image by using edge detection and morphological operations, and obtain the corresponding target carrier static visual edge structure features at each scale;
[0034] Based on the corresponding target carrier static visual edge structure features at each scale, perform static visual key point extraction on the corresponding target static visual sub-images at each scale to extract the corresponding corner points and edge intersection points, and obtain the key point set of the target carrier static visual image;
[0035] Based on the key point set of the target carrier static visual image and using scale-invariant feature transform and speeded-up robust features, perform static motion feature analysis on the static visual image to obtain the static motion features of the target carrier, including the shape, texture, and color features corresponding to the target carrier.
[0036] Further, the dynamic motion feature analysis of the dynamic visual image described in step S24 includes the following steps:
[0037] Take each dynamic visual image as a spatio-temporal cube, and analyze the gray-scale change and gradient change between adjacent frame dynamic visual images in the time dimension;
[0038] Perform motion cue statistical analysis on each dynamic vision image based on the gray-scale change and gradient change between adjacent-frame dynamic vision images to obtain the dynamic motion direction of the target carrier and the dynamic motion speed of the target carrier;
[0039] Based on the dynamic motion direction of the target carrier and the dynamic motion speed of the target carrier, and combined with the sparse optical flow method based on the Lucas-Kanade algorithm, perform dynamic motion feature analysis on the dynamic vision image corresponding to the target carrier to obtain the dynamic motion features of the target carrier, including the motion trajectory and speed change corresponding to the target carrier during the motion process.
[0040] Further, step S3 includes the following steps:
[0041] Step S31: Search for the nearest neighbor feature point pairs of the static and dynamic motion vision features of the target carrier to obtain the motion matching feature point pairs of the target carrier;
[0042] Step S32: Based on the motion matching feature point pairs of the target carrier and combined with triangulation and the principle of epipolar geometry, perform spatio-temporal feature point correlation analysis on the target carrier to analyze the triangle angle and side length change relationship formed by the feature points in the adjacent-frame images of the target carrier, and obtain the spatio-temporal correlation relationship between the feature points of the target carrier;
[0043] Step S33: Based on the spatio-temporal correlation relationship between the feature points of the target carrier, perform motion speed and angle statistical analysis on the motion trajectory corresponding to the target carrier to obtain the motion speed and motion rotation angle corresponding to the target carrier on the three spatial axes;
[0044] Step S34: Based on the motion speed and motion rotation angle corresponding to the target carrier on the three spatial axes, perform visual motion estimation on the target carrier to obtain the visual motion state variables of the target carrier, including the visual motion translation vector of the target carrier and the visual motion rotation matrix of the target carrier.
[0045] Further, step S31 includes the following steps:
[0046] Perform motion vision depth deconstruction on the static and dynamic motion vision features of the target carrier, use Fourier transform and wavelet transform to analyze the frequency components and spatial distribution characteristics of the static features, and analyze the change trend corresponding to different time points of the dynamic features by establishing a time series model to obtain the static and dynamic feature deconstruction data of the target carrier;
[0047] Based on the static and dynamic characteristics of the target carrier, deconstruct the data, construct the spatio-temporal encoding feature point index for each feature point in the static and dynamic motion visual features of the target carrier, adopt the corresponding encoding method based on the Hilbert curve for spatial encoding to map the three-dimensional spatial coordinates of the feature points into codes, and generate corresponding time codes according to the time sequence of the appearance of the feature points for time encoding, so as to obtain the static and dynamic visual index features of the target carrier after spatio-temporal encoding;
[0048] According to the static and dynamic visual index features of the target carrier after spatio-temporal encoding, perform feature point neighborhood division, adopt the Delaunay triangulation part to divide the corresponding static and dynamic visual feature points of the target carrier into corresponding triangular neighborhoods, define large-scale neighborhoods for feature points located at the edges and corners of the target carrier with obvious feature changes, and define small-scale neighborhoods for general feature points, to obtain the multi-scale feature point neighborhood division result of the target carrier;
[0049] Based on the multi-scale feature point neighborhood division result of the target carrier, screen the nearest neighbor candidate points for the feature points corresponding to the corresponding scale neighborhoods in the static and dynamic motion visual features of the target carrier, calculate the corresponding energy values and visual similarities between each feature point in the feature point neighborhood, and screen out the feature points with high energy values and high visual similarities as the nearest neighbor candidate points, so as to obtain the motion matching feature point pairs of the target carrier.
[0050] Further, the specific visual motion translation vector of the target carrier described in step S34 is:
[0051] ;
[0052] ;
[0053] ;
[0054] ;
[0055] Among them, represents the visual motion translation vector of the target carrier, represents the translation component along the X-axis, represents the translation component along the Y-axis, represents the translation component along the Z-axis, represents the motion measurement time, represents the corresponding motion speed of the target carrier on the X-axis, represents the corresponding motion speed of the target carrier on the Y-axis, represents the corresponding motion speed of the target carrier on the Z-axis.
[0056] Further, the specific visual motion rotation matrix of the target carrier described in step S34 is:
[0057] The rotation matrix corresponding to rotation about the X-axis is: ;
[0058] The rotation matrix corresponding to rotation about the Y-axis is: ;
[0059] The rotation matrix corresponding to rotation about the Z-axis is: ;
[0060] Among them, represents the motion rotation angle corresponding to the target carrier on the X-axis, represents the motion rotation angle corresponding to the target carrier on the Y-axis, represents the motion rotation angle corresponding to the target carrier on the Z-axis.
[0061] Furthermore, step S4 includes the following steps:
[0062] Step S41: Analyze the dual-sensor motion parameter characteristics of the target carrier's visual motion state variables and the target inertial navigation motion sequence to obtain a dual-sensor motion parameter characteristic data set;
[0063] Step S42: Perform dual-sensor data fusion on the target carrier's visual motion state variables and the target inertial navigation motion sequence based on the dual-sensor motion parameter characteristic data set, so as to use the motion measurement of the inertial navigation sensor in a short time to assist in fusing the motion parameter estimation of the visual sensor, and obtain a dual-sensor motion parameter fusion vector;
[0064] Step S43: Based on the dual-sensor motion parameter fusion vector and using the extended Kalman filter to perform visual-inertial calibration on the error measurement corresponding to the inertial navigation sensor, taking the error parameters corresponding to the inertial navigation sensor as state variables, including gyroscope drift and accelerometer zero bias, and taking the measurement values of the visual sensor and the inertial sensor in the dual-sensor motion parameter fusion vector as observation variables, and continuously updating the state variables through the filtering algorithm to realize the real-time estimation and calibration compensation of the error measurement corresponding to the inertial navigation sensor, so as to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
[0065] Furthermore, the process of the dual-sensor motion parameter characteristic analysis in step S41 is specifically to analyze the translation error range between the translation vector corresponding to the visual sensor and the motion trajectory in different directions, analyze the rotation angle accuracy corresponding to the rotation matrix in different directions, and use the dynamic model to analyze the response characteristics of the angular velocity measured by the gyroscope and the acceleration measured by the accelerometer of the motion measurement parameters corresponding to the inertial navigation sensor in different motion states, and evaluate its zero bias stability and scale factor error.
[0066] Advantages of the present invention:
[0067] The vision-based inertial navigation calibration method proposed by the present invention, compared with the prior art, the beneficial effects of the present application are as follows: By installing and deploying vision sensors and inertial navigation sensors, the visual image data and inertial navigation motion data of the target are measured in real time on the target carrier. Through timestamp annotation and alignment of these data, the synchronization of visual images and inertial navigation data in the same time dimension is achieved. This synchronization processing ensures the timeliness and accuracy of sensor data and is the basis for subsequent data fusion and motion estimation. First of all, the vision sensor can provide rich external environment information for the target, such as image, video data, etc. The inertial navigation sensor measures the acceleration and angular velocity of the target through the accelerometer and gyroscope, providing accurate dynamic information for the motion state of the target. Since the data acquisition frequencies and measurement dimensions of these two types of sensors are different, relying solely on either of these sensors may not be able to comprehensively describe the motion of the target. Therefore, by aligning timestamps and fusing visual images and inertial navigation data, more accurate motion estimation can be achieved, thereby providing a basis for subsequent static and dynamic feature extraction, feature point matching, and error correction, and laying a solid foundation for subsequent multi-sensor fusion and motion state estimation. Secondly, through the fusion analysis of static and dynamic motion features of the target visual image at each time point, the visual image sequence contains information such as the position, attitude, and shape change of the target in the dynamic environment, and this information is usually manifested as the time change of visual features. By fusing static and dynamic motion features, the motion information of the target at different time points can be effectively extracted. Static features usually refer to the fixed attributes such as the geometric shape and texture of the target, while dynamic features represent the behaviors such as the motion trajectory, speed, and rotation of the target that change over time. Fusing these two can improve the recognition and modeling accuracy of the target motion state. The key to this step lies in the accurate extraction and fusion of static and dynamic information in the target image sequence, enabling the inertial navigation system to better understand the motion trajectory and its behavioral characteristics of the target in three-dimensional space. By comprehensively analyzing static and dynamic features, the details of the target motion can be more accurately captured in a complex environment, thereby providing a more accurate visual basis for subsequent feature point matching and motion estimation.Then, by performing a nearest-neighbor feature point pair search on the static and dynamic motion visual features of the target carrier and based on this, performing visual motion estimation on the target carrier. In the target visual image sequence, the feature points can be corner points, edges, or other image regions with unique identifiers in the image. These feature points can reflect the position changes of the target in three-dimensional space. By matching these feature points, the target visual images at different time points can be linked together, thereby obtaining the motion trajectory of the target. Using the matched feature points and combining with the motion estimation method based on the geometric model, the visual motion state of the target can be estimated, usually including the translation vector and the rotation matrix. The translation vector describes the movement of the target in space, while the rotation matrix depicts the attitude change of the target. Through this process, the motion trajectory of the target can be accurately inferred in a complex dynamic environment, and on this basis, the next error correction and fusion can be carried out, thereby improving the accuracy of target state estimation. Finally, through visual-inertial calibration, the errors of the inertial navigation sensors of the target carrier are corrected. There will be certain errors in the inertial navigation sensors themselves, such as the bias of the accelerometer and the noise of the gyroscope. These errors will affect the estimation accuracy of the system for the target motion state. By combining the visual motion state variables of the target carrier with the data of the inertial navigation sensors and performing visual-inertial calibration, these errors can be effectively identified and compensated. The visual motion state variables provide the translation vector and the rotation matrix of the target. This information can be compared with the measurement results of the inertial navigation sensors. By optimizing the algorithm to model and correct the sensor errors, the errors of the inertial navigation sensors can be accurately compensated, thereby improving the positioning and navigation accuracy of the overall system and significantly enhancing the motion state tracking and navigation performance of the target carrier. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Other features, objects, and advantages of the present invention will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:
[0069] Figure 1 It is a schematic diagram of the step flow of the visual-based inertial navigation calibration method of the present invention;
[0070] Figure 2 For Figure 1 it is a detailed step flow schematic diagram of step S1 in
[0071] Figure 3 For Figure 1 it is a detailed step flow schematic diagram of step S2 in DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0073] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0074] It should be understood that although the terms "first", "second", etc. may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed associated items.
[0075] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a vision-based inertial navigation calibration method, and the method includes the following steps:
[0076] Step S1: Install and deploy corresponding vision sensors and inertial navigation sensors on the target carrier, and use the vision sensors and inertial navigation sensors to measure the target vision image data and target inertial navigation motion data corresponding to the target carrier in real time; add corresponding timestamps to align the target vision image data and target inertial navigation motion data in the time dimension to obtain corresponding target vision image sequences and target inertial navigation motion sequences in the same time dimension;
[0077] Step S2: Perform static and dynamic motion feature fusion analysis on the target vision images corresponding to each time point in the target vision image sequence to obtain the static and dynamic motion vision features of the target carrier;
[0078] Step S3: Conduct a nearest neighbor feature point pair search on the static and dynamic motion visual features of the target carrier to obtain the motion matching feature point pairs of the target carrier; based on the motion matching feature point pairs of the target carrier, perform visual motion estimation on the target carrier to obtain the visual motion state variables of the target carrier, including the visual motion translation vector of the target carrier and the visual motion rotation matrix of the target carrier.
[0079] Step S4: Based on the visual motion state variables of the target carrier and the target inertial navigation motion sequence, perform visual-inertial calibration on the error measurement corresponding to the inertial navigation sensor to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
[0080] In the embodiment of the present invention, please refer to Figure 1 As shown, it is a schematic diagram of the step flow of the visual-inertial calibration method based on vision of the present invention. In this example, the visual-inertial calibration method based on vision includes the following steps:
[0081] Step S1: Install and deploy the corresponding visual sensor and inertial navigation sensor on the target carrier, and use the visual sensor and inertial navigation sensor to measure the target visual image data and target inertial navigation motion data corresponding to the target carrier in real time; perform time dimension alignment on the target visual image data and target inertial navigation motion data by adding the corresponding time stamps to obtain the corresponding target visual image sequence and target inertial navigation motion sequence in the same time dimension.
[0082] In the embodiment of the present invention, for the target carrier, according to its structure and motion characteristics, the visual sensor (such as a binocular camera) and the inertial navigation sensor (such as a high-precision gyroscope and accelerometer) are respectively installed in appropriate positions. For the visual sensor, it is installed at the front end of the target carrier using a special bracket to ensure the field of view coverage; the inertial navigation sensor is installed at a position close to the center of gravity of the carrier and with little vibration to ensure the measurement accuracy. After installation, start the sensors to start real-time measurement. The visual sensor captures the surrounding environment of the target carrier at the set frame rate to generate the target visual image data; the inertial navigation sensor continuously measures information such as the acceleration and angular velocity of the carrier to obtain the target inertial navigation motion data. During the data acquisition process, when the synchronization signal generator sends a trigger signal, record the accurate time stamp and add it to the target visual image data and target inertial navigation motion data respectively. Subsequently, sort and match the data according to the time stamp. If the time stamps are inconsistent, fine-tune them using linear interpolation. Finally, obtain the target visual image sequence and target inertial navigation motion sequence in the same time dimension.
[0083] Step S2: Perform static and dynamic motion feature fusion analysis on the target visual image corresponding to each time point in the target visual image sequence to obtain the static and dynamic motion visual features of the target carrier.
[0084] In an embodiment of the present invention, for each target visual image corresponding to a time point in the target visual image sequence, the static and dynamic attributes are judged by comparing the images of adjacent time frames before and after this time point. By using Canny edge detection and gray-level co-occurrence matrix to analyze the changes in target edges and texture differences respectively, if there are differences, it is a dynamic visual image, otherwise it is a static visual image. For the static visual image, first perform Gaussian blur and downsampling to construct an image pyramid, then use Canny edge detection and morphological operations to extract the edge contour and regional shape, use Harris corner detection to extract key points, and finally use Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF) to analyze shape, texture and color features. For the dynamic visual image, construct a spatio-temporal cube, analyze the gray-level and gradient changes of adjacent frames, and based on this, perform motion cue statistics to obtain the motion direction and speed, combine the sparse optical flow method of the Lucas-Kanade algorithm to analyze the motion trajectory and speed changes, and finally associate and fuse the static and dynamic motion features to finally obtain the static and dynamic motion visual features of the target carrier.
[0085] Step S3: Perform a nearest neighbor feature point pair search on the static and dynamic motion visual features of the target carrier to obtain the motion matching feature point pairs of the target carrier; perform visual motion estimation on the target carrier based on the motion matching feature point pairs of the target carrier to obtain the visual motion state variables of the target carrier, including the visual motion translation vector of the target carrier and the visual motion rotation matrix of the target carrier;
[0086] In an embodiment of the present invention, by extracting feature points from the static and dynamic motion visual features of the target carrier, calculating the Euclidean distance between the feature points using a feature descriptor (such as the SIFT descriptor), setting a distance threshold, when the Euclidean distance is less than the threshold, the two feature points are determined as a pair of nearest neighbor feature points, search and match all feature points, and exclude unstable or ambiguous feature point pairs, so as to obtain the motion matching feature point pairs of the target carrier. Based on these matching feature point pairs, combining the principles of triangulation and epipolar geometry, analyze the changes in the triangle angles and side lengths formed by the feature points in adjacent frame images, determine the spatio-temporal association between the feature points, calculate the displacement and angle changes of the feature points between adjacent frames according to the spatio-temporal association, project the displacement onto the three spatial axes to obtain the motion speed, and statistically analyze the rotation angle changes. Construct the visual motion translation vector of the target carrier with the motion speed, and use the Rodriguez formula to transform according to the rotation axis and rotation angle to obtain the visual motion rotation matrix of the target carrier, so as to obtain the visual motion state variables of the target carrier.
[0087] Step S4: Perform visual-inertial calibration on the error measurement corresponding to the inertial navigation sensor based on the visual motion state variables of the target carrier and the target inertial motion sequence to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
[0088] In an embodiment of the present invention, by combining the target carrier visual motion state variables (including the visual motion translation vector and the visual motion rotation matrix) with the target inertial navigation motion sequence, the state vector of the extended Kalman filter is defined. The gyroscope drift and the accelerometer zero bias are used as state variables, and the visual measurement values in the target carrier visual motion state variables and the inertial navigation measurement values in the target inertial navigation motion sequence are used as observation variables. In the prediction step of the extended Kalman filter, the state variables at the next moment are predicted based on the dynamic model of the inertial navigation sensor. In the update step, the error covariance between the observation variables and the predicted state variables is calculated, the Kalman gain is calculated through the Kalman gain formula, and then the state variables are adjusted. The prediction and update steps are continuously repeated to achieve real-time estimation of the error measurement of the inertial navigation sensor. According to the estimation result, the error parameters of the inertial navigation sensor are compensated, and finally the corresponding error calibration compensation result of the inertial navigation sensor is obtained, such as the compensated gyroscope drift value and the accelerometer zero bias value.
[0089] Further, step S1 includes the following steps:
[0090] Step S11: Install and deploy corresponding visual sensors and inertial navigation sensors on the target carrier;
[0091] Step S12: Use the visual sensor and the inertial navigation sensor to perform synchronous real-time measurement on the target carrier in a hardware synchronous triggering manner through the built-in synchronous signal generator to obtain the target visual image data and the target inertial navigation motion data;
[0092] Step S13: Clean and standardize the target visual image data and the target inertial navigation motion data. For the target visual image data, contrast enhancement, noise removal, and pixel standardization are performed, while for the target inertial navigation motion data, duplicate data, outliers, missing value compensation, and motion parameter standardization are removed to obtain the target visual standard image data and the target inertial navigation standard data;
[0093] Step S14: Align the target visual standard image data and the target inertial navigation standard data in the time dimension by adding corresponding timestamps to obtain the corresponding target visual image sequence and target inertial navigation motion sequence in the same time dimension.
[0094] As an embodiment of the present invention, refer to Figure 2 shown in Figure 1 is the detailed step flow schematic diagram of step S1 in
[0095] Step S11: Install and deploy corresponding visual sensors and inertial navigation sensors on the target carrier;
[0096] In the embodiment of the present invention, by analyzing in detail the structure and motion characteristics of the target carrier, for a vision sensor, if the target carrier is a small unmanned aerial vehicle (UAV), considering the field of view requirements during its flight, a binocular camera with a wide-angle lens is selected and installed at the front end of the UAV using a special bracket, ensuring that the screws are tightened to guarantee a firm installation and an angle that can cover the main field of view in front of the UAV. For an inertial navigation sensor, such as a high-precision gyroscope and accelerometer, according to the center-of-gravity position and vibration characteristics of the UAV, it is installed at a position close to the center of gravity and with less vibration, and fixed using welding and glue to ensure that the sensor can accurately sense the motion state of the UAV.
[0097] Step S12: Using the vision sensor and the inertial navigation sensor, perform synchronous real-time measurement on the target carrier in a hardware synchronization trigger manner through the built-in synchronization signal generator to obtain target visual image data and target inertial navigation motion data.
[0098] In the embodiment of the present invention, by connecting the built-in synchronization signal generator to the vision sensor and the inertial navigation sensor through signal lines, a stable hardware connection structure is formed. The synchronization signal generator emits synchronization trigger signals at preset time intervals and frequencies. This signal is simultaneously transmitted to the vision sensor and the inertial navigation sensor. When the vision sensor receives the trigger signal, it immediately activates the shooting function to capture the visual image around the target carrier and stores the image data in the data storage module in the form of digital signals. At the same time, when the inertial navigation sensor receives the trigger signal, it starts measuring motion parameters such as the acceleration and angular velocity of the target carrier and records these data in real time, finally obtaining target visual image data and target inertial navigation motion data.
[0099] Step S13: Perform data cleaning and standardization on the target visual image data and the target inertial navigation motion data. For the target visual image data, perform contrast enhancement, noise removal, and pixel standardization, while for the target inertial navigation motion data, remove duplicate data, outliers, missing value compensation, and motion parameter standardization to obtain target visual standard image data and target inertial navigation standard data.
[0100] In the embodiments of the present invention, for the target visual image data, the histogram equalization algorithm is used to enhance the contrast of the image, making the brightness distribution of the image more uniform. The median filtering algorithm is adopted to remove the salt-and-pepper noise in the image. This algorithm replaces the gray value of each pixel point with the median of the gray values of the pixels in its neighborhood. Then, the pixel values of the image are normalized by subtracting the pixel mean and dividing by the pixel standard deviation, so that the pixel values conform to the standard normal distribution. For the target inertial navigation motion data, duplicate data is found and deleted through data comparison. Statistical methods, such as the method based on the standard deviation, are used to identify and remove outliers. For missing values, linear interpolation is used for compensation. Finally, the motion parameters are standardized so that their mean is 0 and the standard deviation is 1, and finally the target visual standard image data and the target inertial navigation standard data are obtained.
[0101] Step S14: Align the target visual standard image data and the target inertial navigation standard data in the time dimension by adding corresponding timestamps, so as to obtain the corresponding target visual image sequence and target inertial navigation motion sequence in the same time dimension.
[0102] In the embodiments of the present invention, during the data acquisition process, the synchronization signal generator records the accurate timestamp while sending out the trigger signal, and adds this timestamp to the corresponding target visual standard image data and target inertial navigation standard data respectively. During subsequent processing, based on the timestamp, the two groups of data are sorted and matched. If the timestamps are not exactly the same, the linear interpolation method is used to fine-tune the data, so that the target visual standard image data and the target inertial navigation standard data correspond one by one in the time dimension, and finally the corresponding target visual image sequence and target inertial navigation motion sequence in the same time dimension are obtained.
[0103] Furthermore, step S11 includes the following steps:
[0104] Conduct material and dynamics analysis on the target carrier to obtain the corresponding structural material, motion characteristics, and environmental characteristics of the target carrier;
[0105] In the embodiments of the present invention, the structural material of the target carrier is detected by using material analysis instruments. For example, a spectral analyzer is used to determine the elemental composition of each part of the carrier, and then its material type is clarified. Through dynamic testing equipment, the actual motion of the target carrier is simulated, and its acceleration, speed, displacement and other parameters are measured to analyze its motion characteristics, including the smoothness and periodicity of the motion. At the same time, the environment where the target carrier is located is monitored on-site. A temperature and humidity sensor is used to measure the temperature and humidity of the environment, a barometer is used to measure the air pressure, and an anemometer is used to measure the wind speed, etc., so as to obtain the temperature, humidity, air pressure, wind speed and other characteristics of the environment where the target carrier is located.
[0106] Preferably, a sensor adaptation planning analysis is performed on the visual sensor and the inertial navigation sensor based on the structural material, motion characteristics and environmental characteristics of the target carrier, so as to adapt and plan the type, accuracy requirements and protection level of the visual sensor and the inertial navigation sensor, and obtain the target carrier sensor adaptation planning result;
[0107] In the embodiment of the present invention, if the target carrier has a complex structure, a fast movement speed, and a large change in ambient light, and considering the need to accurately obtain three-dimensional spatial information, a binocular camera is planned to be used as a visual sensor, and its accuracy requirement is set to millimeter level, and the protection level is set to IP67 to adapt to complex environments. For inertial navigation sensors, due to the drastic changes in carrier movement, high-precision gyroscopes and accelerometers are selected, and the accuracy requirements reach 0.01° / h and 0.001m / s². The protection level is set to be able to withstand a certain degree of vibration and impact. If the target carrier moves relatively smoothly and the environment is relatively simple, a monocular camera combined with a laser radar can be planned as a visual sensor, and the accuracy and protection level can be determined according to actual needs. In this way, the sensor adaptation planning analysis is completed, and finally the target carrier sensor adaptation planning result is obtained.
[0108] Preferably, based on the target carrier sensor adaptation planning result, the corresponding visual sensor and inertial navigation sensor to be installed are subjected to sensor installation topology design, so as to determine the corresponding optimal installation angle and position of the visual sensor according to the change of the motion field of view corresponding to the target carrier, and ensure that the target area is covered to the greatest extent under various motion states, and the inertial navigation sensor is installed at the position with the least vibration and the easiest to sense motion according to the corresponding dynamic characteristics of the target carrier, so as to generate the target carrier sensor installation topology structure;
[0109] In an embodiment of the present invention, by using a computer to simulate the change of the field of view of the target carrier in different motion states for the visual sensor, for example, if the target carrier is an aircraft, its take-off, cruising, landing and other states are simulated, and the optimal installation angle and position of the visual sensor are determined by analyzing the simulation results. For example, the binocular camera is installed at the front end of the aircraft with an elevation angle of 30° and a horizontal angle of 60° to ensure that the front target area can be covered to the greatest extent under various flight attitudes. For the inertial navigation sensor, according to the dynamic model of the target carrier, the part with the smallest vibration and accurate perception of movement is found, such as installing a gyroscope and an accelerometer near the center of gravity of the aircraft, and finally a sensor installation topology structure of the target carrier is generated.
[0110] Preferably, the corresponding vision sensor and inertial navigation sensor are installed on the target carrier based on the target carrier sensor installation topology structure and are subjected to sensing calibration, so as to calibrate the internal parameters corresponding to the vision sensor in combination with the optical imaging principle, including the focal length and the distortion coefficient, and use the gravity and turntable calibration equipment to calibrate the zero bias and scale factor corresponding to the inertial navigation sensor according to the gravity field and the rotation motion law, so as to implement the calibration work corresponding to the vision sensor and the inertial navigation sensor.
[0111] In the embodiment of the present invention, according to the target carrier sensor installation topology structure, a dedicated installation tool is used to accurately install the vision sensor and the inertial navigation sensor on the target carrier. For the vision sensor, using the optical imaging principle, by shooting a standard checkerboard image and applying the Zhang Zhengyou calibration method to calculate the focal length and distortion coefficient of the vision sensor, and then adjusting and calibrating it. For the inertial navigation sensor, the target carrier is placed on the gravity and turntable calibration equipment, and according to the direction and magnitude of the gravity field and the rotation motion law of the turntable, the zero bias and scale factor of the gyroscope and accelerometer are measured and calculated, and then calibrated and adjusted, and finally the calibration work of the two sensors is realized.
[0112] Further, the vision sensor includes a binocular camera or a monocular camera combined with a lidar.
[0113] Further, the inertial navigation sensor includes a gyroscope and an accelerometer.
[0114] Further, step S12 includes the following steps:
[0115] A synchronous trigger topology structure is carried out between the synchronous signal generator built in the target carrier as the core node and the vision sensor and the inertial navigation sensor in a star structure, and the feedback signals corresponding to each sensor are aggregated to the synchronous signal generator by using a bus structure to realize two-way communication, so as to generate a target carrier hardware synchronous trigger topology structure;
[0116] In the embodiment of the present invention, in the target carrier, when the synchronous signal generator is used as the core node and connected in a star structure, independent lines are led out from the synchronous signal generator to connect the vision sensor and the inertial navigation sensor respectively, and each line has stable electrical characteristics to ensure the accuracy of signal transmission. At the same time, a bus structure is adopted to connect the feedback signal lines of the vision sensor and the inertial navigation sensor to the same bus, and then the bus is connected to the synchronous signal generator. In this way, the trigger signal sent by the synchronous signal generator can be independently and accurately transmitted to each sensor, and the feedback signal of the sensor can also be aggregated to the synchronous signal generator through the bus to realize two-way communication, and finally form a target carrier hardware synchronous trigger topology structure. For example, signal lines of specific specifications are used for connection to ensure the stability and anti-interference ability of signal transmission.
[0117] Preferably, based on the target vehicle hardware synchronous triggering topology, a hardware synchronous triggering method is adopted to send synchronous triggering signals to the corresponding vision sensor and inertial navigation sensor, so as to generate target vehicle synchronous triggering signal pulses.
[0118] In the embodiment of the present invention, after the target vehicle hardware synchronous triggering topology is built, the synchronous signal generator is started. The synchronous signal generator generates synchronous triggering signals according to a preset time interval and signal intensity. These signals are respectively transmitted to the vision sensor and inertial navigation sensor through a star-shaped line. The synchronous triggering signal has a specific pulse width and frequency to ensure accurate triggering of the sensor to work. For example, the pulse width of the signal is set to 100 microseconds and the frequency is 10 Hz. During the signal transmission process, the electrical characteristics of the line and signal enhancement equipment are used to ensure the signal intensity and stability, thereby generating target vehicle synchronous triggering signal pulses, so that the vision sensor and inertial navigation sensor can start working at the same moment.
[0119] Preferably, the target vehicle synchronous triggering signal pulse is used to control the corresponding vision sensor and inertial navigation sensor, and the transmission delay of the synchronous signal pulse from the generator to the sensor is measured. At the same time, based on the transmission delay, the corresponding measurement time is finely adjusted to perform synchronous real-time measurement on the target vehicle, so as to obtain target visual image data and target inertial navigation motion data.
[0120] In the embodiment of the present invention, after the target vehicle synchronous triggering signal pulse reaches the vision sensor and inertial navigation sensor, the sensor starts to respond and work. To measure the transmission delay of the synchronous signal pulse from the generator to the sensor, a high-precision time measurement device is set at the sensor end. When the sensor receives the signal pulse, the time is recorded and compared with the time when the signal generator sends the signal to calculate the transmission delay. For example, the signal generator sends a signal at time t1, and the vision sensor receives the signal at time t2. The transmission delay is t2 - t1. According to the measured transmission delay, the measurement time of the sensor is finely adjusted. If the transmission delay of the vision sensor is large, its measurement time is advanced by the corresponding delay time to ensure that the two sensors can perform synchronous measurement. After such synchronous real-time measurement, the vision sensor collects target visual image data, and the inertial navigation sensor collects target inertial navigation motion data.
[0121] Further, step S2 includes the following steps:
[0122] Step S21: Obtain the target visual images corresponding to the adjacent time frames before and after at each time point through the target visual images corresponding to the target visual image sequence at each time point.
[0123] Step S22: Perform target carrier change analysis on the target visual images corresponding to the adjacent time frames before and after at this time point, to compare and analyze the target edge and texture difference changes between the target visual images corresponding to this time point and the adjacent time frames before and after, so as to obtain the target visual image difference changes between each time point and the adjacent time frames before and after;
[0124] Step S23: Based on the target visual image difference changes between each time point and the adjacent time frames before and after, perform static and dynamic image judgment on the target visual images corresponding to the corresponding time points. If there are difference changes between each time point and the adjacent time frames before and after, determine the target visual image corresponding to this time point as a dynamic visual image; if there are no difference changes between each time point and the adjacent time frames before and after, determine the target visual image corresponding to this time point as a static visual image;
[0125] Step S24: Perform static motion feature analysis on the static visual images to obtain the static motion features of the target carrier; perform dynamic motion feature analysis on the dynamic visual images to obtain the dynamic motion features of the target carrier;
[0126] Step S25: Perform static and dynamic motion feature fusion analysis on the static motion features of the target carrier and the dynamic motion features of the target carrier to obtain the static and dynamic motion visual features of the target carrier.
[0127] As an embodiment of the present invention, referring to Figure 3 shown, for Figure 1 the detailed step flow schematic diagram of step S2 in
[0128] Step S21: Obtain the target visual images corresponding to the adjacent time frames before and after at this time point through the target visual images corresponding to each time point in the target visual image sequence;
[0129] In the embodiment of the present invention, the target visual image sequence is stored in chronological order, and each time point corresponds to one frame of target visual image. For a certain specific time point t in the sequence, the corresponding target visual image is I_t. Through time indexing, directly obtain the target visual image I_{t - 1} corresponding to the previous frame and the target visual image I_{t + 1} corresponding to the next frame at this time point. For example, if the target visual image sequence has a total of 100 frames, when the time point t = 50, through index operation, read the 49th frame image from the storage as I_{49}, the 50th frame image as I_{50}, and the 51st frame image as I_{51}, and obtain the target visual images corresponding to the adjacent time frames before and after at each time point in this way.
[0130] Step S22: Based on the target visual images corresponding to the adjacent time frames before and after at this time point, perform target carrier change analysis on the target visual images corresponding to the corresponding time point, so as to compare and analyze the target edge and texture difference changes between the target visual images at this time point and the adjacent time frames before and after, so as to obtain the target visual image difference changes between each time point and the adjacent time frames before and after;
[0131] In the embodiment of the present invention, for each time point t, the target visual image I_t is compared with I_{t - 1} and I_{t + 1}. First, the Canny edge detection algorithm is used to extract the target edges of I_t, I_{t - 1} and I_{t + 1} respectively. Then, the differences between the edge images are calculated. For example, by calculating the absolute value of the difference of the corresponding pixel positions, the number and distribution of the difference pixels are counted. For the texture difference, the gray-level co-occurrence matrix method is used to calculate the texture features of I_t, I_{t - 1} and I_{t + 1}, such as contrast, correlation, etc. Then, the differences of these texture feature values are compared. The edge difference and the texture difference are combined to finally obtain the target visual image difference changes between each time point and the adjacent time frames before and after.
[0132] Step S23: Based on the target visual image difference changes between each time point and the adjacent time frames before and after, perform static and dynamic image judgment on the target visual images corresponding to the corresponding time point. If there are difference changes between each time point and the adjacent time frames before and after, the target visual image corresponding to this time point is determined as a dynamic visual image; if there are no difference changes between each time point and the adjacent time frames before and after, the target visual image corresponding to this time point is determined as a static visual image;
[0133] In the embodiment of the present invention, by setting an edge difference threshold and a texture difference threshold, for each time point t, if the number of edge difference pixels between the target visual image I_t and I_{t - 1}, I_{t + 1} exceeds the edge difference threshold, or the difference of the texture feature values exceeds the texture difference threshold, it is determined that there are difference changes between this time point and the adjacent time frames before and after, and I_t is determined as a dynamic visual image. On the contrary, if neither the number of edge difference pixels nor the difference of the texture feature values exceeds the corresponding threshold, it is determined that there are no difference changes, and I_t is determined as a static visual image. For example, the edge difference threshold is set to 100 pixels, and the texture feature difference threshold is set to 0.1. When the number of edge difference pixels between I_t and I_{t - 1} is 150, then I_t is a dynamic visual image.
[0134] Step S24: Perform static motion feature analysis on the static visual image to obtain the static motion features of the target carrier; perform dynamic motion feature analysis on the dynamic visual image to obtain the dynamic motion features of the target carrier.
[0135] In the embodiment of the present invention, for the static visual image, first perform Gaussian blur and downsampling operations on it to construct image pyramids of different scales. At each scale, use Canny edge detection and morphological operations to extract the target edge contour and regional shape, and use the Harris corner detection algorithm to extract corners and edge intersections to obtain a key point set. Then, use the Scale-Invariant Feature Transform (SIFT) and Speeded-Up Robust Features (SURF) algorithms to analyze the shape, texture, and color features of the target carrier to obtain the static motion features of the target carrier. For the dynamic visual image, construct it into a spatio-temporal cube, analyze the gray-scale and gradient changes of adjacent frames in the time dimension, and perform motion cue statistical analysis based on these changes to obtain the dynamic motion direction and speed of the target carrier. Combine the sparse optical flow method based on the Lucas-Kanade algorithm to analyze the motion trajectory and speed change of the target carrier, and finally obtain the dynamic motion features of the target carrier.
[0136] Step S25: Perform static and dynamic motion feature fusion analysis on the static motion features of the target carrier and the dynamic motion features of the target carrier to obtain the static and dynamic motion visual features of the target carrier.
[0137] In the embodiment of the present invention, by fusing the shape, texture, and color features in the static motion features of the target carrier with the motion direction, speed, motion trajectory and other features in the dynamic motion features of the target carrier. For the shape feature, analyze whether the target shape changes during the dynamic motion process. If there is a change, record the change situation and rules. For the texture feature, observe the change situation of the texture during the dynamic motion, such as the flow or deformation of the texture. In terms of the color feature, statistically analyze the distribution and change of the color during the dynamic process, and associate the dynamic motion information with the static feature information. For example, combine the motion trajectory with the shape and color features of the target to comprehensively analyze the state and changes of the target carrier at different times, and finally obtain the static and dynamic motion visual features of the target carrier.
[0138] Further, the static motion feature analysis of the static visual image described in step S24 includes the following steps:
[0139] Perform Gaussian blur and downsampling operations on the static visual image to obtain the corresponding target static visual sub-images at each scale;
[0140] In the embodiment of the present invention, for a static visual image, Gaussian filtering is performed using a Gaussian filter with a kernel size set to 5x5 and a standard deviation of 1.0. The Gaussian filter is applied to each pixel of the image, and new pixel values are calculated through a convolution operation to smooth the image and reduce the influence of noise. Then, a downsampling operation is performed, adopting an interlaced sampling method, that is, one pixel is selected every other row and column, reducing the size of the image to half of the original. The above Gaussian blur and downsampling processes are repeated to construct image pyramids of different scales. For example, if the initial image is 512x512 pixels, a sub-image of 256x256 pixels is obtained after the first processing, and a sub-image of 128x128 pixels is obtained after the second processing, and so on, finally obtaining the corresponding target static visual sub-images at each scale.
[0141] Preferably, target edge structure feature analysis is performed on the target static visual sub-images corresponding to each scale to extract the target edge contour and regional shape corresponding to the static visual image by using edge detection and morphological operations, obtaining the target carrier static visual edge structure features corresponding to each scale;
[0142] In the embodiment of the present invention, for the target static visual sub-images of each scale, the Canny edge detection algorithm is used for edge detection. First, Gaussian smoothing is performed on the image, then the gradient magnitude and direction of the image are calculated, followed by non-maximum suppression to remove non-edge pixels, and finally, true edge pixels are determined through double-threshold processing. After obtaining the edge image, morphological operations are used for optimization. The dilation operation is used to fill small gaps in the edge to make the edge more continuous; then the erosion operation is used to remove small noise points around the edge. By finding connected regions, the edge contour and regional shape of the target are extracted. For example, for a sub-image containing a circular target, the circular edge contour and its regional shape can be accurately extracted through these operations, finally obtaining the target carrier static visual edge structure features corresponding to each scale.
[0143] Preferably, based on the target carrier static visual edge structure features corresponding to each scale, static visual key point extraction is performed on the target static visual sub-images corresponding to each scale to extract the corresponding corner points and edge intersection points, obtaining the key point set of the target carrier static visual image;
[0144] In an embodiment of the present invention, by using the Harris corner detection algorithm to extract corners based on the static visual edge structure features of the target carrier at each scale, the Harris corner detection algorithm calculates the gray-scale change rate of the image in different directions to find areas with large gray-scale changes, and these areas are the corners. For the extraction of edge intersection points, it is determined by analyzing the intersection of edges in the edge image. Specifically, each pixel in the edge image is traversed, and the edge situation of its surrounding pixels is checked. When it is found that multiple edges intersect here, this point is marked as an edge intersection point. The corners and edge intersection points extracted at each scale are summarized to obtain the key point set of the static visual image of the target carrier. For example, 20 corners and 10 edge intersection points are extracted from a sub-image at one scale, and such points at all scales are combined to form a complete key point set, and finally the key point set of the static visual image of the target carrier is obtained.
[0145] Preferably, based on the key point set of the static visual image of the target carrier and using scale-invariant feature transform and speeded-up robust features to perform static motion feature analysis on the static visual image to obtain the static motion features of the target carrier, including the shape, texture, and color features corresponding to the target carrier.
[0146] In an embodiment of the present invention, based on the key point set of the static visual image of the target carrier, scale-invariant feature transform (SIFT) and speeded-up robust features (SURF) algorithms are used to perform static motion feature analysis. For the SIFT algorithm, a local area is constructed around each key point, the gradient magnitude and direction within this area are calculated, and a feature descriptor is generated. The SURF algorithm extracts features around the key point by calculating the integral image of the image. Through these feature descriptors, the shape features of the target carrier are analyzed. For example, it is judged whether the target is circular, rectangular, or other shapes by the distribution and connection relationship of the feature points. For the texture features, the change law of the gradient in the feature descriptor is analyzed to judge features such as the thickness and density of the texture. In terms of color features, combined with the pixel color information at the position where the key point is located, the distribution and proportion of colors are statistically analyzed, so as to obtain the shape, texture, and color features corresponding to the target carrier, and finally the static motion features of the target carrier are obtained.
[0147] Further, the dynamic motion feature analysis of the dynamic visual image described in step S24 includes the following steps:
[0148] Each dynamic visual image is regarded as a spatio-temporal cube, and the gray-scale change and gradient change between adjacent frame dynamic visual images are analyzed in the time dimension;
[0149] In an embodiment of the present invention, a series of dynamic visual images are arranged in chronological order to construct a spatio-temporal cube, where the spatial dimension is composed of the length and width of the images, and the time dimension is reflected by the number of frames of the images. For two adjacent frames of dynamic visual images in the time dimension, first calculate the gray-scale difference of their corresponding pixel points. For example, for the previous frame image and the subsequent frame image at the same coordinates subtract the pixel gray-scale values to obtain the gray-scale change value . At the same time, use the Sobel operator to calculate the gradients of the two frames of images in the horizontal and vertical directions respectively, and then calculate the difference of the gradients at the corresponding positions between adjacent frames to analyze the gradient change situation. Through these operations, comprehensively understand the gray-scale and gradient change characteristics of adjacent frame images in the time dimension.
[0150] Preferably, based on the gray-scale change and gradient change between adjacent frame dynamic visual images, perform motion cue statistical analysis on each dynamic visual image to obtain the dynamic motion direction of the target carrier and the dynamic motion speed of the target carrier;
[0151] In an embodiment of the present invention, through motion cue statistics based on the gray-scale change and gradient change of adjacent frame dynamic visual images, regions with large gray-scale changes often indicate obvious motion of the target carrier. For the gradient change, its direction and magnitude can reflect the motion trend of the target. By statistically analyzing the gray-scale and gradient changes of multiple regions in the image, determine the approximate motion direction of the target carrier. For example, if the gray-scale value of a certain region continuously increases and the gradient direction points to a specific direction, it can be judged that the target moves in that direction. When calculating the motion speed, use the amplitude of the gray-scale change and the time interval to estimate. Assume that the time interval between adjacent frames is , and the gray-scale change of a certain region is . By establishing an empirical model ( is a coefficient determined according to experiments) between the gray-scale change and the motion speed, calculate the motion speed of the target carrier in different regions, and comprehensively obtain the overall dynamic motion speed of the target carrier based on the situations of each region.
[0152] Preferably, based on the dynamic motion direction of the target carrier and the dynamic motion speed of the target carrier and combined with the sparse optical flow method based on the Lucas-Kanade algorithm, perform dynamic motion feature analysis on the dynamic visual image corresponding to the target carrier to obtain the dynamic motion features of the target carrier, including the motion trajectory and speed change corresponding to the target carrier during the motion process.
[0153] In the embodiment of the present invention, based on the dynamic motion direction and dynamic motion speed of the target carrier, the sparse optical flow method based on the Lucas-Kanade algorithm is used to analyze the dynamic vision image. First, some representative feature points, such as corner points, are selected in the initial frame image. Then, according to the assumption of the Lucas-Kanade algorithm, that is, the pixels in the local area have the same motion, the gray information of the adjacent frame images around these feature points is used to calculate the optical flow. The optical flow represents the motion vector of the feature points between adjacent frames, and its direction and magnitude respectively correspond to the motion direction and speed of the target. By tracking the motion of these feature points in multiple frame images and connecting the positions of the feature points in each frame, the motion trajectory of the target carrier is obtained. At the same time, the change of the optical flow magnitude of each feature point in different frames is recorded to analyze the speed change of the target carrier, and finally the dynamic motion characteristics of the target carrier are comprehensively obtained.
[0154] Further, step S3 includes the following steps:
[0155] Step S31: Search for the nearest neighbor feature point pairs of the static and dynamic motion vision features of the target carrier to obtain the motion matching feature point pairs of the target carrier;
[0156] In the embodiment of the present invention, by extracting the static and dynamic motion vision feature points from the multi-frame image sequence of the target carrier, for example, using the SIFT (Scale-Invariant Feature Transform) algorithm to extract the feature points and their feature descriptors. For each feature point in each frame image, search for its nearest neighbor feature point in the adjacent frame image, calculate the Euclidean distance of the feature descriptors between the feature points. The smaller the distance, the more similar the features. Set a distance threshold. When the Euclidean distance is less than this threshold, these two feature points are determined as a pair of nearest neighbor feature points. Perform such search and matching operations on all feature points, excluding those feature point pairs with unstable or ambiguous matches, and finally obtain the motion matching feature point pairs of the target carrier. For example, in an image frame containing 1000 feature points, after matching and screening, 500 pairs of motion matching feature point pairs are obtained.
[0157] Step S32: Based on the motion matching feature point pairs of the target carrier and combined with triangulation and the principle of epipolar geometry, perform spatio-temporal feature point correlation analysis on the target carrier to analyze the triangle angle and side length change relationship formed by the feature points in the adjacent frame images of the target carrier, and obtain the spatio-temporal correlation relationship between the feature points of the target carrier;
[0158] In the embodiment of the present invention, by using the target carrier motion-matching feature point pairs and combining the principle of triangulation, for the matching feature point pairs in adjacent frame images, according to the internal and external parameters of the camera, a triangle composed of the feature points and the camera optical center is constructed. By using the known camera parameters and the pixel coordinates of the feature points in the image, the side lengths and angles of the triangle are calculated. At the same time, based on the principle of epipolar geometry, the epipolar constraint relationship between the feature points in adjacent frame images is determined, that is, the epipolar line where the feature points appear in another frame image. The changes in the angles and side lengths of the triangles formed by the feature points in adjacent frame images are analyzed, such as the increase or decrease of the side lengths and the increase or decrease of the angles. According to these change relationships, the spatio-temporal correlation relationship between the feature points is established, and the corresponding relationship of the feature points in different times and spaces is clarified, such as how a certain feature point moves and changes in adjacent frames. Finally, the spatio-temporal correlation relationship between the target carrier feature points is obtained.
[0159] Step S33: Based on the spatio-temporal correlation relationship between the target carrier feature points, perform statistical analysis on the motion speed and angle of the motion trajectory corresponding to the target carrier, so as to obtain the motion speed and motion rotation angle corresponding to the target carrier on the three spatial axes.
[0160] In the embodiment of the present invention, by determining the displacement and angle changes of each feature point between adjacent frames according to the spatio-temporal correlation relationship between the target carrier feature points, for each feature point, its position change between adjacent frames is calculated, and combined with the time interval, the motion speed of the feature point is obtained. The motion speeds of all feature points are projected onto the three spatial axes (X, Y, Z axes), and statistical analysis is performed on the projected speed values, such as calculating the average value, to obtain the motion speed of the target carrier on the three spatial axes. At the same time, the rotation situation of the feature points between adjacent frames is analyzed. By calculating the direction change of the feature points, the rotation angle is obtained. Statistical analysis is performed on the rotation angles of all feature points to obtain the motion rotation angle of the target carrier in space. For example, after analyzing the motion of 100 feature points, the average motion speed of the target carrier on the X axis is 0.5 m / s, on the Y axis is 0.3 m / s, on the Z axis is 0.2 m / s, and the motion rotation angle is 10 degrees. Finally, the motion speed and motion rotation angle corresponding to the target carrier on the three spatial axes are obtained.
[0161] Step S34: Based on the motion speed and motion rotation angle corresponding to the target carrier on the three spatial axes, perform visual motion estimation on the target carrier to obtain the visual motion state variables of the target carrier, including the visual motion translation vector and the visual motion rotation matrix of the target carrier.
[0162] In the embodiments of the present invention, by constructing a visual motion translation vector of the target carrier according to the motion speeds corresponding to the target carrier on the three spatial axes, and taking the motion speeds on the X, Y, and Z axes as the three components of the translation vector respectively. For example, the translation vector , where , , are the motion speeds on the X, Y, and Z axes respectively. For the motion rotation angle of the target carrier, according to the rotation axis and the rotation angle, it is converted into a rotation matrix using the Rodriguez formula. The rotation axis can be determined by analyzing the rotation direction of the feature points, and the rotation angle is the previously obtained motion rotation angle. Through these calculations, the visual motion state variables of the target carrier are obtained, that is, the visual motion translation vector of the target carrier and the visual motion rotation matrix of the target carrier, which are used for subsequent inertial navigation calibration and other operations. Among them, the visual motion translation vector of the target carrier is specifically:
[0163] ;
[0164] ;
[0165] ;
[0166] ;
[0167] Among them, represents the visual motion translation vector of the target carrier, represents the translation component along the X axis, represents the translation component along the Y axis, represents the translation component along the Z axis, represents the motion measurement time, represents the motion speed corresponding to the target carrier on the X axis, represents the motion speed corresponding to the target carrier on the Y axis, represents the motion speed corresponding to the target carrier on the Z axis; Further, the visual motion rotation matrix of the target carrier is specifically:
[0168] The rotation matrix corresponding to the rotation around the X axis is: ;
[0169] The rotation matrix corresponding to the rotation around the Y axis is: ;
[0170] The rotation matrix corresponding to the rotation around the Z axis is: ;
[0171] Among them, represents the motion rotation angle corresponding to the target carrier on the X axis, represents the motion rotation angle corresponding to the target carrier on the Y axis, It represents the corresponding motion rotation angle of the target carrier on the Z-axis, and finally obtains the visual motion state variables of the target carrier.
[0172] Furthermore, step S31 includes the following steps:
[0173] Perform motion vision depth deconstruction on the static and dynamic motion vision features of the target carrier. For static features, use Fourier transform and wavelet transform to analyze their frequency components and spatial distribution characteristics. For dynamic features, establish a time series model to analyze the corresponding change trends at different time points, so as to obtain the static and dynamic feature deconstruction data of the target carrier.
[0174] In the embodiment of the present invention, for the static features of the target carrier, the visual image of the target carrier is regarded as a two-dimensional signal, and Fourier transform is performed on it to convert the image from the spatial domain to the frequency domain. By analyzing the spectrogram, the distribution of different frequency components is determined. For example, the low-frequency components may correspond to the overall contour of the image, and the high-frequency components correspond to the detailed texture of the image. Then, wavelet transform is used to decompose the image into sub-bands of different scales and directions, and further analyze the distribution characteristics of static features at different spatial positions, such as the spatial distribution of features such as edges and corners. For dynamic features, collect a sequence of visual images of the target carrier over a period of time, use the position, gray value, etc. of feature points as the observed values of the time series, establish an autoregressive integrated moving average (ARIMA) model, and predict the feature change trends at future time points based on historical observed values. For example, by analyzing the position changes of feature points at different time points, determine the motion speed and direction changes of the target carrier, and finally organize and obtain the static and dynamic feature deconstruction data of the target carrier.
[0175] Preferably, based on the static and dynamic feature deconstruction data of the target carrier, construct a spatio-temporal encoded feature point index for each feature point in the static and dynamic motion vision features of the target carrier. In terms of spatial encoding, use the encoding method corresponding to the Hilbert curve to map the three-dimensional spatial coordinates of the feature point into an encoding, and in terms of time encoding, generate the corresponding time encoding according to the time sequence of the appearance of the feature point, so as to obtain the static and dynamic visual index features of the target carrier after spatio-temporal encoding.
[0176] In the embodiment of the present invention, by deconstructing data based on the static and dynamic characteristics of the target carrier to determine the three-dimensional spatial coordinates of each feature point, for spatial coding, the Hilbert curve coding method is adopted to divide the three-dimensional space into several small cube grids, each small cube corresponding to a unique Hilbert coding value. According to the three-dimensional coordinates of the feature point, the small cube where it is located is determined, and the Hilbert coding value of this small cube is used as the spatial coding of the feature point. In terms of time coding, the time point when the feature point first appears in the image sequence is recorded, and an increasing integer is assigned to each feature point in chronological order as the time coding. For example, the time coding of the first feature point to appear is 1, the second is 2, and so on. The spatial coding and time coding are combined to obtain the static and dynamic visual index features of the target carrier after spatio-temporal coding, which facilitates the subsequent rapid retrieval and analysis of feature points.
[0177] Preferably, according to the static and dynamic visual index features of the target carrier after spatio-temporal coding, the feature point neighborhood is divided, and the Delaunay triangulation part is used to divide the corresponding static and dynamic visual feature points of the target carrier into corresponding triangular neighborhoods. For feature points located at the edges and corners of the target carrier and with obvious feature changes, large-scale neighborhoods are defined, while for general feature points, small-scale neighborhoods are defined, obtaining the multi-scale feature point neighborhood division result of the target carrier.
[0178] In the embodiment of the present invention, by using the spatial coding information in the static and dynamic visual index features of the target carrier after spatio-temporal coding, the three-dimensional spatial coordinates of the feature point are obtained, and the Delaunay triangulation algorithm is used to connect all feature points into non-overlapping triangles, such that no other feature points are contained within the circumcircle of each triangle. For feature points located at the edges and corners of the target carrier, by analyzing the distribution and changes of the surrounding feature points, it is determined that their feature changes are obvious. For example, the gray value changes significantly around the edge feature points, and the gradient changes significantly in multiple directions at the corner feature points. Large-scale neighborhoods are defined for these feature points, such as a circular area centered on this feature point with a radius of 5 feature point spacings. For general feature points, that is, feature points with relatively gentle surrounding feature changes, small-scale neighborhoods are defined, such as a circular area centered on this feature point with a radius of 2 feature point spacings, finally obtaining the multi-scale feature point neighborhood division result of the target carrier.
[0179] Preferably, based on the multi-scale feature point neighborhood division result of the target carrier, the nearest neighbor candidate points are screened for the feature points corresponding to the internal corresponding scale neighborhoods of the static and dynamic motion visual features of the target carrier. By calculating the energy values and visual similarities corresponding to each feature point within the feature point neighborhood and screening out the feature points with high energy values and high visual similarities as the nearest neighbor candidate points, the motion matching feature point pairs of the target carrier are obtained.
[0180] In an embodiment of the present invention, according to the multi-scale feature point neighborhood division result of the target carrier, for the corresponding scale neighborhood of each feature point, the energy value and visual similarity between each feature point in the neighborhood are calculated. For the calculation of the energy value, the change rate of the gray value of the feature point is used as a measure of energy. For example, the gradient amplitude of the gray value within a certain range around the feature point is calculated. The larger the gradient amplitude, the higher the energy value. For the calculation of visual similarity, the Euclidean distance between the feature descriptors of the feature points (such as SIFT descriptors) is used as a measure. The smaller the Euclidean distance, the higher the visual similarity. For each feature point, feature points with an energy value higher than a set threshold (such as the top 30% in terms of energy value ranking) and a visual similarity higher than a set threshold (such as the top 30% in terms of Euclidean distance ranking) in its neighborhood are selected as the nearest neighbor candidate points. These nearest neighbor candidate points are combined in pairs to finally obtain the target carrier motion matching feature point pairs for subsequent motion analysis and inertial navigation calibration.
[0181] Further, step S4 includes the following steps:
[0182] Step S41: Analyze the dual-sensor motion parameter characteristics of the target carrier visual motion state variables and the target inertial navigation motion sequence to obtain a dual-sensor motion parameter characteristic data set;
[0183] In an embodiment of the present invention, for the visual sensor, based on the motion trajectory of the target carrier in the actual scene. For example, in a set rectangular motion trajectory, the translation vectors of the visual sensor in the X, Y, and Z directions are recorded. The real motion trajectory is obtained through a high-precision measurement device, and the translation error range between the measured translation vectors of the visual sensor and the real trajectory in different directions is calculated. For the rotation matrix, during the rotation in different directions, the actual rotation angles are recorded using a precise angle measurement instrument and compared with the rotation angles corresponding to the rotation matrix output by the visual sensor to obtain the rotation angle accuracy. For the inertial navigation sensor, using dynamic models such as Newton's second law, the gyroscope and accelerometer are installed on a test bench simulating different motion states (such as uniform linear motion, accelerated linear motion, uniform circular motion, etc.). In each motion state, the angular velocity measured by the gyroscope and the acceleration measured by the accelerometer are recorded. Through long-term observation and data analysis, the zero-bias stability of the gyroscope, that is, the drift situation of the measured value in the stationary state, is evaluated; the scale factor error of the accelerometer is calculated. For example, in the case of a known standard acceleration, the difference between the measured value of the accelerometer and the standard value is compared. Finally, a dual-sensor motion parameter characteristic data set is compiled.
[0184] Step S42: Perform dual-sensor data fusion on the target carrier's visual motion state variables and the target inertial navigation motion sequence based on the dual-sensor motion parameter characteristic dataset, so as to utilize the motion measurements of the inertial navigation sensor within a short period of time to assist in fusing the motion parameter estimation of the visual sensor, and obtain a dual-sensor motion parameter fusion vector;
[0185] In the embodiment of the present invention, during the dual-sensor data fusion process, since the inertial navigation sensor has high measurement accuracy within a short period of time, taking a single acceleration linear motion of the target carrier within a short period of time as an example, the accelerometer of the inertial navigation sensor can quickly and accurately measure the acceleration value, and the gyroscope can also accurately measure the angular velocity. According to the dual-sensor motion parameter characteristic dataset, the error conditions of the translation vector and rotation matrix of the visual sensor in this short-time rapid motion state are understood. The acceleration and angular velocity information measured by the inertial navigation sensor is converted into displacement and angle information through kinematic formulas. For example, using integral operations, the displacement increment within a short period of time is calculated based on the acceleration value measured by the accelerometer, and the current position is obtained by combining the initial position information. Then, this position and angle information are fused with the translation vector and rotation matrix measured by the visual sensor. By using the weighted average method, different weights are assigned to the data of the inertial navigation sensor and the visual sensor according to the accuracy performance of each sensor in the dual-sensor motion parameter characteristic dataset under different motion states. For example, when moving rapidly within a short period of time, the data weight of the inertial navigation sensor is set to 0.7, and the data weight of the visual sensor is set to 0.3. After weighted fusion of the two data, a dual-sensor motion parameter fusion vector is finally obtained.
[0186] Step S43: Perform visual-inertial calibration on the error measurement corresponding to the inertial navigation sensor based on the dual-sensor motion parameter fusion vector by using the extended Kalman filter, taking the error parameters corresponding to the inertial navigation sensor as state variables, including gyroscope drift and accelerometer zero bias, and taking the measurement values of the visual sensor and the inertial navigation sensor in the dual-sensor motion parameter fusion vector as observation variables. At the same time, the state variables are continuously updated through the filtering algorithm to realize real-time estimation and calibration compensation of the error measurement corresponding to the inertial navigation sensor, so as to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
[0187] In the embodiment of the present invention, by defining the state vector of the extended Kalman filter, the gyroscope drift and accelerometer zero bias are used as state variables. For example, the state vector is set as , where is the gyroscope drift, is the accelerometer zero bias. The measured values of the vision sensor (such as the translation vector, rotation matrix) and the inertial navigation sensor (such as acceleration, angular velocity) in the dual-sensor motion parameter fusion vector are used as the observation variables. In the prediction step of the extended Kalman filter, according to the dynamic model of the inertial navigation sensor, the state variables at the next moment are predicted. For example, according to the error models of the gyroscope and accelerometer, the gyroscope drift and accelerometer zero bias at the next moment are predicted. In the update step, the error covariance between the observation variables and the predicted state variables is calculated, and the state variables are adjusted through the Kalman gain. For example, using the Kalman gain formula , where is the predicted error covariance,[[]]END]] is the observation matrix,[[]]END]] is the observation noise covariance, and the Kalman gain is calculated. Then, the updated state variables are calculated according to the above parameters, where is the observed value. By continuously repeating the prediction and update steps, the real-time estimation of the error measurement of the inertial navigation sensor is realized, and finally the error calibration compensation result corresponding to the inertial navigation sensor is obtained, such as the compensated gyroscope drift value and accelerometer zero bias value.[[]]END]]
[0188] Furthermore, the process of analyzing the characteristics of the dual-sensor motion parameters in step S41 is specifically as follows: for the translation vector corresponding to the vision sensor, analyze the translation error range between it and the motion trajectory in different directions; for the rotation matrix, analyze the rotation angle accuracy corresponding to it in different directions; and for the motion measurement parameters corresponding to the inertial navigation sensor, use the dynamic model to analyze the response characteristics of the angular velocity measured by the gyroscope and the acceleration measured by the accelerometer in different motion states, and evaluate its zero bias stability and scale factor error.[[]]END]]
[0189] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the application documents within the present invention.[[]]END]]
[0190] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features invented herein.[[]]END]]
Claims
1. A vision-based inertial navigation calibration method, characterized in that: The following steps are involved: Step S1: installing and deploying corresponding visual sensors and inertial navigation sensors on the target carrier, and using the visual sensors and inertial navigation sensors to measure the target visual image data and target inertial navigation motion data corresponding to the target carrier in real time; aligning the target visual image data and the target inertial navigation motion data in time dimension by adding corresponding timestamps, so as to obtain the corresponding target visual image sequence and target inertial navigation motion sequence in the same time dimension; Step S2: Perform static and dynamic motion feature fusion analysis on the target visual image corresponding to the target visual image sequence at each time point to obtain static and dynamic motion visual features of the target carrier; Step S3: searching for the nearest neighbor feature point pairs of the static and dynamic motion visual features of the target carrier to obtain the target carrier motion matching feature point pairs; Based on the target carrier motion matching feature point pair, the target carrier visual motion is estimated to obtain the target carrier visual motion state variable, which includes the target carrier visual motion translation vector and the target carrier visual motion rotation matrix; wherein step S3 includes the following steps: Step S31: searching for the nearest neighbor feature point pairs of the static and dynamic motion visual features of the target carrier to obtain the target carrier motion matching feature point pairs; Step S32: performing a temporal and spatial feature point correlation analysis on the target carrier based on the target carrier motion matching feature point pairs and combining triangulation and epipolar geometry principles to analyze the relationship between the angle and side length changes of the triangles formed by the feature points in adjacent frame images of the target carrier, and obtain the temporal and spatial correlation relationship between the feature points of the target carrier; Step S33: performing a statistical analysis of the movement speed and angle of the movement trajectory corresponding to the target carrier based on the spatiotemporal correlation between the characteristic points of the target carrier, so as to obtain the movement speed and movement rotation angle corresponding to the target carrier on the three spatial axes; Step S34: performing visual motion estimation on the target carrier based on the corresponding motion speed and motion rotation angle of the target carrier on the three axes of space to obtain the visual motion state variable of the target carrier, including the visual motion translation vector of the target carrier and the visual motion rotation matrix of the target carrier; Step S4: Based on the visual motion state variables of the target carrier and the target inertial navigation motion sequence, visual-inertial navigation calibration is performed on the error measurement corresponding to the inertial navigation sensor to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
2. The vision-based inertial navigation calibration method according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: installing and deploying corresponding visual sensors and inertial navigation sensors on the target carrier; Step S12: Using a built-in synchronization signal generator and a hardware synchronization trigger mode to synchronously measure the target carrier using a visual sensor and an inertial navigation sensor to obtain target visual image data and target inertial navigation motion data; Step S13: cleaning and standardizing the target visual image data and the target inertial navigation motion data, so as to enhance contrast, remove noise and standardize pixels for the target visual image data, and remove duplicate data, outliers, compensate for missing values and standardize motion parameters for the target inertial navigation motion data, so as to obtain target visual standard image data and target inertial navigation standard data; Step S14: aligning the target visual standard image data and the target inertial navigation standard data in time dimension by adding corresponding timestamps to obtain the corresponding target visual image sequence and target inertial navigation motion sequence in the same time dimension.
3. The vision-based inertial navigation calibration method according to claim 2, characterized in that: Step S11 includes the following steps: Conduct material and dynamic analysis on the target carrier to obtain the corresponding structural material, motion characteristics and environmental characteristics of the target carrier; Based on the structural material, motion characteristics and environmental characteristics of the target carrier, the sensor adaptation planning analysis of the visual sensor and the inertial navigation sensor is carried out to adapt the type, accuracy requirements and protection level of the visual sensor and the inertial navigation sensor to obtain the target carrier sensor adaptation planning results; Based on the target carrier sensor adaptation planning results, the corresponding visual sensors and inertial navigation sensors to be installed are designed for sensor installation topology, so as to determine the corresponding optimal installation angle and position of the visual sensor according to the change of the motion field of view of the target carrier, and ensure that the target area is covered to the greatest extent under various motion states. For the inertial navigation sensor, it is installed at the position with the least vibration and the easiest to sense motion according to the corresponding dynamic characteristics of the target carrier, so as to generate the target carrier sensor installation topology structure; Based on the target carrier sensor installation topology, the corresponding visual sensors and inertial navigation sensors are installed on the target carrier and sensor calibration is performed. The corresponding internal parameters of the visual sensors, including focal length and distortion coefficient, are calibrated in combination with the optical imaging principle. Gravity and turntable calibration equipment is used to calibrate the corresponding zero bias and scale factor of the inertial navigation sensor according to the gravity field and the law of rotational motion, so as to realize the corresponding calibration work of the visual sensors and inertial navigation sensors.
4. The vision-based inertial navigation calibration method according to claim 3, characterized in that: The visual sensor includes a binocular camera or a monocular camera combined with a laser radar.
5. The vision-based inertial navigation calibration method according to claim 3, characterized in that: The inertial navigation sensor includes a gyroscope and an accelerometer.
6. The vision-based inertial navigation calibration method according to claim 2, characterized in that: Step S12 includes the following steps: The built-in synchronization signal generator of the target carrier is used as the core node to synchronize the trigger topology with the visual sensor and the inertial navigation sensor in a star structure, and the feedback signals corresponding to each sensor are aggregated to the synchronization signal generator using a bus structure to achieve two-way communication, so as to generate the hardware synchronization trigger topology of the target carrier; Based on the hardware synchronous trigger topology of the target carrier, a hardware synchronous trigger method is used to send a synchronous trigger signal to the corresponding visual sensor and inertial navigation sensor to generate a target carrier synchronous trigger signal pulse; The target carrier is synchronized with the trigger signal pulse response to control the corresponding visual sensor and inertial navigation sensor, and the transmission delay of the synchronization signal pulse from the generator to the sensor is measured. At the same time, the corresponding measurement time is fine-tuned based on the transmission delay to perform synchronous real-time measurement of the target carrier to obtain the target visual image data and target inertial navigation motion data.
7. The vision-based inertial navigation calibration method according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: obtaining the target visual images corresponding to the adjacent time frames before and after the time point through the target visual image corresponding to the target visual image sequence at each time point; Step S22: performing target carrier change analysis between target visual images corresponding to the corresponding time points based on the target visual images corresponding to the previous and next adjacent time frames at the time point, so as to compare and analyze the target edge and texture difference changes between the target visual images corresponding to the previous and next adjacent time frames at the time point, so as to obtain the target visual image difference changes between each time point and the previous and next adjacent time frames; Step S23: Based on the difference change of the target visual image between each time point and the adjacent time frames, the target visual image corresponding to the corresponding time point is judged as a static or dynamic image. If there is a difference change between each time point and the adjacent time frames, the target visual image corresponding to the time point is determined as a dynamic visual image; if there is no difference change between each time point and the adjacent time frames, the target visual image corresponding to the time point is determined as a static visual image. Step S24: performing static motion feature analysis on the static visual image to obtain static motion features of the target carrier; performing dynamic motion feature analysis on the dynamic visual image to obtain dynamic motion features of the target carrier; Step S25: performing static and dynamic motion feature fusion analysis on the static motion features of the target carrier and the dynamic motion features of the target carrier to obtain static and dynamic motion visual features of the target carrier.
8. The vision-based inertial navigation calibration method according to claim 7, characterized in that: The static motion feature analysis of the static visual image in step S24 includes the following steps: Perform Gaussian blur and down-sampling operations on the static visual image to obtain the corresponding target static visual sub-image at each scale; The target edge structure feature analysis is performed on the target static visual sub-image corresponding to each scale, so as to extract the target edge contour and regional shape corresponding to the static visual image by edge detection and morphological operation, and obtain the static visual edge structure feature of the target carrier corresponding to each scale; Based on the static visual edge structure features of the target carrier at each scale, the static visual key points of the target static visual sub-image at each scale are extracted to extract the corresponding corner points and edge intersections, and obtain the key point set of the static visual image of the target carrier; Based on the key point set of the static visual image of the target carrier, the static motion feature analysis of the static visual image is performed by using scale-invariant feature transformation and accelerated robust features to obtain the static motion features of the target carrier, including the shape, texture and color features corresponding to the target carrier.
9. The vision-based inertial navigation calibration method according to claim 7, characterized in that: The dynamic motion feature analysis of the dynamic visual image described in step S24 includes the following steps: Treat each dynamic visual image as a space-time cube, and analyze the grayscale changes and gradient changes between adjacent frames of dynamic visual images in the time dimension; Based on the grayscale changes and gradient changes between the dynamic visual images of adjacent frames, the motion clues between the dynamic visual images are statistically analyzed to obtain the dynamic motion direction and the dynamic motion speed of the target carrier; Based on the dynamic motion direction and dynamic motion speed of the target carrier and combined with the sparse optical flow method based on the Lucas-Kanade algorithm, the dynamic motion characteristics of the target carrier corresponding to the dynamic visual image are analyzed to obtain the dynamic motion characteristics of the target carrier, including the corresponding motion trajectory and speed changes of the target carrier during the movement.
10. The vision-based inertial navigation calibration method according to claim 1, characterized in that: Step S31 includes the following steps: Performing motion visual deep deconstruction on the static and dynamic visual features of the target carrier, using Fourier transform and wavelet transform to analyze the frequency components and spatial distribution characteristics of static features, and establishing a time series model to analyze the corresponding change trends at different time points for dynamic features, so as to obtain the static and dynamic feature deconstruction data of the target carrier; Based on the static and dynamic feature deconstruction data of the target carrier, a spatiotemporal coding feature point index is constructed for each feature point in the static and dynamic motion visual feature of the target carrier, so that the three-dimensional spatial coordinates of the feature point are mapped into a code using a coding method based on the Hilbert curve in spatial coding, and the corresponding time code is generated according to the time sequence of the feature points in time coding, so as to obtain the static and dynamic visual index feature of the target carrier after spatiotemporal coding; According to the static and dynamic visual index features of the target carrier after spatiotemporal coding, the neighborhood of the feature points is divided, so that the corresponding static and dynamic visual feature points of the target carrier are divided into corresponding triangular neighborhoods based on the Delaunay triangulation part. For the feature points at the edge and corner of the target carrier and with obvious feature changes, a large-scale neighborhood is defined, and for the general feature points, a small-scale neighborhood is defined, so as to obtain the neighborhood division result of the multi-scale feature points of the target carrier; Based on the multi-scale feature point neighborhood division result of the target carrier, the feature points corresponding to the corresponding scale neighborhood within the static and dynamic motion visual features of the target carrier are screened for the nearest neighbor candidate points, so as to obtain the target carrier motion matching feature point pairs by calculating the corresponding energy values and visual similarities between each feature point in the feature point neighborhood.
11. The vision-based inertial navigation calibration method according to claim 1, characterized in that: The target carrier visual motion translation vector described in step S34 is specifically: ; ; ; ; in, represents the target carrier visual motion translation vector, represents the translation component along the X axis, represents the translation component along the Y axis, represents the translation component along the Z axis, represents the motion measurement time, Indicates the movement speed of the target carrier on the X-axis. Indicates the movement speed of the target carrier on the Y axis. Indicates the movement speed of the target carrier on the Z axis.
12. The vision-based inertial navigation calibration method according to claim 1, characterized in that: The target carrier visual motion rotation matrix described in step S34 is specifically: The corresponding rotation matrix around the X axis is: ; The corresponding rotation matrix around the Y axis is: ; The corresponding rotation matrix around the Z axis is: ; in, Indicates the corresponding motion rotation angle of the target carrier on the X-axis. Indicates the corresponding motion rotation angle of the target carrier on the Y axis. Indicates the corresponding movement rotation angle of the target carrier on the Z axis.
13. The vision-based inertial navigation calibration method according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: performing dual-sensor motion parameter characteristic analysis on the target carrier visual motion state variables and the target inertial navigation motion sequence to obtain a dual-sensor motion parameter characteristic data set; Step S42: performing dual sensor data fusion on the visual motion state variables of the target carrier and the target inertial navigation motion sequence based on the dual sensor motion parameter characteristic data set, so as to use the motion measurement of the inertial navigation sensor in a short time to assist in estimating the motion parameters corresponding to the fusion visual sensor, and obtain a dual sensor motion parameter fusion vector; Step S43: Based on the dual-sensor motion parameter fusion vector and using the extended Kalman filter, the error measurement corresponding to the inertial navigation sensor is calibrated with vision-inertial navigation, so that the error parameters corresponding to the inertial navigation sensor are used as state variables, including gyroscope drift and accelerometer zero bias, and the measurement values of the visual sensor and the measurement values of the inertial navigation sensor in the dual-sensor motion parameter fusion vector are used as observation variables. At the same time, the state variables are continuously updated through the filtering algorithm to realize real-time estimation and calibration compensation of the error measurement corresponding to the inertial navigation sensor, so as to obtain the error calibration compensation result corresponding to the inertial navigation sensor.
14. The vision-based inertial navigation calibration method according to claim 13, characterized in that: The process of dual-sensor motion parameter characteristic analysis described in step S41 is specifically to analyze the translation error range between the translation vector corresponding to the visual sensor and the motion trajectory in different directions, analyze the rotation angle accuracy corresponding to the rotation matrix in different directions, and use the dynamic model to analyze the response characteristics of the angular velocity measured by the gyroscope and the acceleration measured by the accelerometer under different motion states for the motion measurement parameters corresponding to the inertial navigation sensor, and evaluate its zero bias stability and scale factor error.
Citation Information
Patent Citations
Method and apparatus for calibrating parameters of visual-inertial system, and electronic device and medium
WO2022100189A1
Visual inertial odometry method that contains self-calibration and is based on keyframe sliding window filtering
WO2023155258A1