A Coupling Method, Device and Equipment for Multiple Sensors
Through feature extraction and photometric calibration of multi-sensor data, the problem of insufficient accuracy and stability of multi-sensor system mapping is solved, efficient tight coupling and multi-scale map construction are achieved, and real-time mapping construction capabilities are improved.
Patent Information
- Application Number
- CN202510494566.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The existing multi-sensor systems have significant differences in data characterization in real-time mapping and positioning, high real-time requirements, unstable accuracy of the fusion algorithm, and the sensor calibration process relies on hardware and is susceptible to error accumulation, resulting in insufficient mapping accuracy and stability.
By acquiring multi-sensor data, extracting image features, point cloud features and IMU features, performing photometric calibration and spatial synchronization processing, building a factor graph model for tight coupling, and realizing multi-scale map construction.
It improves the mapping accuracy and coupling speed of multi-sensor systems, reduces the burden on the central processor, and enhances the real-time mapping capability in complex scenarios.
Smart Images

Figure CN120027782B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of real-time mapping and positioning, and particularly to a coupling method, device, and equipment for multi-sensors. Background Art
[0002] With the continuous development of laser detection and ranging technology, high-precision mapping technology has become a key link. In the processes of aerial navigation, geophysical surveying and mapping, and autonomous vehicle navigation, etc., integrating multiple sensors such as visual SLAM (visual sensor) and 3D lidar through artificial intelligence algorithms to construct a map can provide targeted solutions for various fields such as disaster management, energy exploration, transportation facilities, ecological protection, and robotics.
[0003] In this development background, the prior art usually performs multi-modal data fusion based on multiple sensors such as visual SLAM, lidar, and IMU (inertial measurement unit), improving the mapping accuracy. Although the multi-sensor system improves the reliability of data acquisition compared with the single-sensor system or the single-lidar system and solves the problem of single-sensor failure. However, the multi-sensor system also has problems such as significant differences in data representation, high real-time requirements, and unstable accuracy of fusion algorithms. In some complex scenarios, especially in real-time dynamic scenarios, there are still problems such as insufficient mapping and positioning accuracy and poor robustness.
[0004] In addition, data calibration and photometric calibration are required for multi-sensors before fusion. The prior art usually uses pulse sensors for time calibration, which has a strong dependence on hardware. However, any tiny error in the calibration process will be accumulated and amplified during long-term operation, affecting the accuracy and stability of subsequent mapping. Traditional photometric calibration relies on calibration objects and cannot be applied to cameras of different models. If an unknown camera is used to shoot videos or the environmental light changes significantly, it may take a long time for photometric calibration or even photometric calibration cannot be performed. How to perform efficient data processing and tight coupling among multi-sensors, and how to make full use of the advantages of each sensor to achieve high-precision and real-time mapping is still in the process of continuous exploration. Summary of the Invention
[0005] Embodiments of the present application provide a multi-sensor coupling method, apparatus, and device, which solve the problems in the prior art that the calibration and coupling of multiple sensors are inaccurate and the mapping accuracy is relatively low. By acquiring multi-sensor data and performing preprocessing, camera images, point cloud data, and IMU data are obtained. An image pyramid is constructed to detect feature points in the camera image, image features are obtained, photometric calibration is performed, a photometric model is constructed to achieve photometric consistency of the camera image, and key image features are extracted therefrom. According to the region segmentation method, the point cloud data is classified to obtain point cloud features. Feature points of the IMU data are extracted and jointly used with the key image features and point cloud features for subsequent construction of a factor graph model for multi-sensor coupling to complete the construction of a multi-scale map. It realizes real-time and efficient tight coupling of multiple sensors, constructs a multi-scale map, provides interaction between the user and the robot, and solves the problems in the prior art that real-time mapping cannot be performed in complex scenarios and the accuracy is relatively low.
[0006] In a first aspect, embodiments of the present application provide a multi-sensor coupling method, including: acquiring multi-sensor data; wherein, the multi-sensor data includes camera images, point cloud data, and IMU data; extracting feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features; performing photometric calibration on the image features to obtain an optimized camera image and extracting key image features therefrom; performing spatial synchronization processing on the key image features, the point cloud features, and the IMU features to obtain tightly coupled multi-sensor data; and constructing a multi-scale map based on the tightly coupled multi-sensor data.
[0007] In a possible implementation manner, the extracting feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features includes: extracting the image features of the camera image based on a feature point detection algorithm; performing region division on the point cloud data according to the region segmentation method to obtain the point cloud features; preprocessing the IMU data and extracting features therefrom to obtain the IMU features.
[0008] In a possible implementation manner, the extracting the image features of the camera image based on a feature point detection algorithm includes: constructing an image pyramid of the camera image by downsampling to determine key points therein; determining the directions of the key points according to the first moment; and rotating the sampling window by an angle according to the directions of the key points to construct a feature region to obtain the image features.
[0009] In a possible implementation, the region division of the point cloud data according to the region division method to obtain the point cloud features includes: dividing the point cloud data according to a preset threshold to obtain segmentation blocks; determining the curvature value of the point cloud data in the segmentation blocks; classifying the point cloud data based on the curvature value to obtain corner points and plane points; and performing downsampling processing on the plane points to obtain the point cloud features.
[0010] In a possible implementation, the photometric calibration of the image features includes: constructing a bidirectional optical flow tracking chain based on the image features to determine feature correspondence; determining the initial parameters of the camera image according to the camera imaging principle, where the initial parameters include a response function, vignetting, and exposure time; constructing a photometric model based on the feature correspondence and the initial parameters; defining an energy function based on the photometric model to obtain an optimized photometric model; and obtaining an optimized camera image based on the optimized photometric model and extracting key image features therefrom.
[0011] In a possible implementation, the constructing a photometric model based on the feature correspondence and the initial parameters includes: constructing a camera response function model, performing dimensionality reduction processing on the camera response function to determine the principal component vector; extracting the brightness attenuation coefficient of the camera image and fitting a vignetting model using a radial attenuation model; modeling the exposure time through an affine transfer function to obtain an exposure time model; and constructing a photometric model based on the camera response function model, the vignetting model, and the exposure time model.
[0012] In a possible implementation, the spatial synchronization processing of the key image features, the point cloud features, and the IMU features to obtain tightly coupled multi-sensor data includes: performing spatial calibration processing on the key image features, the point cloud features, and the IMU features to construct a factor graph model; and performing tight coupling optimization according to the factor graph model to obtain tightly coupled multi-sensor data.
[0013] In a possible implementation, the performing tight coupling optimization according to the factor graph model includes: initializing the factor nodes in the factor graph model; calculating the residuals and Jacobian matrices of the factor nodes to construct a normal equation; and iteratively updating the normal equation until the norm of the residuals is less than a preset iteration threshold and / or the update value of the updated nodes is less than a preset iteration threshold to obtain updated nodes.
[0014] In a second aspect, an embodiment of the present application provides a coupling device for multiple sensors, including: a data acquisition module that acquires multi-sensor data; wherein the multi-sensor data includes camera images, point cloud data, and IMU data; a feature extraction module that extracts feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features; a photometric calibration module that performs photometric calibration on the image features to obtain an optimized camera image and extracts key image features therefrom; a feature coupling module that performs spatial synchronization processing on the key image features, the point cloud features, and the IMU features to obtain tightly coupled multi-sensor data; and a mapping module that constructs a multi-scale map based on the tightly coupled multi-sensor data.
[0015] In a third aspect, an embodiment of the present application provides a device for executing a coupling method for multiple sensors. The device includes: a processor; a memory for storing processor-executable instructions; when the processor executes the executable instructions, the method as described in the first aspect or any possible implementation manner of the first aspect is implemented.
[0016] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0017] By adopting a coupling method, device, and device for multiple sensors, the embodiments of the present application solve the problems of inaccurate real-time calibration and coupling between multiple existing sensors and low mapping accuracy. Feature extraction is performed on multi-sensor data to obtain image features, point cloud features, and IMU features, reducing the amount of calculation and resource consumption for subsequent mapping. Photometric calibration is performed on the image features, and a two-way optical flow tracking chain is constructed to reversely optimize the photometric model to achieve photometric consistency processing and extract key image features therefrom. Spatial synchronization processing is performed on the key image features, point cloud features, and IMU features to obtain tightly coupled multi-sensor data, and finally a multi-scale map is constructed. The features of different types of sensors are efficiently extracted, improving the final coupling speed and accuracy, reducing the burden on the central processing unit, and improving the mapping accuracy. Description of the Drawings
[0018] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required for describing the embodiments of the present application or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a flowchart of a coupling method for multiple sensors provided by an embodiment of the present application;
[0020] Figure 2Schematic diagram of a coupling device for multiple sensors provided by an embodiment of the present application. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0022] The following explanations are made for some technologies involved in the embodiments of the present application to facilitate understanding. It should be considered that they are only exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, the descriptions of some well-known functions and structures are omitted below.
[0023] Figure 1 is a flowchart of a coupling method for multiple sensors provided by an embodiment of the present application, including steps 101 to 105. Figure 1 It is only an execution order shown in the embodiment of the present application and does not represent the only execution order of the coupling method for multiple sensors. Under the condition that the final result can be achieved, Figure 1 the steps shown can be executed in parallel or reversed, as follows.
[0024] Step 101: Obtain multi-sensor data; wherein, the multi-sensor data includes camera images, point cloud data, and IMU data.
[0025] In the embodiment of the present application, the multiple sensors include, but are not limited to, cameras, 3D lidars, and IMU sensors, and the multiple sensors can be installed on a robot. Continuous images or videos are captured by the camera to capture environmental information. The camera can select 100 digital cameras of different models, such as monocular and binocular cameras, such as GMSL2 cameras. And different frame rates, resolutions, and exposure times are set to adapt to different lighting conditions, and about 15,210 images are collected in total as camera images, including the internal parameters and distortion coefficients of the camera. The 3D lidar is used to obtain three-dimensional spatial information, usually stored in the form of original data packets, including information such as the rotation angle, distance value, and reflection intensity of the laser beam. The distance between the object and the sensor is measured by emitting laser pulses and receiving the reflected signals, and through preprocessing steps such as denoising, thinning, and coordinate transformation, point cloud data is generated. The IMU (Inertial Measurement Unit) is used to provide the motion state information of the object to obtain IMU data. Exemplarily, an IMU sensor of BMI088 can be selected (the acceleration is ±16 , with an angular velocity of ±2000 ). Among them, the IMU data includes the real-time attitude and acceleration of the target object. In addition, when using cameras, lidars, and IMUs, they need to be calibrated in advance to ensure the accuracy of the collected data.
[0026] In a possible implementation, the multi-sensor data is first subjected to time synchronization processing to ensure that the synchronized data can accurately reflect the environmental states of each sensor at the same moment. Since abnormal situations such as frame loss or timestamp errors may occur during the sensor transmission process, it is necessary to arrange the multi-sensor data in chronological order to construct a time index. According to the acquisition moment of the 3D lidar, find the corresponding position in the timeline of other sensors and lock the data of the previous and next frames.
[0027] Specifically, since each sensor has an independent clock source, the prior art usually aligns the timestamps of each sensor at the initial moment for time synchronization data processing. And a pulse generator is added to trigger all sensors by the pulse generator, and the current clock is corrected each time it is triggered to eliminate the cumulative error of the clock source and unify the clock source at the hardware level. Or use existing sensors such as GNSS (Global Navigation Satellite System) with built-in pulse generators to correct according to the satellite atomic clock to further improve the accuracy. However, in actual applications, the acquisition times of each sensor are different. Exemplarily, in Kitti (a computer vision algorithm evaluation dataset), both the lidar data and the IMU data are at the ten-level, such as 20Hz. However, the acquisition time of the 3D lidar is often dozens of milliseconds slower than that of the IMU each time, and the IMU data is usually more high-frequency, such as 200Hz. The time difference will accumulate through continuous acquisition, resulting in the problem of time difference drift.
[0028] In the embodiment of the present application, all sensor clocks are aligned through PTP (Precision Time Protocol) to achieve alignment at the sensor hardware level. The IMU data (200Hz) is interpolated to the timestamp of the camera image (30Hz) to complete the synchronization of high-frequency data to low-frequency data, and then the point cloud data (10Hz) is extrapolated and predicted to complete the interpolation alignment of low-frequency to high-frequency data.
[0029] Specifically, taking the received point cloud data as the reference data, when the point cloud data is received once, the acquisition moment of the current point cloud data is used as the pre-inserted time point, so as to obtain the equivalent information of other sensors except the 3D lidar at the same moment as the time point. The data of the previous and next frames before and after the time point are obtained by establishing a time index, and linear interpolation calculation is performed according to the acquisition moments of the previous and next frames of data and the time point to obtain the equivalent information and complete the time synchronization alignment.
[0030] Step 102: Extract feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features.
[0031] Based on the feature point detection algorithm, extract the image features of the camera image. In the embodiment of the present application, according to the feature point detection algorithm, extract the features of the camera image to obtain image features. Among them, the feature point detection algorithm includes:
[0032] Construct an image pyramid of the camera image through downsampling and determine the key points therein.
[0033] Exemplarily, the feature point detection algorithm can use ORB (a computer vision algorithm). Construct a Gaussian image pyramid with 7 levels, and perform downsampling at a ratio of 1 / 2 for the detection of scale-space key points. Use FAST (a fast corner detection algorithm) to determine the threshold parameters of the feature points from different levels of the image pyramid. Among them, the threshold parameter can be taken as 9. Since the gray value can intuitively reflect the feature that the pixel value changes sharply from light to dark, the gray value is divided by the threshold parameter to obtain the N feature points with the strongest response, and they are selected as the key points. Through the ORB algorithm, the key points of the camera images with different sizes in the image pyramid can be efficiently determined, thereby achieving partial scale invariance.
[0034] Determine the direction of the key point according to the first moment.
[0035] In the neighborhood centered on the key point and with the scale as the radius, use the first moment (intensity centroid method) to calculate the direction of the key point according to the direction calculation formula. Among them, the scale is the Gaussian scale and can be taken as 1.5 , represents the scale parameter and can be taken as 1.6. The direction calculation formula is as follows:
[0036] ;
[0037] ;
[0038] In the formula, represents the coordinates of the centroid, , , represents the first moment, represents the angle between the key point and the centroid, used to represent the direction of the key point, represents taking the tangent.
[0039] Among them, the coordinates of the centroid represent the position where the average intensity of the gray value is taken in the neighborhood. By making a vector from the key point to the centroid, the direction of the key point can be obtained.
[0040] The sampling window is rotated by an angle according to the direction of the key point to construct a feature region and obtain image features. Regardless of the orientation of the target camera image, the ORB algorithm can create the same feature vector for the key point, detect the same key point in the camera image rotated at any angle, and achieve rotation invariance. In the defined sampling window around a given key point, 256 pairs of pixel points are randomly selected to construct a 256-dimensional bit vector. Then, these randomly selected pixel point pairs are rotated according to the direction angle of the key point so that the direction of the randomly selected pixel point pairs is consistent with the direction of the key point. Finally, the brightness of the randomly selected pixel point pairs is compared and 1 and 0 are assigned accordingly to create the corresponding feature vector. The order of 1 and 0 varies according to the specific key point and the pixel region around it, representing the intensity pattern around the key point. Therefore, multiple feature vectors can be used to identify a larger region, that is, the feature region. In the feature region, a feature descriptor is generated, and the image features of the camera image are extracted.
[0041] The point cloud data is partitioned into regions according to the region segmentation method to obtain point cloud features.
[0042] Specifically, according to a preset threshold, the point cloud data is partitioned into regions to obtain segmentation blocks. In the embodiments of the present application, according to the characteristics of the point cloud data, a preset threshold for height can be set. The points with height values less than the preset threshold are initially partitioned into the segmentation blocks of the road point cloud region, and the points with height values equal to or greater than the preset threshold are partitioned into the segmentation blocks of the non-road point cloud region. In the embodiments of the present application, since the accuracy of the sensor is relatively high, the preset threshold can be taken as -0.05 to 0.1 meters for differentiation. In addition, those skilled in the art can also use a dynamic adjustment strategy to automatically increase or decrease the preset threshold according to the point cloud density to adapt to different environments.
[0043] The curvature value of the point cloud data is determined in the segmentation block. Specifically, in the segmentation block, each point in the point cloud data is traversed, and based on the sum of the squares of the distance differences between the current point and the 5 points before and after it, and this result is used as the curvature value of the point. By initializing the point cloud label of the point, it is divided into an unprocessed point set and an unclassified point set, and the curvature value and index of each point are stored in cloudSmoothness (point cloud smooth structure) for subsequent construction of an index for sorting. The markOccludedPoints (feature point occlusion marking function) is used to mark the occluded points and the points parallel to the light beam to avoid affecting the accuracy of feature extraction. Here, being parallel to the light beam means that if the distance differences between the current point and the 5 points before and after it are greater than a preset ratio, it is considered that the current point is parallel to the light beam and is marked as processed at the same time.
[0044] Classify the point cloud data based on the curvature values to obtain corner points and planar points. Extract the points in the segmentation block through the extractFeatures (corner point and planar point extraction function), and then classify the points according to the curvature values.
[0045] Specifically, take the point with the highest curvature value as the starting point for judgment until all the points in the unprocessed point set and the unclassified point set are processed. If the starting point is not occluded and the curvature value is greater than the edge threshold, mark it as a corner point and add it to the corner point cloud set. If the curvature value is less than the surface threshold, mark it as a planar point and add it to the planar point cloud set. Among them, the edge threshold can be taken as the average value of the point cloud density, and the surface threshold is taken as 0.1. To avoid the influence of noise points, set an extraction upper limit, and extract at most 20 corner points and planar points respectively each time. Mark the points adjacent to the corner points and planar points as processed points from the unprocessed points to prevent the points from being selected repeatedly.
[0046] Perform downsampling processing on the planar points to obtain point cloud features. In addition, the point cloud features can also be described through the FPFH (Fast Point Feature Histogram) algorithm, and combined with geometric and intensity information to generate point cloud feature descriptors. This reduces the redundant data of the point cloud and improves the subsequent processing efficiency.
[0047] Preprocess the IMU data and extract the features therein to obtain IMU features.
[0048] Specifically, use a low-pass filter, such as a Butterworth filter, to remove the high-frequency noise in the IMU data, and calculate the zero bias value according to the bias estimation value. Integrate the angular velocity and acceleration in the coordinate system to calculate the relative rotation angle, velocity and displacement, and then calculate statistical quantities such as their mean, variance and peak value for feature statistics. Extract the frequency domain information therein through the FFT (Fast Fourier Transform) algorithm to obtain IMU features.
[0049] Step 103: Perform photometric calibration on the image features to obtain an optimized camera image, and extract the key image features therein.
[0050] Construct a bidirectional optical flow tracking chain based on the image features to determine the feature correspondence.
[0051] In the embodiments of the present application, based on the key points and feature descriptors in the image features, a two-way optical flow tracking chain is constructed using Lucas-Kanade (dense optical flow estimation method). Since the directions of the key points have been obtained using the ORB algorithm, the optical flow vectors of the key points between consecutive frames can be directly estimated. Through forward optical flow tracking, the positions of these key points in the next frame can be predicted, and the key points can be mapped from the previous frame to the next frame. By adding the optical flow vector of each key point to the position of the corresponding key point in the previous frame, the inter-frame motion prior is realized, and a forward optical flow tracking chain from front to back is formed. On the basis of forward optical flow tracking, reverse optical flow tracking is performed on the predicted key point positions, tracking backward from the next frame to the previous frame, and a reverse optical flow tracking chain from back to front is formed. At the same time, the accuracy of forward optical flow tracking is verified. By comparing the results of forward and reverse optical flow tracking, incorrect tracking points can be eliminated, false matches can be reduced, and reverse verification can be completed. The results of forward and reverse optical flow tracking are combined to perform two-way optical flow tracking, and a two-way optical flow tracking chain is constructed.
[0052] Specifically, the accuracy of target tracking can be verified through feature correspondence. By setting a matching threshold, in the embodiments of the present application, the matching threshold is set to 1-3 pixels. Each key point is traversed. If the results of forward and reverse optical flow tracking match, it is considered that the tracking of this key point is accurate, and it is included in the tracking chain. If the results of forward and reverse optical flow tracking do not match, it is considered that the tracking of this key point may be incorrect; the position error between the forward and reverse tracking results is calculated by Euclidean distance. If the position error is less than the set matching threshold, this key point is included in the tracking chain; if the position error is greater than the set matching threshold, this key point is eliminated. For each key point that meets the matching conditions, the results of its forward and reverse optical flow tracking are regarded as the corresponding positions of the same key point in different frames, so that the feature correspondence can be confirmed, and thus the effective tracking of the target can be realized, and the accuracy and robustness of the tracking can be improved.
[0053] According to the camera imaging principle, the initial parameters of the camera image are determined; among them, the initial parameters include the response function, vignetting, and exposure time.
[0054] In the embodiments of the present application, according to the camera imaging principle, the light emitted by the light source onto the target object will be reflected onto the camera lens. After passing through the lens, the light intensity will change and then be transmitted to the camera's image sensor, where energy accumulation occurs within a certain period of time, and the corresponding illumination intensity is obtained through the processing of the response function. Based on DSO (a computer vision algorithm of the sparse direct method), the initial parameters can be initially determined. The light reflected from the object surface is usually called radiance, while the light emitted onto the camera sensor is usually called irradiance. The exposure time is usually stored in the camera image in the metadata form of EXIF (Exchangeable Image File Format). EXIF records various setting information of the camera during shooting, including aperture value, exposure time, ISO sensitivity, focal length, shooting time, etc.
[0055] Based on the feature correspondence relationship and the initial parameters, a photometric model is constructed.
[0056] According to the initial parameters, the camera response function, vignetting function, and exposure time function are respectively constructed. The radiance is mapped to the image intensity to construct the photometric model. The brightness attenuation coefficient of the camera image is extracted, and the vignetting model is fitted using the radial attenuation model. The exposure time is modeled through the affine transfer function to obtain the exposure time model. Based on the camera response function model, vignetting model, and exposure time model, the photometric model is constructed. Specifically as follows.
[0057] ; where
[0058] ,
[0059] ,
[0060] ;
[0061] In the formula, represents the gray value corresponding to the pixel at position on the camera image collected at time , that is, the radiance, represents the set of pixel position coordinates in the camera image, represents the th acquisition time, represents the camera response function, represents the exposure time at time represents the vignetting function, represents the radiance, 、 represent the vignetting coefficients, which are two-dimensional real vectors, 、 represent the radial radius at position , represents the camera average response function, represents the number of principal component vectors, represents the response coefficient vector, represents the -th principal component vector, represents the exposure time parameter at time represents the photometric offset at time represents the mathematical constant.
[0062] Among them, the vignetting coefficient is used to describe the attenuation law of the brightness of the camera image with the radial radius, . The exposure time parameter is used to describe the change of the sensitivity of the camera to light with time, and can take seconds. The photometric offset is used for photometric correction. Considering that in practical applications, the camera will also have a certain output signal in the absence of light, set , and adjust the brightness of the camera image through the photometric offset. Among them, when tends to 0, it means that no brightness adjustment is performed.
[0063] Specifically, since the two-way optical flow tracking chain has screened out the key points from the pixel points of the camera image, the calculation amount can be reduced to the greatest extent, and the feature correspondence can be established, which can assist in constructing the photometric model more accurately and quickly. The vignetting function is used to correct the vignetting aperture effect and eliminate the brightness difference between different exposure images. By establishing a normalized coordinate system, such as , the normalized radial radius , assuming that the vignetting is symmetric about the center, extract the brightness attenuation coefficient of each pixel point in the camera image, and approximately fit the vignetting function with the radial attenuation model to complete the vignetting modeling of the camera. The exposure function is used to respond at different exposure times. Since some automatic exposure cameras may not be able to obtain the accurate exposure time, establish the parameter model of the exposure time function through the affine transfer function . Among them, the exponential form is to prevent the exposure time from being negative and improve the calculation efficiency at the same time.
[0064] Construct the response function model, perform dimensionality reduction processing on the response function, and determine the principal component vectors. Realize reducing the data dimension while reducing the calculation amount.
[0065] The camera response function is used to convert the light intensity received by the camera sensor into pixel values, and construct the camera response function matrix, where each row of the data matrix represents a camera response function , in PCA (Principal Component Analysis), SVD (Singular Value Decomposition) is used for dimensionality reduction to determine the camera response function. Through SVD decomposition, it is possible to directly obtain the feature matrix after dimensionality reduction without calculating matrices with complex structures and long computation times such as the covariance matrix. The SVD decomposition is as follows.
[0066] ;
[0067] In the formula, represents the camera response function, and represent the left and right singular vector orthogonal matrices respectively, represents the singular value diagonal matrix, represents the transpose of the orthogonal matrix.
[0068] Among them, the singular vector orthogonal matrix can capture the principal components of the row and column spaces. The diagonal elements of the singular value diagonal matrix are arranged in descending order to reflect the importance of each component. Retain the first 4 principal component vectors with the largest amount of information , to capture the non-linear response and achieve dimensionality reduction while retaining key information.
[0069] Define the energy function based on the photometric model to obtain the optimized photometric model. The definition of the energy function is as follows:
[0070] ;
[0071] In the formula, represents the energy function, represents the set of pixel position coordinates in the camera image, represents the valid region of the camera image, represents the predicted radiance of the camera image at the pixel point at position , represents the vignetting function, represents the radiance of the actually captured camera image at the pixel point at position , represents the camera response function, represents the regularization coefficient, Take , represents the regularization term to prevent overfitting.
[0072] ;
[0073] In the formula, represents the predicted camera image, that is, the camera image after photometric correction, represents the inverse function of the camera response function, represents the actually captured camera image, Indicates the exposure time at the moment, indicates the vignetting function, indicates the th moment collected.
[0074] The energy function optimizes the response function and the vignetting function, and solves the response function and the vignetting function through an iterative method, minimizing the error between the observed values of all pixels and the predicted pixel values of the photometric model, so as to make the photometric model accurately fit the actual imaging process to the greatest extent and ensure the photometric consistency of the camera image.
[0075] Calculate the derivatives of the response function, vignetting, and exposure time through the Jacobian matrix to obtain the optimal photometric parameters. Inverse calibrate the photometric model to obtain the first photometric model. The specific calculation of the Jacobian matrix is as follows.
[0076] ;
[0077] , ;
[0078] In the formula, represents the camera response function, represents the inverse function of the camera response function, represents the response coefficient vector, represents the th principal component vector, represents the predicted camera image, , represent the vignetting coefficients, , represents the radial radius, represents the vignetting function.
[0079] In the embodiments of the present application, different from the traditional online photometric calibration method, the parameters can be updated in real time instead of optimizing all the parameters at one time.
[0080] The energy function needs to be constructed based on the color difference of key points, and the colors of each key point in multiple frames of images need to be consistent. Therefore, constructing the energy function based on the photometric model can minimize the color error sum between the actual camera image and the optimized camera image, realize preliminary photometric correction, and obtain the first photometric model.
[0081] The first photometric model can be used to optimize the parameters of the bidirectional optical flow tracking chain, and the bidirectional optical flow tracking chain can also optimize the first photometric model in reverse. Adjust the parameters of the first photometric model according to the tracking results, and perform iterative optimization until convergence to obtain the optimized photometric model and the final bidirectional optical flow tracking chain.
[0082] Based on the optimized photometric model, an optimized camera image is obtained, and key image features therein are extracted.
[0083] The camera image is corrected by the optimized photometric model to reduce the brightness inconsistency problem caused by the non-linearity of the camera response or the vignetting effect, and an optimized camera image with photometric consistency after photometric correction is obtained. Feature extraction is performed on the optimized camera image through OpenCV (a cross-platform computer vision library) to obtain key image features.
[0084] Step 104: Perform spatial synchronization processing on the key image features, point cloud features, and IMU features to obtain tightly coupled multi-sensor data.
[0085] Perform spatial calibration processing on the key image features, point cloud features, and IMU features to construct a factor graph model.
[0086] Based on the obtained key image features, point cloud features, and IMU features, the accurate coordinates of the feature points in the camera image can be obtained, and the external parameters are determined through coordinate transformation. The external parameters are added as variables to the factor graph to construct a factor graph model, and the objective function is jointly optimized according to the sensor factors and the loop detection factors. Specifically, sensor factors and their nodes are designed, where the sensor factors are visual factors, lidar factors, and IMU factors.
[0087] In the embodiment of the present application, the visual factor is composed of a reprojection error factor and a photometric error factor. Among them, the reprojection error factor is determined by using the error between the projection position of the key point in the camera image and its position in the optimized photometric model, and the photometric error factor can be determined by directly comparing the pixel brightness differences of the camera images by the optimized photometric model.
[0088] The lidar factor is composed of a point cloud matching factor and a plane factor. Among them, the point cloud features are registered according to the ICP algorithm (a point cloud registration algorithm) to constrain the poses of adjacent frames. The point cloud matching factor obtains the constrained poses of adjacent frames according to the ICP matching error. The point cloud plane features are extracted, and the plane factor is determined according to the constrained poses of adjacent frames.
[0089] Specifically, the point cloud features are divided into two groups through the ICP algorithm, and the feature registration is completed by solving the iterative steps to obtain the matching point pairs. The iterative steps include that if the error reduction amount between the current iteration and the previous iteration is less than the preset error value or the maximum iteration number is reached, the iteration ends; otherwise, the iteration continues with the current iteration as the initial value. The efficiency of the nearest neighbor search is accelerated by the k-d tree, and a robust estimator is introduced to reduce the influence of outliers on the registration, improving the global convergence of the ICP algorithm. Using the matching point pairs, the relative pose transformation matrix between the features is calculated through the transformation estimation method, that is, the constrained poses of adjacent frames are obtained.
[0090] The IMU factor uses the integration of IMU data to calculate the relative motion between two adjacent frames, constructs motion constraints to obtain the IMU pre-integration factor, and is used to predict the poses of the camera and the 3D lidar.
[0091] Optionally, the loop closure factor identifies loop closures based on point cloud matching, adds pose constraints, and is used to optimize the factor graph.
[0092] The relative pose transformation matrix is used to describe the relative positions and orientations between different sensor coordinate systems, and can merge features obtained from different perspectives or at different times into a unified coordinate system. Based on the relative pose transformation matrix, the relative positions and poses of the camera, the 3D lidar, and the IMU sensor can be determined in reverse. Exemplarily, using Kalibr (a calibration toolbox), the extrinsic matrix is estimated through the hand-eye calibration method, and the spatial calibration between the camera and the IMU sensor is completed according to the rotation matrix and the translation vector. Use calibration tools such as Autoware (an automatic calibration software) or manual feature matching.
[0093] Specifically, an overdetermined equation is constructed based on the relative pose transformation matrix, and a non-linear optimization such as the Gauss-Newton algorithm is used to solve the extrinsic matrix, realizing the spatial calibration process of the camera, the 3D lidar, and the IMU sensor, that is, the spatial synchronization process.
[0094] Corresponding sensor factors are added to each key frame. When a loop closure is detected, a loop closure factor is added, the active nodes of the last 5 frames are retained, and the old frames are marginalized to reduce the computational amount, completing the construction of the factor graph.
[0095] Tightly coupled optimization is performed according to the factor graph model to obtain tightly coupled multi-sensor data.
[0096] Specifically, the factor nodes in the factor graph model are initialized. The residuals and Jacobian matrices of the factor nodes are calculated to construct the normal equation. The normal equation is iteratively updated until the norm of the residuals is less than a preset iteration threshold, and / or the update value of the node is less than a preset iteration threshold, obtaining the updated nodes. Among them, the preset iteration threshold can be taken .
[0097] Exemplarily, a non-linear least squares algorithm such as the Levenberg-Marquardt algorithm is used to optimize the factor graph model.
[0098] Using the results of the spatial calibration process as a reference, all factor nodes are initialized. The residuals for the current variable node values are calculated, and the Jacobian matrix of the residuals with respect to the variable nodes is found. The normal equations are constructed using the residuals and Jacobian matrix. The updated values of the variable nodes are obtained by solving the normal equations, and the iterative update is completed by updating the variable node values. The damping factor is adjusted based on the reduction in the residuals. If the residual reduction is significant, it is reduced to accelerate convergence; if the residual reduction is not significant, it is increased to ensure convergence. The Levenberg-Marquardt algorithm achieves multi-sensor data fusion by minimizing the residuals between different sensor observations. Simultaneously, robust kernel functions such as the Huber kernel (a kernel function used in regression analysis) or the Cauchy kernel are used to suppress outliers, resulting in tightly coupled multi-sensor data. For example, in visual-inertial fusion, the Levenberg-Marquardt algorithm can simultaneously optimize the camera pose and IMU bias, minimizing both the visual reprojection error and the IMU pre-integration error. In addition, the old poses are marginalized using Schur complement to keep the computational complexity constant.
[0099] Step 105: Construct a multi-scale map based on the tightly coupled multi-sensor data.
[0100] A multiscale map is constructed based on tightly coupled multi-sensor data. This multiscale map can be optimized using filtering algorithms, such as the Kalman filter, to eliminate noise and errors. As the multi-sensor movement acquires new data, the multiscale map can be continuously updated and refined, enabling better interaction between the user and the multiscale map, i.e., the robot.
[0101] Based on the depth estimate, the factor graph model generates initial map points and associates them with feature points from the tightly coupled multi-sensor data. A bag-of-words model is used for scene recognition. A joint optimization objective function is constructed based on sensor factors and closed-loop factors, which can be solved using a least-squares algorithm. iSAM2 (an incremental optimizer) is used for real-time incremental updates, optimizing the factor graph model to avoid repeated optimization of the entire map. Furthermore, multi-scale map construction is completed based on the optimized map points in the optimized factor graph model.
[0102] Testing was conducted on datasets such as Euroc (an indoor aircraft navigation dataset). The scenes used were laboratory building lobbies or complex lighting areas with numerous light sources and ground reflections. The optimized photometric model used in this application achieved photometric consistency in camera image processing, reducing the standard deviation of feature point brightness across frames by 62% compared to traditional DSO photometric calibration. This reduced photometric calibration time from minutes to milliseconds while maintaining accuracy. See Table 1 for details on the comparative experimental results.
[0103] Table 1 Comparative experimental results
[0104]
[0105] In Table 1, Visual SLAM is traditional pure visual SLAM, SLAM-IMU is traditional two sensors, the tightly coupled multi-sensor is the camera, 3D lidar and IMU sensors of the present application, and ATE represents the absolute trajectory error.
[0106] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on routine or non-creative labor. The step sequence listed in this embodiment is only one way among the execution sequences of numerous steps and does not represent the only execution sequence. When the actual device or client product executes, it can be executed in the method sequence shown in this embodiment or the drawings or executed in parallel (such as in an environment of parallel processors or multi-threaded processing).
[0107] As Figure 2 shown, the embodiment of the present application also provides a coupling device 200 for multi-sensors. The device includes: a data acquisition module 201, a feature extraction module 202, a photometric calibration module 203, a feature coupling module 204, and a mapping module 205.
[0108] The data acquisition module 201 acquires multi-sensor data; wherein, the multi-sensor data includes camera images, point cloud data, and IMU data.
[0109] The feature extraction module 202 extracts feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features.
[0110] The photometric calibration module 203 performs photometric calibration on the image features to obtain an optimized camera image and extracts key image features therefrom.
[0111] The feature coupling module 204 performs spatial synchronization processing on the key image features, point cloud features, and IMU features to obtain tightly coupled multi-sensor data.
[0112] The mapping module 205 constructs a multi-scale map based on the tightly coupled multi-sensor data.
[0113] Some modules in the device described in this application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0114] The devices or modules illustrated in the above application embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions and described separately. When implementing the embodiments of this application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, the module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.
[0115] The methods, devices or modules described in this application can be implemented in the form of computer-readable program code. The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium that stores computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to be implemented in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers to achieve the same function. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and the structures within the hardware component.
[0116] An embodiment of the present application also provides a device for executing a coupling method of multiple sensors. The device includes: a processor; a memory for storing processor-executable instructions; when the processor executes the executable instructions, the method described in the embodiment of the present application is implemented.
[0117] An embodiment of the present application also provides a non-volatile computer-readable storage medium, on which a computer program or instruction is stored. When the computer program or instruction is executed, the method described in the embodiment of the present application is implemented.
[0118] In addition, in each embodiment of the present application, each functional module may be integrated into one processing module, or each module may exist independently, or two or more modules may be integrated into one module.
[0119] The above storage medium includes but is not limited to Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory may be used to store computer program instructions.
[0120] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or can also be reflected in the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the method described in each embodiment or some parts of the embodiments of the present application.
[0121] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. All or part of the present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on.
[0122] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting the present application; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the present application.
Claims
1. A coupling method for multi-sensors, characterized in that, Including: Obtain multi-sensor data; wherein, the multi-sensor data includes camera images, point cloud data, and IMU data; Extract feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features; Perform photometric calibration on the image features to obtain an optimized camera image, and extract key image features therefrom; Construct a bidirectional optical flow tracking chain based on the image features to determine feature correspondence; Determine the initial parameters of the camera image according to the camera imaging principle; wherein, the initial parameters include a response function, vignetting, and exposure time; Construct a photometric model based on the feature correspondence and the initial parameters; construct a camera response function model, perform dimensionality reduction processing on the camera response function to determine the principal component vector; extract the brightness attenuation coefficient of the camera image, and fit the vignetting model using a radial attenuation model; model the exposure time through an affine transfer function to obtain an exposure time model; construct the photometric model based on the camera response function model, the vignetting model, and the exposure time model; Define an energy function based on the photometric model to obtain an optimized photometric model; use the first photometric model to optimize the parameters of the bidirectional optical flow tracking chain, and the bidirectional optical flow tracking chain optimizes the first photometric model in reverse; adjust the parameters of the first photometric model according to the tracking results and perform iterative optimization until convergence to obtain the optimized photometric model and the final bidirectional optical flow tracking chain; Based on the optimized photometric model, obtain an optimized camera image, and extract key image features therefrom; Perform spatial synchronization processing on the key image features, the point cloud features, and the IMU features to obtain tightly coupled multi-sensor data; Construct a multi-scale map based on the tightly coupled multi-sensor data.
2. The method according to claim 1, characterized in that The extracting feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features includes: Extract the image features of the camera image based on a feature point detection algorithm; Perform regional division on the point cloud data according to a regional segmentation method to obtain the point cloud features; Preprocess the IMU data and extract features therefrom to obtain the IMU features.
3. The method according to claim 2, wherein The extracting the image features of the camera image based on a feature point detection algorithm includes: Construct an image pyramid of the camera image through downsampling to determine key points therein; Determine the direction of the key points according to the first moment; Rotate the sampling window by an angle according to the direction of the key points to construct a feature region to obtain the image features.
4. The method according to claim 2, characterized in that, The performing regional division on the point cloud data according to a regional segmentation method to obtain the point cloud features includes: Perform regional division on the point cloud data according to a preset threshold to obtain segmentation blocks; Determine the curvature value of the point cloud data in the segmentation blocks; Classify the point cloud data based on the curvature value to obtain corner points and plane points; Perform downsampling processing on the plane points to obtain the point cloud features.
5. The method according to claim 1, wherein Performing spatial synchronization processing on the key image features, the point cloud features, and the IMU features to obtain tightly coupled multi-sensor data includes: Performing spatial calibration processing on the key image features, the point cloud features, and the IMU features to construct a factor graph model; Performing tightly coupled optimization according to the factor graph model to obtain tightly coupled multi-sensor data.
6. The method according to claim 5, characterized in that, The performing tightly coupled optimization according to the factor graph model includes: Initializing factor nodes in the factor graph model; Calculating the residuals and Jacobian matrices of the factor nodes to construct a normal equation; Iteratively updating the normal equation until the norm of the residuals is less than a preset iteration threshold, and / or the updated value of the factor nodes is less than a preset iteration threshold to obtain updated nodes.
7. A coupling device for multiple sensors, characterized in that, Including: A data acquisition module that acquires multi-sensor data; wherein, the multi-sensor data includes camera images, point cloud data, and IMU data; A feature extraction module that extracts feature points from the multi-sensor data to obtain image features, point cloud features, and IMU features; A photometric calibration module that performs photometric calibration on the image features to obtain an optimized camera image and extracts key image features therefrom; Constructing a bidirectional optical flow tracking chain based on the image features to determine feature correspondences; Determining initial parameters of the camera image according to the camera imaging principle; wherein, the initial parameters include a response function, vignetting, and exposure time; Constructing a photometric model based on the feature correspondences and the initial parameters; constructing a camera response function model, performing dimensionality reduction processing on the camera response function to determine a principal component vector; extracting a brightness attenuation coefficient of the camera image, fitting a vignetting model using a radial attenuation model; modeling the exposure time through an affine transfer function to obtain an exposure time model; constructing the photometric model based on the camera response function model, the vignetting model, and the exposure time model; Defining an energy function based on the photometric model to obtain an optimized photometric model; using the first photometric model to optimize the parameters of the bidirectional optical flow tracking chain, and the bidirectional optical flow tracking chain optimizes the first photometric model in reverse; adjusting the parameters of the first photometric model according to the tracking results for iterative optimization until convergence to obtain the optimized photometric model and the final bidirectional optical flow tracking chain; Based on the optimized photometric model, obtaining an optimized camera image and extracting key image features therefrom; A feature coupling module that performs spatial synchronization processing on the key image features, the point cloud features, and the IMU features to obtain tightly coupled multi-sensor data; A mapping module that constructs a multi-scale map based on the tightly coupled multi-sensor data.
8. An apparatus for performing a coupling method of multi-sensors, characterized in that, Including: A processor; A memory for storing processor-executable instructions; When the processor executes the executable instructions, the method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Multi-Camera / Lidar / IMU-based multi-sensor SLAM method
CN111983639A