Multi-sensor coupling method, device and equipment

Through feature extraction and photometric calibration of multi-sensor data, combined with spatial synchronization processing and factor graph model, the problem of inaccurate multi-sensor calibration and coupling is solved, and high-precision and real-time graph building effects are achieved.

CN120027782AActive Publication Date: 2025-05-23XIAN DASHENG TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510494566.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the prior art, the calibration and coupling of multiple sensors are inaccurate, resulting in low graph building accuracy, especially in complex scenarios and real-time dynamic scenarios, which have problems such as insufficient accuracy and poor robustness.

Method used

By acquiring the data of multiple sensors for preprocessing, the feature points in camera images, point cloud data and IMU data are extracted, photometric calibration and spatial synchronization are performed, and a factor graph model is constructed to achieve tight coupling of multiple sensors.

Benefits of technology

Real-time and efficient multi-sensor data processing and tight coupling are realized, improving the accuracy and robustness of graph construction, and real-time graph construction in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120027782A_ABST
    Figure CN120027782A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-sensor coupling method, device and equipment, and relates to the technical field of real-time mapping and positioning. The method comprises the following steps: acquiring multi-sensor data; wherein the multi-sensor data comprises a camera image, point cloud data and IMU data; extracting feature points in the multi-sensor data to obtain image features, point cloud features and IMU features; performing photometric calibration on the image features to obtain an optimized camera image, and extracting key image features in the optimized camera image; performing spatial synchronization processing on the key image features, the point cloud features and the IMU features to obtain tight coupling multi-sensor data; based on tight coupling multi-sensor data, a multi-scale map is constructed for mapping. According to the method, real-time cooperative work of a plurality of sensors is realized, and the mapping precision and accuracy of the target scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of real-time mapping and positioning technology, and in particular to a multi-sensor coupling method, device and equipment. Background Art

[0002] With the continuous development of laser detection and ranging technology, high-precision mapping technology has become a key link. In the process of aerial navigation, geophysical mapping and autonomous vehicle navigation, the integration of multiple sensors such as visual SLAM (visual sensor) and 3D laser radar through artificial intelligence algorithms to build maps can provide targeted solutions for various fields such as disaster management, energy surveying, transportation facilities, ecological protection and robotics technology.

[0003] In this development context, existing technologies usually perform multimodal data fusion based on multiple sensors such as visual SLAM, LiDAR and IMU (Inertial Measurement Unit) to improve mapping accuracy. Although multi-sensor systems improve the reliability of data acquisition compared to single-sensor systems or single LiDAR systems and solve the problem of single sensor failure, multi-sensor systems also have problems such as significant differences in data representation, high real-time requirements, and unstable fusion algorithm accuracy. In some complex scenarios, especially in real-time dynamic scenarios, there are still problems such as insufficient mapping positioning accuracy and poor robustness.

[0004] In addition, multiple sensors need to be calibrated for data and photometric calibration before fusion. Existing technologies usually use pulse sensors for time calibration, which is highly dependent on hardware. However, any small errors in the calibration process will be accumulated and amplified in the long run, affecting the accuracy and stability of subsequent mapping. Traditional photometric calibration relies on calibration objects and cannot be applied to cameras of different models. If an unknown camera is used to shoot a video, or the ambient light changes significantly, photometric calibration may take a long time or even be impossible. How to efficiently process data and tightly couple multiple sensors, and how to fully utilize the advantages of each sensor to achieve high-precision, real-time mapping are still under continuous exploration. Summary of the invention

[0005] The embodiments of the present application solve the problems of inaccurate calibration and coupling of multiple sensors and low mapping accuracy in the prior art by providing a multi-sensor coupling method, device and equipment. Camera images, point cloud data and IMU data are obtained by acquiring data from multiple sensors and performing preprocessing. An image pyramid is constructed to detect feature points in the camera image, and image features are obtained for photometric calibration. A photometric model is constructed to achieve photometric consistency of the camera image, and key image features are extracted therein. Point cloud data is classified according to a region segmentation method to obtain point cloud features. Feature points of the IMU data are extracted and used together with key image features and point cloud features for subsequent construction of a factor graph model, multi-sensor coupling is performed, and multi-scale map construction is completed. It realizes real-time and efficient tight coupling of multiple sensors, construction of multi-scale maps, and interaction between users and robots, solving the problem that the prior art cannot perform real-time mapping and low accuracy in complex scenes.

[0006] In a first aspect, an embodiment of the present application provides a multi-sensor coupling method, comprising: acquiring multi-sensor data; wherein the multi-sensor data includes camera images, point cloud data and IMU data; extracting feature points in the multi-sensor data to obtain image features, point cloud features and IMU features; performing photometric calibration on the image features to obtain an optimized camera image, and extracting key image features therein; performing spatial synchronization processing on the key image features, the point cloud features and the IMU features to obtain tightly coupled multi-sensor data; and constructing a multi-scale map based on the tightly coupled multi-sensor data.

[0007] In a possible implementation, the extracting of feature points in the multi-sensor data to obtain image features, point cloud features, and IMU features includes: extracting the image features of the camera image based on a feature point detection algorithm; performing region segmentation on the point cloud data according to a region segmentation method to obtain the point cloud features; and preprocessing the IMU data and extracting features therein to obtain the IMU features.

[0008] In a possible implementation, the extracting the image features of the camera image based on a feature point detection algorithm includes: constructing an image pyramid of the camera image by downsampling to determine key points therein; determining the direction of the key points according to the first-order moment; and rotating the sampling window according to the direction of the key points to construct a feature area to obtain the image features.

[0009] In a possible implementation, the point cloud data is divided into regions according to a region segmentation method to obtain the point cloud features, including: dividing the point cloud data into regions according to a preset threshold to obtain segmented blocks; determining the curvature value of the point cloud data in the segmented blocks; classifying the point cloud data based on the curvature value to obtain corner points and plane points; and downsampling the plane points to obtain the point cloud features.

[0010] In a possible implementation, the photometric calibration of the image features includes: constructing a bidirectional optical flow tracking chain based on the image features to determine feature correspondences; determining initial parameters of the camera image according to the camera imaging principle; wherein the initial parameters include a response function, a vignetting angle, and an exposure time; constructing a photometric model based on the feature correspondences and the initial parameters; defining an energy function based on the photometric model to obtain an optimized photometric model; obtaining an optimized camera image based on the optimized photometric model, and extracting key image features therein.

[0011] In a possible implementation, the photometric model is constructed based on the feature correspondence and the initial parameters, including: constructing a camera response function model, performing dimensionality reduction processing on the camera response function, and determining a principal component vector; extracting the brightness attenuation coefficient of the camera image, and fitting a vignetting model using a radial attenuation model; modeling the exposure time through an affine transfer function to obtain an exposure time model; and constructing a photometric model based on the camera response function model, the vignetting model, and the exposure time model.

[0012] In a possible implementation, the key image features, the point cloud features and the IMU features are spatially synchronized to obtain tightly coupled multi-sensor data, including: spatially calibrating the key image features, the point cloud features and the IMU features to construct a factor graph model; and performing tight coupling optimization according to the factor graph model to obtain tightly coupled multi-sensor data.

[0013] In a possible implementation, the tightly coupled optimization based on the factor graph model includes: initializing the factor nodes in the factor graph model; calculating the residual and Jacobian matrix of the factor nodes to construct normal equations; iteratively updating the normal equations until the norm of the residual is less than a preset iteration threshold, and / or the update value of the update node is less than a preset iteration threshold, to obtain an updated node.

[0014] In a second aspect, an embodiment of the present application provides a multi-sensor coupling device, comprising: a data acquisition module for acquiring multi-sensor data; wherein the multi-sensor data includes camera images, point cloud data and IMU data; a feature extraction module for extracting feature points in the multi-sensor data to obtain image features, point cloud features and IMU features; a photometric calibration module for performing photometric calibration on the image features to obtain an optimized camera image and extract key image features therefrom; a feature coupling module for performing spatial synchronization processing on the key image features, the point cloud features and the IMU features to obtain tightly coupled multi-sensor data; and a mapping module for constructing a multi-scale map based on the tightly coupled multi-sensor data.

[0015] In a third aspect, an embodiment of the present application provides a device for executing a multi-sensor coupling method, the device comprising: a processor; a memory for storing processor executable instructions; when the processor executes the executable instructions, it implements the method described in the first aspect or any possible implementation method of the first aspect.

[0016] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: The embodiments of the present application solve the problems of inaccurate real-time calibration and coupling between multiple sensors and low mapping accuracy in the prior art by adopting a multi-sensor coupling method, device and equipment. Feature extraction is performed on multiple sensor data to obtain image features, point cloud features and IMU features, reducing the amount of calculation and resource consumption for subsequent mapping. The image features are photometrically calibrated, and a bidirectional optical flow tracking chain is constructed to reversely optimize the photometric model to achieve photometric consistency processing and extract key image features. Key image features, point cloud features and IMU features are spatially synchronized to obtain tightly coupled multi-sensor data, and finally a multi-scale map is constructed. The features of different categories of sensors are efficiently extracted, so that the final coupling speed and accuracy are improved, while also reducing the burden on the central processing unit and improving mapping accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flowchart of a multi-sensor coupling method provided in an embodiment of the present application; Figure 2 A schematic diagram of the structure of a multi-sensor coupling device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0020] The following describes some of the techniques involved in the embodiments of the present application to facilitate understanding, and they should be considered as merely exemplary. Therefore, it should be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, some descriptions of well-known functions and structures are omitted in the following description.

[0021] Figure 1 It is a flowchart of a multi-sensor coupling method provided in an embodiment of the present application, including steps 101 to 105. Figure 1 This is only an execution order shown in the embodiment of the present application, and does not represent the only execution order of the multi-sensor coupling method. If the final result can be achieved, Figure 1 The steps shown may be performed in parallel or in reverse, as follows.

[0022] Step 101: Acquire multi-sensor data; wherein the multi-sensor data includes camera images, point cloud data and IMU data.

[0023] In the embodiment of the present application, the multi-sensor includes but is not limited to a camera, a 3D laser radar and an IMU sensor, and the multi-sensor can be installed on the robot. The camera is used to capture continuous images or videos to capture environmental information. The camera can use 100 digital cameras of different models, such as monocular and binocular cameras, such as GMSL2 cameras. And different frame rates, resolutions and exposure times are set to adapt to different lighting conditions. A total of about 15,210 images are collected as camera images, including the internal parameters and distortion coefficients of the camera. The 3D laser radar is used to obtain three-dimensional spatial information, which is usually stored in the form of raw data packets, including information such as the rotation angle, distance value, and reflection intensity of the laser beam. The distance between the object and the sensor is measured by emitting laser pulses and receiving reflected signals, and point cloud data is generated through the pre-processing steps of denoising, thinning, and coordinate transformation. IMU (inertial measurement unit) is used to provide the motion state information of the object and obtain IMU data. Exemplarily, the IMU sensor of BMI088 (acceleration is ±16 , angular velocity is ±2000 ). Among them, IMU data includes the real-time attitude and acceleration of the target object. In addition, when using cameras, lidar and IMU, they need to be calibrated in advance to ensure the accuracy of the collected data.

[0024] In one possible implementation, the multi-sensor data is first time-synchronized to ensure that the synchronized data can accurately reflect the environmental status of each sensor at the same time. Since abnormal situations such as frame loss or timestamp errors may occur during sensor transmission, it is necessary to arrange the multi-sensor data in chronological order and construct a time index. According to the 3D lidar acquisition time, the corresponding position is found in the timeline of other sensors, and the data of the previous and next two frames are locked.

[0025] Specifically, since each sensor has an independent clock source, the existing technology usually aligns the timestamps of each sensor at the initial moment to perform time-synchronized data processing. A pulse generator is added so that all sensors are triggered by the pulse generator. Each trigger corrects the current clock once to eliminate the cumulative error of the clock source and unify the clock source in hardware. Or use existing sensors with built-in pulse generators such as GNSS (Global Navigation Satellite System) to calibrate according to the satellite atomic clock to further improve accuracy. However, in actual applications, the acquisition time of each sensor is not the same. For example, in Kitti (a computer vision algorithm evaluation data set), both lidar data and IMU data are at the ten-bit level, such as 20Hz. However, the acquisition time of 3D lidar is often tens of milliseconds slower than the acquisition time of IMU, and IMU data is usually higher frequency, such as 200Hz. Time differences will accumulate through continuous acquisition, resulting in time difference drift problems.

[0026] In the embodiment of the present application, all sensor clocks are aligned through PTP (Precision Time Protocol) to achieve sensor hardware level alignment. The IMU data (200Hz) is interpolated to the timestamp of the camera image (30Hz) to complete the synchronization of high-frequency data to low-frequency data, and then the point cloud data (10Hz) is extrapolated and predicted to complete the interpolation alignment of low-frequency to high-frequency data.

[0027] Specifically, the received point cloud data is used as the reference data. Each time the point cloud data is received, the acquisition time of the current point cloud data is used as the pre-insertion time point, so as to obtain the equivalent information of other sensors at the same time point except the 3D laser radar. By establishing a time index, the two frames of data before and after the time point are obtained. According to the acquisition time and time point of the two frames of data before and after, linear interpolation calculation is performed to obtain equivalent information and complete time synchronization alignment.

[0028] Step 102: Extract feature points from multi-sensor data to obtain image features, point cloud features, and IMU features.

[0029] Based on the feature point detection algorithm, the image features of the camera image are extracted. In the embodiment of the present application, the features of the camera image are extracted according to the feature point detection algorithm to obtain the image features. Among them, the feature point detection algorithm includes: An image pyramid is constructed by downsampling the camera image and key points are determined.

[0030] Exemplarily, the feature point detection algorithm can use ORB (a computer vision algorithm). Construct a 7-level Gaussian image pyramid and downsample it at a ratio of 1 / 2 to detect scale space key points. Use FAST (a fast corner detection algorithm) to determine the threshold parameters of feature points from different levels of the image pyramid, where the threshold parameter can be 9. Since the grayscale value can intuitively reflect the characteristics of the pixel value changing sharply from light to dark, the grayscale value is divided using the threshold parameter to obtain the N feature points with the strongest response, which are selected as key points. The ORB algorithm can efficiently determine the key points of camera images of different sizes in the image pyramid, thereby achieving partial scale invariance.

[0031] Determine the orientation of key points based on the first-order moment.

[0032] The direction of the key point is calculated using the first-order moment (intensity centroid method) according to the direction calculation formula in a neighborhood with the key point as the center and the scale as the radius. The scale is a Gaussian scale, which can be 1.5 , represents the scale parameter, which can be 1.6. The direction calculation formula is as follows: ; ; In the formula, represents the coordinates of the center of mass, , , represents the first-order moment, Represents the angle between the key point and the centroid, used to indicate the direction of the key point. Indicates taking the tangent.

[0033] The coordinates of the centroid represent the location of the average intensity of the grayscale value in the neighborhood. By drawing a vector from the key point to the centroid, the direction of the key point can be obtained.

[0034] The sampling window is rotated according to the direction of the key point to construct the feature region and obtain the image features. Regardless of the direction of the target camera image, the ORB algorithm can create the same feature vector for the key point, and the same key point can be detected in the camera image rotated at any angle to achieve rotation invariance. By randomly selecting 256 pixel pairs in the sampling window defined around the given key point, a 256-dimensional bit vector is constructed. These random pixel pairs are then rotated according to the direction angle of the key point so that the direction of the random pixel pairs is consistent with the direction of the key point. Finally, the brightness of the random pixel pairs is compared and 1 and 0 are assigned accordingly to create the corresponding feature vector. The order of 1 and 0 varies according to the specific key point and the pixel area around it, indicating the intensity pattern around the key point, so multiple feature vectors can be used to identify a larger area, namely the feature area. In the feature area, feature descriptors are generated to extract the image features of the camera image.

[0035] The point cloud data is divided into regions according to the region segmentation method to obtain point cloud features.

[0036] Specifically, according to a preset threshold, the point cloud data is divided into regions to obtain segmented blocks. In an embodiment of the present application, a preset threshold for height can be set according to the characteristics of the point cloud data. Points with height values ​​less than the preset threshold are preliminarily divided into segmented blocks of the road point cloud area, and points with height values ​​equal to or greater than the preset threshold are divided into segmented blocks with non-road point cloud areas. In an embodiment of the present application, the accuracy of the sensor is relatively high, so the preset threshold can be -0.05 to 0.1 meters for distinction. In addition, those skilled in the art can also use a dynamic adjustment strategy to automatically increase or decrease the preset threshold according to the point cloud density to adapt to different environments.

[0037] Determine the curvature value of the point cloud data in the segmented block. Specifically, in the segmented block, traverse each point in the point cloud data, and based on the sum of the squares of the distance differences between the current point and the 5 points before and after it, use the result as the curvature value of the point. By initializing the point cloud label of the point, it is divided into a set of unprocessed points and a set of unclassified points, and the curvature value and index of each point are stored in cloudSmoothness (point cloud smoothing structure) for subsequent index construction and sorting. Use markOccludedPoints (feature point shielding marking function) to mark the occluded points and points parallel to the light beam to avoid affecting the accuracy of feature extraction. Among them, beam parallelism means that if the distance difference between the 5 points before and after the current point and the current point is greater than the preset ratio, the current point is considered to be parallel to the light beam and is marked as processed.

[0038] Based on the curvature value, the point cloud data is classified to obtain corner points and plane points. The points in the segmented blocks are extracted through extractFeatures (corner point and plane point extraction function), and then the points are classified according to the curvature value.

[0039] Specifically, the point with the highest curvature value is used as the starting point for judgment until all points in the unprocessed point set and the unclassified point set are processed. If the starting point is not blocked and the curvature value is greater than the edge threshold, it is marked as a corner point and added to the corner point cloud set. If the curvature value is less than the surface threshold, it is marked as a plane point and added to the plane point cloud set. Among them, the edge threshold can be taken as the average value of the point cloud density, and the surface threshold is 0.1. In order to avoid the influence of noise points, an upper limit of extraction is set, and a maximum of 20 corner points and plane points are extracted each time. The points adjacent to the corner points and plane points are marked from unprocessed points to processed points to prevent the points from being selected repeatedly.

[0040] Downsample the plane points to obtain point cloud features. In addition, the point cloud features can be described by the FPFH (Fast Point Feature Histogram) algorithm, and the point cloud feature descriptor can be generated by combining the geometric and intensity information. This reduces the redundant data of the point cloud and improves the efficiency of subsequent processing.

[0041] Preprocess the IMU data and extract its features to obtain IMU features.

[0042] Specifically, a low-pass filter, such as Butterworth, is used to remove high-frequency noise in the IMU data, and the zero bias value is calculated based on the bias estimate. The angular velocity and acceleration are integrated in the coordinate system to calculate the relative rotation angle, velocity and displacement, and then the mean, variance and peak values ​​are calculated for feature statistics. The frequency domain information is extracted through the FFT (Fast Fourier Transform) algorithm to obtain the IMU features.

[0043] Step 103: photometrically calibrate the image features to obtain an optimized camera image and extract key image features therein.

[0044] A bidirectional optical flow tracking chain is constructed based on image features to determine feature correspondence.

[0045] In an embodiment of the present application, a bidirectional optical flow tracking chain is constructed based on the key points and feature descriptors in the image features using Lucas-Kanade (dense optical flow estimation method). Since the direction of the key points has been obtained using the ORB algorithm, the optical flow vector of the key points between consecutive frames can be directly estimated. Through forward optical flow tracking, the positions of these key points in the next frame can be predicted, and the key points can be mapped from the previous frame to the next frame. The inter-frame motion prior is realized by adding the optical flow vector of each key point to the position of the corresponding key point in the previous frame, forming a forward optical flow tracking chain from front to back. On the basis of forward optical flow tracking, the predicted key point position is tracked in reverse optical flow, and reverse tracking is performed from the next frame to the previous frame to form a reverse optical flow tracking chain from back to front. At the same time, the accuracy of forward optical flow tracking is verified. By comparing the results of forward and reverse optical flow tracking, erroneous tracking points can be eliminated, mismatching can be reduced, and reverse verification can be completed. The results of forward and reverse optical flow tracking are combined to perform bidirectional optical flow tracking and construct a bidirectional optical flow tracking chain.

[0046] Specifically, the accuracy of target tracking can be verified by feature correspondence. By setting a matching threshold, the matching threshold is set to 1-3 pixels in the embodiment of the present application. Traverse each key point, if the results of the forward and reverse optical flow tracking match, the tracking of the key point is considered to be accurate, and it is included in the tracking chain. If the results of the forward and reverse optical flow tracking do not match, it is considered that the tracking of the key point may be wrong; the position error between the forward and reverse tracking results is calculated by Euclidean distance, if the position error is less than the set matching threshold, the key point is included in the tracking chain; if the position error is greater than the set matching threshold, the key point is removed. For each key point that meets the matching conditions, the results of its forward and reverse optical flow tracking are regarded as the corresponding positions of the same key point in different frames, and the feature correspondence can be confirmed, thereby achieving effective tracking of the target, which can improve the accuracy and robustness of tracking.

[0047] According to the camera imaging principle, the initial parameters of the camera image are determined; the initial parameters include response function, vignetting and exposure time.

[0048] In an embodiment of the present application, according to the camera imaging principle, the light emitted by the light source to the target object will be reflected on the camera lens. After passing through the lens, the brightness will change and then be transmitted to the image sensor of the camera. Energy accumulation will be formed within a certain period of time, and the corresponding light intensity will be obtained through processing of the response function. Based on DSO (a computer vision algorithm of sparse direct method), the initial parameters can be preliminarily determined. The light reflected from the surface of the object is usually called radiant brightness, and the light emitted to the camera sensor is usually called radiant illumination. The exposure time is usually stored in the camera image in the form of metadata of EXIF ​​(Exchangeable Image File Format). EXIF ​​records various settings information of the camera during shooting, including aperture value, exposure time, ISO sensitivity, focal length, shooting time, etc.

[0049] Based on the feature correspondence and initial parameters, a photometric model is constructed.

[0050] According to the initial parameters, the camera response function, vignetting function and exposure time function are constructed respectively. According to the mapping of the radiation brightness to the image intensity, the photometric model is constructed. The brightness attenuation coefficient of the camera image is extracted, and the vignetting model is fitted using the radial attenuation model. The exposure time is modeled through the affine transfer function to obtain the exposure time model. The photometric model is constructed based on the camera response function model, the vignetting model and the exposure time model. The details are as follows.

[0051] ;in, , , ; In the formula, Indicated in The position on the camera image collected at the moment is The gray value corresponding to the pixel point is the radiation brightness. represents the pixel position coordinate set in the camera image, Indicates the collected a moment, represents the camera response function, express The exposure time of the moment, represents the vignetting function, represents the radiance, , represents the vignetting coefficient, which is a two-dimensional real vector. , Indicated in The radial radius at position, represents the average camera response function, represents the number of principal component vectors, represents the response coefficient vector, Indicates principal component vectors, express The exposure time parameters at the moment, express The photometric offset at time, Represents a mathematical constant.

[0052] Among them, the vignetting coefficient is used to describe the attenuation law of the brightness of the camera image with the radial radius. . Exposure time parameters It is used to describe the change of the camera's sensitivity to light over time. Seconds. Photometric offset Used for photometric correction. Considering that in actual applications, the camera will also have a certain output signal in the absence of light, set , the brightness of the camera image is adjusted by the photometric offset, where When it approaches 0, it means no brightness adjustment is performed.

[0053] Specifically, since the bidirectional optical flow tracking chain has already screened out the key points from the pixels of the camera image, the amount of calculation can be reduced to the greatest extent, and the feature correspondence can be established, which can assist in building the photometric model more accurately and quickly. The vignetting function is used to correct the aperture effect of the vignetting and eliminate the brightness difference between images with different exposures. By establishing a normalized coordinate system, such as , normalized radial radius , assuming that the vignetting is symmetrical about the center, extract the brightness attenuation coefficient of each pixel in the camera image, and use the radial attenuation model to approximate the vignetting function , complete the camera's vignetting modeling. The exposure function is used to respond to different exposure times. Since some automatic exposure cameras may not be able to obtain accurate exposure times, the exposure time function is established through the affine transfer function. Parameter model of . Among them, The exponential form prevents the exposure time from being negative and improves computational efficiency.

[0054] Construct a response function model, perform dimensionality reduction on the response function, and determine the principal component vector, thereby reducing the data dimension while reducing the amount of calculation.

[0055] The camera response function is used to convert the light intensity received by the camera sensor into pixel values ​​and construct The camera response function matrix of the data matrix, where each row of the data matrix represents a camera response function In PCA (principal component analysis), SVD (singular value decomposition) is used for dimensionality reduction to determine the camera response function. Through SVD decomposition, it is possible to directly obtain the feature matrix after dimensionality reduction without calculating the covariance matrix and other complex and lengthy matrices. The specific SVD decomposition is as follows.

[0056] ; In the formula, represents the camera response function, and denote the left and right singular vector orthogonal matrices respectively, represents a diagonal matrix of singular values, Represents the transpose of an orthogonal matrix.

[0057] The singular vector orthogonal matrix can capture the principal components of the row and column space, and the diagonal elements of the singular value diagonal matrix are arranged in descending order to reflect the importance of each component. The first four principal component vectors with the largest amount of information are retained. , to capture nonlinear responses and achieve dimensionality reduction while retaining key information.

[0058] Based on the photometric model, the energy function is defined to obtain the optimized photometric model. The energy function is defined as follows: ; In the formula, represents the energy function, represents the pixel position coordinate set in the camera image, represents the effective area of ​​the camera image, Indicates that the pixel point is at position The predicted radiance of the camera image at , represents the vignetting function, Indicates that the pixel point is at position The radiance of the camera image actually captured at represents the camera response function, represents the regularization coefficient, Pick , Represents a regularization term to prevent overfitting.

[0059] ; In the formula, represents the predicted camera image, i.e., the photometrically corrected camera image, represents the inverse function of the camera response function, represents the actual camera image captured, Indicates The exposure time of the moment, represents the vignetting function, Indicates the collected a moment.

[0060] The energy function optimizes the response function and the vignetting function, solves the response function and the vignetting function through an iterative method, minimizes the error between the observed values ​​of all pixels and the predicted pixel values ​​of the photometric model, and makes the photometric model accurately fit the actual imaging process to the greatest extent, ensuring the photometric consistency of the camera image.

[0061] The response function, the derivative of the vignetting angle and the exposure time are calculated by the Jacobian matrix to obtain the optimal photometric parameters. The photometric model is reversely calibrated to obtain the first photometric model. The Jacobian matrix calculation is as follows.

[0062] ; , ; In the formula, represents the camera response function, represents the inverse function of the camera response function, represents the response coefficient vector, Indicates principal component vectors, represents the predicted camera image, , represents the vignetting coefficient, , represents the radial radius, Represents the vignetting function.

[0063] In the embodiment of the present application, unlike the traditional method for realizing online photometric calibration, various parameters can be updated in real time instead of optimizing all parameters at one time.

[0064] The energy function needs to be constructed based on the color difference of key points, and the color of each key point in multiple frames needs to be consistent. Therefore, constructing the energy function based on the photometric model can minimize the color error between the actual camera image and the optimized camera image, achieve preliminary photometric correction, and obtain the first photometric model.

[0065] The first photometric model can be used to optimize the parameters of the bidirectional optical flow tracking chain, and the bidirectional optical flow tracking chain can also reversely optimize the first photometric model. The parameters of the first photometric model are adjusted according to the tracking results, and iterative optimization is performed until convergence to obtain the optimized photometric model and the final bidirectional optical flow tracking chain.

[0066] Based on the optimized photometric model, the optimized camera image is obtained and the key image features are extracted.

[0067] The camera image is corrected by optimizing the photometric model to reduce the brightness inconsistency caused by the nonlinearity of the camera response or the vignetting effect, and an optimized camera image with photometric consistency is obtained after photometric correction. The optimized camera image is feature extracted through OpenCV (a cross-platform computer vision library) to obtain key image features.

[0068] Step 104: Perform spatial synchronization processing on key image features, point cloud features and IMU features to obtain tightly coupled multi-sensor data.

[0069] The key image features, point cloud features and IMU features are spatially calibrated to construct a factor graph model.

[0070] According to the acquired key image features, point cloud features and IMU features, the accurate coordinates of the feature points in the camera image can be obtained, and the external parameters can be determined by coordinate transformation. The external parameters are added as variables to the factor graph, and the factor graph model is constructed. According to the sensor factor and the closed-loop detection factor, the objective function is jointly optimized for optimization. Specifically, the sensor factor and its nodes are designed, where the sensor factor is the visual factor, the lidar factor and the IMU factor.

[0071] In the embodiment of the present application, the visual factor is composed of a reprojection error factor and a photometric error factor. The reprojection error factor is determined by using the error between the projected position of the key point in the camera image and the position in the optimized photometric model, and the photometric error factor can be determined by directly comparing the pixel brightness difference of the camera image with the optimized photometric model.

[0072] The LiDAR factor is composed of a point cloud matching factor and a plane factor. The point cloud features are registered according to the ICP algorithm (a point cloud registration algorithm) to constrain the adjacent frame poses. The point cloud matching factor is obtained based on the ICP matching error to constrain the adjacent frame poses. The point cloud plane features are extracted and the plane factor is determined based on the constrained adjacent frame poses.

[0073] Specifically, the point cloud features are divided into feature points through the ICP algorithm, and initially divided into two groups. The iterative steps are performed to complete the feature registration and obtain the matching point pairs. The iterative steps include: if the error reduction between the current iteration and the previous iteration is less than the preset error value or the maximum number of iterations is reached, the iteration is terminated; otherwise, the iteration is continued with the current iteration as the initial value. The efficiency of the nearest neighbor search is accelerated by the kd tree, and the robust estimator is introduced to reduce the impact of outliers on the registration, thereby improving the global convergence of the ICP algorithm. Using the matching point pairs, the relative posture transformation matrix between the features is calculated by the transformation estimation method, that is, the constrained adjacent frame posture is obtained.

[0074] The IMU factor integrates the relative motion between two adjacent frames using IMU data, constructs motion constraints, and obtains the IMU pre-integration factor, which is used to predict the position and posture of the camera and 3D lidar.

[0075] Optionally, the loop closure factor identifies loop closures based on point cloud matching and adds pose constraints for optimizing the factor graph.

[0076] The relative attitude transformation matrix is ​​used to describe the relative position and orientation between different sensor coordinate systems, and can merge features acquired from different perspectives or at different times into a unified coordinate system. The relative position and attitude of the camera, 3D lidar, and IMU sensor can be reversely determined based on the relative attitude transformation matrix. For example, Kalibr (calibration toolbox) is used to estimate the extrinsic parameter matrix through the hand-eye calibration method, and the spatial calibration between the camera and the IMU sensor is completed based on the rotation matrix and translation vector. Use calibration tools such as Autoware (an automatic calibration software) or manual feature matching.

[0077] Specifically, an overdetermined equation is constructed based on the relative posture transformation matrix, and nonlinear optimization, such as the Gauss-Newton algorithm, is used to solve the extrinsic parameter matrix to achieve spatial calibration processing of the camera, 3D lidar and IMU sensor, that is, spatial synchronization processing.

[0078] Add the corresponding sensor factor to each key frame. When a closed loop is detected, add the closed loop factor, retain the active nodes of the last 5 frames, marginalize the old frames to reduce the amount of calculation, and complete the factor graph construction.

[0079] Tightly coupled optimization is performed according to the factor graph model to obtain tightly coupled multi-sensor data.

[0080] Specifically, the factor nodes in the factor graph model are initialized. The residual and Jacobian matrix of the factor nodes are calculated to construct the normal equation. The normal equation is iteratively updated until the norm of the residual is less than a preset iteration threshold and / or the update value of the update node is less than the preset iteration threshold, and the update node is obtained. The preset iteration threshold can be .

[0081] Illustratively, a nonlinear least squares algorithm, such as the Levenberg-Marquardt algorithm, optimizes the factor graph model.

[0082] Initialize all factor nodes with the result of spatial calibration as reference. Calculate the residual under the current variable node value, and find the Jacobian matrix of the residual to the variable node. Use the residual and Jacobian matrix to construct the normal equation, and obtain the updated value of the variable node by solving the normal equation. Update the value of the variable node to complete the iterative update. Adjust the damping factor according to the reduction of the residual. If the residual decreases significantly, reduce it to speed up the convergence; if the residual decreases insignificantly, increase it to ensure convergence. The Levenberg-Marquardt algorithm achieves multi-sensor data fusion by minimizing the residuals between different sensor observation data. At the same time, robust kernel functions such as Huber kernel (a kernel function used for regression analysis) or Cauchy kernel (Cauchy kernel function) are used to suppress outliers and obtain tightly coupled multi-sensor data. For example, in visual-inertial fusion, the Levenberg-Marquardt algorithm can simultaneously optimize the camera pose and IMU bias, so that both the visual reprojection error and the IMU pre-integration error are minimized. In addition, the old poses are marginalized using Schur complement to keep the computational complexity constant.

[0083] Step 105: Construct a multi-scale map based on the tightly coupled multi-sensor data.

[0084] Based on the tightly coupled multi-sensor data, a multi-scale map is constructed. To optimize the multi-scale map, a filtering algorithm, such as the Kalman filter algorithm, can be used to eliminate noise and errors. As the multi-sensor moves to obtain new data, the multi-scale map can be continuously updated and improved, and the interaction between the user and the multi-scale map, i.e., the robot, can be better realized.

[0085] Based on the factor graph model, the initial map points are generated according to the depth estimation, and are associated with the feature points in the tightly coupled multi-sensor data. The bag-of-words model is used for scene recognition. The joint optimization objective function is constructed based on the sensor factor and the closed-loop factor, which can be solved by the least squares algorithm. iSAM2 (an incremental optimizer) is used to achieve real-time incremental updates, and the factor graph model is optimized to avoid repeated optimization of the entire map. At the same time, based on the optimized map points in the optimized factor graph model, multi-scale map construction is completed.

[0086] Tests were conducted on datasets such as Euroc (an indoor aircraft navigation dataset). The scene can be selected from the lobby of the laboratory building or a complex lighting area with a large number of light sources and ground reflected light in the scene. The optimized photometric model of this application has a 62% reduction in the standard deviation of the brightness of feature points across frames compared to the traditional DSO photometric calibration in terms of photometric consistency of camera image processing, and the photometric calibration time is compressed from minutes to milliseconds while ensuring accuracy. See Table 1 for details, the comparison experiment results table.

[0087] Table 1 Comparative experimental results

[0088] In Table 1, visual SLAM refers to the traditional pure visual SLAM, SLAM-IMU refers to the two traditional sensors, the tightly coupled multi-sensor refers to the camera, 3D lidar and IMU sensor of this application, and ATE refers to the absolute trajectory error.

[0089] Although the present application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative labor. The order of steps listed in this embodiment is only one way of executing the order of many steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in sequence or in parallel according to the method shown in this embodiment or the accompanying drawings (for example, in a parallel processor or multi-threaded processing environment).

[0090] like Figure 2 As shown, the embodiment of the present application further provides a multi-sensor coupling device 200 . The device comprises: a data acquisition module 201 , a feature extraction module 202 , a photometric calibration module 203 , a feature coupling module 204 and a mapping module 205 .

[0091] The data acquisition module 201 acquires multi-sensor data, wherein the multi-sensor data includes camera images, point cloud data and IMU data.

[0092] The feature extraction module 202 extracts feature points from the multi-sensor data to obtain image features, point cloud features and IMU features.

[0093] The photometric calibration module 203 performs photometric calibration on the image features to obtain an optimized camera image and extract key image features therein.

[0094] The feature coupling module 204 performs spatial synchronization processing on key image features, point cloud features and IMU features to obtain tightly coupled multi-sensor data.

[0095] The mapping module 205 constructs a multi-scale map based on the tightly coupled multi-sensor data.

[0096] Some modules in the apparatus described in the present application can be described in the general context of computer executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0097] The devices or modules described in the above application embodiments can be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above devices are described in various modules according to their functions. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware. Of course, the module that implements a certain function can also be implemented by combining multiple sub-modules or sub-units.

[0098] The methods, devices or modules described in this application can be implemented in the form of computer-readable program codes. The controller can be implemented in any appropriate manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program codes (such as software or firmware) that can be executed by the (micro)processor, logic gates, switches, application-specific integrated circuits (English: Application Specific Integrated Circuit; Abbreviation: ASIC), programmable logic controllers and embedded microcontrollers. Examples of controllers include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20 and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program codes, the controller can be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included in it for implementing various functions can also be regarded as structures within the hardware component. Or even, the means for realizing various functions may be regarded as both a software module for realizing the method and a structure within a hardware component.

[0099] An embodiment of the present application further provides a device for executing a multi-sensor coupling method, the device comprising: a processor; a memory for storing processor executable instructions; when the processor executes the executable instructions, the method described in the embodiment of the present application is implemented.

[0100] The embodiments of the present application also provide a non-volatile computer-readable storage medium on which a computer program or instruction is stored. When the computer program or instruction is executed, the method described in the embodiments of the present application is implemented.

[0101] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist independently, or two or more modules may be integrated into one module.

[0102] The above storage media include but are not limited to random access memory (RAM), read-only memory (ROM), cache, hard disk drive (HDD) or memory card. The memory can be used to store computer program instructions.

[0103] It can be seen from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of the present application can be essentially or partly reflected in the prior art in the form of a software product, or it can be reflected in the implementation process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.

[0104] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. All or part of this application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, mobile communication terminals, multi-processor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0105] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit the present application. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present application.

Claims

1. A multi-sensor coupling method, characterized in that: include: Acquire multi-sensor data; wherein the multi-sensor data includes camera images, point cloud data and IMU data; Extracting feature points from the multi-sensor data to obtain image features, point cloud features and IMU features; Performing photometric calibration on the image features to obtain an optimized camera image, and extracting key image features therein; Performing spatial synchronization processing on the key image features, the point cloud features and the IMU features to obtain tightly coupled multi-sensor data; Based on the tightly coupled multi-sensor data, a multi-scale map is constructed.

2. The method according to claim 1, characterized in that The extracting feature points from the multi-sensor data to obtain image features, point cloud features and IMU features includes: Extracting the image features of the camera image based on a feature point detection algorithm; Divide the point cloud data into regions according to a region segmentation method to obtain the point cloud features; The IMU data is preprocessed and features are extracted therein to obtain the IMU features.

3. The method according to claim 2, characterized in that The step of extracting the image features of the camera image based on a feature point detection algorithm includes: constructing an image pyramid of the camera image by downsampling, and determining key points therein; Determine the direction of the key point according to the first-order moment; The sampling window is rotated according to the direction of the key point to construct a feature area and obtain the image feature.

4. The method according to claim 2, characterized in that: The step of dividing the point cloud data into regions according to the region segmentation method to obtain the point cloud features includes: According to a preset threshold, the point cloud data is divided into regions to obtain segmented blocks; Determining a curvature value of the point cloud data in the segmented block; Classifying the point cloud data based on the curvature value to obtain corner points and plane points; Downsampling is performed on the plane points to obtain the point cloud features.

5. The method according to claim 1, characterized in that The photometric calibration of the image features comprises: Building a bidirectional optical flow tracking chain based on the image features to determine feature correspondence; According to the camera imaging principle, the initial parameters of the camera image are determined; wherein the initial parameters include response function, dark angle and exposure time; Based on the feature correspondence and the initial parameters, construct a photometric model; Based on the photometric model, an energy function is defined to obtain an optimized photometric model; Based on the optimized photometric model, an optimized camera image is obtained, and key image features are extracted therein.

6. The method according to claim 5, characterized in that The step of constructing a photometric model based on the feature correspondence and the initial parameters includes: Constructing a camera response function model, performing dimensionality reduction processing on the camera response function, and determining a principal component vector; Extracting the brightness attenuation coefficient of the camera image, and fitting a vignetting model using a radial attenuation model; Modeling the exposure time through an affine transfer function to obtain an exposure time model; A photometric model is constructed based on the camera response function model, the vignetting model and the exposure time model.

7. The method according to claim 1, characterized in that The spatial synchronization processing of the key image features, the point cloud features and the IMU features to obtain tightly coupled multi-sensor data includes: Performing spatial calibration processing on the key image features, the point cloud features, and the IMU features to construct a factor graph model; Tightly coupled optimization is performed according to the factor graph model to obtain tightly coupled multi-sensor data.

8. The method according to claim 7, characterized in that The performing tightly coupled optimization according to the factor graph model comprises: Initializing factor nodes in the factor graph model; Calculating the residual and Jacobian matrix of the factor nodes and constructing the normal equation; The normal equation is iteratively updated until the norm of the residual is less than a preset iteration threshold, and / or the update value of the update node is less than a preset iteration threshold, to obtain an update node.

9. A multi-sensor coupling device, characterized in that: include: A data acquisition module, which acquires multi-sensor data; wherein the multi-sensor data includes camera images, point cloud data and IMU data; A feature extraction module extracts feature points from the multi-sensor data to obtain image features, point cloud features and IMU features; A photometric calibration module performs photometric calibration on the image features to obtain an optimized camera image and extract key image features therein; A feature coupling module performs spatial synchronization processing on the key image features, the point cloud features and the IMU features to obtain tightly coupled multi-sensor data; A mapping module constructs a multi-scale map based on the tightly coupled multi-sensor data.

10. A device for performing a multi-sensor coupling method, characterized in that: include: processor; a memory for storing processor-executable instructions; When the processor executes the executable instructions, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Multi-Camera / Lidar / IMU-based multi-sensor SLAM method

    CN111983639A

  • Image luminosity calibration method and device and computer readable storage medium

    CN113052909A

  • Tight coupling SLAM method based on coding dimension raising fusion

    CN115628739A

  • Stable mapping positioning method and system based on multi-sensor fusion

    CN118067109A

  • SLAM method and system for fusing IMU and visual data based on point-line visual features

    CN118424266A