Automatic Calibration Method and Device, Computer-Readable Storage Medium, and Terminal

Through the automatic calibration method, the feature matching and iterative optimization of point cloud and image data are used to solve the problems of high cost, low efficiency and insufficient accuracy of joint calibration of laser sensors and image sensors, and efficient and accurate multi-sensor calibration is achieved.

CN114519681BActive Publication Date: 2025-07-18SHANGHAI XIANTU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111677841.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-18
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

The existing joint calibration technology of laser sensors and image sensors has problems such as high labor costs, low flexibility and efficiency, and insufficient calibration accuracy.

Method used

By determining the original point cloud data and original image data collected by the same time stamp for the same scene, the point cloud depth continuous line features and image depth continuous line features are extracted respectively, and the two-dimensional pixel points set is determined using the preset rotation matrix and translation vector initial values, and the rotation matrix and translation vector are optimized through the iterative algorithm to achieve optimal value to achieve automatic calibration.

Benefits of technology

It realizes fully automatic calibration of the site and equipment without additional calibration, reduces labor costs, improves calibration flexibility and accuracy, enhances the density of point cloud data, and ensures the accuracy of feature point matching and calibration results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519681B_ABST
    Figure CN114519681B_ABST
Patent Text Reader

Abstract

An automatic calibration method, device, computer-readable storage medium, and terminal. The method includes: obtaining original point cloud data and original image data collected for the same regional scene with the same timestamp; respectively determining point cloud depth continuous line features and image depth continuous line features according to the original point cloud data and the original image data; determining a two-dimensional pixel point set of the point cloud depth continuous line features according to a preset initial rotation matrix value and an initial translation vector value; for each two-dimensional pixel point in the two-dimensional pixel point set, searching for neighboring pixel points from the image depth continuous line features to obtain a neighboring pixel point set and a normal vector set of the neighboring pixel points; determining an optimal value of the rotation matrix and an optimal value of the translation vector based on the neighboring pixel point set and the normal vector set. The present invention can achieve full-automatic calibration between multiple sensors, reduce the calibration cost, and improve the calibration efficiency and calibration accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-sensor external parameter calibration, and in particular to an automatic calibration method and device, a computer-readable storage medium, and a terminal. Background Art

[0002] With the rapid development of artificial intelligence technology, the technology of motor vehicle autonomous driving is in the ascendant. LiDAR and camera are two important sensors in the sensor suite of autonomous driving, and each of them has its own advantages and disadvantages. LiDAR emits a laser beam to detect the target, then compares the received reflected signal with the transmitted signal, and obtains the precise distance, direction, height, shape and other characteristics of the target after appropriate processing. It has high performance in identifying the contour of obstacles and can provide the main vehicle with the most direct reflection of the true morphological characteristics of the target. LiDAR sensors are also more robust to weather, but their disadvantage is that they cannot depict the texture and color information of objects. In contrast, the resolution of cameras is high, and the images collected can provide rich texture and color information, but the images collected by cameras are easily affected by bad weather conditions, and the distance accuracy is poor, which is not suitable for estimating information such as object shape, position, speed and acceleration.

[0003] Due to the above-mentioned defects of using a single sensor, the current mainstream technology is to fuse the information sensed by multiple sensors. The data fusion of the lidar and the camera requires the data collected by both to be converted to a unified coordinate system through coordinate transformation, so that the point cloud coordinate points of the lidar can be converted to the two-dimensional coordinate system of the camera's image. The process of determining the coordinate transformation relationship is the external parameter calibration process of different sensors. Accurate calibration results can provide more stable and good judgment results for the later information fusion perception, thereby providing safer protection for the entire unmanned vehicle system.

[0004] In the prior art, multi-sensor joint calibration often requires the use of a calibration device to assist in calibration. For example, the most common method is to use markers with specific patterns such as checkerboards and QR codes for calibration, that is, to capture the feature points on the specific pattern in the image data collected by the laser radar and the camera, and match the feature points one by one, so as to calculate the 6-DOF conversion matrix of the laser radar and the camera. This calibration technology depends on specific devices and calibration sites, and can only rely on offline manual auxiliary calibration, which has high labor costs, low flexibility and efficiency. In addition, in the prior art, the feature points on the commonly used calibration devices are relatively concentrated, resulting in serious overfitting of the 6-DOF conversion matrix; or, because the calibration devices are not large in size, only close-range calibration can be performed, and the obtained calibration results are not suitable for some wide scenes; furthermore, the prior art mostly uses single-frame point cloud data, resulting in sparse point cloud data and inability to accurately match feature points. The above factors will cause the calibration results to be insufficiently accurate.

[0005] Therefore, there is an urgent need for an automatic calibration method that can achieve full-automatic calibration of the point cloud data collected by a laser sensor and the image data collected by an image sensor, reduce the calibration cost, and improve the calibration efficiency and accuracy. Summary of the Invention

[0006] The technical problem solved by the present invention is the problems of high labor cost, low flexibility and efficiency, and insufficient calibration accuracy in the existing combined calibration technology of laser sensors and image sensors.

[0007] To solve the above problems, an embodiment of the present invention provides an automatic calibration method, including the following steps: determining the original point cloud data and the original image data collected for the same regional scene at the same timestamp; determining the point cloud depth continuous line feature according to the original point cloud data, and determining the image depth continuous line feature according to the original image data; determining a two-dimensional pixel point set of the point cloud depth continuous line feature according to a preset initial rotation matrix value and a preset initial translation vector value; for each two-dimensional pixel point in the two-dimensional pixel point set, finding the nearest neighboring pixel point to the two-dimensional pixel point from the image depth continuous line feature, and determining the normal vector of the neighboring pixel point, so as to obtain a neighboring pixel point set and a normal vector set of the neighboring pixel points; based on the neighboring pixel point set and the normal vector set, determining the optimal value of the rotation matrix and the optimal value of the translation vector.

[0008] Optionally, determining the point cloud depth continuous line feature according to the original point cloud data includes: determining multiple frames of point cloud data collected for the same regional scene within a preset time period before and / or after the acquisition timestamp of the original point cloud data; constructing cumulative dense point cloud data for the multiple frames of point cloud data, and the positive direction of the coordinate system origin of the cumulative dense point cloud data is the same as the positive direction of the coordinate system origin of the original point cloud data; extracting the point cloud depth continuous line feature from the cumulative dense point cloud data.

[0009] Optionally, extracting the point cloud depth continuous line feature from the cumulative dense point cloud data includes: downsampling and rasterizing the cumulative dense point cloud data to obtain multiple grids; for the point cloud data in each grid, using the random sample consensus algorithm RANSAC for fitting to obtain multiple planes; taking the lines where the multiple planes intersect as the point cloud depth continuous line feature.

[0010] Optionally, determining the image depth continuous line feature according to the original image data includes: removing distortion from the original image data and converting it into a grayscale image; using the Laplace operator to perform edge extraction on the grayscale image to obtain the image depth continuous line feature.

[0011] Optionally, the following formula is used to determine the set of two-dimensional pixel points of the point cloud depth continuous line feature:

[0012] p = π[(R|t)P];

[0013] where p represents the set of two-dimensional pixel points; π represents the projection function; R represents the rotation matrix; t represents the translation vector; and P represents the point cloud depth continuous line feature.

[0014] Optionally, the search algorithm is selected from: Quad-tree algorithm, Oct-tree algorithm, k-nearest neighbor algorithm.

[0015] Optionally, determining the optimal value of the rotation matrix and the optimal value of the translation vector based on the set of neighboring pixel points and the set of normal vectors includes: constructing a calibration loss function using the set of neighboring pixel points, the set of normal vectors, the initial value of the rotation matrix, and the initial value of the translation vector; using a preset iterative algorithm and a preset termination condition to minimize the calibration loss function to determine the optimal value of the rotation matrix and the optimal value of the translation vector.

[0016] Optionally, the following formula is used to construct the calibration loss function:

[0017] f(x) = 0.5 * norm((transpose(n) * (π(R|t)P) - q));

[0018] where f(x) is used to represent the calibration loss function; norm() represents the vector norm calculation function; transpose() represents the transpose function; n represents the set of normal vectors; π represents the projection function; R represents the rotation matrix; t represents the translation vector; P represents the point cloud depth continuous line feature; and q represents the set of neighboring pixel points.

[0019] Optionally, the iterative algorithm is selected from: gradient descent algorithm, Newton algorithm, Gauss-Newton algorithm, Levenberg-Marquardt algorithm.

[0020] An embodiment of the present invention further provides an automatic calibration device, including:

[0021] An original data acquisition module, configured to determine the original point cloud data and the original image data collected for the same regional scene with the same timestamp; a depth continuous line feature determination module, configured to determine the point cloud depth continuous line feature according to the original point cloud data, and determine the image depth continuous line feature according to the original image data; a two-dimensional pixel point set determination module, configured to determine the two-dimensional pixel point set of the point cloud depth continuous line feature according to a preset initial rotation matrix value and a preset initial translation vector value; a neighboring pixel point set determination module, configured to, for each two-dimensional pixel point in the two-dimensional pixel point set, find the neighboring pixel point closest to the two-dimensional pixel point from the image depth continuous line feature, and determine the normal vector of the neighboring pixel point, so as to obtain a neighboring pixel point set and a normal vector set of the neighboring pixel points; an optimal value calculation module, configured to determine the optimal value of the rotation matrix and the optimal value of the translation vector based on the neighboring pixel point set and the normal vector set.

[0022] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, the steps of the above automatic calibration method are executed.

[0023] An embodiment of the present invention further provides a terminal, including a memory and a processor, where a computer program capable of running on the processor is stored on the memory, and when the processor runs the computer program, the steps of the above automatic calibration method are executed.

[0024] Compared with the prior art, the technical solution of the embodiment of the present invention has the following beneficial effects:

[0025] In an embodiment of the present invention, by determining the original point cloud data and the original image data collected for the same regional scene using the same timestamp; then respectively determining the point cloud depth continuous line feature and the image depth continuous line feature according to the original point cloud data and the original image data; then determining the two-dimensional pixel point set of the point cloud depth continuous line feature according to the preset initial rotation matrix value and the preset initial translation vector value; determining the nearest neighboring pixel point set and the normal vector set of the neighboring pixel points based on the two-dimensional pixel point set; and finally determining the optimal value of the rotation matrix and the optimal value of the translation vector based on the neighboring pixel point set and the normal vector set. Compared with the existing multi-sensor joint calibration technology that relies on specific calibration devices and calibration sites, has high labor costs, low flexibility and efficiency, and insufficient calibration accuracy, the embodiment of the present invention performs feature point matching on the depth continuous line feature of the point cloud data and the depth continuous line feature of the image data collected by different sensors through the preset initial rotation matrix value and the preset initial translation vector value, and then uses an iterative algorithm to determine the optimal value of the rotation matrix and the optimal value of the translation vector. The whole process is based on a fully automatic method, without the need for additional calibration sites and calibration devices, saving the labor cost of calibration and improving the calibration flexibility and calibration accuracy.

[0026] Further, by determining multiple frames of point cloud data collected for the same regional scene within a preset time period before and / or after the acquisition timestamp of the original point cloud data; then constructing cumulative dense point cloud data from the multiple frames of point cloud data, and then extracting the point cloud depth continuous line feature from the cumulative dense point cloud data. The purpose is to enrich the data volume of the point cloud data by superimposing multiple frames of point clouds, solve the problem that the feature points cannot be accurately matched with the image depth continuous line feature due to the excessive sparsity of a single frame of point cloud data, and improve the calibration accuracy.

[0027] Further, when determining the image depth continuous line feature according to the original image data, by removing the distortion of the original image and converting it into a grayscale image and then extracting the image depth continuous line feature from the grayscale image, the extracted feature points can be made more accurate, and thus the subsequent feature point matching and calculation results can be more accurate, improving the calibration accuracy.

[0028] Further, by constructing a calibration loss function and using a preset iterative algorithm and a preset termination condition to minimize the calibration loss function to determine the optimal value of the rotation matrix and the optimal value of the translation vector, the effect of automatically completing the external parameter calibration is achieved. Description of the Drawings

[0029] Figure 1 is the flowchart of the first automatic calibration method in the embodiment of the present invention;

[0030] Figure 2 is Figure 1Flowchart of a specific implementation manner of step S12;

[0031] Figure 3 It is the flowchart of the second automatic calibration method in the embodiments of the present invention;

[0032] Figure 4 It is the structural schematic diagram of an automatic calibration device in the embodiments of the present invention. Specific implementation manner

[0033] As mentioned above, in the field of motor vehicle autonomous driving technology, due to the respective advantages and disadvantages of single sensors, fusing the information sensed by multiple sensors is the mainstream technology currently used in autonomous driving. When multiple sensors are used jointly, different sensors need to be calibrated. For example, the point cloud coordinate points of a lidar need to be converted into the two-dimensional image coordinate system of a camera.

[0034] In the prior art, multi-sensor joint calibration often requires the assistance of a calibration device. For example, the most common method is to use markers with specific patterns such as checkerboards and two-dimensional codes for calibration, that is, to capture the feature points on the specific patterns in the image data collected by the lidar and the camera respectively, and perform one-by-one matching of the feature points, so as to calculate the 6-degree-of-freedom transformation matrix of the lidar and the camera.

[0035] The inventors of the present invention have found through research that the existing calibration technologies rely on specific devices and calibration sites, and often rely on offline manual assistance for calibration, resulting in high labor costs, low flexibility and efficiency. In addition, in the prior art, the feature points on the commonly used calibration devices are relatively concentrated, resulting in serious overfitting of the 6-degree-of-freedom transformation matrix; or, due to the small size of the calibration devices, only short-distance calibration can be performed, and the obtained calibration results are not applicable to some vast scenarios; furthermore, the prior art mostly uses single-frame point cloud data, resulting in sparse point cloud data and unable to accurately match feature points. The above factors will all cause insufficient accuracy of the calibration results.

[0036] In an embodiment of the present invention, the original point cloud data and the original image data collected for the same regional scene at the same timestamp are determined; then, a point cloud depth continuous line feature and an image depth continuous line feature are respectively determined according to the original point cloud data and the original image data; next, a two-dimensional pixel point set of the point cloud depth continuous line feature is determined according to a preset initial rotation matrix value and a preset initial translation vector value; a nearest neighboring pixel point set and a normal vector set of the neighboring pixel points are determined based on the two-dimensional pixel point set; finally, an optimal value of the rotation matrix and an optimal value of the translation vector are determined based on the neighboring pixel point set and the normal vector set. Compared with the existing multi-sensor joint calibration technology, which relies on specific calibration devices and calibration sites, has high labor costs, low flexibility and efficiency, and insufficient calibration accuracy, the embodiment of the present invention performs feature point matching on the depth continuous line feature of the point cloud data and the depth continuous line feature of the image data collected by different sensors through the preset initial rotation matrix value and the preset initial translation vector value, and then uses an iterative algorithm to determine the optimal value of the rotation matrix and the optimal value of the translation vector. The whole process is based on a fully automatic method, without the need for additional calibration sites and calibration devices, saving the labor cost of calibration and improving the calibration flexibility and calibration accuracy.

[0037] To make the above objects, features, and beneficial effects of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.

[0038] Refer to Figure 1 , Figure 1 which is a flowchart of the first automatic calibration method according to an embodiment of the present invention. The first automatic calibration method may include steps S11 to S15:

[0039] Step S11: Determine the original point cloud data and the original image data collected for the same regional scene at the same timestamp;

[0040] Step S12: Determine a point cloud depth continuous line feature according to the original point cloud data, and determine an image depth continuous line feature according to the original image data;

[0041] Step S13: Determine a two-dimensional pixel point set of the point cloud depth continuous line feature according to a preset initial rotation matrix value and a preset initial translation vector value;

[0042] Step S14: For each two-dimensional pixel point in the two-dimensional pixel point set, find the nearest neighboring pixel point to the two-dimensional pixel point from the image depth continuous line feature, and determine the normal vector of the neighboring pixel point, so as to obtain a neighboring pixel point set and a normal vector set of the neighboring pixel points;

[0043] Step S15: Determine the optimal values of the rotation matrix and the optimal values of the translation vector based on the set of adjacent pixel points and the set of normal vectors.

[0044] In the specific implementation of step S11, the point cloud data can be used to indicate a set of sampling points with spatial coordinates in a three-dimensional coordinate system. The point cloud data can have geometric position information and can also have attribute information such as color. The original point cloud data can be collected by a laser sensor at a certain timestamp for a certain area scene. For example, it can be collected by a lidar, or it can be obtained from a certain acquisition module after being collected by a laser sensor.

[0045] The image data can be used to indicate a set of gray values of each pixel represented by numerical values, and this data set can be represented in a two-dimensional coordinate system. An image in the real world is generally represented by the intensity and spectrum (color) of light at each point on the image. When converting image information into data information, the image must be decomposed into many small regions, which are called pixels and can be represented by a numerical value for its gray level. By sequentially extracting the information of each pixel, a discrete array can be used to represent a continuous image. The original image data can be collected by an image sensor at the same timestamp for the same area scene. For example, it can be collected by an image sensor such as a camera or a camera, or it can be obtained from a certain acquisition module after being collected by an image sensor.

[0046] It can be understood that the reason for determining to collect the original point cloud data and the original image data at the same timestamp is that in the field of motor vehicle driverless, sensors such as lidars and cameras are often placed at the same position directly in front of the main vehicle. During the movement of the main vehicle, the specific area scenes sensed by the sensors at different times are also changing, and the data collected by the same sensor at different times is different. Therefore, by controlling the collection timestamp to be consistent and the area scene targeted during collection to be consistent, it can be ensured that the original point cloud data and the original image data obtained reflect or point to the same object in the real world, thereby ensuring the accuracy of the subsequent calibration results.

[0047] It should be noted that in the specific implementation, before determining the original point cloud data and the original image data, the self-calibration of the laser sensor should have been completed, and the internal parameter calibration of the image sensor has also been completed.

[0048] In the specific implementation of step S12, the depth continuous line feature can be a set of edge feature points used to describe the edge information in the scene. The point cloud depth continuous line feature can be a set of edge feature points extracted from the point cloud data, and the image depth continuous line feature can be a set of edge feature points extracted from the image data.

[0049] Refer toFigure 2 , Figure 2 is Figure 1 a flowchart of a specific implementation manner of step S12 in the method. Determining the point cloud depth continuous line feature according to the original point cloud data may include steps S121 to S125. Each step is described below.

[0050] In step S121, multiple frames of point cloud data collected for the same regional scene within a preset time period before and / or after the acquisition timestamp of the original point cloud data are determined.

[0051] In step S122, cumulative dense point cloud data is constructed for the multiple frames of point cloud data.

[0052] It can be understood that in specific implementations, due to the occlusion of the laser scanning beam by objects, it may not be possible to obtain the three-dimensional point cloud of the entire object through a single scan. In addition, since a single frame of point cloud data may be too sparse, it may not be possible to accurately obtain the features of a real object only through a single frame of point cloud data. Therefore, it is often necessary to scan the object from different positions and angles to obtain multiple frames of point cloud data. For example, multiple lidars can be used simultaneously, or a high-beam lidar can be used, that is, a lidar with a 360-degree detection field of view (denoted as the lidar system) can be used to collect multiple frames of point cloud data for the same object, and then the multiple frames of point cloud data are fused to construct cumulative dense point cloud data to enhance the point cloud data.

[0053] Specifically, there are many methods for constructing cumulative dense point cloud data for the multiple frames of point cloud data. In some non-limiting examples, methods such as point cloud stitching, registration, and merging can be used to transform the point clouds at different positions to the same position through the information of the overlapping parts, so that the positive direction of the coordinate system origin of the cumulative dense point cloud data is the same as the positive direction of the coordinate system origin of the original point cloud data.

[0054] In the embodiments of the present invention, cumulative dense point cloud data is constructed by using multiple frames of point cloud data, and then the point cloud depth continuous line feature is extracted from the cumulative dense point cloud data. The purpose is to enhance and enrich the data volume of the point cloud data by superimposing multiple frames of point clouds, solve the problem that a single frame of point cloud data is too sparse to comprehensively reflect the features of a real object and cannot accurately match the feature points with the image depth continuous line feature, and can improve the calibration accuracy.

[0055] In step S123, the cumulative dense point cloud data is downsampled and rasterized to obtain multiple grids.

[0056] Among them, downsampling may refer to sampling (extracting) the cumulative dense point cloud data at intervals of several sample values. In this way, the new data sequence obtained is the downsampling of the original sequence. It can be understood that the purpose of downsampling is to extract a part of the effective data from the cumulative dense point cloud data; rasterization is to process the cumulative dense point cloud data using a grid to obtain multiple grids, and the point cloud data in each grid represents a small area of space. The purpose of rasterizing the point cloud data is to homogenize the point cloud data.

[0057] In step S124, for the point cloud data in each grid, the Random Sample Consensus (RANSAC) algorithm is used for fitting to obtain multiple planes.

[0058] In step S125, the lines where the multiple planes intersect are used as the point cloud depth continuous line features.

[0059] Among them, the Random Sample Consensus (RANSAC) algorithm is an iterative algorithm that can estimate the parameters of a mathematical model from a set of observation data sets containing "outliers" through an iterative method. The basic assumption of RANSAC is that the data consists of "inliers". For example, the distribution of the data can be explained by some model parameters; "outliers" are used to indicate data that does not fit the model, and "inliers" are used to indicate data that fits the model. A simple example is to find a suitable 2D line from a set of observation data. Assume that the observation data contains inliers and outliers, where the inliers are approximately passed through by the line, and the outliers are far from the line. The input of RANSAC is a set of observation data, a parameterized model that can explain or fit the observation data, and some reliable parameters. RANSAC achieves the goal by repeatedly selecting a random subset of the data, and the selected subset is assumed to be inliers. This process is repeated a fixed number of times. Each time the generated model is either discarded because there are too few inliers or selected because it is better than the existing model.

[0060] Fitting, also known as curve fitting, commonly known as drawing a curve, is a representation method that substitutes existing data into a mathematical formula through mathematical methods. Scientific and engineering problems can obtain several discrete data through methods such as sampling and experiments. Based on these data, we often hope to obtain a continuous function (that is, a curve) or a more dense discrete equation that fits the known data. This process is called fitting.

[0061] Specifically, for the specific steps of determining the point cloud depth continuous line features according to the original point cloud data, reference can be made to the previous text for execution, and details will not be elaborated here.

[0062] Further, in a specific implementation, determining the image depth continuous line feature according to the original image data includes: removing distortion from the original image data and converting it into a grayscale image; using a Laplace operator to perform edge extraction on the grayscale image to obtain the image depth continuous line feature.

[0063] Among them, image distortion may refer to image distortion caused by factors such as lens manufacturing precision and assembly process deviation, that is, the errors in spectral characteristics and geometric characteristics between the image and the real scene it reflects, namely radiation error and geometric error. The former is manifested as the distortion of the image in grayscale; the latter is manifested as the deformation in geometric relationship. Therefore, in a specific implementation, the original image data collected by the image sensor usually needs to perform radiation correction and geometric correction on the image according to the distortion parameters, and this process is to remove distortion.

[0064] The grayscale image corresponds to a color image and is also called a gray-scale image. The white and black are divided into several levels according to a logarithmic relationship, which is called grayscale (the grayscale is divided into 256 levels from 0 to 255), and the image represented by grayscale is called a grayscale image.

[0065] The Laplace operator is a second-order differential operator in an n-dimensional Euclidean space, defined as the divergence (▽·f) of the gradient (▽f). The Laplace operator is a classic image edge enhancement operator, and it can detect the edge features of an image by sharpening the image. Among them, the role of image sharpening processing is to enhance the gray-scale contrast, so that the blurred image becomes clearer; the essence of image blurring is that the image is subjected to an average operation or an integral operation, so the inverse operation can be performed on the image. For example, the differential operation can highlight the details of the image and make the image clearer. Since the Laplace is a differential operator, its application can enhance the region with sudden gray-scale changes in the image and weaken the region with slow gray-scale changes. Therefore, for sharpening processing, the Laplace operator can be selected to process the original image to generate an image describing the sudden gray-scale changes, and then the Laplace image is superimposed on the original image to generate a sharpened image.

[0066] Edge extraction may refer to the processing of the picture contour in digital image processing. For the boundary of the image, that is, the place where the gray-scale value changes violently, it can be defined as an edge or an inflection point, and the inflection point can refer to the point where the concavity and convexity of the function change, that is, the point where the second derivative is zero. The basic idea of the edge extraction is first to use an edge enhancement operator, such as the Laplace operator mentioned above, to highlight the local edges in the image, and then define the "edge strength" of the pixel, and extract the edge point set by setting a threshold. The edge extraction of an image includes two basic contents: (1) Use an edge operator to extract the edge point set reflecting the gray-scale change. (2) Eliminate some boundary points or fill in the boundary discontinuity points in the edge point set, and connect these edges into a complete line.

[0067] In the embodiments of the present invention, after removing distortion from the original image and converting it into a grayscale image, the Laplacian operator is then used to extract edges from the grayscale image to obtain the continuous line feature of the image depth, which can make the extracted feature points more accurate, and further make the subsequent feature point matching and operation results more accurate, improving the calibration accuracy.

[0068] Continue to refer to Figure 1 , in the specific implementation of step S13, a two-dimensional pixel point set of the continuous line feature of the point cloud depth is determined according to a preset initial rotation matrix and a preset initial translation vector.

[0069] Among them, the rotation matrix may refer to a matrix that changes the direction of a vector but does not change its magnitude when multiplying a vector. In a three-dimensional coordinate system, rotation is performed around a certain axis. The translation vector is also called the vector of translation transformation. The rotation matrix can be used to represent the rotation relationship between different data sets, and the translation vector can be used to represent the translation transformation relationship between different data sets.

[0070] Furthermore, the following formula is used to determine the two-dimensional pixel point set of the continuous line feature of the point cloud depth:

[0071] p = π[(R|t)P];

[0072] Among them, p represents the two-dimensional pixel point set; π represents the projection function; R represents the rotation matrix; t represents the translation vector; P represents the continuous line feature of the point cloud depth.

[0073] Among them, the projection function can be a mapping function that converts points in a three-dimensional space coordinate system into points in a two-dimensional space coordinate system. In the specific implementation of the embodiments of the present invention, according to the preset initial rotation matrix and the preset initial translation vector, and by using the projection function π, the two-dimensional pixel point set p corresponding to the continuous line feature P of the point cloud depth can be determined.

[0074] In the specific implementation of step S14, for each two-dimensional pixel point in the two-dimensional pixel point set, the nearest neighboring pixel point to the two-dimensional pixel point is searched from the continuous line feature of the image depth, and the normal vector of the neighboring pixel point is determined to obtain a set of neighboring pixel points and a set of normal vectors of the neighboring pixel points.

[0075] Furthermore, in some non-limiting embodiments, the search algorithm for searching the nearest neighboring pixel point to the two-dimensional pixel point from the continuous line feature of the image depth can be selected from: Quad-tree algorithm, Oct-tree algorithm, k-nearest neighbor algorithm. It can also be selected from other search algorithms that can search for neighboring points, and the embodiments of the present invention do not limit the search algorithm used.

[0076] Among them, the search for the nearest neighboring pixel points can be to search for one or more pixel points with the closest distance to the target pixel coordinates in a given two-dimensional image pixel set. Among them, the quadtree is a tree-like data structure that evenly divides the 2D space and manages the objects in the space. The quadtree is a data structure, a data structure in which each node has at most four subtrees. In the two-dimensional space, the plane pixels can be repeatedly divided into four parts, and the depth of the tree is determined by the complexity of the picture, computer memory, and graphics. The quadtree algorithm performs matching searches by continuously dividing the records to be searched into 4 parts until only one record remains.

[0077] The octree is a tree-like data structure used to describe three-dimensional space. Each node of the octree represents a cubic volume element, and each node has eight child nodes. The sum of the volume elements represented by the eight child nodes is equal to the volume of the parent node. The basic principle of the octree algorithm is as follows: Step (1), set the maximum recursion depth; Step (2), find the maximum size of the scene and create the first cube with this size; Step (3), sequentially throw the unit elements into the cube that can contain them and has no child nodes; Step (4), if the maximum recursion depth is not reached, divide it into eight equal parts, and then distribute all the unit elements contained in the cube to the eight child cubes; Step (5), if it is found that the number of unit elements assigned to the child cube is not zero and is the same as that of the parent cube, then the subdivision of the child cube stops, because according to the space segmentation theory, the distribution obtained from the subdivided space must be less. If the number is the same, the number will still be the same no matter how it is cut, which will cause an infinite cutting situation. Step (6), repeat the above step of sequentially throwing the unit elements into the cube that can contain them and has no child nodes until the maximum recursion depth is reached.

[0078] The k-nearest neighbor algorithm is to, given a training data set, find the k nearest instances or k neighbors of a new input instance in the training data set. If the majority of these k instances belong to a certain class, then classify the input instance into this class. Among them, k is a positive constant greater than or equal to 1, and the specific value of k can be set according to the requirements of different application scenarios. For example, k can be 5, and k can also be 10. This algorithm is more suitable for the automatic classification of class domains with a relatively large sample size, while it is more likely to cause misclassification for those class domains with a relatively small sample size.

[0079] In a specific implementation, for each two-dimensional pixel point in the two-dimensional pixel point set, the process of using a search algorithm to find the nearest neighboring pixel points in the image depth continuous line feature to obtain the corresponding set of neighboring pixel points and the set of normal vectors of the neighboring pixel points is the process of matching feature point pairs. It can be understood that the more accurate the feature point matching is, the more accurate the values of the rotation matrix and the translation vector calculated by the iterative algorithm in the subsequent steps will be, and the more stable and good the determination result can be provided for the later information fusion perception, thereby providing a more secure guarantee for the entire unmanned vehicle system.

[0080] In the specific implementation of step S15, determining the optimal value of the rotation matrix and the optimal value of the translation vector based on the set of neighboring pixel points and the set of normal vectors includes: constructing a calibration loss function using the set of neighboring pixel points, the set of normal vectors, the initial value of the rotation matrix, and the initial value of the translation vector; using a preset iterative algorithm and a preset termination condition to minimize the calibration loss function and determine the optimal value of the rotation matrix and the optimal value of the translation vector.

[0081] Further, the calibration loss function is constructed using the following formula:

[0082] f(x) = 0.5 * norm((transpose(n) * (π(R|t)P) - q));

[0083] where f(x) is used to represent the calibration loss function; norm() represents a vector norm calculation function; transpose() represents a transpose function; n represents the set of normal vectors; π represents a projection function; R represents a rotation matrix; t represents a translation vector; P represents the point cloud depth continuous line feature; and q represents the set of neighboring pixel points. Further, in some non-limiting embodiments, the iterative algorithm can be selected from: gradient descent algorithm, Newton algorithm, Gauss-Newton algorithm, Levenberg-Marquardt algorithm. In some other non-limiting embodiments, the iterative algorithm can also be selected from other algorithms that can calculate the optimal value of a variable through iteration.

[0084] where the preset termination condition can be a preset maximum number of iterations or a preset minimum function value of the calibration loss function.

[0085] Iterative algorithms such as the gradient descent algorithm, Newton's algorithm (Newton-Raphson method), Gauss-Newton algorithm, and Levenberg-Marquardt algorithm are algorithms that can continuously derive new values of variables from old values. Corresponding to the iterative method is the direct method (or called the one-time solution method), that is, solving the problem at once. The iterative algorithm is a basic method for a computer to solve problems. It takes advantage of the fast computing speed of the computer and its suitability for repetitive operations, and allows the computer to repeatedly execute a set of instructions (or certain steps). Each time these instructions (or these steps) are executed, a new value of the variable is derived from its original value.

[0086] Taking the gradient descent algorithm as an example, its iterative calculation process is to solve for the minimum value along the direction of the gradient descent (it can also solve for the maximum value along the direction of the gradient ascent). Generally, when the gradient vector is 0, it means that an extreme point is reached, and at this time, the magnitude of the gradient is also 0. Therefore, in specific implementation, if the gradient descent algorithm is used to solve for the optimal values of the rotation matrix and the translation vector, the termination condition of the algorithm iteration can be set to the magnitude of the gradient vector approaching 0 or a suitable number of iterations.

[0087] In specific implementation, the basic principles of using Newton's algorithm, Gauss-Newton algorithm, and Levenberg-Marquardt algorithm to solve for the optimal values of variables are similar to the above, that is, they all find the point where the derivative approaches 0 infinitely (at this time the function converges) through successive iterations, and then determine the values of the rotation matrix and the translation vector at this time as the optimal values of the rotation matrix and the translation vector. Here, the specific solution process of the algorithm will not be elaborated.

[0088] In the embodiment of the present invention, by constructing a calibration loss function and using a preset iterative algorithm and a preset termination condition, the calibration loss function is minimized to determine the optimal values of the rotation matrix and the translation vector, thereby achieving the effect of fully automatically completing the external parameter calibration.

[0089] Refer to Figure 3 , Figure 3 is the flow chart of the second automatic calibration method in the embodiment of the present invention. The second automatic calibration method may include steps S31 to S36, and each step will be described below.

[0090] In step S31, the original point cloud data and the original image data collected for the same regional scene at the same timestamp are determined.

[0091] In step S32, the point cloud depth continuous line feature is determined according to the original point cloud data, and the image depth continuous line feature is determined according to the original image data.

[0092] In step S33, a projection function is used to determine a set of two-dimensional pixel points of the point cloud depth continuous line feature according to a preset initial value of the rotation matrix and a preset initial value of the translation vector, and the set of two-dimensional pixel points of the point cloud depth continuous line feature is determined.

[0093] Among them, the formula used in step S33 is as follows:

[0094] p = π[(R|t)P];

[0095] Among them, p represents the set of two-dimensional pixel points; π represents the projection function; R represents the rotation matrix; t represents the translation vector; P represents the point cloud depth continuous line feature.

[0096] In step S34, the Oct-tree algorithm is used. For each two-dimensional pixel point in the set of two-dimensional pixel points, the nearest neighboring pixel point is found from the image depth continuous line feature, and the normal vector of the neighboring pixel point is determined to obtain a set of neighboring pixel points and a set of normal vectors of the neighboring pixel points.

[0097] In step S35, a calibration loss function is constructed based on the set of neighboring pixel points, the set of normal vectors, the initial value of the rotation matrix, and the initial value of the translation vector.

[0098] Among them, the formula used to construct the calibration loss function is as follows:

[0099] f(x) = 0.5 * norm((transpose(n) * (π(R|t)P) - q));

[0100] Among them, f(x) is used to represent the calibration loss function; norm() represents the vector norm calculation function; transpose() represents the transpose function; n represents the set of normal vectors; π represents the projection function; R represents the rotation matrix; t represents the translation vector; P represents the point cloud depth continuous line feature; q represents the set of neighboring pixel points.

[0101] In step S36, the Levenberg-Marquardt algorithm and a preset termination condition are used to minimize the calibration loss function, and the optimal value of the rotation matrix and the optimal value of the translation vector are determined.

[0102] In specific implementation, for more detailed content about steps S31 to S36, please refer to the previous text for execution, and details are not described here again.

[0103] Refer to Figure 4 , Figure 4 is a schematic structural diagram of an automatic calibration device in an embodiment of the present invention. The automatic calibration device may include:

[0104] An original data acquisition module 41, configured to determine original point cloud data and original image data collected for the same regional scene with the same timestamp;

[0105] A depth continuous line feature determination module 42, configured to determine a point cloud depth continuous line feature according to the original point cloud data, and determine an image depth continuous line feature according to the original image data;

[0106] A two-dimensional pixel point set determination module 43, configured to determine a two-dimensional pixel point set of the point cloud depth continuous line feature according to a preset initial rotation matrix value and a preset initial translation vector value;

[0107] A neighboring pixel point set determination module 44, configured to, for each two-dimensional pixel point in the two-dimensional pixel point set, find the neighboring pixel point closest to the two-dimensional pixel point from the image depth continuous line feature, and determine the normal vector of the neighboring pixel point, so as to obtain a neighboring pixel point set and a normal vector set of the neighboring pixel points;

[0108] An optimal value calculation module 45, configured to determine an optimal value of the rotation matrix and an optimal value of the translation vector based on the neighboring pixel point set and the normal vector set.

[0109] For the principle, specific implementation and beneficial effects of this automatic calibration device, please refer to the foregoing and Figures 1 to 3 the related description of the automatic calibration method shown herein, which will not be elaborated herein.

[0110] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is processed and run, it executes the steps of the above automatic calibration method. The computer-readable storage medium may include a non-volatile memory or a non-transitory memory, and may also include an optical disc, a mechanical hard disk, a solid-state drive, etc.

[0111] Specifically, in the embodiments of the present invention, the processor may be a central processing unit (CPU for short), and the processor may also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), field programmable gate arrays (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0112] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM for short), a programmable read-only memory (PROM for short), an erasable programmable read-only memory (EPROM for short), an electrically erasable programmable read-only memory (EEPROM for short), or a flash memory. The volatile memory may be a random access memory (RAM for short), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM for short) are available, such as static random access memory (SRAM for short), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM for short), double data rate synchronous dynamic random access memory (DDR SDRAM for short), enhanced synchronous dynamic random access memory (ESDRAM for short), synchronous link dynamic random access memory (SLDRAM for short), and direct rambus random access memory (DR RAM for short).

[0113] An embodiment of the present invention further provides a terminal, including a memory and a processor. A computer program capable of running on the processor is stored on the memory. When the processor runs the computer program, it executes the steps of the above-mentioned automatic calibration method. The terminal may include, but is not limited to, terminal devices such as mobile phones, computers, and tablets, and may also be a server, a cloud platform, etc.

[0114] It should be understood that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article indicates that the associated objects before and after are in an "or" relationship.

[0115] In the embodiments of the present application, "a plurality of" refers to two or more.

[0116] The descriptions such as the first and the second in the embodiments of the present application are only for schematic and distinguishing the described objects, without an order, and do not represent a special limitation on the number of devices in the embodiments of the present application, and cannot constitute any limitation on the embodiments of the present application.

[0117] It should be noted that the sequence numbers of the steps in this embodiment do not represent the limitation of the execution sequence of each step.

[0118] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.

Claims

1. An automatic calibration method, characterized in that, including: determining the original point cloud data and the original image data collected for the same regional scene using the same timestamp; determining the point cloud depth continuous line feature according to the original point cloud data, and determining the image depth continuous line feature according to the original image data; determining a two-dimensional pixel point set of the point cloud depth continuous line feature according to a preset initial rotation matrix value and a preset initial translation vector value; for each two-dimensional pixel point in the two-dimensional pixel point set, finding the nearest neighboring pixel point to the two-dimensional pixel point from the image depth continuous line feature, and determining the normal vector of the neighboring pixel point, so as to obtain a neighboring pixel point set and a normal vector set of the neighboring pixel points; determining the optimal value of the rotation matrix and the optimal value of the translation vector based on the neighboring pixel point set and the normal vector set; wherein, determining the optimal value of the rotation matrix and the optimal value of the translation vector based on the neighboring pixel point set and the normal vector set includes: constructing a calibration loss function using the neighboring pixel point set, the normal vector set, the initial rotation matrix value, and the initial translation vector value; using a preset iterative algorithm and a preset termination condition to minimize the calibration loss function, and determining the optimal value of the rotation matrix and the optimal value of the translation vector.

2. The method according to claim 1, wherein Determining the point cloud depth continuous line feature according to the original point cloud data includes: determining multiple frames of point cloud data collected for the same regional scene within a preset time period before and / or after the acquisition timestamp of the original point cloud data; constructing cumulative dense point cloud data for the multiple frames of point cloud data, and the positive direction of the coordinate system origin of the cumulative dense point cloud data is the same as the positive direction of the coordinate system origin of the original point cloud data; extracting the point cloud depth continuous line feature from the cumulative dense point cloud data.

3. The method according to claim 2, wherein Extracting the point cloud depth continuous line feature from the cumulative dense point cloud data includes: performing downsampling and rasterization on the cumulative dense point cloud data to obtain multiple grids; for the point cloud data in each grid, using the random sample consensus algorithm RANSAC for fitting to obtain multiple planes; taking the lines where the multiple planes intersect as the point cloud depth continuous line feature.

4. The method according to claim 1, wherein Determining the image depth continuous line feature according to the original image data includes: removing distortion from the original image data and converting it into a grayscale image; using the Laplace operator to perform edge extraction on the grayscale image to obtain the image depth continuous line feature.

5. The method according to claim 1, wherein using the following formula to determine the two-dimensional pixel point set of the point cloud depth continuous line feature: ; Among them, represents the set of two-dimensional pixel points; represents the projection function; represents the initial value of the rotation matrix; represents the initial value of the translation vector; represents the point cloud depth continuous line feature.

6. The method according to claim 1, wherein the search algorithm for finding the nearest neighboring pixel point to the two-dimensional pixel point from the image depth continuous line feature is selected from: Quad-tree algorithm, Oct-tree algorithm, k-nearest neighbor algorithm.

7. The method according to claim 1, characterized in that constructing the calibration loss function using the following formula: ; wherein, is used to represent the calibration loss function; represents a vector norm calculation function; represents a transpose function; represents the set of normal vectors; represents a projection function; represents a rotation matrix; represents a translation vector; represents the point cloud depth continuous line feature; represents the set of adjacent pixel points.

8. The method according to claim 1, wherein the iterative algorithm is selected from: gradient descent algorithm, Newton algorithm, Gauss-Newton algorithm, Levenberg-Marquardt algorithm.

9. An automatic calibration device, characterized in that, including: an original data determination module, configured to determine the original point cloud data and the original image data collected for the same regional scene using the same timestamp; A depth continuous line feature determination module, configured to determine a point cloud depth continuous line feature according to the original point cloud data and an image depth continuous line feature according to the original image data; A two-dimensional pixel point set determination module, configured to determine a two-dimensional pixel point set of the point cloud depth continuous line feature according to a preset initial rotation matrix value and a preset initial translation vector value; A neighboring pixel point set determination module, configured to, for each two-dimensional pixel point in the two-dimensional pixel point set, find a neighboring pixel point closest to the two-dimensional pixel point from the image depth continuous line feature and determine a normal vector of the neighboring pixel point, so as to obtain a neighboring pixel point set and a neighboring pixel point normal vector set; An optimal value calculation module, configured to determine an optimal value of a rotation matrix and an optimal value of a translation vector based on the neighboring pixel point set and the normal vector set; Wherein, the optimal value calculation module executes the steps of: constructing a calibration loss function by using the neighboring pixel point set, the normal vector set, the initial rotation matrix value, and the initial translation vector value; using a preset iterative algorithm and a preset termination condition to minimize the calibration loss function and determine the optimal value of the rotation matrix and the optimal value of the translation vector.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by a processor, it executes the steps of the automatic calibration method according to any one of claims 1 to 8.

11. A terminal, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored on the memory, characterized in that, When the processor runs the computer program, it executes the steps of the automatic calibration method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Calibration method and device, computer equipment and storage medium

    CN113077523A

  • Camera and laser radar calibration method and system based on end-to-end and medium

    CN113160330A