Camera-lidar calibration method and device
Through neural networks, the data position changes of cameras and lidars are predicted and calibrated, and the time and space mismatch between heterogeneous sensors is solved, and efficient data synchronization and calibration of autonomous driving systems is achieved.
Patent Information
- Application Number
- CN202410576833.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-05-10
- Publication Date
- 2025-05-30
AI Technical Summary
In autonomous driving technology, time mismatch and space mismatch between cameras and lidars lead to difficulties in data synchronization and calibration, and the prior art requires additional hardware support.
By designing a neural network, using the data acquired by the camera and lidar, the ground edge images and road marking point clouds are extracted, the position changes of the road marking point clouds relative to the ground edge images are predicted, and the calibration matrix is generated to calibrate the camera and lidar.
Time synchronization and spatial calibration between the camera and lidar are achieved without additional hardware, improving the data accuracy and reliability of autonomous driving systems.
Smart Images

Figure CN120065184A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit of Korean Patent Application No. 10 - 2023 - 0171455, filed on November 30, 2023, which is hereby incorporated by reference herein in its entirety. Technical Field
[0003] The present disclosure relates to methods and apparatuses for camera - lidar calibration. Background Art
[0004] Recently, with the continuous development of computer vision technology based on deep neural networks in autonomous driving technology, various artificial intelligence (AI) models for object detection, semantic segmentation, depth map estimation, or lane detection have been explored and studied.
[0005] Specifically, methods for performing autonomous driving by inputting data obtained by a camera and data obtained by a lidar into an AI model have been continuously explored and studied. However, since a camera and a lidar are different sensors from each other, the camera and the lidar can be installed at different positions in a vehicle, and thus a spatial calibration procedure is required. To solve the spatial mismatch between heterogeneous sensors, internal parameters and external parameters can be used. In addition, the camera and the lidar may be different from each other at the time point of acquiring data or at the time point of completing data processing, and thus time synchronization work is required between the camera and the lidar. To solve the time mismatch between heterogeneous sensors, specific hardware can be used, or timestamps related to the time point of acquiring data of each sensor can be used. Summary of the Invention
[0006] The present disclosure relates to methods and apparatuses for camera - lidar calibration, and in a specific embodiment, relates to a neural network design for calibration between a camera and a lidar, and methods and apparatuses for camera - lidar calibration.
[0007] Embodiments provided by the present disclosure can solve problems that occur in the prior art while maintaining the advantages achieved by the prior art.
[0008] One aspect of the present disclosure provides a method for compensating for a time mismatch between a camera and a lidar through a neural network and an apparatus for the method.
[0009] Another aspect of the present disclosure provides a method for time - synchronizing between a camera and a lidar and an apparatus for the method, the method and the apparatus using a camera, a lidar, and a processor provided inside a vehicle without additional hardware.
[0010] Another aspect of the present disclosure provides a method for designing a neural network to perform temporal calibration on a camera and a lidar and a device for the method, and a device for the method.
[0011] Another aspect of the present disclosure provides a method for substituting a temporal mismatch between a camera and a lidar into a spatial mismatch between the camera and the lidar and a device for the method.
[0012] The technical problems to be solved by the present disclosure are not limited to the above problems, and those skilled in the art to which the present disclosure pertains will clearly understand any other technical problems not mentioned herein from the following description.
[0013] According to an embodiment of the present disclosure, a method for camera-lidar calibration may include: obtaining, by an acquisition device, an image captured by a camera at a specific time point and lidar point cloud projected onto a camera coordinate system at the specific time point; extracting, by an extraction device, a ground edge image corresponding to an edge of the ground from the image; extracting, by the extraction device, a road marking point cloud from the lidar point cloud, the road marking point cloud representing a point cloud of a road marking on the ground; generating, by a generation device, a first change indicating a predicted position change by predicting a position change of the road marking point cloud relative to the ground edge image by using a neural network; and calibrating, by a calibration device, the camera and the lidar based on the first change.
[0014] According to an embodiment, the method may further include: obtaining, by the acquisition device, a training ground edge image; obtaining, by the acquisition device, a training road marking point cloud matching the training ground edge image; generating, by the generation device, a training transformed point cloud by transforming the training road marking point cloud by using a second change; generating, by the generation device, a third change indicating a predicted position change by predicting a position change of the training transformed point cloud relative to the training ground edge image via a neural network; and training, by a training device, the neural network by performing regression analysis based on the second change and the third change.
[0015] According to an embodiment, extracting the road marking point cloud may include: extracting, by the extraction device, a ground point cloud corresponding to the ground from the lidar point cloud; and extracting, by the extraction device, a road marking point cloud having an intensity greater than or equal to a specific intensity from the ground point cloud.
[0016] According to an embodiment, extracting the ground point cloud may include extracting, by the extraction device, the ground point cloud by a ground estimation algorithm.
[0017] According to an embodiment, extracting the ground point cloud may include extracting, by the extraction device, the ground point cloud by a normal estimation algorithm.
[0018] According to an embodiment, extracting a ground edge image may include: generating a grayscale image by a generating device by transforming an image into a grayscale image; extracting an edge image by an extracting device by applying an edge filtering algorithm to the grayscale image; and extracting a ground edge image by the extracting device by setting a ground of the edge image as a region of interest (RoI) and cropping a remaining region of the edge image other than the region of interest.
[0019] According to an embodiment, extracting a ground edge image may include: extracting a ground edge image by an extracting device by extracting a part of feature points of an image via a neural network.
[0020] According to an embodiment, the first transformation may include elements of a rotation matrix and elements of a translation matrix, and calibration may include calibrating a camera and a lidar by a calibration device by multiplying coordinates of a road marking point cloud by an inverse matrix of the first transformation.
[0021] According to an embodiment, the neural network may be a convolutional neural network (CNN) or a multi-layer perceptron (MLP).
[0022] According to an embodiment, a device for camera-lidar calibration may include a camera; a lidar; an acquisition device configured to acquire an image captured by the camera at a specific time point and a lidar point cloud projected onto a camera coordinate system at the specific time point; an extraction device that extracts a ground edge image corresponding to an edge of the ground from the image and extracts a road marking point cloud representing a point cloud of a road marking on the ground from the lidar point cloud; a generation device configured to generate a first transformation indicating a predicted position change by predicting a position change of the road marking point cloud relative to the ground edge image via a neural network; and a calibration device for calibrating the camera and the lidar based on the first transformation.
[0023] According to an embodiment, the acquisition device may acquire a training ground edge image and acquire a training road marking point cloud matching the training ground edge image, the generation device may generate a training transformed point cloud by transforming the training road marking point cloud by means of a second transformation, and generate a third transformation indicating a predicted position change by predicting a position change of the training transformed point cloud relative to the training ground edge image via a neural network, and the device may further include a training device configured to train the neural network by performing regression analysis based on the second transformation and the third transformation.
[0024] According to an embodiment, the extraction device may extract a ground point cloud corresponding to the ground from the lidar point cloud, and extract a road marking point cloud having an intensity greater than or equal to a specific intensity from the ground point cloud.
[0025] According to an embodiment, the extraction device may extract the ground point cloud by a ground estimation algorithm.
[0026] According to an embodiment, the extraction device may extract ground point clouds through a normal estimation algorithm.
[0027] According to an embodiment, the generation device may generate a grayscale image by converting an image into a grayscale image. The extraction device may extract an edge image by applying an edge filtering algorithm to the grayscale image, and extract a ground edge image by setting the ground of the edge image as a region of interest (RoI) and cropping the remaining region of the edge image except the region of interest.
[0028] According to an embodiment, the extraction device may extract a ground edge image by extracting a part of the feature points of an image via a neural network.
[0029] According to an embodiment, the first variation may include elements of a rotation matrix and elements of a translation matrix, and the calibration device may calibrate the camera and lidar by multiplying the coordinates of the road marker point cloud by the inverse matrix of the first variation.
[0030] According to an embodiment, the neural network may be a convolutional neural network (CNN) or a multi-layer perceptron (MLP).
[0031] The features of the present disclosure briefly described are provided as exemplary aspects of the detailed description of the present disclosure to be described below, and the scope of the present disclosure is not limited thereto. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] From the following detailed description in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent:
[0033] Figure 1 is a flowchart showing a method for camera-lidar calibration according to an embodiment of the present disclosure;
[0034] Figure 2 is a view showing a preprocessing process of data in a method for camera-lidar calibration according to an embodiment of the present disclosure;
[0035] Figure 3 is a view showing a method for camera-lidar calibration according to an embodiment of the present disclosure;
[0036] Figure 4 is a flowchart showing a method for training a neural network for camera-lidar calibration according to an embodiment of the present disclosure;
[0037] Figure 5 is a view showing a method for training a neural network for camera-lidar calibration according to an embodiment of the present disclosure;
[0038] Figure 6 is a view showing a method for training a neural network for camera-lidar calibration according to an embodiment of the present disclosure;
[0039] Figure 7A is a view showing a method for camera-lidar calibration according to an embodiment of the present disclosure;
[0040] Figure 7B is a view showing a method for camera-lidar calibration according to an embodiment of the present disclosure;
[0041] Figure 8 is a block diagram showing an apparatus for camera-lidar calibration according to an embodiment of the present disclosure; and
[0042] Figure 9 is a block diagram showing a computing system for performing camera-lidar calibration according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0043] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily reproduce the inventive concept. However, the present disclosure can be implemented in various forms and is not limited to the embodiments described herein.
[0044] In the following description of the present disclosure, details of related known configurations or functions may be omitted if they may obscure the subject matter of the present disclosure. In addition, for the purpose of clearly describing the present disclosure, parts not related to the present disclosure are omitted, and throughout the specification, like reference numerals will be assigned to like parts.
[0045] When a certain component is "coupled to", "connected to" or "joined to" another component, a certain component may be directly coupled to or connected to another component, and a third component may be electrically "coupled", "connected" or "joined" between a certain component and another component. It should be understood that when used herein, the terms "comprises", "comprising", "includes" and / or "including" specify the presence of the stated elements and / or components, but do not preclude the presence or addition of one or more other elements and / or components. In addition, when a certain component includes or has another component, the other component is not excluded, but rather the other component is also included, unless otherwise specifically stated.
[0046] In the present disclosure, the terms "first" and "second" are used to distinguish one component from another. Unless otherwise specifically stated, the terms "first" and "second" do not limit the order or importance between components. Thus, within the scope of the present disclosure, a "first component" according to an embodiment may be referred to as a "second component" according to another embodiment, and a "second component" according to an embodiment may be referred to as a "first component" according to another embodiment.
[0047] According to the present disclosure, components are provided to be distinct from each other to clearly describe the features of each component, and do not mean that the components are separated from each other. In other words, multiple components are integrated and implemented in one hardware or software unit. Alternatively, one component is split and implemented in multiple hardware or software units. Thus, unless otherwise specifically stated, even embodiments having components disposed integrally or separately from each other also fall within the scope of the present disclosure.
[0048] In the present disclosure, the components described according to various embodiments do not refer to essential components, and some components may be provided selectively. Thus, embodiments including a subset of the components described according to an embodiment are included in the present disclosure. In addition, embodiments including other components in addition to the components described according to various embodiments fall within the scope of the present disclosure.
[0049] In the present disclosure, the positional relationships (e.g., upper, lower, left, or right) indicated in this specification are provided for illustrative purposes only. When the drawings of the present disclosure are shown in reverse, the positional relationships described in the present disclosure may be interpreted in reverse.
[0050] In the present disclosure, each of the phrases such as "A or B", "at least one of A and B", "at least one of A or B", "at least one of A, B, or C", "at least one of A, B, and C", or "at least one of A, B, or C", etc. may include any one or all combinations of the items arranged together in the relevant phrase of the phrase.
[0051] Hereinafter, embodiments of the present disclosure will be described with reference to Figures 1 to 9 description of the embodiments of the present disclosure.
[0052] Figure 1 is a flowchart showing a method for camera-lidar calibration according to an embodiment of the present disclosure.
[0053] A camera and a lidar are mutually different sensors. Synchronized operation is required between heterogeneous sensors so that the heterogeneous sensors can be used together. For example, a camera and a lidar need to be synchronized with each other spatially and temporally.
[0054] Spatial mismatch between heterogeneous sensors may be caused by differences in the mounting positions among sensors disposed inside a vehicle and by a mismatch between the inherent coordinate systems of the sensors. The spatial mismatch between the heterogeneous sensors can be compensated by estimating internal parameters and external parameters.
[0055] Temporal mismatch between heterogeneous sensors may be caused by differences among the sensors in terms of a time point for acquiring data or a time point for completing data processing. For example, a lidar sensor may be a physically rotating sensor, so data at a time point when the rotation angle of the lidar sensor is 0° and data at a time point when the rotation angle of the lidar sensor is 360° may be different from a camera image. Specifically, assume that the lidar takes a time period of 0.9 seconds to rotate 360°, and an image is obtained by the camera at the time point after 0.9 seconds. Then, the point cloud of the lidar (lidar point cloud) at the time point after 0.9 seconds may match the image of the camera, but the lidar point cloud at 0.1 second may not match the camera image at 0.9 seconds.
[0056] Temporal mismatch between heterogeneous sensors results in using additional hardware for time synchronization or using time stamps related to the time points for acquiring data of each sensor. However, estimating the motion of lidar points requires more complex techniques compared to compensating for the spatial mismatch between sensors. In addition, when using new hardware to solve the temporal mismatch between heterogeneous sensors, problems of additional costs and reliability of the hardware may occur.
[0057] According to an embodiment of the present disclosure, a method for camera-lidar calibration does not require additional hardware to solve the temporal mismatch between a camera and a lidar. In addition, according to an embodiment of the present disclosure, in the method for camera-lidar calibration, the temporal mismatch between the camera and the lidar is replaced with a spatial mismatch between the camera and the lidar, thereby solving the problem of temporal mismatch.
[0058] See Figure 1, according to an embodiment of the present disclosure, in a method for camera-lidar calibration, an image captured by a camera at a specific time point and lidar point cloud projected onto a camera coordinate system at the specific time point may be obtained in S110. Specifically, an acquisition device may obtain the image captured by the camera at the specific time point and the lidar point cloud projected onto the camera coordinate system. In the method for camera-lidar calibration, by obtaining the lidar point cloud projected onto the camera coordinate system from the lidar coordinate system, the spatial mismatch between the camera and the lidar can be solved. However, if the camera and the lidar are not time-calibrated, even if the image and the lidar point cloud projected onto the camera coordinate system are obtained at the same time point, when the lidar point cloud projected onto the camera coordinate system is projected onto the image captured by the camera, the lidar point cloud projected onto the camera coordinate system may also be mismatched with the image captured by the camera.
[0059] According to the method for camera-lidar calibration, in S120, a ground edge image corresponding to the edge of the ground may be extracted from the image. Specifically, an extraction device may extract the ground edge image corresponding to the edge of the ground from the image. Additionally, according to the method for camera-lidar calibration, the image may be transformed into a grayscale image. Furthermore, according to the method for camera-lidar calibration, an edge image may be extracted by applying an edge filtering algorithm to the grayscale image. Moreover, according to the method for camera-lidar calibration, the ground in the edge image may be set as a region of interest (RoI), and the remaining region in the edge image except the RoI may be cropped, thereby extracting the ground image. Additionally, according to the method for camera-lidar calibration, the ground edge image may be extracted from the image by a neural network.
[0060] According to the method for camera-lidar calibration, in S130, a road marking point cloud representing the point cloud of road markings on the ground may be extracted from the lidar point cloud. Specifically, an extraction device may extract the road marking point cloud from the lidar point cloud, which represents the point cloud of road markings on the ground.
[0061] The lidar point cloud projected onto the camera coordinate system may include static point cloud (e.g., ground, or point cloud for a building) and dynamic point cloud (e.g., point cloud of a vehicle or a pedestrian). Due to the time mismatch between the camera and the lidar, a consistent position transformation is applied to the static point cloud mismatched with the image, such that the image matches the transformed point cloud, thereby solving the time mismatch between the camera and the lidar. In other words, since the time mismatch between the camera and the lidar is replaced by the spatial mismatch between the camera and the lidar, it is easy to solve the time mismatch between the camera and the lidar.
[0062] According to the method for camera-lidar calibration, ground point clouds corresponding to the ground can be extracted from the lidar point cloud. In other words, according to the method for camera-lidar calibration, at least a part of the static point cloud can be extracted from the lidar point cloud. In addition, according to the method for camera-lidar calibration, a ground estimation algorithm can be used to extract only the ground point cloud corresponding to the ground from the lidar point cloud. Additionally, according to the method for camera-lidar calibration, the ground point cloud can be extracted by a normal estimation algorithm. For example, the method for camera-lidar calibration can include extracting the ground point cloud from the lidar point cloud by extracting points having values that allow the normal vector to approach (0, 0, 1). However, the normal vector is not limited to (0, 0, 1). For example, various vectors can be used as the normal vector as long as these vectors correspond to the normal of the ground.
[0063] In addition, according to the method for camera-lidar calibration, road marking point clouds having an intensity greater than or equal to a specific intensity can be extracted from the ground point cloud. Specifically, road markings can typically be represented by a brighter color (such as white), and the part of the ground point cloud corresponding to the road marking can have an intensity greater than that of the ground (e.g., asphalt). Therefore, if point clouds having an intensity greater than or equal to a specific intensity are filtered from the ground point cloud, the point cloud of the road marking can be extracted.
[0064] According to the method for camera-lidar calibration, in S140, the position change of the road marking point cloud relative to the ground edge image is predicted by a neural network, and a first change representing the predicted position change is generated. Specifically, the generating device can predict the position change of the road marking point cloud relative to the ground edge image by a neural network and generate a first change indicating the predicted position change. Since the camera and the lidar are not calibrated in time, the ground edge image and the road marking point cloud may not match each other. Therefore, according to the method for camera-lidar calibration, the neural network is trained to predict the position change of the road marking point cloud relative to the ground edge image. The first change indicating the predicted position change can be represented in the form of a matrix. Specifically, the first change can have a matrix including elements of a rotation matrix and elements of a translation matrix.
[0065] Formula 1
[0066]
[0067] Including the element r of Equation 1 11 to r 33 and the element t 1 to t 3 The matrix can be an embodiment for the first change. Among the elements of the first change, r 11 to r33 can correspond to the elements of a rotation matrix, and t 1 to t 3 can correspond to the elements of a translation matrix.
[0068] The neural network can be a convolutional neural network (CNN) or a multi-layer perceptron (MLP), but the present disclosure is not limited thereto.
[0069] To train the neural network, according to the method of camera-lidar calibration, a ground edge image for training (or a training ground edge image) can be obtained. In addition, according to the method of camera-lidar calibration, a road marking point cloud for training (or a training road marking point cloud) can be obtained to match the training ground edge image. Additionally, according to the method of camera-lidar calibration, the training road marking point cloud is transformed by a second transformation to generate a transformed point cloud for training (or a training transformed point cloud). Additionally, to perform data augmentation of training data, the training road marking point cloud that matches the training ground edge image is transformed by a second transformation to generate a training transformed point cloud. In this case, the second transformation can be arbitrarily determined within a preset range.
[0070] Furthermore, according to the method for camera-lidar calibration, the neural network predicts the position change of the training transformed point cloud relative to the training ground edge image, and generates a third transformation indicating the predicted position change.
[0071] Additionally, according to the method of camera-lidar calibration, regression analysis can be performed based on the second transformation and the third transformation, thereby training the neural network. Details related to the training of the neural network will be described later.
[0072] According to the method of camera-lidar calibration, in S150, the camera and the lidar can be calibrated based on the first transformation. Specifically, the calibration device can calibrate the camera and the lidar based on the first transformation. According to the method of camera-lidar calibration, the coordinates of the points of the road marking point cloud can be multiplied by the inverse matrix of the first transformation to calibrate the camera and the lidar. Additionally, the coordinates of the points of the road marking point cloud that do not match the ground edge image can be multiplied by the inverse matrix of the first transformation so that the ground edge image matches the reversed road marking point cloud. Since the ground edge image matches the reversed road marking point cloud, the camera and the lidar can be calibrated in time.
[0073] Figure 2 is a view showing a preprocessing process of data in the method for camera-lidar calibration according to an embodiment of the present disclosure.
[0074] In order to perform time calibration on a camera and a lidar, it is necessary to preprocess the data obtained by the camera and the data obtained by the lidar.
[0075] Generally, the driving environment is dynamic. Thus, if a transformation is uniformly applied to the lidar point cloud projected onto the camera coordinate system, the mismatch between the lidar point cloud and the image captured by the camera may not be resolved. As described above, the lidar point cloud projected onto the camera coordinate system may include a static point cloud (e.g., the point cloud of the ground, buildings) and a dynamic point cloud (e.g., the point cloud of a vehicle or a pedestrian). Applying a consistent transformation to the static point cloud can resolve the time mismatch between the static point cloud and the image. Thus, the time mismatch between the camera and the lidar can be resolved by extracting the parts corresponding to the ground from the lidar point cloud and the image captured by the camera. However, regarding the entire point cloud corresponding to the ground and the entire ground image, it may be difficult to find feature points to determine whether the entire point cloud corresponding to the ground matches the entire ground image.
[0076] For the ground, feature points can be easily extracted from lane or road markings. Specifically, the intensity of the point cloud of the painted part on asphalt can be greater than that of asphalt. Therefore, the road marking point cloud 213, which is the point cloud corresponding to the road marking, can be extracted from the lidar point cloud 211 projected onto the camera coordinate system. The lidar point cloud projected onto the camera coordinate system can be used to resolve the spatial mismatch between the camera and the lidar.
[0077] By applying a ground estimation algorithm to the lidar point cloud 211, the ground point cloud 212 can be extracted from the lidar point cloud 211 projected onto the camera coordinate system. In addition, although not shown, the ground point cloud 212 can be extracted from the lidar point cloud 211 projected onto the camera coordinate system by applying a normal estimation algorithm to the lidar point cloud 211.
[0078] If only the ground point cloud is extracted as described above, feature points for matching with the image may not be found. Therefore, the road marking point cloud 213 can be extracted from the ground point cloud 212. Specifically, the road marking point cloud 213 can be extracted by extracting the point cloud with an intensity greater than or equal to a specific intensity from the ground point cloud 212.
[0079] In order to compare with the road marking point cloud 213, a ground edge image 222, which is an image corresponding to the edge of the ground, can be extracted from the image 221 captured by the camera. In order to extract the ground edge image 222 from the image 221, an image processing method for converting the image into a grayscale image and extracting edges from the grayscale image can be applied. Additionally, the ground edge image can be directly extracted through a neural network.
[0080] Figure 3 It is a view showing a method for camera - lidar calibration according to an embodiment of the present disclosure.
[0081] Refer to Figure 3 , the road marking point cloud 310 and the ground edge image 320 can be input into a neural network. In addition, the neural network can predict a first transformation T_pred corresponding to the positional difference between the road marking point cloud 310 and the ground edge image 320. As described above, the first transformation T_pred can be a matrix including the elements R_pred of the rotation matrix and the elements t_pred of the translation matrix.
[0082] The first transformation can be predicted once by the neural network, or can be predicted several times. Specifically, the neural network can predict at least a part of the change (positional change) in the position between the road marking point cloud 310 and the ground edge image 320 several times. For example, the neural network can predict the (1 - 1)th transformation, which is at least a part of the positional change between the road marking point cloud 310 and the ground edge image 320. In addition, the neural network can predict the (1 - 2)th transformation after predicting the (1 - 1)th transformation, which is at least a part of the positional change. In addition, the neural network can predict the (1 - 3)th transformation after predicting the (1 - 2)th transformation, which is at least a part of the positional change. Assuming that the (1 - 3)th transformation is predicted, the equation "First transformation = (the (1 - 1)th transformation) * (the (1 - 2)th transformation) * (the (1 - 3)th transformation)" can be established.
[0083] An inverse transformation is performed on the road marking point cloud 310 based on the predicted first transformation, thereby obtaining a road marking point cloud 330 that matches the ground edge image 320. In other words, the camera and the lidar can be calibrated in time. Specifically, according to the method for camera - lidar calibration, the points of the road marking point cloud are multiplied by the inverse matrix of the first transformation, thereby calibrating the camera and the lidar.
[0084] Figure 4 It is a flowchart showing a method for training a neural network for camera - lidar calibration according to an embodiment of the present disclosure.
[0085] Refer to Figure 4 , in the method for camera - lidar calibration according to an embodiment of the present disclosure, the neural network can be trained to predict the positional difference between the ground edge image and the road marking point cloud.
[0086] Specifically, according to the method for camera - lidar calibration, in S410, a ground edge image for training (training ground edge image) can be obtained. Specifically, the acquisition device can obtain the training ground edge image.
[0087] In addition, according to the method for camera-lidar calibration, in S420, a road marking point cloud for training (training road marking point cloud) that matches the training ground edge image can be obtained. Specifically, the obtaining device can obtain the training road marking point cloud that matches the training ground edge image.
[0088] In addition, according to the method for camera-lidar calibration, in S430, the training road marking point cloud can be transformed by a second transformation to generate a training transformed point cloud. Specifically, the generating device can transform the training road marking point cloud by a second transformation to generate the training transformed point cloud. According to the method for camera-lidar calibration, data for training can be generated by intentionally transforming the training road marking point cloud that matches the training ground edge image to enhance the data.
[0089] According to the method for camera-lidar calibration, in S440, a third transformation indicating a predicted position change can be generated by predicting, via a neural network, the position change of the training transformed point cloud relative to the training ground edge image. Specifically, the generating device can predict, via the neural network, the position change of the training transformed point cloud relative to the training ground edge image to generate the third transformation indicating the predicted position change. The third transformation can be a value predicted by the neural network and can be different from the second transformation.
[0090] According to the method for camera-lidar calibration, in S450, the neural network can be trained by performing regression analysis based on the second transformation and the third transformation. Specifically, the training device can train the neural network by performing regression analysis based on the second transformation and the third transformation.
[0091] Figure 5 FIG. is a view showing a manner of training a neural network to perform camera-lidar calibration according to an embodiment of the present disclosure.
[0092] Reference Figure 5 , the neural network according to an embodiment of the present disclosure can be designed to predict a third transformation T_pred based on the training transformed road marking point cloud 520 and the training ground edge image 530. The third transformation can be a matrix corresponding to the position difference between the training transformed point cloud 520 and the training ground edge image 530. Specifically, the third transformation can be a matrix including elements R_pred of a rotation matrix and elements t_pred of a translation matrix.
[0093] According to an embodiment of the present disclosure, in a method for camera-lidar calibration, a training road marking point cloud 510 can be transformed into a training transformed point cloud 520 for data augmentation. Specifically, the training transformed point cloud 520 can be generated by transforming the training road marking point cloud 510 with a second transformation T_GT. The second transformation can be a matrix including elements R_GT of a rotation matrix and elements t_GT of a translation matrix, and the second transformation is similar to a third transformation.
[0094] Specifically, according to the method for camera-lidar calibration, a neural network can be trained by performing regression analysis based on the second transformation and the third transformation.
[0095] Figure 6 is a view showing a manner of training a neural network to perform camera-lidar calibration according to an embodiment of the present disclosure.
[0096] Reference Figure 6 , according to an embodiment of the present disclosure, in a method for camera-lidar calibration, a ground estimation algorithm can be applied to lidar point clouds (projected point clouds) that are matched with an image and projected onto a camera coordinate system, thereby generating a training ground point cloud P_ground. In addition to the ground estimation algorithm, a normal estimation algorithm can be applied to extract the training ground point cloud from the training lidar point cloud.
[0097] According to the method for camera-lidar calibration, a training road marking point cloud P_roadmark can be extracted from the training ground point cloud by performing intensity filtering (reflection filtering).
[0098] In addition, according to the method for camera-lidar calibration, a training transformed point cloud P_perturb can be generated by transforming the training road marking point cloud with a random second transformation T_arbi.
[0099] In addition, according to the method for camera-lidar calibration, a training image (RGB image) captured by a camera can be converted into a training grayscale image I_grayscale.
[0100] In addition, according to the method for camera-lidar calibration, a training ground edge image I_edge can be generated by extracting edges of the training grayscale image and cropping portions other than the ground as the RoI.
[0101] In addition, a neural network can be trained by receiving inputs of the training transformed point cloud and the training ground edge image such that a third transformation T_pred is predicted. Specifically, the neural network can be trained by performing regression analysis based on the second transformation and the third transformation.
[0102] Figure 7AIt is a view showing a method for camera-lidar calibration according to an embodiment of the present disclosure.
[0103] Referring to Figure 7A , it can be recognized that the training road marking point cloud and the training edge image are expressed. According to an embodiment of the present disclosure, in the method for camera-lidar calibration, although the training ground edge image matching the training road marking point cloud is used, the edge image of the vehicle and the edge image of the ground are expressed together so that Figure 7A .
[0104] Referring to Figure 7A , it can be recognized that the training edge image matches the training road marking point cloud, as shown by reference numeral 710.
[0105] Figure 7B It is a view showing a method for camera-lidar calibration according to an embodiment of the present disclosure.
[0106] Referring to Figure 7B , it can be recognized that the training transformation point cloud and the training edge image are expressed. According to an embodiment of the present disclosure, in the method for camera-lidar calibration, although the training ground edge image matching the training road marking point cloud is used, the edge image of the vehicle and the edge image of the ground are expressed so that Figure 7B .
[0107] Referring to Figure 7B , it can be recognized that the training edge image does not match the training road marking point cloud, as shown by reference numeral 720. Specifically, according to an embodiment of the present disclosure, in the method for camera-lidar calibration, data augmentation can be performed by changing the training road marking point cloud matching the training edge image through a second transformation. In other words, the training transformation point cloud is generated by transforming the training road marking point cloud using the second transformation. Through the transformation based on the second transformation, the training transformation point cloud may not match the training edge image.
[0108] Figure 8 It is a block diagram showing a device for camera-lidar calibration according to an embodiment of the present disclosure.
[0109] According to an embodiment of the present invention, the device 100 for camera-lidar calibration may include a camera 110, a lidar 120, an acquisition device 130, an extraction device 140, a generation device 150, a calibration device 160, and a training device 170.
[0110] The acquisition device 130 may be configured to acquire an image captured by the camera 110 at a specific time point and a lidar point cloud projected onto the camera coordinate system at the specific time point. However, when the camera and the lidar are not time-calibrated, even if the acquisition device 130 acquires an image at a specific time point and a lidar point cloud projected onto the camera coordinate system at the specific time point, if the lidar point cloud projected onto the camera coordinate system is projected onto the image captured by the camera, the lidar point cloud projected onto the camera coordinate system may not match the image captured by the camera.
[0111] In addition, the acquisition device 130 may be configured to acquire a training ground edge image and acquire a training road marking point cloud that matches the training ground edge image.
[0112] The extraction device 140 may extract a ground edge image corresponding to the edge of the ground from the image, and may extract a road marking point cloud of the point cloud indicating road markings on the ground from the lidar point cloud.
[0113] Specifically, the extraction device 140 may be configured to extract a ground point cloud corresponding to the ground from the lidar point cloud, and extract a road marking point cloud having an intensity greater than or equal to a specific intensity from the ground point cloud.
[0114] In addition, the extraction device 140 may be configured to extract a ground point cloud through a ground estimation algorithm.
[0115] In addition, the extraction device 140 may be configured to extract a ground point cloud through a normal estimation algorithm.
[0116] In addition, the extraction device 140 may be configured to extract an edge image through an edge filtering algorithm for a grayscale image, set the ground of the edge image as the RoI, and crop the remaining area of the edge image except for the RoI, thereby extracting a ground edge image.
[0117] In addition, the extraction device 140 may be configured to extract a ground edge image by extracting a part of the feature points of the image via a neural network.
[0118] The generation device 150 may be configured to generate a first change indicating a predicted position change by predicting, via a neural network, a position change of the road marking point cloud relative to the ground edge image. Specifically, the first change may be a matrix including elements of a rotation matrix and elements of a translation matrix.
[0119] In addition, the generation device 150 may be configured to generate a training transformed point cloud by transforming the training road marking point cloud using a second change, and predict a position change of the training transformed point cloud of the training ground edge image through a neural network, thereby generating a third change indicating the position change.
[0120] In addition, the generation device 150 may be configured to generate a grayscale image by transforming an image into a grayscale image.
[0121] The calibration device 160 may be configured to calibrate the camera 110 and the lidar 120 based on the first transformation.
[0122] In addition, the calibration device 160 may be configured to calibrate the camera and the lidar by multiplying the coordinates of the points of the road marking point cloud by the inverse matrix of the first transformation.
[0123] Specifically, the training device 170 may train the neural network by performing regression analysis based on the second transformation and the third transformation.
[0124] Figure 9 is a block diagram of a computing system for performing a method for camera-lidar calibration according to an embodiment of the present disclosure.
[0125] The acquisition device 130, the extraction device 140, the generation device 150, the calibration device 160, and the training device 170 may correspond to processors.
[0126] See Figure 9 , a method for camera-lidar calibration according to an embodiment of the present disclosure may be implemented by a computing system. The computing system 1000 may include at least one processor 1100, a memory 1300, a user interface input device 1400, a user interface output device 1500, a storage device 1600, and a network interface 1700 that are connected to each other via a bus 1200.
[0127] The processor 1100 may be a central processing unit (CPU) or a semiconductor device for processing instructions stored in the memory 1300 and / or the storage device 1600.
[0128] Each of the memory 1300 and the storage device 1600 may include various types of volatile or non-volatile storage media. For example, the memory 1300 may include a read-only memory (ROM) and a random access memory (RAM) (1320).
[0129] Therefore, the operations of the methods or algorithms described in connection with the embodiments disclosed in the present disclosure may be directly implemented by hardware modules, software modules, or a combination thereof executed by the processor 1100. The software modules may reside on a storage medium (i.e., the memory 1300 and / or the storage device 1600), such as RAM, flash memory, ROM, erasable and programmable ROM (EPROM), electrically EPROM (EEPROM), registers, a hard disk, a removable disk, or a CD-ROM.
[0130] As described above, according to the present disclosure, the time mismatch between the camera and the lidar can be compensated by using a neural network.
[0131] According to the present disclosure, a camera, a lidar, and a processor disposed inside a vehicle can be used to perform time synchronization between the camera and the lidar without additional hardware.
[0132] According to the present disclosure, a neural network for time calibration of the camera and the lidar can be designed and the neural network can be trained.
[0133] According to the present disclosure, the time mismatch between the camera and the lidar can be replaced with a spatial mismatch between the camera and the lidar, and the spatial mismatch can be solved, thereby solving the time mismatch between the camera and the lidar.
[0134] According to the present disclosure, the point cloud corresponding to the ground can be effectively processed, and a data set can be automatically formed to perform a quantitative evaluation of semantic segmentation for estimating the depth distance of road markings for road information such as lanes or road surfaces.
[0135] The effects obtained according to the present disclosure are not limited to the above effects, and those skilled in the art to which the present disclosure pertains will clearly understand any other technical effects not mentioned herein from the detailed description.
[0136] In the foregoing, although the present disclosure has been described with reference to exemplary embodiments and the drawings, the present disclosure is not limited thereto, but various modifications and changes can be made by those skilled in the art to which the present disclosure pertains without departing from the spirit and scope of the present disclosure claimed in the appended claims. Therefore, the exemplary embodiments of the present disclosure are provided to explain the spirit and scope of the present disclosure, but do not limit them, such that the spirit and scope of the present disclosure are not limited by the embodiments. The scope of the present disclosure should be interpreted based on the appended claims, and all technical concepts within the scope equivalent to the claims should be included within the scope of the present disclosure.
Claims
1. A method for camera-lidar calibration, the method comprising: Obtain an image captured by a camera at a specific time point and a lidar point cloud captured by a lidar and projected onto a camera coordinate system at the specific time point; extracting a ground edge image corresponding to an edge of the ground from the image; Extracting a road marking point cloud representing road markings on the ground from the laser radar point cloud; Predicting a position change of the road marking point cloud relative to the ground edge image by a neural network to generate a first change indicating a predicted position change; as well as The camera and the lidar are calibrated based on the first change.
2. The method according to claim 1, further comprising: Get training ground edge images; Acquire a training road marking point cloud that matches the training ground edge image; Transforming the training road marking point cloud by a second transformation to generate a training transformed point cloud; generating a third change indicating a predicted position change by predicting a position change of the training transformed point cloud relative to the training ground edge image via the neural network; and The neural network is trained by performing a regression analysis based on the second variation and the third variation.
3. The method according to claim 1, wherein: Extracting the road marking point cloud includes: Extracting a ground point cloud corresponding to the ground from the laser radar point cloud; and The road marking point cloud having an intensity greater than or equal to a specific intensity is extracted from the ground point cloud.
4. The method according to claim 3, wherein: Extracting the ground point cloud includes: The ground point cloud is extracted by a ground estimation algorithm.
5. The method according to claim 3, wherein: Extracting the ground point cloud includes: The ground point cloud is extracted by a normal estimation algorithm.
6. The method according to claim 1, wherein: Extracting the ground edge image includes: generating a grayscale image by converting the image into a grayscale image; extracting an edge image by applying an edge filtering algorithm to the grayscale image; and The ground edge image is extracted by setting the ground of the edge image as a region of interest (RoI) and cropping a remaining region of the edge image except the RoI.
7. The method according to claim 1, wherein: Extracting the ground edge image includes extracting the ground edge image by extracting a portion of feature points of the image through the neural network.
8. The method according to claim 1, wherein: The first change includes: elements of a rotation matrix and elements of a translation matrix, and Wherein, the calibration comprises: The coordinates of the road marking point cloud are multiplied by the inverse matrix of the first change to calibrate the camera and the lidar.
9. The method according to claim 1, wherein: The neural network is a convolutional neural network (CNN) or a multi-layer perceptron (MLP).
10. A device for camera-lidar calibration, the device comprising: Camera; LiDAR; processor; an acquisition device configured to acquire an image captured by the camera at a specific time point and a laser radar point cloud captured by the laser radar and projected onto a camera coordinate system at the specific time point; an extraction device configured to extract a ground edge image corresponding to an edge of the ground from the image, and to extract a road marking point cloud from the laser radar point cloud, wherein the road marking point cloud represents a point cloud of road markings on the ground; A generating device configured to generate a first change indicating a predicted position change by predicting a position change of the road marking point cloud relative to the ground edge image through a neural network; as well as A calibration device is configured to calibrate the camera and the lidar based on the first change.
11. The device according to claim 10, wherein: The acquisition device is used for: Acquire a training ground edge image, and acquire a training road marking point cloud matching the training ground edge image, Wherein, the generating device is configured as follows: transforming the training road marking point cloud by a second transformation to generate a training transformed point cloud, and generating a third change indicating a predicted position change by predicting a position change of the training transformed point cloud relative to the training ground edge image via the neural network, and Wherein, the device also includes: A training device is configured to train the neural network by performing regression analysis based on the second change and the third change.
12. The device according to claim 10, wherein: The extraction device is configured to extract a ground point cloud corresponding to the ground from the laser radar point cloud; and extract the road marking point cloud with an intensity greater than or equal to a specific intensity from the ground point cloud.
13. The device according to claim 12, wherein: The extraction device is configured as follows: The ground point cloud is extracted by a ground estimation algorithm.
14. The apparatus according to claim 12, wherein: The extraction device is configured to extract the ground point cloud by a normal estimation algorithm.
15. The apparatus according to claim 10, wherein: The generating means is configured to generate a grayscale image by converting the image into a grayscale image; and Wherein, the extraction device is configured as follows: extracting an edge image by applying an edge filtering algorithm to the grayscale image; and The ground edge image is extracted by setting the ground of the edge image as a region of interest (RoI) and cropping the remaining region of the edge image except the RoI.
16. The apparatus according to claim 10, wherein: The extracting device is configured to extract the ground edge image by extracting a portion of feature points of the image via the neural network.
17. The apparatus according to claim 10, wherein: The first change includes elements of a rotation matrix and elements of a translation matrix; and Wherein, the calibration device is configured to calibrate the camera and the lidar by multiplying the coordinates of the road marking point cloud by the inverse matrix of the first change.
18. The apparatus according to claim 10, wherein: The neural network is a convolutional neural network (CNN) or a multi-layer perceptron (MLP).
19. A device for camera-lidar calibration, the device comprising: Camera; LiDAR; A processing system comprising a processor and a non-transitory computer-readable storage medium storing a program to be executed by the processor, the program comprising instructions for: Get the image captured by the camera at a specific point in time; Acquire a laser radar point cloud captured by the laser radar and projected onto a camera coordinate system at the specific time point; extracting a ground edge image corresponding to an edge of the ground from the image; extracting a road marking point cloud representing road markings on a road surface from the lidar point cloud; Predicting a position change of the road marking point cloud relative to the ground edge image by a neural network to generate a first change indicating a predicted position change; as well as The camera and the lidar are calibrated based on the first change.
20. The apparatus of claim 19, wherein: The program further includes instructions for: Get training ground edge images; Acquire a training road marking point cloud that matches the training ground edge image; Transforming the training road marking point cloud by a second transformation to generate a training transformed point cloud; generating a third change indicating a predicted position change by predicting a position change of the training transformed point cloud relative to the training ground edge image via the neural network; and The neural network is trained by performing a regression analysis based on the second variation and the third variation.
Citation Information
Patent Citations
Bicyclic heteroaromatic inhibitors of KLK5
KR1020230171455A