Point Cloud Continuous Frame Annotation Method and System Based on Multi-Position Image Mapping Assistance

By obtaining the camera's internal and external parameters, calculating the transformation matrix and mapping point cloud points to the image coordinate system, the recognition and labeling problems of distant objects or unclear outlines in point cloud labels are solved, and the labeling accuracy and machine learning recognition capabilities are improved.

CN116596955BActive Publication Date: 2025-07-08COWA TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310200426.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-07-08
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

During the point cloud labeling process, distant or unclear objects are difficult to accurately identify and label through multi-angle pictures, resulting in insufficient labeling accuracy.

Method used

By acquiring and calibrating the camera's internal and external parameters, continuously collecting point cloud and image data, calculating the rough and accurate transformation matrix, mapping point cloud points into the image coordinate system, and rendering and annotating.

Benefits of technology

Improve the labeling accuracy of objects at distant or unclear contours, enhance the ability of machine learning to identify sparse, long-distance and obscured objects, and improve the accuracy and consistency of labeling results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116596955B_ABST
    Figure CN116596955B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for continuous frame annotation of point clouds assisted by multi-position image mapping, including: obtaining and calibrating the internal and external parameters of a camera; while continuously collecting multiple frames of point cloud data, collecting corresponding matching image data at the same position, and thus obtaining a rough transformation matrix between any two frames of point clouds; performing refined registration through point cloud algorithm optimization, and then representing a point A in the nth frame of point cloud as a coordinate A' in the corresponding camera coordinate system of the (n + i)th frame of point cloud; mapping the coordinate A' to the pixel coordinate system of the (n + i)th frame of point cloud to obtain the pixel coordinates of A' in multiple cameras, and if within the size range of the current camera imaging picture, rendering the specified pixel points, and thus annotating the contour or surface of a specific object in the nth frame of point cloud in the corresponding multiple images of the (n + i)th frame of point cloud; otherwise, no rendering is required. The present invention can accurately identify and annotate objects that are far away or have unclear contours.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of point cloud annotation and segmentation annotation. Specifically, it relates to a method and system for continuous frame annotation of point clouds assisted by multi-position image mapping. Background Art

[0002] Point cloud annotation is widely used in fields such as artificial intelligence and autonomous driving, and the annotation results are used to train machines to recognize certain specific objects. Point cloud annotation mainly includes 3D detection box annotation, semantic segmentation annotation, instance segmentation annotation, etc. To improve the annotation efficiency and accuracy, multi-frame point cloud data is often continuously collected at different positions in the same scene, and multi-angle images associated with each frame of point cloud are synchronously collected to assist in judgment by users, and annotation is carried out in the form of continuous frames.

[0003] Currently, during the process of point cloud annotation, for a frame of point cloud data, multi-angle images taken at the same position as the point cloud collection are often used for reference to more intuitively observe an object in the current single frame of point cloud, so as to perform more accurate point cloud annotation. However, while using a lidar to collect point cloud data at a certain position, multiple cameras are used to collect multiple images in each direction at the current position, and the images are numbered to determine the orientation of the images relative to the collection device. During the process of annotating the point cloud, when the annotator cannot directly judge the object type through the point cloud contour, a 3D cube box can be used to select a spatial area, or a spatial polygon can be used to select a specific point set in the point cloud, and then the contour of the 3D cube or the selected point set is mapped to the 2D image collected simultaneously with the point cloud. The user determines the type of the annotation object by observing the image in this direction. For target objects that actually exist in the point cloud but cannot be clearly observed in the image due to problems such as too far distance, being blocked, and insufficient resolution, the annotator will not be able to accurately classify them. For example, because the street lamps are far from the position where the point cloud is scanned and the street lamps are relatively dense, it is impossible to accurately distinguish which specific street lamp in the point cloud mapped to the image from the images taken at the same time. Similarly, when other smaller or blocked objects appear in the distance, they cannot be accurately annotated.

[0004] In summary, there is a need in the market for a method and system for continuous frame annotation of point clouds assisted by multi-position image mapping that can accurately identify and annotate distant or unclear objects during the 3D point cloud continuous frame annotation process. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the purpose of the present invention is to provide a method and system for continuous frame annotation of point clouds assisted by multi-position image mapping.

[0006] According to a method for continuous frame annotation of point clouds assisted by multi-position image mapping provided by the present invention, it includes:

[0007] Step S1: Obtain and calibrate the internal and external parameters of the camera;

[0008] Step S2: Continuously collect multiple frames of point cloud data. While collecting each frame of point cloud data, collect corresponding and matching image data at the same position;

[0009] Step S3: Obtain a rough transformation matrix between any nth frame of point cloud and the (n + i)th frame of point cloud according to the image data;

[0010] Step S4: Refine the registration through point cloud algorithm optimization to obtain an accurate transformation matrix;

[0011] Step S5: Represent the point A in the nth frame of point cloud as coordinate A' in the corresponding same camera coordinate system of the (n + i)th frame of point cloud through the accurate transformation matrix;

[0012] Step S6: Map the coordinate A' to the pixel coordinate system of the (n + i)th frame of point cloud to obtain the pixel coordinates of A' in multiple cameras, and determine whether the pixel coordinates are within the size range of the current camera's imaging picture. If so, trigger Step S7; if not, it means that A' does not exist in the picture and there is no need to render;

[0013] Step S7: Indicate that A' exists in the picture, and render the specified pixel points, and then mark the contour or surface of the specific object in the nth frame of point cloud in the corresponding multiple images of the (n + i)th frame of point cloud.

[0014] Preferably, the obtaining of the internal and external parameters of the camera includes respectively performing mapping calibration between the world coordinate and the pixel coordinate for multiple cameras carried by the acquisition device to obtain the internal parameters of each camera, and at the same time calibrating the external parameters between multiple cameras carried on the same acquisition device.

[0015] Preferably, Step S2 further includes: While collecting point cloud data, use an odometer or GPS and IMU to record the pose parameters of the acquisition device.

[0016] Preferably, Step S3 includes:

[0017] Step S3.1: Extract multiple points in the (n + i)th frame of point cloud, and the distance between the multiple points is greater than a preset minimum distance threshold d;

[0018] Step S3.2: Search for points with the same features as the multiple points in the nth frame of point cloud, and select one point from the points as the corresponding point of the (n + i)th frame in the nth frame;

[0019] Step S3.3: Calculate the rough transformation matrix of the corresponding points.

[0020] Preferably, the point cloud optimization algorithm in step S4 includes using the ICP algorithm or the SVD algorithm.

[0021] According to a point cloud continuous frame annotation system assisted by multi-position picture mapping provided by the present invention, it includes:

[0022] Module M1: Obtain and calibrate the internal and external parameters of the camera;

[0023] Module M2: Continuously collect multiple frames of point cloud data. While collecting each frame of point cloud data, collect corresponding matching image data at the same position;

[0024] Module M3: Obtain a rough transformation matrix between any nth frame of point cloud and the (n + i)th frame of point cloud according to the image data;

[0025] Module M4: Perform refined registration through point cloud algorithm optimization to obtain an accurate transformation matrix;

[0026] Module M5: Represent the point A in the nth frame of point cloud as the coordinate A' in the corresponding same camera coordinate system of the (n + i)th frame of point cloud through the accurate transformation matrix;

[0027] Module M6: Map the coordinate A' to the pixel coordinate system of the (n + i)th frame of point cloud to obtain the pixel coordinates of A' in multiple cameras, and determine whether the pixel coordinates are within the size range of the current camera imaging picture. If so, trigger Module M7; if not, it means that A' does not exist in the picture and there is no need to render;

[0028] Module M7: Indicate that A' exists in the picture and render the specified pixel point, and then label the contour or surface of a specific object in the nth frame of point cloud in the corresponding multiple images of the (n + i)th frame of point cloud.

[0029] Preferably, obtaining the internal and external parameters of the camera includes performing mapping calibration of world coordinates and pixel coordinates for multiple cameras carried by the acquisition device respectively to obtain the internal parameters of each camera, and at the same time calibrating the external parameters between multiple cameras carried on the same acquisition device.

[0030] Preferably, Module M2 further includes: while collecting point cloud data, using an odometer or GPS and IMU to record the pose parameters of the acquisition device.

[0031] Preferably, Module M3 includes:

[0032] Module M3.1: Extract multiple points in the (n + i)th frame of point cloud, and the distance between the multiple points is greater than a preset minimum distance threshold d;

[0033] Module M3.2: Find points with the same features among the multiple points in the n-th frame of the point cloud, and select one point from these points as the corresponding point of the (n + i)-th frame in the n-th frame;

[0034] Module M3.3: Calculate the rough transformation matrix of the corresponding point.

[0035] Preferably, the point cloud optimization algorithm in the module M4 includes using the ICP algorithm or the SVD algorithm.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] 1. For an object scanned by the lidar in a certain frame but not captured by the auxiliary camera, when the annotator makes annotations, they can use the images captured by the auxiliary camera in other frames, combined with the mapping from the 3D space points in the current frame to the 2D images in other frames, to annotate the data more precisely.

[0038] 2. The richer and more accurate annotation results of the present invention can enable the machine to improve its ability to recognize objects in sparse, distant, occluded and other scenarios during the learning and training process, and thus perform more stably, accurately and efficiently in applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects and advantages of the present invention will become more apparent:

[0040] Figure 1 It is a schematic diagram of the working process of the present invention;

[0041] Figure 2 It is a diagram showing the effect of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0043] Embodiment 1

[0044] According to a method for continuous frame annotation of point clouds assisted by multi-position picture mapping provided by the present invention, as Figure 1 and Figure 2 shown, it includes:

[0045] Step S1: Obtain and calibrate the internal and external parameters of the camera. The obtaining of the internal and external parameters of the camera includes performing mapping calibration of world coordinates and pixel coordinates respectively for multiple cameras mounted on the acquisition device to obtain the internal parameters of each camera, and at the same time calibrating the external parameters between multiple cameras mounted on the same acquisition device.

[0046] Step S2: Continuously acquire multiple frames of point cloud data. While acquiring each frame of point cloud data, acquire corresponding and matching image data at the same position. This step also includes, while acquiring point cloud data, using an odometer or GPS and IMU to record the pose parameters of the acquisition device. Specifically, continuously acquire M frames of point cloud data in the same scene. While each frame of point cloud data is generated, match P pictures that can cover all angles of the acquisition device at the same position. Continuously acquire M frames of point cloud data, as well as the P pictures corresponding to each frame of point cloud and the vehicle pose.

[0047] Step S3: Obtain a rough transformation matrix between any nth frame of point cloud and the (n + i)th frame of point cloud according to the image data. Specifically, first extract multiple points in the (n + i)th frame of point cloud, and the distance between the multiple points is greater than a preset minimum distance threshold d. Then find points with the same features as the multiple points in the nth frame of point cloud, and select one point from these points as the corresponding point of the (n + i)th frame in the nth frame. Finally, calculate the rough transformation matrix of the corresponding points.

[0048] Step S4: Refine the registration through point cloud algorithm optimization to obtain an accurate transformation matrix. The point cloud optimization algorithm in Step S4 includes using the ICP algorithm, Iterative Closest Point, iterative closest point or SVD algorithm. Specifically, taking the ICP algorithm as an example, assume that the point cloud N is Q, and N + i is P. Perform registration on any frame of the Q and P point clouds. Arbitrarily select a point in the point cloud Q, denoted as Q i , find a point with the shortest Euclidean distance from it in the point cloud P, denoted as P i ; the method for finding the point with the shortest Euclidean distance from it includes using a KD-Tree for searching.

[0049] At this time, Q i and P i are corresponding points, obtain the transformation matrix, and after multiple iterations, finally obtain the most ideal transformation matrix to make the two point clouds coincide. The formula is as follows:

[0050]

[0051] Wherein, R represents the rotation transformation matrix, t represents the translation vector, and k represents the number of iterations. Since the lidar and the camera are mounted on the same acquisition device and the relative position relationship remains consistent, the external parameters and the homogeneous transformation matrix T corresponding to the camera in the spatial transformation from the nth frame to the (n + i)th frame can be obtained based on t and R.

[0052] Step S5: Represent the point A in the nth frame of the point cloud as the coordinate A' in the corresponding camera coordinate system of the (n + i)th frame of the point cloud through the precise transformation matrix. Specifically, a world coordinate system is established based on the optical center of a certain camera in the nth frame. Assuming that there is an arbitrary homogeneous vector a with vertex A in the space of the nth frame of the point cloud, the camera coordinate a' of point A in the (n + i)th frame based on the same camera can be obtained through the external parameter T, i.e., a' = Ta.

[0053] Step S6: Map the coordinate A' to the pixel coordinate system of the (n + i)th frame of the point cloud to obtain the pixel coordinates of A' in multiple cameras, and determine whether the pixel coordinates are within the size range of the current camera's imaging picture. If so, trigger Step S7; if not, it means that A' does not exist in the picture and no rendering is required. After obtaining a', the imaging plane coordinates and pixel coordinates of point A can be obtained through the camera internal parameters, so as to map the spatial point A in the nth frame of the point cloud to the 2D image of the (n + i)th frame, realizing the mapping of spatial points across frames and solving the problem of being unable to distinguish labeled objects.

[0054] Step S7: Indicate that A' exists in the picture, and render the specified pixel points, and then label the contour or surface of the specific object in the nth frame of the point cloud in the corresponding multiple images of the (n + i)th frame of the point cloud.

[0055] The present invention also provides a point cloud continuous frame annotation system assisted by multi-position picture mapping. Those skilled in the art can implement the point cloud continuous frame annotation system assisted by multi-position picture mapping by executing the step flow of the point cloud continuous frame annotation method assisted by multi-position picture mapping. That is, the point cloud continuous frame annotation method assisted by multi-position picture mapping can be understood as the preferred implementation manner of the point cloud continuous frame annotation system assisted by multi-position picture mapping.

[0056] Embodiment 2

[0057] A point cloud continuous frame annotation system assisted by multi-position picture mapping according to the present invention includes:

[0058] Module M1: Obtain and calibrate the internal and external parameters of the camera. The obtaining of the internal and external parameters of the camera includes respectively performing mapping calibration between the world coordinate and the pixel coordinate for multiple cameras mounted on the acquisition device to obtain the internal parameters of each camera, and at the same time calibrating the external parameters between multiple cameras mounted on the same acquisition device.

[0059] Module M2: Continuously collect multiple frames of point cloud data. While collecting each frame of point cloud data, corresponding and matching image data is collected at the same location. The module M2 further includes: While collecting point cloud data, the pose parameters of the collection device are recorded using an odometer or GPS and IMU.

[0060] Module M3: Obtain a rough transformation matrix between any n-th frame of point cloud and the (n + i)-th frame of point cloud according to the image data. The module M3 includes: Module M3.1: Extract multiple points in the (n + i)-th frame of point cloud, and the distance between the multiple points is greater than a preset minimum distance threshold d. Module M3.2: Search for points with the same features as the multiple points in the n-th frame of point cloud, and select one point from the points as the corresponding point of the (n + i)-th frame in the n-th frame. Module M3.3: Calculate the rough transformation matrix of the corresponding points.

[0061] Module M4: Refine the registration through point cloud algorithm optimization to obtain an accurate transformation matrix. The point cloud optimization algorithm in the module M4 includes using the ICP algorithm or the SVD algorithm.

[0062] Module M5: Through the accurate transformation matrix, represent the point A in the n-th frame of point cloud in the corresponding camera coordinate system of the (n + i)-th frame of point cloud as coordinate A'.

[0063] Module M6: Map the coordinate A' to the pixel coordinate system of the (n + i)-th frame of point cloud to obtain the pixel coordinates of A' in multiple cameras, and determine whether the pixel coordinates are within the size range of the current camera imaging picture. If so, trigger module M7; if not, it means that A' does not exist in the picture and there is no need to render.

[0064] Module M7: Indicate that A' exists in the picture, and render the specified pixel points, and then mark the contour or surface of a specific object in the n-th frame of point cloud in the corresponding multiple images of the (n + i)-th frame of point cloud.

[0065] Those skilled in the art know that in addition to implementing the system, device and its various modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to make the system, device and its various modules provided by the present invention be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc. to implement the same program. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as either a software program for implementing the method or the structure within the hardware component.

[0066] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be arbitrarily combined with each other.

Claims

1. A method for continuous frame annotation of point clouds assisted by multi-position image mapping, characterized in that, Including: Step S1: Obtain and calibrate the internal and external parameters of the camera; Step S2: Continuously collect multiple frames of point cloud data. While collecting each frame of point cloud data, collect corresponding and matching image data at the same location; Step S3: Obtain a rough transformation matrix between any nth frame of point cloud and the (n + i)th frame of point cloud according to the image data; Step S4: Refine the registration through point cloud algorithm optimization to obtain an accurate transformation matrix; Step S5: Represent the point A in the nth frame of point cloud as coordinate A' in the corresponding same camera coordinate system of the (n + i)th frame of point cloud through the accurate transformation matrix; Step S6: Map the coordinate A' to the pixel coordinate system of the (n + i)th frame of point cloud to obtain the pixel coordinates of A' in multiple cameras, and determine whether the pixel coordinates are within the size range of the current camera's imaging picture. If so, trigger Step S7; if not, it means that A' does not exist in the picture and no rendering is required; Step S7: Indicate that A' exists in the picture, and render the specified pixel points, thereby annotating the contour or surface of a specific object in the nth frame of point cloud in the corresponding multiple images of the (n + i)th frame of point cloud.

2. The method for continuous frame annotation of point cloud based on multi-position picture mapping assistance according to claim 1, wherein The obtaining and calibrating of the internal and external parameters of the camera includes performing mapping calibration between the world coordinates and pixel coordinates for multiple cameras carried by the acquisition device to obtain the internal parameters of each camera, and simultaneously calibrating the external parameters between multiple cameras carried on the same acquisition device.

3. The method for continuous frame annotation of point cloud based on multi-position picture mapping assistance according to claim 1, wherein Step S2 further includes: While collecting point cloud data, use an odometer or GPS and IMU to record the pose parameters of the acquisition device.

4. The method for labeling consecutive frames of point cloud based on multi-position picture mapping assistance according to claim 1, wherein Step S3 includes: Step S3.1: Extract multiple points in the (n + i)th frame of point cloud, and the distance between the multiple points is greater than a preset minimum distance threshold d; Step S3.2: Search for points with the same features as the multiple points in the nth frame of point cloud, and select one point from the points as the corresponding point of the (n + i)th frame in the nth frame; Step S3.3: Calculate the rough transformation matrix of the corresponding points.

5. The method for continuously labeling point cloud frames based on multi-position picture mapping assistance according to claim 1, wherein The point cloud optimization algorithm in Step S4 includes using the ICP algorithm or the SVD algorithm.

6. A point cloud continuous frame annotation system based on multi-position picture mapping assistance, characterized in that, Including: Module M1: Obtain and calibrate the internal and external parameters of the camera; Module M2: Continuously collect multiple frames of point cloud data. While collecting each frame of point cloud data, collect corresponding and matching image data at the same location; Module M3: Obtain a rough transformation matrix between any nth frame of point cloud and the (n + i)th frame of point cloud according to the image data; Module M4: Refine the registration through point cloud algorithm optimization to obtain an accurate transformation matrix; Module M5: Represent the point A in the nth frame of point cloud as coordinate A' in the corresponding same camera coordinate system of the (n + i)th frame of point cloud through the accurate transformation matrix; Module M6: Map the coordinate A' to the pixel coordinate system of the (n + i)th frame of point cloud to obtain the pixel coordinates of A' in multiple cameras, and determine whether the pixel coordinates are within the size range of the current camera's imaging picture. If so, trigger Module M7; if not, it means that A' does not exist in the picture and no rendering is required; Module M7: Indicates the existence of A' in the said picture, and renders specified pixel points, thereby marking the contour or surface of a specific object in the n-th frame of point cloud in the corresponding multiple images of the (n + i)-th frame of point cloud.

7. The point cloud continuous frame annotation system based on multi-position picture mapping assistance according to claim 6, characterized in that The acquisition and calibration of the internal and external parameters of the camera include mapping and calibrating the world coordinates and pixel coordinates respectively by multiple cameras mounted on the acquisition device to obtain the internal parameters of each camera, and simultaneously calibrating the external parameters between multiple cameras mounted on the same acquisition device.

8. The point cloud continuous frame annotation system based on multi-position picture mapping assistance according to claim 6, characterized in that, The said Module M2 further includes: while collecting point cloud data, using an odometer or GPS and IMU to record the pose parameters of the acquisition device.

9. The point cloud continuous frame annotation system based on multi-position picture mapping assistance according to claim 6, characterized in that, The said Module M3 includes: Module M3.1: Extract multiple points in the (n + i)-th frame of point cloud, where the distance between the multiple points is greater than a preset minimum distance threshold d; Module M3.2: Search for points with the same features as the multiple points in the n-th frame of point cloud, and select one point from these points as the corresponding point of the (n + i)-th frame in the n-th frame; Module M3.3: Calculate the rough transformation matrix of the corresponding point.

10. The point cloud continuous frame annotation system based on multi-position picture mapping assistance according to claim 6, wherein The point cloud optimization algorithm in the said Module M4 includes using the ICP algorithm or the SVD algorithm.

Citation Information

Patent Citations

  • Image labeling method and device based on point cloud, equipment and storage medium

    CN115222588A

  • Image tracking method and device, equipment and storage medium

    CN115439627A