Point cloud coloring processing method, device, equipment and storage medium
By synchronizing the LiDAR and camera clocks in the SLAM scanning device, determining the timestamp of the pixel row based on the exposure and readout duration, and using the camera pose data for point cloud projection and coloring, the problem of mismatched point cloud colors is solved, coloring accuracy is improved and costs are reduced.
Patent Information
- Application Number
- CN202510192701.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-02-20
AI Technical Summary
In SLAM scanning equipment, the rolling shutter effect caused by the camera capturing images during dynamic movement leads to a problem where the point cloud color does not correspond to the actual color. Existing technologies using global shutter cameras are costly and have low resolution.
By equipping a laser scanning device with a lidar and a camera, clock synchronization is achieved. The timestamp of each pixel row is determined based on the exposure time and pixel readout time of the target image. Point cloud projection processing is performed using camera pose data and intrinsic parameters to determine the target pixel of each point and perform color processing.
It improves the accuracy of point cloud shading, reduces color differences caused by the rolling shutter effect, and eliminates the need for a high-cost global shutter camera.
Smart Images

Figure CN120125790B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to three-dimensional reconstruction technology and point cloud processing technology, and in particular to a point cloud coloring processing method, apparatus, device and storage medium. Background Technology
[0002] Simultaneous Localization and Mapping (SLAM) scanning equipment typically includes a lidar unit and a camera. The lidar unit performs laser scanning of the surrounding environment to acquire 3D spatial information and generate corresponding point clouds, while the camera acquires images of the surrounding environment to colorize the point cloud generated by the lidar unit using the color values of each point in the image.
[0003] During 3D reconstruction, the SLAM scanning equipment continuously moves during data acquisition, resulting in the camera also acquiring images while in dynamic motion. This leads to the "rolling shutter effect." Specifically, because each row of pixels in the image is exposed individually, each row has a readout time delay, and the exposure time for the entire image can be tens of milliseconds. During this time, the camera may move a few centimeters or rotate a few degrees, thus changing the camera pose data. Therefore, projecting and coloring the point cloud using a method that maps one image to one camera pose data will result in a mismatch between the point cloud colors and the actual colors. Related technologies typically use global shutter cameras for image acquisition to address this issue; however, global shutter cameras are expensive and have low resolution and poor color reproduction. Summary of the Invention
[0004] This disclosure provides a point cloud coloring processing method, apparatus, device, and storage medium that can improve the accuracy of point cloud coloring.
[0005] One aspect of this disclosure provides a point cloud coloring processing method applied to a laser scanning device, the laser scanning device being equipped with a lidar and a camera, the lidar being used to generate point clouds for a target scene, the camera being used to acquire images of the target scene, and the lidar and the camera having synchronized clocks; the method includes:
[0006] Based on the exposure time and pixel readout time of the target image, the timestamp of each pixel row in the target image is determined. The exposure time is the time required for the camera to generate a row of pixels, and the pixel readout time is the time required for the camera's processor to read the pixel information of the target image row by row.
[0007] The camera pose data corresponding to each pixel row is determined based on the timestamp of each pixel row. The camera pose data is used to characterize the position and orientation of the camera in the world coordinate system when generating the corresponding row of pixels.
[0008] Based on the camera pose data and the camera intrinsic parameters of the camera, projection processing is performed on each point in the point cloud to determine the target pixel corresponding to each point; wherein, the target pixel is the projection point of the point under the camera pose data corresponding to the target pixel row, and the target pixel row is the pixel row where the target pixel is located.
[0009] The point cloud is colored using the target pixel corresponding to each point.
[0010] Optionally, the step of projecting each point in the point cloud based on the camera pose data and the camera's intrinsic parameters to determine the target pixel corresponding to each point includes:
[0011] Using a preset pixel row in the target image as the initial projection row, and based on the camera pose data corresponding to the initial projection row and the camera intrinsic parameters, calculate the projection coordinates of the point in the image plane coordinate system where the initial projection row is located;
[0012] In response to the fact that the number of rows corresponding to the projected coordinates is the same as the number of rows of the initial projected rows or the difference in the number of rows is less than the difference threshold, the pixel at the projected coordinates is determined as the target pixel corresponding to the point;
[0013] In response to the difference between the number of rows corresponding to the projected coordinates and the number of rows of the initial projected row being greater than or equal to the difference threshold, the pixel row where the previously calculated projected coordinates are located is used as the iterative projected row in the next iteration calculation. Based on the camera pose data corresponding to the iterative projected row and the camera intrinsic parameters, the projected coordinates of the point in the image plane coordinate system where the iterative projected row is located are calculated until the difference between the number of rows corresponding to the calculated projected coordinates and the number of rows of the iterative projected row is less than the difference threshold.
[0014] Optionally, calculating the projection coordinates of the point in the image plane coordinate system where the iterative projection row is located, based on the camera pose data corresponding to the iterative projection row and the camera intrinsic parameters, includes:
[0015] Based on the camera pose data and the first coordinates of the point in the world coordinate system, a first coordinate transformation is performed on the point to obtain the second coordinates of the point in the camera coordinate system corresponding to the camera pose data; based on the second coordinates and the camera intrinsic parameters, a second coordinate transformation is performed on the point to obtain the projected coordinates, wherein the camera intrinsic parameters are used to characterize the mapping relationship between the camera coordinate system and the image plane coordinate system.
[0016] Optionally, determining the camera pose data corresponding to each pixel row based on the timestamp of each pixel row includes:
[0017] For each pixel row in the target image, compare the timestamp of the pixel row with the acquisition time of each camera pose data in the camera pose list;
[0018] In response to the existence of a target acquisition time with the same timestamp, the camera pose data of the target acquisition time is determined as the camera pose data corresponding to the pixel row;
[0019] In response to the absence of the target acquisition time, interpolation calculation is performed based on the camera pose data corresponding to at least two sets of acquisition times adjacent to the timestamp to obtain the camera pose data corresponding to the pixel row.
[0020] Optionally, the laser scanning device is also equipped with an inertial measurement unit;
[0021] The method of comparing the timestamp of each pixel row in the target image with the acquisition time of each camera pose data in the camera pose list before the acquisition time of each pixel row is described.
[0022] Obtain the first pose list acquired by the inertial measurement unit and the second pose list acquired by the lidar, wherein the first pose data in the first pose list is used to characterize the position and attitude of the inertial measurement unit in the world coordinate system, and the second pose data in the second pose list is used to characterize the position and attitude of the lidar in the world coordinate system.
[0023] Based on the first pose list and the first extrinsic parameter of the camera relative to the inertial measurement unit, a first camera pose list is obtained, and based on the second pose list and the second extrinsic parameter of the camera relative to the lidar, a second camera pose list is obtained. The first extrinsic parameter is used to characterize the relative position and relative attitude of the camera relative to the inertial measurement unit in the world coordinate system, and the second extrinsic parameter is used to characterize the relative position and relative attitude of the camera relative to the lidar in the world coordinate system.
[0024] Using the second camera pose list as a priori constraint and the pre-integration interpolation result of the first camera pose list as a relative factor constraint, a global joint optimization is performed on the first camera pose list and the second camera pose list to obtain the camera pose list.
[0025] Optionally, the laser scanning device has an independent clock;
[0026] Before acquiring the first pose list collected by the inertial measurement unit and the second pose list collected by the lidar, the method further includes:
[0027] Based on a preset clock synchronization method, the independent clocks are used to synchronize the internal clocks of the lidar, the camera, and the inertial measurement unit, respectively, to obtain the acquisition time of each first pose data in the first pose list, the acquisition time of each second pose data in the second pose list, and the acquisition time of the target image.
[0028] Optionally, determining the timestamp of each pixel row in the target image based on the exposure time and pixel readout time of the target image includes:
[0029] The timestamp of each pixel row is determined based on the acquisition time of the target image, the exposure duration, the pixel readout duration, and the total number of rows of the target image;
[0030] Wherein, the total number of rows is the number of pixel rows in the target image, the timestamp of the j-th pixel row is the sum of the acquisition time of the target image, half of the exposure time, and the readout time of the j-th pixel row, the readout time of the j-th pixel row is the product of j and the average readout time, and the average readout time is the quotient of the pixel readout time and the total number of rows, where j is a positive integer.
[0031] Another aspect of this disclosure provides a point cloud coloring processing apparatus applied to a laser scanning device, the laser scanning device being equipped with a lidar and a camera, the lidar being used to generate point clouds for a target scene, the camera being used to acquire images of the target scene, and the lidar and the camera having synchronized clocks; the apparatus includes:
[0032] The first determining module is used to determine the timestamp of each pixel row in the target image based on the exposure time and pixel readout time of the target image, wherein the exposure time is the time required for the camera to generate a row of pixels, and the pixel readout time is the time required for the camera's processor to read the pixel information of the target image row by row;
[0033] The second determining module is used to determine the camera pose data corresponding to each pixel row based on the timestamp of each pixel row. The camera pose data is used to characterize the position and orientation of the camera in the world coordinate system when generating the corresponding row of pixels.
[0034] The third determining module is used to perform projection processing on each point in the point cloud based on the camera pose data and the camera intrinsic parameters of the camera to determine the target pixel point corresponding to each point; wherein, the target pixel point is the projection point of the point under the camera pose data corresponding to the target pixel row, and the target pixel row is the pixel row where the target pixel point is located.
[0035] A coloring processing module is used to color the point cloud using the target pixel corresponding to each point.
[0036] In another aspect of this disclosure, an electronic device is provided, comprising:
[0037] Memory, used to store computer programs;
[0038] A processor is configured to execute a computer program stored in the memory, wherein, when the computer program is executed, it implements the methods described above.
[0039] In another aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the methods described above.
[0040] In another aspect of this disclosure, a computer program is provided, including computer program instructions that, when executed by a processor, implement the method described above.
[0041] Based on the embodiments of this disclosure, when the LiDAR and the camera are synchronized, the timestamp of each pixel row is calculated according to the exposure time and pixel readout time of the target image, thereby determining the camera pose data when the camera collects each row of pixels, and then determining the true projection position of each point in the point cloud in the target image. That is, when the projection point under the camera pose data corresponding to a certain pixel row falls exactly on that pixel row, it can be determined that the projection point at this time is the target pixel of that point. Using the target pixel corresponding to each point to color the point cloud can reduce the difference between the point cloud color and the actual color caused by the rolling shutter effect, improve the accuracy of point cloud coloring, and eliminate the need to use a global shutter camera, resulting in lower cost.
[0042] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0043] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0044] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description, wherein:
[0045] Figure 1 This is a flowchart of one embodiment of the point cloud coloring processing method disclosed herein;
[0046] Figure 2 This is a flowchart of another embodiment of the point cloud coloring processing method of this disclosure;
[0047] Figure 3 This is a flowchart of another embodiment of the point cloud coloring processing method of this disclosure;
[0048] Figure 4 This is a schematic diagram of the structure of an embodiment of the point cloud coloring processing apparatus of this disclosure;
[0049] Figure 5 This is a schematic diagram of another embodiment of the point cloud coloring processing apparatus of this disclosure;
[0050] Figure 6 This is a schematic diagram of the structure of an application embodiment of the electronic device disclosed herein. Detailed Implementation
[0051] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.
[0052] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0053] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0054] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0055] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0056] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0057] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0058] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0059] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0060] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0061] Figure 1 This is a flowchart illustrating a point cloud coloring processing method provided in an exemplary embodiment of this disclosure. The point cloud coloring processing method of this embodiment can be implemented using a laser scanning device equipped with a LiDAR and a camera. The laser scanning device may include, but is not limited to, SLAM devices, Visual SLAM (VSLAM), and 3D stereo scanners. This disclosure uses a SLAM device as an example. The camera is used to acquire images of the target scene, and the LiDAR is used to acquire images of the target scene. The camera and LiDAR are clock-synchronized. A point cloud is a dataset of points in three-dimensional space, capable of representing the shape of objects in that three-dimensional space. The LiDAR emits a large number of laser signals into the target scene. The laser signals are reflected by the surfaces of objects in the target scene, generating reflection signals which are received by the LiDAR. Based on the reflection signals, the LiDAR can calculate the position information of each point on the object surface in the world coordinate system corresponding to the target scene, thereby generating a point cloud corresponding to the target scene. The point cloud includes the position information of each point in the target scene. Furthermore, the position information, color values, laser reflection intensity, etc., of each point can be associated and stored in the point cloud file corresponding to the point cloud to improve the point cloud information.
[0062] like Figure 1 As shown, the method includes the following steps:
[0063] Step 101: Determine the timestamp of each pixel row in the target image based on the exposure time and pixel readout time of the target image.
[0064] The exposure time is the time required for the camera to generate a line of pixels. This exposure time is determined by the camera's shutter speed and can be obtained directly from the shutter speed set when the target image is acquired. The pixel readout time is the time required for the camera's processor to read the pixel information of the target image line by line, that is, the time required from the start of reading the first line of pixels until the last line of pixels is read. This pixel readout time is a fixed preset time and can be obtained directly from the camera information.
[0065] Since the exposure and readout duration of a row of pixels is short, typically a few milliseconds, the embodiments of this disclosure can regard a row of pixels as starting and ending the exposure at the same time. That is, a row of pixels corresponds to a camera pose, thereby determining the camera pose data corresponding to each row of pixels by obtaining the timestamp of each row of pixels.
[0066] Step 102: Determine the camera pose data corresponding to each pixel row based on the timestamp of each pixel row.
[0067] Among them, camera pose data is used to characterize the position and orientation of the camera in the world coordinate system when generating the corresponding row of pixels.
[0068] In one possible implementation, pose data acquisition devices such as LiDAR and Inertial Measurement Unit (IMU) can be used to acquire the camera's own pose data. Then, based on the extrinsic parameters between the camera and the pose data acquisition device, the camera pose data at various times can be calculated. This disclosure does not limit the source of the camera pose data. Camera pose data may include, for example, at least one of the following: the camera's coordinates in the world coordinate system, Euler angles, and direction cosine matrix.
[0069] Step 103: Based on the camera pose data and the camera's intrinsic parameters, perform projection processing on each point in the point cloud to determine the target pixel corresponding to each point.
[0070] Here, the target pixel is the projection point of the point under the camera pose data corresponding to the target pixel row, and the target pixel row is the pixel row where the target pixel is located.
[0071] Camera pose data can be used for point cloud projection, calculating the coordinates of each point in the point cloud in the camera coordinate system, and then converting this to the two-dimensional coordinates (i.e., projected coordinates) of the point cloud in the image coordinate system of the target image. Different pixel rows correspond to different camera poses, therefore, the actual image coordinate systems of different pixel rows are also different. By determining the corresponding target pixel based on the projected coordinates of a point in the image coordinate system, the color information of the target pixel can be used to colorize that point.
[0072] Optionally, a target pixel is defined as the pixel in the point cloud whose projected coordinates in the camera pose data corresponding to a certain pixel row are exactly located in that pixel row (or the difference in row number between the target pixel and the pixel row is less than a threshold, such as less than 2). For each point in the point cloud, the target pixel can be found by calculating the projected coordinates of that point in the camera pose data corresponding to each pixel row.
[0073] Optionally, the point cloud can be selected from the point cloud within the camera's field of view. For example, the point cloud region that may be projected into the target image can be calculated based on the camera pose corresponding to the first pixel row, the camera pose of the last pixel row, and the camera's field of view.
[0074] Step 104: Use the target pixel corresponding to each point to perform coloring processing on the point cloud.
[0075] In one possible implementation, for a point in a point cloud, the color value of its corresponding target pixel can be directly saved as the color value of that point. Alternatively, the corresponding target pixel and its surrounding neighboring pixels can be selected as a coloring window, and the average color value of each pixel in the coloring window can be saved as the color value of that point. This disclosure does not limit the specific method for calculating color values.
[0076] Based on the embodiments of this disclosure, when the LiDAR and the camera are synchronized, the timestamp of each pixel row is calculated according to the exposure time and pixel readout time of the target image, thereby determining the camera pose data when the camera collects each row of pixels, and then determining the true projection position of each point in the point cloud in the target image. That is, when the projection point under the camera pose data corresponding to a certain pixel row falls exactly on that pixel row, it can be determined that the projection point at this time is the target pixel of that point. Using the target pixel corresponding to each point to color the point cloud can reduce the difference between the point cloud color and the actual color caused by the rolling shutter effect, improve the accuracy of point cloud coloring, and eliminate the need to use a global shutter camera, resulting in lower cost.
[0077] In one possible implementation, in order to speed up the convergence of iterative calculations and improve the efficiency of querying the target pixel corresponding to each point in the point cloud, projection can be performed first onto the middle row of the target image. Step 103 can specifically include the following steps:
[0078] Step 103a: Using a preset pixel row in the target image as the initial projection row, calculate the projection coordinates of the point in the image plane coordinate system where the initial projection row is located based on the camera pose data and camera intrinsic parameters corresponding to the initial projection row.
[0079] Optionally, the initial projection row can be selected as the middle row of the target image. For example, when the target image has n pixel rows, the n / 2 pixel row or the (n+1) / 2 pixel row can be selected as the initial projection row.
[0080] Based on the camera pose data and camera intrinsic parameters corresponding to the initial projection row, the projected coordinates of the point in the image plane coordinate system where the initial projection row is located are calculated. Specifically, firstly, based on the camera pose data corresponding to the initial projection row and the point's first coordinate in the world coordinate system, a first coordinate transformation is performed on the point to obtain the second coordinate of the point in the camera coordinate system corresponding to the initial projection row. Then, based on the second coordinate and the camera intrinsic parameters, a second coordinate transformation is performed on the point to obtain the projected coordinates of the point in the image plane coordinate system where the initial projection row is located.
[0081] Based on the positional relationship between the projected coordinates and the initial projected row, select to execute the following steps 103b and 103c.
[0082] Step 103b: In response to the fact that the number of rows corresponding to the projected coordinates is the same as the number of rows of the initial projected rows or the difference in the number of rows is less than the difference threshold, the pixel at the projected coordinates is determined as the target pixel corresponding to the point.
[0083] Optionally, the projection coordinates of a point in the image plane coordinate system where the initial projection row (m-th row) is located are (u, v), where v represents the row number of the pixel row where the projection coordinates are located. If the row number of the pixel row where the projection coordinates are located is the same as the row number of the initial projection row or the difference in row number is less than the difference threshold (2), that is, v = m or v < (m ± 2), then it can be considered that the projection point of the point in the camera pose corresponding to the initial projection row is exactly located in the initial projection row, that is, the pixel point at the projection coordinates is the target pixel point of the point in the target image.
[0084] Step 103c: In response to the difference between the number of rows corresponding to the projection coordinates and the number of rows of the initial projection row being greater than or equal to the difference threshold, the pixel row where the projection coordinates were calculated in the previous iteration is used as the iterative projection row in the next iteration calculation. Based on the camera pose data and camera intrinsic parameters corresponding to the iterative projection row, the projection coordinates of the point in the image plane coordinate system where the iterative projection row is located are calculated until the difference between the number of rows corresponding to the calculated projection coordinates and the number of rows of the iterative projection row is less than the difference threshold.
[0085] Optionally, the projection coordinates of a point in the image plane coordinate system where the initial projection row (the m-th row) is located are (u, v), where v represents the row number of the pixel row where the projection coordinates are located. If the difference between the row number of the pixel row where the projection coordinates are located and the row number of the initial projection row is greater than or equal to the difference threshold (2), i.e., v≥(m±2), it means that the point should correspond to a pixel in another pixel row acquired by the camera under other camera poses. Therefore, it is necessary to continue iterative calculation based on the camera pose data of other pixel rows until the target pixel is found (i.e., the pixel corresponding to the projection coordinates when the projection coordinates under the camera pose data corresponding to a certain pixel row are exactly located in that pixel row).
[0086] Optionally, the projection coordinates of the point in the camera pose data corresponding to each pixel row can be calculated sequentially according to a preset direction (such as gradually moving away from the initial projection row), and it can be determined whether the projection coordinates are located in the pixel row or whether the difference in the number of rows is less than the difference threshold. Alternatively, the pixel row where the projection coordinates of the previous calculation are located can be the iterative projection row in the next iteration calculation. The projection coordinates can be calculated based on the camera pose data and camera intrinsic parameters corresponding to the iterative projection row, and it can be determined whether the projection coordinates are located in the pixel row or whether the difference in the number of rows is less than the difference threshold.
[0087] Indicatively, the target image consists of 9 pixel rows. The 5th pixel row is used as the initial projection row. The projection coordinates of the i-th point in the point cloud are calculated based on the camera pose data of the 5th pixel row, resulting in projection coordinates (11, 8). The row difference is 3, which is greater than the difference threshold of 2. Then, the projection coordinates of the i-th point are calculated based on the camera pose data of the 8th pixel row, resulting in projection coordinates (6, 7). The row difference is 1, which is less than the difference threshold of 2. Therefore, the 6th pixel in the 7th row can be determined as the target pixel of the i-th point in the point cloud.
[0088] If the projected coordinates are not within the coordinate range of the target image, it means that the camera failed to capture the content corresponding to that point. In this case, the iterative calculation and coloring process for that point can be terminated directly, and the projection calculation for the next point in the point cloud can continue.
[0089] Optionally, the projection process involves transforming the coordinates of points in the point cloud from the world coordinate system to the camera coordinate system, and then to the image plane coordinate system. Step 103d above, "calculating the projected coordinates of the points in the image plane coordinate system where the iterative projection rows are located, based on the camera pose data and camera intrinsic parameters corresponding to the iterative projection rows," may specifically include the following steps:
[0090] Based on the camera pose data and the point's first coordinates in the world coordinate system, a first coordinate transformation is performed on the point to obtain its second coordinates in the camera coordinate system corresponding to the camera pose data. Based on the second coordinates and camera intrinsic parameters, a second coordinate transformation is performed on the point to obtain the projected coordinates. The camera intrinsic parameters are used to characterize the mapping relationship between the camera coordinate system and the image plane coordinate system. Camera intrinsic parameters are parameter matrices used to describe the internal properties of the camera, including focal length, principal point (optical center) coordinates, distortion coefficients, etc. Camera intrinsic parameters are usually determined during camera calibration and do not change over time, so they can be directly read from a pre-stored parameter file.
[0091] For illustration, assume that the camera pose data of the j-th pixel row in the target image is T. j The first coordinate of a point in the point cloud in the world coordinate system is P. world First, based on the camera pose data and the point's first coordinate in the world coordinate system, a first coordinate transformation is performed on the point to obtain the point's second coordinate P in the camera coordinate system corresponding to the camera pose data. j Then, based on the camera intrinsic parameters and the second coordinate, a second coordinate transformation is performed on the point to obtain the projected coordinates.
[0092] To illustrate, the second coordinate can be obtained by multiplying the first coordinate and the reciprocal of the camera pose data, as shown in the following formula:
[0093]
[0094] As an illustration, the projected coordinates can be obtained by performing matrix multiplication on the second coordinate and the camera intrinsic parameters. The calculation formula is as follows:
[0095]
[0096] in, For projected coordinates, The second coordinate, f is the camera's intrinsic parameter. x f y c represents the focal length along the x-axis and the focal length along the y-axis, respectively. x c y These represent the offsets of the origin of the image plane coordinate system relative to the origin of the camera coordinate system.
[0097] Optionally, the target in this embodiment can be pre-distorted, so the coordinate transformation process can disregard camera distortion. The formula for calculating camera distortion removal is as follows:
[0098] x″=x′(1+k1r 2 +k2r 4 +k3r 6 )+2p1x′y′+p2(r 2 +2x′ 2 ) Formula (3)
[0099] y″=y′(1+k1r 2 +k2r 4 +k3r 6 )+2p2x′y′+p1(r 2 +2y′ 2 ) Formula (4)
[0100] Where (x′, y′) represents the coordinates of a pixel in the image without camera distortion, (x″, y″) represents the coordinates of a pixel in the image with camera distortion, r represents the distance from the pixel to the center of the image, k1, k2, and k3 are radial distortion correction coefficients, and p1 and p2 are tangential distortion correction coefficients.
[0101] Based on the embodiments of this disclosure, by setting the middle row of the target image as the initial projection row, and continuing iterative calculation based on the row number indicated by the calculated projection coordinates until the target pixel corresponding to the point in the point cloud is determined, the number of iterations can be reduced, the efficiency of determining the target pixel can be improved, and thus the point cloud coloring efficiency can be improved.
[0102] In one possible implementation, the camera pose data at each data acquisition time can be pre-calculated to obtain a camera pose list, thereby finding the camera pose data corresponding to each pixel row based on the timestamp of each pixel row in the target image. For example... Figure 2 As shown, step 102 above may specifically include the following steps:
[0103] Step 201: For each pixel row in the target image, compare the timestamp of the pixel row with the acquisition time of each camera pose data in the camera pose list.
[0104] In the camera pose list, each camera pose data point corresponds one-to-one with its acquisition time. Optionally, the acquisition frequency and start time of information from units such as cameras and LiDAR may differ. Therefore, the camera pose list may contain camera pose data acquired at the same time as a pixel row, or it may not contain such data. By comparing the timestamp of the pixel row with the acquisition time of each camera pose data point in the camera pose list, step 202 or step 203 can be performed based on the query results.
[0105] Step 202: In response to the existence of a target acquisition time with the same timestamp, the camera pose data at the target acquisition time is determined as the camera pose data corresponding to the pixel row.
[0106] If the camera pose list contains a target acquisition time that has the same timestamp as the pixel row, then the camera pose data at the target acquisition time is determined as the camera pose data corresponding to that pixel row.
[0107] Step 203: In response to the absence of a target acquisition time, interpolation calculation is performed based on the camera pose data corresponding to at least two sets of acquisition times adjacent to the timestamp to obtain the camera pose data corresponding to the pixel row.
[0108] If there is no target acquisition time in the camera pose list that matches the timestamp of the pixel row, the camera pose data corresponding to that pixel row can be calculated by interpolation.
[0109] As an illustration, if there is no target acquisition time with the same timestamp as the pixel row in the camera pose list, the camera pose data corresponding to the timestamp of the pixel row can be calculated by linear interpolation. Let the timestamp of the j-th pixel row be t. j The acquisition time t in the camera pose list was queried. k and t k+1 Between these points, the corresponding camera pose data is T. k and T k+1 Then the camera pose data T of the j-th pixel row j The calculation formula is as follows:
[0110] T j =T k ·(t k+1 -t j ) / (t k+1 -t k )+T k+1 ·(t j -t k+1 ) / (t k+1 -t k ) Formula (5)
[0111] In one possible implementation, since the camera itself cannot acquire pose data, pose data can be acquired through other devices in the laser scanning equipment. Then, the camera pose data is obtained by converting the extrinsic parameters between these devices and the camera. Optionally, the laser scanning equipment also includes an inertial measurement unit (IMU). The IMU acquires its own pose data (first pose data) at a first preset frequency, and the lidar acquires its own pose data (second pose data) at a second preset frequency. Figure 3As shown, the method provided in this embodiment further includes the following steps:
[0112] Step 301: Obtain the first pose list acquired by the inertial measurement unit and the second pose list acquired by the lidar.
[0113] The first pose data in the first pose list is used to characterize the position and attitude of the inertial measurement unit in the world coordinate system, and the second pose data in the second pose list is used to characterize the position and attitude of the lidar in the world coordinate system.
[0114] To illustrate, the pose acquisition frequency of the inertial measurement unit is 200Hz, and the pose acquisition frequency of the lidar is 10Hz.
[0115] Step 302: Based on the first pose list and the first extrinsic parameters of the camera relative to the inertial measurement unit, obtain the first camera pose list, and based on the second pose list and the second extrinsic parameters of the camera relative to the lidar, obtain the second camera pose list.
[0116] The first extrinsic parameter is used to characterize the relative position and attitude of the camera relative to the inertial measurement unit in the world coordinate system. For example, it may include, but is not limited to, at least one of the translation matrix and rotation matrix of the camera relative to the inertial measurement unit. The second extrinsic parameter is used to characterize the relative position and attitude of the camera relative to the lidar in the world coordinate system. For example, it may include, but is not limited to, at least one of the translation matrix and rotation matrix of the camera relative to the lidar.
[0117] Optionally, the camera, inertial measurement unit, and lidar are fixedly installed in the laser scanning device. The first and second extrinsic parameters are fixed and can be read directly from a pre-stored file.
[0118] Illustratively, the first camera pose data can be obtained by performing matrix multiplication based on the first pose data and the first extrinsic parameters. The parameter matrix corresponding to the first pose data consists of a translation matrix representing the position of the inertial measurement unit (IMU) in the world coordinate system and a rotation matrix representing the azimuth. The parameter matrix corresponding to the first extrinsic parameters consists of a translation matrix representing the relative position of the camera with respect to the IMU in the world coordinate system and a rotation matrix representing the relative azimuth. Illustratively, the formula for calculating the first camera pose data is as follows:
[0119] T camera→world1 =T imu→world T camera→imu Formula (6)
[0120] Among them, T camera→world1 For the first camera pose data, T imu→world For the first pose data, Tcamera→imu It is the first external parameter.
[0121] Correspondingly, matrix multiplication can be performed based on the second pose data and the second extrinsic parameters to obtain the second camera pose data. The parameter matrix corresponding to the second pose data consists of a translation matrix representing the position of the lidar in the world coordinate system and a rotation matrix representing the azimuth. Similarly, the parameter matrix corresponding to the second extrinsic parameters consists of a translation matrix representing the relative position of the camera relative to the lidar in the world coordinate system and a rotation matrix representing the relative azimuth. The calculation formula for the second camera pose data is illustrated below:
[0122] T camera→world2 =T lidar→world T camera→lidar Formula (7)
[0123] Among them, T camera→world2 For the pose data of the second camera, T lidar→world For the second pose data, T camera→lidar It is the second external parameter.
[0124] Step 303: Using the second camera pose list as a priori constraint and the pre-integration interpolation result of the first camera pose list as a relative factor constraint, perform global joint optimization on the first camera pose list and the second camera pose list to obtain the camera pose list.
[0125] Optionally, an optimization algorithm can be employed. This algorithm inputs the second camera pose list as a priori constraints, specifies the covariance matrix, and performs pre-integration interpolation using the first camera pose data between two adjacent second camera pose data points. The pre-integration interpolation result is then used as a relative factor constraint for global joint optimization to solve for the camera pose list. Illustratively, the optimization algorithm can utilize libraries such as GTASM (Georgia Tech Smoothing and Mapping) or Ceres.
[0126] In one possible implementation, the laser scanning device has an independent clock, and the camera, lidar, and inertial measurement unit are pre-synchronized based on the independent clock. Prior to step 301 above, the method provided in this disclosure embodiment further includes the following steps:
[0127] Based on the preset clock synchronization method, independent clocks are used to synchronize the internal clocks of the lidar, the camera, and the inertial measurement unit, respectively, to obtain the acquisition time of each first pose data in the first pose list, the acquisition time of each second pose data in the second pose list, and the acquisition time of the target image.
[0128] In illustrative purposes, clock synchronization methods may include, but are not limited to, at least one of the following: pulse per second (PPS) synchronization method, Global Navigation Satellite System (GNSS) synchronization method, and Precision Time Protocol (PTP) synchronization method.
[0129] The independent clock is a high-precision clock source (e.g., a high-precision active crystal oscillator) that exists independently of the camera, inertial measurement unit (IMU), and lidar, providing stable clock information with minimal error. Optionally, the time provided by the independent clock can be the standard time provided by the National Time Service Center, or an independently provided time that can be used as the trigger moment at any time (e.g., the moment when the scanning equipment starts operating as time 0). For example, when using the PPS synchronization method, the independent clock can send PPS signals to the camera, IMU, and lidar according to a preset cycle, controlling the internal clocks of the camera, IMU, and lidar to return to zero. The camera, IMU, and lidar respectively collect information and add internal timestamps. The independent clock can determine the time offset value corresponding to the internal clocks of the camera, IMU, and lidar based on the internal timestamps carried in the information sent by the camera, IMU, and lidar. Thus, based on the time of the independent clock and the internal offset value, the acquisition time of each frame of image and the acquisition time of each frame of pose data can be calculated.
[0130] Based on the embodiments of this disclosure, pose data is acquired using an inertial measurement unit and a lidar, and global joint optimization is performed by combining the two pose data, which can improve the accuracy of camera pose data. Furthermore, before data acquisition, the camera, inertial measurement unit, and lidar are synchronized using independent clocks, which can avoid data calculation errors caused by time differences.
[0131] In one possible implementation, the timestamp of each pixel row in the target image can be determined based on the acquisition time of the target image, the exposure time, and the pixel readout time. Step 101 above may specifically include the following steps:
[0132] The timestamp of each pixel row is determined based on the acquisition time, exposure time, pixel readout time, and total number of rows of the target image.
[0133] Wherein, the total number of rows is the number of pixel rows in the target image, the timestamp of the j-th pixel row is the sum of the acquisition time of the target image, half of the exposure time, and the readout time of the j-th pixel row, the readout time of the j-th pixel row is the product of j and the average readout time, the average readout time is the quotient of the pixel readout time and the total number of rows, and j is a positive integer.
[0134] Indicative, the timestamp t of the j-th pixel row. j The calculation formula is as follows:
[0135]
[0136] Among them, t start t represents the acquisition time of the target image, n represents the total number of rows in the target image, and t represents the acquisition time of the target image. read t represents the pixel readout time of the target image. exp This represents the exposure duration for one pixel row in the target image. Since camera exposure is triggered by an electrical signal, the timestamp at which each image begins to be captured can be obtained, i.e., the image acquisition time t. start Since the camera will immediately start exposing the next pixel row j+1 after exposing one pixel row j, and will read the pixel information of pixel row j at the same time, it is only necessary to add the readout time of the first j pixel rows and half of the exposure time of the j-th pixel row (taking the median value) to the acquisition time of the target image to regard it as the timestamp of the j-th pixel row.
[0137] Figure 4 A structural block diagram of a point cloud coloring processing apparatus provided in an exemplary embodiment of this disclosure is shown. The point cloud coloring processing apparatus is applied to a laser scanning device, which includes a lidar and a camera. The lidar is used to generate point clouds for a target scene, and the camera is used to acquire images of the target scene. The lidar and the camera are clock-synchronized. The point cloud coloring processing apparatus includes:
[0138] The first determining module 401 is used to determine the timestamp of each pixel row in the target image based on the exposure time and pixel readout time of the target image, wherein the exposure time is the time required for the camera to generate a row of pixels, and the pixel readout time is the time required for the camera's processor to read the pixel information of the target image row by row.
[0139] The second determining module 402 is used to determine the camera pose data corresponding to each pixel row based on the timestamp of each pixel row determined by the first determining module 401. The camera pose data is used to characterize the position and orientation of the camera in the world coordinate system when generating the corresponding row of pixels.
[0140] The third determining module 403 is used to perform projection processing on each point in the point cloud based on the camera pose data and camera intrinsic parameters determined by the second determining module 402, and determine the target pixel corresponding to each point; wherein, the target pixel is the projection point of the point under the camera pose data corresponding to the target pixel row, and the target pixel row is the pixel row where the target pixel is located.
[0141] The coloring processing module 404 is used to color the point cloud using the target pixel points corresponding to each point determined by the third determining module 403.
[0142] Optionally, in one possible implementation, the third determining module 403 described above can also be used for:
[0143] Using a preset pixel row in the target image as the initial projection row, and based on the camera pose data and camera intrinsic parameters corresponding to the initial projection row, calculate the projection coordinates of the point in the image plane coordinate system where the initial projection row is located.
[0144] In response to the fact that the number of rows corresponding to the projected coordinates is the same as the number of rows in the initial projection row or the difference in the number of rows is less than the difference threshold, the pixel at the projected coordinates is determined as the target pixel corresponding to the point.
[0145] If the difference between the number of rows corresponding to the projected coordinates and the number of rows in the initial projected row is greater than or equal to the difference threshold, the pixel row where the previously calculated projected coordinates are located is used as the iterative projected row in the next iteration. Based on the camera pose data and camera intrinsic parameters corresponding to the iterative projected row, the projected coordinates of the point in the image plane coordinate system where the iterative projected row is located are calculated until the difference between the number of rows corresponding to the calculated projected coordinates and the number of rows in the iterative projected row is less than the difference threshold.
[0146] Optionally, in one possible implementation, the third determining module 403 described above can also be used for:
[0147] Based on the camera pose data and the first coordinate of the point in the world coordinate system, the first coordinate transformation is performed on the point to obtain the second coordinate of the point in the camera coordinate system corresponding to the camera pose data.
[0148] Based on the second coordinate and camera intrinsic parameters, the point is transformed into the second coordinate to obtain the projected coordinate. The camera intrinsic parameters are used to characterize the mapping relationship between the camera coordinate system and the image plane coordinate system.
[0149] Optionally, in one possible implementation, the second determining module 402 described above can also be used for:
[0150] For each row of pixels in the target image, compare the timestamp of the pixel row with the acquisition time of each camera pose data in the camera pose list;
[0151] In response to the existence of a target acquisition time with the same timestamp, the camera pose data at the target acquisition time is determined as the camera pose data corresponding to the pixel row;
[0152] In response to the absence of a target acquisition time, interpolation calculations are performed based on camera pose data corresponding to at least two acquisition times adjacent to the timestamp to obtain camera pose data corresponding to the pixel row.
[0153] Optionally, in one possible implementation, the laser scanning device also includes an inertial measurement unit; such as Figure 5 As shown, the point cloud coloring processing apparatus may further include:
[0154] The first acquisition module 501 is used to acquire the first pose list collected by the inertial measurement unit and the second pose list collected by the lidar. The first pose data in the first pose list is used to characterize the position and attitude of the inertial measurement unit in the world coordinate system, and the second pose data in the second pose list is used to characterize the position and attitude of the lidar in the world coordinate system.
[0155] The second acquisition module 502 is used to obtain a first camera pose list based on the first pose list obtained by the first acquisition module 501 and the first extrinsic parameter of the camera relative to the inertial measurement unit, and to obtain a second camera pose list based on the second pose list and the second extrinsic parameter of the camera relative to the lidar. The first extrinsic parameter is used to characterize the relative position and relative attitude of the camera relative to the inertial measurement unit in the world coordinate system, and the second extrinsic parameter is used to characterize the relative position and relative attitude of the camera relative to the lidar in the world coordinate system.
[0156] The optimization module 503 is used to perform global joint optimization on the first camera pose list and the second camera pose list, using the second camera pose list as a prior constraint and the pre-integration interpolation result of the first camera pose list as a relative factor constraint, to obtain the camera pose list.
[0157] Optionally, in one possible implementation, the laser scanning device has an independent clock; the point cloud coloring processing apparatus further includes a clock synchronization module for:
[0158] Based on the preset clock synchronization method, independent clocks are used to synchronize the internal clocks of the lidar, the camera, and the inertial measurement unit, respectively, to obtain the acquisition time of each first pose data in the first pose list, the acquisition time of each second pose data in the second pose list, and the acquisition time of the target image.
[0159] Optionally, in one possible implementation, the first determining module 401 described above can also be used for:
[0160] The timestamp of each pixel row is determined based on the acquisition time, exposure time, pixel readout time, and total number of rows of the target image.
[0161] Wherein, the total number of rows is the number of pixel rows in the target image, the timestamp of the j-th pixel row is the sum of the acquisition time of the target image, half of the exposure time, and the readout time of the j-th pixel row, the readout time of the j-th pixel row is the product of j and the average readout time, the average readout time is the quotient of the pixel readout time and the total number of rows, and j is a positive integer.
[0162] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. For identical, similar, or corresponding parts between embodiments, please refer to the corresponding sections. Since the method, apparatus, and device embodiments are basically corresponding, relevant parts can be referred to accordingly. The methods, apparatus, and devices in the embodiments of this disclosure also correspond to each other in specific implementation and beneficial technical effects; related content can be referred to mutually and will not be repeated here.
[0163] In addition, this disclosure also provides an electronic device, including:
[0164] Memory, used to store computer programs;
[0165] A processor is configured to execute a computer program stored in the memory, wherein when the computer program is executed, it implements the image processing method described in any of the above embodiments of the present disclosure.
[0166] Figure 6 This is a schematic diagram illustrating the structure of an application embodiment of the electronic device disclosed herein. Below, reference is made to… Figure 6 This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them.
[0167] like Figure 6 As shown, the electronic device includes one or more processors and memory.
[0168] A processor can be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and can control other components in an electronic device to perform desired functions.
[0169] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the image processing methods of the various embodiments of this disclosure described above and / or other desired functions.
[0170] In one example, the electronic device may also include input devices and output devices, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0171] In addition, the input device may include, for example, a keyboard, a mouse, etc.
[0172] This output device can output various information to the outside, including determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0173] Of course, for the sake of simplicity, Figure 6 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0174] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps in the image processing methods according to various embodiments of this disclosure as described in the foregoing portion of this specification.
[0175] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0176] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform the steps in the image processing methods according to various embodiments of this disclosure as described in the foregoing portion of this specification.
[0177] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0178] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as ROM, RAM, magnetic disk, or optical disk.
[0179] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0180] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0181] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0182] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0183] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0184] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0185] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A point cloud coloring processing method, characterized in that, The method is applied to a laser scanning device, which is equipped with a lidar and a camera. The lidar is used to generate a point cloud for a target scene, and the camera is used to acquire images of the target scene. The lidar and the camera are synchronized in time. Based on the exposure time and pixel readout time of the target image, the timestamp of each pixel row in the target image is determined. The exposure time is the time required for the camera to generate a row of pixels, and the pixel readout time is the time required for the camera's processor to read the pixel information of the target image row by row. The camera pose data corresponding to each pixel row is determined based on the timestamp of each pixel row. The camera pose data is used to characterize the position and orientation of the camera in the world coordinate system when generating the corresponding row of pixels. Based on the camera pose data and the camera intrinsic parameters of the camera, projection processing is performed on each point in the point cloud to determine the target pixel corresponding to each point; wherein, the target pixel is the projection point of the point under the camera pose data corresponding to the target pixel row, and the target pixel row is the pixel row where the target pixel is located. The point cloud is colored using the target pixel corresponding to each point.
2. The method according to claim 1, characterized in that, The step of projecting each point in the point cloud based on the camera pose data and the camera intrinsic parameters to determine the target pixel corresponding to each point includes: Using a preset pixel row in the target image as the initial projection row, and based on the camera pose data corresponding to the initial projection row and the camera intrinsic parameters, calculate the projection coordinates of the point in the image plane coordinate system where the initial projection row is located; In response to the fact that the number of rows corresponding to the projected coordinates is the same as the number of rows of the initial projected rows or the difference in the number of rows is less than the difference threshold, the pixel at the projected coordinates is determined as the target pixel corresponding to the point; In response to the difference between the number of rows corresponding to the projected coordinates and the number of rows of the initial projected row being greater than or equal to the difference threshold, the pixel row where the previously calculated projected coordinates are located is used as the iterative projected row in the next iteration calculation. Based on the camera pose data corresponding to the iterative projected row and the camera intrinsic parameters, the projected coordinates of the point in the image plane coordinate system where the iterative projected row is located are calculated until the difference between the number of rows corresponding to the calculated projected coordinates and the number of rows of the iterative projected row is less than the difference threshold.
3. The method according to claim 2, characterized in that, The step of calculating the projection coordinates of the point in the image plane coordinate system where the iterative projection row is located, based on the camera pose data corresponding to the iterative projection row and the camera intrinsic parameters, includes: Based on the camera pose data and the first coordinates of the point in the world coordinate system, the point is transformed to obtain the second coordinates of the point in the camera coordinate system corresponding to the camera pose data. Based on the second coordinates and the camera intrinsic parameters, the point is transformed to obtain the projected coordinates. The camera intrinsic parameters are used to characterize the mapping relationship between the camera coordinate system and the image plane coordinate system.
4. The method according to any one of claims 1 to 3, characterized in that, The process of determining the camera pose data corresponding to each pixel row based on the timestamp of each pixel row includes: For each pixel row in the target image, compare the timestamp of the pixel row with the acquisition time of each camera pose data in the camera pose list; In response to the existence of a target acquisition time with the same timestamp, the camera pose data of the target acquisition time is determined as the camera pose data corresponding to the pixel row; In response to the absence of the target acquisition time, interpolation calculation is performed based on the camera pose data corresponding to at least two sets of acquisition times adjacent to the timestamp to obtain the camera pose data corresponding to the pixel row.
5. The method according to claim 4, characterized in that, The laser scanning device is also equipped with an inertial measurement unit; The method of comparing the timestamp of each pixel row in the target image with the acquisition time of each camera pose data in the camera pose list before the acquisition time of each pixel row is described. Obtain the first pose list acquired by the inertial measurement unit and the second pose list acquired by the lidar, wherein the first pose data in the first pose list is used to characterize the position and attitude of the inertial measurement unit in the world coordinate system, and the second pose data in the second pose list is used to characterize the position and attitude of the lidar in the world coordinate system. Based on the first pose list and the first extrinsic parameter of the camera relative to the inertial measurement unit, a first camera pose list is obtained, and based on the second pose list and the second extrinsic parameter of the camera relative to the lidar, a second camera pose list is obtained. The first extrinsic parameter is used to characterize the relative position and relative attitude of the camera relative to the inertial measurement unit in the world coordinate system, and the second extrinsic parameter is used to characterize the relative position and relative attitude of the camera relative to the lidar in the world coordinate system. Using the second camera pose list as a priori constraint and the pre-integration interpolation result of the first camera pose list as a relative factor constraint, a global joint optimization is performed on the first camera pose list and the second camera pose list to obtain the camera pose list.
6. The method according to claim 5, characterized in that, The laser scanning device has an independent clock; Before acquiring the first pose list collected by the inertial measurement unit and the second pose list collected by the lidar, the method further includes: Based on a preset clock synchronization method, the independent clocks are used to synchronize the internal clocks of the lidar, the camera, and the inertial measurement unit, respectively, to obtain the acquisition time of each first pose data in the first pose list, the acquisition time of each second pose data in the second pose list, and the acquisition time of the target image.
7. The method according to any one of claims 1 to 3, characterized in that, The determination of the timestamp for each pixel row in the target image based on the exposure time and pixel readout time of the target image includes: The timestamp of each pixel row is determined based on the acquisition time of the target image, the exposure duration, the pixel readout duration, and the total number of rows of the target image; Wherein, the total number of rows is the number of pixel rows in the target image, the timestamp of the j-th pixel row is the sum of the acquisition time of the target image, half of the exposure time, and the readout time of the j-th pixel row, the readout time of the j-th pixel row is the product of j and the average readout time, and the average readout time is the quotient of the pixel readout time and the total number of rows, where j is a positive integer.
8. A point cloud coloring processing device, characterized in that, An apparatus for use in laser scanning equipment, wherein the laser scanning equipment is equipped with a lidar and a camera, the lidar is used to generate point clouds for a target scene, and the camera is used to acquire images of the target scene; the lidar and the camera are synchronized by a clock; the apparatus includes: The first determining module is used to determine the timestamp of each pixel row in the target image based on the exposure time and pixel readout time of the target image, wherein the exposure time is the time required for the camera to generate a row of pixels, and the pixel readout time is the time required for the camera's processor to read the pixel information of the target image row by row; The second determining module is used to determine the camera pose data corresponding to each pixel row based on the timestamp of each pixel row. The camera pose data is used to characterize the position and orientation of the camera in the world coordinate system when generating the corresponding row of pixels. The third determining module is used to perform projection processing on each point in the point cloud based on the camera pose data and the camera intrinsic parameters of the camera to determine the target pixel point corresponding to each point; wherein, the target pixel point is the projection point of the point under the camera pose data corresponding to the target pixel row, and the target pixel row is the pixel row where the target pixel point is located. A coloring processing module is used to color the point cloud using the target pixel corresponding to each point.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program stored in the memory, wherein when the computer program is executed, it implements the method described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-7.
Citation Information
Patent Citations
Point cloud map generation method and device, and electronic equipment
CN111784834A
Image-based color dense point cloud generation method and device
CN115578522A