Image and point cloud multi-modal data enhancement method
Through the multimodal data enhancement method of image and point cloud, combined with true sample pasting enhancement and bidirectional enhancement, and feature alignment, the problem of inability to align the internal and external parameter mapping relationship between image and point cloud data, and the accuracy of accurate alignment and detection model is improved.
Patent Information
- Application Number
- CN202510276737.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-13
AI Technical Summary
When the existing multimodal fusion method enhances the operation between image and point cloud data, the internal and external parameter mapping relationship between spatial point cloud and plane images cannot be accurately aligned, resulting in dislocation.
The multimodal data enhancement method of image and point cloud is used to generate a true sample set through instance segmentation network, and true sample pasting enhancement (GTSP) and bidirectional enhancement (TrAug) are performed, and feature alignment is performed to ensure the accurate alignment of the internal and external parameter mapping relationships.
When enhancing operations between images and point cloud data, the precise alignment of internal and external parameter mapping relationships is achieved, which avoids misalignment and improves the accuracy of the detection model.
Smart Images

Figure CN120147382A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing. Specifically, it relates to a method for enhancing multi-modal data of images and point clouds. Background Art
[0002] In order to achieve robust and accurate environmental perception, intelligent vehicles are usually equipped with multiple sensors. Among them, cameras and lidars are widely used in 3D object detection due to the complementarity of their data. Cameras provide rich color and texture information in the form of high-resolution two-dimensional pixels, while lidars provide spatial position and size shape information in the form of sparse three-dimensional point clouds. Given the differences between the data of these two types of sensors, how to effectively integrate image information to strengthen point cloud data and thus improve the accuracy of the detection model has become a crucial research topic.
[0003] In supervised deep learning models, the quantity and diversity of training data are undoubtedly the key factors affecting model performance. In the existing datasets widely used in the field of autonomous driving, there are significant differences in the quantities of cars, pedestrians, and cyclists. Usually, the number of cars is the largest. Directly using such an unbalanced dataset to train the model, the model often overfits to the categories with a large number (such as cars), and thus tends to ignore the categories with a small number (pedestrians and cyclists) during prediction. However, preparing such a dataset is often a time-consuming and costly task because all such samples must be manually labeled. For this reason, various data augmentation techniques have emerged, artificially increasing the quantity and diversity of training data, reducing the labeling cost while improving the generalization performance of the model.
[0004] The defect of the existing technology is that there are already many single-modal data augmentation methods for images or point clouds. Common operations include rotation, scaling, CutMix, and flipping, etc. However, when it comes to multi-modal fusion methods, the data augmentation operations become complex, and the types of their applications are relatively few. When the existing multi-modal fusion methods perform augmentation operations between image and point cloud data, the internal and external parameter mapping relationships between the spatial point cloud and the planar image cannot be accurately aligned, resulting in misalignment phenomena. Summary of the Invention
[0005] Aiming at the problem that when the existing multi-modal fusion methods perform augmentation operations between image and point cloud data, the internal and external parameter mapping relationships between the spatial point cloud and the planar image cannot be accurately aligned, resulting in misalignment phenomena, the present invention provides a method for enhancing multi-modal data of images and point clouds.
[0006] To achieve the above technical objectives, the technical solution adopted by the present invention is as follows:
[0007] A method for enhancing multi-modal data of images and point clouds, comprising the steps:
[0008] S1. Collect image data through a camera, collect point cloud data through a lidar, and combine them to obtain a training dataset;
[0009] S2. Generate a ground truth sample set containing all image ground truth blocks and their corresponding point cloud blocks from the training dataset through an instance segmentation network;
[0010] S3. Paste the qualified ground truth sample set (GTSP) into the original training dataset through occlusion threshold judgment and bounding box collision analysis;
[0011] S4. Process the image and point cloud data using a bidirectional enhancement module (TrAug) to obtain enhanced images and point clouds;
[0012] S5. Align the features of the enhanced images and point clouds.
[0013] Further, the detailed steps of step S3 include:
[0014] Judge whether there is an intersection in the image and point cloud ground truth samples through occlusion threshold judgment and bounding box collision analysis; check whether there is an intersection in the bounding boxes of the image and point cloud ground truth samples in the Bev view.
[0015] Delete the image and point cloud ground truth samples with intersections;
[0016] Guided by depth information, paste the qualified samples into the current frame of training data to increase the number of ground truth samples in the image and point cloud of this frame.
[0017] Further, the detailed steps of step S4 include:
[0018] Perform a forward transformation on the image data to obtain an enhanced image;
[0019] Perform a reverse transformation on the point cloud data to obtain an enhanced point cloud.
[0020] Further, the detailed steps of the forward transformation of the image data:
[0021] Perform random scaling on the image data to obtain scaled image data;
[0022] Perform random flipping on the scaled image data to obtain an enhanced image.
[0023] Further, the detailed steps of the reverse transformation of the point cloud data:
[0024] Perform random rotation on the point cloud data to obtain rotated point cloud data;
[0025] Perform random translation on the rotated point cloud data to obtain translated point cloud data;
[0026] Randomly flip the translated point cloud data to obtain the enhanced point cloud.
[0027] During this process, the positive transformation parameters experienced by the point cloud and the image are detailedly recorded, such as the angle of random rotation around the z-axis, the translation amounts in three directions, etc. For the translation and rotation transformations of the point cloud, it is actually to transform it to a new coordinate system to generate new point coordinates. Since the camera coordinate system remains unchanged during this process, these operations will not affect the captured image, so there is no need to apply the same operations to the image data.
[0028] In the positive transformation enhancement of the point cloud, TrAug first introduces the transformation through random rotation around the z-axis. To ensure the diversity and rationality of the transformation, the rotation angle is set in the range [-45°, 45°]. To introduce slight position changes to enhance the generalization ability of the model, the point cloud is then randomly translated. The ratio of the translation amount to its own size is controlled in the interval [0.95, 1.05] compared to the original size in the three axes. Finally, the point cloud is randomly flipped in the x direction of the camera coordinate system to further enrich the variation form of the data. For the image data, TrAug mainly performs random scaling and random left-right flipping operations. The scaling ratio range is set in [0.8, 1.2], which can not only maintain the basic features of the image but also introduce certain scale changes. At the same time, the random left-right flipping operation can enhance the robustness of the model to image direction changes. During the data augmentation process, the transformation parameters of each data sample will be detailedly recorded so that in the subsequent fusion step, the inverse transformation of the point cloud and the positive transformation of the image can be accurately performed.
[0029] Enhancement of GTSP based on ground truth paste consistency:
[0030] First, collect the ground truth sample pairs of the image and the point cloud;
[0031] Select the ground truth sample pairs of the image and the point cloud from the sample set, and then paste them into the training data of the current frame according to the rules; First, check whether there is an intersection of the bounding boxes in the Bev view to judge whether there will be collisions between the ground truth sample targets in the 3D point cloud scene. Those illegal targets will be discarded to ensure the quality of the data.
[0032] During the multi-modal fusion process, the potential occlusion in the corresponding 2D image view needs to be processed.
[0033] Generate a ground truth sample set containing all the image ground truth patches and their corresponding point cloud patches from multiple training datasets;
[0034] Then, randomly select a series of ground truth sample pairs from this set during the training process and perform screening by combining occlusion threshold judgment and bounding box collision analysis;
[0035] Then, guided by the depth information, samples that meet the requirements are pasted into the training data of the current frame according to certain rules, making the newly generated data as close as possible to the real scene, increasing the number of ground truth samples in the point cloud of this frame, simulating objects existing in different environments as much as possible, and increasing the diversity of samples.
[0036] Furthermore, feature alignment includes the alignment of the lidar coordinate system and the camera coordinate system, the alignment transformation between the camera and the image, and the coordinate transformation between the image and the pixel.
[0037] Furthermore, the alignment of the lidar coordinate system and the camera coordinate system:
[0038] Let the lidar coordinate system be SP l , and the camera coordinate system be SP c . There is a point P in the current three-dimensional space. The coordinate of P l in SP l is (x l , y l , z l ), and the coordinate of P c in SP c is (x c , y c , z c ). By applying a translation and rotation transformation to the SP l coordinate system and transforming it to the SP c camera coordinate system, the three-dimensional space transformation formula of point P is:
[0039] P c = R * P l + T (1)
[0040] where R is an orthogonal rotation matrix and T is a three-dimensional translation vector matrix.
[0041] Furthermore, the alignment transformation between the camera and the image:
[0042] The camera coordinate system SP c is a three-dimensional coordinate system centered on the camera and is used to describe the position of an object. In a monocular camera, the origin of the camera coordinate system is located at the projection center O c of the lens. The x-axis is parallel to the x-axis of the image coordinate system SP img , the y-axis is parallel to the y-axis of the image coordinate system, and the z-axis coincides with the optical axis of the lens and points to the projection plane. According to the optical imaging principle of the camera, the relationship from the camera coordinate system to the image coordinate system is a perspective projection. Using similar triangles, assuming f is the focal length of the camera, and the point P c (x c , y c , z c)With the projection point P in the image coordinate system img (x i ,y i ,z i ) coordinate relation formula:
[0043]
[0044] Written as a matrix multiplication in homogeneous coordinate form is:
[0045]
[0046] Furthermore, the coordinate conversion between the image and the pixel:
[0047] The pixel coordinate SP uv is a two-dimensional array, whose rows and columns precisely identify the position of each pixel point in the image, without involving any physical units, and is called the uv coordinate system. The image coordinate system SP img is used to represent the physical position of the pixel points in the image. The origin of this coordinate system is located at the center point of the picture, the x-axis is horizontal to the right and parallel to the u-axis of the pixel coordinate system; the y-axis is vertical downward and parallel to the v-axis of the pixel coordinate system.
[0048] Assume that the physical size of each pixel in the image coordinate system is dx wide and dy high. The coordinate system SP established with the upper left corner of the image as the coordinate origin uv and the coordinate system SP established with the center of the imaging plane img The conversion relationship is as follows:
[0049]
[0050] where (uo, vo) is the image coordinate system, SP img The coordinates of the origin O of the coordinate system in the pixel coordinate system:
[0051] Written as a matrix multiplication in homogeneous coordinate form is:
[0052]
[0053] By combining equations (3), (5), and (7), the transformation alignment formula from the point cloud coordinates to the pixel coordinates can be obtained:
[0054]
[0055] where is the camera internal parameter matrix, is the external parameter matrix.
[0056] Furthermore, the origin of the lidar coordinate system is translated to the origin of the camera coordinate system, and then by rotating counterclockwise around the axis by different angles, the two coordinate systems are made to coincide.
[0057] Rotation has a total of three degrees of freedom, namely rotation around the x, y, and z axes. According to the rotation angles, rotation matrices R x , R y , R z can be obtained in three directions respectively. The total rotation matrix is expressed as their product. Depending on the order of rotation around the axes, the resulting rotation matrix is also different. Rotating in the order of first around the z-axis, then the y-axis, and finally the x-axis, the rotation matrix R = R x * R y * R z . Then the three-dimensional space transformation formula can be expressed as:
[0058]
[0059] Finally, the transformation process is written in homogeneous coordinate form as:
[0060]
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] By performing ground truth sample paste augmentation (GTSP) and two-way augmentation (TrAug) on the collected images and point cloud data and then performing feature alignment, when performing augmentation operations between the images and point cloud data, the effect of accurately aligning the internal and external parameter mapping relationships between the spatial point cloud and the planar image is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is the overall flowchart of a method for enhancing multi-modal data of images and point clouds in an embodiment of the present invention;
[0064] Figure 2 is the overall schematic diagram of multi-modal enhancement in steps S3 and S4 in an embodiment of the present invention;
[0065] Figure 3 is the detailed flowchart of the forward transformation of image data and the reverse transformation of point cloud data in step S4 in an embodiment of the present invention.
[0066] Figure 4 is the schematic diagram of the alignment and transformation of the camera and the image in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] For the convenience of those skilled in the art, the present invention will be further described below in conjunction with the embodiments and the drawings. The content mentioned in the embodiments does not limit the present invention.
[0068] As Figure 1 shown, this embodiment provides a method for enhancing multi-modal data of images and point clouds, including the steps:
[0069] S1. Collect image data through a camera, collect point cloud data through a lidar, and combine them to obtain a training dataset;
[0070] S2. Generate a ground truth sample set containing all image ground truth blocks and their corresponding point cloud blocks from the training dataset through an instance segmentation network;
[0071] S3. Paste the qualified ground truth sample set (GTSP) into the original training dataset through occlusion threshold judgment and bounding box collision analysis;
[0072] S4. Process the image and point cloud data using a bidirectional enhancement module (TrAug) to obtain enhanced images and point clouds;
[0073] S5. Align the features of the enhanced images and point clouds.
[0074] As Figure 2 shown, the detailed steps of step S3 include:
[0075] Check whether there is an intersection in the image and the point cloud ground truth sample through occlusion threshold judgment and bounding box collision analysis; check whether there is an intersection in the bounding boxes of the image and the point cloud ground truth sample in the Bev view.
[0076] Delete the image and point cloud ground truth samples with intersections;
[0077] Guided by depth information, paste the qualified samples into the current frame of training data to increase the number of ground truth samples in the image and point cloud of this frame.
[0078] Enhancement of GTSP based on ground truth paste consistency:
[0079] First, collect the ground truth sample pairs of images and point clouds;
[0080] Select the ground truth sample pairs of images and point clouds from the sample set, and then paste them into the training data of the current frame according to the rules; first, check whether there is an intersection in the Bev view of the bounding boxes to judge whether there will be a collision between the ground truth sample targets in the 3D point cloud scene, and those illegal targets will be discarded to ensure the data quality.
[0081] During the multi-modal fusion process, it is necessary to handle potential occlusions in the corresponding 2D image view.
[0082] Generate a ground truth sample set containing all image ground truth blocks and their corresponding point cloud blocks from multiple training datasets;
[0083] Then, randomly select a series of ground truth sample pairs from this set during the training process, and perform screening in combination with occlusion threshold judgment and bounding box collision analysis;
[0084] Then, guided by the depth information, samples that meet the requirements are pasted into the training data of the current frame according to certain rules, making the newly generated data as close as possible to the real scene, increasing the number of ground-truth samples in the point cloud of this frame, simulating objects existing in different environments as much as possible, and increasing the diversity of samples.
[0085] As Figure 3 shown, the detailed steps of step S4 include:
[0086] Perform a forward transformation on the image data to obtain an enhanced image;
[0087] Perform an inverse transformation on the point cloud data to obtain an enhanced point cloud.
[0088] Detailed steps of the forward transformation of the image data:
[0089] Randomly scale the image data to obtain the scaled image data;
[0090] Randomly flip the scaled image data to obtain an enhanced image.
[0091] Detailed steps of the inverse transformation of the point cloud data:
[0092] Randomly rotate the point cloud data to obtain the rotated point cloud data;
[0093] Randomly translate the rotated point cloud data to obtain the translated point cloud data;
[0094] Randomly flip the translated point cloud data to obtain an enhanced point cloud.
[0095] During this process, the forward transformation parameters experienced by the point cloud and the image are recorded in detail, such as the angle of random rotation around the z-axis, the translation amounts in three directions, etc. For the translation and rotation transformations of the point cloud, it is actually to transform it to a new coordinate system to generate new point coordinates. Since the camera coordinate system remains unchanged during this process, these operations will not affect the captured image, so there is no need to apply the same operations to the image data.
[0096] In the forward transformation enhancement of point clouds, TrAug first introduces transformations by randomly rotating around the z-axis. To ensure the diversity and rationality of the transformations, the rotation angle is set in the range [-45°, 45°]. To introduce slight position changes to enhance the generalization ability of the model, the point cloud is then randomly translated. The ratio of the translation amount to its own size is controlled in the interval [0.95, 1.05] compared to the original size in the three axes. Finally, the point cloud is randomly flipped in the x direction of the camera coordinate system to further enrich the variation form of the data. For image data, TrAug mainly performs random scaling and random left-right flipping operations. The scaling ratio range is set in [0.8, 1.2], which can not only maintain the basic features of the image but also introduce certain scale changes. At the same time, the random left-right flipping operation can enhance the robustness of the model to image direction changes. During the data augmentation process, the transformation parameters of each data sample will be recorded in detail so that in the subsequent fusion step, the inverse transformation of the point cloud and the forward transformation of the image can be accurately performed.
[0097] As Figure 4 shown, feature alignment includes the alignment of the lidar coordinate system and the camera coordinate system, the alignment transformation between the camera and the image, and the coordinate transformation between the image and the pixel.
[0098] Alignment of the lidar coordinate system and the camera coordinate system:
[0099] Let the lidar coordinate system be SP l , and the camera coordinate system be SP c . There is a point P in the current three-dimensional space. The coordinate of P in SP l is P l =(x l , y l , z l ), and the coordinate of P in SP c is P c =(x c , y c , z c ). By applying a translation and rotation transformation to the SP l coordinate system and transforming it to the SP c camera coordinate system, the three-dimensional space transformation formula of point P is:
[0100] P c = R * P l + T (1)
[0101] where R is an orthogonal rotation matrix and T is a three-dimensional translation vector matrix.
[0102] Alignment transformation between the camera and the image:
[0103] The camera coordinate system SP cIt is a camera-centered three-dimensional coordinate system used to describe the position of an object. In a monocular camera, the origin of the camera coordinate system is located at the projection center O of the lens c where. The x-axis is parallel to the x-axis of the image coordinate system SP img , the y-axis is parallel to the y-axis of the image coordinate system, and the z-axis coincides with the optical axis of the lens and points to the projection plane. According to the optical imaging principle of the camera, the relationship from the camera coordinate system to the image coordinate system is a perspective projection. Using similar triangles, assuming f is the focal length of the camera, and the point P c (x c ,y c ,z c ) in the camera coordinate system and the projection point P img (x i ,y i ,z i ) in the image coordinate system, the coordinate relationship formula is:
[0104]
[0105] Written as a matrix multiplication in homogeneous coordinate form:
[0106]
[0107] Coordinate conversion between image and pixel:
[0108] The pixel coordinate SP uv is a two-dimensional array whose rows and columns precisely identify the position of each pixel point in the image, without any physical units, and is called the uv coordinate system. The image coordinate system SP img is used to represent the physical position of the pixel points in the image. The origin of this coordinate system is located at the center point of the picture. The x-axis is horizontal to the right and parallel to the u-axis of the pixel coordinate system; the y-axis is vertical downward and parallel to the v-axis of the pixel coordinate system.
[0109] Assume that the physical size of each pixel in the image coordinate system is dx wide and dy high. The coordinate system SP uv established with the upper left corner of the image as the coordinate origin and the coordinate system SP img established with the center of the imaging plane have the following conversion relationship:
[0110]
[0111] where (uo, vo) is the coordinate of the origin O of the image coordinate system SP img in the pixel coordinate system:
[0112] Written as a matrix multiplication in homogeneous coordinate form:
[0113]
[0114] Combining equations (3), (5), and (7) gives the transformation alignment formula from point cloud coordinates to pixel coordinates:
[0115]
[0116] where is the camera intrinsic matrix, is the extrinsic matrix.
[0117] Furthermore, the origin of the lidar coordinate system is translated to the origin of the camera coordinate system, and then by rotating counterclockwise around the axis by different angles, the two coordinate systems are made to coincide.
[0118] And there are 3 degrees of freedom for rotation, namely rotation around the x, y, and z axes. According to the rotation angles, rotation matrices R x , R y , R z can be obtained respectively in the three directions. And the total rotation matrix is expressed as their product. Depending on the order of rotation around the axes, the final obtained rotation matrix is also different. Rotating in the order of z-axis, y-axis, and x-axis first, the rotation matrix R = R x * R y * R z . Then the three-dimensional space transformation formula can be expressed as:
[0119]
[0120] Finally, the transformation process is written in homogeneous coordinate system form as:
[0121]
[0122] Compared with the prior art, the present invention has the following beneficial effects:
[0123] By performing ground truth sample paste enhancement (GTSP) and two-way enhancement (TrAug) on the collected images and point cloud data and then performing feature alignment, when performing enhancement operations between the image and the point cloud data, the effect of accurately aligning the internal and external parameter mapping relationships between the spatial point cloud and the planar image is achieved.
[0124] The above has introduced in detail a method for enhancing multi-modal data of images and point clouds provided by the present application. The description of the specific embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for multimodal data enhancement of images and point clouds, characterized in that: Includes steps: S1, collect image data through the camera, collect point cloud data through the lidar, and combine them to obtain a training data set; S2, generate a set of true value samples containing all image truth blocks and their corresponding point cloud blocks from the training dataset through the instance segmentation network; S3, pasting the qualified true value sample set into the original training data set through occlusion threshold judgment and bounding box collision analysis; S4, using a bidirectional enhancement module to process the image and point cloud data to obtain enhanced images and point clouds; S5. Perform feature alignment on the enhanced image and point cloud.
2. The method for enhancing multimodal data of images and point clouds according to claim 1, characterized in that: The detailed steps of step S3 include: Through occlusion threshold judgment and bounding box collision analysis, whether there is an intersection between the image and the point cloud true value sample; check whether there is an intersection between the bounding boxes in the image and the point cloud true value sample in the Bev perspective. Delete the images and point cloud ground truth samples that have intersections; Guided by the depth information, samples that meet the requirements are pasted into the current frame training data to increase the number of true value samples in the frame image and point cloud.
3. The method for enhancing multimodal data of images and point clouds according to claim 1, characterized in that: The detailed steps of step S4 include: Perform a forward transformation on the image data to obtain an enhanced image; Perform a reverse transformation on the point cloud data to obtain the enhanced point cloud.
4. The method for enhancing multimodal data of images and point clouds according to claim 3, characterized in that: Detailed steps for forward transformation of image data: Randomly scale the image data to obtain scaled image data; The scaled image data is randomly flipped to obtain the enhanced image.
5. The method for enhancing multimodal data of images and point clouds according to claim 3, characterized in that: Detailed steps for reverse transformation of point cloud data: Randomly rotate the point cloud data to obtain the rotated point cloud data; Randomly translate the rotated point cloud data to obtain translated point cloud data; The translated point cloud data is randomly flipped to obtain the enhanced point cloud.
6. The method for enhancing multimodal data of images and point clouds according to claim 2, 4 or 5, characterized in that: Feature alignment includes the alignment of the LiDAR coordinate system and the camera coordinate system, the alignment transformation of the camera and the image, and the coordinate transformation of the image and the pixel.
7. The method for enhancing multimodal data of images and point clouds according to claim 6, characterized in that: Align the laser radar coordinate system and the camera coordinate system: Assume the laser radar coordinate system is SP l , the camera coordinate system is SP c , there is a point P in the current three-dimensional space, in SP l The coordinates P in l is (x l ,y l ,z l ), in SP c The coordinates P in c is (x c ,y c ,z c ), through SP l The coordinate system is transformed into SP by applying translation and rotation transformation c Camera coordinate system, three-dimensional space transformation formula of point P: P c =R*P l +T (1) Where R is the orthogonal rotation matrix, T is the three-dimensional translation vector matrix; By rotating the two coordinate systems counterclockwise around the axis at different angles, the two coordinate systems are made to coincide. The rotation is performed in the order of first rotating around the z-axis, then the y-axis, and then the x-axis. The rotation matrix R = R x *R y *R z , then the three-dimensional space transformation formula can be expressed as: Finally, the transformation process is written in the form of a homogeneous coordinate system:
8. The method for enhancing multimodal data of images and point clouds according to claim 7, characterized in that: Alignment transformation of camera and image: Based on the principle of similar triangles, let f be the focal length of the camera, and point P in the camera coordinate system c (x c ,y c ,z c ) and the projection point P in the image coordinate system img (x i ,y i ,z i )’s coordinate relationship: The matrix multiplication written in homogeneous coordinate form is:
9. The method for enhancing multimodal data of images and point clouds according to claim 8, characterized in that: Coordinate transformation of image and pixel: Let the physical size of each pixel in the image coordinate system be d x Width, d y High; the coordinate system SP established with the upper left corner of the image as the origin uv The coordinate system SP established with the imaging plane as the center img The conversion relationship is as follows: Where (uo, vo) is the image coordinate system, SP img The coordinates of the origin O in the pixel coordinate system; The matrix multiplication written in homogeneous coordinate form is: Combining equations (3), (5), and (7), we can get the transformation alignment formula from point cloud coordinates to pixel coordinates: in is the camera intrinsic parameter matrix, is the external parameter matrix.