Point cloud processing method and device, excavator, storage medium and program product

By synchronizing the time of point cloud frames and image frames in the excavator operation scene, and using image semantic segmentation technology to set dimension information on point cloud data, the problem of degraded sensor perception ability is solved, and efficient point cloud data processing and excavator control are achieved.

CN120014276APending Publication Date: 2025-05-16JIANGSU XCMG STATE KEY LAB TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510128870.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art improves the sensor perception capability of excavator, the calculation complexity is high, the efficiency is low, and the sensor perception capability is reduced in harsh environments, affecting control accuracy and robustness.

Method used

By synchronizing the time of the point cloud frame and image frame of the excavator operation scene, and using image semantic segmentation technology to set dimension information on the point cloud data based on the semantic label of the image frame, the precise alignment and effective fusion of the image and the point cloud are achieved.

Benefits of technology

It improves the sensor's perception ability, reduces the computational complexity, improves the accuracy of point cloud frames, and enhances the control accuracy and robustness of the excavator.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014276A_ABST
    Figure CN120014276A_ABST
Patent Text Reader

Abstract

The invention provides a point cloud processing method and device, an excavator, a storage medium and a computer program product, and the method comprises the steps: carrying out the time synchronization processing of a point cloud frame and an image frame corresponding to an operation scene of the excavator; performing calibration processing on the point cloud frame based on the coordinate transformation matrix; performing image semantic segmentation processing on the image frame, and determining a semantic tag corresponding to each pixel point of the image frame; projecting the point cloud frame to a two-dimensional image plane of the image frame, and determining a corresponding pixel point of each three-dimensional coordinate point of the point cloud frame in the image frame; and according to the semantic tags of the corresponding pixel points, setting dimension information of the three-dimensional coordinate points so as to control the excavator according to the point cloud frame. According to the method, the effects of dust removal and the like on the point cloud data can be efficiently achieved, the sensing ability of the sensor is improved, the calculation complexity is low, the efficiency is high, the accuracy of the collected point cloud frame is improved, and the control accuracy and robustness of the excavator are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of excavators, and in particular to a point cloud processing method, device, excavator, storage medium and computer program product. Background Art

[0002] With the advancement of artificial intelligence and Internet of Things technology, the intelligence level of excavators has been continuously improved, and intelligent excavator technology has developed rapidly in recent years. In order to accurately control the excavator, the excavator uses multi-sensor fusion technologies such as laser radar (LiDAR), visual sensor (Camera), millimeter wave radar (Radar), etc., which can use the point cloud, image and other data collected by the sensor to automatically or auxiliary control the excavator. During the working process of the excavator, the excavation operation itself will generate a large amount of dust, which will not only pollute the environment, but also affect the perception ability of the excavator's sensors. In addition, under harsh environmental conditions such as rain and snow, the perception ability of the excavator's sensors will also decrease. Due to the decrease in the perception ability of the sensors, it is difficult to maintain high accuracy and robustness for the control of the excavator; the existing methods to improve the perception ability of sensors usually rely on complex calculations or a large amount of labeled data, with high computational complexity and low efficiency. Summary of the invention

[0003] The present disclosure provides a point cloud processing method, device, excavator, storage medium and computer program product.

[0004] According to a first aspect of the present disclosure, a point cloud processing method is provided, comprising: performing time synchronization processing on point cloud frames and image frames corresponding to an excavator working scene, wherein the point cloud frames are collected by a laser radar device, and the image frames are collected by an image acquisition device; determining a coordinate conversion matrix between a first coordinate system of the laser radar device and a second coordinate system of the image acquisition device, and calibrating the point cloud frames based on the coordinate conversion matrix; performing image semantic segmentation processing on the image frames, and determining semantic labels corresponding to each pixel point of the image frame; projecting the point cloud frame onto a two-dimensional image plane of the image frame, and determining corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame; and setting dimension information of the three-dimensional coordinate points according to the semantic labels of the corresponding pixel points, so as to perform control processing of the excavator according to the point cloud frame.

[0005] Optionally, the time synchronization processing of the point cloud frame and the image frame corresponding to the excavator operation scene includes: obtaining first timestamp information corresponding to the point cloud frame and second timestamp information corresponding to the image frame; determining the minimum time difference information between the first timestamp information and the second timestamp information; and performing time synchronization processing on the point cloud frame and the image frame according to the minimum time difference information.

[0006] Optionally, the coordinate transformation matrix includes a rotation matrix and a translation matrix; determining the coordinate transformation matrix between the first coordinate system of the laser radar device and the second coordinate system of the image acquisition device, and calibrating the point cloud frame based on the coordinate transformation matrix includes: jointly calibrating the laser radar device and the image acquisition device to determine the rotation matrix and the translation matrix; based on the rotation matrix and the translation matrix, transforming the coordinates of the point cloud frame to calibrate the point cloud frame.

[0007] Optionally, performing image semantic segmentation processing on the image frame to determine the semantic labels corresponding to each pixel point of the image frame includes: inputting the image frame into a trained image semantic segmentation model to obtain a semantic segmentation processing result of the image frame, wherein the semantic segmentation processing result includes a semantic segmentation area and a corresponding semantic area label; and determining the semantic labels corresponding to each pixel point of the image frame based on the semantic segmentation area and the semantic area label.

[0008] Optionally, training samples are generated based on historical image frames, semantic segmentation regions of the historical image frames and corresponding semantic region labels; and the image semantic segmentation model is trained using the training samples to obtain a trained image semantic segmentation model.

[0009] Optionally, projecting the point cloud frame onto the two-dimensional image plane of the image frame and determining the corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame includes: determining the intrinsic parameter matrix and the extrinsic parameter matrix of the image acquisition device; performing coordinate transformation processing on the three-dimensional coordinate points according to the intrinsic parameter matrix and the extrinsic parameter matrix to obtain two-dimensional coordinate information of the three-dimensional coordinate points; and determining the corresponding pixel points of the three-dimensional coordinate points in the image frame according to the two-dimensional coordinate information.

[0010] Optionally, setting the dimension information of the three-dimensional coordinate point according to the semantic label of the corresponding pixel point includes: adding an additional dimension to the three-dimensional coordinate point; determining the corresponding pixel point of the three-dimensional coordinate point and obtaining the semantic label of the corresponding pixel point; setting the value of the additional dimension of the three-dimensional coordinate as the semantic label of the corresponding pixel point.

[0011] Optionally, before performing image semantic segmentation processing on the image frame, image filtering processing is performed on the image frame, and conditional filtering processing is performed on the point cloud frame.

[0012] Optionally, the image filtering process includes at least one of mean filtering, median filtering, and Gaussian filtering; the conditional filtering process includes at least one of distance filtering, height filtering, angle filtering, and intensity filtering.

[0013] According to a second aspect of the present disclosure, a point cloud processing device is provided, comprising: a synchronization processing module, used for performing time synchronization processing on point cloud frames and image frames corresponding to an excavator working scene, wherein the point cloud frames are collected by a laser radar device, and the image frames are collected by an image acquisition device; a calibration processing module, used for determining a coordinate conversion matrix between a first coordinate system of the laser radar device and a second coordinate system of the image acquisition device, and performing calibration processing on the point cloud frames based on the coordinate conversion matrix; a semantic processing module, used for performing image semantic segmentation processing on the image frames, and determining semantic labels corresponding to each pixel point of the image frame; a point cloud projection module, used for projecting the point cloud frame onto a two-dimensional image plane of the image frame, and determining corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame; a dimension setting module, used for setting dimension information of the three-dimensional coordinate point according to the semantic label of the corresponding pixel point, so as to perform control processing of the excavator according to the point cloud frame.

[0014] According to a third aspect of the present disclosure, a point cloud processing device is provided, comprising: a memory; and a processor coupled to the memory, wherein the processor is configured to execute the control method for the operation of the cold storage system as described above based on instructions stored in the memory.

[0015] According to a fourth aspect of the present disclosure, there is provided an excavator, comprising the point cloud processing device as described above.

[0016] According to a fifth aspect of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are executed by a processor as described above.

[0017] According to a sixth aspect of the present disclosure, a computer program product is provided, wherein the computer program product stores computer instructions, and the computer instructions are executed by a processor according to the method described above.

[0018] The point cloud processing method, device, excavator, storage medium and computer program product disclosed in the present invention realize accurate alignment of image and point cloud by performing time synchronization and space synchronization processing on point cloud frames and image frames; through image semantic segmentation technology, dimension information is set for point cloud data according to semantic labels of pixels of image frames, so as to realize effective fusion of image processing and point cloud processing, and can efficiently realize dust removal and other effects on point cloud data, thereby improving the perception ability of sensors, having low computational complexity and high efficiency, improving the accuracy of collected point cloud frames, and improving the control accuracy and robustness of the excavator. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] By describing the embodiments of the present disclosure in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present disclosure will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other purposes and advantages of the present disclosure will be further described below in conjunction with specific embodiments and with reference to the accompanying drawings. In the accompanying drawings, the same or corresponding technical features or components will be represented by the same or corresponding figure marks.

[0020] Figure 1 is a flowchart of some embodiments of the point cloud processing method according to the present disclosure;

[0021] Figure 2 is a schematic diagram of a flow chart of time synchronization processing in some embodiments of the point cloud processing method according to the present disclosure;

[0022] Figure 3 is a schematic diagram of a process of performing point cloud calibration processing in some embodiments of the point cloud processing method according to the present disclosure;

[0023] Figure 4 A schematic diagram of a process of determining a semantic label in some embodiments of the point cloud processing method according to the present disclosure;

[0024] Figure 5 is a schematic diagram of a process of setting dimension information in some embodiments of the point cloud processing method according to the present disclosure;

[0025] Figure 6 Schematic diagram of modules of some embodiments of the point cloud processing device according to the present disclosure;

[0026] Figure 7 Schematic diagrams of modules of other embodiments of the point cloud processing device according to the present disclosure;

[0027] Figure 8 Schematic diagrams of modules of some further embodiments of the point cloud processing device according to the present disclosure. DETAILED DESCRIPTION

[0028] Exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. For the sake of clarity and conciseness, not all features of the embodiments are described in the specification. However, it should be understood that many implementation-specific settings must be made in the process of implementing the embodiments in order to achieve the specific goals of the developer, for example, to meet those restrictions related to the device and the service, and these restrictions may vary depending on the implementation. In addition, it should be understood that although the development work may be very complex and time-consuming, it is only a routine task for those skilled in the art who benefit from the content of this disclosure.

[0029] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.

[0030] Those skilled in the art can understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate the necessary logical order between them.

[0031] It should also be understood that in the embodiments of the present disclosure, “plurality” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.

[0032] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.

[0033] In addition, the term "and / or" in the present disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present disclosure generally indicates that the associated objects before and after are in an "or" relationship.

[0034] It should also be understood that the description of the various embodiments in the present disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced to each other, and for the sake of brevity, they will not be described one by one.

[0035] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.

[0036] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0037] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.

[0038] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0039] In addition, in order to avoid obscuring the present disclosure due to unnecessary details, only the processing steps and / or device structures closely related to at least the scheme according to the present disclosure are shown in the drawings, and other details that are not closely related to the present disclosure are omitted. It should also be noted that similar reference numerals and letters in the drawings indicate similar items, and therefore once an item is defined in one drawing, it does not need to be discussed again for subsequent drawings.

[0040] Figure 1 is a flow chart of some embodiments of the point cloud processing method according to the present disclosure, such as Figure 1 As shown:

[0041] Step S101, performing time synchronization processing on the point cloud frames and image frames corresponding to the excavator operation scene, wherein the point cloud frames are collected by using a laser radar device, and the image frames are collected by using an image acquisition device.

[0042] A laser radar device and an image acquisition device can be installed on the excavator. The laser radar device can be a variety of laser radars, and the image acquisition device can be a variety of cameras. When the excavator is working, the laser radar device is used to scan the excavator working scene and collect point cloud frames; the image acquisition device is used to collect images of the excavator working scene and obtain image frames. By performing time synchronization processing on the point cloud frame and image frame corresponding to the excavator working scene, the timestamps of the image frame and the point cloud frame can be aligned to ensure the time consistency of the image frame and the point cloud frame.

[0043] Step S102, determining a coordinate conversion matrix between a first coordinate system of the laser radar device and a second coordinate system of the image acquisition device, and calibrating the point cloud frame based on the coordinate conversion matrix.

[0044] By determining the coordinate transformation matrix between the lidar device and the image acquisition device and calibrating the point cloud frame, spatial synchronization processing of the image frame and the point cloud frame can be achieved.

[0045] Step S103, perform image semantic segmentation on the image frame to determine the semantic label corresponding to each pixel of the image frame. The semantic label is a category label, and the semantic label can be a label such as a mining truck, a bucket, a material, or dust.

[0046] Step S104 , projecting the point cloud frame onto the two-dimensional image plane of the image frame, and determining the corresponding pixel point in the image frame for each three-dimensional coordinate point of the point cloud frame.

[0047] Step S105, according to the semantic label of the corresponding pixel point, the dimension information of the three-dimensional coordinate point is set to control the excavator according to the point cloud frame. A variety of existing methods can be used to control the excavator according to the point cloud frame, for example, to assist the excavator in making a decision on the control operation according to the point cloud frame.

[0048] The point cloud processing method disclosed in the present invention realizes precise alignment of images and point clouds by performing time synchronization and space synchronization processing on point cloud frames and image frames; through image semantic segmentation technology, dimension information is set for point cloud data according to semantic labels of pixels of image frames, thereby realizing effective fusion of image processing and point cloud processing, and being able to efficiently realize dust removal and other effects on point cloud data, thereby improving the perception capability of sensors, solving the problem that existing methods for improving the perception capability of sensors are highly complex and computationally intensive, improving the accuracy of collected point cloud frames, and improving the control accuracy and robustness of excavators.

[0049] Figure 2 FIG. 1 is a flow chart of time synchronization processing in some embodiments of the point cloud processing method according to the present disclosure, such as Figure 2 As shown:

[0050] Step S201, obtaining first timestamp information corresponding to the point cloud frame and second timestamp information corresponding to the image frame.

[0051] Step S202: determine the minimum time difference information between the first timestamp information and the second timestamp information.

[0052] Step S203: performing time synchronization processing on the point cloud frame and the image frame according to the minimum time difference information.

[0053] The first timestamp information of the point cloud frame can be used as a reference to achieve the timestamp alignment between the image frame and the point cloud frame through time soft synchronization to perform time synchronization processing. A variety of methods can be used to calculate the difference between the first timestamp information corresponding to the point cloud frame and the second timestamp information corresponding to the image frame, determine the point cloud frame and image with the smallest time difference, and adjust the second timestamp of the camera to the first timestamp of the laser radar, so as to achieve time alignment, that is, adjust the timestamp of the camera to the moment closest to the timestamp of the laser radar, which can ensure that the image frame collected by the camera is closest to the time of the point cloud frame collected by the laser radar.

[0054] For example, the first timestamp information of the four point cloud frames collected by the laser radar device is (in milliseconds): [100, 200, 300, 400]; the second timestamp information of the four image frames collected by the image acquisition device is (in milliseconds): [95, 198, 297, 396]. The time difference is calculated for each corresponding point cloud frame and image frame to obtain a time difference list: time difference = |first timestamp information-second timestamp information|; the obtained time difference list is: [5, 2, 3, 4]; in the time difference list, the minimum time difference information is determined to be 2, corresponding to the second point cloud frame and image frame; the second point cloud frame is selected as the synchronization benchmark, that is, the camera timestamp 198ms is closest to the laser radar timestamp 200ms.

[0055] A variety of existing methods can be used to adjust the second timestamp of the camera to the first timestamp of the lidar, and the timestamp of the camera can be adjusted to align with the timestamp of the lidar. For example, this can be achieved by interpolation or direct adjustment of the timestamp. For example, the lidar timestamp is 200ms and the camera timestamp is 198ms, and the camera timestamp is adjusted to 200ms. The time difference is calculated as 200 - 198 = 2ms, and the camera timestamp is adjusted to increase the camera timestamp by 2ms to make it consistent with the lidar timestamp. The adjusted camera timestamp = original camera timestamp + time difference, that is, the adjusted camera timestamp = 198 + 2 = 200ms.

[0056] Figure 3 FIG. 1 is a flow chart of performing point cloud calibration processing in some embodiments of the point cloud processing method disclosed herein. Figure 3 As shown:

[0057] Step S301, jointly calibrate the laser radar device and the image acquisition device to determine the rotation matrix and the translation matrix.

[0058] Joint calibration processing refers to the accurate calibration of the relative position and attitude between different sensors. Joint calibration of LiDAR and camera refers to solving the coordinate transformation matrix of the camera relative to the LiDAR to achieve accurate spatial alignment.

[0059] Step S302 , based on the rotation matrix and the translation matrix, the coordinates of the point cloud frame are transformed to calibrate the point cloud frame.

[0060] There are many methods that can be used to calibrate the point cloud frame, i.e., spatial synchronization. The existing external parameter calibration toolkit under the ROS (Robot Operating System) system can be used to jointly calibrate the lidar and camera to solve the coordinate transformation matrix (including rotation matrix and translation matrix) of the camera compared to the lidar. There are many existing methods that can be used to transform the coordinates of the point cloud frame based on the rotation matrix and translation matrix to calibrate the point cloud frame.

[0061] For example, software A is an open source autonomous driving software based on the ROS system. You can use software A to complete the joint calibration of the lidar and camera. Software A can be Autoware, etc. Use a checkerboard calibration plate to ensure that it is clearly visible in the field of view of the lidar and camera. The size of the calibration plate is known and remains fixed during calibration. Start the lidar and camera and ensure that their timestamps are synchronized. You can use the ROS time synchronization mechanism to ensure that the timestamps of the two are consistent.

[0062] Use the rosbag tool of ROS to record the point cloud data collected by the lidar and the image data collected by the camera. Move the calibration board multiple times at different positions and angles to ensure that enough scene changes are covered. Start software A and load the corresponding configuration file. Use the calibration tool of software A to load the recorded rosbag file into software A. In the interface of software A, use the joint calibration tool.

[0063] When calibrating, import the calibration board parameters and enter the physical size of the calibration board (such as the side length of each square on the calibration board). Detect the calibration board, and use the joint calibration tool of software A to automatically detect the checkerboard calibration board in the image and extract the corner point information. For the lidar point cloud, the joint calibration tool is used to fit the position and posture of the checkerboard according to the point cloud data. If the automatic detection result is not ideal, you can manually adjust the position and posture of the calibration board to ensure accurate detection. Generate the calibration result. Software A calculates the coordinate transformation matrix of the camera relative to the lidar (including the rotation matrix and the translation matrix). The output coordinate transformation matrix is ​​usually a 4x4 transformation matrix T, which contains the rotation matrix R and the translation vector T.

[0064] The calibration results can be verified by projecting the LiDAR point cloud onto the camera image using the calibration results to check whether the point cloud is aligned with the image features. If the point cloud and image features can be aligned, the calibration is successful; otherwise, recalibration or parameter adjustment is required. Use the coordinate transformation matrix T obtained by calibration to transform the LiDAR point cloud frame (the point cloud frame in the first coordinate system) to the camera coordinate system (the second coordinate system): P_camera = T * P_lidar, where P_lidar is the coordinate of the three-dimensional coordinate point of the point cloud frame, and P_camera is the coordinate of the three-dimensional coordinate point of the transformed point cloud frame; the coordinates of the point cloud frame are transformed by the rotation matrix and the translation matrix to calibrate the point cloud frame, and the obtained P_camera is the coordinate after the transformation of the coordinates P_lidar of the point cloud frame.

[0065] Figure 4 FIG. 1 is a flow chart of determining semantic labels in some embodiments of the point cloud processing method according to the present disclosure, such as Figure 4 As shown:

[0066] Step S401: input the image frame into the trained image semantic segmentation model to obtain the semantic segmentation processing result of the image frame, wherein the semantic segmentation processing result includes the semantic segmentation area and the corresponding semantic area label.

[0067] Semantic segmentation is to classify each pixel or point of the image, that is, to divide the image into different regions and assign a semantic region label (category label) to each region.

[0068] Step S402: Determine a semantic label corresponding to each pixel of the image frame according to the semantic segmentation region and the semantic region label.

[0069] For example, input image frame B into the trained image semantic segmentation model to obtain the semantic segmentation processing result of image frame B, which includes semantic segmentation region A and semantic segmentation region B, as well as semantic region label C corresponding to semantic segmentation region A and semantic region label D corresponding to semantic segmentation region B; semantic region label C is a forklift label, and semantic region label D is a dust label. The semantic label of each pixel in the semantic segmentation region A of image frame B is determined to be semantic region label C, and the semantic label of each pixel in the semantic segmentation region B of image frame B is determined to be semantic region label D.

[0070] The image semantic segmentation model can be a variety of deep learning models, such as FCN (Fully Convolutional Network) model, U-Net model, R-CNN (Regions with CNN features), STDC (Short-Term Dense Concatenate) model, Transformer model, etc. The image semantic segmentation model can be used to perform semantic segmentation on image frames, identify various objects in image frames, and assign semantic labels.

[0071] The image semantic segmentation model can be pre-trained, and training samples can be generated based on historical image frames, semantic segmentation regions of historical image frames, and corresponding semantic region labels; the image semantic segmentation model is trained using the training samples to obtain a trained image semantic segmentation model. The image semantic segmentation model can be trained using a variety of existing model training methods.

[0072] For example, the image semantic segmentation model is the STDC model; collect historical image data, collect image datasets containing different categories of excavation scenes, and the image datasets include multiple historical image frames. Use existing annotation tools to annotate the historical image frames pixel by pixel, annotate each pixel with a category label (i.e., a semantic label), determine the semantic segmentation area of ​​the historical image frame and the corresponding semantic area label, and generate training samples.

[0073] Construct the STDC model, which consists of multiple short-term dense connection modules, which can improve the feature extraction capability and reduce the amount of calculation. The input layer of the STDC model receives the input image (for example, an RGB image of size 1280x720); the feature extraction layer of the STDC model includes multiple convolutional layers and short-term dense connection modules, which are used to gradually extract multi-scale features; the upsampling layer of the STDC model is used to upsample the low-resolution feature map to the original image size and output the pixel-by-pixel classification result.

[0074] When training the model, you can preprocess the training samples and normalize the input images so that their pixel values ​​are in the range of [0, 1]. You can use data augmentation techniques (such as random cropping, flipping, rotation, etc.) to expand the data set and improve the generalization ability of the model. The loss function of the STDC model includes cross-entropy loss or weighted cross-entropy loss. The weighted cross-entropy loss can handle the problem of class imbalance.

[0075] Use pre-trained weights or randomly initialize model parameters; iterate the training data set and divide the training samples into training set, validation set, and test set. Iterate the training set, taking a batch of data each time for forward propagation and back propagation; input images through the model for forward propagation to obtain prediction results; calculate the loss function and compare it with the true label; update the model parameters through back propagation. Verify the model performance. At the end of each iteration, use the validation set to evaluate the model performance, record the loss and accuracy, and save the model weights that perform best on the validation set during training.

[0076] When using the trained image semantic segmentation model for semantic segmentation processing, load the trained STDC model, input the image to be segmented, and generate the final semantic segmentation map through the STDC model. By inputting the image frame into the trained image semantic segmentation model, the pixel-by-pixel classification result, that is, the semantic segmentation processing result of the image frame, is obtained. The semantic segmentation processing result includes the semantic segmentation area and the corresponding semantic area label. The semantic area label includes labels such as mining truck, bucket, material, and dust. The semantic area label can be mapped to a color map and superimposed on the original image for visual display.

[0077] In some embodiments, various methods can be used to determine the corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame. For example, the internal parameter matrix and the external parameter matrix of the image acquisition device are determined, and the three-dimensional coordinate points are subjected to coordinate conversion processing according to the internal parameter matrix and the external parameter matrix to obtain the two-dimensional coordinate information of the three-dimensional coordinate points; and the corresponding pixel points of the three-dimensional coordinate points in the image frame are determined according to the two-dimensional coordinate information.

[0078] Using the camera's intrinsic and extrinsic matrix, each 3D coordinate point of the point cloud frame can be projected onto the 2D image plane of the image frame and the corresponding pixel coordinates can be obtained. For example, when projecting a 3D coordinate point (X, Y, Z) of the point cloud frame onto the 2D image plane of the image frame, a transformation matrix is ​​constructed: the camera intrinsic matrix K and the extrinsic matrix [R|T] are combined into a complete projection matrix P: P = K * [R|T]; the 3D coordinate point (X, Y, Z) is converted to homogeneous coordinates (X, Y, Z, 1), and then multiplied by the projection matrix P: p = P * [X, Y, Z, 1]^T to obtain the 2D coordinate information of the 3D coordinate point (X, Y, Z).

[0079] The obtained two-dimensional coordinate information can be divided by its fourth component (depth value) to obtain the pixel coordinates (u, v) on the 2D image plane: u = p[0] / p[2]; v = p[1] / p[2]; after the three-dimensional coordinate point is transformed according to the intrinsic parameter matrix and the extrinsic parameter matrix to obtain the two-dimensional coordinate information (u, v) of the three-dimensional coordinate point, the corresponding pixel point of the three-dimensional coordinate point in the image frame can be determined according to the two-dimensional coordinate information (u, v).

[0080] Figure 5 FIG. 1 is a flow chart of setting dimension information in some embodiments of the point cloud processing method according to the present disclosure, such as Figure 5 As shown:

[0081] Step S501, adding an additional dimension to the three-dimensional coordinate point.

[0082] Step S502, determining the corresponding pixel point of the three-dimensional coordinate point, and obtaining the semantic label of the corresponding pixel point.

[0083] Step S503: setting the value of the additional dimension of the three-dimensional coordinate as the semantic label of the corresponding pixel point.

[0084] The point cloud frame can be projected onto the two-dimensional image plane of the image frame according to the intrinsic parameter matrix and extrinsic parameter matrix of the camera. For each three-dimensional coordinate point of the projected point cloud frame, the corresponding pixel position of each three-dimensional coordinate point of the point cloud frame in the image frame can be determined according to its two-dimensional coordinate information, and the semantic segmentation results (semantic labels, etc.) at the corresponding pixel positions can be assigned to each three-dimensional coordinate point of the point cloud frame; the semantic labels of each three-dimensional coordinate point of the point cloud frame can be added as an additional dimension.

[0085] For example, a three-dimensional coordinate point (X, Y, Z) of a point cloud frame is projected onto the two-dimensional image plane of an image frame to obtain the two-dimensional coordinate information (u, v) of the three-dimensional coordinate point; the corresponding pixel position on the image frame is found based on the projected pixel coordinates (u, v), the semantic label (category label) of the pixel position is obtained on the image, and the semantic label is added as an additional dimension of the three-dimensional coordinate point (X, Y, Z) of the point cloud frame.

[0086] In some embodiments, before image semantic segmentation is performed on the image frame, image filtering can be performed on the image frame and conditional filtering can be performed on the point cloud frame. Image filtering can remove noise from the image data and improve image contrast through grayscale equalization; conditional filtering can be performed on the point cloud frame to perform a preliminary screening of the point cloud data and remove three-dimensional coordinate points that do not meet specific rules (distance, height, angle, intensity, etc.) through conditional filtering to align with the camera field of view.

[0087] Image filtering includes at least one of mean filtering, median filtering, and Gaussian filtering. A variety of mean filtering, median filtering, and Gaussian filtering methods can be used. Mean filtering is to use a sliding window to calculate the average value of pixels in the window and replace the center pixel value; median filtering is to use a sliding window to calculate the median of pixels in the window and replace the center pixel value; Gaussian filtering is to use a Gaussian kernel to perform a convolution operation to smooth the image. Grayscale equalization is to enhance the image contrast through histogram equalization.

[0088] The conditional filtering process includes at least one of distance filtering, height filtering, angle filtering, and intensity filtering. The conditional filtering process is to remove the three-dimensional coordinate points of the point cloud frame that do not meet the requirements according to specific rules (such as distance, height, angle, intensity, etc.) to improve the quality of the point cloud data.

[0089] Distance filtering: In some application scenarios, it is necessary to focus on point cloud data within a certain range. For example, in the application of intelligent hydraulic excavators, only the point cloud data near the excavation area needs to be focused. When performing distance filtering, a maximum distance distance_threshold is set, such as 5 meters; the point cloud data of the point cloud frame is [(1, 2, 3), (4, 5, 6), (7, 8, 9)], and the distance threshold is 5 meters. Calculate the distance of each point: the distance of point (1, 2, 3) is sqrt(1^2 + 2^2 + 3^2) ≈ 3.74 meters, which meets the condition; the distance of point (4, 5, 6) is sqrt(4^2 + 5^2 + 6^2) ≈ 8.77 meters, which does not meet the condition; the distance of point (7, 8, 9) is sqrt(7^2 + 8^2 + 9^2) ≈ 13.93 meters, which does not meet the condition; through distance filtering, finally retain point (1, 2, 3).

[0090] Height filtering: In some application scenarios, point cloud data within a certain height range is required. For example, in the application of intelligent hydraulic excavators, only point cloud data near the ground is needed. When performing height filtering, set a minimum height min_height and a maximum height max_height, such as 0 meters to 3 meters; retain points whose Z coordinates are within the height threshold range: the point cloud data is [(1, 2, 3), (4, 5, 6), (7, 8, 9)], and the height threshold is 0 meters to 3 meters; the height of point (1, 2, 3) is 3 meters, which meets the condition; the height of point (4, 5, 6) is 6 meters, which does not meet the condition; the height of point (7, 8, 9) is 9 meters, which does not meet the condition; after height filtering, point (1, 2, 3) is finally retained.

[0091] Angle filtering: In some application scenarios, point cloud data within a specific viewing angle range is required. For example, in the application of intelligent hydraulic excavators, only the point cloud data within the front viewing angle is required. When performing angle filtering, set a viewing angle range, such as ±45 degrees in front. Calculate the angle: For each point (X, Y, Z), calculate its angle relative to the reference direction (such as the front). The vector dot product formula can be used to calculate the angle; reference_vector is the reference direction vector (such as [1,0, 0] means the front), point_vector is the direction vector of the point cloud point [X, Y, Z]; the point cloud data is [(1,2, 3), (4, 5, 6), (7, 8, 9)], the angle threshold is ±45 degrees, calculate the angle of each point: the angle of point (1, 2, 3) is arccos((1*1 + 2*0 + 3*0) / sqrt(1^2 + 2^2 + 3^2)) ≈ 0.955 radians (about 54.7 degrees), which does not meet the conditions; the angle of point (4, 5, 6) is arccos((4*1 + 5*0 + 6*0) / sqrt(4^2+ 5^2 + 6^2)) ≈ 0.723 radians (about 41.4 degrees), which meets the requirements; the angle of point (7, 8, 9) is arccos((7*1 + 8*0 + 9*0) / sqrt(7^2 + 8^2 + 9^2)) ≈ 0.675 radians (about 38.7 degrees), which meets the requirements; through angle filtering, points (4, 5, 6) and (7, 8, 9) are finally retained.

[0092] Intensity filtering: In some application scenarios, point cloud data with a certain reflection intensity is required. For example, in the application of intelligent hydraulic excavators, objects of different materials need to be distinguished. When performing intensity filtering, set a minimum intensity min_intensity and a maximum intensity max_intensity, such as 10 to 100. Retain points with intensities within the threshold range: The point cloud data is [(1, 2, 3, 50), (4, 5, 6, 150), (7, 8, 9, 80)], and the intensity threshold is 10 to 100. The intensity of point (1, 2, 3, 50) is 50, which meets the condition; the intensity of point (4, 5, 6, 150) is 150, which does not meet the condition; the intensity of point (7, 8, 9, 80) is 80, which meets the condition; through intensity filtering, points (1, 2, 3, 50) and (7, 8, 9, 80) are finally retained.

[0093] The point cloud processing method in the above embodiment achieves precise alignment of image and point cloud data by using time synchronization and space synchronization technology, and sets dimension information for point cloud data according to the semantic labels of pixels in the image frame through image semantic segmentation technology, thereby achieving effective fusion of image processing and point cloud processing, and being able to efficiently achieve dust removal and other effects on point cloud data, thereby reducing computational complexity, improving computational efficiency, and improving sensor perception capabilities, solving the problem that existing methods for improving sensor perception capabilities are complex and computationally intensive, improving the accuracy of collected point cloud frames, and improving the control accuracy and robustness of excavators; it can be extended to other technical fields such as self-driving cars, drones, and industrial robots.

[0094] In some embodiments, Figure 6 As shown, the present disclosure provides a point cloud processing device 60, including a synchronization processing module 61, a calibration processing module 62, a semantic processing module 63, a point cloud projection module 64 and a dimension setting module 65. The synchronization processing module 61 performs time synchronization processing on the point cloud frames and image frames corresponding to the excavator operation scene, wherein the point cloud frames are collected by the laser radar device, and the image frames are collected by the image acquisition device.

[0095] The calibration processing module 62 determines the coordinate conversion matrix between the first coordinate system of the laser radar device and the second coordinate system of the image acquisition device, and performs calibration processing on the point cloud frame based on the coordinate conversion matrix. The semantic processing module 63 performs image semantic segmentation processing on the image frame to determine the semantic label corresponding to each pixel point of the image frame.

[0096] The point cloud projection module 64 projects the point cloud frame onto the two-dimensional image plane of the image frame to determine the corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame. The dimension setting module 65 sets the dimension information of the three-dimensional coordinate point according to the semantic label of the corresponding pixel point, so as to control the excavator according to the point cloud frame.

[0097] In some embodiments, the synchronization processing module 61 obtains first timestamp information corresponding to the point cloud frame and second timestamp information corresponding to the image frame; the synchronization processing module 61 determines the minimum time difference information between the first timestamp information and the second timestamp information; the synchronization processing module 61 performs time synchronization processing on the point cloud frame and the image frame according to the minimum time difference information.

[0098] The coordinate transformation matrix includes a rotation matrix and a translation matrix; the calibration processing module 62 performs joint calibration processing on the laser radar device and the image acquisition device to determine the rotation matrix and the translation matrix; the calibration processing module 62 transforms the coordinates of the point cloud frame based on the rotation matrix and the translation matrix to calibrate the point cloud frame.

[0099] The semantic processing module 63 inputs the image frame into the trained image semantic segmentation model to obtain the semantic segmentation processing result of the image frame, wherein the semantic segmentation processing result includes the semantic segmentation area and the corresponding semantic area label; the semantic processing module 63 determines the semantic label corresponding to each pixel point of the image frame according to the semantic segmentation area and the semantic area label.

[0100] The point cloud projection module 64 determines the intrinsic parameter matrix and the extrinsic parameter matrix of the image acquisition device; performs coordinate transformation processing on the three-dimensional coordinate points according to the intrinsic parameter matrix and the extrinsic parameter matrix to obtain the two-dimensional coordinate information of the three-dimensional coordinate points; the point cloud projection module 64 determines the corresponding pixel points of the three-dimensional coordinate points in the image frame according to the two-dimensional coordinate information.

[0101] The dimension setting module 65 adds an additional dimension to the three-dimensional coordinate point; determines the corresponding pixel point of the three-dimensional coordinate point and obtains the semantic label of the corresponding pixel point; and sets the value of the additional dimension of the three-dimensional coordinate as the semantic label of the corresponding pixel point.

[0102] In some embodiments, Figure 7 As shown, the present disclosure provides a point cloud processing device 60 ′, which includes not only all the modules of the point cloud processing device 60 , but also a model training module 66 and a filtering processing module 67 .

[0103] The model training module 66 is used to generate training samples based on the historical image frames, the semantic segmentation regions of the historical image frames and the corresponding semantic region labels; the model training module 66 uses the training samples to train the image semantic segmentation model to obtain a trained image semantic segmentation model. The filtering processing module 67 performs image filtering processing on the image frames and conditional filtering processing on the point cloud frames before performing image semantic segmentation processing on the image frames.

[0104] In some embodiments, Figure 8 As shown, the present disclosure provides a point cloud processing device, which may include a memory 82, a processor 81, a communication interface 83, and a bus 84. The memory 82 is used to store instructions, the processor 81 is coupled to the memory 82, and the processor 81 is configured to execute the above-mentioned point cloud processing method based on the instructions stored in the memory 82.

[0105] The memory 82 may be a high-speed RAM memory, a non-volatile memory, etc., or a memory array. The memory 82 may also be divided into blocks, and the blocks may be combined into virtual volumes according to certain rules. The processor 81 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the point cloud processing method of the present disclosure.

[0106] In some embodiments, the present disclosure provides an excavator, including a point cloud processing device as in any of the above embodiments. The excavator may be a variety of excavators, such as an intelligent hydraulic excavator.

[0107] In some embodiments, the present disclosure provides a computer-readable storage medium storing computer instructions, which implement the methods in any of the above embodiments when executed by a processor.

[0108] Computer readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive enumeration) of readable storage media can include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0109] The embodiments of the present disclosure may also be a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, the processor executes the steps of the method according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.

[0110] The point cloud processing method, device, excavator, storage medium and computer program product in the above-mentioned embodiments realize precise alignment of the image and the point cloud by performing time synchronization and space synchronization processing on the point cloud frame and the image frame; through the image semantic segmentation technology, the dimension information is set for the point cloud data according to the semantic labels of the pixels of the image frame, so as to realize the effective fusion of image processing and point cloud processing, and can efficiently realize dust removal and other effects on the point cloud data, thereby improving the perception ability of the sensor, solving the problem that the existing methods for improving the perception ability of the sensor are complex and computationally intensive, improving the accuracy of the collected point cloud frames, and improving the control accuracy and robustness of the excavator, thereby improving the user experience.

[0111] The basic principles of the present disclosure are described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. are required by each embodiment of the present disclosure. In addition, the specific details disclosed above are only for the purpose of illustration and ease of understanding, and are not limitations. The above details do not limit the present disclosure to the necessity of adopting the above specific details to be implemented.

[0112] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0113] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including," "comprising," "having," and the like are open words, referring to "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or," and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.

[0114] It should also be noted that in the apparatus, device and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.

[0115] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

[0116] The above description has been given for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, it will be appreciated by those skilled in the art that the above embodiments are merely illustrative and do not limit the scope of the present disclosure. It will be appreciated by those skilled in the art that the above embodiments may be combined, modified or replaced without departing from the scope and essence of the present disclosure.

Claims

1. A point cloud processing method, comprising: Performing time synchronization processing on the point cloud frames and image frames corresponding to the excavator operation scene, wherein the point cloud frames are collected by using a laser radar device, and the image frames are collected by using an image acquisition device; Determine a coordinate conversion matrix between a first coordinate system of the laser radar device and a second coordinate system of the image acquisition device, and perform calibration processing on the point cloud frame based on the coordinate conversion matrix; Performing image semantic segmentation processing on the image frame to determine a semantic label corresponding to each pixel point of the image frame; Projecting the point cloud frame onto the two-dimensional image plane of the image frame, and determining corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame; According to the semantic label of the corresponding pixel point, the dimension information of the three-dimensional coordinate point is set to perform control processing of the excavator according to the point cloud frame.

2. The method of claim 1, wherein: The time synchronization processing of the point cloud frame and the image frame corresponding to the excavator operation scene includes: Acquire first timestamp information corresponding to the point cloud frame and second timestamp information corresponding to the image frame; Determine minimum time difference information between the first timestamp information and the second timestamp information; According to the minimum time difference information, time synchronization processing is performed on the point cloud frame and the image frame.

3. The method of claim 1, wherein: The coordinate transformation matrix includes a rotation matrix and a translation matrix; determining the coordinate transformation matrix between the first coordinate system of the laser radar device and the second coordinate system of the image acquisition device, and calibrating the point cloud frame based on the coordinate transformation matrix includes: Performing joint calibration processing on the laser radar device and the image acquisition device to determine the rotation matrix and the translation matrix; Based on the rotation matrix and the translation matrix, the coordinates of the point cloud frame are transformed to calibrate the point cloud frame.

4. The method of claim 1, wherein: The performing image semantic segmentation processing on the image frame to determine the semantic label corresponding to each pixel point of the image frame includes: Inputting the image frame into a trained image semantic segmentation model to obtain a semantic segmentation processing result of the image frame, wherein the semantic segmentation processing result includes a semantic segmentation region and a corresponding semantic region label; According to the semantic segmentation area and the semantic area label, a semantic label corresponding to each pixel point of the image frame is determined.

5. The method of claim 4, comprising: Generate training samples according to historical image frames, semantic segmentation regions of the historical image frames, and corresponding semantic region labels; The image semantic segmentation model is trained using the training samples to obtain a trained image semantic segmentation model.

6. The method of claim 1, wherein: The step of projecting the point cloud frame onto the two-dimensional image plane of the image frame and determining corresponding pixel points of each three-dimensional coordinate point of the point cloud frame in the image frame comprises: Determining an intrinsic parameter matrix and an extrinsic parameter matrix of the image acquisition device; According to the internal parameter matrix and the external parameter matrix, coordinate conversion processing is performed on the three-dimensional coordinate point to obtain two-dimensional coordinate information of the three-dimensional coordinate point; The corresponding pixel point of the three-dimensional coordinate point in the image frame is determined according to the two-dimensional coordinate information.

7. The method of claim 1, wherein: The step of setting the dimension information of the three-dimensional coordinate point according to the semantic label of the corresponding pixel point comprises: Adding an additional dimension to the three-dimensional coordinate point; Determine the corresponding pixel point of the three-dimensional coordinate point, and obtain the semantic label of the corresponding pixel point; The value of the additional dimension of the three-dimensional coordinate is set as the semantic label of the corresponding pixel point.

8. The method according to any one of claims 1 to 7, comprising: Before performing image semantic segmentation processing on the image frame, performing image filtering processing on the image frame, and performing conditional filtering processing on the point cloud frame.

9. The method of claim 8, wherein: The image filtering process includes at least one of mean filtering, median filtering, and Gaussian filtering; The conditional filtering process includes at least one of distance filtering, height filtering, angle filtering, and intensity filtering.

10. A point cloud processing device, comprising: A synchronization processing module, used for performing time synchronization processing on the point cloud frame and the image frame corresponding to the excavator operation scene, wherein the point cloud frame is collected by a laser radar device, and the image frame is collected by an image acquisition device; A calibration processing module, used to determine a coordinate conversion matrix between a first coordinate system of the laser radar device and a second coordinate system of the image acquisition device, and to perform calibration processing on the point cloud frame based on the coordinate conversion matrix; A semantic processing module, used to perform image semantic segmentation processing on the image frame to determine a semantic label corresponding to each pixel point of the image frame; A point cloud projection module, used to project the point cloud frame onto the two-dimensional image plane of the image frame, and determine the corresponding pixel point of each three-dimensional coordinate point of the point cloud frame in the image frame; A dimension setting module is used to set the dimension information of the three-dimensional coordinate point according to the semantic label of the corresponding pixel point, so as to perform control processing of the excavator according to the point cloud frame.

11. A point cloud processing device, comprising: Memory; and a processor coupled to the memory, wherein the processor is configured to execute the control method for the operation of the cold storage system according to any one of claims 1 to 9 based on instructions stored in the memory.

12. An excavator, comprising: A point cloud processing device as claimed in claim 10 or 11.

13. A computer-readable storage medium storing computer instructions, wherein the computer instructions are executed by a processor to perform the method according to any one of claims 1 to 9.

14. A computer program product, wherein the computer program product stores computer instructions, wherein the computer instructions are executed by a processor to perform the method according to any one of claims 1 to 9.