Point cloud data completion processing method, self-moving device and storage medium
By acquiring target segmentation images and target radar point cloud data, and performing grid division and plane fitting processing, the missing location points are filled in using visual information. This solves the shortcomings of pure LiDAR and pure vision solutions, and improves the accuracy and completeness of drivable area identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN MAMMOTION INNOVATION CO LTD
- Filing Date
- 2026-01-12
- Publication Date
- 2026-04-17
AI Technical Summary
Pure LiDAR solutions have insufficient point cloud density beyond a specified distance, making it easy to miss detections in drivable areas. Pure vision solutions are easily affected by lighting conditions, leading to errors in drivable area identification. How to balance the completeness and accuracy of drivable area identification has become an urgent problem to be solved.
By acquiring target segmentation images and target radar point cloud data, target point cloud projection maps are obtained, and grid division is performed. Missing location points are filled in based on plane fitting processing and visual information. Visual effective information is used to fill in the point cloud failure location points in the drivable area, and corrections are made for local terrain differences to reduce the impact of lighting environment.
It improves the completeness and accuracy of drivable area identification, enhances the accuracy of drivable area completion in complex outdoor environments, and improves the completeness and accuracy of point cloud data.
Smart Images

Figure CN121883791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, specifically to a method for completing point cloud data, a self-moving device, and a computer-readable storage medium. Background Technology
[0002] With the rapid development of intelligent and digital technologies, intelligent technologies are being applied to an increasing number of fields. For example, intelligent driving is being used in intelligent robots (such as intelligent lawnmowers, cleaning robots, and delivery robots) and intelligent vehicles. Typically, intelligent driving relies on environmental perception based on technologies such as LiDAR and cameras. For instance, an outdoor intelligent lawnmower needs to use LiDAR to perceive the environment and generate a drivable area to ensure operational effectiveness. However, pure LiDAR solutions rely on dense point clouds, while sparse LiDAR (e.g., 16 lines or less) has insufficient point cloud density beyond a specified distance (e.g., 2 meters), leading to missed detections of drivable areas. Pure vision solutions are easily affected by lighting conditions (e.g., cloudy days, shadows), resulting in errors in drivable area identification. Therefore, how to balance the completeness and accuracy of drivable area identification has become an urgent problem to be solved. Summary of the Invention
[0003] This application provides a point cloud data completion processing method, a self-moving device, and a computer-readable storage medium, which can improve the completeness and accuracy of drivable area identification.
[0004] Firstly, this application provides a method for processing point cloud data completion, the method comprising: Acquire the target segmentation image and the target radar point cloud data in the image coordinate system as a projection map of the target point cloud; The target camera image corresponding to the segmented target image is divided into multiple grid regions in the field of view of three-dimensional space. Based on the target segmentation image and the target point cloud projection map, the target radar point cloud data is allocated to the corresponding grid regions; Based on the allocated point cloud data of each grid region, a plane fitting process is performed on each grid region to obtain the planar features of each grid region; Based on the drivable area segmented from the target segmentation image and the missing point cloud locations in the target point cloud projection image, the missing point locations in the target radar point cloud data are determined. Based on the planar features, point cloud completion processing is performed on the points to be completed.
[0005] Secondly, this application also provides a self-moving device, which includes a processing unit and a memory. The memory stores a computer program, and when the processing unit calls the computer program in the memory, it executes any of the point cloud data completion processing methods provided in this application.
[0006] Thirdly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processing unit to execute the point cloud data completion processing method.
[0007] In this application, firstly, by determining the missing point cloud locations in the target radar point cloud data based on the segmented drivable area in the target segmentation image and the missing point cloud locations in the target point cloud projection image, the missing point cloud locations in the drivable area can be filled in using effective visual information, thereby improving the integrity of the point cloud and thus improving the completeness of the drivable area identification. Secondly, since planar fitting processing of each grid region can correct for local terrain differences, the planar features of each grid region are obtained by using the allocated point cloud data of each grid region to perform planar fitting processing on each grid region. These features are then used to fill in the missing point cloud locations (i.e., the missing point cloud locations in the drivable area). This allows for correction and filling in of local terrain differences, ensuring that the filled point cloud reflects the actual ground height, thereby improving the accuracy of point cloud filling, improving the reconstruction accuracy of the drivable area, and thus improving the identification accuracy of the drivable area. Thirdly, by performing point completion processing based on planar features, rather than directly using visual recognition results for point cloud completion, the problem of errors in drivable area recognition caused by the susceptibility of visual information to lighting conditions (such as cloudy days or shadows) is reduced. This improves the accuracy of drivable area completion in complex outdoor environments, enhances the reconstruction accuracy of drivable areas, and consequently improves the overall accuracy of drivable area recognition. Therefore, the completeness and accuracy of drivable area recognition can be improved. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic block diagram of the structure of a self-moving device provided in an embodiment of this application; Figure 2 This is a flowchart illustrating a point cloud data completion processing method provided in an embodiment of this application; Figure 3 This is an illustrative diagram illustrating the conversion of target radar point cloud data into a target point cloud projection map in an embodiment of this application; Figure 4 This is a schematic diagram illustrating the division of the field of view region provided in the embodiments of this application; Figure 5 This is another schematic diagram illustrating the division of the field of view region provided in the embodiments of this application; Figure 6 This is another schematic diagram illustrating the division of the field of view region provided in the embodiments of this application; Figure 7 This is an illustrative diagram illustrating the allocated point cloud data for determining a grid region in an embodiment of this application; Figure 8 This is an illustrative diagram illustrating the point cloud allocation grid region provided in the embodiments of this application; Figure 9 This is an illustrative diagram illustrating the position points to be filled in provided in the embodiments of this application. Detailed Implementation
[0010] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0011] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0012] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0013] To enable any person skilled in the art to implement and use this application, the following description is provided. In this description, details are set forth for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be implemented without using these specific details. In other instances, well-known processes will not be described in detail to avoid obscuring the description of the embodiments of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in the embodiments of this application.
[0014] This application provides a method for completing point cloud data, a self-moving device, and a computer-readable storage medium.
[0015] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0016] Figure 1 This is a schematic block diagram of the structure of a self-moving device provided in an embodiment of this application.
[0017] like Figure 1 As shown, the self-moving device 100 includes a processing unit 101, a memory 102, a camera 103, and a lidar 104. The processing unit 101, memory 102, camera 103, and lidar 104 are connected via a bus, such as an I2C (Inter-integrated Circuit) bus.
[0018] Specifically, the processing unit 101 provides computing and control capabilities to support the operation of the entire self-moving device 100. The processing unit 101 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0019] Specifically, the memory 102 can be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a portable hard drive, etc.
[0020] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the embodiments of this application, and does not constitute a limitation on the self-moving device to which the embodiments of this application are applied. A specific self-moving device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0021] The processing unit 101 is configured to run a computer program stored in the memory 102, and implement any of the point cloud data completion processing methods provided in this application embodiment when executing the computer program. For example, the processing unit 101 is configured to run a computer program stored in the memory 102, and can implement the following steps when executing the computer program: The process involves: acquiring a target segmentation image and a target radar point cloud projection map in an image coordinate system; dividing the target camera image corresponding to the target segmentation image into a grid within the field of view in three-dimensional space to obtain multiple grid regions; allocating the target radar point cloud data to the corresponding grid regions based on the target segmentation image and the target point cloud projection map; performing planar fitting processing on each grid region based on the allocated point cloud data of each grid region to obtain the planar features of each grid region; determining the missing point locations in the target radar point cloud data based on the drivable area segmented from the target segmentation image and the missing point locations in the target point cloud projection map; and performing point cloud completion processing on the missing point locations based on the planar features.
[0022] Camera 103 is used to acquire images, such as images of the target camera.
[0023] The lidar 104 is used to acquire point cloud data images, such as acquiring target radar point cloud data.
[0024] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the self-moving device described above can be referred to the corresponding process in the following embodiment of the point cloud data completion processing method, and will not be repeated here.
[0025] The following will be based on Figure 1 Taking the self-moving device shown as the execution subject of the point cloud data completion processing method as an example, the point cloud data completion processing method provided in this application embodiment will be described in detail. For simplicity and ease of description, the execution subject will be omitted in subsequent method embodiments. It should be noted that... Figure 1 The scenarios described are only used to explain the point cloud data completion processing method provided in the embodiments of this application, but do not constitute a limitation on the application scenarios of the point cloud data completion processing method provided in the embodiments of this application.
[0026] Please see Figure 2 , Figure 2 This is a flowchart illustrating a point cloud data completion processing method provided in an embodiment of this application. The point cloud data completion processing method includes steps 201-206, wherein: 201. Obtain the target segmentation image and the target radar point cloud data in the image coordinate system as a projection of the target point cloud.
[0027] The target point cloud projection map and the target segmentation image are synchronized in time.
[0028] The target radar point cloud data refers to the radar point cloud data to be supplemented in each frame of radar point cloud data collected by the self-moving device during its operation. For example, the self-moving device can be equipped with a LiDAR, such as a low-cost sparse LiDAR, and the radar point cloud data can be collected by the LiDAR of the self-moving device.
[0029] Among them, the point cloud projection map refers to the image obtained by projecting radar point cloud data onto the image coordinate system.
[0030] Among them, the target point cloud projection map refers to the image obtained by projecting the target radar point cloud data onto the image coordinate system.
[0031] Here, the target camera image refers to the image captured during the driving process, whose acquisition time coincides with the acquisition time of the target radar point cloud data. For example, the self-moving device can be equipped with a camera, such as a monocular or binocular camera, and can acquire camera images through the self-moving device's camera.
[0032] Image segmentation refers to semantic segmentation of camera images to obtain images containing one or more drivable regions.
[0033] Among them, target segmentation image refers to the image obtained by semantic segmentation of target camera image, which contains one or more drivable areas.
[0034] Depending on the specific business application scenario, step 201 can be implemented in various ways. For example, it includes the following methods: <1> to <4> : <1> In some embodiments, two queues are used to maintain each frame of camera image and each frame of radar point cloud data collected during the movement of the self-moving device. In this case, step 201 may specifically include: adding each frame of camera image collected during the movement of the self-moving device to a first preset queue according to the acquisition time; adding each frame of radar point cloud data collected during the movement of the self-moving device to a second preset queue according to the acquisition time; when the acquisition time of the first preset queue's head camera image and the second preset queue's head radar point cloud data is consistent, extracting the head camera image as the target camera image and extracting the head radar point cloud data as the target radar point cloud data; performing semantic segmentation on the target camera image using a trained semantic segmentation model to obtain the target segmented image; and performing coordinate transformation based on the target radar point cloud data to obtain the target point cloud projection map. For example, taking an outdoor intelligent lawn mower as an example, the outdoor intelligent lawn mower is equipped with a monocular camera module and a low-cost sparse LiDAR. During execution, the device is first activated. A monocular camera acquires images of the lawn scene at a frequency of 80ms / frame, with each frame associated with a timestamp (accurate to milliseconds). A sparse LiDAR acquires point cloud data at a frequency of 100ms / frame, also recording timestamps. Each frame of camera imagery (each lawn scene image can be considered a single camera image) is stored in a first preset queue according to its acquisition time. Each frame of LiDAR point cloud data (each point cloud data frame can be considered a single LiDAR point cloud data frame) is stored in a second preset queue according to its acquisition time. The two queues can buffer a maximum of 10 frames. The system continuously monitors the timestamp of the first element in each queue, allowing for an error of ±5ms. If the timestamp of the first image in the first preset queue matches the timestamp of the first point cloud in the second preset queue, the first camera image is extracted as the target camera image, and the first LiDAR point cloud data is extracted as the target LiDAR point cloud data. Then, the corresponding first elements of both queues are deleted. For the target camera image, a trained semantic segmentation model (trained on the lawn scene dataset) is used for semantic segmentation to extract drivable areas such as the lawn and stone path, resulting in the target segmented image. For the target radar point cloud data, motion distortion is first removed using an IMU (Inertial Measurement Unit), then the data is projected onto the camera coordinate system according to the calibrated radar extrinsic parameters. After normalization, it is projected onto the image coordinate system based on the camera intrinsic parameters, filtering out distant point clouds and generating a target point cloud projection map. Thus, the obtained target point cloud projection map contains multiple point cloud location points, which correspond to multiple valid points in the target radar point cloud data, effectively generating a radar valid point cloud mask.
[0035] Furthermore, after acquiring camera images, pre-calibrated camera internal parameter calibration files (such as camera focal length, principal point coordinates, and distortion coefficients) can be used to perform image correction on the camera images to remove image distortion.
[0036] Furthermore, after acquiring radar point cloud data, the lidar of the self-moving device can also perform motion distortion removal on the radar point cloud data. For example, by using a high-frequency IMU, the line beam point cloud data from different times in a single radar frame can be unified to the same time.
[0037] In some implementations, the process of "performing coordinate transformation based on the target radar point cloud data to obtain the target point cloud projection map" can specifically include: projecting the target radar point cloud data into the camera coordinate system of the self-moving device according to the calibration radar to camera extrinsic parameters of the self-moving device, to obtain the camera point cloud data of the target radar point cloud data; and projecting the camera point cloud data into the image coordinate system of the self-moving device according to the calibration camera intrinsic parameters of the self-moving device, to obtain the target point cloud projection map of the target radar point cloud data. The calibration radar to camera extrinsic parameters are used to indicate the transformation relationship from the self-moving device's lidar coordinate system to the self-moving device's camera coordinate system, as shown in Formula 1. First, each point in the target radar point cloud data can be transformed from the lidar coordinate system to the camera coordinate system according to Formula 1, thereby realizing the projection of the target radar point cloud data into the camera coordinate system of the self-moving device to obtain the camera point cloud data of the target radar point cloud data. Then, referring to the transformation relationship from the camera coordinate system to the normalized coordinate system, as shown in Formula 2, each point in the camera point cloud data is transformed from the camera coordinate system to the normalized coordinate system, thereby obtaining the normalized point cloud data of the target radar point cloud data; then referring to the transformation relationship from the normalized coordinate system to the image coordinate system (i.e., the pixel coordinate system), as shown in Formula 3, each point in the normalized point cloud data is transformed from the normalized coordinate system to the image coordinate system (i.e., the pixel coordinate system), thereby obtaining the target point cloud projection map of the target radar point cloud data.
[0038] Formula 1 In Formula 1, x c y c z c These are the point cloud coordinates in the transformed camera coordinate system; r 11 ~r 33 t1 represents the rotation parameters of the x / y / z axes from the lidar coordinate system to the camera coordinate system; t1~t3 represents the translation parameters of the x / y / z axes from the lidar coordinate system to the camera coordinate system. All of these parameters can be obtained through joint calibration of the lidar and camera of the self-moving device. l y l z l These are the point cloud coordinates in the lidar coordinate system.
[0039] Formula 2 In Formula 2, xn y n These represent the x-axis and y-axis coordinates in the normalized coordinate system, respectively.
[0040] Formula 3 In Formula 3, (u,v) represents the coordinate values in the image coordinate system (i.e., pixel coordinate system); where f and (u0,v0) are the intrinsic parameters of the calibrated camera, f is the focal length of the camera, and (u0,v0) are the coordinates of the camera's optical center in the image coordinate system (i.e., pixel coordinate system). Through the above transformation, the target radar point cloud data can be converted to the image pixel plane to generate an effective radar point cloud mask, and radar point clouds exceeding the image boundary size are filtered out. For example, please refer to... Figure 3 , Figure 3 This is an illustrative diagram illustrating the conversion of target radar point cloud data into a target point cloud projection map in an embodiment of this application. For a certain frame of target radar point cloud data, such as Figure 3 As shown in (a), after transformations using formulas 1 to 3, a target point cloud projection image is obtained, which has filtered out radar point clouds exceeding the image boundary size. Figure 3 As shown in (b), since point clouds a1, a2, b1, c1, and c2 in the target radar point cloud data exceed the image boundary, they will be filtered out in the target point cloud projection map, and the point clouds within the image boundary range will be retained.
[0041] In some embodiments, taking an outdoor lawn scene as an example, a large number of outdoor lawn scene images can be collected. Then, the objects in the outdoor lawn scene images are classified and segmented (mainly divided into lawn, stone path, dirt path, cement path, gravel path, and other objects). Using the labeled outdoor lawn semantic segmentation data, a preset semantic segmentation model is trained. The weights of the semantic segmentation model with the smallest training error in the training results are saved, thus obtaining a trained semantic segmentation model. This trained semantic segmentation model can then be used to perform semantic segmentation on camera images to obtain segmented images. For example, by training the semantic segmentation model, semantic segmentation can be performed on a target camera image to obtain the target segmented image.
[0042] Furthermore, to ensure that the acquisition time of the camera image at the head of the first preset queue is consistent with the acquisition time of the radar point cloud data at the head of the second preset queue, the method may further include: detecting whether the time of the head element of the first preset queue is consistent with the time of the head element of the second preset queue; if the time of the head element of the first preset queue is inconsistent with the time of the head element of the second preset queue, and the time of the head element of the first preset queue is earlier than the time of the head element of the second preset queue, then removing the head element from the first preset queue; if the time of the head element of the first preset queue is inconsistent with the time of the head element of the second preset queue, and the time of the head element of the first preset queue is earlier than the time of the head element of the second preset queue, then removing the head element from the second preset queue; if the time of the head element of the first preset queue is consistent with the time of the head element of the second preset queue, then determining that the acquisition time of the camera image at the head of the first preset queue is consistent with the acquisition time of the radar point cloud data at the head of the second preset queue. For example, taking an outdoor smart lawnmower as an example, the first preset queue stores each frame of camera images, and the second preset queue stores each frame of radar point cloud data, both queues being sorted according to the acquisition time. The system reads the acquisition timestamp (T1) of the camera image at the head of the first preset queue and the acquisition timestamp (T2) of the radar point cloud data at the head of the second preset queue in real time, using ±5ms as the time consistency threshold. If |T1-T2|≤5ms, the time of the head element is determined to be consistent, confirming that the acquisition time of the head camera image and the head radar point cloud data is consistent, and can be subsequently extracted as the target camera image and target radar point cloud data. If |T1-T2|>5ms and T1<T2, it means that the time of the head element in the first preset queue is earlier, and the head element removal process is performed on the first preset queue, deleting the frame of camera image, and continuing to detect new head element times. If |T1-T2|>5ms and T1>T2, it means that the time of the head element in the second preset queue is earlier, and the head element removal process is performed on the second preset queue, deleting the frame of radar point cloud data, and the detection continues in a loop until a head element pair with consistent times is found.
[0043] <2> In some embodiments, two queues are used to maintain segmented images of each frame of camera images and radar point cloud data of each frame acquired during the movement of the mobile device. In this case, step 201 may specifically include: adding the segmented images of each frame of camera images acquired during the movement of the mobile device to a first preset queue according to the acquisition time; adding the radar point cloud data of each frame acquired during the movement of the mobile device to a second preset queue according to the acquisition time; when the acquisition time of the segmented image at the head of the first preset queue is the same as that of the radar point cloud data at the head of the second preset queue, extracting the segmented image at the head of the queue as the target segmented image, and extracting the radar point cloud data at the head of the queue as the target radar point cloud data; performing coordinate transformation based on the target radar point cloud data to obtain the target point cloud projection map.
[0044] In some embodiments, the process of "adding segmented images of each frame of camera images acquired during the movement of the self-moving device to a first preset queue according to the acquisition time" can be as follows: acquiring each frame of camera images acquired during the movement of the self-moving device; performing semantic segmentation on each frame of the camera images by training a semantic segmentation model to obtain segmented images of each frame of the camera images; and adding the segmented images of each frame of the camera images to the first preset queue according to the acquisition time. For example, "acquiring a target segmented image" can be as follows: acquiring a target camera image acquired at a target time during the movement of the self-moving device; performing semantic segmentation on the target camera image by training a semantic segmentation model to obtain the target segmented image.
[0045] Furthermore, to ensure that the acquisition time of the segmented image at the head of the first preset queue is consistent with the acquisition time of the radar point cloud data at the head of the second preset queue, the method may further include: detecting whether the time of the head element of the first preset queue is consistent with the time of the head element of the second preset queue; if the time of the head element of the first preset queue is inconsistent with the time of the head element of the second preset queue, and the time of the head element of the first preset queue is earlier than the time of the head element of the second preset queue, then removing the head element from the first preset queue; if the time of the head element of the first preset queue is inconsistent with the time of the head element of the second preset queue, and the time of the head element of the first preset queue is earlier than the time of the head element of the second preset queue, then removing the head element from the second preset queue; if the time of the head element of the first preset queue is consistent with the time of the head element of the second preset queue, then determining that the acquisition time of the segmented image at the head of the first preset queue is consistent with the acquisition time of the radar point cloud data at the head of the second preset queue. For example, taking an outdoor smart lawnmower as an example, the first preset queue stores segmented images of each frame of camera images, and the second preset queue stores radar point cloud data of each frame. Both queues are sorted according to the acquisition time. The system reads the acquisition timestamp (denoted as T1) of the segmented image at the head of the first preset queue and the acquisition timestamp (denoted as T2) of the radar point cloud data at the head of the second preset queue in real time, using ±5ms as the time consistency judgment threshold. If |T1-T2|≤5ms, the time of the head element is determined to be consistent, and it is determined that the acquisition time of the head segmented image and the head radar point cloud data is consistent, and it can be extracted as the target segmented image and target radar point cloud data. If |T1-T2|>5ms and T1<T2, it means that the time of the head element in the first preset queue is earlier. The head element removal process is performed on the first preset queue, the segmented image of that frame is deleted, and the detection of the new head element time continues. If |T1-T2|>5ms and T1>T2, it means that the time of the first element of the second preset queue is earlier. Perform the removal of the first element of the second preset queue, delete the radar point cloud data of that frame, and continue to detect in a loop until a pair of first elements with the same time is found.
[0046] <3> In some embodiments, two queues are used to maintain the point cloud projection maps of each frame of camera images and each frame of radar point cloud data collected during the movement of the mobile device. In this case, step 201 may specifically include: adding each frame of camera images collected during the movement of the mobile device to a first preset queue according to the acquisition time; adding the point cloud projection maps of each frame of radar point cloud data collected during the movement of the mobile device to a second preset queue according to the acquisition time; when the acquisition time of the camera image at the head of the first preset queue is the same as that of the point cloud projection map in the second preset queue, extracting the camera image at the head of the queue as the target camera image, and extracting the point cloud projection map at the head of the queue as the target point cloud projection map; and performing semantic segmentation on the target camera image by training a semantic segmentation model to obtain the target segmented image.
[0047] In some embodiments, "adding the point cloud projection images of each frame of radar point cloud data collected during the movement of the self-moving device to a second preset queue according to the acquisition time" may specifically include: acquiring each frame of radar point cloud data collected during the movement of the self-moving device; performing coordinate transformation based on each frame of radar point cloud data to obtain the point cloud projection image of each frame of radar point cloud data; and adding the point cloud projection images of each frame of radar point cloud data to the second preset queue according to the acquisition time. The specific implementation of "performing coordinate transformation based on each frame of radar point cloud data to obtain the point cloud projection image of each frame of radar point cloud data" is similar to "performing coordinate transformation based on the target radar point cloud data to obtain the target point cloud projection image," and can be referred to the preceding description for details, which will not be repeated here.
[0048] Furthermore, to ensure that the acquisition time of the camera image at the head of the first preset queue is consistent with the acquisition time of the point cloud projection map at the head of the second preset queue, the method may further include: detecting whether the acquisition time of the head element of the first preset queue is consistent with the acquisition time of the head element of the second preset queue; if the acquisition time of the head element of the first preset queue is inconsistent with the acquisition time of the head element of the second preset queue, and the acquisition time of the head element of the first preset queue is earlier than the acquisition time of the head element of the second preset queue, then removing the head element from the first preset queue; if the acquisition time of the head element of the first preset queue is inconsistent with the acquisition time of the head element of the second preset queue, and the acquisition time of the head element of the first preset queue is earlier than the acquisition time of the head element of the second preset queue, then removing the head element from the second preset queue; if the acquisition time of the head element of the first preset queue is consistent with the acquisition time of the head element of the second preset queue, then determining that the acquisition time of the camera image at the head of the first preset queue is consistent with the acquisition time of the point cloud projection map at the head of the second preset queue. For example, taking an outdoor smart lawnmower as an example, the first preset queue stores each frame of camera images, and the second preset queue stores each frame of radar point cloud data point cloud projection images. Both queues are sorted according to the acquisition time. The system reads the acquisition timestamp (denoted as T1) of the camera image at the head of the first preset queue and the acquisition timestamp (denoted as T2) of the point cloud projection image at the head of the second preset queue in real time, using ±5ms as the time consistency judgment threshold. If |T1-T2|≤5ms, the time of the head element is determined to be consistent, confirming that the acquisition time of the head camera image and the head point cloud projection image is consistent, and it can be extracted as the target camera image and target point cloud projection image. If |T1-T2|>5ms and T1<T2, it means that the time of the head element in the first preset queue is earlier. The head element removal process is performed on the first preset queue, deleting the frame of camera image, and continuing to detect the time of the new head element. If |T1-T2|>5ms and T1>T2, it means that the time of the first element of the second preset queue is earlier. Perform the removal of the first element of the second preset queue, delete the point cloud projection map of that frame, and continue to check in a loop until a pair of first elements with the same time is found.
[0049] <4> In some embodiments, two queues are used to maintain segmented images of each frame of camera images acquired during the movement of the mobile device and point cloud projection maps of each frame of radar point cloud data, respectively. In this case, step 201 may specifically include: adding the segmented images of each frame of camera images acquired during the movement of the mobile device to a first preset queue according to the acquisition time; adding the point cloud projection maps of each frame of radar point cloud data acquired during the movement of the mobile device to a second preset queue according to the acquisition time; when the acquisition time of the segmented image at the head of the first preset queue is the same as that of the point cloud projection map at the head of the second preset queue, extracting the segmented image at the head of the queue as the target segmented image, and extracting the point cloud projection map at the head of the queue as the target point cloud projection map.
[0050] Furthermore, to ensure that the acquisition time of the segmented image at the head of the first preset queue is consistent with the acquisition time of the point cloud projection image at the head of the second preset queue, the method may further include: detecting whether the acquisition time of the head element of the first preset queue is consistent with the acquisition time of the head element of the second preset queue; if the acquisition time of the head element of the first preset queue is inconsistent with the acquisition time of the head element of the second preset queue, and the acquisition time of the head element of the first preset queue is earlier than the acquisition time of the head element of the second preset queue, then removing the head element from the first preset queue; if the acquisition time of the head element of the first preset queue is inconsistent with the acquisition time of the head element of the second preset queue, and the acquisition time of the head element of the first preset queue is earlier than the acquisition time of the head element of the second preset queue, then removing the head element from the second preset queue; if the acquisition time of the head element of the first preset queue is consistent with the acquisition time of the head element of the second preset queue, then determining that the acquisition time of the segmented image at the head of the first preset queue is consistent with the acquisition time of the point cloud projection image at the head of the second preset queue. For example, taking an outdoor smart lawnmower as an example, the first preset queue stores segmented images of each frame of camera images, and the second preset queue stores point cloud projection images of each frame of radar point cloud data. Both queues are sorted according to the acquisition time. The system reads the acquisition timestamp (denoted as T1) of the segmented image at the head of the first preset queue and the acquisition timestamp (denoted as T2) of the point cloud projection image at the head of the second preset queue in real time, using ±5ms as the time consistency judgment threshold. If |T1-T2|≤5ms, the time of the head element is determined to be consistent, and it is determined that the acquisition time of the head segmented image and the head point cloud projection image is consistent, which can be extracted as the target segmented image and the target point cloud projection image. If |T1-T2|>5ms and T1<T2, it means that the time of the head element in the first preset queue is earlier. The head element removal process is performed on the first preset queue, deleting the segmented image of that frame, and continuing to detect the time of the new head element. If |T1-T2|>5ms and T1>T2, it means that the time of the first element of the second preset queue is earlier. Perform the removal of the first element of the second preset queue, delete the point cloud projection map of that frame, and continue to check in a loop until a pair of first elements with the same time is found.
[0051] 202. Divide the target camera image corresponding to the target segmentation image into a grid in the field of view region of three-dimensional space to obtain multiple grid regions.
[0052] The field of view (FOV) is the area of the target camera image within the corresponding field of view in three-dimensional space. The size of the FOV is determined by the field of view of the camera acquiring the target image and the camera's detection range. The FOV refers to the range of angular dimensions of the real-world environment that the camera can effectively perceive and cover, while the detection range refers to the range of spatial distances of the real-world environment that the camera can effectively perceive and cover. For example, ... Figure 4 As shown, Figure 4 This is a schematic diagram illustrating the division of the field of view region provided in the embodiments of this application. The field of view region of the target camera image in three-dimensional space is a region with a field of view of 60° and a detection distance of 3 meters.
[0053] There are multiple ways to implement step 202, including, for example, the following methods (1) to (3): (1) In some embodiments, the field of view area is divided into grids according to a preset block distance and a preset block angle to obtain multiple grid areas.
[0054] The specific values of the preset block distance and the preset block angle can be set according to the actual business scenario requirements. In this embodiment, the specific values of the preset block distance and the preset block angle are specified. For example, ... Figure 4 As shown, the preset segmentation distance can be 0.5 meters, and the preset segmentation angle can be 20°. Therefore, for a field of view area with a 60° field of view and a detection distance of 3 meters, the grid can be divided into 3 × 6 = 18 grid areas. Furthermore, the resulting grid areas can be numbered to distinguish each grid area and the adjacency relationships between them.
[0055] (2) In some embodiments, the field of view is divided into multiple grid regions according to a preset block distance. For example, as Figure 5 As shown, the preset block distance can be 0.5 meters. Therefore, for a field of view area with a field of view of 60° and a detection distance of 3 meters, the grid can be divided into 6 grid areas.
[0056] (3) In some embodiments, the field of view is divided into multiple grid regions according to a preset segmentation angle. For example, as shown in the figure... Figure 6 As shown, the preset segmentation angle can be 20°. Therefore, for a field of view area with a field of view of 60° and a detection distance of 3 meters, the grid can be divided into 3 grid areas.
[0057] 203. Based on the target segmentation image and the target point cloud projection map, allocate the target radar point cloud data to the corresponding grid regions.
[0058] Since the point cloud form used for planar fitting of each grid region is different, there are multiple ways to implement step 203, including, for example, the following methods (1) to (2): (1) In some embodiments, a planar fitting is performed using the first point cloud to be assigned within the grid area, thus assigning the first point cloud to be assigned to the grid area. The first point cloud to be assigned refers to the point cloud in the target radar point cloud data that falls into the target point cloud projection map and falls into the drivable area. In this case, step 203 may specifically include the following steps 2031A~2033A: 2031A. From the target radar point cloud data, obtain the location points of the point clouds that fall into the target point cloud projection map and fall into the drivable area, and obtain each of the first point clouds to be assigned.
[0059] Please refer to Figure 7 , Figure 7 This is an illustrative diagram illustrating the allocated point cloud data for determining a grid region in an embodiment of this application. For example... Figure 7 As shown in the target point cloud projection in (a), the point cloud locations in the projection include points a3-a8, b2-b8, and c3-c8. Except for points a3-a8, b2-b8, and c3-c8, the remaining points can be considered as missing points in the point cloud. It is understandable that... Figure 7 The point cloud locations and missing point locations shown are for illustrative purposes only and do not represent the actual number and distribution of point clouds. Figure 7 As shown in the target segmentation image in (b), semantic segmentation of the target camera image can extract drivable areas (such as lawns, flagstone paths, dirt roads, cement roads, gravel roads, etc.) from the target camera image. Figure 7 (c) shows the overlay effect of the target point cloud projection map and the target segmentation image. The point clouds in the target radar point cloud data that fall within the target point cloud projection map and are within the drivable area are: the point clouds corresponding to positions a4-a8, b2-b8, and c5-c8. Therefore, the point clouds corresponding to positions a4-a8, b2-b8, and c5-c8 can be used as the first point cloud to be assigned, as follows: Figure 7 As shown in (d).
[0060] 2032A. Perform region allocation processing on the first point cloud to be allocated to obtain the allocation grid region of the first point cloud to be allocated.
[0061] Please refer to Figure 8 , Figure 8 This is an illustrative diagram illustrating the point cloud allocation grid region provided in the embodiments of this application. For example, step 2032A may specifically include the following steps A1 to A3: A1. Obtain the index distance and index angle of the first point cloud to be assigned.
[0062] For example, the field of view area is divided into multiple grid areas according to a preset block distance and a preset block angle; the first camera coordinates of the first point cloud to be assigned can be obtained; the index distance is determined based on the first camera coordinates and the preset block distance according to a preset distance index formula (refer to formula 4); the index angle is determined based on the first camera coordinates and the preset block angle according to a preset angle index formula (refer to formula 5).
[0063] Here, the first camera coordinates refer to the coordinate values of the first point cloud to be assigned in the camera coordinate system of the self-moving device.
[0064] For example, assuming the preset distance index formula is as shown in Formula 4 and the preset angle index formula is as shown in Formula 5, please refer to... Figure 4 , Figure 7 and Figure 8 The field of view is divided into multiple grid regions according to a preset block distance (e.g., 0.5 meters) and a preset block angle (e.g., 20°). (For example, 18 grid regions are obtained, numbered 1, 2, ..., 18 from bottom to top and left to right). Figure 7 The first point cloud to be assigned (such as point clouds a4-a8, b2-b8, and c5-c8, corresponding to position points a4-a8, b2-b8, and c5-c8 respectively) is determined in the image. For example, for point cloud a4, the camera coordinates of point cloud a4 in the camera coordinate system (i.e., the first camera coordinates) can be extracted and substituted into the preset distance index formula shown in Formula 4 to calculate the index distance of point cloud a4 (e.g., 2.2 meters); substituted into the preset angle index formula shown in Formula 5, the index angle of point cloud a4 can be calculated (e.g., 5°). Similarly, for each first point cloud to be assigned, a similar process can be performed to obtain its index distance d and index angle a.
[0065] Formula 4 In Formula 4, d is the index distance of the point cloud obtained from the camera coordinates (x, y, z). seg The preset block distance is used; int represents the integer value of the result.
[0066] Formula 5 In Formula 5, 'a' represents the index angle of the point cloud obtained from the camera coordinates (x, y, z). seg The preset segmentation angle is ; int is the integer value for the result; arctan is the arctangent function.
[0067] A2. Obtain the mark distance range and mark angle range for each of the grid regions.
[0068] A3. Based on the marked distance range of each grid region, the marked angle range of each grid region, the index distance, and the index angle, the allocation grid region of the first point cloud to be allocated is obtained by matching.
[0069] To facilitate understanding, let's continue with the example from step A1. For instance, in step 202, the marked distance range and marked angle range of the 18 grid regions are as follows: Figure 4 As shown, the index distance of the first point cloud to be assigned (e.g., the index distance of point cloud a4 is 2.2 meters) can be compared with the marked distance range of each grid region, and the index angle of the first point cloud to be assigned (e.g., the index angle of point cloud a4 is 5°) can be compared with the marked angle range of each grid region to determine the marked distance range (e.g., 2 meters to 2.5 meters) and the marked angle range (e.g., 0° to 20°) into which the first point cloud to be assigned falls. The grid region corresponding to the marked distance range and the marked angle range into which the first point cloud to be assigned falls (e.g., grid region numbered 13) is used as the assigned grid region for the first point cloud to be assigned. Similarly, for each first point cloud to be assigned determined in the steps above, an assigned grid region can be determined for it. For example, refer to... Figure 4 and Figure 8 The grid regions assigned to point clouds a4-a8 are grid regions 13, 10, 7, 4, and 1, respectively; the grid regions assigned to point clouds b2-b8 are grid regions 17, 14, 11, 8, 5, and 2, respectively; and the grid regions assigned to point clouds c5-c8 are grid regions 12, 9, 6, and 3, respectively.
[0070] 2033A. When all the first point clouds to be allocated have been allocated, the allocated point cloud data of each grid region is obtained based on the allocation grid region of each of the first point clouds to be allocated.
[0071] Please refer to Figure 4 and Figure 8 To make it easier to understand, let's continue with the example above. For instance, the allocated point cloud data for grid region 1 is {a8}, the allocated point cloud data for grid region 2 is {b8}, the allocated point cloud data for grid region 3 is {c8}, and so on. Among them, the allocated point cloud data for grid regions 15, 16, and 18 are all empty.
[0072] (2) In some embodiments, a plane fitting is performed using the second point cloud to be assigned within the grid area, thus assigning the second point cloud to be assigned to the grid area. The second point cloud to be assigned refers to the point cloud in the target radar point cloud data that falls into the target point cloud projection map. In this case, step 203 may specifically include the following steps 2031B~2033B: 2031B. From the target radar point cloud data, obtain the second point cloud to be assigned for each point cloud location that falls into the target point cloud projection map.
[0073] Please refer to Figure 7 ,like Figure 7 As shown in the target point cloud projection diagram in (a), the point cloud locations in the target point cloud projection diagram include location points a3-a8, b2-b8, and c3-c8. Except for location points a3-a8, b2-b8, and c3-c8, the other location points can be regarded as missing point cloud locations. The point clouds in the target radar point cloud data that fall into the target point cloud projection diagram are the point clouds corresponding to location points a3-a8, b2-b8, and c3-c8. At this time, the point clouds corresponding to location points a3-a8, b2-b8, and c3-c8 can be used as the first point cloud to be assigned.
[0074] 2032B. Perform region allocation processing on the second point cloud to be allocated to obtain the allocation grid region of the second point cloud to be allocated.
[0075] 2033B. When all the second point clouds to be allocated have been allocated, the allocated point cloud data of each grid region is obtained based on the allocation grid region of each second point cloud to be allocated.
[0076] The implementation of steps 2032B to 2033B is similar to that of steps 2032A to 2033A. For details, please refer to the relevant explanations above. They will not be repeated here.
[0077] 204. Based on the allocated point cloud data of each grid region, perform plane fitting processing on each grid region to obtain the planar features of each grid region.
[0078] (1) In some embodiments, multiple grid regions can be traversed, and the plane of the currently traversed region can be fitted based on the average height of the allocated point cloud data of the currently traversed region; the plane normal vector (such as the unit direction vector) of the plane of the currently traversed region can be obtained as the plane feature of the currently traversed region. For example, taking an outdoor lawn mower as an example, please refer to... Figure 4 The field of view is divided into 18 grid regions with a preset block spacing of 0.5 meters and a preset block angle of 20°. The preset number threshold is set to 10, and the preset plane normal vector is (0,1,0). The system traverses each grid region, and based on the average height of the point cloud data of each grid region assigned to the current traversed region, it uses RANSAC to fit a plane and extracts the plane normal vector (e.g., (0.02,0.98,0.15)) as the plane feature of the current traversed region.
[0079] (2) In some embodiments, multiple grid regions can be traversed, and the number of point clouds in the currently traversed region can be detected. If the number of point clouds is greater than a preset threshold, the plane of the currently traversed region is fitted based on the average height of the point clouds in the currently traversed region. The plane normal vector of the plane of the currently traversed region is obtained as the plane feature of the currently traversed region. If the number of point clouds is less than or equal to the preset threshold, and the surrounding grid regions of the currently traversed region have plane features, the average plane feature value of the surrounding grid regions of the currently traversed region is obtained as the plane feature of the currently traversed region. If the number of point clouds is less than or equal to the preset threshold, and the surrounding grid regions of the currently traversed region do not have plane features, the camera height of the target segmentation image and the preset plane normal vector are used as the plane feature of the currently traversed region.
[0080] The specific values of the preset plane normal vector and the preset quantity threshold can be set according to the actual business scenario requirements. Here, there are no restrictions on the specific values of the preset plane normal vector and the preset quantity threshold. For example, the preset plane normal vector can be a normal vector perpendicular to the ground (0,1,0), and the preset quantity threshold can be 10.
[0081] Among them, the surrounding grid region refers to the grid region surrounding the current traversed region (i.e., the grid region adjacent to the current traversed region) among the multiple grid regions divided for the field of view.
[0082] Among them, "the surrounding grid regions of the currently traversed region possess planar features" means that one or more grid regions among all the surrounding grid regions of the currently traversed region possess valid planar features, for example, referencing... Figure 4 The surrounding grid regions of grid region 5 (i.e., grid region number 5) may include grid regions 1, 2, 3, 4, 6, 7, 8, and 9. If one or more of grid regions 1, 2, 3, 4, 6, 7, 8, and 9 have planar features, then the surrounding grid regions of the currently traversed region are considered to have planar features.
[0083] Wherein, "no planar features exist in the surrounding grid regions of the currently traversed region" means that none of the surrounding grid regions of the currently traversed region have valid planar features. For example, referencing... Figure 4 The surrounding grid regions of grid region 18 (i.e., grid region number 18) may include grid regions 14, 15, and 17. If grid regions 14, 15, and 17 do not have planar features, then it is assumed that the surrounding grid regions of the currently traversed region do not have planar features.
[0084] For example, taking a self-moving device like an outdoor smart lawn mower as an example, please refer to... Figure 4The field of view is divided into 18 grid regions with a preset block spacing of 0.5 meters and a preset block angle of 20°. The preset number threshold is set to 10, and the preset plane normal vector is (0,1,0). The system traverses each grid region and first detects the number of point cloud data allocated to the currently traversed region. If the effective number of point cloud data in a certain grid region (such as number 5) is 15 (>10), the average height of the point cloud in this set is calculated, and the plane is fitted using RANSAC. The plane normal vector of the plane (such as (0.02,0.98,0.15)) is extracted as the planar feature of the grid region (such as number 5). If a grid region (e.g., number 11) has 6 valid point cloud values (≤10), and at least one of its surrounding grid regions 7, 8, 9, 10, 12, 13, 14, and 15 has valid planar features (e.g., surrounding grid regions 7, 8, 9, 10, 12, and 13 have planar features), then the average of the planar features of these surrounding grid regions is taken as the planar feature of that grid region (e.g., number 11). If a grid region (e.g., number 18) has 3 valid point cloud values (≤10), and its surrounding grid regions do not have planar features (i.e., surrounding grid regions 14, 15, and 17 do not have planar features), then the camera height of the target segmentation image (e.g., 1.2 meters) and the preset plane normal vector (0, 1, 0) are taken as the planar feature of that region. After traversal, all grid regions have obtained their corresponding planar features.
[0085] 205. Based on the drivable area segmented from the target segmentation image and the missing point cloud locations in the target point cloud projection image, determine the missing point locations in the target radar point cloud data.
[0086] For example, please refer to Figure 7 and Figure 9 , Figure 9 This is an illustrative diagram illustrating the position points to be filled in provided in the embodiments of this application, such as... Figure 7 As shown, assuming the point cloud locations in the target point cloud projection map include locations a3-a8, b2-b8, and c3-c8, all other locations except a3-a8, b2-b8, and c3-c8 can be considered as missing point cloud locations. Traverse each location point corresponding to the target point cloud projection map; if a location is a missing point cloud location and is within the drivable area, such as... Figure 9 As shown by point P1, this location is designated as the point to be filled in. If this location is a missing point in the point cloud and is outside the drivable area, such as... Figure 9 As shown at point P2, skip this point and continue traversing to the next point. If the point is a location in the point cloud and is within the drivable area, then... Figure 9If point a4 is shown, then skip that point and continue traversing to the next point. This process can be repeated to find all the points in the target radar point cloud data that need to be filled in.
[0087] 206. Based on the planar features, perform point cloud completion processing on the points to be completed.
[0088] For example, step 206 may specifically include the following steps 2061 to 2064: 2061. Obtain the second camera coordinates of the position point to be filled.
[0089] Here, the second camera coordinates refer to the coordinate values of the point to be filled in the camera coordinate system of the mobile device. In step 205, the point to be filled in is found based on the target point cloud projection map, so the image coordinates of the point to be filled in (i.e., the coordinate values of the point to be filled in the image coordinate system) can be obtained. Referring to the transformation relationship of formulas 2 and 3, the coordinate values of the point to be filled in the camera coordinate system of the mobile device can be determined, thus obtaining the second camera coordinates of the point to be filled in.
[0090] 2062. Determine the three-dimensional coordinates of the position point to be filled based on the second camera coordinates and the preset camera height.
[0091] The camera height of the self-moving device is known. The preset camera height can be the camera height of the self-moving device. Assuming that the self-moving device is on a horizontal plane, with the camera height of the self-moving device known, the three-dimensional coordinates (such as (x_ipm, y_ipm, z_ipm)) of the position to be filled can be calculated by using the second camera coordinates (such as P(x_2d, y_2d)) of the position to be filled, referring to the following formula 6.
[0092] Formula 6 In Formula 6, f and (u0,v0) are the intrinsic parameters of the calibrated camera, f is the focal length of the camera, (u0,v0) are the coordinates of the camera's optical center in the image coordinate system (i.e., the pixel coordinate system), and (x_ipm, y_ipm, z_ipm) are the three-dimensional coordinates of the point to be filled.
[0093] 2063. Based on the three-dimensional coordinates, determine the grid region where the position point to be filled is located.
[0094] For example, firstly, based on the 3D coordinates of the point to be filled, the index distance and index angle of the point to be filled are obtained; the specific implementation is similar to step A1, and can be referred to the relevant explanations above, so it will not be repeated here. Then, the marked distance range and marked angle range of each grid region are obtained; the specific implementation is similar to step A2, and can be referred to the relevant explanations above, so it will not be repeated here. Next, based on the marked distance range and marked angle range of each grid region, the index distance and index angle of the point to be filled are matched to obtain the grid region where the point to be filled is located; the specific implementation is similar to step A3, and can be referred to the relevant explanations above, so it will not be repeated here.
[0095] 2064. Based on the planar features of the grid region, perform point cloud completion processing on the points to be completed.
[0096] For example, the planar features of the grid region include the planar normal vector of the grid region. First, based on the planar features of the grid region and the preset pitch angle formula, the pitch angle of the grid region is determined. For example, the pitch angle (θ) in the camera coordinate system (x to the right, y down, z forward) is defined as the angle of rotation of the camera around its own x-axis. The preset pitch angle formula is shown in Formula 7. The core principle is that the pitch angle is determined by the angle between the projection of the normal vector on the yz plane and the camera's y-axis (when rotating around the x-axis, the normal vector only shifts in the yz direction). Substituting the planar normal vector (nx, ny, nz) of the grid region into the preset pitch angle formula shown in Formula 7, the pitch angle θ of the grid region is calculated.
[0097] θ=arcsin(nz)=arctan(nz,-ny)Formula 7 In Formula 7, θ is the pitch angle of the grid region, and (nx, ny, nz) is the plane normal vector of the grid region.
[0098] Based on the planar features of the grid region and the preset roll angle formula, the roll angle of the grid region is determined. For example, the roll angle (φ) in the camera coordinate system (x to the right, y down, z forward) is defined as the angle of rotation of the camera around its own z-axis. The preset roll angle formula is shown in Formula 8. The core principle is that the roll angle is determined by the angle between the projection of the normal vector on the xy plane and the camera's y-axis (when rotating around the z-axis, the normal vector only shifts in the xy direction). Substituting the plane normal vector (nx, ny, nz) of the grid region into the preset roll angle formula shown in Formula 8, the pitch angle φ of the grid region is calculated.
[0099] φ=arcsin(nx)=arctan(nx,-ny)Formula 8 In Formula 8, φ is the roll angle of the grid region, and (nx,ny,nz) is the plane normal vector of the grid region.
[0100] Then, based on the average point cloud height of the grid region, the pitch angle of the grid region, and the roll angle of the grid region, the three-dimensional coordinates are corrected to obtain the corrected coordinates. Based on the corrected coordinates, point cloud completion processing is performed on the missing point. For example, the correction formula is shown in Formula 9. The three-dimensional coordinates (x_ipm, y_ipm, z_ipm), the second camera coordinates (x_2d, y_2d) of the point to be filled, the average point cloud height d_mean(xm, ym, zm) of the grid region, the pitch angle θ of the grid region, and the roll angle φ of the grid region are substituted into Formula 9 to obtain the corrected coordinates (x_ipm_cor, y_ipm_cor, z_ipm_cor). Adding the corrected coordinates (x_ipm_cor, y_ipm_cor, z_ipm_cor) to the target radar point cloud data can complete the missing point cloud points. Therefore, by utilizing the local planar features fitted to the corresponding grid region (i.e., the average point cloud height, pitch angle, and roll angle of the grid region), the 3D coordinates can be corrected. This can correct the 3D coordinates (x_ipm, y_ipm, z_ipm) obtained by inverse perspective mapping (IPM), solve the coordinate estimation error problem of uneven drivable areas (such as slopes, undulating grass, etc.), and thus improve the accuracy of point cloud completion.
[0101] Formula 9 In Formula 9, f is the focal length of the camera, (u0, v0) are the coordinates of the camera optical center in the image coordinate system (i.e., pixel coordinate system), (x_ipm, y_ipm, z_ipm) are the three-dimensional coordinates of the point to be filled, (x_2d, y_2d) are the second camera coordinates of the point to be filled, d_mean(xm, ym, zm) is the average point cloud height of the grid region, and (x_ipm_cor, y_ipm_cor, z_ipm_cor) are the corrected coordinates.
[0102] It is evident that combining the target segmentation image with the target point cloud projection map, dividing the field of view into multiple grid regions, and using RANSAC to fit the plane within the local grid regions can more accurately describe the true height and normal vector of the local surface, rather than relying on the global horizontal assumption. This can effectively improve the accuracy of inverse perspective transformation projection on uneven surfaces such as slopes and uneven lawns, making the completed point cloud more consistent with the actual terrain and improving the accuracy and terrain adaptability of the completed drivable area. By acquiring the target radar point cloud projection map and the target segmentation image, point cloud completion processing can be performed. Image semantic information and effective sparse radar point clouds can be used to guide the completion. For invalid radar points (i.e., missing points in the point cloud), the inverse perspective transformation projection (i.e., based on the average point cloud height, pitch angle, and roll angle of the grid region, the three-dimensional coordinates are corrected by local geometric constraints (i.e., the planar features fitted to the local grid region) to obtain the corrected coordinates (x_ipm_cor, y_ipm_cor, z_ipm_cor)) rather than completely relying on point cloud density or complex generation models. This reduces the requirement for point cloud density, enabling effective operation in low-beam LiDAR scenarios. As a result, self-moving devices can be adapted to low-cost sparse LiDAR, reducing the cost burden for consumers who purchase self-moving devices with high-end dense LiDAR (such as high-end dense LiDAR lawnmowers). By applying the solution of this application to lawnmowers, the drivable area of the lawnmower can be supplemented, solving the problems of "missed mowing" (sparse point cloud missing drivable areas) and "incorrect mowing" (low supplementation accuracy leading to the identification of non-lawn areas), alleviating the problems of lawnmowers getting stuck and colliding in complex lawns (with undulations and many obstacles), and improving mowing efficiency.
[0103] As can be seen from the above, firstly, by determining the missing point cloud locations in the target radar point cloud data based on the segmented drivable area in the target segmentation image and the point cloud projection map, the missing point cloud locations in the drivable area can be filled in using effective visual information, thereby improving the integrity of the point cloud and thus the completeness of the drivable area identification. Secondly, since planar fitting processing of each grid region can correct for local terrain differences, the planar features of each grid region are obtained by using the allocated point cloud data of each grid region for planar fitting processing. This feature is then used to fill in the missing point cloud locations (i.e., the missing point cloud locations in the drivable area). Corrections can be made for local terrain differences before filling in the missing points, ensuring that the filled point cloud reflects the actual ground height, thereby improving the accuracy of point cloud filling, the reconstruction accuracy of the drivable area, and ultimately the identification accuracy of the drivable area. Thirdly, by performing point completion processing based on planar features, rather than directly using visual recognition results for point cloud completion, the problem of errors in drivable area recognition caused by the susceptibility of visual information to lighting conditions (such as cloudy days or shadows) is reduced. This improves the accuracy of drivable area completion in complex outdoor environments, enhances the reconstruction accuracy of drivable areas, and consequently improves the overall accuracy of drivable area recognition. Therefore, the completeness and accuracy of drivable area recognition can be improved.
[0104] Those skilled in the art will understand that all or part of the steps in the above-described point cloud data completion processing method can be completed by instructions, or by controlling related hardware through instructions. These instructions can be stored in a computer-readable storage medium and loaded and executed by the processing unit.
[0105] Therefore, embodiments of this application provide a computer-readable storage medium storing multiple computer programs that can be loaded by a processing unit to execute any of the point cloud data completion processing methods provided in embodiments of this application.
[0106] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0107] In the above embodiments of the self-movable device and computer-readable storage medium, the descriptions of each embodiment have different focuses. For parts not described in detail in a particular embodiment, please refer to the relevant descriptions of other embodiments. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes and beneficial effects of the computer-readable storage medium, self-movable device, and their corresponding units described above can be referred to the description of the point cloud data completion processing method in the above embodiments, and will not be repeated here.
[0108] The above provides a detailed description of a point cloud data completion processing method, a self-moving device, and a computer-readable storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A processing method for point cloud data completion, characterized in that, The method includes: Acquire the target segmentation image and the target radar point cloud data in the image coordinate system as a projection map of the target point cloud; The target camera image corresponding to the segmented target image is divided into multiple grid regions in the field of view of three-dimensional space. Based on the target segmentation image and the target point cloud projection map, the target radar point cloud data is allocated to the corresponding grid regions; Based on the allocated point cloud data of each grid region, a plane fitting process is performed on each grid region to obtain the planar features of each grid region; Based on the drivable area segmented from the target segmentation image and the missing point cloud locations in the target point cloud projection image, the missing point locations in the target radar point cloud data are determined. Based on the planar features, point cloud completion processing is performed on the points to be completed.
2. The processing method of point cloud data completion according to claim 1, characterized in that, The step of allocating the target radar point cloud data to corresponding grid regions based on the target segmentation image and the target point cloud projection image includes: From the target radar point cloud data, obtain the location points of the point cloud that fall into the target point cloud projection map and fall into the drivable area, and obtain each of the first point clouds to be assigned. Perform region allocation processing on the first point cloud to be allocated to obtain the allocation grid region of the first point cloud to be allocated; When all the first point clouds to be allocated have been allocated, the allocated point cloud data of each grid region is obtained based on the allocation grid region of each first point cloud to be allocated.
3. The processing method of point cloud data completion according to claim 1, characterized in that, The step of performing planar fitting processing on each grid region based on the allocated point cloud data of each grid region to obtain the planar features of each grid region includes: The multiple grid regions are traversed, and the plane of the currently traversed region is fitted based on the average height of the point cloud data of the allocated point cloud data of the currently traversed region; Obtain the plane normal vector of the plane in the currently traversed region, and use it as the plane feature of the currently traversed region.
4. The processing method of point cloud data completion according to claim 3, characterized in that, The method further includes: Detect the number of point cloud data points allocated in the currently traversed region; The fitting of the plane of the currently traversed region based on the average height of the allocated point cloud data of the currently traversed region includes: If the number of point clouds is greater than a preset threshold, then the plane of the current traversed region is fitted based on the average height of the point cloud data of the currently traversed region.
5. The point cloud data completion processing method according to claim 4, characterized in that, The method further includes: If the number of point clouds is less than or equal to a preset threshold, and the surrounding grid regions of the current traversed region have planar features, then the average value of the planar features of the surrounding grid regions of the current traversed region is obtained as the planar features of the current traversed region.
6. The point cloud data completion processing method according to claim 4, characterized in that, The method further includes: If the number of point clouds is less than or equal to a preset number threshold, and the surrounding grid region of the current traversed region does not have planar features, then the camera height of the target segmented image and the preset plane normal vector are used as the planar features of the current traversed region.
7. The point cloud data completion processing method according to claim 1, characterized in that, The point cloud completion process based on the planar features for the points to be completed includes: Obtain the second camera coordinates of the position points to be filled; The three-dimensional coordinates of the point to be filled are determined based on the second camera coordinates and the preset camera height. Based on the three-dimensional coordinates, determine the grid region where the point to be filled is located; Based on the planar features of the grid region, point cloud completion processing is performed on the points to be completed.
8. The point cloud data completion processing method according to claim 7, characterized in that, The point cloud completion processing for the points to be completed based on the planar features of the grid region includes: Based on the planar characteristics of the grid region and the preset pitch angle formula, the pitch angle of the grid region is determined; Based on the planar characteristics of the grid region and the preset roll angle formula, the roll angle of the grid region is determined; Based on the average point cloud height of the grid region, the pitch angle of the grid region, and the roll angle of the grid region, the three-dimensional coordinates are corrected to obtain the corrected coordinates; Based on the corrected coordinates, point cloud completion processing is performed on the points to be filled.
9. The point cloud data completion processing method according to claim 1, characterized in that, The acquisition of the target segmentation image and the target radar point cloud data projected onto the target point cloud in the image coordinate system includes: The segmented images of each frame of camera image captured during the movement of the mobile device are added to the first preset queue according to the acquisition time. Each frame of radar point cloud data collected during the movement of the mobile device is added to the second preset queue according to the collection time. When the acquisition time of the segmented image at the head of the first preset queue is the same as that of the radar point cloud data at the head of the second preset queue, the segmented image at the head of the queue is extracted as the target segmented image, and the radar point cloud data at the head of the queue is extracted as the target radar point cloud data. Based on the target radar point cloud data, coordinate transformation is performed to obtain the target point cloud projection map.
10. The point cloud data completion processing method according to claim 9, characterized in that, The step of performing coordinate transformation based on the target radar point cloud data to obtain the target point cloud projection map includes: According to the calibration radar of the self-moving device and the camera extrinsic parameters, the target radar point cloud data is projected onto the camera coordinate system of the self-moving device to obtain the camera point cloud data of the target radar point cloud data. According to the calibration camera intrinsic parameters of the self-moving device, the camera point cloud data is projected onto the image coordinate system of the self-moving device to obtain the target point cloud projection map of the target radar point cloud data.
11. The point cloud data completion processing method according to claim 1, characterized in that, The step involves dividing the camera image corresponding to the target segmentation image into a grid within the field of view in three-dimensional space, resulting in multiple grid regions, including: The field of view is divided into multiple grid regions by dividing the field of view into grids according to a preset block distance and / or a preset block angle.
12. A self-moving device, characterized in that, The device includes a camera, a lidar, a processing unit, and a memory, wherein the memory stores a computer program, and the processing unit executes the point cloud data completion processing method as described in any one of claims 1 to 11 when it calls the computer program in the memory.
13. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processing unit to execute the point cloud data completion processing method according to any one of claims 1 to 11.