Method for perorming point cloud dynamic object filtering on basis of fusion of camera and lidar
Through the camera lidar fusion method, combined with visual object detection and clustering segmentation, the problem of dragging caused by dynamic objects in SLAM is solved, which improves the positioning accuracy of mobile robots in dynamic scenarios and reduces the calculation cost.
Patent Information
- Application Number
- PCT/CN2024/088148
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-04-17
- Publication Date
- 2025-06-26
AI Technical Summary
The existing SLAM method assumes that the environment is static, dynamic objects cause shadowing in the map, affecting positioning and path planning, and the existing deep learning algorithms are costly for point cloud processing and are difficult to implement on mobile robot platforms.
The point cloud dynamic object filtering method is adopted with camera lidar fusion, and the point cloud is acquired through lidar and the camera is acquired. The point cloud is projected into the image using external parameters, and dynamic objects are filtered out by combining visual object detection and cluster segmentation.
It improves the positioning accuracy of autonomous mobile robots in dynamic scenarios, effectively alleviates the impact of dynamic objects on map construction, reduces the computing power requirement, and achieves efficient real-time dynamic point cloud filtering.
Smart Images

Figure CN2024088148_26062025_PF_FP_ABST
Abstract
Description
Point cloud dynamic object filtering method based on camera lidar fusion Technical Field
[0001] The present invention relates to a point cloud dynamic object filtering method based on camera laser radar fusion, and belongs to the field of mobile robot autonomous navigation. Background Art
[0002] With the advancement of science and technology and the increasing demands of society, robots have gradually become part of people's lives, with their applications expanding across a wider range and playing an increasingly important role in all aspects of human life. Currently, with the booming development of artificial intelligence, the mobile robotics industry is also experiencing rapid growth. An increasing number of mobile robot products are moving from experimental to practical applications, such as restaurant service robots and household sweeping robots in indoor scenarios, and campus cleaning robots and warehouse logistics robots in outdoor scenarios. This has led to higher demands for the autonomy and intelligence of mobile robots. In particular, robots offer irreplaceable advantages in specialized scenarios, such as post-disaster search and rescue, high-altitude operations, extreme environments like high and low temperatures and strong radiation, and deep space planetary exploration, assisting humans in completing tasks. The emergence of robots has significantly increased productivity, and continuously improving their intelligence and autonomy is a key development direction for the future of robotics.
[0003] For mobile robots to achieve unmanned and intelligent operation, reliable autonomous navigation is crucial. Autonomous navigation has been a hot research area, and Simultaneous Localization and Mapping (SLAM) technology is key to solving this problem. With precise and stable SLAM algorithms, a robot can determine its own position and perceive the surrounding structure for mapping. Combined with motion planning algorithms, it can search for a navigable path to its target, thereby completing autonomous navigation.
[0004] Using sensors like lidar and cameras, robots can reconstruct the spatial structure of their surroundings online. However, most existing SLAM methods assume that the environment they perceive is static and time-invariant, a assumption that severely limits the applicability of SLAM systems in real-world scenarios. Dynamic objects can cause significant artifacts in the resulting map, negatively impacting subsequent map-based positioning and path planning. Dynamic objects can also cause deviations between current observations and the prior map, potentially leading to positioning drift or failure.
[0005] Some existing technologies alleviate the phenomenon of decreased robot positioning accuracy in dynamic scenarios by installing external reflectors in the environment. This method requires the installation of reflectors at fixed locations in the external environment as landmarks in advance, which does not fully utilize the performance of the mobile robot's own sensors and is limited by the environment.
[0006] Furthermore, with the rapid development of deep learning, neural networks have become a new means of processing large-scale 3D point cloud data. The ideal method for filtering out dynamic points online is to directly input the current frame's point cloud and output the dynamic points end-to-end through a 3D point cloud object detection or semantic segmentation algorithm. However, due to the disorder, sparsity, scale invariance, and large size of point cloud information, existing deep learning algorithms for processing point clouds are computationally expensive and typically require high-performance professional graphics cards to meet the computing resource requirements. This is often difficult to achieve for platforms like mobile robots, which have limited loads and computing resources.
[0007] Summary of the Invention
[0008] The present invention provides a point cloud dynamic object filtering method based on camera-lidar fusion, aiming to solve at least one of the technical problems existing in the prior art.
[0009] The technical solution of the present invention relates to a method for filtering dynamic objects in a point cloud. The method is based on a data acquisition platform equipped with a camera and a laser radar. The method includes the following steps:
[0010] S100, acquiring a real-time point cloud through the laser radar, and synchronously acquiring a real-time image through the camera; wherein the camera and the laser radar are pre-calibrated with external parameters using an external parameter automatic calibration method;
[0011] S200, obtaining a dynamic object detection frame in the image through a deep learning-based visual target detection network;
[0012] S300, projecting the point cloud into the image according to the extrinsic parameters obtained by pre-calibration;
[0013] S400 , after filtering out the point cloud within the dynamic object detection frame in combination with the point cloud projected onto the image, cluster and segment the point cloud within the dynamic object detection frame and then filter out the point cloud.
[0014] Furthermore, the external parameter calibration of the camera and the lidar includes the following steps:
[0015] S110, extracting depth-continuous point cloud edge features from the original point cloud, and extracting image edge features from the original image;
[0016] S120, matching the point cloud edge features with the image edge features;
[0017] S130: Construct a nonlinear optimization equation based on the successfully found matching feature points.
[0018] Furthermore, in step S110, the edge features in the original image are extracted using the Canny edge detection algorithm, which includes the following steps:
[0019] S111, accumulating the original point cloud data obtained by the laser radar for a period of time to form a relatively dense point cloud;
[0020] S112, voxelizing the point cloud;
[0021] S113 , fitting a plane to each voxel and retaining adjacent plane pairs that form a certain angle therein, and extracting edge features of the point cloud by solving the intersection lines of the adjacent plane pairs.
[0022] Further, step S120 includes the following steps:
[0023] S121. Projecting the point cloud edge features onto the image based on the initial extrinsic parameters:
[0024] Where, L P i Represents a point cloud edge feature point; p i =(u i ,v i ) represents the point obtained by projecting the above point cloud edge feature points onto the image plane; represents the initial extrinsic rotation from the radar coordinate system to the camera system; represents the initial extrinsic translation from the radar coordinate system to the camera coordinate system; π(P) represents the pinhole model of the camera; and f(P) represents the distortion model of the camera. Initial extrinsic parameters are usually obtained by manual measurement.
[0025] S122, through K nearest neighbor search in p i Find the k nearest image edge feature points around, recorded as a set To obtain the direction vectors of the k image edge feature points; wherein the direction vectors of the k image edge feature points are calculated by the following formula:
[0026] In the formula, the mean value of the edge feature pixels of the above image is q i And the covariance matrix is S i ;
[0027] The covariance matrix Si The eigenvector corresponding to the maximum eigenvalue is set as the direction vector of the image edge feature, so as to further filter and match the point cloud edge features and the image edge features according to the approximation degree of the direction vector.
[0028] Furthermore, in step S130, the nonlinear optimization equation constructed for joint calibration is expressed as follows:
[0029] Where N represents the number of matches between point cloud edge features and image edge features.
[0030] Furthermore, in step S300, the pixel coordinates [vv] of the point cloud projected onto the image plane are T Obtained by the following calculation:
[0031] In the formula, represents any point in a frame of point cloud data collected by the laser radar; represents the external parameter between the laser radar and the camera; represents the internal parameter of the camera.
[0032] The technical solution of the present invention also relates to a computer-readable storage medium having program instructions stored thereon, and the above-mentioned method is implemented when the program instructions are executed by a processor.
[0033] The technical solution of the present invention also relates to a point cloud dynamic object filtering system based on camera-lidar fusion, wherein the system includes a computer device containing the above-mentioned computer-readable storage medium.
[0034] Furthermore, the point cloud dynamic object filtering system also includes a data acquisition platform, which includes a camera and a laser radar, and the camera and the laser radar are connected through a 3D printed structural part.
[0035] The beneficial effects of the present invention are as follows:
[0036] The present invention is based on a point cloud dynamic object filtering method based on camera-lidar fusion. It combines two different sensors, camera and lidar, to perform dynamic object filtering in point clouds. This can improve the positioning accuracy of autonomous mobile robots in dynamic scenes and effectively alleviate the impact of dynamic objects on mapping. The present invention uses lidar and camera as the smallest sensor units. The lidar model used is the Livox AVIA solid-state lidar, which has a similar FoV to that of a camera. By performing target detection on the image, it avoids the use of deep learning algorithms to directly detect point clouds, reducing computing power requirements. The visual target detection algorithm used can be lightweight and deployed on the computing platform carried by the robot end. The entire real-time dynamic point cloud filtering process can achieve high computing efficiency. Therefore, it can be easily added as a preprocessing process to the laser SLAM algorithm, which can effectively reduce the phenomenon of dynamic objects producing ghosting during mapping. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] FIG1 is a basic flow chart of a method for filtering out dynamic objects in a point cloud according to the present invention.
[0038] FIG2 is a schematic structural diagram of a mobile robot data acquisition platform according to an embodiment of the present invention.
[0039] FIG3 is a schematic structural diagram of a handheld data acquisition device according to an embodiment of the present invention.
[0040] FIG4 is a schematic diagram of residual optimization of point cloud edge features and image edge features according to an embodiment of the present invention.
[0041] FIG5 a is a diagram showing the projection effect of using initial extrinsic parameters to estimate the projection effect according to an embodiment of the present invention.
[0042] FIG5 b is a projection effect diagram of the extrinsic parameters after calibration according to an embodiment of the present invention.
[0043] FIG6 is a schematic diagram of visual target detection results according to an embodiment of the present invention.
[0044] FIG. 7 is a schematic diagram of the association between point cloud and image information according to an embodiment of the present invention.
[0045] FIG8 is a diagram showing the dynamic point cloud clustering and segmentation results according to an embodiment of the present invention.
[0046] FIG9 a is a diagram showing the result of image target detection using the method according to an embodiment of the present invention.
[0047] FIG9 b is a diagram showing the dynamic point cloud segmentation result using the method according to an embodiment of the present invention.
[0048] FIG9 c is a diagram of the FAST-LIO online mapping result according to an embodiment of the present invention.
[0049] FIG9 d is a diagram showing the online mapping result after real-time dynamic point cloud filtering using the method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The following will provide a clear and complete description of the concept, specific structure and technical effects of the present invention in conjunction with the embodiments and drawings to fully understand the purpose, scheme and effects of the present invention.
[0051] It should be noted that, unless otherwise specified, when a feature is referred to as being "fixed" or "connected" to another feature, it may be directly fixed or connected to the other feature, or it may be indirectly fixed or connected to the other feature. The singular forms "a", "said" and "the" used herein are also intended to include the plural forms, unless the context clearly indicates otherwise. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art. The terms used in this specification are only for describing specific embodiments and are not intended to limit the invention. The term "and / or" used herein includes any combination of one or more related listed items.
[0052] Should be understood that, although the present disclosure may adopt the term first, second, third etc. to describe various elements, these elements should not be limited to these terms.These terms are only used to distinguish the elements of the same type from each other.For example, without departing from the scope of the present disclosure, the first element may also be referred to as the second element, and similarly, the second element may also be referred to as the first element.The use of any and all examples or exemplary language ("for example", "such as" etc.) provided herein is only intended to better illustrate embodiments of the present invention, and unless otherwise required, will not impose limitations on the scope of the present invention.
[0053] 1 to 3 , the technical solution of the present invention is based on a data acquisition platform. The sensors used in the data acquisition platform include cameras and lidars. Point cloud information is collected in real time by the lidars, and image information is collected synchronously by the cameras. The connection between the lidars and the cameras uses 3D-printed structural parts to achieve a rigid connection, thereby enabling external parameter calibration of the camera and lidar and achieving time synchronization between the two.
[0054] In some specific embodiments of the present invention, the lidar system utilizes the new Livox Avia solid-state lidar, which boasts lightweight and compact dimensions, making it easily mountable on mobile robots. Furthermore, the Avia solid-state lidar features a non-repeating scanning mode, resulting in a petal-shaped point cloud. Its field of view coverage significantly increases over time, with the longer the accumulation time, the denser the point cloud. Its maximum detection range reaches 450 meters. In non-repeating scanning mode, the Avia solid-state lidar achieves a FoV of up to 70.4° × 77.23°, while the vertical FoV of traditional mechanical multi-line lidars typically does not exceed 40°. The Avia's large FoV allows for the acquisition of point clouds covering a larger scene at once. Its long-range capability allows it to better capture distant details of the environment. Furthermore, due to its high FoV overlap with the camera, it retains more laser information during fusion, achieving better fusion results. Furthermore, the Avia lidar incorporates a built-in BMI088 inertial measurement unit, eliminating the need for a separate IMU sensor. Furthermore, Livox officially provides comprehensive ROS driver support for easy software development. In some specific embodiments of the present invention, the camera used in these embodiments is an MVSUA133GC color industrial camera, which supports adjustment of multiple camera parameters such as brightness, contrast, and frame rate, and also supports soft triggering to facilitate time synchronization with the LiDAR. It should be noted that the methods of these embodiments are also applicable to other combinations of LiDARs and industrial cameras that meet the requirements of a large forward-looking FoV.
[0055] In one embodiment, referring to FIG2 , the data acquisition platform of the method of the present invention is provided with a set of mobile robot data acquisition experimental platform and a set of handheld device data acquisition platform. Furthermore, the mobile robot data acquisition experimental platform is provided with a sensor bracket, on which necessary equipment such as a power supply, a sensor, and an industrial computer are placed. Referring to FIG1 , the mobile robot acquisition experimental platform of the embodiment of the present invention includes a sensor bracket, a camera, and a laser radar. Four pulleys are connected to the lower side of the sensor bracket. The sensor bracket is a multi-layer bracket for placing equipment such as a power supply, a sensor, and an industrial computer. The laser radar is arranged on the top surface of the sensor, and the camera is arranged on the side of the laser radar and is rigidly connected to the laser radar. The camera is rigidly fixed to the laser radar through a 3D printed structural part, so that the camera and the radar have as large a common viewing area as possible to facilitate the association of point cloud and image information.
[0056] Furthermore, for the rapid data collection work in some indoor scenes, see Figure 3, the method of the present invention adopts a small handheld device data collection platform to achieve portability of data collection. Specifically, the handheld device data acquisition platform includes a sensor bracket, a camera and a laser radar. The sensor bracket includes a mounting plate, a connecting frame and a handle. The connecting frame and the handle are both connected to the bottom surface of the mounting plate, and the connecting frame is connected to the middle of the mounting plate. The inner cavity of the connecting frame and the top surface of the mounting plate are used to place the equipment. The laser radar is arranged on the top surface of the sensor, and the camera is arranged on the side of the laser radar and is rigidly connected to the laser radar. The camera is rigidly fixed to the laser radar through a 3D printed structural part.
[0057] 1 to 9 , in some embodiments, the method for filtering dynamic objects from a point cloud based on camera-lidar fusion according to the present invention includes at least the following steps:
[0058] S100, acquiring a real-time point cloud through a laser radar, and synchronously acquiring a real-time image through a camera; wherein the camera and the laser radar are pre-calibrated with external parameters using an external parameter automatic calibration method;
[0059] S200, obtaining a dynamic object detection frame in the image through a deep learning-based visual target detection network;
[0060] S300, projecting the point cloud into the image according to the extrinsic parameters obtained by pre-calibration;
[0061] S400 , after filtering out the point cloud within the dynamic object detection frame by combining the point cloud projected onto the image, the point cloud within the dynamic object detection frame is clustered and segmented and then filtered out.
[0062] The present invention is based on a method for filtering dynamic objects from point clouds using a camera-lidar fusion method. This method uses a combination of two different sensors, a camera and a lidar, to filter dynamic objects from point clouds. This method can improve the positioning accuracy of autonomous mobile robots in dynamic scenes and effectively mitigate the impact of dynamic objects on mapping. The present invention uses lidar and cameras as the smallest sensor units. The lidar model used is the Livox AVIA solid-state lidar, which has a similar FoV to that of a camera. By performing target detection on images, the method avoids using deep learning algorithms to directly detect point clouds, reducing computing power requirements. The visual target detection algorithm employed can be lightweight and deployed on the computing platform carried by the robot. The entire real-time dynamic point cloud filtering process can achieve high computational efficiency, and can therefore be conveniently incorporated into the laser SLAM algorithm as a preprocessing process, effectively mitigating the phenomenon of dynamic objects producing ghosting during mapping.
[0063] In one embodiment, the method of the embodiment of the present invention adopts a targetless external parameter automatic calibration method. Before performing information fusion, the external parameter calibration is performed on the sensor data collected in their respective coordinate systems to convert the data into the same coordinate system for processing, thereby realizing the joint calibration of the camera and the lidar. The traditional calibration method is to obtain the external parameters between the camera and the lidar from the mechanical drawings of the print, which is easily affected by the precision error of the print processing and the assembly error, and the measured mechanical external parameters are usually inaccurate. The targetless external parameter calibration method adopted in the embodiment of the present invention can realize robust and accurate external parameter automatic calibration between the lidar and the camera without relying on a specific target (such as a calibration plate), and the calibration result has good accuracy in an environment with rich structural features.
[0064] Specifically, the non-repetitive scanning characteristics of the Livox lidar are first utilized to accumulate the raw data of the lidar for a period of time to form a relatively dense point cloud, and then the point cloud is voxelized into a fixed size. In some specific embodiments, voxelization is usually set to 1m×1m×1m in outdoor scenes. A plane is fitted to each voxel and adjacent plane pairs forming a certain angle are retained. By solving the intersection of the adjacent plane pairs, the depth-continuous point cloud edge cloud features are extracted. In some specific embodiments, the RANSAC algorithm is used to fit the plane for each voxel. Furthermore, in some specific embodiments, the edge features in the image are extracted using the Canny edge detection algorithm commonly used in the field of digital image processing.
[0065] First, the point cloud edge features are projected onto the image based on the initial external parameters:
[0066] Where, L P i Represents a point cloud edge feature point; p i =(u i ,v i ) represents the point obtained by projecting the above point cloud edge feature points onto the image plane; represents the initial extrinsic rotation from the radar coordinate system to the camera system; represents the initial extrinsic translation from the radar coordinate system to the camera coordinate system; π(P) represents the pinhole model of the camera; and f(P) represents the distortion model of the camera. It should be noted that in some specific implementations, the initial extrinsic parameters are typically obtained by manual measurement.
[0067] Then, the edge features of the point cloud and the edge features of the image are further matched. i Find the k nearest image edge feature points around, recorded as a set Calculate the direction vectors of the k image edge feature points according to the following formula:
[0068] In the formula, the mean value of the edge feature pixels of the above image is q i And the covariance matrix is S i . Among them, the covariance matrix S i The eigenvector corresponding to the maximum eigenvalue of is the direction vector of the image edge feature, so that the point cloud edge features and the image edge features can be further screened and matched according to the approximation degree of the direction vector.
[0069] Refer to Figure 4, where blue represents image edge features, red represents point cloud edge features, and green represents the correspondence between image edge features and point cloud edge features. Based on the successfully found matching feature points, the following nonlinear optimization problem is constructed to associate the point cloud edge features with the image edge features at the same time. The nonlinear optimization equation used for joint calibration is expressed as follows:
[0070] Where N represents the number of matches between point cloud edge features and image edge features.
[0071] In the method of the embodiment of the present invention, the extrinsic parameters obtained from the mechanical model are used as the initial values for optimization. By utilizing the automatic differentiation function of the Ceres library, the nonlinear optimization problem can converge quickly. The point cloud is projected onto the image using the extrinsic parameters. The comparison before and after calibration is shown in Figures 5a and 5b.
[0072] In one application embodiment, the camera of the present invention uses a Maideweishi color industrial camera that supports both software and hardware triggering, and an Avia lidar that collects point cloud data at a frequency of 10Hz. Every time the lidar acquires a frame of point cloud data, it sends a trigger signal to enable the camera to capture the image, thereby achieving soft synchronization between the camera and the lidar. Therefore, when each sensor is independently packaged and operates according to its own clock reference and has different sampling frequencies, the method embodiment of the present invention can synchronize the data from each sensor in time. Furthermore, the synchronization time error of the method embodiment of the present invention is within 2ms.
[0073] In one embodiment, the method of the embodiment of the present invention obtains a dynamic object detection frame in the image through a visual target detection network. Taking advantage of the fact that the semantic information contained in color images is richer than that of point clouds, and in the field of computer vision, the research history of deep learning target detection methods for images is longer, and the tool chain for quantization, cutting, and deployment acceleration of neural network models is more mature, the method of the present invention uses a deep learning target detection algorithm (YOLO V5) and deploys it on a computing platform carried by a mobile robot to perform target detection on the image in real time. It has the advantages of both fast speed and high accuracy and is very suitable for engineering applications. Furthermore, the method of the present invention uses a pre-trained model provided by YOLO v5 under the coco dataset, which contains common object categories in life. First, the pre-trained model parameters are exported as an onnx file, and the official Nvidia Tensor RT tool is used to quantize the fp16 model, and finally the model is deployed on the vehicle computing device.
[0074] Furthermore, the mobile robot data acquisition platform constructed in this embodiment of the present invention uses a microcomputer as the computing device. With a network input resolution of 640×640, the deployed model can achieve an inference speed of 250fps, meeting real-time requirements. Furthermore, the pre-trained model provided by YOLO v5 also exhibits good generalization, achieving excellent detection results in multiple experimental scenarios. See Figure 6 for an example of YOLO v5 visual object detection.
[0075] In one embodiment, after obtaining the pixel position of the dynamic object in the image (dynamic object detection box), the method of the embodiment of the present invention needs to find the corresponding point cloud information after projection. It should be noted that the lidar used in the method of the embodiment of the present invention is a Livox AVIA solid-state lidar, whose FOV is 70.4° horizontally and 77.2° vertically. The camera is directly and rigidly fixed to the lidar through 3D printed structural parts, so that the camera and lidar have as large a common view area as possible, which facilitates the association of point cloud and image information.
[0076] Specifically, the precise extrinsic parameters between the camera and the LiDAR are obtained through pre-processing sensor calibration. Any point in a frame of point cloud data collected by the LiDAR is converted into homogeneous coordinates and multiplied successively with the extrinsic parameters between the camera and the LiDAR and the camera intrinsic parameters to obtain the pixel coordinates [uv] projected onto the image plane. T , the calculation process is as follows:
[0077] Where, represents any point in a frame of point cloud data collected by the lidar; represents the external parameter between the lidar and the camera; represents the intrinsic parameter of the camera.
[0078] Furthermore, after projecting the point cloud onto the image, RGB color values are assigned based on their distance. A maximum and minimum threshold are set for the point distance, and JET color mapping is performed on points within the required distance range. Close points appear red, distant points appear blue, and points beyond the maximum distance threshold appear black. The final projection effect is shown in Figure 6. As shown in Figure 6, after time synchronization and extrinsic calibration, the point cloud information collected at the same time can be well associated with the image information.
[0079] In one embodiment, the method of the embodiment of the present invention obtains a dynamic object detection frame in the image through a visual target detection network, combines it with the point cloud projected onto the image to screen out the point cloud within the dynamic object detection frame, and then clusters and segments the point cloud within the dynamic object detection frame and then filters it out. This is beneficial for overcoming the problem of high computational cost of deep learning algorithms for processing point clouds caused by the disorder, sparsity, scale invariance and large scale of point cloud information.
[0080] Specifically, after the prediction of the visual target detection network, the detection frame of the dynamic object can be obtained in the image. Combined with the point cloud projected onto the image, all point clouds belonging to the dynamic object detection frame can be screened out. However, since the detection frame output by YOLO v5 is rectangular, the corresponding point cloud in each image detection frame includes static points belonging to the environment in addition to the dynamic object. If the point cloud belonging to each detection frame is simply filtered out, it will inevitably cause "false positives" of static points. When the number of dynamic objects is too large and occupies most of the sensor's field of view, this "false positives" will cause too few static points, resulting in a reduction in the front-end registration accuracy. The embodiment of the present invention adopts a method of clustering and segmenting the point cloud belonging to the dynamic object detection frame to retain the static points belonging to the environment as much as possible, thereby reducing the "false positives" of static points.
[0081] Furthermore, the method of the embodiment of the present invention adopts the existing DBSCAN clustering algorithm, which is a density-based clustering algorithm that can cluster point clouds distributed in space according to their density. It is insensitive to noise and uses the centroid of the point cloud within the candidate box as a screening condition to segment point clouds that truly belong to dynamic objects. The segmentation effect is shown in Figure 8.
[0082] The present invention experimentally verifies the point cloud dynamic object filtering method based on camera lidar fusion. The point cloud dynamic object filtering algorithm based on camera and lidar fusion of the embodiment of the present invention is lightweight and efficient. After testing, on the NUC11 computing platform, the average processing time per frame of image target detection plus point cloud clustering segmentation is only 10.7ms, which can be easily integrated into the front-end odometer of the mobile robot as a preprocessing step. Further, referring to Figures 9a, 9b, 9c and 9d, it is shown that the beneficial effect of using this method as a SLAM preprocessing step for mapping. According to the figure, after using the algorithm of the present invention as a preprocessing step, the mobile robot can effectively alleviate the ghosting phenomenon caused by dynamic objects when building a map.
[0083] It should be appreciated that the method steps in the embodiments of the present invention can be implemented or executed by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable memory. The method can use standard programming techniques. Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system. However, if desired, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can be run on a programmed application-specific integrated circuit.
[0084] Furthermore, the operations of the processes described herein may be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, by hardware, or a combination thereof. The computer program includes a plurality of instructions that can be executed by one or more processors.
[0085] Further, the method can be implemented in any type of computing platform that is operably connected to a suitable computer, including but not limited to a personal computer, a minicomputer, a mainframe, a workstation, a network or distributed computing environment, a separate or integrated computer platform, or in communication with a charged particle tool or other imaging device, etc. Various aspects of the present invention can be implemented as machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into a computing platform, such as a hard disk, an optical read and / or write storage medium, an RSM, a ROM, etc., so that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the process described herein. In addition, the machine-readable code, or portions thereof, can be transmitted over a wired or wireless network. When such media includes instructions or programs that implement the steps described above in conjunction with a microprocessor or other data processor, the invention described herein includes these and other different types of non-transitory computer-readable storage media. When programmed according to the methods and techniques of the present invention, the present invention can also include the computer itself.
[0086] The computer program can be applied to input data to perform the functions described herein, thereby converting the input data to generate output data that is stored in a non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the converted data represents a physical and tangible object, including a specific visual depiction of the physical and tangible object produced on the display.
[0087] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the aforementioned embodiments. As long as the technical effects of the present invention are achieved by the same means, any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention. Within the scope of protection of the present invention, various modifications and variations of the technical solutions and / or implementation methods are possible.
Claims
1. A point cloud dynamic object filtering method, characterized in that: Based on a data acquisition platform, the data acquisition platform is provided with a camera and a laser radar, and the method comprises the following steps: S100, acquiring a real-time point cloud through the laser radar, and synchronously acquiring a real-time image through the camera; wherein the camera and the laser radar use an external parameter automatic calibration method to pre-calibrate external parameters; S200, obtaining a dynamic object detection frame in the image through a deep learning-based visual object detection network; S300, projecting the point cloud into the image according to the external parameters obtained by pre-calibration; S400 , after filtering out the point cloud within the dynamic object detection frame in combination with the point cloud projected onto the image, the point cloud within the dynamic object detection frame is clustered and segmented and then filtered out.
2. The method according to claim 1, characterized in that The external parameter calibration of the camera and the laser radar comprises the following steps: S110, extracting point cloud edge features from the original point cloud, and extracting image edge features from the original image; S120, matching the point cloud edge features with the image edge features; S130: Construct a nonlinear optimization equation based on the successfully found matching feature points.
3. The method according to claim 2, characterized in that In step S110, the edge features in the original image are extracted using the Canny edge detection algorithm, which includes the following steps: S111, accumulating the original point cloud data acquired by the laser radar for a period of time to form a relatively dense point cloud; S112, voxelizing the point cloud; S113, fitting a plane for each voxel and retaining adjacent plane pairs that form a certain angle therebetween, and extracting point cloud edge features by solving the intersection lines of the adjacent plane pairs.
4. The method according to claim 3, characterized in that: The step S120 includes the following steps: S121, projecting the point cloud edge features onto the image according to the initial external parameters: In the formula, L P i Represents a point cloud edge feature point; p i =(u i , v i ) represents the point obtained by projecting the edge feature points of the above point cloud onto the image plane; represents the initial external parameter rotation from the radar coordinate system to the camera system; represents the initial extrinsic translation from the radar coordinate system to the camera system; π(P) represents the pinhole model of the camera; f(P) represents the distortion model of the camera; S122, through K nearest neighbor search in p i Find the k nearest image edge feature points around, recorded as a set To obtain the direction vectors of the k image edge feature points; wherein the direction vectors of the k image edge feature points are calculated by the following formula: In the formula, the mean value of the edge feature pixels of the above image is q i And the covariance matrix is S i ; The covariance matrix S i The eigenvector corresponding to the maximum eigenvalue of is set as the direction vector of the image edge feature, so as to further filter and match the point cloud edge features and the image edge features according to the approximation degree of the direction vector.
5. The method according to claim 4, characterized in that in, In step S130, the nonlinear optimization equation constructed for joint calibration is expressed as follows: Where N represents the number of matches between point cloud edge features and image edge features.
6. The method according to claim 5, characterized in that in, In step S300, the point cloud is projected onto the pixel coordinates [uv] of the image plane. T Obtained by the following calculation: In the formula, represents any point in a frame of point cloud data collected by the laser radar; represents the external parameter between the laser radar and the camera; represents the internal parameter of the camera. 7 . A computer-readable storage medium having program instructions stored thereon, wherein the program instructions, when executed by a processor, implement the method according to any one of claims 1 to 6 .
8. A point cloud dynamic object filtering system based on camera laser radar fusion, characterized in that: include: A computer device comprising a computer readable storage medium according to claim 7.
9. The point cloud dynamic object filtering system according to claim 8, characterized in that: It also includes a data acquisition platform, which includes a camera and a laser radar, and the camera and the laser radar are connected through a 3D printed structural part.
Citation Information
Patent Citations
Bayesian optimization-based solid-state laser radar and camera self-calibration method
CN115032614A
SLAM (Simultaneous Localization and Mapping) method for eliminating dynamic target by combining vision and laser radar
CN116643291A
Dynamic object filtering method based on 4D millimeter wave radar
CN116978009A
Point cloud dynamic object filtering method based on camera and laser radar fusion
CN117611809A
Multi-sensor object detection fusion system and method using point cloud projection
US11403860B1
Cited By
Navigation method and system for green cutting robot
CN121089710A
Three-dimensional point cloud dynamic object deletion method and three-dimensional reconstruction equipment
CN121353612A
Water surface moving target detection method, recording medium and system
CN121437862A
Camera parameter calibration method and system based on multi-sensor fusion
CN121527194A
Live working method and system based on laser radar and visual identification fusion
CN121722123A