Methods for simplifying a point cloud that represents an environment

Converting point clouds into 3D shapes like cuboids simplifies environmental data processing, reducing computational burden and enhancing collision detection and movement planning efficiency.

DE102024208757A1Pending Publication Date: 2026-03-19ROBERT BOSCH GMBH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102024208757
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Processing large point clouds for environmental information is computationally intensive and time-consuming, which hinders efficient collision detection and movement planning in applications like robotics and autonomous driving.

Method used

A method to simplify point clouds by converting them into a dataset of 3D shapes, such as cuboids, defined by grid cells and heights, reducing data complexity while preserving relevant information for collision calculations.

Benefits of technology

Significantly reduces processing time and computational requirements for collision analysis and movement planning, enabling efficient robot operations and simplified terrain representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method for simplifying a point cloud representing an environment, comprising: providing a data set representing a point cloud based on a coordinate system (400), which has been acquired from the environment, in particular by means of at least one sensor, wherein the coordinate system (400) defines a grid plane (404) with grid cells (410); determining, based on the data set, one or more 3D shapes, wherein each 3D shape (430, 432) has a base (418) and a height (416), the base corresponding to one or more contiguous grid cells of the grid plane, and the height being determined based on points (420) of the point cloud, which correspond to the one or more contiguous grid cells; and providing a simplified data set representing the one or more 3D shapes for further use.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention relates to a method for simplifying a point cloud representing an environment, a computing unit and a computer program for carrying it out, and a robot. Background of the invention

[0002] In various fields, it may be necessary to detect and recognize objects in an environment. This can be done using so-called point clouds, which map the environment including the objects to be detected. Disclosure of the invention

[0003] According to the invention, a method for simplifying a point cloud representing an environment, a computing unit and a computer program for carrying it out, as well as a robot with the features of the independent claims, are proposed. Advantageous embodiments are the subject of the dependent claims and the following description.

[0004] The invention generally deals with the acquisition of an environment using a sensor in order to obtain a point cloud. A typical sensor that can be used to acquire a point cloud representing the environment is a so-called lidar sensor, laser scanner, or even a fringe projection system. Such a point cloud is generally a set of points in 3D, where the points typically correspond to reflection points on surfaces in the environment, i.e., in particular, objects. However, environmental information acquired using other sensors can also be considered a point cloud, e.g., other depth sensors or possibly even a 3D camera, in which individual pixels then possess depth information. Point clouds that have been created and calculated, for example, using software algorithms from individual camera images and / or data acquired by camera sensors are also conceivable.

[0005] The invention will be explained in more detail below, primarily in connection with robots that can be used, for example, in production facilities, logistics, or other systems. Types of such robots, often also referred to as kinematics, include, for example, Cartesian robots, SCARA robots, and articulated robots. Such robots can be used to grasp and / or move objects (i.e., items or parts). A typical application is to remove an object from a box containing, for example, a large number of (identical or different) objects and, for example, place it somewhere or possibly assemble it. For this purpose, such a robot has, for example, a gripper or, more generally, an end effector. Such robots can also be referred to as manipulators or manipulation devices; the grasping, possibly with movement, can also be described as manipulation.

[0006] In order for the robot to grasp and / or move an object in a work environment, e.g. in a box, or in or on other carriers (e.g. a container or a pallet) using the end effector, the robot, in particular its end effector, must perform a movement sequence to move the end effector to the position of the object and then, e.g. after the object has been grasped, to move the end effector away from the position of the object again, e.g. to a desired storage position.

[0007] For a successful grip, the exact pose for the gripper – hereinafter also referred to as the gripping pose – can be determined beforehand. Likewise, the movement sequence should be defined in such a way that the end effector or any other part of the robot does not touch any other objects in the environment during execution. For this purpose, environmental information, such as the aforementioned point cloud of the environment (scene; in the context of robots, this can also be referred to as a work environment), and especially of any existing support structure, such as a box, can be acquired and provided. Sensors such as the aforementioned lidar sensors (laser scanners) or other depth sensors can be used for this. The environmental information then specifically includes one or possibly several objects located in the work environment.

[0008] Based on environmental information, such as a point cloud, a gripping pose can be determined. From this, the robot's movement sequence can be defined, allowing for the calculation of control information for robot movement, which can then be implemented by a control unit or other processing unit. However, a movement sequence can also be determined without specifying the gripping pose.

[0009] The sensor data or environmental information is provided as a point cloud or a data set; this environmental information then typically requires further processing to capture surfaces and disturbance contours of the environment. Using this information, for example, data models of complex algorithms can be updated to calculate evasive maneuvers, part positioning, or obstacle avoidance. However, processing the point clouds is computationally intensive and time-consuming due to the typically large amount of data.

[0010] As mentioned earlier, there are other use cases where the environment and objects within it are captured and represented as a point cloud, and these problems also arise in these cases. For example, in 3D games, this includes simplifying terrain and collision detection. In navigation, route planning, and autonomous driving, the goal might be to simplify the terrain or its representation. In particular, applications are relevant where point clouds represent terrain, relief, topography, surfaces (or combinations thereof) that then need to be simplified. This is especially relevant for collision detection, as it reduces the amount of data and simplifies the calculations.

[0011] Against this background, the present invention proposes a simplification of the point cloud that depicts an environment. Simplification here refers in particular to a reduction in the amount of data, while preserving as much of the content relevant to the application as possible.

[0012] It should be mentioned again at this point that this type of simplification of a point cloud can be applied not only in the specific field of robotics, but also in other areas where such point clouds are processed, such as autonomous (or automated or semi-automated) driving or the other application areas mentioned.

[0013] This involves providing a dataset that represents a point cloud based on a coordinate system, captured from the environment, particularly by at least one sensor. The coordinate system allows coordinates to be assigned to individual points. For illustrative purposes, a Cartesian coordinate system will be used below; a point is then defined by its x, y, and z coordinates (i.e., as a translational vector relative to the origin). However, it goes without saying that other coordinate systems can also be used.

[0014] The coordinate system defines a grid plane with grid cells. In the case of the Cartesian coordinate system, the grid plane can be, for example, the xy-plane; the grid plane with grid cells can also be referred to as a grid. The grid cells themselves could then, for example, have the shape of squares with a side length of 1 in the x-direction and 1 in the y-direction (possibly with a corresponding scale). It should be understood that this is only an example. Even in the Cartesian coordinate system, the shape of the grid cells does not necessarily have to be square, but can also be rectangular, for example. In other coordinate systems, such a grid cell can also have the shape of a parallelogram or another shape altogether.

[0015] Based on the dataset, one or more 3D shapes are then determined, each with a base and a height. The base corresponds to one or more contiguous grid cells of the grid plane. In the example mentioned above, the base would therefore be, for instance, square or rectangular. In the case of multiple contiguous grid cells, the base could also be something other than square or rectangular, although the use of uniform bases is preferred, as will be explained in more detail later.

[0016] The height is determined based on points in the point cloud that correspond to one or more contiguous grid cells when projected onto the grid plane, meaning the projection onto the grid cell(s) lies within the relevant grid cell(s). The coordinate system assigns each point a height relative to the grid plane; in the aforementioned Cartesian coordinate system, this is, for example, the value of the z-coordinate. The specific value for the height can then be, for example, the highest value of the relevant points or an average value (e.g., an arithmetic mean).

[0017] The base area and height, in the coordinate system, allow for the definition of a unique 3D shape. When using a Cartesian coordinate system, the 3D shapes can then be, in particular, cuboids. However, with non-rectangular coordinate axes, they can also be, for example, parallelepipeds.

[0018] Unlike the large number of individual points (from the point cloud), the amount of data required to represent the 3D shapes is significantly smaller, since a certain number of points are combined to form a 3D shape, such as a cuboid. This simplified representation is sufficient, especially for recognizing objects in the environment, for example, regarding the rendering of object contours, while the necessary processing power and time are considerably lower.

[0019] In particular, this method greatly simplifies collision calculations; for example, a movement sequence for a device such as the aforementioned robot can be determined, with the simplified data set being used to prevent collisions between the device and an object. In this case, a highly detailed representation of the environment is either unnecessary or only minimally important.

[0020] The proposed approach is primarily intended to simplify surface topologies; undercut volumes are not considered. As mentioned, the point cloud is preferably simplified using cuboids representing the surface. A cuboid representation of objects is particularly advantageous for calculating collisions between objects, as the computation times are significantly shorter than those for calculating more complex objects such as 3D meshes. This simplification thus considerably reduces the processing time for collision analysis.

[0021] The simplified dataset, representing one or more 3D shapes, is then made available for further use. This could be, for example, to determine a movement sequence for the robot to grasp and / or move an object in the environment using an end effector.

[0022] In one embodiment, one or more 3D shapes are determined by defining several individual 3D shapes based on the dataset. For each individual 3D shape, a base area and a height are determined, where the base area corresponds to a grid cell of the grid plane, and where the height is determined based on points in the point cloud, which correspond to the grid cell by projection. In other words, corresponding 3D shapes (also called individual 3D shapes), such as cuboids, are initially determined based only on individual grid cells.

[0023] One or at least one of the multiple 3D shapes is then determined based on at least two related individual 3D shapes. This allows the data set to be reduced even further. Alternatively, at least one of the individual 3D shapes can be used as the one or at least one of the multiple 3D shapes. This is the case, for example, when certain individual 3D shapes cannot be combined with at least one other.

[0024] Determining one or at least one of several 3D shapes, based on at least two connected individual 3D shapes, can be done in particular as follows. For at least some of the individual 3D shapes, it is checked whether two or more individual 3D shapes, whose heights are within specified tolerances, are connected. The tolerances can be chosen as needed, e.g., as absolute or relative values. If this is the case, one or at least one of the several 3D shapes is determined based on the connected individual 3D shapes. The height of a 3D shape generated from two or more connected individual 3D shapes can be, for example, an average of the heights of the individual 3D shapes or the highest height of the individual 3D shapes.

[0025] In this way, for example, each individual 3D shape can be checked to see if it can be combined with one or more other related individual 3D shapes. However, it is advantageous if several related individual 3D shapes are only combined into one 3D shape if the (then new) 3D shape corresponds to a predefined type of 3D shape, in particular the same type as the individual 3D shape. If the individual 3D shapes are cuboids, it can be stipulated that the (new or combined) 3D shapes must also be cuboids. Such 3D shapes can be represented with just a few data points, such as the height and two points of the base, while more complex shapes, such as L-shapes or the like, require more data for representation.

[0026] A computing unit according to the invention (i.e., generally a system for data processing), e.g., a control unit or a control unit of a robot, or a central server or other computing system, is, in particular in terms of programming, equipped to carry out a method according to the invention.

[0027] The invention also relates to a robot configured to receive control information as described above. In addition, or alternatively, the robot comprises a computing unit according to the invention. Furthermore, the robot comprises, in particular, a control unit and a drive unit for moving the robot. The robot may also have at least one sensor for acquiring environmental information, e.g., a camera and / or a lidar sensor.

[0028] Implementing a method according to the invention in the form of a computer program or computer program product with program code for carrying out all method steps is also advantageous, as this incurs particularly low costs, especially if an executing control unit is already available for other tasks. Finally, a machine-readable storage medium is provided with a computer program stored on it as described above. Suitable storage media or data carriers for providing the computer program are, in particular, magnetic, optical, and electrical storage media, such as hard drives, flash memory, EEPROMs, DVDs, etc. Downloading a program via computer networks (Internet, intranet, etc.) is also possible. Such a download can be wired or wireless (e.g., via a WLAN network, a 3G, 4G, 5G, or 6G connection, etc.).

[0029] Further advantages and embodiments of the invention will become apparent from the description and the accompanying drawing.

[0030] The invention is schematically illustrated in the drawing using an exemplary embodiment and is described below with reference to the drawing. Brief description of the drawings Fig. Figure 1 schematically shows a robot to illustrate the invention. Fig. Figure 2 schematically shows a container to illustrate the invention. Fig. Figure 3 schematically shows the sequence of a procedure in one embodiment. Fig. 4a and Fig. Figure 4b schematically shows details to explain the invention. embodiment(s) of the invention

[0031] In Fig. Figure 1 schematically and exemplarily illustrates a robot 100 to explain the invention. By way of example, the robot 100 has stand and arm components 102, 104, 106, so-called axes, which are each movably and movably connected by means of joints 112, 114.

[0032] Furthermore, the robot 100 has an end effector 108, e.g., a gripper. The end effector 108 is movably and reversibly connected to the arm component 106 by means of a joint 116.

[0033] Furthermore, the robot 100 has a drive system 120, shown only schematically here, as well as a computing unit 122 designed, for example, as a control or regulation unit. This allows the drive system 120 to be controlled, for example, using control information, in order to move the robot according to a desired sequence of movements. This can include, for example, moving the axes relative to each other by means of the joints, but also rotating the axes themselves, provided that appropriate drives are available.

[0034] It should be noted that the robot 100 is only used here as an example for illustrative purposes. A robot for grasping and / or moving objects in containers can also be designed differently, for example, using an end effector that can only move linearly along several different rails.

[0035] Furthermore, a box 132 is shown in a work environment 130, containing, for example, an object 140. The robot 100 can now be controlled, for example, in such a way that it grasps and / or moves the object 140 using the end effector 108, in particular also taking it out of the box and, for example, placing it elsewhere.

[0036] Furthermore, an example sensor 124 is shown, which could be, for example, a camera or a lidar sensor. Both types of sensors can also be used. Likewise, several identical sensors can be used. As already mentioned, the use of a lidar sensor or another time-of-flight sensor (TOF sensor) or a 3D fringe projection scanner is particularly useful for obtaining a point cloud as environmental information. However, it is also conceivable to capture 3D and / or 2D images using a camera. It is also conceivable to capture only 2D images using a camera. The sensor 124 can, for example, be arranged in a suitable way in the work environment, e.g., on a ceiling and thus separately from the robot 100. The sensor 124 could, however, also be part of the robot and, for example, be arranged on the arm component 106 or the end effector 108.

[0037] The sensor 124 can now detect the working environment 130, and in particular the box 132 and its interior, including the object 140. Based on the environmental information and sensor data obtained in this way, a movement sequence can be created for the robot to grasp and / or move the object 140 using the end effector 108.

[0038] In Fig. 2 is a box 232, comparable to box 132 according to Fig. Figure 1 shows examples of various objects 240, 241, 242, 243, and 244. A typical task for a robot is to remove as many objects as possible from the box. It is advisable to first remove the easiest or safest object to grasp, and so on, while adhering to certain guidelines, such as not covering forbidden areas with the end effector. It can also be seen that the objects can differ from one another, for example, being cuboid or cylindrical, and that the objects can also be randomly ordered.

[0039] An example movement sequence is shown for grasping object 240 and removing it from the box. This includes an approach path, or first part of the movement sequence 251, by which the end effector moves towards object 240 in order to grasp it. Additionally, a retraction path, or second part of the movement sequence 252, is shown, along which the end effector can be moved with object 240 out of the box.

[0040] Together, the first partial movement sequence 251 and the second partial movement sequence 252 form a complete movement sequence. It can be seen that the movement sequence is such that the end effector must be in a specific pose – a grasping pose – in order to properly grasp the object 240.

[0041] To find, for example, the optimal gripping pose, but also a collision-free movement sequence, it is useful to capture the environment, i.e., the box containing the objects, using the sensor. The resulting point cloud, as indicated by 270, can then be simplified to find the gripping pose as quickly and computationally efficiently as possible. This simplification of the point cloud will be explained in more detail below.

[0042] In Fig. Figure 3 schematically illustrates the sequence of a procedure in one embodiment. It also refers to the Fig. 4a and Fig. Reference is made to section 4b, which shows a container or box and objects, and which will explain the various steps in more detail.

[0043] In step 300, a data set 302 is provided, which, based on a coordinate system, represents a point cloud that has been captured from the environment by at least one sensor. The coordinate system defines a grid plane with grid cells. Such a point cloud is in Fig. 2 with 270 shown for box 232 with the objects contained therein.

[0044] In Fig. Figure 4a shows, for further explanation, a coordinate system 400 with origin 402, specifically a Cartesian coordinate system with axes x, y, z. A grid plane defined by the coordinate system 400 is labeled 404 and corresponds here to the xy-plane. The grid plane 404 has grid cells, one of which is labeled 410. The grid cells have, for example, a length of 412 in the x-direction and a length (or width) of 414 in the y-direction. These could, for example, each be a unit length.

[0045] Furthermore, in Fig. Figure 4a shows a set of 420 points, which may be part of a point cloud acquired by a lidar sensor. The coordinate system allows x, y, and z coordinates to be assigned to each point; a corresponding vector for a point is denoted by 422 as an example.

[0046] In step 310, one or more 3D shapes, e.g., cuboids, are determined based on the dataset. Specifically, in step 320, several individual 3D shapes, e.g., also cuboids, are determined based on the dataset. For each individual cuboid, a base area and a height are determined, where the base area corresponds to a grid cell of the grid plane, and where the height is determined based on points in the point cloud, which correspond to the grid cell by projection.

[0047] This is in Fig. Figure 4a shows an example of a cuboid 430. The base 418 of the cuboid corresponds to a grid cell of the grid plane, here, for example, the grid cell in the lower left. A vertex of this grid cell, or base, is labeled 432. The length and width of the base correspond to the length 412 and width 414 of the grid cell, respectively. The height of the cuboid 430, labeled 416 here, is determined from the points 420 shown, which are part of the (not shown here) point cloud that, by projection (into the grid plane or xy-plane), correspond to the grid cell serving as the base. The specific value for the height 416 could, for example, be the arithmetic mean of the z-values ​​of these points or the highest z-value of these points.

[0048] Examples include: Fig. 4a shows further cuboids that can be determined in the same way. For example, a cuboid can be determined for each grid cell in which (when projected) at least one point of the point cloud lies. However, it is also conceivable that this is only done for grid cells in which (when projected) at least two, or at least five, or generally a predetermined minimum number of points lie.

[0049] In step 330, it is then checked, for at least some of the individual 3D shapes or cuboids, whether two or more individual 3D shapes or cuboids, whose heights match within specified tolerances, are connected.

[0050] This is in Fig. 4b illustrates, in which the coordinate system 400 is again used. Fig. 4a and, besides the individual cuboid 430, some other cuboids are shown (these other cuboids do not correspond to those from Fig. 4a).

[0051] For this purpose, one can start, for example, with the individual cuboid 430 or the corresponding grid cell to check whether adjacent individual cuboids have a height within the tolerances. The height of the individual cuboid 430 is again shown as 416, with a minimum height of 416.1 and a maximum height of 416.2 shown as tolerances. Depending on the scaling and application, various values ​​or ranges are possible. While this might be in the (possibly double-digit) centimeter range in the field of autonomous driving, it might be in the (possibly sub-) millimeter range for handling technology (e.g., the described gripping of objects).

[0052] The individual cuboid 430 is also labelled 1,1 here as an example. Based on this, further individual cuboids are labelled 1,2, 1,3, 1,4 along the y-axis, and 2,1, 3,1 along the x-axis. The other individual cuboids shown are labelled 2,2, 3,2, 3,2, 3,3. The adjacent individual cuboids of individual cuboid 430 (or 1,1) are initially individual cuboids 1,3, 2,2, and 2,1. Their heights are within the specified tolerances. Individual cuboids 1,1, 1,2, 2,1, and 2,2 can thus be combined to form a (new) cuboid 432 (only indicated here).

[0053] The height of individual cuboid 1.3, for example, would still be within the tolerances, while the heights of individual cuboids 1.4, 2.3, 3.1, 3.2, and 3.3 would not. Therefore, cuboid 432 will not be extended further, because simply adding individual cuboid 1.3 would result in a cuboid that is no longer complete. However, it is conceivable, for example, that individual cuboids 2.3 and 3.3 could be combined into a (new) cuboid.

[0054] It would also be conceivable that, for example, the individual cuboids 1,1, 1,2 and 1,3 could be combined into one (new) cuboid, and likewise the individual cuboids 2,1 and 2,2.

[0055] To decide which individual cuboids to combine, one can, for example, start with an individual cuboid (here cuboid 320 or 1.1) and check every possibility of forming a (new) cuboid. The variant that combines the most individual cuboids can then be chosen. It is also conceivable to choose the variant that creates the largest contiguous base (if the grid cells are not all the same size, this could lead to a different result).

[0056] This can be carried out until it has been checked for each individual cuboid whether it can be combined with at least one other individual cuboid.

[0057] In step 332, one or at least one of the several 3D shapes or cuboids are then determined based on the connected individual 3D shapes or cuboids.

[0058] Individual 3D shapes or cuboids that have not been combined with at least one other can then be used directly as a 3D shape or cuboid, step 322.

[0059] In step 340, a simplified data set 342, representing one or more 3D shapes, is provided for further use. In step 350, for example, a movement sequence for the robot can then be determined, such as with regard to... Fig. 2 already explained, in order to be able to move the robot according to control information determined in step 360.

Claims

[1] Methods for simplifying a point cloud representing an environment, comprising: Providing (300) a data set (302) which represents a point cloud based on a coordinate system (400) which has been detected in particular by means of at least one sensor (124) from the environment (130), wherein the coordinate system (400) defines a grid plane (404) with grid cells (410); Determine (310), based on the data set, one or more 3D shapes (430, 432), wherein each 3D shape has a base area (418) and a height (416), wherein the base area corresponds to one or more connected grid cells of the grid plane, and wherein the height is determined based on points of the point cloud, which points correspond to the one or more connected grid cells by projection; Providing (340) a simplified data set representing one or more 3D shapes for further use. [2] Method according to claim 1, wherein the determining (310), based on the data set, comprises one or more 3D shapes: Determine (320), based on the dataset, several individual 3D shapes (430), wherein for each individual 3D shape a base area (418) and a height (416) are determined, the base area corresponding to a grid cell of the grid plane, and the height being determined based on points of the point cloud which correspond to the grid cell by projection; and Determining one or at least one of the several 3D shapes, based on at least two connected individual 3D shapes, and / or using (322) at least one of the individual 3D shapes as the one or at least one of the several 3D shapes. [3] Method according to claim 2, wherein determining one or at least one of the several 3D shapes, based on at least two connected individual 3D shapes, comprises: Check (330) for at least some of the individual 3D shapes whether two or more individual 3D shapes, whose heights are within specified tolerances, are connected, and If that is the case, determine (332) one or at least one of the several 3D shapes based on the related individual 3D shapes. [4] Method according to claim 3, wherein the multiple 3D shapes are of the same type. [5] Method according to one of the preceding claims, wherein one or each of the several 3D shapes corresponds to a parallelepiped, in particular a cuboid. [6] Method according to one of the preceding claims, for determining a movement sequence for a device, wherein the simplified data set is used to avoid a collision of the device with an object. [7] Method according to claim 6, for determining the movement sequence for a robot (100) as a device to move an end effector of the robot to an object in the environment and to grasp and / or move the object by means of the end effector (108), comprising: Determine (350) the motion sequence based on the simplified data set; and Providing the movement sequence. [8] The method of claim 7, further comprising: Determine (360°), based on the movement sequence, control information for moving the robot, and Providing control information and / or moving the robot based on the control information. [9] Computing unit (122) comprising means for carrying out the method according to any of the preceding claims. [10] Robot (100) which is configured to receive control information determined by a method according to claim 8, and / or with a computing unit according to claim 9, and with a drive system and a control or regulation unit for controlling the drive system, with an end effector having at least two gripping means for gripping and / or moving an object, and preferably with at least one sensor, in particular a camera, for capturing environmental information of an environment. [11] Computer program comprising instructions which, when the program is executed by a computer, cause it to execute the method according to claims 1 to 8. [12] Computer-readable storage medium on which the computer program according to claim 11 is stored.

Citation Information

Patent Citations

  • Method and device for estimating directional information conveyed by a free-space gesture to determine user input at a human-machine interface

    DE102018121317A1