Computer-implemented method for estimating scene flow for dynamic scenes from point clouds
The method addresses the inefficiency and inaccuracy of current scene flow estimation techniques by classifying points as static or dynamic without clustering, enabling more accurate and efficient scene flow estimation in dynamic environments.
Patent Information
- Application Number
- PCT/DE2024/100961
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-22
AI Technical Summary
Current methods for scene flow estimation in dynamic scenes from point clouds require object analysis through clustering, which is inefficient and less accurate for non-rigid motions.
A computer-implemented method that classifies three-dimensional points as static or dynamic without the need for object analysis, using a processing unit with segmentation, self-motion, and scene flow estimation units to generate scene flow vectors for both static and dynamic points.
This method enables more accurate estimation of scene flow, particularly for non-rigid motions, by processing environmental data in real-time without the need for clustering, thus improving the efficiency and accuracy of scene interpretation in dynamic environments.
Smart Images

Figure DE2024100961_22052025_PF_FP_ABST
Abstract
Description
[0001] COMPUTER-IMPLEMENTED SCENE FLOW ESTIMATION METHOD FOR DYNAMIC SCENES FROM POINT CLOUDS
[0002] TECHNICAL FIELD
[0003] The invention disclosed here lies in the technical field of scene flow estimation for dense, dynamic scenes from point clouds using self-motion, segmentation, and neural networks. The invention is aimed at accelerated scene interpretation, particularly for use in vehicles.
[0004] BACKGROUND
[0005] Current assistance systems for motor vehicles use environmental data to alert the driver to certain events outside the vehicle. This environmental data particularly includes three-dimensional data recorded by sensors on the vehicle. The three-dimensional data is available, for example, as a point cloud, where each point is part of an object outside the vehicle. Such points can be recorded, for example, using Lidar (light detection and ranging) sensors. The points lie in a coordinate system whose origin is the respective sensor. If multiple sensors are used, the coordinates of the sensors can be referenced to a common origin, for example a point lying between the sensors or any selected sensor. The axes of the point cloud coordinate system can point in the direction of travel and in the opposite direction, as well as vertically upwards and at right angles to the direction of travel.In general, a coordinate system that corresponds to a predefined vehicle coordinate system can be used for the point cloud. The acquisition and processing of such environmental data ideally takes place in real time, because the data is subject to extensive movement and because immediate presentation of results to the driver is desired. Therefore, the data is processed by devices embedded in the vehicle rather than by transmitting it to external devices.
[0006] The movement to which the data is subject includes, firstly, the movement of objects relative to the vehicle due to the vehicle's own movement, secondly, the movement of objects relative to the vehicle due to the objects' own movement, and thirdly, the movement of objects without them changing their position. The movement of objects relative to the vehicle due to the vehicle's own movement affects all recorded objects, since all data is measured from the moving vehicle. The movement of objects relative to the vehicle due to the objects' own movement affects other objects, such as vehicles or pedestrians, that are also moving. The movement of objects that do not change their position affects things like trees and wind turbines that move in the wind, or people and animals that can remain in one place and still move. The types of movement listed can overlap or accumulate.In general, data movement includes any movement in the scene or of the recording sensor that results in a relative change in the position of a point with respect to the sensor or a chosen origin of a vehicle's coordinate system.
[0007] In order to classify the movement of objects, methods are known in the prior art in which measured points are first assigned to objects, for example by clustering. The objects are then viewed as a whole and subjected to a motion analysis. For this purpose, point clouds can be measured at different points in time. Each point cloud is subjected to object analysis (clustering). Subsequently, corresponding objects from multiple point clouds are identified. Only then are changes in their position and / or shape determined and their movement analyzed. In contrast to the prior art, the invention estimates the scene flow for all dynamic points individually. Object analysis (clustering) is not necessary. Therefore, a more accurate estimation is possible for non-rigid movements. With rigid movements, all points of an object undergo the same Euclidean transformation.Non-rigid movements allow deformations within an object, as is typical for articulated objects, for example in the movement of a pedestrian.
[0008] SUMMARY
[0009] The invention disclosed here relates to a computer-implemented method, comprising: receiving environmental data by a processing unit, wherein the environmental data contains three-dimensional points; classifying, by a segmentation unit of the processing unit, the three-dimensional points as static or dynamic points; creating, by a self-motion unit of the processing unit, a transformation matrix for the static points; converting, by a conversion unit of the processing unit, the static points into scene flow vectors based on the transformation matrix; estimating, by a scene flow estimation unit of the processing unit, scene flow vectors for the dynamic points; and outputting the scene flow vectors of the static and dynamic points.
[0010] Embodiments of the invention further relate to a device, comprising: a processing unit configured to receive environmental data of the device, wherein the environmental data contains three-dimensional points; a segmentation unit of the processing unit, wherein the segmentation unit is configured to classify the three-dimensional points as static or dynamic points; an ego-motion unit of the processing unit, wherein the ego-motion unit is configured to create a transformation matrix for the static points; a conversion unit of the processing unit configured to convert the static points into scene flow vectors based on the transformation matrix; a scene flow estimation unit of the processing unit configured to estimate scene flow vectors for the dynamic points;and a unit configured to output the scene flow vectors of the static and dynamic points. Further embodiments relate to a computer-readable medium having instructions stored thereon that, when executed by a processor, perform one of the methods disclosed herein.
[0011] SHORT DESCRIPTION OF THE CHARACTERS
[0012] Figure 1 shows a method according to the invention.
[0013] Figure 2 shows a device according to the invention.
[0014] Figure 3 shows a data model for embodiments of the invention.
[0015] DETAILED DESCRIPTION
[0016] The invention comprises the provision of environmental data comprising three-dimensional points from the surroundings of a device, wherein the device can be attached or arranged in or on a motor vehicle, a bicycle, or even the clothing of a person. The environmental data can be acquired by one or more sensors, for example lidar sensors or radar or other sensors for measuring three-dimensional scenes, and are received by a processing unit. This processing unit can be implemented in the form of a suitable hardware circuit or as a process running on a processor. The processor can be part of the on-board electronics of a car or bicycle or can run on a mobile device, such as a smartphone or notebook.
[0017] The processing unit has a segmentation unit, which can also be implemented as a hardware circuit or as an additional process. The segmentation unit serves to segment the environmental data into static and dynamic points. A static point is a point whose perceived position changes due to the device's own movement, but which itself does not move. A dynamic point undergoes the same position change but also exhibits actual movement.
[0018] The result of the segmentation unit is preferably generated as a mask. The mask is preferably implemented as a one-dimensional vector that contains as many components as the point cloud has points. Alternatively, multidimensional data structures can be used to represent the mask, such as a multidimensional array that contains as many dimensions as the point cloud. Each component of the vector or data structure encodes whether the associated point with the same index is static or dynamic. For example, a dynamic point can be encoded with the value one and a static point with the value zero, or vice versa. The segmentation unit is preferably designed as a neural network, but implementations with conventional classifiers are also possible. The distinction between static and dynamic points can be made based on multiple point clouds that were recorded at different times.For example, neighborhoods can be analyzed for the points of several point clouds. Based on these neighborhoods, correspondences between points in different point clouds can be determined and their movements can be classified into two types of movement—static and dynamic. Alternatively, a distinction can be made between static and dynamic points based solely on the point cloud at the same time point. For example, the neighborhoods of points often exhibit features that indicate the presence of more or less movement. For example, based on neighborhoods, a blurriness of points can be determined that differs from other points and therefore suggests the presence of blurriness due to movement.
[0019] The processing unit further comprises a self-motion unit, which can also be configured as a circuit or as a process. The self-motion unit determines a transformation matrix for all points, or alternatively only for the static points, in order to determine the movement of the static points relative to the sensor(s). This movement is generally a three-dimensional Euclidean transformation that includes translation and rotation. Preferably, the three-dimensional points comprise two or more point clouds recorded at different times, and the classification is performed separately for each of these point clouds. In this case, the transformation matrix can be determined using known methods that analyze the change in the static points in these point clouds.
[0020] The transformation matrix is applied to the static points by a conversion unit, which is also implemented as a circuit or as a running process. The resulting points of this application, combined with the static points, generate a vector field of scene flow vectors that models the self-motion of the vehicle or person. The conversion unit is also a component of the disclosed device or is connected to it.
[0021] The processing unit also has a unit for estimating the scene flow of the dynamic points. This measure can also be performed using two or more point clouds measured at different times. The estimation includes, for example, an analysis of the dynamic points or both the dynamic and static points from at least two point clouds, whereby corresponding points in the point clouds are determined. The corresponding points can be determined based on technical characteristics of the points and / or based on a cross-correlation of two or more point clouds. The characteristics can include neighborhoods of the points within the same point cloud, their three-dimensional positions, their intensities, and / or their colors. Points in different point clouds correspond to each other if their technical characteristics are similar.This similarity can be determined based on a correlation between the points in both point clouds and a cost function. Pairs of corresponding points are stored as scene flow vectors.
[0022] Finally, the scene flow vectors of the static and dynamic points are output, for example, in a graphic displayed on a screen. The dynamic points can be highlighted compared to the static points, for example, using a different color and / or intensity. The steps described are preferably implemented using neural networks that are trained prior to the application described above. Training is performed using training data containing multiple point clouds from different points in time. The training data also includes a predefined segmentation, for example, in the form of a mask.
[0023] For example, the segmentation unit can comprise a first and a second neural network. The first neural network extracts the technical features of a point cloud already described. For example, the technical features include neighborhoods of points. For example, for each point, a number of other points in the same point cloud that lie within a certain radius can be determined, or a predefined number of nearest points can simply be determined. The second neural network segments the points based on the features. For example, areas of points can be determined whose geometry indicates dynamic or static situations. For example, vehicles or people usually have typical geometric shapes; their dynamics can be inferred based on the position of such entities.For example, a vehicle detected to the left of the recording sensor and traveling in the opposite direction to the recording device can be classified as dynamic. Alternatively, dynamic or static properties can be determined by a correspondence analysis of points in two point clouds from different points in time.
[0024] The classification / segmentation of the points can be repeated sequentially at multiple resolutions of the environment data. This step can, for example, be performed by the aforementioned second neural network based on the neighborhoods produced by the first neural network. The different resolutions can be created by multiple coarsening of the environment data, for example by summarizing, in particular averaging, several neighboring points. Alternatively, different point densities can be used for each resolution. The classification preferably begins with the coarsest resolution and is repeated with progressively finer resolutions. Training the neural networks involves comparing the output of the second neural network with the specified segmentation. The result of the comparison is returned to the first and / or second neural network via backpropagation.The networks repeat the segmentation with modified weightings until a satisfactory result is achieved. Such a result can be achieved if the deviations from the specified segmentation are only minor, for example, if they are below a specified threshold. During training, the backpropagation to the first neural network can be partially interrupted by a zero gradient so that it receives its training signal only from the second neural network. By selectively intervening in the backpropagation, the first and second neural networks are optimized for point classification.
[0025] In one embodiment, the above segmentation is performed in parallel on multiple point clouds measured at different times. This parallel processing initially requires two neural networks, the first of which determines the neighborhoods of points, and the second of which performs the actual segmentation at different resolutions. During training, the parallel segmentations are repeated until a satisfactory segmentation result is achieved. In one embodiment, the parallel networks use identical parameters.
[0026] In parallel to the first and second neural networks, a third neural network can also examine neighborhoods of the point cloud or point clouds. The functionality of the third neural network is similar to that of the first network; however, the parameters of the third network are not optimized for segmentation, but rather for the subsequent determination of corresponding points in multiple point clouds by a fourth neural network, as well as a dynamic motion analysis based on this. This optimization therefore affects the points classified as dynamic by the second neural network. Using the results of the second neural network, the results of the first network for static points are selected based on the mask; the results of the third network for dynamic points are selected analogously.The selected points are then combined, for example, using an OR function or element-wise addition, and provided as input values to the fourth neural network. The fourth neural network receives these results from all point clouds analyzed in parallel and examines them for correspondences.
[0027] The correspondences include both static and dynamic point correspondences. These can be further differentiated from one another using the mask. Thus, a transformation matrix can be developed either from all correspondences or, alternatively, only from the correspondences of static points. This transformation matrix can be applied to the static points of a first point cloud to model the movement of the static points—in effect, the sensor's own motion—in the form of a vector field. In parallel, a second vector field can be determined from all correspondences, or alternatively, only from the dynamic points, which models the points' own motion.
[0028] In one embodiment, the transformation matrix can alternatively or additionally be determined from a known distance traveled by the vehicle between two point cloud images. This distance can be determined, for example, based on the known speed and steering angle of the vehicle.
[0029] Figure 1 shows an example of a method 100 according to the invention for determining movement in a scene.
[0030] In step 110, the method 100 receives environmental data from one or more sensors, which may be mounted in or on a vehicle or on a person's clothing. The environmental data includes three-dimensional points whose origin corresponds to the position of the sensor; with multiple sensors, the points may use a point between the sensors, for example, a geometric center, or the position of a selected sensor as the origin. The points may all be recorded at the same time or at multiple times.
[0031] In step 120, the method 100 classifies the points into static and dynamic points. The result is preferably represented as a mask / vector whose components are assigned, for example, zero or one to identify static or dynamic points, or vice versa. In step 120, for example, the neighborhoods of each point are first determined, i.e., those points closest to the point. Furthermore, other properties can be determined, such as color, intensity, sharpness. Based on a typification of neighborhoods, an initial assessment can be made as to whether a point is static or dynamic. Alternatively, points from different points in time can be examined in this step to distinguish static from dynamic points. For example, the points of a first point cloud can be correlated with the points of a second point cloud to determine correspondences between the points.Based on the distances of the determined correspondence pairs and a threshold value, the points of these correspondence pairs can be classified as static or dynamic.
[0032] In step 130, the method creates a transformation matrix for the static points. The transformation matrix is preferably developed from the already determined correspondence pairs, using either all correspondence pairs or, based on the mask, only the static points or their correspondence pairs. The transformation matrix is applied to the static points in step 150 to obtain vectors for representing the scene flow resulting from the self-motion of the sensor(s).
[0033] In step 140, a scene flow is created for the dynamic points. Step 140 can be performed in parallel to steps 130 and 150. Step 140 includes, for example, the creation of correspondence pairs of points from two or more different points in time, provided these were not already created in an earlier step. The correspondence pairs can be examined with regard to their relative position in order to estimate the movement for dynamic and static points. The creation of the correspondence pairs can include the application of a cost minimization function to determine the associated points from different points in time based on previously determined neighborhood features. Finally, vectors (scene flow vectors) that represent the dynamic movement are formed from the correspondence pairs of the dynamic or all points.
[0034] In step 160, the scene flow vectors of the static and dynamic points are output. The scene flow vectors of the static points from step 150 and the scene flow vectors of the dynamic points from step 140 are selected based on the result from step 120. The scene flow vectors can be displayed to a user, for example a vehicle driver or another person, for example on a display. Alternatively, only the points and not the vectors are displayed. The static and dynamic points or their vectors can be highlighted differently. For example, static points can be visually receded into the background by lower intensities or predefined or uniform color values, while dynamic points can be highlighted by stronger intensities or color values. In one embodiment, only the dynamic points or their vectors are displayed.
[0035] Figure 2 shows an exemplary device 200 for implementing method 100. Device 200 can be arranged as a circuit in a vehicle. In one embodiment, device 200 is implemented exclusively in program code that runs on a computer or mobile device (smartphone, notebook).
[0036] The device 200 comprises a processing unit 210 that receives environmental data, for example from one or more sensors, as described above in step 110. The processing unit 210 comprises a segmentation unit 220 that identifies static and dynamic points in the environmental data according to step 120 and labels them, for example, in a mask. The processing unit 210 also comprises an ego-motion unit 230 that creates a transformation matrix for the static points according to step 130. As shown, the ego-motion unit 230 can be part of the processing unit 210 or at least coupled to it.
[0037] The device 200 further comprises a conversion unit 240 which is configured to apply the transformation matrix developed by the self-motion unit 230 to the static points according to step 150 and thus to model and represent the self-motion of the sensor(s) by scene flow vectors.
[0038] Furthermore, the device 200 comprises a scene flow estimation unit (scene flow unit) 250, which subjects the dynamic points to a motion analysis according to step 140 and models the determined movements into scene flow vectors. The scene flow vectors of the static points and the dynamic points are finally combined and output by an output unit 260; step 160.
[0039] The components shown in Figure 2 essentially work with neural networks. Figure 3 shows an example configuration of such neural networks and their training using training data.
[0040] Figure 3 shows two point clouds P and Q, acquired at different times and containing three-dimensional points. In addition to the coordinates of these points, they may contain brightness values, color values, timestamps, and other point properties. Point clouds P and Q are first analyzed independently of each other, namely by neural networks 310, 320, and 330, and by neural networks 340, 350, and 360.
[0041] A first neural network 320 extracts features, such as neighborhood information, from the point cloud P. These features are fed as input values to a second neural network 330. The second neural network 330 generates a segmentation mask from these features, which classifies each point of the point cloud P as static or dynamic. The second neural network 330 can use different resolutions of the point cloud for this purpose, starting with the coarsest resolution. At each resolution, the neighborhoods of a point determined by the first neural network 320 are used to determine whether the point is static or dynamic.
[0042] The training of the two networks 320 and 330 is implemented by a comparison with training data comprising a predetermined segmentation mask. In one embodiment, the training data also contains scene flow vectors. The second neural network 330 outputs a first mask 312 in which the static points ("bg") are marked. The mask 312 is compared with the predetermined segmentation mask. If the two masks differ, the deviation is encoded in the form of a scalar value and sent back to the first and second neural networks 320 and 330 via backpropagation. The scalar value can be generated, for example, by forming a cross-entrophy. After receiving the scalar value, the neural network 320 changes its operation such that the scalar value is reduced in subsequent comparisons.These training steps are repeated until mask 312 corresponds to the specified mask or at least no longer approaches it. The neural networks 320 and 330 together form the segmentation unit 220 shown in Figure 2.
[0043] A third neural network 310 analyzes neighborhoods of the point cloud P. The analysis of the neighborhoods is essentially analogous to the procedure of the second neural network 320, but the parameters of the third neural network 310 are optimized according to different criteria. In the case of the second neural network 320, optimization is carried out with regard to the detection of static points, and in the case of the third neural network 310, optimization is carried out with regard to dynamic points. These different optimizations are trained by a further analysis that includes a fourth neural network 370 and the parallel analysis of the second point cloud Q. The adaptation of the third neural network is controlled via the estimated scene flow of the fourth neural network 370. In one embodiment, the first and second neural networks are trained exclusively for segmentation using the aforementioned zero gradient.The optimization of the third neural network 310 is controlled by the fourth neural network 370 and further neural networks behind it, which serve to estimate scene flow vectors for the dynamic points, as explained further below.
[0044] The second point cloud Q is analyzed, analogously to the first point cloud, by further neural networks 340 and 350 with regard to neighborhoods and static / dynamic segmentation. In parallel, neighborhoods are also analyzed by the further neural network 360. The neural network 350 generates a segmentation mask, which is compared with a predefined mask. Analogous to the above description, a scalar value is formed from the deviation of the mask 314 and the predefined segmentation mask, for example, using cross-entropy, and supplied to the first and second neural networks 340 and 350 for further optimization by backpropagation. The optimization of the neural networks 340 and 350 is adapted, as previously explained, by using a zero gradient so that the network 340 receives the backpropagation training signal only from network 350 by comparing the segmentation masks.
[0045] The optimized masks 311 and 312 now define the dynamic ("fg") and the static ("bg") points of the point cloud P, respectively. The optimized masks 313 and 314 define the dynamic and the static points of the point cloud Q, respectively. The masks 311 and 312 are combined with each other, for example, by means of an element-wise addition 316; the combining particularly includes the use of the underlying technical features of the respective points. Similarly, the masks 313 and 314 are combined by means of addition 315.
[0046] For example, the masks can be combined for each of the resolutions considered according to the following equation:
[0047] HFk- Mf .k ■ Fcontext.k + (1 Mfg,k) ' F(F_encoder, k), where: k: index of the respective resolution;
[0048] HFk: “Hybrid Features”; a data structure containing a combination of output values from the neural networks 3io and 320;
[0049] Mfg ,k: (binary) mask of the points detected as dynamic at resolution k;
[0050] F C ontext,k: technical features determined by the third neural network 310 at resolution k;
[0051] 1 - Mf g ,k: inverted mask of the dynamic points at resolution k;
[0052] F e ncoder,k: technical characteristics determined by the first neural network 320 at resolution k
[0053] - 1 -: Operator that sets the gradient of its operand to zero gradient.
[0054] The results of the additions 315 and 316 are used as input parameters for the fourth neural network 370. The fourth neural network 370 determines pairs of corresponding points from both input parameters using a cost minimization function, for example, by correlating the RFk of the neural networks 310 and 320 with the RFk of the neural networks 340 and 360. The result of the neural network 370 can in turn be compared with a predetermined set of correspondences and delivered to the upstream neural networks 310 and 360 by way of backpropagation.
[0055] Further neural networks can be coupled downstream of the fourth neural network 370. These networks include, in particular, a neural network for determining scene flow vectors of static points and a neural network for determining scene flow vectors of dynamic points (both networks not shown). The neural network for determining scene flow vectors of static points can be an essential component of the self-motion unit 230 shown in Figure 2. The neural network for determining scene flow vectors of dynamic points can be an essential component of the scene flow unit 250 shown in Figure 2.
[0056] The neural network for determining scene flow vectors of static points determines a transformation matrix that models the self-motion of the sensor(s) (step 130 of method 100), while the neural network for determining scene flow vectors of dynamic points determines scene flow vectors directly from the correspondences determined above (step 140 of method 100). Once completed, the transformation matrix can be applied to the static points to define the actual self-motion of the sensor(s).
[0057] In particular, the output values of the neural network for determining scene flow vectors of dynamic points can be compared with training data. Analogous to the above explanations, deviations from the training data can be summarized in a scalar value, which is fed to the neural networks 310, 360, and 370 for optimization. In one embodiment, the optimization of the networks 310, 360, and 370 can also be performed with training data that does not contain scene flow vectors. For this purpose, two physical properties are exploited. The first property describes that neighboring points undergo similar motion. The second property maps the shape preservation and visibility of the scene by comparing the second point cloud with the first point cloud after estimated motion. Each of the two properties can be summarized in a scalar value, which is fed to the networks 310, 360, and 370 for optimization.A scalar value of the first property can be calculated using a smoothing function, for example. A scalar value of the second property can be calculated using a Chamfer function, for example. Both scalar values can be combined, for example, by addition or a normalization rule, and fed to the networks as a single scalar value.
Claims
CLAIMS 1. A computer-implemented method comprising: Receiving environmental data by a processing unit, wherein the environmental data contains three-dimensional points; Classifying, by a segmentation unit of the processing unit, the three-dimensional points as static or dynamic points; Creating, by an egomotion unit of the processing unit, a transformation matrix for the static points; Converting, by a conversion unit of the processing unit, the static points into scene flow vectors using the transformation matrix; Estimating, by a unit of the scene flow estimation processing unit, scene flow vectors for the dynamic points; and Output the scene flow vectors of the static and dynamic points.
2. The method according to claim 1, wherein the classifying, creating, transforming and estimating are implemented by a respective neural network, and wherein the method further comprises: Training the neural networks using training data, where the training data includes point clouds and segmentation into static and dynamic points. 3- Method according to claim 1 or 2, wherein the environmental data contain first and second three-dimensional points, the first three-dimensional points being recorded at a first time and the second three-dimensional points being recorded at a second time.
4. The method of claim 3, wherein classifying the three-dimensional points comprises identifying correspondences between the first three-dimensional points on the one hand and the second three-dimensional points on the other hand, and wherein the classification is repeated successively in a plurality of different resolutions of the environmental data.
5. The method of claim 4, wherein the results of the second neural network are multiplied by a zero gradient before merging.
6. Device comprising: A processing unit configured to receive environmental data of the device, wherein the environmental data contains three-dimensional points; a segmentation unit of the processing unit, wherein the segmentation unit is configured to classify the three-dimensional points as static or dynamic points; a self-motion unit of the processing unit, wherein the self-motion unit is configured to create a transformation matrix for the static points; a conversion unit of the processing unit configured to convert the static points into scene flow vectors based on the transformation matrix; a unit of the scene flow estimation processing unit configured to estimate scene flow vectors for the dynamic points; and a unit configured to output the scene flow vectors of the static and dynamic points.
7. The apparatus of claim 6, wherein the classifying, creating, transforming, and estimating are implemented by a respective neural network, and wherein the apparatus further comprises means configured to train the neural networks using training data, wherein the training data comprises point cloud segmentation into static and dynamic points.
8. The apparatus of claim 6 or 7, wherein the environmental data includes first and second three-dimensional points, the first three-dimensional points being recorded at a first time and the second three-dimensional points being recorded at a second time.
9. The apparatus of claim 8, wherein classifying the three-dimensional points comprises identifying correspondences between the first three-dimensional points on the one hand and the second three-dimensional points on the other hand, and wherein the classification is repeated successively in a plurality of different resolutions of the environmental data.
10. A computer-readable medium having stored thereon instructions which, when executed by a processor, perform the method of any one of claims 1 to 5.