A method for generating an at least three-dimensional representation

WO2026176055A1PCT designated stage Publication Date: 2026-08-27ROBERT BOSCH GMBH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/054704
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2026-02-20
Publication Date
2026-08-27

Smart Images

  • Figure EP2026054704_27082026_PF_FP_ABST
    Figure EP2026054704_27082026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method (100) for generating an at least three-dimensional representation of at least one scene in an environment (60) of an ego-vehicle (50) using LiDAR data, the LiDAR data resulting from at least one LiDAR sensor (51) of the ego-vehicle (50), the method (100) comprising: - Providing (101) an ego motion indicator for an ego motion of the ego-vehicle (50), - Detecting (102) moving objects within the at least one scene based on the LiDAR data, - Estimating (103) a ground plane within the at least one scene based on a ground plane segmentation based on the LiDAR data, - Determining (104) a scene flow of the detected moving objects based on the provided ego motion indicator and the estimated ground plane, thereby estimating a motion of the detected moving objects based on the LiDAR data, - Performing (105) a synchronization of camera data with the LiDAR data using the determined scene flow and the provided ego motion indicator, the camera data resulting from at least one camera (52) of the ego-vehicle (50), - Generating (106) the at least three-dimensional representation based on the synchronization.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] R.416809

[0002] - 1 -

[0003] Description

[0004] Title

[0005] A method for generating an at least three-dimensional representation

[0006] The invention relates to a method for generating an at least three-dimensional representation. Furthermore, the invention relates to a computer program, an apparatus, and a storage medium for this purpose.

[0007] State of the art

[0008] It is known that 3D ground truth data is essential for robotics, self-driving cars, and mapping technologies. Combining LiDAR and multi-camera sensor systems allows for large-scale data collection and comprehensive 3D environment reconstruction.

[0009] However, sensors like cameras, LiDARs, IMUs (Inertial Measurement Units), and odometry operate at different frequencies. Directly generating 3D ground truth from raw data may result in misaligned information. To achieve high-quality 3D data suitable for production use, sensor synchronization (especially cameras relative to LiDAR) is crucial.

[0010] However, synchronizing moving objects consistently across 3D coordinates and cameras presents a significant challenge.

[0011] While existing methods like Chodosh et al.'s (2023) "Re-Evaluating LiDAR Scene Flow for Autonomous Driving" provide robust pipelines for estimating moving objects, they lack proper synchronization to cameras, accumulation of multiple point clouds, and denser representation of sparse areas. The pipeline may depend on high-quality static calibration, including extrinsic parameters (sensor positions relative to the vehicle / robot) and intrinsic parameters (sensor properties).R.416809

[0012] - 2 -

[0013] Disclosure of the invention

[0014] According to aspects of the invention a method with the features of claim 1, a computer program with the features of claim 11, a data processing apparatus with the features of claim 12 as well as a computer-readable storage medium with the features of claim 13 are provided. Further features and details of the invention are disclosed in the respective dependent claims, the description and the drawings. Features and details described in the context to the method according to the invention also correspond to the computer program according to the invention, the data processing apparatus according to the invention as well as the computer-readable storage medium according to the invention, and vice versa in each case.

[0015] According to an aspect of the invention a method for generating an at least three-dimensional (3D) representation of at least one scene in an environment of an egovehicle using LiDAR data is provided. In other words, the method may generate 3D-data (i.e. the representation) that represents at least one scene from the surroundings (i.e. the environment) of a vehicle (referred to as the 'ego-vehicle').

[0016] The LiDAR data may result from at least one LiDAR sensor of the ego vehicle. Therefore, the ego-vehicle may be the vehicle that comprises the at least one LiDAR sensor for providing the LiDAR data.

[0017] The process carried out within the method may involve creating a visual or spatial model that depicts the at least one scene. It may be a goal of the method to capture, process, and / or render the environment and objects in this environment like other vehicles or pedestrians in a form that allows for detailed spatial awareness, which may be used for applications like navigation, object detection, or advanced driver-assistance systems.

[0018] The method may comprise the provision of an ego motion indicator for an ego motion of the ego-vehicle. The ego motion indicator may quantify a velocity of the ego-vehicle and / or specify a trajectory of the ego-vehicle. The provision of the ego motion indicator may also be possible by estimating an ego motion of the ego-vehicle based on the LiDAR data, especially in case the ego-motion is not directly available. The ego motion indicator may also be referred to as ego motion value or ego motion measurement result or ego motion estimator.R.416809

[0019] -3-

[0020] The method may further comprise the detection of moving objects of the at least one scene based on the LiDAR data.

[0021] The method may optionally comprise the estimation of a ground plane of the at least one scene based on a ground plane segmentation based on the LiDAR data. The ground plane segmentation may be used to identify and separate the ground or flat surface (the "ground plane") from other objects in the scene. This may preferably be carried out by methods that are already known for this purpose.

[0022] The method may further comprise the determination of a scene flow of the detected moving objects. The scene flow may therefore be referred to as moving objects scene flow. The determination of the scene flow may be based on the provided ego motion indicator and / or the estimated ground plane.

[0023] The scene flow may be configured as a representation of the motion of the detected objects or of surfaces within a three-dimensional space over time, derived from consecutive frames or measurements of the LiDAR data.

[0024] This process of determining the scene flow may also include an ego motion compensation and / or ground plane removal, particularly from the LiDAR data.

[0025] It is possible that the determination of the scene flow allows for an estimation of a motion, particularly velocities, of the detected moving objects based on the LiDAR data. The moving objects may also be referred to as dynamic objects.

[0026] The method may further comprise the performance of a synchronization of camera data with the LiDAR data using the determined scene flow and / or the provided ego motion indicator. The camera data may result from at least one camera of the ego vehicle.

[0027] It is possible that the determined scene flow provides misaligned information about the moving objects, and the synchronization is therefore used to correct the misaligned information. Particularly after the ego-motion compensation, dynamic objects are misaligned.R.416809

[0028] -4-

[0029] The method may further comprise the generation of the at least three-dimensional representation based on the synchronization. This may also include the generation of ground truth based on the synchronization.

[0030] It is therefore possible, using the method according to the invention to generate a detailed 3D representation of a scene around an ego vehicle using LiDAR and camera data synchronized by estimated ego motion and scene flow. This allows, for example, for accurate object detection, tracking, and understanding of the surrounding environment, facilitating applications like autonomous driving and robotics.

[0031] The method according to the invention may be regarded as a multi-step approach: determining ego motion, detecting moving objects, estimating a ground plane, calculating scene flow, synchronizing camera data with LiDAR, and finally generating a 3D representation based on this synchronized information. This comprehensive pipeline yields high-quality depth data enriched by semantic context for a more robust understanding of the scene.

[0032] The goal of the method may also be understood as a minimization problem, how to align the misaligned moving objects better. The method may particularly be used to synchronize the LiDAR data in the form of asynchronously recorded LiDAR data, odometry, IMU, and cameras to generate dense 3D ground truth.

[0033] Aspects of the invention may be characterized by a processing pipeline to generate dense depth maps in image space given a sequence of sensor and particularly LiDAR points. This pipeline may comprise odometry estimation with ICP, and / or LiDAR scene flow estimation, and / or accumulation of the LiDAR point cloud over time (also referred to as LiDAR accumulation), and / or synchronization from the point cloud to the camera, and / or semantic segmentation post-processing to remove hidden points, and / or depth completion based on both camera and LiDAR data. ICP stands for Iterative Closest Point.

[0034] In contrast to many solutions of the prior art, the invention may provide consistent scene flow over multiple frames, which allows LiDAR accumulation over a higher number of LiDAR scans. It may provide complete 3D structure by using a meshing step. It may also offer an approach for synchronization between cameras and LiDAR by usingR.416809

[0035] -5-

[0036] velocity and time intervals and may provide an approach to get dense, consistent depth maps for every camera in the given sensor setup.

[0037] Another advantage includes at least one of the following: improved optimization speed for scene flow estimation, accurate synchronization between LiDAR and camera, a map of surroundings, a dense 3D mesh of the 3D environment with moving object meshes instances, 3D dense depth map ground truth per each camera, moving / static objects and their velocities for each camera.

[0038] The generated ground truth may save costs for labelling 3D objects since they can be labelled in pure LiDAR and later reprojected to cameras by using the synchronization technique. It also enables labelling in 3D space. The invention may also provide automatic 3D ground truth at scale, which enables cost-efficient autonomous systems' data loop.

[0039] It is possible that, based on the generated at least three-dimensional representation, at least one of the following procedures is carried out:

[0040] A depth completion, thereby estimating depth information from the at least three- dimensional representation, the depth information preferably providing distance information corresponding to the moving objects and / or further objects within the at least one scene relative to a position of the ego-vehicle.

[0041] A semantic segmentation refinement for correcting incorrectly detected ground plane information and / or adding ground plane information about the ground plane (of the at least one scene) in close range to the ego-vehicle, particularly, where the LiDAR data does not provide sufficient information.

[0042] In other words, it is possible to further refine the generated 3D representation by incorporating depth completion techniques. This allows for estimation of depth information, providing distance details about moving objects and other scene elements relative to the ego-vehicle's position.

[0043] Additionally, semantic segmentation refinement can be employed to correct inaccuracies in ground plane detection and supplement ground plane information in close proximity to the ego-vehicle, where LiDAR data might be insufficient.R.416809

[0044] -6-

[0045] It is further possible that the ego motion indicator is specific for the ego motion and indicates and particularly quantifies at least a velocity of the ego-vehicle and / or trajectory of the ego-vehicle. It is particularly possible to determine and quantify the ego-vehicle's velocity and trajectory using various sensors, such as GPS, inertial measurement units (IMU), or odometers. This information can be used to provide a precise ego motion indicator for the method, enabling accurate detection and tracking of moving objects within the scene. The specific quantification of the ego-vehicle's motion can significantly improve the accuracy and reliability of the generated 3D representation by ensuring correct alignment and relative motion estimations.

[0046] It is also possible that the provision of the ego motion indicator comprises estimating the ego motion of the ego-vehicle based on the LiDAR data, particularly in case the egomotion is not already available. In other words, it is possible that the method can estimate the ego motion of the vehicle using LiDAR data if this information is not directly available. This estimation could be based on analysing the changes in LiDAR point cloud readings between consecutive scans, allowing the system to determine the vehicle's movement and position. This approach offers robustness by relying on sensor data directly acquired by the ego-vehicle, potentially overcoming limitations associated with external motion tracking systems

[0047] It is further possible that the determination of the scene flow of the detected moving objects comprises an ego motion compensation and / or ground plane removal from the LiDAR data. It is possible to improve the accuracy of the scene flow calculation by compensating for the ego motion of the ego-vehicle and / or removing the influence of the ground plane from the LiDAR data. By applying these techniques, the method can more accurately track the movement of objects in the scene, even in complex environments with moving vehicles and pedestrians. The removal of the ground plane and / or scene flow creation may be carried out by using a neural scene flow prior (NSFP) approach.

[0048] It is also possible that the at least three-dimensional representation comprises a depth map, the depth map preferably providing distance information corresponding to the moving objects and / or further objects within the at least one scene relative to a position of the ego vehicle, the depth map particularly being refined using the camera data. TheR.416809

[0049] -7-

[0050] depth map can therefore be refined based on the synchronisation and particularly the synchronized camera data, leading to more accurate distance measurements.

[0051] It is further possible that the LiDAR data comprises a sparse at least three-dimensional point cloud. At least one of the following steps may also be carried out:

[0052] Densification of the sparse point cloud, particularly using a depth completion, preferably based on at least one deep learning model trained on paired LiDAR and camera training images,

[0053] Constructing an at least three-dimensional mesh representing the environment based on the point cloud,

[0054] Projecting the mesh onto the camera's image plane, thereby creating a denser representation of the represented scene,

[0055] Applying ray casting to remove hidden points from the point cloud.

[0056] These methods enhance the quality and completeness of the LiDAR data, enabling more accurate scene reconstruction and object detection. Sparse may refer to gaps between individual point of the point cloud.

[0057] It is possible for the method according to the invention to be used in a vehicle. The vehicle may, for example, be configured as a motor vehicle and / or passenger vehicle and / or at least partially automated / autonomous vehicle. The vehicle can have a vehicle device, e.g. for providing an autonomous driving function and / or a driver assistance system. The vehicle device can be designed to control and / or accelerate and / or brake and / or steer the vehicle at least partially automatically. The vehicle device may be part of an autonomous driving system.

[0058] It is possible that the generated at least three-dimensional representation is used for the generation of a data basis, particularly of training data for the training of a machine learning model, and preferably for providing ground truth data for the training, particularly for the classification of sensor data on the basis of pixel values of the sensor data.

[0059] Furthermore, it is possible that an autonomous driving system is configured and / or controlled based on the generated at least three-dimensional representation. TheR.416809

[0060] -8 -

[0061] autonomous driving system may be used to perform autonomous driving of a technical system such as a vehicle or robot or the like.

[0062] The representation generated by a method according to the invention may be used for providing training data for a machine learning system. The machine learning system, in particular in the form of a machine learning model, is preferably trained for classification and in particular for object detection. The training can be intended to train the machine learning system / the machine learning model by means of a training data set for classification, in particular for image classification, of image data such as digital images on the basis of pixels and / or pixel values, preferably edges or pixel attributes (of the image data). The image data or digital images can, for example, result from a recording by at least one sensor, preferably at least one camera, preferably of a vehicle and particularly preferably of a camera environment and / or vehicle environment during driving (the vehicle). The classification can be intended to recognise objects in an environment depicted by the image data or digital images and / or to recognise a traffic scene.

[0063] Based on the representation generated by a method according to the invention, at least one control action, preferably for a vehicle or for another technical system, may be initiated and / or carried out. The control action may comprise at least one of the following: braking, steering, accelerating, overtaking manoeuvres, emergency braking, activation of an alarm system, activation of a hazard warning system, activation of a direction indicator, light control, or the like.

[0064] Based on the representation generated by a method according to the invention, a classification can be carried out and / or used to recognise an obstacle, for example, regardless of whether it is directly in the direction of travel or next to it. Depending on the location (e.g. depending on the expected vehicle trajectory), a corresponding control action such as braking or swerving can be initiated.

[0065] The 'classification' and 'image classification' can also include 'object detection' or 'object detection in images'. In particular, this means classifying whether or not there are objects in certain areas of the image. In addition, the terms 'classification' and 'image classification' can also refer to 'semantic segmentation', in particular in the form of pixel-by-pixel classification.R.416809

[0066] -9 -

[0067] In another aspect of the invention, a computer program may be provided, in particular a computer program product, comprising instructions which, when the computer program is executed by at least one computer, cause the computer to carry out the method according to the invention. Thus, the computer program according to the invention can have the same advantages as have been described in detail with reference to a method according to the invention.

[0068] In another aspect of the invention, an apparatus for data processing may be provided, which is configured to execute the method according to the invention. As the apparatus, for example, at least one computer can be provided which executes the computer program according to the invention. The computer may include at least one processor that can be used to execute the computer program. Also, a non-volatile data memory may be provided in which the computer program may be stored and from which the computer program may be read by the processor for being carried out.

[0069] According to another aspect of the invention a computer-readable storage medium may be provided which comprises the computer program according to the invention and / or instructions which, when executed by at least one computer, cause the computer to carry out the steps of the method according to the invention. The storage medium may be formed as a data storage device such as a hard disk and / or a non-volatile memory and / or a memory card and / or a solid state drive. The storage medium may, for example, be integrated into the computer.

[0070] Furthermore, the method according to the invention may be implemented as a computer-implemented method. Alternatively or additionally, at least one of the disclosed method steps may be computer-implemented and / or automated.

[0071] Further advantages, features and details of the invention will be apparent from the following description, in which embodiments of the invention are described in detail with reference to the drawings. In this context, the features mentioned in the claims and in the description may each be essential to the invention individually or in any combination. Showing:R.416809

[0072] - 10-

[0073] Fig. 1: A method, computer program, a storage medium and apparatus according to embodiments of the invention.

[0074] Fig. 2: An overview of the approach according to embodiments of the invention.

[0075] Fig. 3: Further details of embodiments of the invention.

[0076] Fig. 4: Overview of the ground plane detection.

[0077] Fig. 5-9 Further details of embodiments of the invention.

[0078] Fig. 1 shows a method, computer program, a storage medium and apparatus according to embodiments of the invention. Particularly, Fig. 1 shows a method 100 according to embodiments of the invention for generating an at least three-dimensional representation of at least one scene in an environment 60 of an ego-vehicle 50 using LiDAR data. The LiDAR data may origin from at least one LiDAR sensor 51 of the egovehicle 50.

[0079] The method 100 may comprise, according to a first method step 101, providing an ego motion indicator for an ego motion of the ego-vehicle 50. Furthermore, according to a second method step 102, moving objects may be detected within the at least one scene based on the LiDAR data. According to a third method step 103, a ground plane within the at least one scene may be estimated based on a ground plane segmentation based on the LiDAR data. Then, according to a fourth method step 104, a scene flow of the detected moving objects may be detected based on the provided ego motion indicator and the estimated ground plane, thereby estimating a motion of the detected moving objects based on the LiDAR data. According to a fifth method step 105, a synchronization of camera data with the LiDAR data may be carried out using the determined scene flow and the provided ego motion indicator. The camera data may result from at least one camera 52 of the ego-vehicle 50, as shown in Fig. 1. Then, according to a sixth method step 106, a generating of the at least three-dimensional representation based on the synchronization may take place.

[0080] In Fig. 2, embodiments of the invention are shown where the LiDAR flow pipeline comprises at least two, but preferably up to five steps 201-210.R.416809

[0081] - 11 -

[0082] According to a first step 201, also referred to as ICP Odometry Step, ICP-based odometry estimation and ground plane detection is performed. The second step 202 may involve the scene flow estimation and the estimation of the three-dimensional moving objects. In the third step 203, a synchronization between the camera data and LiDAR data is achieved. The fourth step 204 focuses on depth completion. Finally, in the fifth step 205, semantic refinement can be applied. The block 210 stands for an additional output of the scene flow estimation, where moving objects can be segmented, and / or three-dimensional boxes can be estimated. In this context, ICP can stand for Iterative Closest Point.

[0083] In Fig. 3, further details of embodiments of the invention are shown. According to an ICP Odometry Step 303, this step may take as input a sequence of at least two LiDAR scans 301 and static calibration transformation from LiDAR to the ego vehicle (see 302). The output 304 can be a set of 6 DoF (Degrees of Freedom) poses (rotation and translation) with respect to the first LiDAR scan, expressed in the ego vehicle's coordinate system.

[0084] Fig. 4 illustrates the ground plane detection module 402 according to embodiments of the invention. It takes a single LiDAR scan 401 as input and produces two outputs: a semantic segmentation map 403 indicating whether each point belongs to the ground or not, and a generated ground mesh 404 representing the detected ground surface. The semantic segmentation map 403 can also be referred to as ground plane segmentation, represented in Fig. 5 with the reference sign 501.

[0085] Fig. 5 illustrates the Scene Flow Estimation module according to embodiments of the invention, a component in tracking object movement. It takes as input the ground plane segmentation 501 from the previous step, a set of LiDAR scans 502, and odometry information 503 obtained from the ICP odometry step, along with calibration data 505.

[0086] The calibration data 505 has been represented in Fig. 3 with the reference sign 302 and the odometry information 503 with 304.

[0087] The Scene Flow Estimation module's objective is particularly to predict scene flow network weights (see 506 in Fig. 5), motion segmentation for each point (see 507 in Fig.

[0088] 5), and 3D bounding boxes for moving objects (see 508 in Fig. 5). The module mayR.416809

[0089] - 12 -

[0090] operate in four distinct steps: First, a LiDAR preprocessing 504 may be used to align consecutive LiDAR scans 502 using the obtained ICP poses 503 (compensating for ego motion) and to remove ground plane points and irrelevant 3D points to focus on relevant scene information for flow estimation. Second, a scene flow estimation 506 may be used to estimate 3D vectors 511 representing the motion between two consecutive aligned LiDAR scans 502. Third, a motion segmentation 507 may be used to leverage the estimated scene flow. This stage may segment moving objects from static background elements within the LiDAR data (see 512). Then, a 3D Box detection 508 may be finally used to identify and generate 3D bounding boxes 513 around detected moving objects, providing a spatial representation of their location and extent.

[0091] Fig. 6 shows the that the module according to further embodiments of the invention takes as input a sequence of LiDAR scans (at least 3), LiDAR poses and calibration, as well as LiDAR semantic segmentation from the ground plane detection step. It also takes camera calibration and timestamps of the rolling shutter. It outputs a depth map 608, motion segmentation 609, global pose 610 of the camera in ego vehicle coordinates aligned with the LiDAR, and moving objects' velocity 612. In Fig. 6, the reference sign 601 refers to the sequence of the lidar scans, 602 to the set of 6 DoF poses with respect to the first lidar scan, 603 the motion segmentation, 604 to the camera static calibration (transformation to the ego vehicle), 605 to the LiDAR static calibration, 606 to the camera timestamp, intrinsics, and the projection model, 607 to the camera lidar synchronization (lidar distortion), 608 to the depth map, 609 to the motion segmentation, 610 to the cameras global pose, 611 to the aligned LiDAR, and 612 to the moving objects velocity.

[0092] Fig. 7 shows the depth completion module 705 according to further embodiments of the invention. As input, it takes LiDAR poses 702 and calibration 704, camera poses 703 and calibration 706, depth map 707, semantic segmentation map 707, and ground mesh 701. It produces densified depth by mesh 708, combined depth with sparse elevated objects, meshed ground semantic segmentation 709, and normal 710.

[0093] Fig. 8 shows the semantic refinement module 804. As input, it takes densified depth 801, normal 803, semantic segmentation 802, and an RGB image. It processes the semantic segmentation CNN with the RGB image and, using the provided semantic segmentation mask, adds ground plane and sky maximum depth values to the depthR.416809

[0094] - 13 -

[0095] map. The output refers to a refined depth 805, refined segmentation 806 and refined normal 807.

[0096] Fig. 9 shows former improvements with only sparse LiDAR supervision (left column) compared with semi-supervision (right column). The point cloud view is displayed at an angle to highlight the improvements observed with and without self-supervision.

[0097] Reference sign 901 refers to "Wrong / distorted depth on nearby objects with few or no LiDAR depth". 902 refers to "Improved flat surfaces with homogeneous textures and nearby objects". 903 refers to "Incorrect depth for sky regions". 904 refers to "Correct elevated objects and improved sky pixel depth". 905 refers to "Straight guardrails but contains noisy depth". 906 refers to "Improved guardrails".

[0098] In self-driving car systems or robotics, a complete 3D map representation is essential for reliable and safe navigation. Data-driven approaches require large-scale 3D ground truth data to enable autonomous systems to navigate the world reliably. Embodiments of the invention facilitates the creation of 3D ground truth, saving development time and costs, and allowing for the delivery of cheaper solutions with full autonomous functionality.

[0099] The overall pipeline is shown on the Fig. 1 and 2 and may comprise the following steps: According to Step 1, LiDAR ICP odometry may estimate the ego motion of the vehicle (Fig. 3) and ground plane segmentation estimates the ground plane (Fig. 4). The second step particularly aligns multiple LiDARs by using the estimated ego motion and removes the ground plane to estimate scene flow, moving objects segmentation, and 3d moving objects boxes (Fig. 5). The third step may synchronize the camera with LiDAR by using velocities estimated from scene flow and reprojects points from LiDAR to the camera (Fig. 6). The fourth step may densify the sparse LiDAR points in the camera view by using meshing and then applies ray casting to get only visible points within the camera (Fig. 7). Depth completion could also be done by other popular methods in literature like CompletionFormer. Approaches like CompletionFormer propagate the sparse measurements to the entire image to obtain dense depth prediction. The fifth step may further refine depth maps in the camera by using semantic segmentation from a foundational network to remove incorrectly projected hidden points from LiDAR (Fig. 8).R.416809

[0100] - 14-

[0101] The first step, as particularly represented in Fig. 3, may involve estimating ICP odometry between LiDAR scans to obtain an ego motion trajectory for subsequent alignments. Concurrently, ground plane segmentation is performed (see Fig. 4). The ground plane detection segments the point cloud into elevated and ground points (and additional labels as described below), producing a mesh of the ground.

[0102] The ground plane detection process may segment the point cloud as follows:

[0103] 1. Ego vehicle points (those LiDAR points touching ego vehicle)

[0104] 2. Ground plane

[0105] 3. Elevated objects used for scene flow estimation. Points below ground and above a certain height and distance may be masked to reduce memory consumption during scene flow computation.

[0106] 4. Unlabelled objects (points not falling into any of the above categories)

[0107] 5. Ground plane outliers (points that do not fit the ground after mesh creation)

[0108] In Fig. 5, according to the second step, it is the goal of estimating 3D vectors 511 of motion of every point between n-LiDARs (here n > 2) and extract moving objects instances. The goal of this step is to do alignments of the n-point clouds in a way that improve the efficiency of the scene flow estimation by following approach:

[0109] 1. Removing the ground plane and non-related points by using LiDAR segmentation masks from the ground plane detection from LiDAR scans. This step aims to reduce memory consumption by concentrating on the scene flow estimation only on the relevant part of the point cloud which is in the interval. As the result of the preprocessing, the number of points for scene flow estimation are reduced and the scene flow estimation is enforced mainly for the moving objects, since already ego motion with ICP odometry are already separately estimated in meters.

[0110] 2. Aligning two (or more) consecutive point clouds with respect to each other by ego motion compensation.

[0111] 3. Estimating the scene flow F from the target point cloud Starget to the source point clouds SSOurce_i, Ssource_2, ... , Ssource_n based on learnable model 0. The learnable model can be represented as multilayer perceptron neural network, or the occupancy grid with learnable features at the corners, or hash table or any combination of these representations. The optimization objective is:R.416809

[0112] - 15-

[0113] >

[0114]

[0115] The scene flow F = g(p, 6) is represented implicitly. Based on the constant velocity assumption, the target point cloud Starget may be transformed to be as close as possible to the given closest point cloud by timestamp from the list

[0116]

[0117] SSOurce_2, ... , SSOurce_n-

[0118] Then the velocity may be computed as ratio between scene flow Velocity = F / At, where At - time delta between target point cloud Starget and corresponding closest source point cloud.

[0119] Given the velocities per each point, the time interval between target and each and every source point cloud can be computed and given constant velocity assumption corresponded scene flow to every point cloud can be computed from the listSsotlrCe_i, Ssource_2r, Ssource_n-

[0120] As loss function, the Chamfer distance function can be used:

[0121]

[0122] or distance transform.

[0123] Additionally, rigid body loss may be used: Clustering can be applied to get over segmented clusters of points for each cluster to enforce rigid consistency. The distances can be computed for every point to its neighbours in cluster. This distance could be preserved after propagation with scene flow. The scene flow vectors should be same. Also, occlusions can be explicitly modelled.

[0124] The next step could comprise moving object segmentation: Moving from static objects can be separated by using following assumptions:

[0125] 1. In aligned LiDAR all static objects should have zero scene flow.

[0126] 2. All dynamic objects should have scene flow more than threshold (e.g. ~ 2 m / s). 3. Refinement with clustering if majority of the points in cluster dynamic all points in cluster dynamic; otherwise, static.R.416809

[0127] - 16-

[0128] Given the motion segmentation, 3d boxes can be obtained for each moving object as following:

[0129] 1. Cluster masked moving objects point cloud.

[0130] 2. For each cluster compute 3D box and use as orientation mean scene flow vector.

[0131] 3. Track 3D boxes for the shape refinement.

[0132] The synchronization of the LiDAR with given camera is the process of the shifting of the 3D points which were observed in different time to a position which correspond to a single timestamp of the camera image. Non-solid-state LiDAR (e.g. Velodyne) creates measurement scan by rotation. During the scanning of the environment, the ego vehicle with the LiDAR is also moving. The motion of ego vehicle creates distortion of the LiDAR points. Additional distortion happens due to the motion of the moving objects. Both the effect of the ego vehicle motion and from moving objects should be compensated. The synchronization and reprojection of the point cloud are applied to the camera as follows:

[0133] 1. First, the closest LiDAR to the given camera is found by timestamps,

[0134] 2. For every point of the LiDAR a velocity computed in previous step is obtained, 3. For every point in the LiDAR delta time from LiDAR points timestamps and camera timestamp is computed,

[0135] 4. Given velocity for every point and delta time, moving objects scene flow F = V *At can be computed, here F is the scene flow, V is the velocity for given point, At is the delta timestamp for given point,

[0136] 5. In order to do ego motion compensation, a given ego motion pose for the current LiDAR from ICP Odometry step can be obtained, thereby finding the closest ego pose timestamp to this LiDAR so that: curr_lidar_timestamp <camera_timestamp < closest_ego_pose_timestamp (3) or closest_ego_pose_timestamp < camera_timestamp < curr_lidar_timestamp (4),

[0137] 6. Slerp can be used to interpolate ego pose for the given camera timestamp.

[0138] 7. Finally, matrix multiplication can be used to transform all points from closest LiDAR to the camera ego pose.

[0139] 8. A camera model can be used to reproject 3d points to the image and create raw depth mapR.416809

[0140] - 17-

[0141] According to the fourth step, meshing and raycasting can be applied to make a complete depth and also remove hidden points of the LiDAR, which are not visible from camera. According to an alternative with Depth completion, given the sparse measurements from LiDAR aligned with the camera, it is referred to conventional solutions to propagate the sparse measurements to the entire image. One notable work is from Zhang et al., published in CVPR 2023: CompletionFormer: Depth completion with convolution and vision transformers.

[0142] The performance of the CompletionFormer can be improved by the following steps:

[0143] 1) The CompletionFormer can be trained in a semi-supervised mode. The semisupervision includes the self-supervision that is popularly employed in monocular depth estimation literature. The self-supervision leverages the adjacent frames during training as a supervisory signal for view synthesis. Fig. 9 illustrates the observed benefits of selfsupervision.

[0144] 2) To make the CompletionFormer robust to LiDAR-camera cross-calibration issues, it can be trained by augmenting the cross-calibration transformation matrix. The augmentation can be done by explicitly introducing noise in the cross-calibration matrix by introducing small delta errors in roll, pitch, and yaw. The perturbed cross-calibration matrix may be used to warp the input sparse LiDAR depth map to the CompletionFormer.

[0145] In the course of training, the CompletionFormer learns to correct the noisy crosscalibration so that the predicted depths are well aligned with the camera image.

[0146] According to the fifth step, semantic refinement may get as input depth, semantic segmentation and normals and refinement of the ground plane with semantic segmentation may be applied.

[0147] The goal of the approach is to fix incorrect ground plane in close range where the LiDAR does not provide any information (blind spot). To do that foundational semantic segmentation model may be applied on the RGB image and this semantic information is used to refine depth maps: Also semantic refinement puts max depth on sky semantic segmentation class in depth map.R.416809

[0148] - 18-

[0149] The above explanation of the embodiments describes the present invention in the context of examples. Of course, individual features of the embodiments can be freely combined with each other, provided that this is technically reasonable, without leaving the scope of the present invention.

Claims

R.416809- 19 -Claims1. A method (100) for generating an at least three-dimensional representation of at least one scene in an environment (60) of an ego-vehicle (50) using LiDAR data, the LiDAR data resulting from at least one LiDAR sensor (51) of the egovehicle (50), the method (100) comprising:Providing (101) an ego motion indicator for an ego motion of the egovehicle (50),Detecting (102) moving objects within the at least one scene based on the LiDAR data,Estimating (103) a ground plane within the at least one scene based on a ground plane segmentation based on the LiDAR data,Determining (104) a scene flow of the detected moving objects based on the provided ego motion indicator and the estimated ground plane, thereby estimating a motion of the detected moving objects based on the LiDAR data,Performing (105) a synchronization of camera data with the LiDAR data using the determined scene flow and the provided ego motion indicator, the camera data resulting from at least one camera (52) of the ego-vehicle (50),Generating (106) the at least three-dimensional representation based on the synchronization.R.416809- 20-2. The method (100) of claim 1, characterized in that, based on the generated at least three-dimensional representation, at least one of the following procedures is carried out:A depth completion (705), thereby estimating depth information from the at least three-dimensional representation, the depth information preferably providing distance information corresponding to the moving objects and / or further objects within the at least one scene relative to a position of the ego-vehicle (50),A semantic segmentation refinement for correcting incorrectly detected ground plane information and / or adding ground plane information about the ground plane in close range to the ego-vehicle (50), particularly, where the LiDAR data does not provide sufficient information.

3. The method (100) of any one of the preceding claims, characterized in that the ego motion indicator is specific for the ego motion and indicates and particularly quantifies at least a velocity of the ego-vehicle (50) and / or trajectory of the ego-vehicle (50).

4. The method (100) of claim 3, characterized in that the provision (101) of the ego motion indicator comprises estimating the ego motion of the ego-vehicle (50) based on the LiDAR data, particularly in case the ego-motion is not already available.

5. The method (100) of any one of the preceding claims, characterized in that the determination (104) of the scene flow of the detected moving objects comprises an ego motion compensation and ground plane removal from the LiDAR data.

6. The method (100) of any one of the preceding claims, characterized in that the at least three-dimensional representation comprises a depth map, the depth map preferably providing distance information corresponding to the moving objects and / or further objects within the at least one scene relative to a position of the ego vehicle, the depth map particularly being refined using the camera data.R.416809-21-7. The method (100) of any one of the preceding claims, characterized in that the determined scene flow provides misaligned information about the moving objects, and the synchronization is used to correct the misaligned information.

8. The method (100) of any one of the preceding claims, characterized in that the generated at least three-dimensional representation is used for the generation of a data basis, particularly of training data for the training of a machine learning model, and preferably for providing ground truth data for the training, particularly for the classification of sensor data on the basis of pixel values of the sensor data.

9. The method (100) of any one of the preceding claims, characterized in that an autonomous driving system is configured and / or controlled based on the generated at least three-dimensional representation.

10. The method (100) of any one of the preceding claims, characterized in that the LiDAR data comprises a sparse at least three-dimensional point cloud, wherein at least one of the following steps are carried out:Densification of the sparse point cloud, particularly using a depth completion, preferably based on at least one deep learning model trained on paired LiDAR and camera training images,Constructing an at least three-dimensional mesh representing the environment based on the point cloud,Projecting the mesh onto the camera's image plane, thereby creating a denser representation of the represented scene,Applying ray casting to remove hidden points from the point cloud.

11. A computer program (20), comprising instructions which, when the computer program (20) is executed by at least one computer (10), cause the computer (10) to carry out the method (100) of any one of the preceding claims.

12. A data processing apparatus (10), comprising means for carrying out the method (100) of any one of claims 1 to 10.R.416809- 22 -13. A computer-readable storage medium (15) comprising instructions which, when executed by at least one computer (10), cause the computer (10) to carry out the steps of the method (100) of any one of claims 1 to 10.