Method and system for increasing lidar data
By fusing data from multiple sensors to generate simulated driving scenarios, the problem of insufficient driving scenarios in existing technologies has been solved, enabling efficient training and improved safety of autonomous driving systems.
Patent Information
- Application Number
- CN202180074944.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-05
- Filing Date
- 2021-11-04
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2041-11-04
AI Technical Summary
Existing technologies struggle to generate a sufficient number and variety of driving scenarios under limited conditions, resulting in insufficient safety and reliability of autonomous vehicles in actual testing.
By receiving data from multiple sensors, fusing LiDAR point clouds and camera images, the system locates and classifies static objects and dynamic traffic participants, generates simulated driving scenarios, optimizes scenario characteristics through neural networks, and outputs simulated sensor data to expand the training data.
It increases the amount and diversity of training data for autonomous driving systems, enhances the safety and reliability of autonomous vehicles, and can simulate various driving scenarios to test and optimize driving functions.
Smart Images

Figure CN116529784B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a computer-implemented method for generating driving scenarios based on raw lidar data, a computer-readable data carrier, and a computer system. Background Technology
[0002] Autonomous driving promises unprecedented comfort and safety in everyday traffic. Despite significant investments from various companies, existing solutions remain limited to certain conditions or only achieve partially autonomous behavior. One reason for this is the lack of a sufficient number and variety of driving scenarios. Training and testing autonomous driving functions requires vast amounts of mileage to ensure safe operation. Therefore, it is not statistically possible, for example, to prove based on real-world testing on the road that autonomous vehicles are safer than human drivers in terms of fatal accidents.
[0003] Simulation can significantly increase the "mileage traveled". However, modeling suitable driving scenarios in a simulation environment is cumbersome, and the reproduction of recorded sensor data is limited to previously encountered driving scenarios. Summary of the Invention
[0004] Therefore, the objective of this invention is to provide a better method for generating sensor data for driving scenarios; and more importantly, to easily extend some variations to existing driving scenarios.
[0005] This task is solved by a method for generating simulated scenarios for land vehicles according to the present invention, a computer-readable data carrier according to the present invention, and a computer system according to the present invention.
[0006] Therefore, a computer-implemented method for generating simulated scenarios for vehicles, particularly land vehicles, is provided, comprising the following steps:
[0007] • Receive raw data, which includes multiple successive lidar point clouds, multiple successive camera images, and multiple successive velocity and / or acceleration data.
[0008] • Merge multiple lidar point clouds over a specific area into a common coordinate system to generate a merged point cloud.
[0009] • Locate and classify one or more static objects within the merged point cloud.
[0010] • Road information is generated based on merged point clouds, one or more static objects, and at least one camera image.
[0011] • Locate and classify one or more dynamic traffic participants within multiple successive lidar point clouds and generate trajectories for said one or more traffic participants.
[0012] • To create a simulation scenario based on the one or more static objects, road information, and the created trajectories for the one or more traffic participants, and
[0013] • Output the simulated scenario.
[0014] Static objects do not change their position over time, while the positions of traffic participants can change dynamically. The term dynamic traffic participants preferably also includes temporarily static traffic participants, such as parked cars, that is, traffic participants that are in motion at a specific point in time but can also remain stationary for a certain period of time.
[0015] The simulation scenario preferably describes continuous driving maneuvers, such as an overtaking maneuver occurring in an environment given by road information and static objects. Depending on the dynamic behavior or trajectories of traffic participants, the simulation scenario may also involve safety-critical driving maneuvers when, for example, there is a collision risk due to an oncoming vehicle during overtaking.
[0016] A specific area can refer to a geographically defined region by GPS coordinates. However, it can also refer to an area defined by recorded sensor data, such as a zone of the surrounding environment detected by environmental sensors.
[0017] The output simulation scenario may include storing one or more files on a data carrier and / or storing information in a database. The information in the files or database can then be accessed arbitrarily and frequently to generate, for example, sensor data for virtual driving tests. Therefore, existing driving scenarios can be used to test different autonomous driving functions and / or to simulate different environmental sensors. Alternatively, it can be specified that sensor data simulating existing driving scenarios is output directly.
[0018] The method of this invention focuses on lidar data and integrates scene generation steps, thereby ensuring that the simulated scene is not limited to fixed sensor data but includes relative coordinates for objects and traffic participants. Therefore, scenes of interest can be intelligently pre-selected, while existing scenes can also be supplemented.
[0019] In a preferred embodiment of the invention, the raw data includes synthetic sensor data, which is realistically generated by sensors in a simulated environment. Sensor data recorded by a real vehicle, synthetic sensor data, or a combination of recorded and synthetic sensor data can be used as input data for the method according to the invention.
[0020] A preferred embodiment of the invention provides an end-to-end pipeline with defined interfaces for all tools and operations / computations. This allows for the utilization of synergies between different tools, such as using scenario-based testing to enrich simulated scenarios or simulated sensor data generated by these scenarios.
[0021] In a preferred embodiment, the invention further includes a step of refining the simulated scene before outputting the simulated scene, particularly by refining at least one trajectory and / or adding at least one additional dynamic traffic participant. The refinement can be arbitrary or adaptive to achieve the desired characteristics of the scene. The advantage of this is that by adding simulated scenes, particularly those for key scenarios, sensor data for multiple scenarios can be simulated, and thus the amount of training data can be increased. This allows the user to train their model with a larger amount of relevant data, which, for example, leads to better perception algorithms.
[0022] In a preferred embodiment, the steps of refining and outputting the simulated scene are repeated, wherein another refinement is applied before outputting the simulated scene, thus compiling a set of simulated scenes or a group of simulated scenes. Each simulated scene may have metadata, such as descriptions of how many pedestrians appear in the simulated scene and / or cross the road. In particular, metadata can be derived from the location and identification of static objects and / or traffic participants. In addition, descriptions of road orientation, such as curve parameters, or descriptions of the surrounding environment can also be added to the simulated scene as metadata. The final data set preferably includes both the original data and extended synthetic point clouds. In a preferred embodiment, the computer used to generate the scene is connected to and / or includes a database server, wherein existing simulated scenes are stored in the database, and thus existing scenes can be used to supplement the group of simulated scenes.
[0023] In a preferred embodiment, at least one characteristic of a group of simulated scenarios is determined, and modified simulated scenarios are added to the group of simulated scenarios until the desired characteristic is met. The characteristic may in particular relate to a minimum number of simulated scenarios having specific features. As a characteristic of the group of simulated scenarios, for example, different types of simulated scenarios, such as city center scenarios, highway scenarios, and / or scenarios in which a predetermined object appears, or scenarios in which the predetermined object appears at at least a predetermined frequency, may be required. As a characteristic, it may also be required that a predetermined portion of the simulated scenarios in the group of simulated scenarios leads to a predetermined traffic situation, i.e., for example, describing an overtaking process and / or causing a collision hazard.
[0024] In a preferred embodiment, determining the characteristics of the group of simulated scenarios includes analyzing each modified simulated scenario using at least one neural network and / or performing at least one simulation of the modified simulated scenario. Performing the simulation can, for example, guarantee that the modified simulated scenario results in a collision hazard.
[0025] In a preferred embodiment, the characteristic is associated with at least one feature of the simulated scenario, particularly a representative characteristic of the statistical distribution of the simulated scenario, and the set of simulated scenarios is expanded to obtain a desired statistical distribution of the simulated scenario. The characteristic may, for example, describe whether and / or how many objects of a predetermined category appear in the simulated scenario. It may also be predetermined as a representative characteristic, with the set of simulated scenarios being large enough to allow testing to be performed with a predetermined confidence level. Thus, a predefined number of scenarios can be provided, for example, for different object categories or traffic participant categories, to provide sufficient data for machine learning.
[0026] The method preferably includes the steps of receiving a desired sensor configuration, generating simulated sensor data based on the simulated scene and the desired sensor configuration, and outputting the simulated sensor data. The simulated scene includes the spatial relationships of the scene and therefore contains sufficient information to generate sensor data from any environmental sensor.
[0027] The method particularly preferably includes the steps of training a neural network for perception using simulated sensor data and / or testing autonomous driving functions using simulated sensor data.
[0028] In a preferred embodiment, the received raw data has a lower resolution than the simulated sensor data. By first extracting an abstract scene from the raw data, driving scenarios recorded with older, lower-resolution sensors can also be reused in a new system with higher resolution.
[0029] In a preferred embodiment, the simulated sensor data includes multiple camera images. Scenes recorded by a LiDAR sensor can optionally or additionally be converted into images from the camera sensor.
[0030] The present invention also relates to a computer-readable data carrier containing instructions that, when executed by a processor of a computer system, cause the computer system to perform the method according to the present invention.
[0031] Furthermore, the present invention relates to a computer system comprising a processor, a human-machine interface, and non-volatile memory, wherein the non-volatile memory contains instructions that, when executed by the processor, cause the computer system to perform the method according to the present invention.
[0032] The processor relates to a general-purpose microprocessor, which is typically used as the central processing unit of a workstation computer, or it may include one or more processing elements suitable for performing specific computations, such as a graphics processing unit. In alternative embodiments of the invention, the processor may be replaced or supplemented by programmable logic components, such as field-programmable gate arrays, configured to perform a specific number of operations, and / or include an IP core microprocessor.
[0033] The invention will now be explained in more detail with reference to the accompanying drawings. Here, the same parts are designated with the same names. The embodiments shown are strongly illustrative; that is, the spacing and lateral and vertical dimensions are not faithful to scale, and their geometric relationships cannot be deduced without further explanation. Attached Figure Description
[0034] In the picture:
[0035] Figure 1 This is an exemplary diagram of a computer system;
[0036] Figure 2 This is a stereo view of an exemplary lidar point cloud;
[0037] Figure 3 It is an illustrative flowchart of an embodiment of the method according to the invention for generating simulated scenes; and
[0038] Figure 4 This is an example of a bird's-eye view of a synthesized point cloud. Detailed Implementation
[0039] Figure 1 An exemplary implementation of a computer system is shown.
[0040] The illustrated implementation includes a host PC with a monitor DIS and input devices such as a keyboard and a mouse.
[0041] The host PC includes: at least one processor CPU with one or more cores; selectively accessible working memory (RAM); and a number of devices connected to a local bus (such as PCI Express), which exchanges data with the CPU via a bus controller (BC). Devices include, for example, a graphics processing unit (GPU) for controlling a display, a USB controller for connecting peripheral devices, non-volatile memory such as a hard disk drive or solid-state drive, and a network interface (NC). The non-volatile memory may include instructions that, when executed by one or more cores of the CPU, cause the computer system to perform the method according to the invention.
[0042] In one embodiment of the invention, illustrated in the figure, via a programmed cloud, the host may include one or more servers comprising one or more computing units such as processors or FPGAs, wherein the servers are connected via a network to clients including display devices and input devices. Therefore, the method for generating the simulated scene can be performed partially or entirely on a remote server, for example, in a cloud computing setup. As an alternative to PC clients, the graphical user interface of the simulated environment can be displayed on portable computing devices, particularly laptops or smartphones.
[0043] Figure 2 A stereo view of an exemplary point cloud generated by a conventional LiDAR sensor is shown. The raw data has been labeled with bounding boxes surrounding the detected vehicles. The density of measurement points is high on the sides of the object facing the LiDAR sensor, while there are almost no measurement points on the rear sides due to occlusion. Further objects may also consist of only a few points, where possible.
[0044] In addition to LiDAR sensors, vehicles typically also have one or more cameras, receivers for satellite navigation signals (such as GPS), speed sensors (or wheel speed sensors), acceleration sensors, and yaw sensors. This data is preferably also stored during driving and can therefore be taken into account when generating simulated scenarios. Camera data, in addition to its high resolution, usually also provides color information, thus complementing LiDAR data or point clouds well.
[0045] Figure 3 A schematic flowchart of one embodiment of the method according to the invention for generating simulated scenes is shown.
[0046] The input data used to generate the scene is a point cloud recorded / collected at successive time points. Optionally, additional data, such as camera data, can be used to enrich the information in the data set, for example, when GPS data is not accurate enough. For this purpose, algorithms known per se can be used to simultaneously locate and map the scene.
[0047] In step S1 (fusion of lidar point clouds), the lidar point clouds of a specific area are merged or fused into a common coordinate system.
[0048] To construct a temporally valid / coherent scene, scans at different time points are correlated or relative 3D translations and 3D rotations between point clouds are determined. Information such as vehicle range, consisting of 2D translation and yaw or 3D rotation information and determined by vehicle sensors, is used for this purpose, along with satellite navigation data (GPS) consisting of 3D translation information. This is supplemented by lidar range, which provides relative 3D translations and 3D rotations using an Iterative Closest Point (ICP) algorithm. ICP algorithms known per se can be used here, such as those described in the paper "Sparse Iterative Closest Point" by Bouaziz et al., presented at the 2013 European Society for Graphics Geometric Processing Workshop. This information is then fused using graph-based optimization, which weights the given information (using its covariance) and calculates the resulting mileage. An exemplary algorithm for graph-based optimization is described in the paper "Graph-Based SLAM Tutorial" by Grisetti et al., published in the *Journal of Intelligent Transportation Systems* (IEEE, 2(4):31-43, 2010). The calculated mileage can then be used to fuse given sensor data (such as LiDAR, camera, etc.) into a common coordinate system. For labeling static objects, it is appropriate to merge different individual image point clouds into a single registered point cloud.
[0049] In step S2 (Location and classification of static objects), static objects in the registered or merged point cloud are labeled, i.e., located and identified.
[0050] Static object data includes buildings, vegetation, road infrastructure, etc. Each static object in the registered point cloud is labeled manually, semi-automatically, automatically, or by a combination of these methods. In a preferred embodiment, static objects within the registered or merged point cloud are automatically identified and filtered using algorithms known per se. In an alternative embodiment, the host computer can receive annotations from human annotators.
[0051] Using registered or merged, and therefore dense, point clouds offers numerous advantages in annotating static objects. Because there are significantly more points available for an object, determining the correct location and size of each individual object is much simpler. Furthermore, during driving, objects can be viewed from different perspectives, thus providing additional points on the object from LiDAR point clouds from all directions. Therefore, overall, much more accurate annotation of object boundaries can be achieved. In a point cloud from a single viewpoint, only points from that viewpoint are available for annotation. When someone observes an object, for example, from the front, it is difficult to determine the boundaries in the rear portion of the object because there is no information to help estimate the boundaries. Conversely, static objects are defined much more clearly in a merged point cloud that includes records / acquisitions from multiple perspectives.
[0052] In step S3 (generating road information), road information is generated based on the registered point cloud and camera data.
[0053] To generate road information, the point cloud is filtered to identify all points describing the road surface. These points are used to estimate a so-called ground plane, which represents the ground surface for a given lidar point cloud, or, in the case of a registered point cloud, the ground surface for the entire scene. In the next step, color information is extracted from images generated by one or more cameras, and this color information is projected onto the ground plane using intrinsic and extrinsic calibration information, particularly the camera's lens focal length and angle of view.
[0054] Then, using this image, roads are created in a top-down view. First, lane boundaries are identified and labeled. In the second step, so-called road segments and intersections are identified. A road segment is a portion of the road network with a constant number of lanes. In the next step, obstacles on the road, such as safety islands, are incorporated. The subsequent step is to label the road markings. Finally, the individual components are correctly linked together to merge all the components into a single road model that describes the geometry and semantics of the road network for that specific scenario.
[0055] In step S4 (Location and classification of traffic participants), dynamic objects or traffic participants in the successive point clouds are labeled.
[0056] Due to the dynamic behavior of traffic participants, these participants are labeled individually. Because traffic participants are in motion, it is impossible to use registered or merged point clouds; therefore, the labeling of dynamic objects or traffic participants is instead performed in a single image point cloud. Each dynamic object is labeled, that is, located and classified. Dynamic objects or traffic participants in the point cloud can refer to cars, trucks, delivery vehicles, motorcycles, cyclists, pedestrians, and / or animals. The host computer can, in principle, receive the results of manual or computer-aided labeling. In a preferred embodiment, labeling is performed automatically using algorithms known per se, particularly trained deep learning algorithms.
[0057] For annotation, it is appropriate to consider images taken by at least one camera in parallel with the LiDAR sensor, which, due to temporal overlap, necessarily display the same objects (provided the viewpoints overlap accordingly). Compared to sparse LiDAR point clouds, camera images contain more information, especially due to their higher resolution. In LiDAR point clouds, pedestrians at great distances (>100 meters) can be represented by only a single LiDAR point, but are clearly visible in camera images. Therefore, it is appropriate to use known algorithms, such as YOLO, for object recognition in camera images and to associate them with corresponding regions in the LiDAR point cloud. Camera information is therefore extremely helpful in identifying and classifying objects; furthermore, classification based on camera information can also help determine the size of objects. Then, a predetermined standard size is used or pre-defined according to the classification for different traffic participants. This facilitates the reproduction of traffic participants who have never been close enough to the detected vehicle, so that the size can be determined from the dense point distribution on the objects within the LiDAR point cloud.
[0058] After identifying and classifying traffic participants in a single image LiDAR point cloud and, if present, camera images, a temporal chain is determined for each object. Each traffic participant is assigned a unique ID for all individual records / acquisitions, i.e., the point cloud and images in which the traffic participant appears. In a preferred embodiment, the first image in which the traffic participant appears is used to identify and classify the object, and then the corresponding field or label is passed to subsequent images. In an alternative preferred embodiment, traffic participants are tracked through successive images, i.e., the algorithm tracks the traffic participants. Here, identified objects are compared across multiple frames, and when the overlap of faces exceeds a predetermined threshold, i.e., when consistency is identified, they are assigned to the same traffic participant, thus the identified objects have the same ID or unique identification code. Both techniques enable the generation of consistent temporal chains from successive camera images and LiDAR point clouds. Thus, temporal and spatial trajectories are generated for each dynamic object, which can be used to describe the behavior of traffic participants in the simulation. Therefore, dynamic objects are associated along the timeline to obtain the trajectory of traffic participants.
[0059] In step S5 (generating a simulation scene), a recreated scene for the simulation environment is created from static objects, road information, and the trajectories of traffic participants.
[0060] Preferably, the information acquired during the annotation of the raw data in steps S2 to S4 is automatically transmitted to the simulation environment. This information includes, for example, the size, category, and attributes of static objects, and, for traffic participants, their trajectories. This information is appropriately stored in a suitable file exchange format, such as a JSON file; JSON here refers to JavaScript Object Markup, a common language used to describe the scenario.
[0061] First, the road information contained in the JSON file is transmitted to the appropriate road model in the simulation environment. This includes the road geometry and all semantics determined or imported during annotation, such as which lanes merge into which other lanes at the boundaries between lane segments, or which lane to take when crossing an intersection.
[0062] Next, all the labeled static objects are placed in the scene. This involves reading or exporting the classification for each object, along with some additional attributes, from a JSON file. Based on this information, a suitable 3D object is selected from the 3D asset library and placed in its appropriate location within the scene.
[0063] Finally, traffic participants were placed in the scene and made to move according to marked trajectories. Road signs derived from the trajectories were used for this purpose, and were placed in the scene with the necessary timing information. The simulated driver model then reproduced the behavior of the vehicle recorded during the test drive.
[0064] Therefore, a "reproducible scenario" is generated in the simulation environment using road information, static objects, and the trajectories of traffic participants. Here, all traffic participants behave exactly as in the recorded scenario, and the reproducible scenario can be reproduced arbitrarily frequently. This enables a high degree of repeatability of the scenarios recorded in the simulation environment, for example, to check whether a new version of the driving function exhibits erroneous behaviors similar to those during a test drive.
[0065] In a further step, these recreated scenarios are abstracted into "logical scenarios." Logical scenarios are derived from the recreated scenarios by abstracting the specific behaviors of different traffic participants into maneuvers, in which parameters can vary within predetermined ranges for these maneuvers. The parameters of a logical scenario may include, in particular, relative position, speed, acceleration, the starting point for specific behaviors such as lane changes, and relationships between different objects. By deriving or inserting maneuvers with reasonable parameter ranges, it becomes possible to implement variations of the recorded scenarios in a simulated environment. This lays the foundation for subsequent expansion of the simulated scenario set.
[0066] In step S6 (Characteristics OK?), the characteristics of a group of existing scenarios (one or more already created scenarios) are determined and compared with expected values. Depending on whether the expected characteristics are met, proceed to step S7 or S9. Here, one or more characteristics of a single scenario can be observed; for example, it can be required that the group only includes scenarios exhibiting the required characteristics, or representative characteristics of the group, such as frequency distribution, can be determined from simulated scenarios. These can be formulated as simple or comprehensive criteria that the data group must meet.
[0067] Analysis of the data set in terms of desired specifications (“Delta analysis”) may include, but is not limited to, questions such as: How are the different objects distributed in the data set? What is the target distribution of the database inventory for the desired application? In particular, minimum occurrence frequencies for different categories of traffic participants may be required. This can be examined based on metadata or parameters of the simulated scenario.
[0068] During the annotation of the raw data in steps S2 to S4, features suitable for analysis of the data set can be appropriately determined. It is preferable to analyze the raw data using neural networks to discern the distribution of these features in the actual LiDAR point cloud. For this purpose, a series of object recognition networks, object trackers, and attribute recognition networks are used, which automatically identify objects within the scene and preferably add attributes to these objects specific to the application context. Because these networks are originally required to create the simulated scene, only a small additional overhead is incurred. The determined features can be stored separately and assigned to the simulated scene. The automatically identified objects and their attributes can then be used to analyze the features of the initially recorded scene.
[0069] If the raw data already contains groups of multiple scenarios, the characteristics of these raw data groups can be compared to application-specific distributions. These characteristics can include, in particular, the frequency distribution of object categories, lighting, and weather conditions (scene attributes). The comparison results can determine specifications for data enrichment. For example, it may be confirmed that a particular object category or attribute is underrepresented in a given data group and therefore corresponding simulated scenarios with these objects or attributes must be added. An excessively low frequency may be attributable to the fact that the observed objects or attributes are rare in the specific area where the data was recorded and / or happen to be rare at the time of recording.
[0070] However, when analyzing a dataset, it may be necessary to have a certain number of simulated scenarios in which pedestrians cross the road and / or collision hazards occur. Therefore, it is optional to simulate the relevant scenarios to determine further specifications and the selection of useful data for enrichment through scenario-based testing. As a result, preferably, a specification for enriching the data is defined, specifying which scenarios are required.
[0071] Alternatively, scenario-based testing can be used to identify suitable scenarios to refine the specifications for data expansion. When there is particular interest, for example, in key scenarios within a city center, scenario-based testing can be used to identify scenarios with specific performance indicators (KPIs) that meet all requirements. Correspondingly, the expanded dataset can be limited to selected and therefore relevant scenarios, rather than simply arranged by KPI.
[0072] If no specific feature is given (No), then in step S9 (supplementing the simulation scenario), the data set is expanded by changing the simulation scenario.
[0073] In the expansion of the data set performed appropriately in a simulated environment in step S9, the user can define any statistical distribution they wish to achieve in their expanded data set, starting with a distribution automatically determined in the original data. Preferably, this distribution can be achieved by generating a digital twin from an existing, at least partially automatically annotated, scene. This digital twin can then be orchestrated by adding traffic participants with specific object categories and / or different behaviors or modified trajectories. A virtual data detection vehicle equipped with arbitrary sensors is placed in the added, orchestrated simulated scene and then used to generate new, synthetic sensor data. The sensor equipment can be arbitrarily different from the sensor equipment used to record the original data. This not only helps to supplement existing data but also helps to use the recorded scene as a basis for generating sensor data for new / different sensor equipment when the sensor devices of the observed vehicle change.
[0074] Within the simulated environment, the simulated scene group can be expanded not only by adding traffic participants, but also by changing contrast, weather, and / or lighting conditions.
[0075] Optionally, existing scenarios from the scenario database are included in the expansion of the simulation scenario group; these scenarios can be created using earlier recorded raw data. Such a scenario database improves the efficiency of data expansion; scenarios for different application scenarios can be stored in the scenario database and labeled with features. By filtering using the stored features, suitable scenarios can be easily retrieved from the database.
[0076] If the characteristic is satisfied (yes), then in step S7 (outputting sensor data to the simulated scene), simulated sensor data is output to the simulated scene or the entire group of simulated scenes. Simulated sensor data for one or more environmental sensors can be generated using sensor configurations such as altitude above ground and predetermined resolution; in addition to LiDAR data, camera images can also be generated from the simulated scene using camera parameters. In the illustrated embodiment, simulated sensor data is output for the entire group of scenes; each individual scene can typically be output independently, either alternatively or supplementarily, after creation. If the determination of a feature requires scene simulation, then, for example, it needs to be output beforehand.
[0077] By using neural networks, all scenarios can be treated as simulated sensor data or actual sensor data output. Neural networks, such as Generative Adversarial Networks (GANs), can mimic the characteristics of different sensors, including sensor noise. This can simulate the characteristics of the sensor originally used to record the raw data, or it can simulate the characteristics of completely different sensors. Therefore, on the one hand, scenarios as similar as possible to the original data can be used for algorithm training or testing. On the other hand, driving scenarios can be virtually recorded using different LiDAR sensors, but other imaging sensors, especially cameras, can also be used to generate and utilize scenarios that mimic the original data, thus enabling training or testing of algorithms. Conversely, virtual recordings of driving scenarios can also be generated and used using different LiDAR sensors, but also other imaging sensors, such as cameras.
[0078] In step S8 (testing autonomous driving functions using sensor data), one or more output scenarios, that is, self-contained sets of simulated sensor data, are used to test the autonomous driving functions. One or more output scenario alternatives can also be used to train the autonomous driving functions.
[0079] This invention augments lidar sensor data sets by intelligently supplementing missing data from the original data set. By outputting the supplementary data as actual sensor data, this additional data can be directly used to train and / or test autonomous driving functions. Alternatively, it can be specified that data sets be created or used that include both the initial raw data and synthesized sensor data.
[0080] As a specific application scenario, the aforementioned pipeline can be used to convert data from one sensor type into a synthetic point cloud of the same environment, and in scenarios involving another sensor type. Thus, data from, for example, older sensors can be converted into synthetic data representing modern sensors. Older sensor data can be used accordingly to expand currently recorded data with newer sensor types. For this application scenario, which outputs "recycled" sensor data, a processing pipeline is preferably used, specifically including steps S1, S2, S3, S4, S5, and S7.
[0081] Figure 4 An aerial view of a synthetic point cloud generated from sensor data output from a simulated scene is shown. In this view, each traffic participant is represented by a bounding box.
[0082] The method of this invention allows for the supplementation of simulated sensor data with measured sensor data from the entering scenario to adapt to changing simulated scenarios. This enables better training of perception algorithms and allows for more extensive testing of autonomous driving capabilities.
Claims
1. A computer-implemented method for generating simulated scenarios for vehicles, the method comprising the following steps: Receive raw data, which includes multiple successive lidar point clouds, multiple successive camera images, and multiple successive velocity and / or acceleration data. The system locates and classifies one or more dynamic traffic participants within multiple successive lidar point clouds and generates trajectories for the one or more traffic participants. Create simulation scenarios, and Output the simulated scenario. Its characteristics are as follows: The multiple lidar point clouds in a specific area are fused into a common coordinate system to generate a merged point cloud. Locate and classify one or more static objects within the merged point cloud. Road information is generated based on merged point clouds, one or more static objects, and at least one camera image. The simulation scenario is created based on one or more static objects, road information, and the created trajectories for one or more traffic participants.
2. The method according to claim 1, wherein, The method further includes the following step: correcting the simulated scene before outputting the simulated scene.
3. The method according to claim 2, wherein, Repeat the steps of modifying and outputting the simulation scene, and apply another modification before outputting the simulation scene to compile a group of simulation scenes.
4. The method according to claim 3, wherein, Determine at least one characteristic of the simulated scenario group, and add modified simulated scenarios to the simulated scenario group until the desired characteristic is met.
5. The method according to claim 4, wherein, Determining the characteristics of the group of simulation scenarios includes: analyzing each modified simulation scenario and / or performing at least one simulation of the modified simulation scenario using at least one neural network.
6. The method according to claim 4 or 5, wherein, The characteristic is associated with at least one feature of the simulated scenario and expands the group of simulated scenarios to obtain the desired statistical distribution of the simulated scenarios.
7. The method according to any one of claims 1 to 5, wherein, The output of the simulated scenario includes: Receive the desired sensor configuration. Simulated sensor data is generated based on the simulated scenario and the desired sensor configuration, and Output the simulated sensor data.
8. The method according to claim 7, wherein, The method further includes the following steps: Training a neural network for perception using the simulated sensor data, and / or The autonomous driving function was tested using the simulated sensor data.
9. The method according to claim 7, wherein, The received raw data has a lower resolution than the simulated sensor data.
10. The method according to claim 7, wherein, The simulated sensor data includes images from multiple cameras.
11. The method according to claim 1, wherein, The method further includes the step of modifying the simulated scenario by correcting at least one trajectory and / or adding at least one additional dynamic traffic participant before outputting the simulated scenario.
12. A computer-readable data carrier comprising an instruction that, when executed by a processor of a computer system, causes the computer system to perform the method according to any one of claims 1 to 11.
13. A computer system, said computer system comprising a processor, a human-computer interface, and non-volatile memory, wherein, The non-volatile memory contains instructions that, when executed by a processor, cause a computer system to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Realistic 3D virtual world creation and simulation for training automated driving systems
CN109643125A
Voxel based ground plane estimation and object segmentation
CN110770790A