Method for generating a taged training data set
By using LiDAR sensors and RRO prediction neural networks to generate an annotated training data set, the problem of difficulty in effectively generating annotated training data for training ADS perception algorithms in the prior art is solved, and more efficient and consistent data generation is achieved, and the performance of the autonomous driving system is improved.
Patent Information
- Application Number
- CN202411890479.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively generate labeled training data for training autonomous driving systems (ADS) perception algorithms, resulting in poor performance of autonomous driving systems in dealing with complex traffic conditions.
The neural network is predicted by using the frame sequence and road reference object (RRO) captured by the LiDAR sensor, the RRO position data set is predicted in the frame, and the global RRO position data set is populated by matching the subset of the RRO position data of different frames, thereby generating the annotated training data set.
This method can generate labeled training data sets more efficiently and consistently, reduce dependence on human operators and improve the performance of autonomous driving systems under different traffic conditions.
Smart Images

Figure CN120198869A_ABST
Abstract
Description
Technical Field
[0001] The disclosed technology relates to methods and systems for generating an annotated training dataset for training a perception algorithm of an autonomous driving system (ADS) of a vehicle. In particular, but not exclusively, the disclosed technology relates to more efficiently generating the annotated training data. Background Art
[0002] In the past few years, there have been significant leaps in the development within the field of autonomous driving. For example, it is common practice today to use deep neural networks as part of a system for controlling a vehicle or assisting a vehicle driver. In addition to the development within the field of data processing, there have also been improvements in sensors. For instance, by being able to capture reliable sensor data at a high frequency and at a competitive cost, improved input can be provided for the deep neural networks. Thus, through the development within the field of data processing (such as the development of deep neural network technology), combined with more efficient and reliable sensors, a more reliable and robust system can be manufactured for autonomous driving.
[0003] However, even though there are technologies in the form of deep neural networks available today for generating reliable models (for autonomous driving) and improved sensors (such as LiDAR (Light Detection and Ranging) sensors), there are still challenges. One such challenge is how to effectively generate the annotated training data so that the neural networks or other models can be fully trained. This becomes even more important as the demand for autonomous driving systems continues to increase. In other words, in order to handle increasingly diverse traffic situations, more complex models are required. For example, if a model is restricted to being used only on a predetermined road section under certain road conditions, a less complex model can be used compared to a model that can handle general roads without being restricted to certain road conditions.
[0004] Although more complex models for data processing are being used and sensors that generate more data are being used, it is still common for human operators to manually generate the training data. For example, it is common today to generate the training data by presenting the sensor data to an operator and the operator identifying the objects in the sensor data. To make it easier for the operator to identify the objects, different software programs have been developed. Further, to facilitate the operator in identifying objects that can be lane markings, other vehicles, pedestrians, wild animals, etc., the sensor data can be processed before being displayed to the operator.
[0005] Even though autonomous driving systems that can handle various different traffic situations can be produced today, ADS can be further improved by being able to more effectively generate training data for the perception module in an autonomous driving system (ADS). Since accessing large amounts of training data is often a bottleneck, being able to more effectively generate such data provides the possibility for significant improvement of ADS. SUMMARY OF THE INVENTION
[0006] The techniques disclosed herein are aimed at alleviating, mitigating, or eliminating one or more of the above-identified deficiencies and drawbacks in the prior art to address various problems related to generating labeled training data for training the perception algorithms of ADS.
[0007] Aspects and embodiments of the disclosed invention are defined below and in the appended independent and dependent claims.
[0008] A first aspect of the disclosed technique includes a computer-implemented method for generating a labeled training data set for training the perception algorithms of an ADS of a vehicle. Method 300 includes: obtaining a sequence of frames captured by one or more light detection and ranging (LiDAR) sensors, and predicting, by using a road reference object (RRO) prediction neural network, an RRO position data set for each subset of at least one subset of the frames, wherein each RRO position data set respectively includes one or more RRO position data subsets of one or more RROs, and wherein each RRO position data subset is related to the spatial information of one RRO found in the frame. The method further includes matching one or more RRO position data subsets of one frame with one or more RRO position data subsets of at least one other frame to populate a global RRO position data set, wherein the global RRO position data set includes one or more global RRO position data subsets, and wherein each global RRO position data subset in the global RRO position data set has a corresponding RRO position data subset in at least two of the frames, and forming a labeled training data set based on the sequence and the global RRO position data set.
[0009] As the name implies, the RRO position data set and the RRO position data subsets are related to spatial information. However, in addition to this, these may include information about the type of the RRO, such as lane markings or median barriers, the cardinality of the RRO (i.e., the number of elements included), and other attributes that can be linked to the RRO. This information can be linked to the RRO position data set (i.e., all RROs in the frame), or can be separately linked to the RRO position data subsets (i.e., the attributes can be linked to individual RROs in the frame).
[0010] The RRO position data set may include information about all the RROs identified in a frame. A subset of the RRO position data may include information about one of these RROs. In other words, the RRO position data set may include multiple subsets of RRO position data. For example, in the case where a frame includes a single RRO, the RRO position data set includes one subset of RRO position data.
[0011] Even though the use of LiDAR to obtain a sequence of frames has been described above, one or several other types of sensors may also be used. For example, one or several cameras or more generally image-based sensors may be used to obtain the sequence.
[0012] A second aspect of the disclosed technology includes a computer program product including instructions that, when executed by a computing device, cause the computing device to perform a method according to any of the embodiments of the first aspect disclosed herein. Through this aspect of the disclosed technology, there are advantages and preferred features similar to those of the other aspects.
[0013] A third aspect of the disclosed technology includes a (non-transitory) computer-readable storage medium including instructions that, when executed by a computing device, cause the computing device to perform a method according to any of the embodiments of the first aspect disclosed herein. Through this aspect of the disclosed technology, there are advantages and preferred features similar to those of the other aspects.
[0014] As used herein, the term "non-transitory" is intended to describe a computer-readable storage medium (or "memory") that excludes propagating electromagnetic signals, but is not intended to otherwise limit the type of physical computer-readable storage devices encompassed by the phrase computer-readable medium or memory. For example, the term "non-transitory computer-readable medium" or "tangible memory" is intended to encompass storage device types that include, for example, random access memory (RAM) which does not necessarily store information permanently. Program instructions and data stored in a non-transitory form on a tangible computer-accessible storage medium may further be transmitted via a transmission medium or a signal such as an electrical, electromagnetic, or digital signal, which may be conveyed via a communication medium such as a network and / or a wireless link. Thus, as used herein, the term "non-transitory" is a limitation on the medium itself (i.e., tangible, rather than the signal), rather than a limitation on data storage persistence (e.g., RAM versus ROM).
[0015] The fourth aspect of the disclosed technology includes a device for generating an annotated training dataset for training a perception algorithm of an autonomous driving system (ADS) of a vehicle. The device includes a control circuit configured to: obtain a sequence of frames captured by at least one LiDAR sensor, predict a RRO position dataset for each subset of at least one subset of the frames by using a road reference object (RRO) prediction neural network, wherein each RRO position dataset respectively includes one or more RRO position data subsets of one or more RROs, wherein each RRO position data subset is related to the spatial information of one RRO found in the frame, match one or more RRO position data subsets of one frame with one or more RRO position data subsets of at least one other frame to populate a global RRO position dataset, wherein the global RRO position dataset includes one or more global RRO position data subsets, wherein each global RRO position data subset in the global RRO position dataset has a corresponding RRO position data subset in at least two frames in the frames, and form an annotated training dataset based on the sequence and the global RRO position dataset. Through this aspect of the disclosed technology, there are similar advantages and preferred features as in other aspects.
[0016] The fifth aspect of the disclosed technology includes a vehicle including: an ADS including a perception algorithm, at least one LiDAR sensor, and a device for generating an annotated training dataset according to the second aspect for training the perception algorithm of the ADS. Through this aspect of the disclosed technology, there are similar advantages and preferred features as in other aspects.
[0017] The disclosed aspects and preferred embodiments can be appropriately combined with each other in any manner obvious to those of ordinary skill in the art, such that one or more features or embodiments related to one aspect can also be regarded as related to another aspect or an embodiment of another aspect.
[0018] The advantages of some embodiments are that, since the need for manual intervention can be reduced or even completely avoided in some cases, an annotated training dataset can be provided more effectively. In addition to being more efficient, compared with the current common practice of having a large group of human operators who may make mistakes or interpret instructions differently to perform annotation, the annotation of training data can also be carried out in a more consistent manner. Compared with the effect that can be obtained by manually generating training data, the effect of being able to generate training data on a larger scale is that the ADS can be trained at different levels, which in turn provides an improvement in the ADS.
[0019] Further embodiments are defined in the dependent claims. It should be emphasized that, as used in this specification, the terms "comprise / comprising" are used to specify the presence of the recited features, integers, steps, or components. It does not preclude the presence or addition of one or more other features, integers, steps, components, or groups thereof.
[0020] These and other features and advantages of the disclosed technology will be further clarified below with reference to the embodiments described in the following text. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] When combined with the accompanying drawings, the above aspects, features, and advantages of the disclosed technology will be more fully understood by reference to the following illustrative and non - limiting detailed description of exemplary embodiments of the present disclosure, in which:
[0022] Figure 1 Illustrate two vehicles equipped with sensors according to some embodiments.
[0023] Figure 2A Illustrate a scene formed by multiple LiDAR scans according to some embodiments.
[0024] Figure 2B Illustrate a scene Figure 2A marked with different forms of RRO according to some embodiments.
[0025] Figure 3 A flowchart for illustrating a method for generating a labeled training dataset according to some embodiments, the labeled training dataset being used to train the perception algorithm of the ADS of a vehicle.
[0026] Figure 4 For illustrating according to some embodiments Figure 3 Another flowchart of a specific example of the method illustrated by the flowchart in
[0027] Figure 5 A schematic diagram of a vehicle equipped with an ADS according to some embodiments.
[0028] Figure 6 For forming according to some embodiments Figure 5 A schematic diagram of a device that is part of a vehicle, wherein the device is arranged to generate a labeled training dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present disclosure will now be described in detail with reference to the accompanying drawings, in which some example embodiments of the disclosed technology are shown. However, the disclosed technology may be embodied in other forms and should not be construed as limited to the disclosed example embodiments. The disclosed example embodiments are provided to fully convey the scope of the disclosed technology to those skilled in the art. Those skilled in the art will understand that the steps, services, and functions explained herein may be implemented using separate hardware circuits, using software working in conjunction with a programmed microprocessor or a general-purpose computer, using one or more application-specific integrated circuits (ASICs), using one or more field-programmable gate arrays (FPGAs), and / or using one or more digital signal processors (DSPs).
[0030] It will also be recognized that when the present disclosure is described in the form of a method, it may also be embodied in a device including one or more processors and one or more memories coupled to the one or more processors, wherein computer code is loaded to implement the method. For example, in some embodiments, the one or more memories may store one or more computer programs that, when executed by the one or more processors, cause the device to perform the steps, services, and functions disclosed herein.
[0031] It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. It should be noted that, as used in the specification and the appended claims, the articles "a", "an", "the", and "said" are intended to mean that there is one or more elements, unless the context clearly indicates otherwise. Thus, for example, in some contexts, a reference to "a unit" or "the unit" may refer to more than one unit, etc. Further, the words "comprising", "including", and "containing" do not exclude other elements or steps. It should be emphasized that when used in this specification, the term "comprise / comprising" is used to specify the presence of the recited features, integers, steps, or components. It does not exclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. The term "and / or" should be construed to mean "both" as well as each as an alternative.
[0032] It should also be understood that although terms such as first, second, etc. may be used herein to describe various elements or features, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the embodiments, a first signal may be referred to as a second signal, and similarly, a second signal may be referred to as a first signal. The first signal and the second signal are both signals, but they are not the same signal.
[0033] Figure 1 The first vehicle 100a and the second vehicle 100b on the road 102 are generally illustrated from above. As an example, the first vehicle 100a (illustrated as an automobile herein) is provided with a plurality of sensors 104a. If image-based sensors or other sensors with a limited field of view (for example, less than 180 degrees) are used, using a plurality of sensors may be a preferred option. As illustrated, the second vehicle 100b is provided with a single sensor 104b, for example. In the case where the sensor is a LiDAR (Light Detection and Ranging) sensor or other sensors that can scan 360 degrees or at least greater than 180 degrees, using such a single sensor 104b may be a preferred option. Additionally, several LiDAR sensors or a single image-based sensor may also be used. For example, several LiDAR sensors may be used, each LiDAR sensor providing a 120-degree field of view.
[0034] The plurality of sensors 104 of the first vehicle 100a and the sensor 104b of the second vehicle 100b may be configured to identify road reference objects (RROs) 106a to 160c during operation, and based on the detection of these RROs, control steering, acceleration, braking, and any other actions that form part of autonomous driving. The RROs can have different forms. For example, as illustrated, the RROs can include lane markings 106a in the form of solid lines and lane markings 106b in the form of non-solid (dashed) lines, and can also include RROs 106c in the form of a median guardrail or other objects that define the road without being marked on the road 102.
[0035] In addition to Figure 1 the sensors 104a, 104b illustrated above, the vehicles 100a, 100b may also include other sensors that are not specifically arranged to detect or identify the RROs 106a to 160c. For example, in addition to the sensors 104a to 140b, there may be sensors in the form of image sensors provided at the rear end of the vehicle to assist the driver when reversing, or to automatically reverse without driver participation. Additionally, different types of distance sensors (sometimes also referred to as proximity sensors (such as radar)) may be provided on the vehicle and can be used, for example, when a pedestrian approaches the vehicle.
[0036] If LiDAR sensors are used, scans can be performed and these scans can then be aggregated into a single scene 200 as Figure 2A illustrated above. In this single scene 200, different RROs can be identified. For example, as illustrated, RROs 106b in the form of lane markings (dashed lines) and RROs 106a in the form of solid lines can be identified.
[0037] When generating training data for a deep neural network (DNN) or other model that can be trained to automatically detect RROs, a single scenario can be provided to a human operator, such as Figure 2A as illustrated by way of example in, for creating a polyline, for example, representing the RRO in scenario 200. Additionally, the operator is typically also required to identify detections made by the LiDAR sensor that do not represent an RRO. In other words, the operator is typically required to identify not only positives (i.e., the presence of an RRO), but also negatives (i.e., the point cloud that does not represent an RRO).
[0038] To improve the efficiency of generating training data (also referred to herein as the labeled training data set), a so-called RRO prediction neural network can be used. In short, this neural network can be trained to predict (sometimes referred to as detect) the dashed lane markings 106a, the solid lane markings 106b, the median guardrail 106c, and other RROs that can be useful for autonomous driving.
[0039] Once the predictions are made, these predictions can be matched. Matching can include determining the distance between a point cloud in a first frame (i.e., a detection associated with one of the RROs) and the same point cloud in a second frame, and in the case where the distance is below a threshold, this point cloud can be considered to be associated with the actual RRO depicted in the sensor data. On the other hand, in the case where the distance is above the threshold, the point cloud can be excluded as a point cloud associated with the actual RRO depicted in the sensor data (also referred to as a frame or a scan). In other words, in the case where the point cloud cannot be found in both the first frame and the second frame, the point cloud is considered not to be associated with any actual RRO depicted, i.e., it is negative. Another example of matching is to determine the distance by using the nodes of a polygon representing one of the RROs. For example, in the case where the RRO is a lane marking, this can be represented by a polygon, and the nodes of this polygon can be used to determine a point, such as the midpoint of the polygon. This point can be compared with the corresponding point in another frame to determine the distance.
[0040] By correlating these frames (sometimes referred to as scans) with each other, incorrect predictions made by the RRO prediction neural network can be identified. Two frames can be correlated with each other because they are temporally correlated (i.e., the times at which they are captured are close such that one or more RROs are depicted in both frames) and / or because the two frames are spatially correlated (i.e., the two frames are captured by two different sensors with at least partially overlapping fields of view such that one or more RROs are depicted in both frames).
[0041] To track the found matching point clouds, a global RRO position dataset can be used. If matching point clouds are found in the first and second frames, this point cloud can be added to the global RRO position dataset. As an alternative, instead of adding the point cloud, the RROs (e.g., lane markings) extracted from the point cloud can be added. Since several RROs can be depicted in each frame, the RRO position dataset for each frame can include one or more subsets of RRO position data each associated with one RRO found in the frame. In the same way, the global RRO position dataset can include one or more subsets of global RRO position data. Since the global RRO position dataset is populated only if there are point clouds representing RROs in at least two frames, each subset of global RRO position data has a corresponding subset of RRO position data in at least two frames.
[0042] Once the RRO position dataset has been predicted using the RRO prediction neural network and one or more RRO position datasets have been matched such that a global RRO position dataset is formed, a labeled training dataset of a sequence of frames including at least the first and second frames and the global RRO position dataset can be formed. Thus, with the method proposed in this article, the positions of RROs in different frames can be automatically predicted using the RRO prediction neural network instead of manually identifying and labeling different point clouds in different frames to form a labeled training dataset. Since the subsets of RRO position data each associated with one RRO in a frame may include false positives, this can be overcome by matching different subsets of RRO position data in different frames such that mismatched RROs can be removed. By having a large number of frames and having most RROs depicted in multiple frames, the risk of removing true positives can be kept at a low level. In the same way, in the case of false negatives in a frame (i.e., there is no RRO, but an RRO is still (erroneously) reflected by the subset of RRO position data associated with that frame), these must be present in at least two frames in order to be reflected in the global RRO position dataset. As an effect, incorrect detections made in a single frame can be corrected during the matching.
[0043] The tolerance for prediction can be adjusted in accordance with the matching criteria so as to achieve sufficient overall performance. When setting the prediction tolerance and the matching criteria, the frame rate, i.e., the time between two subsequent frames, can be taken into account. For example, in the case of capturing frames at a high frequency, it is highly likely that the RRO is depicted in multiple frames. This, in turn, can ensure that the matching criteria can be such that the RRO should be present in at least three consecutive frames and satisfy the distance threshold condition. As an effect, the prediction tolerance can be set such that there are deliberately created false positives (i.e., a large number of false positives are created), thereby reducing the risk of missing true positives. On the contrary, compared with the case of capturing frames with high-frequency frames, in the case of capturing frames at a low frequency, it is less likely that the RRO is depicted in several frames. As an effect, it may be more difficult to identify and remove false positives during the matching step. For this reason, in this case, the prediction tolerance can be set more strictly compared with high-frequency frame capture.
[0044] In Figure 2B Scenario 200 with an identified RRO is illustrated by way of example. As illustrated, the RRO can have different forms. For example, the RRO can be lane markings 202a to 202e that divide two lanes of a road into a right lane and a left lane. A further example is that the RRO can be solid line markings 204a to 204b that define the side boundaries of a road. Still another example is that the RRO can be shadow area markings 208a to 208f. The identified RRO can be indicated by a polyline that is readily understandable to those skilled in the art.
[0045] Figure 3 A flowchart of a method 300 for generating an annotated training dataset for a perception algorithm (sometimes referred to as a perception module) that can be used to train an ADS is illustrated. The method includes obtaining 302 a sequence of frames captured by one or more LiDAR sensors. Thereafter, by using an RRO prediction neural network, predicting 304 an RRO position dataset for each subset of at least one subset of the frames. Each RRO position dataset can respectively include one or more RRO position data subsets of one or more RROs. Once the prediction is made, the method can further include matching 306 one or more RRO position data subsets of one frame with one or more RRO position data subsets of at least one other frame to populate a global RRO position dataset. The global RRO position dataset can include one or more global RRO position data subsets. Each global RRO position data subset in the global RRO position dataset can have a corresponding RRO position data subset in at least two of the frames. Finally, the method can include forming 308 an annotated training dataset based on the sequence and the global RRO position dataset.
[0046] As described above, each frame may include an RRO position data set, which in turn may include one or more RRO position data subsets. Each RRO position data subset in the RRO position data subset is associated with one RRO found in the frame. The global RRO position data set may include one or more global RRO position data subsets. As the name implies, the global RRO position data set covers several frames. As described above, in order to form a global RRO position data subset of an RRO, the RRO should be predicted by the RRO prediction neural network in at least two of the frames.
[0047] Method 300 can be used for fully automatic annotation, that is, annotation without human intervention, or can be completed as pre-annotation for a human operator. For example, in the latter case, the forming step may involve a human operator confirming or rejecting the annotation suggested by the machine.
[0048] The matching 306 may then include: for each RRO position data subset in one or more RRO position data subsets of a frame, for each RRO position data subset in one or more RRO position data subsets of at least one other frame, determining 310 the distance between the RRO position data subset of the one frame and the RRO position data subset of the at least one other frame, and in the case where the distance is below the minimum distance threshold, populating 312 the RRO position data subset of the one frame into the global RRO position data set. Even though using distance (e.g., minimum distance) has proven to be a viable option, other measures can also be used to determine that the RRO position data subsets of different frames may be associated with the same RRO.
[0049] Optionally, in order to further reduce the risk of false positives in the training data set for annotation, after the matching 306, the method may further include removing 314 outliers from the global RRO position data set. An outlier may be such a global RRO position data subset in the global RRO position data set that is significantly different from the rest of the global RRO position data subsets in the global RRO position data set.
[0050] The method may include aggregating 316 the global RRO position data set by adjusting the global RRO position data subsets placed inside a defined window. For example, the global RRO position data set may be adjusted to align it along a polyline formed by multiple lane markings. The polyline may be formed by averaging the detected positions of the lane markings within the window. The aggregating step 316 may then include obtaining 318 a lane marking tracking score data set from the lane marking tracking device, determining 320 weights based on the lane marking tracking score data set, and adjusting 322 the aggregated global RRO position data set using the weights. By also considering the weights from the lane marking tracking device, the aggregation can be further improved.
[0051] Method 300 is preferably a computer-implemented method performed by a processing system of a vehicle equipped with ADS. The processing system may include, for example, one or more processors and one or more memories coupled to the one or more processors, where the one or more memories store one or more programs that, when executed by the one or more processors, perform the steps, services, and functions of method 300 disclosed herein.
[0052] The executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0053] Figure 4 To illustrate the flowchart of the more detailed method 400 provided by way of example. Consistent with Figure 3 the method 300 illustrated in, a set of inputs forming sequence 402 may be provided. As described above, the inputs may be scans from a LiDAR sensor, but may also be inputs from other types of sensors. An RRO prediction neural network (illustrated herein as a deep neural network 404 trained to detect 3D lane markings) may be used to generate a set of predicted 3D lane markings 406 for each input. This set 406 may be fed into a matching function 408, where the predicted 3D lane markings from different inputs are compared. After performing the matching, a set of matched 3D lane markings 410 may be provided. Each match may be a set of detections, each set of detections from each input. Thereafter, the set of matched 3D lane markings 410 may be fed into an aggregation function 412 that aggregates the matched 3D lane markings into one. For example, if subsets of global RRO position data are placed to be spatially close to each other, these subsets of global RRO position data may be aggregated into one and the same subset of global RRO position data. As input to the aggregation function, a lane marking tracking score data set may be provided. By using this data provided by the lane marking tracking device, historical data may be considered, and thus the position data of the aggregated 3D lane markings may be determined more precisely.
[0054] Figure 5 A schematic illustration of vehicle 1 equipped with ADS including device 10. As used herein, "vehicle" is any form of motorized conveyance. For example, vehicle 1 may be any road vehicle such as, for example, a car, a motorcycle, a (freight) truck, a bus, etc. as illustrated herein.
[0055] Device 10 includes a control circuit 11 and a memory 12. The control circuit 11 may physically include a single circuit device. Alternatively, the control circuit 11 may be distributed over several circuit devices. As an example, device 10 may share its control circuit 11 with other components of vehicle 1 (such as ADS 510). Moreover, device 10 may form part of ADS 510, that is, device 10 may be implemented as a module or feature of the ADS. The control circuit 11 may include one or more processors such as a central processing unit (CPU), a microcontroller, or a microprocessor. The one or more processors may be configured to execute program code stored in the memory 12 to perform various functions and operations of vehicle 1 in addition to the methods disclosed herein. The processor may be or include any number of hardware components for performing data or signal processing or for executing computer code stored in the memory 12. The memory 12 optionally includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 12 may include database components, object code components, script components, or any other type of information structure for supporting the various activities of this specification.
[0056] In the illustrated example, the memory 12 further stores map data 508. The map data 508 may be used, for example, by the ADS 510 of vehicle 1 to perform the automatic functions of vehicle 1. The map data 508 may include high-definition (HD) map data. It is contemplated that even though the memory 12 is illustrated as a separate element from the ADS 510, it may be provided as an integrated element of the ADS 510. In other words, according to an exemplary embodiment, any distributed or local memory device may be used for the implementation of the inventive concept. Similarly, the control circuit 11 may be distributed, for example, such that one or more processors of the control circuit 11 are provided as an integrated element of the ADS 510 or any other system of vehicle 1. In other words, according to an exemplary embodiment, any distributed or local control circuit device may be used for the implementation of the inventive concept. The ADS 510 is configured to perform the functions and operations of the automatic or semi-automatic functions of vehicle 1. The ADS 510 may include a plurality of modules, where each module is responsible for a different function of the ADS 510.
[0057] Vehicle 1 includes a plurality of elements that are typically found in an automatic or semi-automatic vehicle. It should be understood that vehicle 1 may have Figure 5 any combination of the various elements shown. Moreover, in addition to Figure 5In addition to the components shown, vehicle 1 may include further components. Although various components are shown herein as being located inside vehicle 1, one or more components may be located outside vehicle 1. For example, map data may be stored in a remote server and accessed by various components of vehicle 1 via communication system 326. Further, even though various components are depicted herein in certain arrangements, as will be readily understood by those skilled in the art, the various components may be implemented in different arrangements. It should be further noted that the various components may be communicatively connected to each other in any suitable manner. Figure 5 The vehicle 1 shown should be regarded only as an illustrative example, since the components of vehicle 1 may be implemented in several different ways.
[0058] Vehicle 1 further includes a sensor system 520. The sensor system 520 is configured to acquire sensor data regarding the vehicle itself or its surrounding environment. The sensor system 520 may include, for example, a global navigation satellite system (GNSS) module 522 (such as GPS) configured to collect geographical location data of vehicle 1. The sensor system 520 may further include one or more sensors 524. The sensors 524 may be any type of in-vehicle sensors such as cameras, lidar and radar, ultrasonic sensors, gyroscopes, accelerometers, odometers, etc. It should be recognized that the sensor system 520 may also provide the possibility of acquiring sensor data directly or via dedicated sensor control circuitry in vehicle 1.
[0059] Vehicle 1 further includes a communication system 526. The communication system 526 is configured to communicate with external units such as other vehicles (i.e., via vehicle-to-vehicle (V2V) communication protocols), remote servers (such as cloud servers), databases, or other external devices, i.e., vehicle-to-infrastructure (V2I) or vehicle-to-everything (V2X) communication protocols. The communication system 526 may communicate using one or more communication technologies. The communication system 526 may include one or more antennas (not shown). Cellular communication technologies may be used for remote communication such as to a remote server or cloud computing system. Additionally, if the cellular communication technology used has low latency, it may also be used for V2V, V2I, or V2X communication. Examples of cellular radio technologies are GSM, GPRS, EDGE, LTE, 5G, 5GNR, etc., and also include future cellular solutions. However, in some solutions, short- to medium-range communication technologies such as wireless local area network (LAN) (e.g., IEEE 802.11-based solutions) may be used for communication with other vehicles in the vicinity of vehicle 1 or with local infrastructure elements. ETSI is researching cellular standards for vehicle communication, and 5G, for example, is considered a suitable solution due to its low latency and efficient handling of high bandwidth and communication channels.
[0060] The communication system 526 can accordingly provide the possibility of sending outputs to and / or receiving inputs from a remote location (e.g., a remote operator or a control center) via one or more antennas. Moreover, the communication system 526 can be further configured to allow various elements of the vehicle 1 to communicate with each other. As an example, the communication system can provide local network setups such as CAN bus, I2C, Ethernet, fiber optic, etc. The local communication within the vehicle can also be of the wireless type with protocols such as WiFi, LoRa, Zigbee, Bluetooth, or similar medium / short-range technologies.
[0061] The vehicle 1 further includes a maneuvering system 528. The maneuvering system 528 is configured to control the maneuvering of the vehicle 1. The maneuvering system 528 includes a steering module 530 configured to control the heading of the vehicle 1. The maneuvering system 528 further includes a throttle module 532 configured to control the actuation of the throttle of the vehicle 1. The maneuvering system 528 further includes a braking module 334 configured to control the actuation of the brakes of the vehicle 1. The various modules of the maneuvering system 328 can also receive manual inputs from the driver of the vehicle 1 (i.e., from the steering wheel, the accelerator pedal, and the brake pedal respectively). However, the maneuvering system 528 can be communicatively connected to the vehicle's ADS 510 to receive instructions on how the respective modules of the maneuvering system 528 should act. Thus, the ADS 510 can control the maneuvering of the vehicle 1, for example, via the decision-making and control module 518.
[0062] The ADS 510 can include a positioning module 512 or positioning block / system. The positioning module 512 is configured to determine and / or monitor the geographical location and heading of the vehicle 1 and can utilize data from the sensor system 520, such as data from the GNSS module 522. Alternatively or in combination, the positioning module 512 can utilize data from one or more sensors 524. The positioning system can alternatively be implemented as real-time kinematic (RTK) GPS in order to improve accuracy.
[0063] The ADS 510 can further include a perception module 514 or perception block / system 514. The perception module 514 can refer to any well-known module and / or function, for example, included in one or more electronic control modules and / or nodes of the vehicle 1, that is adapted to and / or configured to interpret sensed data related to the driving of the vehicle 1 to identify, for example, obstacles, lanes, relevant signs, appropriate navigation paths, etc. The perception module 514 can thus be adapted to rely on multiple data sources and obtain inputs from multiple data sources such as automotive imaging, image processing, computer vision, and / or in-vehicle networking, in combination with, for example, sensor data from the sensor system 520.
[0064] The positioning module 512 and / or the perception module 514 may be communicatively connected to the sensor system 520 to receive perception data from the sensor system 520. The positioning module 512 and / or the perception module 514 may further transmit control instructions to the sensor system 520.
[0065] Figure 6 Schematic block diagram representation of the device 10 for the training data set 612 for generating annotations, the training data set 612 for training the perception algorithm 514 (also referred to as the perception module 514) of the ADS 510 of the vehicle 1. The device 10 may include control circuitry 11 (e.g., one or more processors) configured to perform the functions of the method 300 disclosed herein, where the functions may be included in a non-transitory computer-readable storage medium 12 or in other computer program products configured to be executed by the control circuitry 11. In other words, the system 10 includes one or more memory storage areas 12 containing program code, and the one or more memory storage areas 12 and the program code are configured to cause the system 10 to perform the method 300 according to any one of the embodiments disclosed herein by one or more processors 11. However, for better elucidation of the embodiments disclosed herein, the control circuitry is represented in Figure 6 as various “modules”, “functions” or “blocks”, each of which is linked to one or more specific functions of the control circuitry.
[0066] As illustrated, for example, a sequence 600 of frames may be obtained from the LiDAR sensor 524’ or any other sensor 524 for generating the sequence. The RRO prediction neural network 602 may be used to predict the RRO position data set 604. By using the matching function 606, the global RRO position data set 608 may be filled. Thereafter, by using the training data formation function 610, the global RRO position data set 608 and the sequence 600 may be formed into an annotated training data set 612.
[0067] Optionally, the aggregation function 614 may be used to transform the global RRO position data set 608 into an aggregated global RRO position data set 616. Consistent with the description involving Figure 3 the aggregation function may include adjusting the global RRO position data set, for example, such that it is aligned with a polyline. To further improve the aggregation function 614, the weights 620 generated by the lane marking tracking device 618 or the lane marking tracking module 618 may be provided to the aggregation function. The lane marking tracking device 618 may form part of the ADS 510 and / or may be external to the vehicle 1.
[0068] As described above, the labeled training dataset 612 can be generated by vehicle 1. After generation, it can be transmitted to a central server or cloud environment so that this dataset can be combined with other datasets from other vehicles into a combined labeled training dataset that can be used to train a perception algorithm (also referred to as a perception module). The advantage of having vehicle 1 generate the labeled training dataset 612 is that this can be done during times when idle processing capacity is available. Additionally, by having multiple vehicles (i.e., a fleet) generate training data, a wide variety of traffic conditions, weather conditions, etc. can be achieved.
[0069] Instead of having vehicle 1 generate the labeled training dataset 612, Figure 6 the device 10 illustrated in can be external to vehicle 1, while the sensor system 520 forms part of vehicle 1. As an effect, a sequence of frames 600 can be transmitted to a server or cloud environment configured to generate the labeled training dataset 612. This setup can be relevant when there is available data transmission capacity in vehicle 1 but limited processing capacity. These two methods can be used in parallel, and as explained above, whether to process the sequence of frames 600 in vehicle 1 or in the server or cloud environment can depend on the available processing capacity and available data transmission capacity in the vehicle. Another option is to have the generation of the labeled training dataset 612 occur only in vehicle 1 or only in the server or cloud environment. In other words, the server can include the device 10. Further, the cloud environment can include one or more servers.
[0070] The present invention has been presented with reference to specific embodiments. However, other embodiments are possible and within the scope of the present invention. Within the scope of the present invention, method steps for performing the method by hardware or software different from those described above can be provided. Thus, according to an exemplary embodiment, a non-transitory computer-readable storage medium storing one or more programs is provided, the one or more programs being configured to be executed by one or more processors of a computer device, the one or more programs including instructions for performing the method according to any one of the above embodiments. Alternatively, according to another exemplary embodiment, a cloud computing system can be configured to execute any method presented herein. The cloud computing system can include distributed cloud computing resources that jointly execute the methods presented herein under the control of one or more computer program products.
[0071] Generally speaking, computer-accessible media can include any tangible or non-transitory storage medium or storage media such as electrical, magnetic, or optical media (e.g., a disk or CD / DVD-ROM coupled to a computer system via a bus). As used herein, the terms "tangible" and "non-transitory" are intended to describe computer-readable storage media (or "memory") that excludes propagating electromagnetic signals, but are not intended to otherwise limit the types of physical computer-readable storage devices encompassed by the phrase computer-readable media or memory. For example, the term "non-transitory computer-readable media" or "tangible memory" is intended to encompass storage device types that do not necessarily store information permanently, including, for example, random access memory (RAM). Program instructions and data stored in a tangible computer-accessible storage medium in non-transitory form can further be transmitted via a transmission medium or a signal such as an electrical signal, an electromagnetic signal, or a digital signal, and the transmission medium or signal can be conveyed via a communication medium such as a network and / or a wireless link.
[0072] (The) processor 11 (associated with device 10) can be or include any number of hardware components for performing data or signal processing or for executing computer code stored in memory 12. Device 10 has an associated memory 12, and memory 12 can be one or more devices for storing data and / or computer code to complete or facilitate the various methods described in this specification. The memory can include volatile memory or non-volatile storage. Memory 12 can include database components, object code components, script components, or any other type of information structure for supporting the various activities of this specification. According to an exemplary embodiment, any distributed or local storage device can be used with the systems and methods of this specification. According to an exemplary embodiment, memory 12 can be communicatively (e.g., via circuitry or any other wired, wireless, or network connection) connected to the processor and includes computer code for performing one or more of the processes described herein.
[0073] Accordingly, it should be understood that parts of the described solutions can be implemented in vehicle 1, in a system located outside vehicle 1, or in a combination of inside and outside the vehicle; for example, implementing a so-called cloud solution in a server communicating with the vehicle. For example, sensor data can be sent to an external system, and the system executes the steps of method 300, method 400 according to any one of the embodiments disclosed herein. Different features and steps of the embodiments can be combined in other combinations than the described steps.
[0074] It should be noted that any reference numerals do not limit the scope of the claims, and the present invention can be implemented at least in part by both hardware and software, and several "devices" or "units" can be represented by the same hardware item.
[0075] Although the accompanying drawings may show a specific order of method steps, the order of the steps may differ from that depicted. Additionally, two or more steps may be performed simultaneously or partially simultaneously. For example, the step of receiving a signal including information about motion and information about the current road scene may be interchanged based on the specific implementation. Such variations will depend on the chosen software and hardware systems as well as the designer's choices. All such variations are within the scope of the present invention. Similarly, software implementations can be accomplished using standard programming techniques based on rule-based logic and other logics to perform various connection steps, processing steps, comparison steps, and decision steps. The embodiments mentioned and described above are given only as examples and should not be limiting to the present invention. Other solutions, uses, purposes, and functions within the scope of the present invention claimed in the patent claims described below should be obvious to those skilled in the art.
Claims
1. A computer-implemented method (300) for generating a labeled training data set (612) for training a perception algorithm (514) of an automated driving system (ADS) (510) of a vehicle (1), the method (300) comprising: obtaining (302) a sequence of frames (600) captured by one or more light detection and ranging (LiDAR) sensors (524'), predicting (304) an RRO position data set (604) for each of at least one subset of the frames (600) using a road reference object RRO prediction neural network (602), wherein each RRO position data set (604) comprises one or more RRO position data subsets for one or more RROs, respectively, wherein each RRO position data subset is associated with spatial information of one RRO found in the frame, matching (306) the one or more RRO position data subsets of one frame with the one or more RRO position data subsets of at least one other frame to populate a global RRO position data set (608), wherein the global RRO position data set (608) includes one or more global RRO position data subsets, wherein each global RRO position data subset in the global RRO position data subsets has a corresponding RRO position data subset in at least two of the frames (600), and The labeled training data set (612) is formed (308) based on the sequence and the global RRO position data set (608).
2. The method according to claim 1, wherein: The matching criterion is that the one or more subsets of RRO position data of one frame and the one or more subsets of RRO position data of at least one other frame are placed within a minimum distance threshold.
3. The method according to claim 2, wherein: The step of matching (306) further comprises: for each RRO position data subset of the one or more RRO position data subsets of the one frame, for each RRO position data subset of the one or more RRO position data subsets of the at least one other frame: determining (310) a distance between the subset of RRO position data of the one frame and the subset of RRO position data of the at least one other frame, In case the distance is below the minimum distance threshold, the subset of RRO position data of the one frame is populated (312) into the global RRO position data set.
4. The method (300) of claim 1, further comprising: Outliers are removed (314) from the global RRO position data set, wherein the outliers are subsets of global RRO position data in the global RRO position data set that are significantly different from remaining subsets of global RRO position data in the global RRO position data set.
5. The method (300) of claim 1, further comprising: The global RRO position data set is aggregated (316) by adjusting the subset of the global RRO position data to be placed inside a defined window.
6. The method (300) of claim 5, wherein: The RRO position data subset is related to lane marking positions, and wherein the step of aggregating (316) the global RRO position data set into an aggregated global RRO position data set further comprises: obtaining (318) a lane marking tracking score dataset from a lane marking tracking device, determining (320) a weight based on the lane marking tracking score dataset, and The aggregated global RRO position data set is adjusted (322) by using the weights.
7. The method according to claim 1, wherein: Each of the frames (600) is scanned from 100 to 360 degrees.
8. The method according to claim 1, wherein: The sequence of frames (600) includes detections made over a period of time.
9. The method according to claim 1, wherein: The sequence of frames (600) includes detections made by more than one LiDAR sensor (524') at a series of locations in three-dimensional space.
10. The method according to claim 1, wherein: The RRO position data subset (604) includes lane marking positions.
11. A method according to any one of the preceding claims, wherein: The RRO position data set (604) includes a position data set of a median guardrail.
12. A computer program product comprising instructions which, when executed by a computing device, cause the computing device to perform the method (300) according to any one of claims 1 to 11.
13. A device (10) for generating a labeled training data set (612), the labeled training data set (612) being used to train a perception algorithm (514) of an automated driving system (ADS) (510) of a vehicle (1), the device comprising a control circuit (11), the control circuit (11) being configured to: obtaining a sequence of frames (600) captured by at least one LiDAR sensor (524'), A RRO position dataset (604) is predicted for each of at least one subset of the frames using a road reference object RRO prediction neural network (602), wherein: Each RRO position data set comprises one or more RRO position data subsets of one or more RROs, wherein each RRO position data subset is associated with spatial information of one RRO found in the frame, matching the one or more RRO position data subsets of one frame with one or more RRO position data subsets of at least one other frame to populate a global RRO position data set, wherein the global RRO position data set comprises one or more global RRO position data subsets, wherein each global RRO position data subset of the global RRO position data subsets has a corresponding RRO position data subset in at least two of the frames, and The labeled training dataset (612) is formed based on the sequence and the global RRO position dataset (608).
14. The device (10) according to claim 13, wherein: The RRO position data subset is related to lane marking positions, wherein the control circuit (11) is further configured to: obtaining a lane marking tracking score dataset from a lane marking tracking device (618), determining a weight based on the lane marking tracking score dataset (620), and The aggregated global RRO position data set (616) is adjusted by using the weights (620).
15. A vehicle (1) comprising: An automated driving system ADS (510), including a perception algorithm (514), at least one light detection and ranging LiDAR sensor (524'), and An apparatus (10) for generating a labeled training data set (616) for training the perception algorithm (514) of the ADS (510) according to claim 13 or 14.