System and method for enhancing remote object labeling based on roadside sensor
By using a distributed data acquisition system in the autonomous driving system, combined with frame alignment and frame overlay technology of on-board and roadside lidars, the problem of small and distant objects labeling in the autonomous driving system is solved, and the point cloud density is improved and the labeling accuracy is improved.
Patent Information
- Application Number
- CN202510124937.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-07
- Filing Date
- 2025-01-26
- Publication Date
- 2025-06-13
AI Technical Summary
In the training of autonomous driving systems, smaller or farther objects are difficult to effectively mark due to occlusion or insufficient point cloud sampling density.
A distributed data acquisition system, including vehicle-mounted and roadside lidar devices, is adopted to enhance the density of object point clouds through frame alignment and frame overlay technology, thereby improving labeling accuracy.
By enhancing point cloud density, the labeling accuracy of small and distant objects is improved, providing more training data to improve the performance of autonomous driving systems.
Smart Images

Figure CN120147147A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to collecting training data for autonomous driving. More specifically, the present disclosure relates to enhancing the annotation of distant object data using measurement data from roadside sensors. Background Art
[0002] Training of autonomous driving systems typically requires collecting a large amount of real-world traffic data. In a typical scenario, a data collection vehicle may be equipped with various sensors (such as cameras, light detection and ranging (LiDAR) devices, radars, global positioning system (GPS) modules, etc.), and can continuously collect data in an actual driving scenario. More specifically, when the data collection vehicle is driving under different traffic and weather conditions, these on-vehicle sensors can obtain information about the surrounding environment.
[0003] The collected information (such as images, 3D point clouds, meshes, etc.) can be automatically annotated by algorithms or manually annotated, and then used as training data for training autonomous driving systems. For example, an algorithm or a human annotator can identify and label objects (such as vehicles, traffic lights / signs, pedestrians, etc.) in an image or a LiDAR frame. Then, the labeled image or LiDAR frame can be used as training data to perform supervised training on machine learning models for object detection and object tracking. However, smaller or more distant objects (such as vehicles or pedestrians) in a traffic scene may be severely occluded by closer objects, or have too few sampled points in the 3D point cloud, making it difficult to annotate these smaller or more distant objects. Summary of the Invention
[0004] Embodiments of the present disclosure may provide a data collection system for collecting training data for autonomous driving applications. The data collection system may include an on-vehicle light detection and ranging (LiDAR) device, a roadside LiDAR device, a frame alignment subsystem, and a frame superposition subsystem installed on a data collection vehicle. The on-vehicle LiDAR device is used to collect a first continuous frame stream while the data collection vehicle is driving. The roadside LiDAR device collects a second continuous frame stream while stationary. The frame alignment subsystem is used to align the second continuous frame stream collected by the roadside LiDAR with the first continuous frame stream collected by the on-vehicle LiDAR in both the time domain and the spatial domain. The frame superposition subsystem is used to superpose the frames in the second continuous frame stream onto the corresponding frames in the first continuous frame stream, thereby facilitating the enhancement of the object point cloud in the first continuous frame stream.
[0005] In a variation of this embodiment, the data collection system may further include a synchronization subsystem for synchronizing the operations of the on-vehicle LiDAR device and the roadside LiDAR device.
[0006] In another variation, the synchronization subsystem may include a high-precision global positioning system (GPS) module.
[0007] In a variation of this embodiment, the data acquisition system may further include an attitude determination subsystem for determining the attitude of the vehicle-mounted lidar device associated with each frame of the first continuous frame stream and the attitude of the roadside lidar device.
[0008] In another variation, the data acquisition system may further include a frame preprocessing subsystem for preprocessing the frames by removing transient objects from the frames.
[0009] In another variation, the attitude determination subsystem may determine the attitude of the vehicle-mounted lidar device associated with the frame by aligning the preprocessed frame with a reference frame in the spatial domain; and the attitude determination subsystem may determine the attitude of the roadside lidar device by aligning a selected frame in the second continuous frame stream with the reference frame in the spatial domain.
[0010] In another variation, the frame alignment subsystem selects frames from the first continuous frame stream, identifies frames in the second continuous frame stream that are temporally aligned with the selected frames, and aligns the identified frames with the selected frames in the spatial domain based on the attitudes of the vehicle-mounted lidar device and the roadside lidar device associated with the selected frames.
[0011] In another variation, aligning the preprocessed frame with the reference frame in the spatial domain may include applying the Iterative Closest Point (ICP) algorithm or obtaining high-precision GPS information associated with the vehicle-mounted lidar device and the roadside lidar device.
[0012] In another variation, the data acquisition system may further include a map building subsystem for building a high-precision map by superimposing the preprocessed frames in the first continuous frame stream collected by the vehicle-mounted lidar.
[0013] In a variation of this embodiment, the roadside lidar acquires the second continuous frame stream from a higher perspective.
[0014] One embodiment provides a method for collecting training data for autonomous driving applications. The method may include: obtaining a first continuous frame stream from a vehicle-mounted light detection and ranging (lidar) device installed on a data acquisition vehicle; obtaining a second continuous frame stream from a roadside lidar device; aligning the second continuous frame stream collected by the roadside lidar with the first continuous frame stream collected by the vehicle-mounted lidar in the time domain and the spatial domain; and superimposing the frames in the second continuous frame stream on the corresponding frames in the first continuous frame stream, thereby facilitating enhancing the point cloud of the objects in the first continuous frame stream. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1Aillustrates an exemplary data acquisition scenario according to the prior art;
[0016] Figure 1B illustrates an example of an annotated image according to the prior art;
[0017] Figure 1C illustrates an example of an annotated LiDAR frame according to the prior art;
[0018] Figure 2 illustrates an example scenario of using roadside sensors to enhance the 3D point cloud of distant objects according to an embodiment of the present application;
[0019] Figure 3 illustrates an exemplary block diagram of a data acquisition system according to an embodiment of the present application;
[0020] Figure 4 presents a flowchart illustrating an exemplary data acquisition and point cloud enhancement process according to an embodiment of the present application;
[0021] Figure 5 illustrates an exemplary computer system according to an embodiment of the present application.
[0022] In the figures, the same reference numerals denote the same components. Detailed Description
[0023] The following description is intended to enable those skilled in the art to make and use the disclosed embodiments and is provided in the context of one or more specific applications and their requirements. Various modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the scope of the disclosed content. Thus, the present invention or various aspects thereof are not intended to be limited to the embodiments shown, but rather should be accorded the widest scope consistent with the disclosed content.
[0024] Overview
[0025] Embodiments of the present disclosure provide a system and method for facilitating enhanced annotation of training samples for autonomous driving applications. More specifically, to facilitate accurate annotation of busy traffic scenes that may contain occluded objects or distant objects, a distributed data acquisition system may include a mobile data acquisition subsystem and one or more fixed data acquisition subsystems. The mobile data acquisition subsystem may include multiple sensors mounted on a vehicle and may be configured to acquire data of traffic scenes during vehicle travel. The fixed data acquisition subsystem may at least include a lidar device located horizontally above the vehicle for acquiring 3D point clouds of objects in the traffic scene. The distributed data acquisition system may also include an alignment subsystem configured to align the data acquired by the mobile subsystem and the fixed subsystem to overlay the acquired data, thereby increasing the point cloud density of small or distant objects and facilitating more accurate annotation.
[0026] Roadside Sensor
[0027] For an autonomous driving system, the ability to detect distant objects (such as vehicles, obstacles, traffic signs, etc.) is crucial because timely detection of distant traffic hazards can provide more reaction time for the autonomous driving system. Training an autonomous driving system to recognize distant objects requires a large amount of labeled training data. However, annotating or labeling distant objects in training samples is challenging.
[0028] Figure 1A An exemplary data acquisition scenario according to the prior art is shown. In Figure 1A , a specially designed data acquisition vehicle 100 may be equipped with multiple sensors, including but not limited to cameras, radars, lidar devices (or simply lidar), global positioning system (GPS) modules, etc. To acquire traffic data, the vehicle 100 may be deployed under various road and weather conditions. When the vehicle 100 is traveling, various sensors may acquire data about the surrounding environment (such as images, 3D point clouds, etc.), such as pedestrians 102, traffic lights 104, vehicles 106, etc.
[0029] The acquired data can be annotated automatically (e.g., through various object detection algorithms) or manually. Figure 1B An example of an annotated image according to the prior art is shown. As Figure 1B shown, distant objects (such as vehicles) may be small and occluded by objects closer to the camera. Accurately detecting distant objects is challenging for both algorithms and humans. Figure 1C An example of an annotated lidar frame according to the prior art is shown. As Figure 1CAs shown, as the distance between the object and the laser increases, the point cloud density of each detected object decreases. The sparsity of the point cloud of distant objects increases, making it difficult to distinguish the shapes of individual objects even for human annotators.
[0030] An existing method for improving the annotation accuracy of distant objects in the collected data (such as images or 3D point clouds) may include using two data collection vehicles, one following the other. The lidar device (referred to as lidar for short) on the front vehicle can collect the light reflected from distant objects, thereby enhancing the 3D point clouds of these distant objects. However, this method requires carefully maintaining the distance between the two moving vehicles, which can be a cumbersome process, and deploying additional data collection vehicles may be costly and laborious. In addition, the sensors on the front vehicle (such as cameras and sensors on lidar) may encounter the same occlusion problems because their perspectives are at the vehicle's horizontal position.
[0031] To enhance the point cloud of distant objects, in some embodiments of the present application, in addition to the lidar installed on the mobile data collection vehicle, the data collection system may further include one or more additional fixed lidars, which are placed at a high position in the traffic scene to collect the light reflected from distant objects. More specifically, the fixed lidar can be closer to the target object than the lidar on the data collection vehicle, and thus can provide more information about the target object (for example, calculate more points). In addition, since the fixed lidar can be placed at a high position, they can provide an overhead view of the target objects on the road, thereby alleviating the occlusion problems faced by the sensors on the data collection vehicle.
[0032] In some embodiments, the fixed lidar can be installed on permanent roadside traffic facilities (for example, installed on the posts supporting traffic lights or signs). For example, an intersection may include multiple traffic lights, and each signal post is installed with a fixed lidar. In a variant, the fixed lidar can be portable. For example, the fixed lidar can be carried by a vehicle parked on an overpass to scan the traffic under the overpass. In another example, the fixed lidar can be installed on the top of a post carried by a vehicle parked by the roadside to scan the passing vehicles.
[0033] Figure 2 An example scenario of using roadside sensors to enhance the 3D point cloud of distant objects according to an embodiment of the present application is shown. In Figure 2In this case, multiple lamp / sensor support structures (such as columns) can be located around the intersection 200. In this example, each corner of the intersection 200 can include a support structure (for example, support structure 202 is located at the lower right corner of the intersection 200, and support structure 204 is located at the upper right corner). A typical support structure can include a column and a cantilever, on which traffic lights or signs are installed. In many modern cities, other types of sensors (such as cameras, radars, etc.) may also be installed or attached to the cantilever to collect important traffic information.
[0034] In some embodiments of the present application, at least one lidar can be installed on the cantilever of each support structure. For example, lidar 206 can be installed on support structure 202, and lidar 208 can be installed on support structure 204. In a variant, not all support structures are equipped with lidars. One or two lidars may be able to perform omnidirectional scanning.
[0035] Figure 2 A data collection vehicle 210 traveling on the road is also shown. As mentioned before, various sensors can be installed on the data collection vehicle 210 to collect data (such as images, point clouds, etc.) about the traffic scene around the data collection vehicle 210. For example, a camera installed on the data collection vehicle 210 can capture images of vehicles 212 and 214 traveling in front of the data collection vehicle 210. Similarly, a lidar installed on the data collection vehicle 210 can obtain 3D point clouds of vehicles 212 and 214. As Figure 2 can be seen, vehicle 214 is in front of vehicle 212 and smaller than vehicle 212, which means that vehicle 214 may be partially blocked by vehicle 212 in the captured image. In addition, both of these vehicles are far from the data collection vehicle 210, which means that the 3D point clouds of these two vehicles collected by the lidar may be sparse. Both of these situations may cause difficulties in annotating the data related to vehicle 214 (even for manual annotation).
[0036] To improve the accuracy of annotation, in some embodiments, at least one roadside lidar (such as lidar 206 or lidar 208) installed on the support structure can also be used to scan the traffic scene to obtain the point cloud of target objects in the scene, including vehicle 212 and vehicle 214. It is worth noting that since the roadside lidar 206 and the roadside lidar 208 are located above the traffic scene and have a relatively high viewing angle, they are less likely to encounter occlusion problems. For example, when viewed from above, vehicle 212 does not occlude vehicle 214. In addition, compared with the data collection vehicle 210, the roadside lidar (such as lidar 206 or lidar 208) is closer to the target objects (such as vehicle 212 and vehicle 214), which means that the point cloud density obtained by the roadside lidar may be higher than the point cloud density obtained by the lidar on the data collection vehicle 210. In some embodiments, the point cloud obtained by the roadside lidar can be combined with the point cloud obtained by the lidar on the data collection vehicle 210 to increase the point cloud density of objects (such as Figure 2 vehicles 212 and 214 in
[0037] In Figure 2 the example shown, the roadside lidar is installed on the traffic signal near the intersection. In practical applications, these lidars can be installed on any existing roadside traffic structures, such as street lights, road signs, etc. According to the actual application scenario, the roadside lidar can also be installed on buildings, trees, viaducts, etc. In the intelligent street or intelligent highway scenario, dedicated roadside sensor installation structures equipped with lidar can be provided. In addition to being installed on fixed traffic structures, the roadside lidar can also be portable, for example, carried by a vehicle parked at a selected location. Importantly, the positions of these roadside lidars are higher than the traffic scene to alleviate the occlusion problem.
[0038] In Figure 2 the example shown, the data collection vehicle 210 is driving on the road and approaching the intersection 200. In practical applications, the data collection vehicle 210 can also stop by the roadside to collect data on the traffic (such as vehicles, pedestrians, etc.) passing through the intersection 200. For example, the data collection vehicle 210 can stop at a location far from the intersection 200 (such as between 150 and 250 meters away from the intersection). In this case, the roadside lidar (such as lidar 206 and lidar 208) can scan the vehicles / pedestrians near the intersection 200 from above to obtain the high-precision point cloud of these objects. Thus, the data collection vehicle 210 can collect training samples including data (such as point cloud) of distant objects, and the roadside lidar can provide enhancement for these data.
[0039] In some embodiments, in addition to lidars (e.g., lidar 206 and lidar 208), one or more cameras can be installed on the roadside support structures (e.g., structures 202 and 204) to provide additional visual information (e.g., images) of the traffic scene. These visual information can also be combined with the visual information collected by the data collection vehicle 210 to improve the annotation accuracy of the collected traffic data. It should be noted that compared with the roadside lidars that can provide high-precision 3D point clouds of distant objects, roadside cameras usually can only qualitatively, rather than quantitatively, enhance the collected data (e.g., images or 3D point clouds). In addition, using cameras may raise privacy issues and sometimes may not be available.
[0040] Data Merging
[0041] The additional data collected by roadside sensors (such as lidars and / or cameras) can be merged with the data collected by the sensors on the data collection vehicle to obtain enhanced data (e.g., point clouds). Before merging the data from different lidars, different data sets should be aligned in the time domain and the space domain. As long as all sensors are synchronized, aligning the data of different sensors in the time domain is relatively easy. In some embodiments, all lidars may rely on the same clock source (e.g., the clock provided by GPS) to trigger laser pulses. The time offset between different lidars can be less than 0.05 seconds.
[0042] Aligning the data of different sensors in the space domain may be more complex. It should be noted that the instantaneous relative pose between the lidar and the target object can determine the instantaneous pose of the 3D point cloud of the target object. More specifically, the 3D point cloud of the target object collected by the lidar at a certain moment is usually represented in the local coordinate system of the lidar at that moment. As the lidar moves or its pose changes (e.g., as the vehicle carrying the lidar moves), the local coordinate system also changes. Similarly, different lidars with different poses may have different local coordinate systems. Before merging the data represented in different coordinate systems, they should be transformed (or converted) into the same coordinate system.
[0043] In some embodiments, a reference coordinate system can be selected to which all point clouds (including point clouds from the same lidar at different times and point clouds from different lidars) can be aligned. Various point cloud alignment algorithms can be used to align point clouds from different coordinate systems. In one embodiment, the Iterative Closest Point (ICP) algorithm can be used to align the point clouds. It is worth noting that the reference coordinate system can be arbitrarily selected. In one embodiment, the reference coordinate system can be determined according to the first frame obtained by an on-vehicle lidar (e.g., a lidar mounted on a data acquisition vehicle). Such a frame can also be referred to as a reference frame. For example, the origin of the reference coordinate system can be a predetermined point of the on-vehicle lidar when acquiring the first frame. The absolute position of this predetermined point (i.e., its position in the global coordinate system) can be determined by an accurate GPS module (e.g., a real-time kinematic (RTK) GPS module) mounted on the data acquisition vehicle. In other words, all subsequent frames obtained by the on-vehicle lidar and frames obtained by the roadside lidar can be aligned according to the reference frame.
[0044] Once the reference coordinate system is determined, other frames (including frames obtained by the same on-vehicle lidar at different times and frames obtained by the roadside lidar) can be aligned to this reference coordinate system. More specifically, a transformation matrix can be derived for each frame to transform the point cloud in each frame from the local coordinate system to the reference coordinate system.
[0045] When the lidar scans a busy traffic scene (e.g., a traffic scene at an intersection), each acquired frame may contain point clouds of many different objects, making it difficult to align the frame to the reference coordinate system. To simplify the alignment process, in some embodiments, the system can remove transient objects (e.g., vehicles and pedestrians) from a specific frame before aligning it to the reference coordinate system. The resulting frame will be simpler and may contain objects that are easy to detect or identify (e.g., landmarks, lane lines, curbs, etc.). The simplified frame can be aligned to the reference coordinate system (e.g., by applying the ICP algorithm) to obtain the pose of the lidar in the reference coordinate system when capturing this frame. Such a pose can be referred to as an instant local pose. Since the on-vehicle lidar moves with the vehicle, each captured frame corresponds to a different instant local pose. On the other hand, the roadside lidar is stationary, which means its pose in the reference coordinate system remains unchanged and can simply be referred to as a local pose. In some embodiments, the local pose of the roadside lidar can be determined by any frame (e.g., the first frame).
[0046] In one embodiment, after aligning multiple subsequent simplified frames captured by an on-vehicle lidar, the system can stack these aligned simplified frames to create a high-precision local map. This high-precision local map can display the positions of non-transient objects, such as lane lines, curbs, crosswalks, etc. In addition to creating the map, the system can also determine a corresponding transformation matrix for each simplified frame. The transformation matrix can be used to transform the point cloud of an object (such as a vehicle, pedestrian, etc.) from a vehicle-based coordinate system to a reference coordinate system to align the point cloud of a specific object with the corresponding point cloud in the reference frame.
[0047] The data collection vehicle can typically be equipped with multiple lidars, each lidar scanning the traffic scene from a different perspective. In some embodiments, the data (such as point cloud) collected by each on-vehicle lidar can be enhanced by the data (such as point cloud) collected by roadside sensors. For example, a transformation matrix can be derived for each on-vehicle lidar, and the frames obtained by the roadside lidar can be aligned to different reference coordinate systems corresponding to different on-vehicle lidars.
[0048] In some embodiments, enhancing the point cloud of a specific object obtained by an on-vehicle lidar in any frame (such as any moment) can include multiple data matching and merging operations. For example, the system can first identify the corresponding frame obtained by the roadside lidar (such as based on the timestamp), and then use object detection techniques to detect the corresponding point cloud in the roadside frame. The system can also align the point cloud obtained by the roadside lidar with the point cloud obtained by the on-vehicle lidar. It should be noted that the instantaneous local pose of the on-vehicle lidar and the local pose of the roadside lidar have been determined in advance, and the point cloud can be aligned according to the transformation matrix corresponding to these two poses. In one example, the transformation matrix between the two poses can be calculated, and the point cloud can be transformed from one pose to another according to the calculated transformation matrix.
[0049] After alignment, the point cloud obtained by the roadside lidar can be stacked on the point cloud obtained by the on-vehicle lidar to create an enhanced point cloud. It should be noted that although the point cloud of more than one object in a given frame can be enhanced, the enhancement of the point cloud of small or distant objects may be more affected due to the sparsity of these point clouds. Then, the frame with the enhanced point cloud can be labeled and used for training purposes. Compared with the frame without point cloud enhancement, labeling the frame with the enhanced point cloud can be more accurate and provide more data. In extreme cases, distant objects that were originally unrecognizable may become recognizable after point cloud enhancement.
[0050] Figure 3An exemplary block diagram of a data acquisition system according to an embodiment of the present application is shown. The data acquisition system 300 may include a plurality of on-vehicle sensors 302, a plurality of roadside sensors 304, a synchronization subsystem 306, a frame preprocessing subsystem 308, a sensor attitude determination subsystem 310, a frame alignment subsystem 312, a map construction subsystem 314, and a point cloud overlay subsystem 316.
[0051] The on-vehicle sensors 302 may include various sensors mounted on the data acquisition vehicle. For example, the on-vehicle sensors 302 may include one or more cameras, one or more lidars, a GPS module, an inertial measurement unit (IMU), etc. The on-vehicle sensors 302 may collect data when the vehicle is moving or parked. In most cases, the viewing angles of the on-vehicle cameras and lidars do not elevate (e.g., at the same horizontal height as the vehicle).
[0052] The roadside sensors 304 may include various sensors mounted on roadside traffic facilities (such as streetlight poles, traffic signs, viaducts, etc.), buildings, or trees. The roadside sensors 304 are fixed. For example, the roadside sensors 304 may include cameras, lidars, and / or GPS modules. The roadside cameras and lidars typically have a higher viewing angle. In one embodiment, each lidar may operate at a speed of approximately 10 hertz, or capture approximately 10 frames per second.
[0053] The synchronization subsystem 306 may be responsible for synchronizing the data captured by different sensors to facilitate data merging. For example, before merging the frames obtained from the roadside lidar with the frames obtained from the on-vehicle lidar, the synchronization subsystem 306 may be used to determine that these two frames are obtained at the same moment, which means they represent the same traffic scene. In some embodiments, the synchronization subsystem 306 may interface with the GPS modules within the on-vehicle sensors 302 and the roadside sensors 304 to obtain the timing information of each frame. Then, the obtained timing information may be used to align the frames from different sensors in the time domain. In some embodiments, the on-vehicle sensors 302 and the roadside sensors 304 may be triggered by the same high-precision GPS signal. For example, the on-vehicle lidar and the roadside lidar may work synchronously such that each time the on-vehicle lidar captures a frame, the roadside lidar also captures a frame simultaneously.
[0054] The frame preprocessing subsystem 308 may be responsible for preprocessing the frames obtained from different lidars (including on-vehicle lidars and roadside lidars). The frame preprocessing subsystem 308 may remove transient objects (such as vehicles, pedestrians, etc.) from each frame. The preprocessed frames may contain easily recognizable objects, such as landmarks, curbs, lane lines, etc.
[0055] The sensor attitude determination subsystem 310 may be responsible for determining the attitude of the camera or lidar corresponding to each frame (relative to the reference coordinate system). In some embodiments, the sensor attitude determination subsystem 310 may obtain real-time GPS information to determine the attitude of the sensor. In a variant, the sensor attitude determination subsystem 310 may determine the current attitude of the lidar by aligning the corresponding preprocessed frame with a reference frame (e.g., the first frame captured by the vehicle-mounted lidar). When the lidar is in motion (e.g., moving with the data acquisition vehicle), the attitude of the lidar in each frame may also be determined iteratively. For example, the attitude of the lidar in a specific frame may be determined by aligning the specific frame with the frame at the previous moment. Since the roadside sensors are stationary, only one iteration is required to determine the attitude of each roadside sensor. For example, the sensor attitude determination subsystem 310 may align any preprocessed frame of the roadside lidar with the reference frame to determine the attitude of the roadside lidar.
[0056] The frame alignment subsystem 312 may be used to align the frames captured by different attitude sensors in the spatial domain. This alignment can make the point cloud of an object in one frame overlap with the point cloud of the same object in another frame. In some embodiments, the frame alignment subsystem 312 may use the ICP algorithm to align the frames. The frame alignment subsystem 312 may also calculate a transformation matrix, which can be used to transform the frame from one sensor attitude to another. In addition to aligning the frames in the spatial domain, in some embodiments, the frame alignment subsystem 312 may also align the corresponding frames captured by different sensors in the time domain based on the output of the synchronization subsystem 306. For example, the frames captured by different lidars may be aligned in the time domain according to the timestamps associated with each frame.
[0057] The map construction subsystem 314 may utilize the aligned preprocessed frames from the vehicle-mounted lidar to construct a high-precision map. Since these preprocessed frames do not contain transient (or time-varying) objects, after they are aligned to the same coordinate system, these preprocessed frames represent the same scene (e.g., a specific section of an intersection, street, or highway, etc.). Therefore, superimposing these aligned preprocessed frames can provide detailed information about the scene (e.g., the positions of lane lines and curbs). The reconstructed map can be used as training data for training the autonomous driving system.
[0058] The point cloud overlay subsystem 316 can be used to overlay the object point cloud (also referred to as the roadside point cloud) obtained by the roadside lidar onto the corresponding point cloud (also referred to as the vehicle-mounted point cloud) obtained by the vehicle-mounted lidar. Before overlaying the point clouds, the frames or point clouds should be aligned in the time domain and the spatial domain. In one embodiment, the point cloud overlay subsystem 316 can identify frames that are aligned in the time domain (e.g., frames captured at the same moment) and determine the instantaneous poses of the lidars that captured these frames. The point cloud overlay subsystem 316 can further calculate the transformation matrix between the two poses. This transformation matrix can be used to transform the point cloud from one coordinate system to another. For example, the point cloud overlay subsystem 316 can apply the transformation matrix to the roadside point cloud to align it with the corresponding vehicle-mounted point cloud. The aligned roadside point cloud can then be overlaid onto the corresponding vehicle-mounted point cloud to provide enhancement to the vehicle-mounted point cloud.
[0059] Figure 4 A flowchart is shown that illustrates an exemplary data acquisition and point cloud enhancement process according to an embodiment of the present application. In one or more embodiments, Figure 4 one or more of the steps in may be executed repeatedly, and / or may be executed in a different order. Thus, Figure 4 the specific arrangement of the steps shown in should not be construed as limiting the scope of the technology.
[0060] During operation, a distributed data acquisition system including multiple lidars can acquire multiple frames related to a dynamic traffic scene (operation 402). In some embodiments, the distributed data acquisition system can include at least one lidar mounted on a data acquisition vehicle and one lidar mounted on a roadside structure, with each lidar acquiring a continuous stream of frames. In one example, the vehicle-mounted lidar can move with the data acquisition vehicle, and when the data acquisition vehicle approaches the roadside lidar from a distance, it can acquire a continuous stream of frames. In this example, objects closer to the roadside lidar will appear farther away in the frames initially acquired by the vehicle-mounted lidar. The frames acquired by each lidar can be processed locally (e.g., by a computing device inside the data acquisition vehicle) or sent to a centralized location for processing. In some embodiments, the centralized location can be a processing unit inside the data acquisition vehicle. More specifically, the roadside lidar can transmit (e.g., via a wireless network) the frames it acquires to the processing unit. In a variant, the centralized location can be a remote server, and both the vehicle-mounted lidar and the roadside lidar can transmit their frames to the remote server.
[0061] Each collected lidar frame can be pre - processed (operation 404). In some embodiments, pre - processing the lidar frame can include removing transient objects (i.e., objects that change position over time) from the frame. For example, non - map objects such as vehicles and pedestrians can be removed from the frame, leaving only the objects that are part of the map in the frame (such as buildings, traffic signs, curbs, etc.).
[0062] The pre - processed frames (including frames collected by on - vehicle lidar and frames collected by roadside lidar) can then be aligned with a reference frame (operation 406). The reference frame can be arbitrarily selected from the frames captured by the on - vehicle lidar. Once the reference frame is selected, the reference coordinate system can be determined accordingly. In one example, the reference frame can be the first frame in the frame stream obtained by the on - vehicle lidar, and the reference coordinate system can be anchored on the on - vehicle lidar. In some embodiments, frame alignment can be performed based on the high - precision GPS data associated with each frame. For example, based on the timestamp of each frame, the system can identify the instantaneous GPS readings of the data collection vehicle and then determine the pose of the on - vehicle lidar. The relative position between the on - vehicle lidar and the data collection vehicle generally remains unchanged. In alternative embodiments, the ICP (Iterative Closest Point) algorithm can be used to perform frame alignment. In further embodiments, the ICP algorithm can be used to align each frame with the frame at the previous moment. Since the pose change of the on - vehicle lidar between consecutive frames is very small, applying the ICP algorithm to consecutive frames can achieve high - precision alignment. Since the roadside lidar is stationary, the frames captured by the roadside lidar can be aligned with the reference frame by selecting any roadside frame and aligning the selected roadside frame with the reference frame. Once the selected frames are aligned, other roadside frames can be aligned in a similar manner.
[0063] The pose of the lidar associated with each frame can be determined according to the frame alignment result (operation 408). As mentioned above, as the data collection vehicle moves, the pose of the on - vehicle lidar changes continuously, while the pose of the roadside lidar remains unchanged. Therefore, each frame captured by the on - vehicle lidar corresponds to a different lidar pose, and all frames captured by the roadside lidar correspond to the same lidar pose.
[0064] The system can stack the aligned and pre - processed frames captured by the on - vehicle lidar to obtain a high - precision local map (operation 410). Such a high - precision local map can be part of the training data.
[0065] The system can select an on-vehicle frame (i.e., a frame captured by an on-vehicle lidar) and identify a roadside frame that is temporally aligned therewith (e.g., a frame captured by a roadside lidar at the same moment) (operation 412). Then, the system can align and superimpose the two frames (i.e., the selected on-vehicle frame and the identified temporally aligned roadside frame) according to the lidar poses corresponding to the two frames (operation 414). It should be noted that both of these two frames contain transient objects (such as vehicles, pedestrians, etc.). Since they are captured at the same moment, the poses of these objects in the same coordinate system (e.g., a local reference coordinate system (LRC)) should be the same. Aligning the frames also means that the 3D point cloud of a specific object in one frame can be aligned with the 3D point cloud of the same object in the other frame. Superimposing the roadside frame on the on-vehicle frame can enhance the 3D point cloud of the object, and this enhancement can have a greater impact on small objects or objects far from the on-vehicle lidar.
[0066] Next, the system can determine whether the current frame is the last frame in the continuous frame stream captured by the on-vehicle lidar (operation 416). If not, the next frame can be selected (operation 412). Otherwise, the process ends. It should be noted that, in Figure 4 the example shown, the superimposition of frames from the two lidars can be performed in a frame-by-frame order. In actual operation, such operations can also be performed in parallel. After the point cloud enhancement, the continuous frame stream captured by the on-vehicle lidar can be sent for annotation. Through the enhancement, the point clouds of small or distant objects are less sparse, making it easier for human annotators or algorithms to annotate these objects.
[0067] Figure 5 An exemplary computer system according to an embodiment of the present application is shown. The computer system 500 includes a processor 502, a memory 504, and a storage device 506. In addition, the computer system 500 can be connected to a peripheral input / output (I / O) user device 510, such as a display device 512, a keyboard 514, a pointing device 516, and a camera / lidar 518. The storage device 506 can store an operating system 520, a data acquisition and alignment system 522, and data 540. In some embodiments, the computer system 500 can be implemented as part of an autonomous driving sample acquisition system.
[0068] The data acquisition and alignment system 522 may include instructions that, when executed by the computer system 500, may cause the computer system 500 or the processor 502 to perform the methods and / or processes described in this disclosure. Specifically, the data acquisition and alignment system 522 may include the following instructions: instructions for receiving frames acquired by multiple lidars (frame reception instructions 524), instructions for synchronizing the lidars (synchronization instructions 526), instructions for preprocessing lidar frames (frame preprocessing instructions 528), instructions for aligning the preprocessed lidar frames (frame alignment instructions 530), instructions for determining the lidar poses associated with each frame (pose determination instructions 532), instructions for constructing a high-precision local map based on the aligned preprocessed frames (map construction instructions 534), and instructions for overlaying the point clouds of the temporally and spatially aligned frames (point cloud overlay instructions 536).
[0069] This disclosure presents a solution to address the annotation problem in lidar frames caused by sparse point clouds of small and distant objects. More specifically, a distributed data acquisition system including on-vehicle sensors and roadside sensors can be used to simultaneously acquire data related to traffic scenes. The distributed data acquisition system may include one or more on-vehicle lidars and at least one roadside lidar with an overhead view of the traffic scene. The lidar pose associated with each frame can be determined based on GPS data or by using the ICP algorithm to align non-transient objects in the lidar frames. The frames of the on-vehicle lidars and the roadside lidar can be aligned in time and space, and the corresponding temporally and spatially aligned frames can be overlaid to enhance the point clouds of objects (especially the point clouds of small and distant objects) in each on-vehicle frame. The frames with enhanced point clouds can be easily annotated and used as training samples for autonomous driving.
[0070] The data structures and program codes described in this detailed description are generally stored on a non-transient computer-readable storage medium, which can be any device or medium capable of storing code and / or data for use by a computer system. Non-transient computer-readable storage media include, but are not limited to, volatile memories; non-volatile memories; electrical, magnetic, and optical storage devices, solid-state drives, and / or other non-transient computer-readable media known now or developed later.
[0071] The methods and processes described in the detailed description can be embodied as code and / or data, which can be stored in the above-mentioned non-transient computer-readable storage medium. When a processor or a computer system reads and executes the code stored on this medium and operates on the data stored on this medium, the processor or the computer system will execute the methods and processes embodied in the form of code and data structures and stored on this medium.
[0072] In addition, the optimized parameters obtained from the above methods and processes can be programmed into hardware modules, including but not limited to Application-Specific Integrated Circuit (ASIC) chips, Field-Programmable Gate Arrays (FPGAs), and other programmable logic devices now known or later developed. When such a hardware module is activated, it executes the methods and processes contained within the module.
[0073] The above embodiments are for illustration and description only and are not intended to be exhaustive or to limit the scope of this disclosure to the forms disclosed. Thus, many modifications and variations will be apparent to those skilled in the art. The scope of this disclosure is defined by the appended claims rather than the preceding disclosure.
Claims
1. A data acquisition system for acquiring training data for an autonomous driving application, the data acquisition system comprising: A vehicle-mounted laser radar device installed on a data collection vehicle; Roadside LiDAR device; Frame alignment subsystem; as well as Frame overlay subsystem; Wherein, the vehicle-mounted laser radar device is used to collect a first continuous frame stream when the data collection vehicle moves; Wherein, the roadside laser radar device is used to collect a second continuous frame stream when remaining stationary; The frame alignment subsystem is used to align the second continuous frame stream collected by the roadside laser radar with the first continuous frame stream collected by the vehicle-mounted laser radar in the time domain and the space domain; and The frame superposition subsystem is used to superimpose frames in the second continuous frame stream onto corresponding frames in the first continuous frame stream, thereby promoting enhancement of the object point cloud in the first continuous frame stream.
2. The data acquisition system according to claim 1, characterized in that: It also includes a synchronization subsystem for synchronizing the operation of the vehicle-mounted laser radar device and the roadside laser radar device.
3. The data acquisition system according to claim 2, characterized in that: The synchronization subsystem includes a high-precision Global Positioning System (GPS) module.
4. The data acquisition system according to claim 1, characterized in that: It also includes a posture determination subsystem for determining the posture of the vehicle-mounted laser radar device and the posture of the roadside laser radar device associated with each frame in the first continuous frame stream.
5. The data acquisition system according to claim 4, characterized in that: A frame pre-processing subsystem is also included for pre-processing the frames by removing transient objects from the frames.
6. The data acquisition system according to claim 5, characterized in that: The pose determination subsystem determines the pose of the onboard laser radar device associated with a frame by aligning the pre-processed frame with a reference frame in the spatial domain; as well as The attitude determination subsystem determines the attitude of the roadside laser radar device by aligning a frame selected from the second continuous frame stream with the reference frame in the spatial domain.
7. The data acquisition system according to claim 6, characterized in that: The frame alignment subsystem is used to: selecting a frame from said first continuous stream of frames; identifying a frame from the second continuous frame stream that is aligned in time domain with the selected frame; as well as Based on the pose of the onboard laser radar device and the pose of the roadside laser radar device associated with the selected frame, the identified frame is aligned with the selected frame in the spatial domain.
8. The data acquisition system according to claim 6, characterized in that: Aligning the pre-processed frame with the reference frame in the spatial domain comprises: Apply the Iterative Closest Point (ICP) algorithm; or Acquire high-precision GPS information associated with the vehicle-mounted laser radar device and the roadside laser radar device.
9. The data acquisition system according to claim 6, characterized in that: It also includes a map building subsystem for building a high-precision map by superimposing pre-processed frames in the first continuous frame stream collected by the vehicle-mounted laser radar.
10. The data acquisition system according to claim 1, characterized in that: The roadside laser radar collects the second continuous frame stream from a high viewing angle.
11. A method for collecting training data for autonomous driving applications, characterized in that: include: Acquire a first continuous frame stream from a vehicle-mounted laser radar device installed on a data collection vehicle; Acquire a second continuous frame stream from a roadside laser radar device; Aligning the second continuous frame stream collected by the roadside laser radar with the first continuous frame stream collected by the vehicle-mounted laser radar in the time domain and the space domain; as well as Frames in the second continuous frame stream are superimposed on corresponding frames in the first continuous frame stream, thereby facilitating enhancement of the object point cloud in the first continuous frame stream.
12. The method according to claim 11, characterized in that It also includes synchronizing the operations of the vehicle-mounted laser radar device and the roadside laser radar device.
13. The method according to claim 12, characterized in that The operations of the vehicle-mounted LiDAR device and the roadside LiDAR device are triggered by the same high-precision Global Positioning System (GPS) signal.
14. The method according to claim 11, characterized in that It also includes determining the posture of the vehicle-mounted laser radar device and the posture of the roadside laser radar device associated with each frame in the first continuous frame stream.
15. The method according to claim 14, characterized in that Also included is pre-processing the frames by removing transient objects from the frames.
16. The method according to claim 15, characterized in that Determining a pose of the onboard lidar device associated with the frame includes aligning the pre-processed frame with a reference frame in a spatial domain; and Determining the pose of the roadside lidar device includes aligning a frame selected from the second continuous frame stream with the reference frame in the spatial domain.
17. The method according to claim 16, characterized in that Aligning the second continuous frame stream with the first continuous frame stream in the time domain and the spatial domain comprises: selecting a frame from a first continuous stream of frames; determining in the time domain a frame in the second continuous frame stream that is aligned with the selected frame; and Based on the pose of the onboard laser radar device and the pose of the roadside laser radar device associated with the selected frame, the identified frame is aligned with the selected frame in the spatial domain.
18. The method according to claim 16, characterized in that The aligning the pre-processed frame with the reference frame in the spatial domain comprises: Apply the Iterative Closest Point (ICP) algorithm; or Acquire high-precision GPS information associated with the vehicle-mounted laser radar device and the roadside laser radar device.
19. The method according to claim 16, characterized in that It also includes constructing a high-precision map by superimposing pre-processed frames in the first continuous frame stream collected by the vehicle-mounted laser radar.
20. The method according to claim 11, characterized in that The roadside laser radar collects the second continuous frame stream from a high viewing angle.