Incremental learning method and apparatus for vehicle perception system
Through the collaborative incremental learning method between vehicles and roadside facilities, the vehicle perception model is continuously trained using the high-performance sensors and computing units of the roadside facilities, which solves the problem of incorrect prediction of the vehicle perception model in complex traffic environments and improves the safety and reliability of intelligent connected vehicles.
Patent Information
- Application Number
- CN202410345620.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
The vehicle perception models in intelligent connected vehicles have difficulty coping with traffic environments outside the distribution of training datasets during actual road driving, resulting in incorrect predictions and increasing the risk of traffic accidents.
Through collaboration between vehicles and roadside facilities, the high-performance sensors and computing units of roadside facilities are used to perform incremental learning on the vehicle perception model. Traffic environment data from the vehicle and roadside perspectives are combined for data alignment and label information transmission to achieve continuous training of the vehicle perception model.
It improves the accuracy and adaptability of vehicle perception models, reduces the risk of misprediction in complex and changing traffic environments, and enhances the safety and reliability of intelligent connected vehicles.
Smart Images

Figure CN120708181A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to intelligent connected vehicle technology, and more specifically to perception technology in intelligent connected vehicles. Background Art
[0002] Driven by technological advances in artificial intelligence, communication networks, and hardware computing power, intelligent connected vehicle (ICV) technology has become a new trend in the automotive industry. ICV technology connects vehicles to the internet, leveraging intelligence and connectivity to provide assisted or autonomous driving capabilities and enhance vehicle safety. Vehicle perception technology is a crucial component of ICV technology. It integrates sensor, information, and computer technologies to form a system that monitors the vehicle's surroundings, roads, and other traffic participants in real time. It utilizes various machine learning models to process and analyze captured data, enabling perception, identification, tracking, and prediction of the traffic environment. This information is then used for assisted or autonomous driving, thereby enhancing the driver's experience and traffic safety.
[0003] Vehicle perception technology can identify other vehicles, obstacles, traffic signs, and the like in the traffic environment through a machine learning model (i.e., a perception model) deployed on the vehicle to perceive the traffic environment. This machine learning model is pre-trained using a traffic environment dataset. However, due to the complexity of the traffic environment and its continuous evolution over time, vehicles equipped with a perception model will inevitably encounter traffic environments outside the distribution of the training dataset during actual road driving. In such cases, the vehicle's perception model may make incorrect predictions / inferences, increasing the risk of traffic accidents.
[0004] Therefore, a method is needed to continuously train perception models for vehicle perception technology in order to further improve the safety and reliability of intelligent connected vehicles. Summary of the Invention
[0005] The following is a brief description of one or more aspects of this disclosure to provide a basic understanding of these aspects. This summary is not an extensive overview of all aspects, nor is it intended to identify the key elements of all aspects, nor is it intended to delineate the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that follows.
[0006] According to one aspect of the present disclosure, a method for incremental learning of a vehicle perception system is provided. The method includes capturing vehicle-perspective traffic environment data using onboard sensors of the vehicle perception system located on a vehicle; sending an auxiliary information request to a roadside facility, the auxiliary information request including metadata about the vehicle-perspective traffic environment data; receiving an auxiliary information response from the roadside facility, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data; and incrementally learning a perceptual deep neural network of the vehicle perception system using the vehicle-perspective traffic environment data and the label information.
[0007] According to one aspect of the present disclosure, a method for incremental learning of a vehicle perception system is provided. The method includes capturing roadside-perspective traffic environment data using roadside sensors located at roadside facilities; inferring the traffic environment based on the roadside-perspective traffic environment data using a perception deep neural network of the roadside facilities; receiving an auxiliary information request from a vehicle, the auxiliary information request including metadata of the vehicle-perspective traffic environment data; determining label information corresponding to the vehicle-perspective traffic environment data based on the metadata of the vehicle-perspective traffic environment data, the metadata of the roadside-perspective traffic environment data, and an inferred result; and sending an auxiliary information response to the vehicle, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data.
[0008] According to one aspect of the present disclosure, a device for incremental learning of a vehicle perception system is provided. The device includes a memory and a processor coupled to the memory. The processor is configured to capture vehicle-perspective traffic environment data using onboard sensors of a vehicle perception system located on a vehicle; send an auxiliary information request to a roadside facility, the auxiliary information request including metadata about the vehicle-perspective traffic environment data; receive an auxiliary information response from the roadside facility, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data; and perform incremental learning on a perceptual deep neural network of the vehicle perception system using the vehicle-perspective traffic environment data and the label information.
[0009] According to one aspect of the present disclosure, a device for incremental learning of a vehicle perception system is provided. The device includes a memory and a processor coupled to the memory. The processor is configured to capture roadside perspective traffic environment data through roadside sensors located at roadside facilities; infer the traffic environment based on the roadside perspective traffic environment data using a perception deep neural network of the roadside facilities; receive an auxiliary information request from a vehicle, the auxiliary information request including meta-information of the vehicle perspective traffic environment data; determine label information corresponding to the vehicle perspective traffic environment data based on the meta-information of the vehicle perspective traffic environment data, the meta-information of the roadside perspective traffic environment data, and an inference result; and send an auxiliary information response to the vehicle, the auxiliary information response including label information corresponding to the vehicle perspective traffic environment data.
[0010] According to one aspect of the present disclosure, a computer-readable medium is provided. The computer-readable medium stores a computer program for incremental learning of a vehicle perception system. When executed by a processor, the computer program causes the processor to capture vehicle-perspective traffic environment data using onboard sensors of a vehicle perception system located in a vehicle; send an auxiliary information request to a roadside facility, the auxiliary information request including metadata of the vehicle-perspective traffic environment data; receive an auxiliary information response from the roadside facility, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data; and perform incremental learning on a perceptual deep neural network of the vehicle perception system using the vehicle-perspective traffic environment data and the label information.
[0011] According to one aspect of the present disclosure, a computer-readable medium is provided. The computer-readable medium stores a computer program for incremental learning of a vehicle perception system. When executed by a processor, the computer program causes the processor to capture roadside-perspective traffic environment data through roadside sensors located at roadside facilities; use the perception deep neural network of the roadside facilities to infer the traffic environment based on the roadside-perspective traffic environment data; receive an auxiliary information request from a vehicle, the auxiliary information request including metadata of the vehicle-perspective traffic environment data; determine label information corresponding to the vehicle-perspective traffic environment data based on the metadata of the vehicle-perspective traffic environment data, the metadata of the roadside-perspective traffic environment data, and the inference result; and send an auxiliary information response to the vehicle, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data.
[0012] According to one aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program for incremental learning of a vehicle perception system. When executed by a processor, the computer program implements capturing vehicle-perspective traffic environment data through onboard sensors of the vehicle perception system located on a vehicle; sending an auxiliary information request to a roadside facility, the auxiliary information request including metadata of the vehicle-perspective traffic environment data; receiving an auxiliary information response from the roadside facility, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data; and performing incremental learning of a perceptual deep neural network of the vehicle perception system using the vehicle-perspective traffic environment data and the label information.
[0013] According to one aspect of the present disclosure, a computer program product is provided. The computer program product includes a computer program for incremental learning of a vehicle perception system. When executed by a processor, the computer program captures roadside perspective traffic environment data through roadside sensors located at roadside facilities; uses a perception deep neural network of the roadside facilities to infer the traffic environment based on the roadside perspective traffic environment data; receives an auxiliary information request from a vehicle, the auxiliary information request including meta-information of the vehicle perspective traffic environment data; determines label information corresponding to the vehicle perspective traffic environment data based on the meta-information of the vehicle perspective traffic environment data, the meta-information of the roadside perspective traffic environment data, and an inference result; and sends an auxiliary information response to the vehicle, the auxiliary information response including label information corresponding to the vehicle perspective traffic environment data.
[0014] It should be noted that one or more of the above aspects include features described in detail below and specifically recited in the claims. The following description and drawings set forth in detail some exemplary features from various aspects. These features are merely indicative of the various ways in which the principles of various aspects may be implemented, and the present disclosure is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The following drawings describe various embodiments of the present disclosure for illustrative purposes only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the methods and structures disclosed herein can be implemented without departing from the spirit and principles of the disclosure described herein.
[0016] Figure 1 A schematic diagram of a vehicle networking system according to an embodiment of the present disclosure is shown.
[0017] Figure 2 A block diagram of a vehicle and roadside facilities according to an embodiment of the present disclosure is shown.
[0018] Figure 3A flowchart of an incremental learning process for a vehicle perception system according to one embodiment of the present disclosure is shown.
[0019] Figure 4 A flowchart of a method for incremental learning of a vehicle perception system performed on a vehicle side according to an embodiment of the present disclosure is shown.
[0020] Figure 5 A flowchart of a method for incremental learning of a vehicle perception system performed on a roadside facility according to an embodiment of the present disclosure is shown.
[0021] Figure 6 A block diagram of an apparatus for incremental learning of a vehicle perception system according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0022] In the following description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present invention. However, one skilled in the relevant art will recognize that the present invention can be practiced without one or more of these specific details, or can be practiced using alternative methods, components, etc. In some instances, well-known structures and operations are not shown or described in detail to avoid unnecessarily obscuring the present invention.
[0023] Figure 1 The schematic diagram of the vehicle networking system according to one embodiment of the present disclosure is shown. The concept of the vehicle networking originates from the Internet of Things, namely the vehicle Internet of Things, which realizes all-round network connection between vehicles and roads, vehicles and vehicles, vehicles and people, and in-vehicle equipment through the new generation of information and communication technology (for example, 5G communication technology). The vehicle networking can use sensing technology to perceive the vehicle status information and traffic environment information, and realize intelligent control of vehicles and intelligent management of traffic through communication networks and intelligent information processing technology. Figure 1 As shown, the vehicle networking system may include at least a roadside facility 110 and a plurality of vehicles 120-160. The roadside facility 110 may be a sensor for monitoring the traffic environment (e.g., road conditions and vehicle driving conditions, etc.), or a roadside unit (RSU) for communicating with vehicles and other roadside facilities. Figure 1 Not shown, the roadside infrastructure may also include an edge computing unit located on the roadside and a remote cloud server. The roadside infrastructure can sense the driving status of vehicles on the road (e.g., vehicles 120-160), obstacles on the road (e.g., roadblock 170), and traffic signs and facilities on the road.
[0024] Roadside facilities can broadcast the sensed information to nearby vehicles to assist in the vehicle's control decisions. Vehicles can also use their own vehicle perception systems to obtain road environment information to provide hazard warnings to the driver or assist in driving decisions. For example, vehicle 120 can sense the roadblock 170 ahead and alert the driver or automatically perform obstacle avoidance operations. Vehicle 130 can sense the vehicle 140 ahead to assist the driver in following vehicle 140. Vehicles 150 and 160 can sense lane lines and assist the driver in staying within the lane. In addition, vehicles 120-160 can also exchange each other's driving information (for example, location, speed, etc.) and their respective perceived traffic environment information through the Internet of Vehicles to expand the vehicle's perception range.
[0025] The connected vehicle (IoV) allows vehicles to communicate in real time with other vehicles, roadside devices, and the internet, providing a safer and more efficient driving experience. However, the resulting data security issues within the IoV not only pose potential safety risks but also threaten the personal information of vehicle owners. Some national and local governments have also enacted laws and regulations related to data protection. While these laws and regulations protect IoV data security, they also restrict the sharing of IoV data across different countries, organizations, and equipment manufacturers.
[0026] Figure 2 FIG2 shows a block diagram of a roadside facility 210 and a vehicle 250 according to an embodiment of the present disclosure. The roadside facility 210 may be Figure 1 The roadside facilities 110 are shown. The vehicle 250 may be Figure 1Vehicles 120-160 are shown. Roadside facilities 210 may include roadside sensors 220. Roadside sensors 220 may include one or more of cameras, lidar, and millimeter-wave radar. Cameras can directly capture two-dimensional (2D) images of the traffic environment, but in poor lighting conditions, cameras may not be able to capture clear images of the road, especially of obstacles that do not actively illuminate. Millimeter-wave radar and lidar can obtain three-dimensional (3D) point cloud images of the traffic environment by actively emitting electromagnetic waves and laser signals and receiving the reflected signals. Millimeter-wave radar signals have good penetration through smoke, dust, rain, fog, and other objects, are highly resistant to interference in adverse weather conditions, and have a detection range of up to 200 meters. Lidar signals have a small laser beam divergence angle and concentrated energy, allowing them to detect longer distances than millimeter-wave radars, have higher sensitivity and resolution, and can detect both low- and high-speed moving targets, but are susceptible to adverse weather conditions. Compared to cameras, millimeter-wave radar and lidar have the disadvantage of not being able to recognize the color of objects (e.g., traffic lights) or the text on road signs. Since different types of sensors have their own advantages and disadvantages, roadside facilities can include different types of sensors and fuse the information captured by multiple sensors to improve the accuracy and reliability of the roadside perception system.
[0027] The roadside facility 210 may also include a computing unit 230 and a roadside unit (RSU) 240. The computing unit 230 may be an edge computing platform co-located or distributed with the roadside sensor 220 and / or the roadside unit 240, or may be a cloud computing platform located remotely and connected to other roadside facility units via a communication network. The computing unit 230 may be configured to implement a perception model for identifying / inferring target objects based on the traffic environment data captured by the roadside sensor 220. The perception model implemented by the computing unit 230 may be a trained deep neural network (DNN). For example, the perception DNN may include a SECOND (Sparsely Embedded Convolutional Detection) target detection model for analyzing and processing point cloud data collected by a radar. The perception DNN may also include a YOLO (You Only Look Once) target detection model for analyzing and processing image data collected by a camera. The roadside unit 240 may communicate with vehicles or other roadside facility units via, for example, a 5G communication network.
[0028] Vehicle 250 may include a vehicle perception system 260 and an onboard unit (OBU) 290. The OBU 290 may be used to communicate with roadside facilities 210 (e.g., roadside unit 240) and the onboard units of other vehicles. Vehicle perception system 260 may include onboard sensors 270 and a computing unit 280. Onboard sensors 270 may include one or more of a camera, a lidar, and a millimeter-wave radar. Due to vehicle cost considerations, the types of onboard sensors 270 are typically fewer than those of roadside sensors 220, and the performance of onboard sensors 270 is typically lower than that of roadside sensors 220. For example, onboard sensors 270 may only include a camera, and the performance of the camera (e.g., resolution, viewing angle, etc.) may be lower than that of a roadside sensor camera. The computing unit 280 may be co-located with the onboard sensors 270 in a module, or may be separately arranged in another module of the vehicle (e.g., a data processing unit (DPU) or a neural network processing unit (NPU). The computing unit 280 may be configured to implement a perception model, such as a SECOND model and / or a YOLO model, for identifying / inferring target objects based on traffic environment data captured by the onboard sensors 270. Due to considerations such as the computing power and power consumption of the computing unit 280, the perception model implemented in the computing unit 280 may have a simpler structure and parameters than the perception model in the roadside infrastructure. Consequently, the accuracy of the resulting inferences may be lower than that of the roadside infrastructure.
[0029] Due to the complexity of the traffic environment and the continuous development and change of the traffic environment over time, the computing unit 280 (i.e., the perception model) in the vehicle 250 may encounter abnormal or completely new traffic environments (i.e., traffic environment data outside the distribution of the training data set of the perception model) during the actual driving process of the vehicle 250. Therefore, in order to improve the performance of the perception model, it is necessary to use newly observed samples (i.e., new traffic environment data) throughout the life cycle of the perception model to continuously perform incremental learning (also known as continuous learning or lifelong learning) on the perception model. Incremental learning is a machine learning method that uses new training data obtained over time to perform dynamic supervised learning on a previously trained model. The goal of incremental learning is to enable the model to learn from new data without forgetting the knowledge it has learned. In other words, through incremental learning, the inferences made by the model can conform to both the distribution of the original training data set and the distribution of the new data set.
[0030] In order to perform incremental learning on perception models deployed on vehicles (e.g., DNNs), it is often possible to continuously train the vehicle's perception model using computing units on roadside facilities (e.g., edge computing platforms, cloud computing platforms). This approach requires transmitting new data samples observed by the vehicle to the roadside facilities. However, as mentioned above, there are data compliance requirements in the Internet of Vehicles. Since different vehicles and roadside facilities may belong to different countries, organizations, manufacturers, etc., transmitting vehicle data to the roadside facilities that perform model training may be restricted.
[0031] Therefore, the present disclosure provides an incremental learning method for a perception model in a vehicle perception system that is assisted by roadside facilities and implemented on the vehicle side.
[0032] Figure 3 FIG. 3 is a flow chart showing an incremental learning process for a vehicle perception system according to an embodiment of the present disclosure. Figure 1 The roadside facilities 110 shown and Figure 2 The roadside facilities 210 are shown. The vehicle 350 may be Figure 1 The vehicles 120-160 shown and Figure 2 Vehicle 250 is shown.
[0033] In step 352, vehicle 350 may capture traffic environment data around the vehicle through the onboard sensors of its vehicle perception system while driving on the road. The traffic environment data may be referred to as vehicle-perspective traffic environment data, and may include information such as obstacles on the road, traffic signs, other traffic participants, etc. In this embodiment, the onboard sensor may be a camera, and the vehicle-perspective traffic environment data may be a 2D image of the traffic environment captured by the camera, and the 2D image may be a frame image in the traffic environment video image captured by the camera. In other embodiments, the onboard sensor may also be other types of sensors (e.g., lidar), or may include multiple different types of sensors, and the vehicle perception system may fuse the data captured by multiple sensors to obtain vehicle-perspective traffic environment data. Vehicle 350 may save the captured traffic environment data to a memory (e.g., a cache).
[0034] Vehicle 350 can record information such as the vehicle's position information, time information, and internal parameters of the on-board sensors when capturing traffic environment data through on-board sensors. This information can be referred to as meta-information of the vehicle's perspective traffic environment data. The vehicle's position information can be the latitude and longitude data obtained from a positioning module (e.g., a GPS module and / or a Beidou module). The time information can be the timestamp of the captured image. The internal parameters of the on-board sensors can include, for example, the focal length and field of view of the camera, etc. In addition, the meta-information of the vehicle's perspective traffic environment data also includes external parameters of the on-board sensors, such as the installation position of the camera relative to the vehicle, the angle relative to the direction of the vehicle head (including vertical angle and horizontal angle), etc. The internal and external parameters of the on-board sensors can be collectively referred to as calibration parameters of the on-board sensors. Vehicle 350 can save the meta-information of the traffic environment data together with the corresponding traffic environment data in a memory (e.g., a cache).
[0035] In step 354, vehicle 350 can utilize the trained deep neural network in the vehicle perception system to infer the traffic environment surrounding the vehicle based on the captured vehicle-perspective traffic environment data. In this embodiment, the perception DNN can perform target detection based on the 2D image captured by the camera to infer objects in front of the road. For example, the perception DNN can infer the presence of pedestrians, vehicles, etc. in the traffic environment in front of the vehicle. The vehicle perception system can further track the detected target objects (e.g., pedestrians) to assist in driving decisions or issue timely reminders to the driver. In other embodiments, the perception DNN can be used for other perception tasks such as obstacle detection, lane detection, and traffic sign detection. The vehicle perception system can include multiple perception models for different perception tasks. Vehicle 350 can also save the inference results of the traffic environment to memory (e.g., cache). Vehicle 350 can repeatedly and continuously perform steps 352 and 354.
[0036] While the vehicle 350 is performing traffic environment perception from the vehicle perspective, in steps 312 and 314 , the roadside facility 310 may also perform traffic environment perception from the roadside perspective.
[0037] Specifically, in step 312, the roadside facility 310 may capture roadside perspective traffic environment data through a roadside sensor. In this embodiment, the roadside sensor may be a lidar, and the roadside perspective traffic environment data may be a 3D point cloud map of the traffic environment captured by the lidar. In other embodiments, the roadside sensor may also be other types of sensors (e.g., a high-definition camera), or may include multiple different types of sensors, and the roadside facility may fuse the data captured by the multiple sensors to obtain roadside perspective traffic environment data. Since roadside sensors generally have higher performance and more types than vehicle-mounted sensors, and the locations where roadside sensors are installed also have better viewing angles than vehicle-mounted sensors, the roadside perspective traffic environment data may have richer and more accurate information than the vehicle perspective traffic environment data.
[0038] When roadside sensors capture traffic environment data, the roadside infrastructure 310 can also record the timestamp corresponding to the traffic environment data and the roadside sensor calibration parameters. For example, the calibration parameters of a lidar (lidar) can include internal parameters and external parameters. Internal parameters include the lidar's photoelectric characteristics, scanning angle, and resolution, while external parameters include the lidar's position and heading angle. The roadside infrastructure 310 can store the roadside viewpoint traffic environment data and corresponding metadata in a memory (e.g., cache).
[0039] In step 314, the roadside facility 310 can utilize the perception deep neural network implemented on its edge computing platform or cloud computing platform to infer the traffic environment around, for example, vehicle 350 based on the captured roadside-perspective traffic environment data. Because the roadside perception DNN can have a more complex network structure and more optimized network parameters than the vehicle-side perception DNN, and the roadside-perspective traffic environment data contains richer information than the vehicle-perspective traffic environment data, the inference results of the roadside facility 310 can be more accurate than the inference results of the vehicle 350. The roadside facility 310 can also save the inference results of the traffic environment to a memory (e.g., a cache). The roadside facility 310 can repeatedly and continuously perform steps 312 and 314.
[0040] In this embodiment, the on-board sensor is a camera and the roadside sensor is a lidar. If the DNN used for the target detection task of the vehicle perception system has not been trained using images obtained under low-light conditions, then the vehicle 350 may not be able to accurately detect the target objects in the traffic environment in a traffic environment with insufficient lighting. However, since the roadside facility 310 uses a lidar, the roadside sensor can accurately detect the target objects in the traffic environment even in the absence of light. In this case, the vehicle 350 can obtain the label information of the target objects in the traffic environment image under low-light conditions with the assistance of the roadside facility 310 as real data (ground truth) for training and learning the perception ability under low-light conditions.
[0041] Specifically, in step 360, vehicle 350 may send an auxiliary information request to roadside facility 310. Vehicle 350 may send the auxiliary information request to the RSUs in surrounding roadside facilities via the OBU in a broadcast or unicast manner. The auxiliary information request may only include the metadata of the vehicle-perspective traffic environment data captured by the vehicle. If it complies with IoV data compliance requirements, the auxiliary information request may also include the vehicle-perspective traffic environment data itself. The auxiliary information request may also include information such as the vehicle identification and roadside facility identification.
[0042] Vehicle 350 may periodically or periodically send auxiliary information requests to roadside facilities. Preferably, vehicle 350 may send auxiliary information requests to roadside facilities when it determines it is encountering an abnormal scenario. In this embodiment, vehicle 350 may send an auxiliary information request to roadside facilities 310 when the confidence level of the inference performed in step 354 falls below a threshold. For example, in low-light scenarios, the vehicle perception system may perform target detection using a perception model based on traffic environment images captured by a camera. The vehicle perception model may infer that a pedestrian may be present ahead on the road and output a confidence level of 55%. In this case, if the threshold for triggering an auxiliary information request is 60%, vehicle 350 may send an auxiliary information request to roadside facilities 310. The confidence level threshold can be set based on different scenarios, for example, 80%. In other embodiments, the vehicle may trigger the sending of an auxiliary information request to roadside facilities when it determines it is encountering abnormal weather (e.g., smog, sandstorms, etc.) or when it enters another area (e.g., another city, high altitude, etc.).
[0043] After the roadside facility 310 receives the auxiliary information request from the vehicle 350 in step 360, it can determine in step 316 the label information corresponding to the vehicle perspective traffic environment data associated with the metadata of the received vehicle perspective traffic environment data based on the metadata of the vehicle perspective traffic environment data in the received auxiliary information request, the metadata of the roadside perspective traffic environment data obtained or stored at the roadside facility, and the inference result obtained in step 314.
[0044] Roadside facility 310 can determine roadside traffic environment data from the same or similar time as the vehicle-perspective traffic environment data based on the timestamp in the metadata of the vehicle-perspective traffic environment data and the timestamp in the metadata of the roadside-perspective traffic environment data. In perception scenarios targeting fixed targets, such as obstacle detection, roadside facility 310 may not need to compare timestamp information when detecting and marking roadblocks that are fixed over a period of time.
[0045] The roadside facility 310 can determine roadside-perspective traffic environment data that covers the field of view of the vehicle-perspective traffic environment data based on the vehicle position information and the relative position, viewing angle, and focal length of the onboard sensor and the vehicle in the metadata of the vehicle-perspective traffic environment data, as well as the roadside sensor position and viewing angle information in the metadata of the roadside-perspective traffic environment data. Next, the roadside facility 310 can convert the roadside-perspective traffic environment data from the roadside perspective to the vehicle perspective through coordinate axis translation and rotation operations based on the position and viewing angle of the onboard sensor and the position and viewing angle of the roadside sensor, thereby aligning and registering the roadside-perspective traffic environment data with the vehicle-perspective traffic environment data. In this embodiment, the onboard sensor is a camera, the vehicle-perspective traffic environment data is a 2D image, and the roadside sensor is a lidar, and the roadside-perspective traffic environment data is a 3D point cloud. Therefore, the roadside facility 310 can project the 3D point cloud of the traffic environment captured by the lidar onto a 2D image to facilitate alignment and registration with the vehicle-perspective traffic environment data.
[0046] By aligning and registering the roadside perspective traffic environment data with the vehicle perspective traffic environment data, the roadside facility 310 can determine the label information corresponding to the vehicle perspective traffic environment data based on the correspondence between the inference result obtained in step 314 and the roadside perspective traffic environment data converted to the vehicle perspective. For different perception tasks, the label information may contain different content. For example, for the target detection task, the label information may include the bounding box position of the target object from the vehicle perspective (the bounding box position may be the position relative to the 2D image) and the target object type corresponding to the bounding box position (for example, pedestrians, vehicles, etc.). For the obstacle detection task, the label information may only include the bounding box position of the obstacle from the vehicle perspective.
[0047] The roadside facility 310 can operate in different processing orders and / or with different data processing units, thereby aligning and registering the roadside perspective traffic environment data with the vehicle perspective traffic environment data to generate label information corresponding to the vehicle perspective traffic environment data based on the inference results corresponding to the roadside perspective traffic environment data. For example, the roadside facility can first perform 3D to 2D projection processing on the roadside perspective traffic environment data corresponding to the field of view range of the vehicle perspective traffic environment data as a whole, and then convert the 2D image of the roadside perspective to the 2D image of the vehicle perspective, or first convert the 3D point cloud image of the roadside perspective to the 3D point cloud image of the vehicle perspective, and then perform 3D to 2D projection. The roadside facility can also directly perform perspective conversion (i.e., coordinate system conversion) and 3D-2D projection based on the bounding box of the detected target object (e.g., vehicle, obstacle, etc.) based on the inference results, thereby mapping the bounding box position of the target object inferred based on the roadside perspective traffic environment data to the vehicle perspective as the label information corresponding to the vehicle perspective traffic environment data.
[0048] In step 362, the roadside facility 310 may send an auxiliary information response to the vehicle 350. The auxiliary information response may include the tag information corresponding to the vehicle-perspective traffic environment data determined in step 316. If the data compliance requirements of the Internet of Vehicles are met, the auxiliary information request may also include the corresponding roadside-perspective traffic environment data and / or metadata of the roadside-perspective traffic environment data. The auxiliary information response may also include information such as the vehicle identification and the roadside facility identification. The roadside facility 310 may send the auxiliary information response to the vehicle 350 via the RSU in a broadcast or unicast manner.
[0049] After receiving the auxiliary information response from the roadside facility 310 in step 316, vehicle 350 may generate new data samples in step 356 based on the locally stored vehicle-perspective traffic environment data and the label information corresponding to the vehicle-perspective traffic environment data in the auxiliary information response, for incremental learning of the perception deep neural network in the vehicle perception system. Vehicle 350 may save the generated new data samples to local memory and, after collecting a certain number of new data samples, perform incremental learning of the perception model using a local computing unit during idle time (e.g., when not driving).
[0050] Figure 4 FIG. 4 is a flow chart showing a method 400 for incremental learning of a vehicle perception system executed on a vehicle side according to an embodiment of the present disclosure. Figure 1 Vehicles 120-160, Figure 2 Vehicles in 250 and Figure 3 The vehicle 350 is executed.
[0051] In block 410, method 400 may capture vehicle-perspective traffic environment data using onboard sensors of a vehicle perception system located in the vehicle. The onboard sensors may include one or more of a camera, a lidar, a millimeter-wave radar, etc., and accordingly, the vehicle-perspective traffic environment data may include a 2D image and / or a 3D point cloud image.
[0052] In block 420, method 400 may send an auxiliary information request to the roadside infrastructure. The auxiliary information request may include metadata about the vehicle-perspective traffic environment data. The metadata about the vehicle-perspective traffic environment data may include, for example, a timestamp, the vehicle's location, and calibration parameters of the onboard sensors. The calibration parameters of the onboard sensors may include external parameters and internal parameters. External parameters may include, for example, the sensor's installation position and angle. Different types of sensors may have different internal parameters. For example, the internal parameters of a camera may include focal length, field of view, and the like.
[0053] In one embodiment, method 400 can utilize the perception deep neural network of the vehicle's perception system to infer the vehicle's traffic environment based on the captured vehicle-perspective traffic environment data, and when the confidence level of the inference is lower than a threshold, send an auxiliary information request to the roadside equipment. In other embodiments, method 500 can send an auxiliary information request to the roadside equipment periodically or based on other triggering events.
[0054] In block 430, method 400 may receive an auxiliary information response from the roadside infrastructure. The auxiliary information response may include label information corresponding to the vehicle-perspective traffic environment data. The label information is determined by the roadside infrastructure based on the roadside-perspective traffic environment data, and by performing 3D-to-2D projection and roadside-perspective-to-vehicle-perspective conversion. Because the sensors and perception models of roadside infrastructure typically have higher performance, the label information inferred by the roadside infrastructure can be considered the ground truth of the perception task results from the vehicle's perspective.
[0055] In block 440, method 400 may utilize the vehicle-perspective traffic environment data and corresponding label information to incrementally learn the perception deep neural network of the vehicle perception system. The perception deep neural network of the vehicle perception system may be used for one or more different perception tasks, such as object detection, obstacle detection, and lane detection. The perception deep neural network may include multiple different DNNs for different perception tasks.
[0056] Figure 5 1 shows a flow chart of a method 500 for incremental learning of a vehicle perception system executed at a roadside facility according to an embodiment of the present disclosure. Figure 1Roadside facilities 110, Figure 2 Roadside facilities 210 and Figure 3 The roadside facilities 310 in the method 500 can assist the vehicle in performing incremental learning on the perception model in the vehicle perception system.
[0057] In block 510, method 500 may capture roadside-perspective traffic environment data using roadside sensors located at roadside facilities. Roadside sensors may include one or more of cameras, lidar, and millimeter-wave radar, and accordingly, the roadside-perspective traffic environment data may include 2D images and / or 3D point clouds. Generally, roadside sensors offer higher performance, a wider variety, and a better viewing angle than onboard sensors. Therefore, roadside-perspective traffic environment data contains richer and more accurate information than vehicle-perspective traffic environment data.
[0058] In block 520, method 500 may use the roadside infrastructure's perception deep neural network to infer the traffic environment based on the roadside-view traffic environment data. The roadside perception deep neural network may be used for one or more different perception tasks, such as object detection, obstacle detection, and lane detection. The perception deep neural network may include multiple different DNNs for different perception tasks. Method 500 may store the roadside-view traffic environment data and the corresponding inference results in local or cloud storage.
[0059] In block 530, method 500 may receive an auxiliary information request from the vehicle. The auxiliary information request may include metadata of the vehicle-perspective traffic environment data. The metadata of the vehicle-perspective traffic environment data may include a timestamp, the vehicle's location, and calibration parameters of the onboard sensors. The calibration parameters of the onboard sensors may include external parameters and internal parameters. External parameters may include, for example, the sensor's installation position and angle. Different types of sensors may have different internal parameters. For example, the internal parameters of a camera may include focal length, field of view, and the like.
[0060] In block 540, method 500 may determine label information corresponding to the vehicle-perspective traffic environment data based on the metadata of the vehicle-perspective traffic environment data, the metadata of the roadside-perspective traffic environment data, and the results inferred in block 520. The metadata of the roadside-perspective traffic environment data may include a timestamp and calibration parameters of the roadside sensor. The calibration parameters of the roadside sensor may include external parameters and internal parameters. External parameters may include, for example, the sensor's mounting position and angle. Different types of sensors may have different internal parameters. For example, the internal parameters of a lidar may include photoelectric characteristics, scanning angle, and resolution. To determine the label information corresponding to the vehicle-perspective traffic environment data, method 500 may include converting the roadside-perspective traffic environment data from a roadside perspective to a vehicle perspective based on the metadata of the vehicle-perspective traffic environment data and the metadata of the roadside-perspective traffic environment data. If the roadside-perspective traffic environment data includes a three-dimensional point cloud captured by a lidar or millimeter-wave radar, method 500 may also include projecting the three-dimensional point cloud of the roadside-perspective traffic environment data into a two-dimensional image.
[0061] In block 550, method 500 may send an auxiliary information response to the vehicle. The auxiliary information response may include label information corresponding to the vehicle-perspective traffic environment data. For example, the label information may include the bounding box position of an obstacle in a two-dimensional image of the vehicle-perspective traffic environment, or the bounding box position and type of a target object, etc. The vehicle may use this label information, together with the corresponding vehicle-perspective traffic environment data, to construct a new data sample for incremental learning on the vehicle side.
[0062] Figure 6 FIG. 6 shows a block diagram of an apparatus 600 for incremental learning of a vehicle perception system according to an embodiment of the present disclosure. Figure 6 As shown, the apparatus 600 may include a memory 610 and at least one processor 620 coupled to the memory 610 .
[0063] In one embodiment, the device 600 may be a vehicle 120-160, 250, 350 or an intelligent module on the vehicle (eg, the vehicle perception system 260 or the computing unit 280). The processor 620 may be configured to execute the above combined Figure 3 Described processing flow and combination Figure 4For example, the processor 920 may be configured to capture vehicle-perspective traffic environment data through onboard sensors of a vehicle perception system located at a vehicle; send an auxiliary information request to a roadside facility, the auxiliary information request including meta-information of the vehicle-perspective traffic environment data; receive an auxiliary information response from the roadside facility, the auxiliary information response including label information corresponding to the vehicle-perspective traffic environment data; and perform incremental learning on a perception deep neural network of the vehicle perception system using the vehicle-perspective traffic environment data and the label information.
[0064] In another embodiment, the device 600 may be a roadside facility 110, 210, 310 or a computing unit 230 in the roadside facility (eg, an edge computing platform or a cloud computing platform). The processor 620 may be configured to execute the above-mentioned Figure 3 Described processing flow and combination Figure 5 For example, the processor 620 may be configured to capture roadside perspective traffic environment data through a roadside sensor located at a roadside facility; infer the traffic environment based on the roadside perspective traffic environment data using a perception deep neural network of the roadside facility; receive an auxiliary information request from a vehicle, the auxiliary information request including meta-information of the vehicle perspective traffic environment data; determine label information corresponding to the vehicle perspective traffic environment data based on the meta-information of the vehicle perspective traffic environment data, the meta-information of the roadside perspective traffic environment data, and an inference result; and send an auxiliary information response to the vehicle, the auxiliary information response including label information corresponding to the vehicle perspective traffic environment data.
[0065] The processor 620 may be a general-purpose processor, or may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. The memory 610 may include a non-volatile memory for storing computer program code implementing the disclosed solution and parameters of the perception model. The memory 610 may also include a cache memory for temporarily storing data received during the execution of the processor (e.g., traffic environment data) and processed data (e.g., inference results).
[0066] The various operations described in conjunction with the present disclosure may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. In one embodiment of the present disclosure, a computer program product for incremental learning of a vehicle perception system may include a computer program for performing the above-mentioned operations in conjunction with the present disclosure. Figure 3 Described processing flow and combination Figure 4-5In another embodiment of the present disclosure, a computer readable medium may store computer program code for incremental learning of a vehicle perception system. When executed by a processor, the computer program code may cause the processor to perform the above-mentioned method. Figure 3 Described processing flow and combination Figure 4-5 The method described herein. Computer-readable media includes both non-transitory computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Any connection may be properly referred to as a computer-readable medium. Other embodiments and implementations are within the scope of this disclosure.
[0067] In addition to what is described herein, various modifications may be made to the disclosed embodiments and implementations of the present invention without departing from the scope of the disclosed embodiments and implementations of the present invention. Therefore, the description and examples herein should be interpreted as illustrative rather than limiting. The scope of the present invention should be measured solely by reference to the claims.
Claims
1. A method for incremental learning of a vehicle perception system, comprising: Capturing traffic environment data from the vehicle's perspective using onboard sensors of the vehicle perception system located in the vehicle; Sending an auxiliary information request to a roadside facility, the auxiliary information request including metadata of the vehicle-view traffic environment data; receiving an auxiliary information response from the roadside facility, the auxiliary information response including label information corresponding to the vehicle-view traffic environment data; as well as The vehicle-perspective traffic environment data and the label information are used to perform incremental learning on the perception deep neural network of the vehicle perception system.
2. The method according to claim 1, wherein The vehicle-mounted sensors include one or more of a camera, a laser radar, and a millimeter-wave radar.
3. The method according to claim 1, wherein The meta information of the vehicle perspective traffic environment data includes: timestamp, the location of the vehicle, and calibration parameters of the onboard sensors; or The position of the vehicle, and calibration parameters of the on-board sensors.
4. The method according to claim 1, wherein The perception deep neural network of the vehicle perception system is used for one or more perception tasks among target detection, obstacle detection, and lane line detection.
5. The method according to claim 1, further comprising: Inferring the traffic environment of the vehicle by the perception deep neural network based on the vehicle-perspective traffic environment data; And among them, The sending of the auxiliary information request to the roadside facility includes: when the confidence level of the inference is lower than a threshold, sending the auxiliary information request to the roadside facility.
6. A method for incremental learning of a vehicle perception system, comprising: Capture traffic environment data from the roadside perspective through roadside sensors located at roadside facilities; Inferring the traffic environment based on the roadside perspective traffic environment data by the perception deep neural network of the roadside facilities; receiving an auxiliary information request from a vehicle, the auxiliary information request including metadata of traffic environment data from the vehicle's perspective; Determining label information corresponding to the vehicle-perspective traffic environment data based on the meta-information of the vehicle-perspective traffic environment data, the meta-information of the roadside-perspective traffic environment data, and the inference result; as well as An auxiliary information response is sent to the vehicle, where the auxiliary information response includes the tag information corresponding to the vehicle-view traffic environment data.
7. The method according to claim 6, wherein: The roadside sensor includes one or more of a camera, a lidar, and a millimeter-wave radar.
8. The method according to claim 6, wherein: The meta information of the vehicle perspective traffic environment data includes: timestamp, the location of the vehicle, and calibration parameters of the onboard sensors; or The position of the vehicle, and calibration parameters of the on-board sensors.
9. The method according to claim 6, wherein: The meta-information of the roadside perspective traffic environment data includes a timestamp and calibration parameters of the roadside sensor.
10. The method according to claim 6, wherein: The perception deep neural network of the roadside facility is used for one or more perception tasks of target detection, obstacle detection, and lane line detection.
11. The method according to claim 6, wherein: The determining, based on the meta information of the vehicle-perspective traffic environment data, the meta information of the roadside-perspective traffic environment data, and the inference result, of label information corresponding to the vehicle-perspective traffic environment data includes: The roadside perspective traffic environment data is converted from a roadside perspective to a vehicle perspective according to the meta information of the vehicle perspective traffic environment data and the meta information of the roadside perspective traffic environment data.
12. The method according to claim 6, wherein: The roadside sensor includes a lidar or a millimeter-wave radar, and wherein, determining the label information corresponding to the vehicle-perspective traffic environment data based on the meta-information of the vehicle-perspective traffic environment data, the meta-information of the roadside-perspective traffic environment data, and the inference result includes: projecting a three-dimensional point cloud map of the roadside-perspective traffic environment data into a two-dimensional image.
13. A device for incremental learning of a vehicle perception system, comprising: Memory; A processor is coupled to the memory and configured to execute the method according to any one of claims 1-12.
14. A computer-readable medium storing a computer program for incremental learning of a vehicle perception system, wherein the computer program, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 12.
15. A computer program product comprising a computer program for incremental learning of a vehicle perception system, the computer program implementing the method according to any one of claims 1 to 12 when executed by a processor.