Detection and classification of traffic signs based on camera-radar fusion

By using a sensing system in autonomous vehicles to acquire camera and radar images and generate features through neural networks for processing, the problem of difficult to quickly and accurately detect and classify traffic signs in the prior art is solved, efficient traffic sign recognition is achieved, and driving safety of autonomous vehicles is improved.

CN120164183APending Publication Date: 2025-06-17WAYMO LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411831397.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-15
Filing Date
2024-12-12
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately detect and classify traffic signs in driving environments, affecting the driving path selection and safety of autonomous vehicles.

Method used

The vehicle's sensing system is used to acquire camera images and radar images, and generate camera features and radar features through a first neural network, combining these features for processing to identify traffic signs in the environment.

Benefits of technology

It realizes rapid and accurate detection and classification of traffic signs without deploying expensive lidar sensors, improving the driving path selection responsiveness and safety of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164183A_ABST
    Figure CN120164183A_ABST
Patent Text Reader

Abstract

The disclosed systems and techniques facilitate efficient detection and classification of traffic signs in a driving environment. The disclosed techniques include obtaining a first set of perspective camera images of an environment and a second set of radar images of the environment using a sensing system of a vehicle. The technique further includes generating one or more camera features characterizing the first set of images using a first neural network, generating one or more radar features characterizing the second set of images using a second neural network, and processing the one or more camera features and the one or more radar features to obtain an identification of the one or more traffic signs in the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification generally relates to autonomous vehicles. More specifically, this specification relates to the rapid and accurate detection and classification of traffic signs in a driving environment. Background Art

[0002] Autonomous (fully or partially self-driving) vehicles (AVs) operate by sensing the external environment using various electromagnetic (e.g., radar and optical) and non-electromagnetic (e.g., audio and humidity) sensors. Some autonomous driving vehicles map a driving path through the environment based on the sensed data. The driving path can be determined based on Global Positioning System (GPS) data and road map data. While GPS and road map data can provide information about the static aspects of the environment (buildings, street layout, road closures, etc.), dynamic information (such as information about other vehicles, pedestrians, street lights, etc.) is obtained from the sensed data collected contemporaneously. The accuracy and safety of the driving path and speed regime selected by the autonomous vehicle depend on the timely and accurate identification of the various objects present in the external environment, and on the ability of the driving algorithm to process information about the environment and provide correct instructions to the vehicle control and powertrain systems. Summary of the Invention

[0003] In one embodiment, a system is disclosed that includes a sensing system of a vehicle and a data processing system. The sensing system is configured to obtain a set of camera images of the environment and a set of radar images of the environment. The data processing system is configured to generate, using a first neural network (NN), one or more camera features characterizing the set of camera images, and to generate, using a second NN, one or more radar features characterizing the set of radar images. The data processing system is further configured to process the one or more camera features and the one or more radar features to obtain an identification of one or more traffic signs in the environment.

[0004] In another embodiment, a method is disclosed that includes obtaining, using a sensing system of a vehicle, a set of camera images of the environment and a set of radar images of the environment. The method further includes generating, using a first NN, one or more camera features characterizing the set of camera images. The method further includes generating, using a second NN, one or more radar features characterizing the set of radar images. The method further includes processing the one or more camera features and the one or more radar features to obtain an identification of one or more traffic signs in the environment.

[0005] In yet another embodiment, an autonomous vehicle is disclosed that includes one or more cameras configured to acquire a set of camera images of the environment and one or more radar sensors configured to acquire a set of radar images of the environment. The autonomous driving vehicle further includes a vehicle data processing system configured to generate, using a first NN, one or more camera features characterizing the set of camera images and to generate, using a second NN, one or more radar features characterizing the set of radar images. The data processing system is further configured to process the one or more camera features and the one or more radar features to obtain an identification of one or more traffic signs in the environment. The autonomous vehicle further includes an autonomous vehicle control system configured to cause the autonomous vehicle to follow a driving path selected in view of the identification of the one or more traffic signs. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present disclosure is illustrated by way of example and not limitation, and can be more fully understood when considered in conjunction with the following detailed description, in which:

[0007] Figure 1 is a diagram showing components of an example vehicle capable of deploying effective sign detection and classification in a driving environment, according to some embodiments of the present disclosure.

[0008] Figure 2 is a diagram showing an example architecture of a sign detection and classification model that can be used for training and deploying a sign detection and classification model capable of detecting and reading traffic signs in a driving environment, according to some embodiments of the present disclosure.

[0009] Figures 3A to 3B illustrates an example operation of a sign detection and classification model capable of effectively detecting and reading traffic signs in a driving environment, according to some embodiments of the present disclosure. Figure 3A illustrates a first part of the SDCM operation, which includes individual processing of camera images and radar images. Figure 3B illustrates a second part of the SDCM operation, which includes processing of combined camera and radar features.

[0010] Figure 4 is a schematic diagram of an example driving environment of a vehicle deploying a sign detection and classification model for detecting and reading traffic signs, according to some embodiments of the present disclosure.

[0011] Figure 5 illustrates an example architecture of a sign detection and classification model enhanced with an additional branch for processing camera features, according to some embodiments of the present disclosure.

[0012] Figure 6Shows another example architecture of a sign detection and classification model enhanced with additional branches processed by camera features according to some embodiments of the present disclosure.

[0013] Figure 7 Shows an example method of deploying a sign detection and classification model that uses a combination of camera and radar images to accurately identify and read traffic signs in a driving environment according to some embodiments of the present disclosure.

[0014] Figure 8 Depicts a block diagram of an example computer device capable of training and / or deploying a sign detection and classification model that uses a combination of camera and radar images to accurately identify and read traffic signs in a driving environment according to some embodiments of the present disclosure. Detailed Description

[0015] Autonomous vehicles or vehicles that deploy various advanced driver assistance features can use multiple sensor modalities to facilitate the detection of objects in the external environment and the prediction of the future trajectories of such objects. Sensors can include radio detection and ranging (radar) sensors, light detection and ranging (lidar) sensors, digital cameras, ultrasonic sensors, position sensors, etc. Different types of sensors can provide different and complementary benefits. For example, radar and lidar emit electromagnetic signals (radio signals or light signals) that are reflected from objects and carry back information about the distance to the object (e.g., determined based on the time of flight of the signal) and the speed of the object (e.g., based on the Doppler shift of the frequency of the reflected signal). Radar and lidar can scan the entire 360-degree field of view by using a series of consecutive sensing frames. A sensing frame can include many reflections that cover the external environment in a dense grid of return points. Each return point can be associated with the distance to the corresponding reflecting object and the radial velocity of the reflecting object (the velocity component along the line of sight).

[0016] Lidar has high spatial resolution due to its sub-micron light wavelength, which allows many closely-spaced return points to be obtained from the same object. This enables accurate detection and tracking of objects once they are within the range of the lidar sensor. Radar sensors are inexpensive, require less maintenance compared to lidar sensors, have a larger working distance range, and have good tolerance for adverse weather conditions. Cameras (e.g., cameras or video cameras) capture a two-dimensional projection of the three-dimensional external space onto an image plane (or some other non-planar imaging surface) and can acquire high-resolution images at both shorter and longer distances.

[0017] Various sensors of the vehicle's sensing system (e.g., lidar, radar, cameras, and / or other sensors such as sonar) capture complementary depictions of objects in the vehicle's environment. The vehicle's perception system identifies objects based on the appearance, motion state, trajectory of the objects, and / or other attributes. For example, lidar can accurately map the shape of one or more objects (using multiple return points) and can further determine the distance to these objects and / or the speed of the objects. Cameras can obtain visual images of the objects. The perception system can map the shapes and positions of various objects in the environment (obtained from lidar data) to their visual depictions (obtained from camera data) and perform multiple computer vision operations such as segmenting (clustering) the sensed data between individual objects (clusters) and identifying the type / brand / model / etc. of each object. The prediction and planning system can track the motion of various objects (including but not limited to position and speed) over multiple times and then extrapolate the previously observed motion into the future. This predicted motion can be used by various vehicle control systems to select a driving path that takes these objects into account, e.g., avoiding the objects, slowing down the vehicle in the presence of the objects, and / or taking some other appropriate action.

[0018] In addition to detecting animate objects, a vehicle's sensing system is also used for the important purpose of identifying various semantic information, such as markings on the pavement (e.g., the boundaries of driving lanes, the position of stop lines, etc.), traffic lights, and traffic signs, including new and temporary signs that do not exist in static (even regularly updated) road map information. The information conveyed by traffic signs can be quite complex. For example, some signs can prescribe driving behavior for all road users (e.g., stop signs, no entry signs, speed limit signs, etc.), some signs can regulate the driving behavior of only certain types of vehicles (e.g., trucks), only vehicles located in certain driving lanes / streets, etc., or only apply at certain times of the day, etc. However, the semantic meaning of a traffic sign can depend significantly not only on the content (picture) of the traffic sign, but also on the position of the sign (e.g., in a multi-lane driving environment), the direction the sign is facing, etc. Therefore, it is important to accurately and timely detect the content and position of the sign. Camera images can capture an accurate depiction of the sign. Such a depiction can be used (e.g., by a computer vision model) to identify the type of the sign (e.g., speed limit sign) and the semantic meaning (value) of the sign (e.g., a 40 mph speed limit). However, camera images are two-dimensional projections of the external environment and do not include explicit depth (distance) information of the depicted objects. Sometimes the distance to an object can be estimated from the image and various metadata (e.g., the focal length of the camera lens that captured the image). However, compared to sensors equipped with ToF functionality (e.g., lidar and / or radar), such an estimate results in significantly less accurate distance determination.

[0019] LiDAR has high imaging resolution, which can be comparable to camera resolution (for dense point clouds) and can potentially be used to detect and read traffic signs. For example, LiDAR return points can be used to determine the exact placement and shape of signs. Additionally, the intensity of LiDAR returns can be used to determine the content of signs. For example, the black portions of a sign can reflect LiDAR signals differently than the white portions of the sign, allowing a computer vision system to read the text of the sign. However, LiDAR sensors are expensive and require complex maintenance. Thus, LiDAR is generally not deployed in driver assistance systems that do not provide full autonomy capabilities (e.g., level 2, 3, and 4 systems). Radar, on the other hand, is much cheaper, requires little maintenance, and is more likely to be deployed in such systems and vehicles. Radar includes ToF capabilities and is able to accurately determine the distance to an object. However, due to the longer wavelength of electromagnetic signals, radar has a much lower resolution than LiDAR. For example, a 24 GHz radar uses electromagnetic waves with a wavelength λ≈1.25 cm and a resolution of approximately Δd≈√λL≈1 m at a distance L = 100 m. Accordingly, while radar sensors can detect reflections of radar signals from an object (a potential sign), the radar resolution may not be sufficient to read the actual semantic content of the sign or determine the shape of the sign (e.g., detect that the object is the octagonal shape of a stop sign).

[0020] Aspects and embodiments of the present disclosure address these and other challenges of modern perception techniques by implementing methods and systems that utilize camera and radar sensing modalities to effectively detect and classify traffic signs. More specifically, each image stream can be processed by a corresponding modality network. For example, camera images can be processed by a camera network, while radar images can be processed by a radar network. The radar network generates a set of radar features (feature vectors, embeddings) associated with specific coordinates x, y of a two-dimensional bird's-eye view (BEV) grid such that the radar features F R(x, y; t) represents the presence (or absence) of a reflected object or a radar image capture at the point x, y of the BEV grid at a given time t. In some embodiments, the radar image can be initially generated in polar (or spherical) coordinates, and subsequently the mapping to grid (Cartesian) coordinates is performed as part of a gather transformation that associates each point of the radar point cloud with a specific location within the BEV grid. Additionally, the radar features can characterize the type of reflection, e.g., distinguish reflections from metallic objects (traffic signs, vehicles, etc.) from reflections from non-metallic objects (e.g., trees, concrete structures, etc.). The coordinates of various reflection points can be determined directly from the radar data (e.g., the distance and azimuth towards the signal reflection point). The camera network can similarly determine camera features F that characterize the visual appearance of the portion of the environment associated with the point x, y of the BEV grid at time t. C (x, y; t). Since the camera image lacks explicit distance (depth) information, the camera network can also perform a lift transform (either together with or after feature generation) that associates the various pixels of the camera image with the points x, y of the BEV grid, which are also associated with the radar return. The lift transform can be performed by estimating the most likely distance associated with a given pixel in the camera image (e.g., the distance to the object or a part of the object depicted by the pixel) or by evaluating the entire distribution of various such possible distances. Accordingly, the camera network can map the camera features to the same BEV grid to which the radar network maps the radar features.

[0021] In some embodiments, the camera features and radar features can then be aggregated into a joint feature {F R (x, y; t), F C(x, y; t)} → F(x, y; t), which can be processed by another model, also referred to herein as a BEV model. The BEV model can include a backbone network that processes one or more joint features, e.g., a stack (tensor) of features corresponding to multiple times t. In some embodiments, the backbone network can feed or provide intermediate outputs to multiple classifier heads that output various classes of traffic signs captured by camera and radar images. For example, a sign detection head can classify various BEV points x, y as sign points or non-sign points and can generate bounding boxes for identified signs in the original camera image. A sign type head can classify detected signs among multiple predefined types, e.g., speed limit signs, stop signs, yield signs, lane signs (e.g., lane merge signs, lane turn signs, etc.), information signs (e.g., highway exit signs), and / or any other type of traffic sign that can be specified during model training. A relevance head can classify detected signs as vehicle-related, e.g., speed limit signs, stop signs, etc., or as vehicle-unrelated, e.g., signs pointing to other types of vehicles (e.g., commercial trucks) or vehicles occupying different parts of the road (e.g., lanes). Various additional classification heads can be trained, e.g., a sign value head that classifies speed limit signs among multiple subtypes (e.g., 20 mph signs, 65 mph signs, etc.).

[0022] In some embodiments, the sign detection and classification model can be an end-to-end (E2E) model where various networks of the model (e.g., camera network, radar network, backbone, classifier heads, etc.) are trained together using suitable ground truth data, which can include actual traffic sign labels (and values, where appropriate), correct distances to signs, associations of signs with a particular lane of travel, etc. In one example, lidar sensor measurements can be used to measure ground truth depth (distance) data, and the correct sign type / value / association can be determined by a human developer.

[0023] The operation of the sign detection and classification model can be repeated periodically, e.g., every second, every few seconds, or every fraction of a second, to track the movement of signs relative to the vehicle. In an autonomous driving system (or a driver assistance system operating in an autonomous or semi-autonomous mode), the detected and classified signs can be passed to a planner module to chart and implement a driving path for the vehicle that is consistent with the traffic signs. In a driver assistance system operating in a driver-controlled mode, the detected and classified signs can be communicated to the driver, e.g., as part of a displayed (dashboard) warning, an acoustic warning, etc.

[0024] Advantages of the described embodiments include, but are not limited to, accurate, reliable, and fast detection and classification without deploying expensive lidar sensors, while using the complementary advantages of different sensing modalities. In particular, for more accurate monitoring of traffic signs in the driving environment, high-resolution but depth-uncertain camera images can be complemented by low-resolution but depth-cognizant radar data. The E2E architecture of the disclosed sign detection and classification model enables fast sign perception. In turn, the reduced latency in the sign detection task leads to improved responsiveness in driving path selection and enhanced safety in driving operations. In some embodiments, the disclosed technology can be used to identify the location and status of traffic lights as well as traffic sign detection and classification.

[0025] As used in this disclosure, a feature vector (embedding) should be understood as any suitable numerical representation of input data, e.g., understood as a vector (string) of any number M of components, which may have integer or floating-point values. A feature vector can be considered a point in an M-dimensional embedding space. The dimension M of the embedding space (defined as part of any relevant model architecture) can be less than the size of the input data (audio frames). During training, the model learns to associate similar sets of training audio frames with similar feature vectors represented by points that are closely situated in the embedding space, and further learns to associate dissimilar sets of training audio frames with points that are more distant from each other in that space. In some embodiments, a separate sound embedding (or an independent set of sound embeddings) can represent a given audio frame.

[0026] In those instances where the description of an embodiment relates to an autonomous vehicle, it should be understood that similar techniques can be used in various driver assistance systems that do not reach the level of a fully autonomous driving system. In some embodiments, the disclosed technology can be used to implement level 2 driver assistance systems such as steering, braking, accelerating, lane centering, adaptive cruise control, etc., as well as other driver supports. In some embodiments, the disclosed technology can be used in level 3 driving assistance systems capable of autonomous driving under limited (e.g., highway) conditions. In such systems, fast and accurate detection and tracking of objects can be used to notify the driver of approaching vehicles and / or other objects, with the driver making the final driving decision (e.g., in a level 2 system), or making certain driving decisions (e.g., in a level 3 system), such as reducing speed, changing lanes, etc., without requesting feedback from the driver.

[0027] Figure 1FIG. is a diagram showing components of an example vehicle 100 capable of deploying effective sign detection and classification in a driving environment according to some specific implementations of the present disclosure. An autonomous vehicle can include a motor vehicle (car, truck, bus, motorcycle, all-terrain vehicle, recreational vehicle, any specialized agricultural or construction vehicle, etc.), an aircraft (airplane, helicopter, drone, etc.), a naval vehicle (ship, boat, yacht, submarine, etc.), or any other self-propelled vehicle capable of operating in an autonomous mode (without manual input or with reduced manual input) (e.g., a robot, a factory or warehouse robotic vehicle, a sidewalk delivery robotic vehicle, etc.).

[0028] The driving environment 101 can include any object (active or inactive) located outside the vehicle 100, such as roads, buildings, trees, shrubs, sidewalks, bridges, mountains, other vehicles, pedestrians, etc. The driving environment 101 can be urban, suburban, rural, etc. In some implementations, the driving environment 101 can be an off-road environment (e.g., a farm or other agricultural land). In some implementations, the driving environment can be an indoor environment, such as the environment of an industrial factory, a shipping warehouse, a hazardous area of a building, etc. In some implementations, the driving environment 101 can be substantially flat, with various objects moving parallel to the surface (e.g., parallel to the ground). In other implementations, the driving environment can be three-dimensional and can include objects capable of moving along all three directions (e.g., balloons, leaves, etc.). Hereinafter, the term "driving environment" should be understood to include all environments in which autonomous movement of a self-propelled vehicle can occur. For example, the "driving environment" can include any possible flight environment of an airplane or the marine environment of a naval ship. The objects in the driving environment 101 can be located at any distance from the vehicle 100, from a short distance of a few feet (or less) to several miles (or more).

[0029] As described herein, in a semi-autonomous or partially autonomous driving mode, even if the vehicle is assisted with one or more driving operations (e.g., steering, braking, and / or accelerating to perform lane centering, adaptive cruise control, advanced driver assistance systems (ADAS), or emergency braking), the human driver is expected to be situationally aware of the vehicle's surroundings and supervise the assisted driving operations. Here, even if the vehicle can perform all driving tasks in certain situations, the human driver is expected to be responsible for taking control as needed.

[0030] Although, for simplicity and conciseness, various systems and methods may be described below in the context of autonomous vehicles, similar techniques may be used in a variety of driver assistance systems that do not reach the level of a fully autonomous driving system. In the United States, the Society of Automotive Engineers (SAE) has defined different levels of automated driving operation to indicate the extent or degree to which a vehicle controls the driving, although different organizations in the United States or other countries may classify the levels differently. More specifically, the disclosed systems and methods may be used in SAE Level 2 (L2) driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, etc. as well as other driver supports. The disclosed systems and methods may be used in SAE Level 3 (L3) driving assistance systems capable of autonomous driving under restricted (e.g., highway) conditions. Similarly, the disclosed systems and methods may be used in vehicles using SAE Level 4 (L4) automated driving systems that operate autonomously in most conventional driving situations and only require occasional attention from a human operator. In all such driver assistance systems, accurate lane estimation may be automatically performed without driver input or control (e.g., when the vehicle is in motion), and result in an improvement in the reliability of vehicle positioning and navigation as well as the overall safety of autonomous, semi-autonomous, and other driver assistance systems. As previously mentioned, in addition to the way the SAE classifies the levels of automated driving operation, other organizations in the United States or other countries may classify the levels of automated driving operation differently. Without limitation, the systems and methods disclosed herein may be used in driver assistance systems defined by the levels of automated driving operation of these other organizations.

[0031] Exemplary vehicle 100 may include a sensing system 110. The sensing system 110 may include various electromagnetic (e.g., optical) and non-electromagnetic (e.g., acoustic) sensing subsystems and / or devices. The sensing system 110 may include radar(s) 112, which may be any system that uses radio or microwave frequency signals to sense objects within the driving environment 101 of the vehicle 100. The radar 112 may be configured to sense the spatial location of the objects (including their spatial dimensions) and the speed of the objects (e.g., using Doppler shift techniques). Hereinafter, "speed" refers to how fast the object is moving (the speed of the object) and the direction of the object's motion. In some embodiments, the sensing system 110 may include lidar 114, which may be a laser-based unit capable of determining the distance to an object in the driving environment 101 and the speed of the object. Each of the radar 112 and lidar 114 may include a coherent sensor, such as a frequency-modulated continuous-wave (FMCW) lidar or radar sensor. For example, the radar 112 may use heterodyne detection for speed determination. In some embodiments, the functions of ToF and coherent radar are combined into a radar unit capable of simultaneously determining the distance to a reflecting object and the radial speed of the reflecting object. Such a unit may be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection) or simultaneously in both modes. In some embodiments, multiple radars 112 or lidars 114 may be mounted on the vehicle 100.

[0032] The lidar 114 may include one or more light sources that generate and emit signals and one or more detectors that detect the signals reflected back from the objects. In some embodiments, the lidar 114 is capable of performing a 360-degree scan in the horizontal direction. In some embodiments, the lidar 114 is capable of performing a spatial scan along both the horizontal and vertical directions. In some embodiments, the field of view may reach 90 degrees in the vertical direction (e.g., scanning at least a portion of the area above the horizon with radar signals). In some embodiments, the field of view may be a complete sphere (composed of two hemispheres).

[0033] The sensing system 110 may further include one or more cameras 118 to capture images of the driving environment 101. The images may be two-dimensional projections of the driving environment 101 (or a portion of the driving environment 101) on the projection surface (flat or non-flat) of the camera. Some of the cameras 118 of the sensing system 110 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101. The sensing system 110 may further include one or more infrared (IR) sensors 119. In some embodiments, the sensing system 110 may further include one or more ultrasonic sensors 116, which may be ultrasonic sonars.

[0034] The sensing data obtained by the sensing system 110 can be processed by the data processing system 120 of the vehicle 100. For example, the data processing system 120 may include a perception and planning system 130. The perception and planning system 130 may be configured to detect and track objects in the driving environment 101 and identify the detected objects. For example, the perception and planning system 130 can analyze the images captured by the camera 118 and can be capable of detecting traffic light signals, road signs, road layouts (e.g., boundaries of lanes, topologies of intersections, designations of parking locations, etc.), the presence of obstacles, and the like. The perception and planning system 130 can also receive radar sensing data (Doppler data and ToF data) and determine the distances to various objects in the environment 101 and the speeds of these objects (radial speed, and in some embodiments, lateral speed, as described below). In some embodiments, the perception and planning system 130 can use the radar data in combination with the data captured by the camera 118, as described in more detail below.

[0035] The perception and planning system 130 monitors how the driving environment 101 evolves over time, e.g., by tracking the positions and speeds of active objects (e.g., relative to the Earth and / or the AV) and predicting how various objects will move in the future within a certain time horizon (e.g., 1 - 10 seconds or longer). The perception and planning system 130 may include a sign detection and classification model (SDCM) 132 that performs end-to-end detection and reading of traffic signs present in the environment 101. The SDCM 132 may include one or more trainable MLMs capable of processing data of multiple modalities (e.g., radar data and camera data).

[0036] The perception and planning system 130 can also receive information from the positioning subsystem 122, which may include a GPS transceiver and / or an inertial measurement unit (IMU) ( Figure 1 (not shown in the figure) configured to obtain information about the position of the AV relative to the Earth and its surrounding environment. The positioning subsystem 122 can use the positioning data (e.g., GPS and IMU data) in combination with the sensing data to help accurately determine the position of the vehicle 100 relative to fixed objects in the driving environment 101 (e.g., roads, lane boundaries, intersections, sidewalks, crosswalks, road signs, curbs, surrounding buildings, etc.), and the positions of the fixed objects can be provided by the map information 124. In some embodiments, the data processing system 120 can receive non-electromagnetic data, such as audio data (e.g., ultrasonic sensor data or data from one or more microphones detecting emergency vehicle sirens), temperature sensor data, humidity sensor data, pressure sensor data, meteorological data (e.g., wind speed and direction, precipitation data), etc.

[0037] Data generated by the perception and planning system 130, the localization subsystem 122, and / or other systems and components of the data processing system 120 can be used by an autonomous driving system such as the vehicle control system (VCS) 140. The VCS 140 can include one or more algorithms that control how the vehicle 100 behaves in various driving situations and environments. For example, the VCS 140 can include a navigation system for determining a global driving route to a destination location. The VCS 140 can also include a driving path selection system for selecting a specific path through the immediate driving environment, which can include selecting a driving lane, navigating around traffic congestion, selecting a place to make a U-turn, selecting a trajectory for a parking maneuver, and so on. The VCS 140 can also include an obstacle avoidance system for safely avoiding various obstacles (rocks, stalled vehicles, pedestrians walking in a hurry, etc.) within the driving environment of the AV. The obstacle avoidance system can be configured to evaluate the size of the obstacle and the trajectory of the obstacle (if the obstacle is moving), and select an optimal driving strategy (e.g., braking, steering, accelerating, etc.) to avoid the obstacle.

[0038] The algorithms and modules of the VCS 140 can generate instructions for various systems and components of the vehicle, such as the powertrain, brakes, and steering 150 devices, vehicle electronics 160, signaling 170, and Figure 1 other systems and components not explicitly shown in the figure. The powertrain, brakes, and steering devices 150 can include an engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanisms, and other systems. The vehicle electronics 160 can include an on-board computer, engine management, ignition, communication systems, automotive, telematics, in-vehicle entertainment systems, and other systems and components. Signaling 170 can include high and low beam headlights, parking lights, turn signals, and reverse lights, horns and sirens, interior lighting systems, dashboard notification systems, passenger notification systems, radio, and wireless network transmission systems, etc. Some instructions output by the VCS 140 can be delivered directly to the powertrain, brakes, and steering devices 150 (or signaling 170), while other instructions output by the VCS 140 are first delivered to the vehicle electronics 160, which generates commands to the powertrain, brakes, and steering devices 150 and / or signaling 170.

[0039] In one example, the VCS 140 can determine to avoid an obstacle identified by the data processing system 120 by decelerating the vehicle until a safe speed is reached and then steering the vehicle around the obstacle. The VCS 140 can output instructions (either directly or via the vehicle electronics 160) to the powertrain, brakes, and steering 150 to: (1) reduce the fuel flow to the engine by modifying the throttle setting to decrease the engine rpm; (2) downshift the driveline to a lower gear via the automatic transmission; (3) engage the braking unit to decrease the speed of the vehicle (while acting in concert with the engine and transmission) until the safe speed is reached; and (4) perform a steering maneuver using the power steering mechanism until the obstacle is safely bypassed. Subsequently, the VCS 140 can output instructions to the powertrain, brakes, and steering 150 to restore the vehicle's previous speed setting.

[0040] "Autonomous vehicle" can include motor vehicles (cars, trucks, buses, motorcycles, all-terrain vehicles, recreational vehicles, any specialized agricultural or construction vehicles, etc.), aircraft (planes, helicopters, drones, etc.), naval vehicles (ships, boats, yachts, submarines, etc.), robotic vehicles (e.g., factory, warehouse, sidewalk delivery robots, etc.), or any other self-propelled vehicle capable of operating in an autonomous mode (without human input or with reduced human input). "Object" can include any entity, item, device, body, or thing (active or inactive) located outside the autonomous vehicle, such as roads, buildings, trees, shrubs, sidewalks, bridges, mountains, other vehicles, docks, embankments, landing runways, animals, birds, or other things.

[0041] Figure 2 is a diagram showing an example architecture 200 that can be used to train and deploy a sign detection and classification model capable of detecting and reading traffic signs in a driving environment according to some embodiments of the present disclosure. The input into the SDCM 132 can include that from the sensing system 110 (e.g., by radar 112, camera 118, and / or other sensors, refer to Figure 1)The acquired data. The acquired data can be provided via the sensing data acquisition module 210, which can decode, preprocess (e.g., denoise, upsample or downsample, etc.), and reformat the data into a format accessible by the SDCM 132. In one exemplary implementation, the sensing data acquisition module 210 can acquire a sequence of camera images 202, e.g., a two-dimensional projection of the driving environment (or a portion thereof) on an array of sensing detectors (e.g., charge-coupled device or CCD detector, complementary metal-oxide semiconductor or CMOS detector, etc.). Each camera image can have pixels of various intensities of one color (for black-and-white images) or multiple colors (for color images). The camera image can be a panoramic image or an image depicting a specific portion of the driving environment. The camera image can include a plurality of pixels. The number of pixels can depend on the resolution of the image. Each pixel can be characterized by one or more intensity values. A black-and-white pixel can be characterized by, for example, one intensity value representing the brightness of the pixel, where a value of 1 corresponds to a white pixel and a value of 0 corresponds to a black pixel (or vice versa). The intensity values can assume continuous (or discrete) values between 0 and 1 (or between any other chosen limits, e.g., 0 and 255). Similarly, a color pixel can be represented by more than one intensity value, such as three intensity values (e.g., if using an RGB color coding scheme) or four intensity values (e.g., if using a CMYK color coding scheme). The camera image can be preprocessed, e.g., downscaled (with multiple pixel intensity values combined into a single pixel value), upsampled, filtered, denoised, etc. The camera image can be in any suitable digital format (JPEG, TIFF, GIG, BMP, CGM, SVG, etc.).

[0042] The sensing data acquisition module 210 may further obtain a radar image 204, which may include a set of return points (point cloud) corresponding to radar beam reflections from various objects in the driving environment. Each return point may be understood as a data unit (pixel) including the coordinates of the reflecting surface, radial velocity data, intensity data, etc. For example, the sensing data acquisition module 210 may provide a radar image 204 including a radar intensity map I(R,θ,φ), where R,θ,φ are a set of spherical coordinates. In some embodiments, Cartesian coordinates, elliptical coordinates, parabolic coordinates, or any other suitable coordinates may be used alternatively. The radar intensity map identifies the intensity of the radar reflections of the individual points in the radar field of view. The coordinates of the object reflecting the radar signal may be based on direction data (e.g., the polarization angle θ and azimuth angle φ in the direction of the lidar parameters) and distance data (e.g., the radial distance R determined based on the flight time of the radar signal). The radar image 204 may also include velocity data of various reflecting objects identified based on the Doppler shift of the detected reflected signals. In some embodiments, the sensing data acquisition module 210 may similarly obtain a lidar image.

[0043] The camera image 202 and / or the radar image 204 may be a large image of the entire driving environment or an image of a smaller part of the driving environment (e.g., the camera image acquired by the forward camera of the sensing system 110). In some embodiments, the sensing data acquisition module 210 may crop the camera image 202 and / or the radar image 204 corresponding to a certain segment around the direction of motion of the vehicle. For example, in an exemplary non-limiting embodiment, since relevant traffic signs are typically located around the driving direction of the vehicle, the sensing data acquisition module 210 may crop the camera image 202 and the radar image 204 into a forward segment that is 200 - 250m long and 20 - 40m wide. The size of the segment may depend on the speed of the vehicle and the type of the driving environment, and may be different for a highway driving environment and a city driving environment. The camera image 202 is processed by the camera network 220, and the radar image 204 is processed by the radar network 222. The camera network 220 generates camera features ( Figure 2 (not shown in the figure), and the radar network 222 generates radar features ( Figure 2 (not shown in the figure). The camera features and the radar features may be associated with a two-dimensional bird's-eye view (BEV) and may be generated using a suitable lifting transformation from the perspective.

[0044] Camera features and radar features can be combined and processed by a BEV model including a BEV backbone 224 and one or more flag classification heads 226. The flag classification head 226 can identify traffic signs, determine the bounding box of the identified sign, the degree of relevance of the identified sign, the value of the sign (or other sign-specific content), etc. The various networks of the SDCM 132 can include convolutional neural networks, recurrent neural networks (RNNs) with one or more hidden layers, fully connected neural networks, long short-term memory neural networks, transformers, Boltzmann machines, etc.

[0045] The output of the SDCM 132 can be provided to a tracker / planner 230, which can be Figure 1 part of the perception and planning system 130. The tracker / planner 230 can track the movement of traffic signs, vehicles, and other objects (e.g., relative to the vehicle). In some embodiments, a suitable motion filter (e.g., a Kalman filter) can be used to track the behavior of signs, vehicles, and other objects identified by the SDCM 132. The Kalman filter calculates the most likely geo-motion data given the measurements obtained (e.g., the output of the SDCM 132), predictions based on a physical model of the object's motion, and some statistical assumptions about the measurement error (e.g., the covariance matrix of the error). The tracker / planner 230 can also select a path for the vehicle consistent with the identified traffic signs and provide instructions to the vehicle control system 140 to implement the selected driving path.

[0046] The training of the SDCM 132 and / or other MLMs can be executed by a training engine 242 hosted by a training server 240, which can be an external server deploying one or more processing devices (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), etc.). The training engine 242 can access a data store 250 storing various training data for training the SDCM 132. In some embodiments, the training data can include camera images 252 acquired by an in-vehicle camera during an actual driving task, and can also include radar images 254 associated with the camera images 252, e.g., radar images of substantially the same area of the corresponding driving environment acquired substantially simultaneously with the camera images. The training data stored by the data store 250 can also include ground truth labels 256, which can include the correct traffic sign classification of the images depicted in the camera images 252 and the radar images 254. For example, the ground truth sign classification can indicate the presence (or absence) of a sign, the type of the sign (e.g., speed limit sign, stop sign, yield sign, prohibition sign, information sign, etc.), the semantic value of the sign (and / or any other variable information contained therein), the sign-affected portion of the road (e.g., lane) in a specific area of the environment depicted by the camera images 252 and the radar images 254. The ground truth sign classification can also include the correct distance to the sign. In some embodiments, a high-resolution lidar sensor can be used to measure such ground truth distances. In some embodiments, the ground truth distance can be determined by a developer manually mapping the radar returns in the radar image 254 to the corresponding objects in the camera image. In some embodiments, the ground truth distance can be determined (with or without input from a human developer) by matching the depiction of a traffic sign in the camera image 254 with the known (e.g., from map information and the recorded vehicle's geographic movement data) location of the traffic sign.

[0047] As Figure 2As shown, training data including training inputs 244 and corresponding target outputs 246 (correct matches for the respective training inputs) can be used to train the SDCM 132. During training, the training engine 242 can retrieve the training data from the data store 250, prepare one or more training inputs 244 and one or more target outputs 246 (ground truth), and use the prepared inputs and outputs to train one or more models, including but not limited to the SDCM 132. The training data can also include mapping data 248 that maps the training inputs 244 to the target outputs 246. During the training of the SDCM 132, the training engine 242 can cause the SDCM 132 to learn patterns in the training data captured by the training input / target output pairs. To evaluate the difference between the training output and the target output 246, the training engine 242 can use various suitable loss functions, such as the mean squared error loss function (e.g., evaluating the deviation from the continuous ground truth, e.g., the distance to a sign), the binary cross-entropy loss function (e.g., evaluating the deviation from binary classification), and / or any other suitable loss function. In some embodiments, the SDCM 132 can be trained by the training engine 242 and then downloaded to the vehicle's perception and planning system 130.

[0048] During the training of the SDCM 132, the training engine 242 can change the parameters (e.g., weights and biases) of the various networks of the SDCM 132 until the model successfully learns to accurately detect traffic signs and read the semantic content of the detected signs. In some embodiments, more than one SDCM 132 can be trained for use under different conditions and for different driving environments. For example, separate SDCM 132s can be trained for street driving and for highway driving. The different trained SDCM 132s can have different architectures (e.g., different numbers of neuron layers and / or different neural connection topologies), different settings (e.g., types and parameters of activation functions, etc.), and can be trained using different sets of hyperparameters.

[0049] According to various embodiments of the present disclosure, the data store 250 can be a persistent memory capable of storing radar images, camera images, and data structures configured to facilitate accurate and rapid identification and verification of sign detection. The data store 250 can be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, magnetic tapes, or hard disk drives, network-attached storage (NAS), storage area network (SAN), etc. Although depicted as separate from the training server 240, in some embodiments, the data store 250 can be part of the training server 240. In some embodiments, the data store 250 can be a network-attached file server, while in other embodiments, the data store 250 can be some other type of persistent storage, such as an object-oriented database, a relational database, etc., which can be hosted by one or more different machines accessible to the server machine or the training server 240 via a network ( Figure 2 not shown).

[0050] Figures 3A to 3B An example operation of a sign detection and classification model capable of effectively detecting and reading traffic signs in a driving environment is shown according to some embodiments of the present disclosure. Figure 3A A first part 300 of the SDCM operation is shown, which includes respective processing of camera images and radar images. As Figure 3A shown, the input data 301 into the SDCM can include the camera image 202 and the radar image 204. Although for the sake of specificity, Figure 3A and Figure 3B the description refers to radar images, in some embodiments, the radar image 204 can be replaced with a lidar image, and the SDCM can be trained to perform sign detection / classification using the lidar image instead of the radar image. In some embodiments, in addition to the radar image, lidar images can also be processed, for example, using a separate lidar network to generate lidar features, Figure 3A not shown).

[0051] When the respective images are captured, each camera image 202 (and similarly, the radar image 204) can be associated with specific times t1, t2, t3. The acquisition of the camera image 202 and the radar image 204 can be synchronized such that the images of the two modalities substantially depict the driving environment at the same time. The camera image 202 can be processed by the camera network 220, and the radar image 204 can be processed by the radar network 222. In some embodiments, each network can separately process the images associated with different times t j associated with.

[0052] The camera network 220 and, similarly, the radar network 222 can have any suitable architecture. In one example, the camera network 220 and / or the radar network 222 can be a deep convolutional neural network, e.g., having a U-net architecture that includes an encoder stage and a decoder stage. Each stage can have multiple convolutional neuron layers and one or more fully connected layers. The convolutional encoder can include any number of filters (kernels) that broaden the perception field and identify features of an image by aggregating relevant information captured by individual units (pixels) of the image and encoding that information via features arranged in a feature map. A sequence of convolutional layers and pooling (e.g., average pooling or max pooling) layers can be used to produce such a feature map. The convolutional layers apply (typically multiple, e.g., dozens, hundreds, or more) filters - matrices of finite size with learned weights - that scan the image to look for certain features in the image. Different kernels can look for different features, e.g., the boundaries of traffic signs, the shape of the signs, the color patterns of the signs, the presence of text in the signs, etc. The kernel can move across the image in steps (strides) smaller than the size of the kernel (e.g., a 5×5 pixel kernel can shift 1, 2, or 3 pixels during each stride), thereby forming a signal for a neural activation function. Then, the subsampling (pooling) operation reduces the dimension of the generated feature map according to the fundamental premise of the convolutional neural network architecture that information about the presence of the target feature is generally more important than accurate knowledge of the feature coordinates. As a result of such multi-layer convolutional and pooling processing, the intermediate representation of the image grows along the feature (channel) dimension but shrinks along the width-height dimension of the image. This reduction accelerates subsequent computations while ensuring the ability of the neural network to process input images of different scales.

[0053] The decoder portion of the camera network 220 and / or the radar network 222 upsamples the feature map generated by the convolutional encoder to gradually increase the resolution while reducing the feature / channel dimension (which can be performed using another set of learned deconvolution kernels), e.g., back to the original (or slightly reduced) dimension of the input image, where the final layer generates the output features. For example, the camera network 220 can generate a camera feature vector F C (x,y;t) 302 that characterizes the visual appearance of the portion of the environment associated with the points x,y of the BEV grid at time t (as captured by the camera image 202). Similarly, the radar network 222 can generate a radar feature vector F R (x,y;t) 304 that characterizes the presence or absence of reflecting objects in the same portion of the environment associated with the same time t (as captured by the radar image 204). However, for illustrative purposes, in Figure 3Adepicts a single camera feature vector F C (x,y;t) 302 (and similarly, a single radar feature vector F R (x,y;t) 304), and each feature vector can be combined into a camera feature tensor. FT C (t) = {F C (x,y;t)} 312. For those locations (locals) in the BEV grid where traffic signs are not detected and marked, the corresponding camera feature vector F C (x,y;t) 302 can have a zero value (or a value close to zero). The feature tensor FT C (t) 312 can have dimensions X×Y×C, where X and Y are the dimensions of the BEV grid 305, and C is the context dimension, which can (before training) be set as part of the camera network 220 architecture. Similarly, the radar feature vector 304 is combined into the radar feature tensor FT R (t) = {F R (x,y;t)} 314. The radar feature tensor FT R (t) 314 can have the same BEV dimensions X and Y, and a context dimension C C different from the context dimension C of the camera feature tensor FT R . For example, given that the camera image 202 captures more diverse types of visual context compared to the radar image 204, the context dimension C of the camera feature vector / tensor R can have a higher dimension than the dimension C of the radar feature vector / tensor R .

[0054] Although in the above example, a convolutional encoder / transposed decoder architecture is used for illustration, the camera network 220 and / or the radar network 222 can have any other suitable architecture. For example, the encoder part of the network can include a recurrent neural network, a long short-term memory (LSTM) neural network, a fully connected network, and / or some combination of such networks. In some embodiments, the camera network 220 and / or the radar network 222 can have a transformer-based architecture, where the encoder part of the network includes one or more self-attention blocks, and the decoder part of the network includes one or more cross-attention blocks (in addition to the self-attention blocks). In some embodiments, the camera network 220 and / or the radar network 222 can include only the encoder part, and the decoder part can be implemented as part of the BEV backbone 224.

[0055] To generate the camera feature vector 302 (and the camera feature tensor 312), the camera network 220 is capable of performing a lifting transformation 310 to compensate for the lack of explicit distance (depth) information in the camera image 202. In some embodiments, the lifting transformation 310 may be performed in two stages. During the first stage, the camera network 220 generates a feature vector f(c) for the pixels w, h of the camera (perspective), w,h , where c enumerates various context components (e.g., c ∈ [1, C]) and further supplements the depth information for the feature vector f(c) w,h , which may be provided by a separate output of the camera network 220. For example, the depth information may include P(d), the probability that a given pixel w, h depicts an object located at a distance d from the camera w,h . Then, the lifting transformation 310 may compute the direct product of each feature vector f(c) w,h with the corresponding depth distribution P(d) w,h to generate a depth-enhanced feature vector f(c, d) for the pixel w, h w,h = f(c)

[0056]

[0057] Then, the depth-enhanced feature vectors f(c, d) computed for each pixel w,h may be combined (still in perspective) into a joint feature tensor for the entire image 202: {f(c, d) w,h} → ft(c, d, w, h). The depth-enhanced feature tensor ft(c, d, w, h) has dimensions C × D × W × H, where W × H are the dimensions of the camera image 202 (in pixels), and D is the dimension of the depth granularity. For example, the distance d may be discretized in D intervals, Δd1, Δd2... Δd D . The intervals Δd i do not need to have equal sizes and may increase with distance, e.g., Δd1 < Δd D . In some embodiments, the last interval Δd D may extend from a certain distance (e.g., 100m, 200m, etc.) to infinite distance.

[0058] The second stage of the lifting transformation 310 may include a two-dimensional (2D) mapping that maps the depth-enhanced feature tensor to a feature tensor in the BEV grid 305, e.g., in Cartesian coordinates ft(c, d, w, h) → FT CIn (c, x, y) or in any other planar coordinate set, such as polar coordinates r and θ in the plane of the ground. More specifically, perspective coordinates d, w, h can be transformed into 3D Cartesian coordinates d, w, h → x, y, z (or 3D cylindrical coordinates d, w, h → r, θ, z), where z is the vertical coordinate (in the direction perpendicular to the ground). The transformation d, w, h → x, y, z can be a projective transformation, parameterized by the focal length of the camera, the direction of the optical axis of the camera, and other similar parameters. In instances where the camera image 202 is acquired by multiple cameras (or a camera with a rotating optical axis), the transformation d, w, h → x, y, z can include multiple projective transformations, e.g., separate transformations for pixels w, h of different cameras (or the same camera pointing in different directions).

[0059] 2D mapping can be used to project the feature tensor expressed in the new coordinates, ft(c, d, w, h) → ft(c, x, y, z), and sum (or average, weighted average, or otherwise aggregate) over different heights z to obtain the feature tensor FT C (c, x, y)312, e.g., FT C (c, x, y) = ∑ i ft(c, x, y, z i )). In some embodiments, the sum over the coordinate z i can be performed with different weights w i assigned to different coordinates z i : FT C (c, x, y) = ∑ i w i ·ft(c, x, y, z i ). For example, larger weights w i are assigned to the pixels of the image object within certain elevations from the ground (e.g., up to a few meters), while lower weights are assigned to other elevations (e.g., to eliminate spurious objects that do not obstruct vehicle movement, such as tree branches, wires, etc.).

[0060] Similarly, the radar features output by the radar network 222 can be mapped to the same BEV grid 305. In instances where the radar network outputs, the distances to various reflection pixels in the radar image 204 can be accurately known (as part of the radar data), such that the distribution P(d) can be a unit of a specific interval of distance Δd j while other intervals are zero. The lifting transformation 310 and the use of the BEV grid 305 eliminate the distortions associated with the perspective views of the camera and radar sensors.

[0061] Then, the camera features and radar features can be aggregated (e.g., cascaded) to obtain joint features. For example, feature vector aggregation can be performed for each BEV grid position x, y to obtain a joint feature vector 320: FT C (x,y,t),F R (x,y,t)] → F(x,y,t). The set of joint feature vectors for various BEV grid positions represents a joint feature tensor FT(t) = {F(x,y;t)}330. Equivalently, the joint feature tensor FT(t)330 represents the combination of the camera feature tensor 312 and the radar feature tensor 314 FT(t) = [FT C (t),FT R (t)].

[0062] In some embodiments, the joint feature tensors associated with multiple timestamps can be aggregated into a feature stack 340, e.g., {FT(t1), FT(t2), …… FT(t M )}. (For the sake of brevity, an example case with M = 3 is shown in Figure 3A ) The times t1, t2, …… t M can be selected by applying a sliding window to the images of the input data 301. For example, during the next round of input data 301 processing by the SDCM132, a feature stack 340 {FT(t 1+S , t 2+S , …… t M+S )} associated with the images obtained at times t 1+S , t 2+S ), …… FT(t M+S )} can be generated, where an appropriately selected stride S, e.g., S = 1, 2, …, etc. In some embodiments, the stride S can be set based on the time of the E2E processing of the input data 301, e.g., to avoid idling of the SDCM 132 due to new input data 301 being provided before the previous input data 301 has been transferred to the feature stack 340 and the feature stack 340 has been processed, while also preventing clogging of the SDCM 132. For example, if the typical time for E2E processing of the SDCM 132 is τ, the stride can be set such that t j+S - t j ≈ τ.

[0063] The generated feature stack 340 can be processed by the BEV backbone 224. In some embodiments, the BEV backbone 224 can include both an encoder and a decoder. In some embodiments, the BEV backbone 224 can include a decoder, while the encoder is implemented as part of the camera network 220 and / or the radar network 222.

[0064] Figure 3B Shows the second part 350 of the SDCM operation, which includes the processing of combined camera and radar features. The BEV backbone 224 feeds the intermediate output to a plurality of signature classification heads 226-n, and the signature classification heads 226-n output various categories of traffic signs captured by the camera and radar images. Each classification head 226-n can output different types of information about the traffic signs captured in the input data 301. For example, the signature detection head 226-1 can identify the area of the external environment where the sign may be located. In some embodiments, the signature detection head 226-1 can classify various BEV points x, y as signature points or non-signature points, for example, using a final binary (sigmoid) classifier, which outputs a floating-point probability w sign (x,y) and w non-sign (x,y) = 1 - w sign (x,y), and then generate a binary prediction for the points x, y based on whether w sign > 0.5 (or any other empirically set threshold, 0.5, 0.75, etc.). In some embodiments, the signature detection head 226-1 can use the inverse lift transformation 306 to map the BEV points x, y back to the perspective coordinates of the camera pixels w, h in order to generate a bounding box for the signs recognized within the camera image.

[0065] The signature type classification head 226-2 can classify the detected traffic signs among a plurality of predefined types, for example, speed limit signs, stop signs, yield signs, lane direction signs (e.g., lane merge signs, lane turn signs, etc.), information signs (e.g., highway exit signs), and / or any other type of traffic sign, such as can be defined during the training of the SDCM 132. For example, the final neuron layer (e.g., softmax layer) of the classification head 226-2 can output the probabilities w1, w2,... w n that the sign belongs to any one of the n defined types (categories). Then, the type of the sign with the highest probability w high can be output as the predicted sign type. The value of the corresponding probability can be used as the confidence level. For example, in an illustrative non-limiting example, where w high ≥ 0.9 corresponds to a high confidence level, 0.7 ≤ w high < 0.9, and w high < 0.7 corresponds to a low confidence level.

[0066] The sign value classification header 226-3 can classify detected traffic signs among multiple sign values (or any other subtype), if such a value is defined for the type identified by the sign type classification header 226-2. For example, a "45 mph" value for a speed limit sign, a "100 m" value for a lane end sign, etc. Selecting from a predetermined number of sign values can be performed similarly to selecting from multiple sign types, e.g., as disclosed in connection with the sign type classification header 226-2.

[0067] The sign relevance classification header 226-4 can classify a detected sign as relevant or irrelevant to the vehicle performing the detection. For example, if the detected sign pertains to other types of vehicles (e.g., commercial trucks) or vehicles occupying different lanes. A binary neuron classifier can be used to obtain the output of the sign relevance classification header 226-4. Figure 3B The sign classification headers 226-n shown are intended to be examples, as various other classification headers can be defined and trained based on the specific driving environment of the vehicle.

[0068] In some embodiments, the output of the sign classification header 226-n can undergo filtering 360 to eliminate duplicate signs, e.g., using non-maximum suppression (NMS), clustering, and / or other techniques. For example, NMS can be used to select the most likely bounding box in instances of multiple (overlapping or non-overlapping) bounding boxes enclosing closely located spatial regions. NMS can include iteratively comparing the probabilities of various bounding boxes and discarding one or more lower-probability bounding boxes at each iteration until the highest-probability bounding box is identified.

[0069] Filtering 360 can generate the finally detected sign 370, including the position and semantic information in the sign (e.g., sign type, sign value, relevance, etc.). The detected sign 370 can be provided to the tracker / planner module 380, which can track the motion of the detected sign relative to the vehicle over time using, for example, a suitable motion tracker such as a Kalman filter. The tracker / planner module 380 can also make driving decisions in view of the detected sign 370. For example, in an autonomous driving system (or a driver assistance system operating in an autonomous or semi-autonomous mode), the tracker / planner module 230 can identify and implement a driving path for the vehicle consistent with the placement and semantic information of the detected sign 370. In a driver assistance system, the tracker / planner module 230 can provide a representation of the detected sign to the driver, e.g., via a dashboard display.

[0070] In one exemplary embodiment, as Figure 3BAs indicated by the blowout part, the detected signs 370 and the environmental state 380 can be used as inputs to the motion planning model 382 to generate one or more predicted trajectories 384 consistent with the traffic signs. The state of the environment 380 can include the position and motion state of the vehicle on which the tracker / planner 230 is deployed (e.g., speed, direction, acceleration / braking, degree of steering, etc.) and the position and motion state of various objects in the environment (e.g., other vehicles, pedestrians, construction equipment, etc.). In some specific implementations, the motion planning model 382 can include a transformer-based neural network that is trained to output the predicted trajectory 384 of the vehicle in view of the traffic signs located within the visible portion of the environment.

[0071] In some embodiments, the predicted trajectory 384 can be represented numerically, e.g., by determining n motion tokens of the predicted motion state of the vehicle at future times t1, t2, … t n In some embodiments, the motion tokens can specify transitions between discrete motion states. For example, a motion token can characterize the acceleration at time t j while the state of the vehicle can include the position and speed of the vehicle, Then, in an exemplary non-limiting embodiment, the state S(t ) of the vehicle at time t j+1 can be obtained by updating the state S(t j+1 ) at time t j using the corresponding motion token j For example, since the coordinates and speed in the vehicle state and the acceleration in the token can be two-dimensional (or three-dimensional) vectors, e.g., separate components along two (three) spatial dimensions. Then, the tracker / planner 230 can select a target trajectory from the predicted trajectories 384 based on one or more objective metrics (e.g., minimizing travel time, minimizing the number of stops, maximizing fuel economy, maintaining at least a minimum distance from the vehicle to other objects, etc. or any combination thereof) for implementation as the driving path of the vehicle. The operation of the motion planning model 382 can be repeated periodically after a certain time interval (e.g., 0.5 seconds, 0.3 seconds, etc.), outputting a new set of predicted trajectories 384 and selecting a new target trajectory.

[0072] ​The sign detection and classification model can be an end-to-end (E2E) model, where various networks of the model (e.g., camera network 220, radar network 222, BEV backbone 224, and sign classification head 226) are trained together using suitable ground truth data, which can include actual traffic sign labels (and values, where appropriate), correct distances to the signs, associations of the signs with specific driving lanes, etc. In one example, lidar sensors can be used to measure ground truth depth (distance) data, and the correct sign type / value / association can be determined by human developers.

[0073] In some embodiments, some of the SDCM networks can be trained in stages, where the camera network 220, BEV backbone 224, and sign classification head 226 are first trained without input from the radar network 222. For example, the corresponding neurons in the input layer of the BEV backbone 224 receive null input. A suitable loss function can be used to perform pre-training of the camera network 220, which, for each training camera image, evaluates the distribution P(d) w,h of the center and the ground truth distance d True (w,h) between the differences. Then, this difference can be backpropagated through the respective neuron layers of the camera network 220 (BEV backbone 224), where the camera network 220 learns to correctly predict the probability of the depth of pixels in the camera depth with target accuracy. Further training can include using the output of the radar network 222 (radar feature tensor 314) as the input to the BEV backbone 224. Such multi-stage training can teach the SDCM to more effectively utilize camera images without over-relying on radar depth data.

[0074] Figure 4 is a schematic diagram of an example driving environment 400 of a vehicle 402 deploying an SDCM 132 to detect and read traffic signs according to some embodiments of the present disclosure. As Figure 4 shown, the vehicle 402 is traveling in the left lane of a two-lane road and uses camera sensors and radar sensors (not shown) to capture images of the driving environment 400. The captured images can be processed by the SDCM 132, e.g., as disclosed in connection with Figure 3A and Figure 3B The SDCM 132 can output one or more classifications of traffic signs detected in the driving environment 400. For example, the SDCM 132 can detect signs 404 to 412 and output the following example classifications of the signs.

[0075] Sign 404:

[0076] Sign detection: Bounding box

[0077] Sign type: Speed limit sign

[0078] Marker value: 30 mph

[0079] Marker relevance: Yes

[0080] Marker 406:

[0081] Marker detection: Bounding box

[0082] Marker type: Lane direction

[0083] Marker value: Straight and Right

[0084] Marker relevance: No

[0085] Marker 408:

[0086] Marker detection: Bounding box

[0087] Marker type: Lane direction

[0088] Marker value: Straight and Left

[0089] Marker relevance: Yes

[0090] Marker 410:

[0091] Marker detection: Bounding box

[0092] Marker type: Yield sign

[0093] Marker value: N / A

[0094] Marker relevance: No (Yes)

[0095] Marker 412:

[0096] Marker detection: Bounding box

[0097] Marker type: Stop sign

[0098] Marker value: N / A

[0099] Marker relevance: No (Yes)

[0100] In these examples, "bounding box" can indicate both the dimensions of the sign enclosure and the direction the sign is facing.

[0101] In instances of stop signs 410 and / or 412, sign relevance can be defined differently according to the specific implementation. For example, in one implementation, signs pointing to other vehicles can be classified as irrelevant. In other specific implementations, some signs pointing to other vehicles can still be classified as relevant. For example, stop signs 410 and / or 412 can indirectly affect the movement of vehicle 402 by causing other vehicles (e.g., light truck 414 and bus 416) to yield to vehicle 402.

[0102] Figure 5 An example architecture 500 of a sign detection and classification model enhanced with additional branches for camera feature processing in accordance with some embodiments of the present disclosure is shown. The processing of camera image 202 and radar image 204 can be performed similar to the operations described in connection with Figure 3A the disclosure, e.g., generating a camera feature tensor 314 (for various timestamps t j ) using a camera network 220 and generating a radar feature tensor 314 (for various timestamps t j ). The camera feature tensor 312 can be combined (for multiple timestamps) with the radar feature tensor 314 to form a feature stack 340, which is then used as an input to the BEV backbone 524.

[0103] The additional branch for camera feature processing can include an auxiliary sign classification model 526 that processes the camera feature tensor 312 to perform a preliminary detection of traffic signs, which does not involve processing the radar image 204 or data derived from the radar image 204 (e.g., the radar feature tensor 314). The auxiliary sign classification model 526 can include a decoder network and one or more classification heads ( Figure 5 not explicitly shown in the figure), which can implement any, some, or all of the functions of the sign classification heads 226-n, e.g., detecting bounding boxes and one or more sign characteristics, such as sign type, value, relevance detection, etc. In one example lightweight implementation, the auxiliary sign classification model 526 can output bounding boxes for various hypothesized signs but not other sign characteristics. Signs detected by the auxiliary sign classification model 526 can be filtered by an auxiliary sign filter 560 to merge multiple detections of the same sign, e.g., using NMS, clustering, and / or other filtering techniques.

[0104] Then, the detected (and filtered) signs and / or sign characteristics can be used as additional inputs to the BEV backbone 524. The BEV backbone 524 can have an architecture similar to that of the BEV backbone 224 as described in connection with Figure 3ASimilar architectures (such as those disclosed) may include, for example, convolutional decoders, transformer-based decoders, etc., but may have different numbers of neural nodes, including nodes in the input neural layer and at least some hidden layers. In some embodiments, the BEV backbone 524 may have the same number of output neurons as the BEV backbone 224, and the output of the BEV backbone 524 may be provided to the flag classification head 226-n for further processing, for example, as described in connection with Figure 3B disclosed. The output of the auxiliary flag classification model 526 can be used to inform the BEV backbone 524 of the possible positions of traffic signs within the camera image 202, thereby making flag detection and classification more efficient. Multiple training modes can be used to perform Figure 5 the training of the SDCM. More specifically, in the first training mode, the auxiliary branch can be turned on, and in the second training mode, the auxiliary branch can be turned off (e.g., by replacing the input in the BEV backbone 524 generated using the auxiliary flag classification model 526 with an empty output). Such multi-mode training can train the SDCM to utilize the preliminary output of the auxiliary flag classification model 526 without relying too much on such output (which may be inaccurate in the absence of depth information derived from the radar image 204).

[0105] Figure 6 Another example architecture 600 of a flag detection and classification model enhanced with an additional branch for processing camera features according to some embodiments of the present disclosure is shown. The example architecture 600 differs from the Figure 5 example architecture 500 in that, before the operation of the lifting transform 310, for example, the input of the additional branch for processing camera features is directly performed in the perspective representation by the perspective flag classification model 626, the input of which includes a set of camera feature vectors in the perspective representation. In such an embodiment, the perspective flag classification model 626 can output 2D bounding boxes within the camera image 202. The perspective flag filtering 660 can merge multiple bounding boxes associated with the same flag, for example, as described above in connection with the auxiliary flag filtering 560. Although an embodiment is shown in Figure 6 where the output of the perspective flag classification model 626 is used as the input to the BEV backbone 524, in other embodiments, such output can be provided as an independent output to the tracker / planner 380 (as schematically shown by the dashed arrow in Figure 6 ).

[0106] Although described in reference to the detection and classification of traffic signs in connection with Figure 3A , Figure 3B and Figures 4 - 6the disclosed technology, but in some embodiments, similar technologies can be used to identify the location and status of traffic lights (e.g., red, green, yellow, stop, go, etc.). In such embodiments, images of traffic lights (and / or images of traffic lights and traffic signs) and additional ground truth for the location and status of such traffic lights can be used to train the SDCM 132 and its various components (e.g., camera model, radar model, BET backbone, and classification head).

[0107] Figure 7 FIG. 700 illustrates an example method for deploying a sign detection and classification model in accordance with some embodiments of the present disclosure, the sign detection and classification model using a combination of camera and radar images to accurately identify and read traffic signs in a driving environment. A processing device having one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more parallel processing units (PPUs), and a memory device communicatively coupled to the CPUs, GPUs, and / or PPUs can execute method 700 and / or each of its various functions, routines, subroutines, or operations. Method 700 can relate to systems and components of a vehicle. In some embodiments, the vehicle can be an autonomous vehicle. In some embodiments, the vehicle can be a driver-operated vehicle equipped with a driver assistance system (e.g., a level 2 or 3 driver assistance system) that provides limited assistance with respect to specific vehicle systems (e.g., steering, braking, acceleration, etc.) or under limited driving conditions (e.g., highway driving). The processing device executing method 700 can execute instructions issued by Figure 1 the perception and planning system 130 (more specifically, the SDCM 132) of the vehicle during driving operations of the vehicle. In certain embodiments, a single processing thread can execute method 700. Alternatively, two or more processing threads can execute method 700, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In an illustrative example, the processing threads implementing method 700 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads implementing method 700 can execute asynchronously relative to each other. Compared to the Figure 7 order shown, some operations of method 700 can be executed in a different order. Some operations of method 700 can be executed concurrently with other operations. Some operations can be optional.

[0108] At block 710, method 700 can include obtaining a first set of images and a second set of images using a sensing system of the vehicle. The first set of images can include one or more perspective camera images of the environment (e.g., Figure 3Athe camera images in 202). The second set of images may include one or more radar images of the environment (e.g., Figure 3A the radar images in 204).

[0109] At block 720, method 700 can include generating, using a first neural network (e.g., camera network 220), one or more camera features (e.g., camera feature tensor 312) representative of the first set of images. As indicated by annotation box 722, generating one or more camera features can include mapping the one or more camera features from a perspective coordinate system to a coordinate system associated with the ground surface (e.g., Figure 3A the BEV grid 305 in).

[0110] At block 730, method 700 can include generating, using a second neural network (e.g., radar network 222), one or more radar features (e.g., radar feature tensor 314) representative of the second set of images.

[0111] At block 740, method 700 can include processing the one or more camera features and the one or more radar features to obtain an identification of one or more traffic signs in the environment. In some embodiments, the identification of a single traffic sign among the one or more traffic signs can include determining the location of the single traffic sign, the type of the single traffic sign, a value associated with the semantic content of the single traffic sign, and / or the relevance of the single traffic sign to the vehicle.

[0112] In some embodiments, the one or more camera features being processed include a first camera feature associated with a first time (e.g., (e.g., FT C (t1)), a second camera feature associated with a second time (e.g., FT C (t2)), and so on. The one or more radar features can similarly include a first radar feature associated with a first time (e.g., FT R (t1)), a second radar feature associated with a second time (e.g., FT R (t2)), and so on. In some embodiments, processing the first / second / etc. camera features and the first / second / etc. radar features can be performed simultaneously.

[0113] As Figure 3A shown in the bottom annotation portion of, processing the one or more camera features and the one or more radar features can include: at block 742, using a third neural network. In some embodiments, the third neural network can include a backbone neural network (e.g., Figure 3A and Figure 3B the BEV backbone 224 in) and one or more classification neural networks (e.g., Figure 3Bin the sign classification header 226-n). In some embodiments, the first neural network, the second neural network, and the third neural network can be trained together, e.g., end-to-end.

[0114] In some embodiments, method 700 can include obtaining a set of expected traffic signs at block 744 and eliminating one or more duplicate traffic signs from the set of expected traffic signs (e.g., using Figure 3B the filtering 360 in) to obtain an identification of one or more traffic signs.

[0115] In some embodiments, method 700 can include processing one or more camera features using a fourth neural network (e.g., the auxiliary sign classification model 526) to obtain an auxiliary identification of at least one traffic sign in the environment. Then, method 700 can further include using the auxiliary identification as an additional input to the third NN.

[0116] In some embodiments, the vehicle can be an autonomous vehicle, and method 700 can further include, at block 750, causing the driving control system of the autonomous vehicle to select a driving path for the autonomous vehicle in view of the identification of one or more traffic signs.

[0117] Figure 8 A block diagram depicting an example computer device 800 capable of training and / or deploying a sign detection and classification model according to some embodiments of the present disclosure, the sign detection and classification model using a combination of camera and radar images to accurately identify and read traffic signs in a driving environment. The example computer device 800 can be connected to other computer devices in a LAN, intranet, extranet, and / or the Internet. The computer device 800 can operate in the capacity of a server in a client-server network environment. The computer device 800 can be a personal computer (PC), a set-top box (STB), a server, a network router, a switch or bridge, or any device capable of executing a set of instructions (sequentially or otherwise) specifying actions to be taken by that device. Further, although only a single example computer device is shown, the term "computer" should also be considered to include any collection of computers that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.

[0118] The example computer device 800 can include a processing device 802 (also referred to as a processor or CPU), a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), a static memory 806 (e.g., flash memory, static random access memory (SRAM), etc.), and an auxiliary memory (e.g., a data storage device 818), which can communicate with each other via a bus 830.

[0119] The processing device 802 (which may include processing logic 803) represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device 802 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The processing device 802 may also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. According to one or more aspects of the present disclosure, the processing device 802 may be configured to execute instructions that perform the method 700 of deploying a sign detection and classification model that uses a combination of camera and radar images to accurately identify and read traffic signs in a driving environment.

[0120] The example computer device 800 may also include a network interface device 808, which may be communicatively coupled to a network 820. The example computer device 800 may also include a video display 810 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), and an acoustic signal generating device 816 (e.g., a speaker).

[0121] The data storage device 818 may include a computer-readable storage medium (or more specifically, a non-transitory computer-readable storage medium) 828, on which a set or sets of executable instructions 822 are stored. According to one or more aspects of the present disclosure, the executable instructions 822 may include executable instructions for performing the method 700 of deploying a sign detection and classification model that uses a combination of camera and radar images to accurately identify and read traffic signs in a driving environment.

[0122] The executable instructions 822 may also reside, completely or at least partially, within the main memory 804 and / or within the processing device 802 during execution by the example computer device 800, and the main memory 804 and the processing device 802 also constitute computer-readable storage media. The executable instructions 822 may also be sent or received via the network interface device 808 over a network.

[0123] Although the computer-readable storage medium 828 is in Figure 8is shown as a single medium, but the term "computer-readable storage medium" shall be considered to include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) that store a set of one or more operating instructions. The term "computer-readable storage medium" shall also be considered to include any medium that is capable of storing or encoding a set of instructions for execution by a machine, which instructions cause the machine to perform any one or more of the methods described herein. Thus, the term "computer-readable storage medium" shall be considered to include, without limitation, solid-state memories as well as optical and magnetic media.

[0124] Some of the foregoing detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing art to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, considered to be a self-consistent sequence of steps leading to a desired result. These steps are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. For the sake of generality, it has proven convenient at times to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.

[0125] However, it should be borne in mind that all such and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. Unless specifically stated otherwise, as will be apparent from the following discussion, it should be understood that throughout the description, discussions using terms such as "identifying," "determining," "storing," "adjusting," "causing," "returning," "comparing," "creating," "stopping," "loading," "copying," "throwing," "replacing," "executing," etc. refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the memories or registers of the computer system or other such information storage, transmission, or display devices.

[0126] Examples of the present disclosure also relate to apparatuses for performing the methods described herein. The apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer system selectively programmed by a computer program stored in the computer system. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of medium suitable for storing electronic instructions, each coupled to the computer system bus.

[0127] The methods and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct a more specialized apparatus to perform the required method steps. The required structure of various of these systems will appear as described in the following description. Further, the scope of the present disclosure is not limited to any particular programming language. It should be understood that various programming languages may be used to implement the teachings of the present disclosure.

[0128] It should be understood that the above description is intended to be illustrative and not restrictive. After reading and understanding the above description, many other example embodiments will be apparent to those skilled in the art. Although the present disclosure describes specific examples, it will be recognized that the systems and methods of the present disclosure are not limited to the examples described herein, but may be practiced with modifications within the scope of the appended claims. Therefore, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. Accordingly, the scope of the present disclosure should be determined with reference to the appended claims and the full scope of equivalents to which such claims are entitled.

Claims

1. A system comprising: A sensing system of a vehicle, the sensing system being configured to acquire: a set of camera images of the environment, and a set of radar images of the environment; as well as The data processing system of the vehicle, the data processing system being configured to: generating one or more camera features characterizing the set of camera images using a first neural network NN, generating, using the second NN, one or more radar signatures characterizing the set of radar images; and The one or more camera features and the one or more radar features are processed to obtain identification of one or more traffic signs in the environment.

2. The system according to claim 1, wherein: To generate the one or more camera features, the data processing system is configured to: The one or more camera features are mapped from the perspective coordinate system to a coordinate system associated with a ground surface.

3. The system according to claim 1, wherein: To process the one or more camera features and the one or more radar features, the data processing system is configured to: The one or more camera features and the one or more radar features are processed using a third NN.

4. The system according to claim 3, wherein: The third NN includes a backbone NN and one or more classification NNs.

5. The system according to claim 3, wherein: The one or more camera features include: a first camera feature associated with the first time, and a second camera characteristic associated with a second time, The one or more radar features include: a first radar signature associated with the first time, and a second radar signature associated with the second time; and Wherein, in order to process the one or more camera features and the one or more radar features using the third NN, the data processing system is configured to: The first camera feature, the second camera feature, the first radar feature, and the second radar feature are processed simultaneously using the third NN.

6. The system according to claim 3, wherein: The data processing system is also configured to: processing the one or more camera features using a fourth NN to obtain an assisted recognition of at least one traffic sign in the environment; and The auxiliary recognition is used as an additional input to the third NN.

7. The system according to claim 3, wherein: The first NN, the second NN and the third NN are trained together.

8. The system according to claim 1, wherein: The identification of a single traffic sign of the one or more traffic signs includes determining one or more of the following: the position of the single traffic sign, the type of the single traffic sign, a value associated with the semantic content of the single traffic sign, or The relevance of the single traffic sign to the vehicle.

9. The system according to claim 1, wherein: To process the one or more camera features and the one or more radar features, the data processing system is configured to: obtaining a set of expected traffic signs; and One or more duplicate traffic signs are eliminated from the set of expected traffic signs to obtain an identification of the one or more traffic signs.

10. The system according to claim 1, wherein: The vehicle is an autonomous vehicle, and wherein the data processing system is further configured to: A driving control system of the autonomous vehicle is caused to select a driving path of the autonomous vehicle in view of the identification of the one or more traffic signs.

11. A method comprising: Use the vehicle's sensing systems to obtain: a set of camera images of the environment, and a set of radar images of the environment; generating one or more camera features characterizing the set of camera images using a first neural network NN, generating, using a second NN, one or more radar signatures characterizing the set of radar images; as well as The one or more camera features and the one or more radar features are processed to obtain identification of one or more traffic signs in the environment.

12. The method according to claim 11, wherein: Generating the one or more camera features comprises: The one or more camera features are mapped from the perspective coordinate system to a coordinate system associated with a ground surface.

13. The method according to claim 11, wherein: Processing the one or more camera features and the one or more radar features includes: The one or more camera features and the one or more radar features are processed using a third NN.

14. The method according to claim 13, wherein: The third NN includes a backbone NN and one or more classification NNs.

15. The method according to claim 13, wherein: The one or more camera features include: a first camera feature associated with the first time, and a second camera characteristic associated with a second time, The one or more radar features include: a first radar signature associated with the first time, and a second radar signature associated with the second time; and Wherein, using the third NN to process the one or more camera features and the one or more radar features comprises: The first camera feature, the second camera feature, the first radar feature, and the second radar feature are processed simultaneously using the third NN.

16. The method according to claim 13, further comprising: processing the one or more camera features using a fourth NN to obtain an aided recognition of at least one traffic sign in the environment; as well as The auxiliary recognition is used as an additional input to the third NN.

17. The method according to claim 11, wherein: The identification of a single traffic sign of the one or more traffic signs includes determining one or more of the following: the position of the single traffic sign, the type of the single traffic sign, a value associated with the semantic content of the single traffic sign, or The relevance of the single traffic sign to the vehicle.

18. The method according to claim 11, wherein: Processing the one or more camera features and the one or more radar features includes: obtaining a set of expected traffic signs; and One or more duplicate traffic signs are eliminated from the set of expected traffic signs to obtain an identification of the one or more traffic signs.

19. The method according to claim 11, wherein: The vehicle is an autonomous vehicle, the method further comprising: A driving control system of the autonomous vehicle is caused to select a driving path of the autonomous vehicle in view of the identification of the one or more traffic signs.

20. An autonomous vehicle comprising: one or more cameras configured to acquire a set of camera images of the environment, one or more radar sensors configured to acquire a set of radar images of the environment; The data processing system of the vehicle, the data processing system being configured to: generating one or more camera features characterizing the set of camera images using a first neural network NN, generating, using the second NN, one or more radar signatures characterizing the set of radar images; and processing the one or more camera features and the one or more radar features to obtain identification of one or more traffic signs in the environment; and An autonomous vehicle control system is configured to: The autonomous vehicle is caused to follow a driving path selected in view of the identification of the one or more traffic signs.