End-to-end detection in area with reduced travelling performance in use for autonomous vehicle
An end-to-end perception system using multi-sensor data processing addresses the challenge of accurately detecting reduced traversability areas in autonomous vehicles, enhancing safety and reliability by predicting vehicle trajectories through emergency and construction zones.
Patent Information
- Application Number
- JP2025041528
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-14
- Publication Date
- 2025-10-14
AI Technical Summary
Autonomous vehicles face challenges in accurately detecting and navigating through areas of reduced traversability, such as emergency scenes and construction zones, due to the reliance on rule-based heuristics that often result in false positives or missed detections, especially since these situations are infrequently encountered during normal driving missions.
An end-to-end perception system leveraging multi-sensor modalities and temporal aggregation of data, using unlabeled data and predictive learning to efficiently generalize to new driving situations, processes camera, radar, and lidar data to classify areas as drivable or non-drivable, and predicts vehicle trajectories.
The system provides accurate, reliable, and fast detection of reduced traversability areas without relying on rule-based detection, improving trajectory selection and driving maneuver safety by using a unified feature set from multiple sensing modalities.
Smart Images

Figure 2025156019000001_ABST
Abstract
Description
[Technical Field]
[0001] This specification relates generally to autonomous vehicles, and more particularly to detecting areas of reduced traversability, including closed lanes, emergency scenes, construction zones, and the like. [Background technology]
[0002] Autonomous (fully or partially self-driving) vehicles (AVs) operate by sensing the external environment with a variety of electromagnetic (e.g., radar and optical) and non-electromagnetic (e.g., sound and humidity) sensors. Some autonomous vehicles chart a driving path through the environment based on the sensed data. The driving path may be determined based on global positioning system (GPS) data and road map data. While the GPS and road map data may provide information about static aspects of the environment (such as buildings, street layouts, road closures, etc.), dynamic information (such as information about other vehicles, pedestrians, and street lights) is obtained from contemporaneously collected sensory data. The accuracy and safety of the driving path, as well as the accuracy and safety of the speed regime selected by the autonomous vehicle, depend on the timely and accurate identification of various objects present in the external environment and the ability of the driving algorithms to process information about the environment and provide correct instructions to the vehicle control and drivetrain. [Brief explanation of the drawings]
[0003] The present disclosure is presented by way of example, and not by way of limitation, and may be more fully understood by reference to the following detailed description when considered in conjunction with the figures in which:
[0004] [Figure 1] FIG. 1 illustrates exemplary vehicle components that may deploy an end-to-end (E2E) perception model for reduced drivability area (RDA) detection and navigation, according to some implementations of the present disclosure. [Figure 2]FIG. 2 illustrates an example system architecture that may be used to train and deploy an E2E perception model capable of detecting RDA in a driving environment, according to some implementations of the present disclosure. [Figure 3A] 3A-3B illustrate an example operation of an end-to-end perception model capable of efficient identification and navigation of RDA in a driving environment, according to some implementations of the present disclosure. Figure 3A illustrates a first portion of the RDA detection model operation, including the individual processing of camera, radar, and lidar images. [Figure 3B] FIG. 3B shows the second part of the RDA detection model operation, which involves processing the combined camera, radar, and lidar features. [Figure 4] FIG. 4 is a schematic diagram of an example driving environment for a vehicle deploying an RDA detection model for identifying no-drive zones, according to some implementations of the present disclosure. [Figure 5] FIG. 5 illustrates an example architecture of another end-to-end perception model capable of identifying travel-restricted areas and predicting the trajectory of an autonomous vehicle in a driving environment, according to some implementations of the present disclosure. [Figure 6] FIG. 6 illustrates an example architecture of a multi-stage E2E perception model capable of identifying RDAs and predicting the trajectory of an autonomous vehicle in a driving environment, according to some implementations of the present disclosure. [Figure 7] FIG. 7 illustrates an example method for deploying an E2E perception model for RDA detection and navigation in a driving environment, according to some implementations of the present disclosure. [Figure 8] FIG. 8 is a block diagram of an example computing device capable of training and / or deploying E2E perception models, including one or more RDA detection models that use a combination of camera, radar, and / or lidar imagery for accurate identification and navigation of no-drive zones in a driving environment, in accordance with some implementations of the present disclosure. Summary of the Invention
[0005] In one implementation, a system is disclosed that includes a vehicle sensing system and a vehicle data processing system. The sensing system is configured to acquire first sensed data associated with a first time set. The first sensed data includes a first set of camera images of an environment, a first set of radar images of the environment, and a first set of lidar images of the environment. The data processing system is configured to generate, using a first neural network (NN), one or more camera features that characterize the first set of camera images. The data processing system is further configured to generate, using a second NN, one or more radar features that characterize the first set of radar images. The data processing system is further configured to generate, using a third NN, one or more lidar features that characterize the first set of lidar images. The data processing system is further configured to process the one or more camera features, the one or more radar features, and the one or more lidar features to obtain an indication of reduced traficability areas (RDA) within the environment.
[0006] In another implementation, a method is disclosed that includes obtaining, using a sensing system of a vehicle, first sensory data associated with a first time set. The first sensory data includes a first set of camera images of an environment, a second set of radar images of the environment, and a third set of lidar images of the environment. The method further includes generating, using a first neural network, one or more camera features that characterize the first set of camera images. The method further includes generating, using a second neural network, one or more radar features that characterize the second set of radar images. The method further includes generating, using a third neural network, one or more lidar features that characterize the third set of lidar images. The method further includes processing the one or more camera features, the one or more radar features, and the one or more lidar features to obtain an indication of RDA in the environment.
[0007] In yet another implementation, an autonomous vehicle is disclosed, including a sensing system, a data processing system, and a configured driving control system. The sensing system is configured to acquire sensory data of multiple sensing modalities. The multiple sensing modalities are selected from at least a camera sensing modality, a radar sensing modality, or a radar sensing modality. The data processing system is configured to generate, using a first neural network, one or more first features characterizing the sensory data of the first sensing modality. The data processing system is further configured to generate, using a second neural network, one or more second features characterizing the sensory data of the second sensing modality. The data processing system is further configured to process, using a third neural network, the one or more first features and the one or more second features to obtain an indication of RDA in the environment of the autonomous vehicle. The first neural network, the second neural network, and the third neural network are jointly trained using training data for each sensing modality of the multiple sensing modalities. The driving control system is configured to select a driving path for the autonomous vehicle taking into account the indication of RDA. DETAILED DESCRIPTION OF THE INVENTION
[0008] Autonomous vehicles or vehicles deploying various advanced driver assistance features can use multiple sensor modalities to facilitate the detection of objects in the external environment and predict the future trajectories of such objects. Sensors can include radio detection and ranging (radar) sensors, light detection and ranging (lidar) sensors, digital cameras, ultrasonic sensors, position sensors, etc. Different types of sensors can offer different, complementary benefits. For example, radar and lidar emit electromagnetic signals (radio or optical signals) that reflect off objects and return information about the object's distance (e.g., determined from the signal's time of flight) and the object's velocity (e.g., via a Doppler shift in the frequency of the reflected signal). Radar and lidar can scan an entire 360-degree view by using a series of consecutive sensing frames. The sensing frames can include multiple reflections that cover the external environment in a dense grid of return points. Each return point can be associated with the distance to the corresponding reflecting object and the reflecting object's line-of-sight velocity (the component of velocity along the line of sight).
[0009] Lidars have high spatial resolution due to their submicron or micron optical wavelengths, which facilitate obtaining many closely spaced return points from the same object. This allows for accurate detection and tracking of objects once they are within the lidar sensor's range. Radar sensors are less expensive, require less maintenance than lidar sensors, have a larger operating range, and are more resistant to adverse weather conditions. Cameras (e.g., photo or video cameras) capture a two-dimensional projection of three-dimensional external space onto an image plane (or other non-planar imaging surface) and can obtain high-resolution images at both close and long distances.
[0010] Various sensors in a vehicle's sensing system (e.g., lidar, radar, cameras, and / or other sensors such as sonar) capture complementary representations of objects in the vehicle's environment. The vehicle's perception system identifies objects based on their appearance, motion state, object trajectory, and / or other characteristics. For example, lidar can accurately map the shape of one or more objects (using multiple return points) and can further determine the distance to and / or velocity of those objects. Cameras can obtain visual images of objects. The perception system can map the shapes and positions of various objects in the environment (obtained from lidar data) to their visual representations (obtained from camera data) and perform several computer vision operations, such as segmenting (clustering) the sensory data among individual objects (clusters), identifying the type / make / model / etc. of individual objects, etc. Prediction and planning systems can track the movements (including, but not limited to, position and velocity) of various objects over multiple times and then infer the future from previously observed movements. This predicted movement can be used by various vehicle control systems to select a driving path that takes these objects into account, e.g., avoiding the objects, slowing the vehicle in the presence of the objects, and / or taking some other suitable action.
[0011] In addition to detecting moving objects, vehicle sensing systems serve the important purpose of identifying various semantic information, such as markings on road pavements (e.g., lane boundaries, stop line locations, etc.), traffic signals, traffic signs, indications of areas where traffic is temporarily closed, or areas where traffic is restricted. For example, emergency (e.g., fire, police, ambulance, environmental hazard, etc.) responders may temporarily close or restrict other drivable areas by, for example, diverting all traffic to a detour or directing traffic into specific lanes, establishing temporary reversible lanes to manage vehicular flow in both directions of traffic, and / or taking any other measures. Such semantically closed or restricted areas are typically not marked on a map, and autonomous vehicles must rely on sensor data to identify and navigate such areas. Similarly, construction crews may begin maintenance operations without prior warning and restrict traffic within the work zone. Furthermore, even when a work zone is marked on a map, in some cases, there may be no viable alternative route, and autonomous vehicles must drive near the work zone. Furthermore, the layout of a construction zone may be constantly changing, e.g., lanes being redirected in different directions, previously open lanes being closed, and closed lanes being reopened. Such closed areas, partially drivable areas, construction zones, etc. are referred to herein as restricted drive areas (RDAs). RDAs may be marked or semantically identified with a diverse set of features (markers) that may be very situation-specific, e.g., a police car or fire engine blocking the street, emergency responders walking on the road, a "Road Closed" (or similar) temporary sign placed to mark the RDA, warning tape deployed across one or more lanes, a fire hose connected to a fire hydrant or fire engine and placed on the pavement, a flare / light marking the boundary of the non-drivable portion of the road, emergency responders or construction workers walking on the road, etc.
[0012] Because there are numerous situations that can trigger RDA, training a reliable machine learning model that can detect all of them is challenging. Collecting a representative training dataset can be challenging, especially since various RDA situations are not frequently encountered during normal driving missions. Therefore, existing techniques rely on rule-based RDA identification, deploying situation-specific heuristics. For example, detecting an accident scene can rely on determining that a police car is blocking a road with its emergency lights on, while identifying RDA due to construction may rely on detecting safety cones, plastic barriers, barricades, and so on. However, rule-based heuristics may not fully capture the broader context of the scene, resulting in false positives or missed RDAs. For example, even a stationary or moving police car can be mistaken for an RDA marker. Similarly, a person wearing safety gear crossing a road other than a crosswalk can be mistaken for a firefighter or construction worker, triggering an undesired response, such as an autonomous vehicle stopping and blocking traffic. Anticipating and formulating a variety of exceptions to the many rule-based heuristics that encompass the virtually infinite number of real-world situations is a daunting task.
[0013] Aspects and implementations of the present disclosure address these and other challenges of modern technology by providing an end-to-end (E2E) perception system that leverages multi-sensor modalities and temporal aggregation of sensed data, using unlabeled data and predictive learning to efficiently generalize to new driving situations. The E2E perception model can use data from multiple sensed modalities, such as camera data, radar data, lidar data, and audio data, as input and generate an output that classifies various elements (e.g., areas mapped to pixels) of the driving environment (e.g., a specific area around an autonomous vehicle) by their likelihood of belonging to a RDA. The elements can then be combined into connected areas (e.g., closed areas, construction areas, etc.) that represent one or more RDAs of the driving environment.
[0014] More specifically, various streams of data, e.g., camera streams, radar streams, lidar streams, etc., can be first processed by respective modality networks, e.g., camera images can be processed by a camera network, radar images can be processed by a radar network, etc. For example, the radar network generates a set of radar features (feature vectors, embeddings) associated with specific coordinates x,y in a two-dimensional bird's-eye view (BEV) grid, thereby generating a radar feature F R (x,y;t) characterizes the presence (or absence) of a reflective object located at point x,y of the BEV grid or capture of radar data (image) at a given time t. In some implementations, the radar data may be initially generated in polar (or spherical) coordinates, with subsequent mapping performed to grid (Cartesian) coordinates as part of a gather transformation that associates various points of the radar point cloud with specific locations in the BEV grid. Furthermore, the radar feature F R (x,y;t) can characterize the type of reflection, for example, distinguishing reflections from metallic objects (traffic signs, vehicles, etc.) from reflections from non-metallic objects (e.g., trees, concrete structures, etc.). Radar feature F R (x,y;t) can further characterize the context of reflection point x,y, e.g., the semantic association of point x,y with respect to various other radar reflection points x',y'. The coordinates of various reflection points can be determined directly from the radar data (e.g., distance and bearing toward the radar signal reflection point). Similarly, the lidar network can generate lidar features F that characterize the type of lidar reflection at point x,y of the BEV grid at time t, the context provided by various other lidar reflection points x',y', etc. L (x,y;t). The camera network can similarly generate camera features F that characterize the visual appearance of the portion of the environment associated with point x,y of the BEV grid at time t. C(x,y;t) can be determined. Because camera images lack explicit distance (depth) information, the camera network can also perform a lift (collection) transform (along with or after feature generation) that associates various pixels in the camera image with points x,y on the BEV grid that are also associated with the radar / lidar returns. The lift transform can be performed by estimating the most likely distance associated with a given pixel in the camera image (e.g., the distance to an object or a portion of an object indicated by the pixel) or by evaluating the entire distribution of various such possible distances. Accordingly, the camera network can map the camera features to the same BEV grid as the radar / lidar network maps the corresponding radar / lidar features.
[0015] Different sensing modalities offer complementary benefits. For example, camera imagery has rich contextual information and captures the scene at both short and long ranges. Lidar data provides high-resolution imaging that is most effective at short to medium ranges. Radar data has lower resolution but can reach long ranges and is robust in adverse weather conditions.
[0016] In some implementations, the camera features, radar features, and lidar features may then be processed by another model, also referred to herein as the BEV backbone model, to generate a unified feature set {F C (x,y;t),F R (x,y;t),F L(x,y;t)} → F(x,y;t). The BEV backbone model can further incorporate the temporal context of the sensing data, e.g., a stack (tensor) of integrated features corresponding to multiple times t: {F(x,y;t1),F(x,y;t2),F(x,y;t3)…}. In some implementations, the BEV backbone network can feed intermediate outputs to multiple classifier (detection) heads that output classes of various BEV elements x,y. For example, an RDA detection head can classify BEV elements x,y as drivable (normal road) or non-drivable points. More specifically, the output of an RDA detection head is expressed as a probability P RDA (x,y), for example, P RDA ≒0 indicates a reliable element, and P RDA ≒1 indicates an element that is definitely not driveable, and P RDA ≒0.5 indicates elements that are equally likely to be drivable or non-drivable (they may be elements on or near the boundary between drivable and non-drivable areas). RDA ≧P T A cluster (or clusters) of elements having a probability greater than 0 indicates an area of the driving environment that is presumably inaccessible to vehicles.
[0017] A separate driving prediction head calculates the likely driving path, e.g., the probability P of an experienced driver passing through various points x, y while passing through the current driving environment. D For example, an E2E perception model can determine that upon encountering a road blocked (e.g., by firefighters), an expert driver will likely make a right / left turn, a U-turn, and / or other driving maneuvers (e.g., wait to pass).
[0018] An E2E perception model can be trained using several techniques, including supervised learning and imitation learning, among others. For example, a sensory data log collected for a training driving environment (during a previous, past driving mission) can be annotated with RDA boundaries, e.g., by a human developer. In training an E2E perception model, such annotations can be used as ground truth for supervised training of an RDA detection head. Imitation (e.g., self-supervised) learning can be used to train a driving prediction head to generate driving probabilities P that mimic driving maneuvers actually performed by expert human drivers during past missions. D It can output a map of (x,y). The backbone network and / or classification head can be trained using a dropout technique, e.g., dropping (replacing with zeros) sensory data for one or more modalities for at least some training epochs so that the E2E perception model learns to more efficiently use the remaining sensory modalities.
[0019] The output of the E2E perception model can be passed to a planner module to chart and implement a vehicle trajectory that is consistent with the identified RDA and / or estimated human-preferred driving path. In a driver assistance system operating in a driver-controlled mode, the detected RDA can be communicated to the driver, e.g., displayed on the dashboard, accompanied by an audio alert, etc. Similarly, the dashboard can display the most likely driving path that the perception model estimates an experienced driver would take.
[0020] Advantages of the described implementation include, but are not limited to, accurate, reliable, and fast detection and mapping of RDA using an end-to-end perception model that does not rely on rule-based detection. As a result, multiple heuristic classifiers can be replaced with a single, more accurate end-to-end perception model, resulting in improved trajectory selection and increased driving maneuver safety.
[0021] As used in this disclosure, a feature vector (embedding) should be understood as any suitable digital representation of input data, for example, a vector (string) of any number M of components, which may have integer or floating-point values. A feature vector may be viewed as a point in an M-dimensional embedding space. The dimension M of the embedding space (defined as part of any suitable model architecture) may be smaller than the size of the input data (camera / radar / lidar images). During training, the model learns to associate similar sets of training input data with similar feature vectors represented by points located closely together in the embedding space, and further learns to associate different sets of training input data with points located further apart in that space.
[0022] In these instances where the description of implementation refers to an autonomous vehicle, it should be understood that similar techniques may be used in various driver assistance systems that do not rise to the level of a fully autonomous driving system. In some embodiments, the disclosed techniques may be used in Level 2 driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, etc., and other driver support. In some embodiments, the disclosed techniques may be used in Level 3 driver assistance systems that are capable of autonomous driving under limited (e.g., highway) conditions. Such systems may use fast and accurate object detection and tracking to alert the driver to approaching vehicles and / or other objects and allow the driver to make final driving decisions (e.g., in Level 2 systems) or certain driving decisions such as slowing down, changing lanes, etc. (e.g., in Level 3 systems) without requiring driver feedback.
[0023] 1 is a diagram illustrating components of an example vehicle 100 in which an E2E perception model for RDA detection and navigation may be deployed, according to some implementations of this disclosure. An autonomous vehicle may include a motor vehicle (such as a car, truck, bus, motorcycle, all-terrain vehicle, recreational vehicle, any specialized agricultural or construction vehicle, etc.), a watercraft (such as a ship, boat, yacht, submarine, etc.), or any other self-propelled vehicle capable of operating in an autonomous driving mode (without or reduced human input) (e.g., a robot, a factory or warehouse robotic vehicle, a sidewalk delivery robotic vehicle, etc.).
[0024] The driving environment 101 may include any objects (moving or non-moving) located outside the vehicle 100, such as roads, buildings, trees, bushes, sidewalks, bridges, mountains, other vehicles, pedestrians, etc. The driving environment 101 may be an urban area, a suburban area, a rural area, etc. In some implementations, the driving environment 101 may be an off-road environment (e.g., agriculture, or farmland). In some implementations, the driving environment may be an indoor environment, such as an industrial plant environment, a shipping warehouse, a hazardous area of a building, etc. In some implementations, the driving environment 101 may be substantially flat, with various objects moving parallel to the surface (e.g., parallel to the ground). In other implementations, the driving environment may be three-dimensional and may include objects capable of moving along all three directions (e.g., balloons, leaves, etc.). Hereinafter, the term "driving environment" should be understood to include all environments in which autonomous movement of a self-propelled vehicle may occur. For example, the "driving environment" may include any possible flight environment of an aircraft or a marine environment of a ship. Objects in the driving environment 101 may be located at any distance from the vehicle 100, from as close as a few feet (or less) to several miles (or more).
[0025] As described herein, in a semi-autonomous or partially autonomous driving mode, the vehicle assists with one or more driving maneuvers (e.g., steering, braking, and / or accelerating to perform lane centering, adaptive cruise control, advanced driver assistance systems (ADAS), or emergency braking), but the human driver is expected to maintain situational awareness of the vehicle's surroundings and supervise the assisted driving maneuvers. Here, the vehicle may perform all driving tasks in a particular situation, but the human driver is expected to be responsible for taking control as needed.
[0026] For simplicity and brevity, various systems and methods will be described below in conjunction with autonomous vehicles, although similar techniques can be used in various driver assistance systems that fall short of fully autonomous driving systems. In the United States, the Society of Automotive Engineers (SAE) defines different levels of automated driving operation to indicate how much or how little control a vehicle has over the driving; however, different organizations in the United States, or elsewhere, may classify the levels differently. More specifically, the disclosed systems and methods can be used in SAE Level 2 (L2) driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. The disclosed systems and methods can be used in SAE Level 3 (L3) driver assistance systems that are capable of autonomous driving under limited (e.g., highway) conditions. Similarly, the disclosed systems and methods can be used in vehicles using SAE Level 4 (L4) automated driving systems, which operate autonomously under most normal driving conditions and require only occasional attention from a human operator. In all such driver assistance systems, accurate lane estimation can be performed automatically without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, and overall safety of autonomous, semi-autonomous, and other driver assistance systems. As noted above, in addition to the way SAE classifies levels of autonomous driving operation, other organizations in the United States or other countries may classify levels of autonomous driving operation differently. Without limitation, the systems and methods disclosed herein may be used in driver assistance systems defined by the levels of autonomous driving operation of these other organizations.
[0027] The exemplary vehicle 100 may include a sensing system 110. The sensing system 110 may include various electromagnetic (e.g., optical) and non-electromagnetic (e.g., audio) sensing subsystems and / or devices. The sensing system 110 may include a radar (or multiple radars) 112, which may be any system that utilizes radio or microwave frequency signals to sense objects within the vehicle's 100 operating environment 101. The radar 112 may be configured to sense both the spatial location of an object and the object's velocity (e.g., using Doppler shift techniques). Hereinafter, "velocity" refers to both how fast an object is moving (object speed) as well as the direction of the object's motion. In some implementations, the sensing system 110 may include a lidar 114, which may be a laser-based unit that can determine the distance to and velocity of objects (including their spatial dimensions) within the operating environment 101. Each of the radar 112 and the lidar 114 may include a coherent sensor, such as a frequency-modulated continuous wave (FMCW) lidar or radar sensor. For example, the radar 112 may use heterodyne detection for velocity determination. In some implementations, ToF and coherent radar functionality are combined into a radar unit that can simultaneously determine both the distance to a reflecting object and the line-of-sight velocity of the reflecting object. Such a unit may be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection), or both modes simultaneously. In some implementations, multiple radars 112 or lidars 114 may be mounted on the vehicle 100.
[0028] The LIDAR 114 may include one or more light sources that generate and emit signals and one or more detectors for signals reflected back from objects. In some implementations, the LIDAR 114 can perform a 360-degree scan in the horizontal direction. In some implementations, the LIDAR 114 may be capable of spatial scanning along both the horizontal and vertical directions. In some implementations, the field of view may be up to 90 degrees vertically (e.g., at least a portion of the area above the horizon is scanned by the LIDAR signal). In some implementations, the field of view may be a full sphere (consisting of two hemispheres).
[0029] The sensing system 110 may further include one or more cameras 118 to capture images of the driving environment 101. The images may be two-dimensional projections of the driving environment 101 (or portions of the driving environment 101) onto the camera's imaging surface (planar or non-planar). Some of the cameras 118 of the sensing system 110 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101. The sensing system 110 may also include one or more infrared (IR) sensors 119. The sensing system 110 may further include one or more microphone sensors 130 116 that may be used to capture audio data for the driving environment, such as, for example, emergency vehicle sirens and other sounds.
[0030] The sensory data obtained by the sensing system 110 may be processed by the data processing system 120 of the vehicle 100. For example, the data processing system 120 may include a recognition and planning system 130. The recognition and planning system 130 may be configured to detect and track objects in the driving environment 101 and recognize the detected objects. For example, the recognition and planning system 130 may analyze images captured by the camera 118 and may be capable of detecting traffic signals, road signs, road layouts (e.g., lane boundaries, intersection topology, parking designations, etc.), the presence of obstacles, etc. The recognition and planning system 130 may further receive radar sensing data (Doppler data and ToF data) and determine the distances to various objects in the environment 101 and the velocities of such objects (radially, as described below, and, in some implementations, laterally). In some implementations, the recognition and planning system 130 may use radar data in combination with data captured by the camera 118, as described in more detail below.
[0031] The perception and planning system 130 monitors how the driving environment 101 evolves over time, for example, by tracking the position and velocity of moving objects (e.g., relative to the Earth and / or the AV) and predicting how various objects will move over a particular time horizon, e.g., 1-10 seconds or more. The perception and planning system 130 may include an RDA detection model 132 that identifies road segments within the restricted-traffic environment 101. The RDA detection model 132 may include one or more trainable MLMs that can process data from multiple modalities, e.g., camera data, radar data, lidar data, audio data, etc.
[0032] The perception and planning system 130 may also receive information from a positioning subsystem 122, which may include a GPS transceiver and / or an inertial measurement unit (IMU) (not shown in FIG. 1 ) configured to obtain information about the position of the AV relative to the Earth and its surroundings. The positioning subsystem 122 may use the positioning data, e.g., GPS data and IMU data, in conjunction with the sensed data to help accurately determine the location of the vehicle 100 relative to fixed objects in the driving environment 101 (e.g., roads, lane boundaries, intersections, sidewalks, crosswalks, road signs, curves, surrounding buildings, etc.), whose locations may be provided by map information 124. In some implementations, the data processing system 120 may receive non-electromagnetic data, such as audio data (e.g., ultrasonic sensor data or data from one or more microphones that detect emergency vehicle sirens), temperature sensor data, humidity sensor data, pressure sensor data, weather data (e.g., wind speed and direction, precipitation data), etc.
[0033] Data generated by perception and planning system 130, position subsystem 122, and / or other systems and components of data processing system 120 may be used by an autonomous driving system, such as vehicle control system (VCS) 140. VCS 140 may include one or more algorithms that control how vehicle 100 should behave in various driving situations and environments. For example, VCS 140 may include a navigation system for determining a global driving route to a destination. VCS 140 may also include a driving path selection system for selecting a particular path through the immediate driving environment, which may include selecting a lane, navigating traffic jams, selecting a location to make a U-turn, selecting a trajectory for a parking maneuver, etc. VCS 140 may also include an obstacle avoidance system for safely avoiding various obstacles in the AV's driving environment (e.g., stones, stranded vehicles, pedestrians crossing the road, etc., ignoring traffic rules and signals). The obstacle avoidance system may be configured to evaluate the size of an obstacle and its trajectory (if the obstacle is moving) and select an optimal driving strategy (e.g., braking, steering, accelerating, etc.) to avoid the obstacle.
[0034] The algorithms and modules of VCS 140 may generate instructions for various systems and components of the vehicle, such as powertrain, braking, and steering 150, vehicle electronics 160, signaling 170, and other systems and components not explicitly shown in FIG. 1 . Powertrain, braking, and steering 150 may include an engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanism, and other systems. Vehicle electronics 160 may include an on-board computer, engine management, ignition, communication systems, car computers, telematics, in-car entertainment systems, and other systems and components. Signaling 170 may include high and low headlights, stop lights, turn signals and taillights, horns and alarms, interior lighting systems, dashboard notification systems, passenger notification systems, radio and wireless network transmission systems, etc. Some of the commands output by VCS 140 may be delivered directly to powertrain, braking, and steering 150 (or signaling 170), while other commands output by VCS 140 are first delivered to vehicle electronics 160, which generates commands to powertrain, braking, and steering 150 and / or signaling 170.
[0035] In one example, VCS 140 may determine that an obstacle identified by data processing system 120 should be avoided by slowing the vehicle until a safe speed is reached and then steering the vehicle around the obstacle. VCS 140 may output commands to powertrain, brakes, and steering 150 (either directly or via vehicle electronics 160) to (1) reduce fuel flow to the engine and lower engine speed by modifying the throttle setting, (2) downshift the drivetrain to a lower gear via an automatic transmission, (3) engage a brake unit to reduce vehicle speed (working in coordination with the engine and transmission) until a safe speed is reached, and (4) use the power steering mechanism to implement a steering maneuver until the obstacle is safely bypassed. VCS 140 may then output commands to powertrain, brakes, and steering 150 to resume the vehicle's previous speed setting.
[0036] An "autonomous vehicle" may include a motor vehicle (such as a car, truck, bus, motorcycle, all-terrain vehicle, recreational vehicle, any specialized agricultural or construction vehicle, etc.), a watercraft (such as a ship, boat, yacht, submarine, etc.), a robotic vehicle (e.g., a factory, warehouse, sidewalk delivery robot, etc.), or any other self-propelled vehicle capable of operating in an autonomous mode (with no human input or reduced human input). An "object" may include any body, item, device, body, or thing (moving or non-moving) located outside the autonomous vehicle, such as a road, building, tree, bush, sidewalk, bridge, mountain, other vehicle, pier, embankment, runway, animal, bird, or other thing.
[0037] FIG. 2 illustrates an example system architecture 200 that may be used to train and deploy an end-to-end perception model capable of detecting RDA in a driving environment, according to some implementations of the present disclosure. Input to the RDA detection model 132 may include data acquired by the sensing system 110 (e.g., by radar 112, lidar 114, camera 118, and / or other sensors, see FIG. 1 ). The acquired data may be provided via a sensed data acquisition module 210, which may decode, preprocess (e.g., denoising, upsampling, downsampling, etc.) the data, and reformat the data into a format accessible to the RDA detection model 132. In one example implementation, the sensed data acquisition module 210 may acquire a series of camera images 202, e.g., two-dimensional projections of the driving environment (or a portion thereof) onto an array of sensed detectors (e.g., charge-coupled device or CCD detectors, complementary metal-oxide semiconductor or CMOS detectors, etc.). Individual camera images may have pixels of varying intensities of one color (in the case of a black-and-white image) or multiple colors (in the case of a color image). The camera image 202 may be a panoramic (360-degree) image or an image showing a specific portion of the driving environment. The camera image 202 may include a large number of pixels. The number of pixels may depend on the resolution of the image. Each pixel may be characterized by one or more luminance values. A black and white pixel may, for example, be characterized by a single luminance value representing the brightness of the pixel, with a value of 1 corresponding to a white pixel and a value of 0 corresponding to a black pixel (or vice versa). The luminance value may assume a continuous (or discrete) value between 0 and 1 (or any other selected limits, for example, 0 to 255). Similarly, a color pixel may be represented by multiple luminance values, for example, three luminance values (e.g., when using an RGB color encoding scheme) or four luminance values (e.g., when using a CMYK color encoding scheme). The camera image 202 may undergo preprocessing, such as downscaling (combining multiple pixel luminance values into a single pixel value), upsampling, filtering, noise removal, etc.The camera image 202 can be in any suitable digital format (JPEG, TIFF, GIG, BMP, CGM, SVG, etc.).
[0038] The sensed data acquisition module 210 can further obtain radar images 204 (and similarly, lidar images 206), which may include a set of return points (point clouds) corresponding to radar (lidar) beam reflections from various objects in the driving environment. Each return point may be understood as a data unit (pixel) including the coordinates of the reflecting surface, line-of-sight velocity data, brightness data, etc. For example, the sensed data acquisition module 210 can provide radar images 204 (and similarly, lidar images 206) including a radar (lidar) intensity map I(R, θ, φ), where R, θ, φ are a set of spherical coordinates. In some implementations, Cartesian, ellipsoidal, parabolic, or other suitable coordinates may be used instead. The radar (lidar) intensity map identifies the intensity of radar (lidar) reflections for various points within the radar (lidar) field of view. The coordinates of objects reflecting radar (lidar) signals can be determined from direction data (e.g., the polar angle θ and azimuth angle φ of the signal transmission direction) and range data (e.g., the radial range R determined from the signal's time-of-flight). The radar image 204 (and similarly the lidar image 206) may further include velocity data for various reflecting objects identified based on the detected Doppler shift of the reflected signals.
[0039] The camera image 202, radar image 204, and / or lidar image 206 may be a large image of the entire driving environment or an image of a smaller portion of the driving environment (e.g., a camera image acquired by a front-facing camera of the sensing system 110). In some implementations, the sensed data acquisition module 210 may crop the camera image 202, radar image 204, and / or lidar image 206 to correspond to a particular section around the vehicle's direction of travel. For example, because the drivable area of interest is typically located around the vehicle's direction of travel, the sensed data acquisition module 210 may, in one example, non-limiting implementation, crop the camera image 202, radar image 204, and / or lidar image 206 to a forward-looking section that is 200-250 meters long and 20-40 meters wide. The size of the section may depend on the vehicle's speed and the type of driving environment, and may differ between highway and urban driving environments. Camera images 202 may be processed by camera network 220, radar images 204 may be processed by radar network 222, and lidar images 206 may be processed by lidar network 224. Camera network 220 generates camera features, radar network 222 generates radar features, and lidar network 224 generates lidar features (these features are not shown in FIG. 2). The camera features, radar features, and lidar features can be associated with a two-dimensional bird's eye view (BEV).
[0040] Any, some, or all of the camera features, radar features, and lidar features may be combined and processed by the BEV backbone 230 and one or more road classification heads 232. The road classification heads 232 may classify pixels of the BEV grid and corresponding sections of the road as normal (drivable) or restricted, as areas that an experienced human driver would drive over or avoid, etc. Various networks of the RDA detection model 132 may include convolutional neural networks, recurrent neural networks (RNNs) with one or more hidden layers, fully connected neural networks, long short-term memory neural networks, Transformers, Boltzmann machines, etc. In some implementations, the RDA detection model 132 may further process the audio data 208 (e.g., collected by one or more microphone sensors 116 of FIG. 1 ) using an audio network 228 that generates audio features used as additional input to the BEV backbone 230. The audio data 208 may include spectrograms, mel-spectrograms, or any other suitable digital representation of the audio (sound) collected from the driving environment.
[0041] The output of the RDA detection model 132 may include an identified drivable area 234 that is provided to a tracker / planner 260, which may be part of the recognition and planning system 130 of FIG. 1 . The tracker / planner 260 may track the movement (e.g., relative to the vehicle) of roadblocks (e.g., emergency vehicles, safety cones, flares, tape, barriers, etc.), traffic signs, other vehicles, and any other objects. In some implementations, the behavior of objects identified by the RDA detection model 132 may be tracked using a suitable motion filter, e.g., a Kalman filter. The Kalman filter calculates the most likely geomotion data by taking into account the obtained measurements (e.g., the output of the RDA detection model 132), predictions made according to a physical model of the object's motion, and some statistical assumptions about the measurement errors (e.g., an error covariance matrix). The tracker / planner 260 may also select a vehicle path that matches the identified traffic sign and provide instructions to the vehicle control system 140 for implementing the selected driving path.
[0042] Training of the RDA detection model 132 and / or other MLMs can be performed by a training engine 242 hosted by a training server 240, which may be an external server deploying one or more processing devices, such as a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), etc. The training engine 242 can access a data store 250 that stores various training data for training the RDA detection model 132. In some implementations, the training data can include camera images 252 acquired during actual driving missions by an onboard camera and can further include radar images 254 and / or lidar images 256 associated with the camera images 252, e.g., radar / lidar images of substantially the same area of the corresponding driving environment acquired substantially simultaneously with the camera images. The training data stored by the data store 250 can further include maneuverability ground truth 258, which can include, for example, the correct identification of areas of the environment where maneuverability is restricted, such as a polygon surrounding the RDA. In some implementations, such ground truth can be determined by a developer manually identifying restricted areas of the environment. The maneuverability ground truth 258 may further include driving trajectories, e.g., sections of the driving environment driven by a vehicle operated by an experienced human driver. In some implementations, such ground truth may be determined from logs of past driving missions.
[0043] 2, the RDA detection model 132 may be trained using training data including training inputs 244 and corresponding target outputs 246 (correct matches for each training input). During training, the training engine 242 may pull training data from a data store 250, prepare one or more training inputs 244 and one or more target outputs 246 (ground truth), and use the prepared inputs and outputs to train one or more models, including, but not limited to, the RDA detection model 132. The training data may also include mapping data 248 that maps the training inputs 244 to the target outputs 246. During training of the RDA detection model 132, the training engine 242 may cause the RDA detection model 132 to learn patterns in the captured training data by training input / target output pairs. To evaluate the difference between the training output and the target output 246, the training engine 242 can use various suitable loss functions, such as a mean squared error loss function (e.g., to evaluate deviation from a continuous ground truth value, e.g., distance to a landmark), a binary cross-entropy loss function (e.g., to evaluate deviation from a binary classification), and / or other suitable loss functions. In some implementations, the RDA detection model 132 can be trained by the training engine 242 and then downloaded onto the vehicle perception and planning system 130.
[0044] During training of the RDA detection model 132, the training engine 242 may modify parameters (e.g., weights and biases) of various networks in the RDA detection model 132 until the model successfully learns to accurately identify RDA and / or correctly predict a vehicle's driving path to avoid RDA. In some implementations, multiple RDA detection models 132 may be trained for use under different conditions and for different driving environments, e.g., separate RDA detection models 132 may be trained for city driving and highway driving. Different trained RDA detection models 132 may have different architectures (e.g., different numbers of neuron layers and neural connections of different topologies), may have different settings (e.g., activation function type and parameters, etc.), and may be trained using different sets of hyperparameters (e.g., number of epochs, learning rate, etc.).
[0045] Data store 250 may be persistent storage capable of storing data structures configured to facilitate accurate and rapid identification and verification of radar images, camera images, and landmark detections according to various implementations of the present disclosure. Data store 250 may be hosted by one or more storage devices, such as main memory, magnetic or optical storage disks, tape, or hard drives, network-attached storage (NAS), storage area networks (SANs), etc. Although illustrated separately from training server 240, in some implementations data store 250 may be part of training server 240. In some implementations, data store 250 may be a network-attached file server, while in other implementations data store 250 may be some other type of persistent storage, such as an object-oriented database, a relational database, etc., that may be hosted by the server machine or one or more different machines accessible to training server 240 over a network (not shown in FIG. 2 ).
[0046] 3A-3B illustrate an example operation of an end-to-end perception model capable of efficient RDA identification and navigation in a driving environment, according to some implementations of the present disclosure. Figure 3A illustrates a first portion 300 of the operation of an RDA detection model (e.g., RDA detection model 132 of Figures 1 and 2). As shown in Figure 3A, first portion 300 processes any, some, or all of the individual modalities of input data 301, such as camera imagery 202, radar imagery 204, lidar imagery 206, audio data (not explicitly shown in Figure 3A), etc.
[0047] Each camera image 202 (and similarly, radar image 204 and lidar image 206) can be associated with a particular time t1, t2, t3, ... when the respective image was captured. The acquisition of camera image 202, radar image 204, and lidar image 206 can be synchronized so that images of multiple modalities depict the driving environment substantially simultaneously. Camera image 202 can be processed by camera network 220, radar image 204 can be processed by radar network 222, and lidar image 206 can be processed by lidar network 224. In some implementations, any, some, or all of networks 220-224 can capture images of multiple times t j In some implementations, any, some, or all of the networks 220-224 process images associated with different times t j The images associated with the
[0048] The camera network 220, the radar network 222, and similarly, the lidar network 224 may have any suitable architecture. In one example, any, some, or all of the networks 220-224 may be or include deep convolutional neural networks using a U-net architecture, which includes an encoder stage and a decoder stage. Each stage may have multiple convolutional neural layers and one or more fully connected layers. A convolutional encoder may include any number of filters (kernels) that identify image features by broadening the recognition field, aggregating relevant information captured by individual units (pixels) of the image, and encoding this information via features arranged in a feature map. Such feature maps may be generated using a series of convolutional and pooling (e.g., average pooling or max pooling) layers. A convolutional layer applies multiple filters (typically tens, hundreds, or more), which are limited-size matrices with trained weights that scan the entire image looking for certain features within the image. Different kernels may look for different features, such as traffic sign boundaries, sign shapes, sign color patterns, the presence of text within a sign, etc. The kernel can be moved across the image in steps (strides) smaller than the kernel dimensions (e.g., a 5x5 pixel kernel can be shifted by 1, 2, or 3 pixels during each step) to form the neural activation function signal. Subsampling (pooling) operations then reduce the dimensionality of the resulting feature map, in accordance with the fundamental premise of convolutional neural network architecture: information about the presence of a target feature is often more important than precise knowledge of the feature's coordinates. As a result of this multi-layer convolution and pooling process, the intermediate representation of the image grows along the feature (channel) dimension but shrinks along the image's width / height dimensions. This reduction speeds subsequent computations while simultaneously ensuring the neural network's ability to process input images at different scales.
[0049] The camera network 220, radar network 222, and similarly, the lidar network 224, can upsample the feature maps produced by the encoders while reducing the feature / channel dimensionality (which may be implemented using another set of learned deconvolution kernels), e.g., gradually increasing the resolution back to the original (or reduced) dimensionality of the input image with a final layer producing output features.
[0050] Although the above examples use a convolutional encoder / deconvolutional decoder architecture as an example, any, some, or all of the camera network 220, radar network 222, and lidar network 224 may have any other suitable architecture. For example, the encoder portions of the networks may include a recurrent neural network, a long short-term memory (LSTM) neural network, a fully connected neural network, and / or some combination of such networks. In some implementations, any, some, or all of the camera network 220, radar network 222, and lidar network 224 may have a Transformer-based architecture with the encoder portion of the network including one or more self-attention blocks and the decoder portion of the network including one or more cross-attention blocks (in addition to the self-attention blocks). In some implementations, the camera network 220 and / or radar network 222 may include an encoder, while the decoder may be implemented as part of the BEV backbone 230 (see FIG. 3B).
[0051] The camera network 220 computes a camera feature vector F that characterizes the visual appearance (captured by the camera image 202) of the portion of the environment associated with point x,y of the BEV grid 305 at time t. CSimilarly, the radar network 222 may generate a radar feature vector F (x,y;t) 302 that characterizes the presence or absence of reflective objects (captured by the radar image 204) in the same portion of the environment associated with the same time t. R (x,y;t) 304. The lidar network 224 can generate a lidar feature vector F that characterizes the type of lidar reflection at point x,y of the BEV grid 305 at time t, given the context provided by various other lidar reflection points x′,y′. L (x,y;t) can be generated. For those positions on the BEV grid 305 where no indicia of various objects of interest are detected, the respective camera feature vectors F C (x,y;t) 302 (and similarly, the radar feature vector F R (x,y;t) 304 and the lidar feature vector F L (x,y;t) 306) may have a zero value (or a value close to zero). For illustrative purposes, a single camera feature vector F C (x,y;t) 302 (and similarly, a single radar feature vector F R (x,y;t) 304 and a single lidar feature vector F L While the BEV grid (x,y;t) 306) is illustrated in FIG. 3A, feature vectors for individual points x,y can be defined for the entire BEV grid 305, and the camera feature tensor F C (t)={F C (x,y;t)}312 yields the feature tensor FT C (t) 312 can have dimensions X×Y×C, where X and Y are the dimensions of the BEV grid 305 and C is the dimension of the context, which can be empirical (based on testing), for example, as part of the camera network 220 architecture (defined before training).
[0052] Similarly, each individual radar feature vector 304 is represented as a radar feature tensor F R (t)={F R (x,y;t)}314. The radar feature tensor FTR (t) 314 has the same BEV dimensions X and Y, and the camera feature tensor FT C (t) The dimension C of the radar context is different from the dimension C of the 312 context. R For example, the dimension C of the context of the camera feature vector / tensor may be larger than the dimension C of the radar feature vector / tensor, considering the more diverse types of visual context that the camera image 202 captures compared to the radar image 204. R It may have a higher dimension than
[0053] Similarly, the individual lidar feature vectors 306 are expressed as a function of the lidar context dimension C L Lidar feature tensor FT with L (t)={F L (x, y; t)} 316. In some implementations, the dimensions of the rider context C L is the dimension C of the radar context R and may be larger than the dimension C of the camera context.
[0054] In one example implementation, to generate the camera feature vector 302 (and, correspondingly, the camera feature tensor 312), the camera network 220 can perform a lift transform 310 to compensate for the lack of explicit distance (depth) information in the camera image 202. In some implementations, the lift transform 310 can be performed in two stages. During the first stage, the camera network 220 computes the feature vector f(c) for the pixels w, h of the camera (perspective view). w,h where c enumerates various context components (e.g., c∈[1,C]) and generates a feature vector f(c) with depth information that may be provided by separate outputs of the camera network 220. w,h For example, depth information can be expressed as a distribution P(d) of the probability that a given pixel w, h represents an object located at a distance d from the camera. w,h The lift transform 310 can then produce the corresponding depth distribution P(d) w,h Using each feature vector f(c) w,hCalculate the Cartesian product of the depth-extended feature vector f(c,d) for pixels w and h. w,h =f(c) w,h ×P(d) w,h can be generated.
[0055] Then, the depth-extended feature vector f(c,d) calculated for each pixel w,h can be combined (still in perspective) into a unified feature tensor for the entire camera image 202: {f(c,d) w,h}→ft(c,d,w,h). The depth-dilated feature tensor ft(c,d,w,h) has dimensions C×D×W×H, where W×H are the dimensions of the camera image 202 (in pixels) and D is the dimension of the depth granularity. For example, the distance d is spaced apart by D, Δd1, Δd2... Δd D can be discretized between the interval Δd i do not need to be of equal size, but rather a distance, e.g., Δd1<Δd D In some implementations, the last interval Δd D may extend from a certain distance (e.g., 100m, 200m, etc.) to an infinite distance.
[0056] The second stage of the lift transformation 310 is, for example, the transformation from Cartesian coordinates, ft(c,d,w,h) to FT CThe transformation d,w,h→x,y,z may be a projective transformation parameterized by the focal length of the camera, the direction of the camera's optical axis, and / or other similar parameters. In cases where the camera images 202 are acquired by multiple cameras (or cameras with rotating optical axes), the transformation d,w,h→x,y,z may include multiple projective transformations, e.g., with separate transformations used for pixels w,h of different cameras (or by the same camera pointing in different directions).
[0057] Using a 2D mapping, we project the feature tensor expressed in the new coordinates, ft(c,d,w,h) → ft(c,x,y,z), and take the sum (or average, weighted average, or aggregation) over the different heights z to obtain the feature tensor FT C (c,x,y)312, e.g., FT C (c,x,y)=Σ i →ft(c,x,y,z i ) can be obtained. In some implementations, the coordinate z i The summation for different coordinates z i The different weights w assigned to i :FT C (c,x,y)=Σi→w i ft(c,x,y,z i ), e.g., with a larger weight w i is assigned to pixels that image objects within a certain elevation above the ground (e.g., up to a few meters), and lower weights are assigned to other elevations (e.g., to eliminate spurious objects such as tree branches, electrical cables, etc. that do not obstruct vehicle movement). The use of the lift transform 310, and the BEV grid 305, eliminates distortions associated with the perspective view of the camera sensor.
[0058] The radar feature vectors 302 output by the radar network 222 may be mapped to the same BEV grid 305. In the case of a radar network output, the distances to the various reflecting pixels in the radar image 204 are plotted as a distribution P(d) over distance Δd j can be known precisely (as part of the radar data) so that it can be one for certain intervals of and zero for other intervals.
[0059] In some implementations, the lidar feature vector 306 may be mapped to the BEV grid 305 in a manner similar to the mapping of the radar feature vector 304. In some implementations, point-to-grid mapping 311 may be applied to the lidar feature vector. More specifically, the cloud of lidar return points may be a dense cloud (e.g., with more points than in the radar cloud). Each position x,y of the BEV grip may be associated with one or more lidar return points (e.g., points closer to this position x,y than other elements x',y' of the grid). The lidar feature vector F L (x,y;t) may be aggregated (eg, averaged) from the individual features associated with the individual return points associated with the positions x,y.
[0060] The camera features, radar features, and / or lidar features can then be aggregated (e.g., concatenated) to obtain a joint feature. For example, feature vector aggregation can be performed for each BEV grid location x,y to obtain a joint feature vector 320: [F C (x,y,t),F R (x,y,t),F L (x,y,t)]→F(x,y,t). The set of integrated feature vectors for the various BEV grid locations represents the integrated feature tensor F T (t) = {F(x,y;t)}. Equivalently, the integrated feature tensor F T (t) 330 is the combination F T (t) = [F T C (t),FTR (t),FT L (t)].
[0061] 3B illustrates a second stage 350 of the operation of the RDA detection model. In some implementations, as shown in FIG. 3B, the integrated feature tensors associated with multiple timestamps are stored in a feature stack 340, e.g., {FT(t1), FT(t2), ... FT(t M )}. (For simplicity, the case of M=3 is shown in FIG. 3B, but the number of timestamps need not be limited.) M may be selected by applying a sliding window to the images of the input data 301. For example, during the processing of the next input data 301 by the RDA detection model 132, at time t 1+S ,t 2+S ,…t M+S The feature stack 340 {FT(t 1+S ),FT(t 2+S ),…FT(t M+S )} may be generated with a suitably chosen stride S, e.g., S=1, 2, etc. In some implementations, the stride S may be set based on the time of E2E processing of the input data 301 to avoid idling of the RDA detection model 132 while still ensuring that the RDA detection model 132 is not blocked, e.g., by providing new input data before previous input data has been converted and processed into the feature stack 340. For example, if a typical time for E2E processing by the RDA detection model 132 is τ, then the stride may be set to t j+S -t j It can be set so that τ≈τ.
[0062] The generated feature stack 340 can be processed by the BEV backbone 230. In some implementations, the BEV backbone 230 can include an encoder (following individual encoders in the camera network 220, radar network 222, and / or lidar network 224) that further encodes the feature stack 340. In some implementations, the BEV backbone 230 can include both an encoder and a decoder. In some implementations, the BEV backbone 230 can be a decoder, while an encoder can be implemented as part of the camera network 220, radar network 222, and / or lidar network 224.
[0063] The BEV backbone 230 can provide intermediate outputs to multiple road classification heads 232-n, which classify various locations x,y among one or more classes. Each classification head 226-n can output a different type of information about the location x,y. For example, the road identification head 232-1 can generate a probability (e.g., a floating-point value) P that a given location x,y is associated with a drivable road (e.g., a lane, an intersection, etc.). R (x,y) can be output. R (x,y) indicates that the location is associated with a non-traversable part of the environment, e.g., a sidewalk, a tree, a structure, a building, etc. In some implementations, the probability P R (x,y) and 1-P R (x,y) may be output by the final binarization (e.g., sigmoid) classification layer of the road graph mapping head 232-1.
[0064] Similarly, the RDA detection head 232-2 detects the probability P RDA (x,y) and the probability of the location being in a drivable (normal, obstacle-free) part of the road, 1-P RDA (x,y) can be output. RDA (x,y) is an empirically set threshold, e.g., P T= 0.4, 0.5, 0.6, 0.75, etc., the location x, y can be classified as a restricted travel location where vehicles should not enter. In some implementations, the RDA detection head 232-2 can use an inverse lift transform 306 to map the BEV point x, y to the viewpoint coordinates of the camera pixel w, h to mark the RDA in the camera image. In some implementations, the RDA detection head 232-2 can further output a label 340 that identifies the type of RDA. For example, a first label can identify whether travelability through the location is restricted due to an emergency (fire, police, accident, environmental hazard, etc.). A second label can identify whether the location is associated with a construction zone, a third label can identify whether the location is occupied by a vehicle or other mobile object, and a fourth label can identify whether the location is occupied by a vulnerable road user (VRU), such as a pedestrian, bicyclist, animal, etc.
[0065] The output of the road classification head 232-n, e.g., P R (x,y;t i ), P RDA (x,y;t i ), label(x,y;t i ) at a particular timestamp t i Each set of identifications may be affected by noise, such as measurement noise, occlusion, lighting fluctuations, etc. To eliminate or reduce such noise, the output may be subjected to temporal smoothing 360 over a certain period of time, such as T timestamp (which may be the most recent timestamp). For example, a simple averaging may be performed on the output of RDA detection head 232-2 (and similarly other outputs),
number
number
[0066] In some implementations, identifying the drivable area 234 may include one or more post-processing operations, for example, as indicated by the dashed callout boxes. More specifically, contours representing the boundaries between the drivable area and the no-drivable area may be extracted. For example, contour extraction 370 may identify a set of locations x,y, where:
number
[0067] The road classification heads 232-n shown in FIG. 3B are intended to serve as examples so that various other classification heads can be defined and trained depending on the particular driving environment of the vehicle.
[0068] The identified drivable area 234 can be provided to a tracker / planner 260, which can track the movement of the identified area over time relative to the vehicle using, for example, a suitable motion tracker such as a Kalman filter. The tracker / planner 260 can make further driving decisions taking the drivable area 234 into account. For example, in an autonomous driving system (or a driver assistance system operating in an autonomous or semi-autonomous mode), the tracker / planner 260 can identify and implement a driving path for the vehicle that is consistent with the location of the drivable area 234. In a driver assistance system, the tracker / planner 260 can provide a representation of the RDA and / or the drivable area 234 to the driver, for example, via a dashboard display.
[0069] In some implementations, the trajectory heuristic 380 can generate a predicted trajectory 382, e.g., a driving path that an experienced driver is likely to select in the current driving environment, once the drivable area 234 has been determined. More specifically, the trajectory heuristic 380 can search for lanes of travel that remain unobstructed when approaching and leaving the RDA. For example, even at an intersection where police have blocked traffic due to an accident, the right-most lane may remain drivable. If such lanes do not exist, the trajectory heuristic 380 can search for drivable lanes before the RDA and for drivable lanes after the RDA; for example, the right-most lane may be available before the intersection, and the left-most lane is available after the intersection. The trajectory heuristic 380 can then chart a predicted trajectory 382 that first follows the left-most lane, then crosses over to the right-most lane. If all travel lanes in the direction of the intended path of travel are impassable, trajectory heuristic 380 may select a drivable area 234 that minimizes deviation from the intended path of travel, for example, trajectory heuristic 380 may select a right turn if the path ahead is blocked, a left turn if both the path ahead and a right turn are blocked, or a U-turn if all three possibilities are unavailable. Various other heuristics may be coded as part of trajectory heuristic 380.
[0070] FIG. 4 is a schematic diagram of an example driving environment 400 of a vehicle 402 deploying an RDA detection model 132 for identifying no-drive zones, according to some implementations of the present disclosure. As shown in FIG. 4, the vehicle 402 is traveling in the left lane of a two-lane road and processes data from multiple sensing modalities (e.g., camera, radar, lidar, etc.) that capture images of the driving environment 400. The driving environment 400 includes the scene of an emergency, for example, caused by a fire in a building 404. A police car 406 is blocking movement ahead of the vehicle 402. The oncoming lane is similarly blocked by a fire truck 408. A fire hose 410 is preventing traffic from turning right. The RDA detection model 132 can process data collected by the vehicle's 402's sensing system and output a drivable zone and an RDA, for example, as disclosed in conjunction with FIGS. 3A and 3B . More specifically, the RDA detection model 132 detects the presence of some or all of a police car 406, a fire engine 408, a fire hose 410, emergency responders 412, and / or other indicia of RDA in the driving environment 400. The RDA detection model 132 may identify a contour 414 (shown by a dotted line) as the boundary of the RDA. The trajectory heuristic 380 (see FIG. 3B ) may further identify the shaded area as a predicted trajectory 382, e.g., a trajectory that an experienced human driver would select through the driving environment 400. In some implementations, the trajectory heuristic 380 may be replaced by a trained trajectory prediction model.
[0071] FIG. 5 illustrates an example architecture of another E2E perception model 500 capable of identifying RDAs and predicting the trajectory of an autonomous vehicle in a driving environment, according to some implementations of the present disclosure. The E2E perception model 500 deploys a trained trajectory prediction model 580 instead of the trajectory heuristic 380 of FIGS. 3A-3B. Some parts of the E2E perception model 500, such as the temporal smoothing 360, are not explicitly shown in FIG. 5. The trajectory prediction model 580 uses the output of various road graph classification heads, such as the probability P that a particular location x,y belongs to various roads in the driving environment. R(x,y) and the probability P that the location is associated with the RDA of the driving environment. RDA (x,y) can be used. In some implementations, as indicated by the dashed arrows (skip connections), additional inputs to the trajectory prediction model 580 can include one or more feature tensors output by any, some, or all of the camera network 220, radar network 222, lidar network 224, and / or other networks not shown in FIG. 5 (e.g., an audio network).
[0072] The trajectory prediction model 580 calculates the probability P of an expert driver driving through various positions x, y while passing through the current driving environment. D (x,y). The predicted trajectory 382 can then output the probability P of driving various positions x,y. D For (x,y), the maximum probability P D The trajectory may be selected based on the location x, y having (x, y). For example, for the driving environment of Figure 4, the trajectory prediction model 580 may determine that the expert driver will make a left turn and follow trajectory 382.
[0073] In some implementations, the trajectory prediction model 580 can be trained separately from the networks 220-224 and / or the BEV backbone 230. For example, after the RDA detection model 132 has been trained to output reliable road graph mapping and RDA / driveable area predictions, the trajectory prediction model 580 can be trained using additional ground truth, including trajectories selected by experienced human drivers during past driving missions. In some implementations, the trajectory prediction model 580 can be trained together (end-to-end) with the RDA detection model 132. For example, such training ground truth may include both road graph / RDA / driveable area annotations from past missions (e.g., sensing logs, maps, etc.) and trajectories selected by human drivers. In some implementations, such training ground truth may include trajectories selected by human drivers in past missions, but not road graph / RDA / driveable area annotations.
[0074] FIG. 6 illustrates an example architecture of a multi-stage E2E perception model 600 capable of identifying RDA and predicting an autonomous vehicle's trajectory in a driving environment, according to some implementations of the present disclosure. In the first stage of the E2E perception model 600, the camera network 220 and radar network 222 process data from their respective modalities and generate feature vectors, as disclosed in connection with FIG. 3A . The camera feature vectors and radar feature vectors can be processed by the BEV backbone 230. The output of the BEV backbone 230 can be processed by one or more road classification heads, such as road graph mapping 232-1. The output of the road graph mapping 232-1, along with the output of the BEV backbone 230, can be used as input to the lidar network 224, which implements the second stage of the E2E perception model 600. The output of the lidar network 224 can then undergo temporal smoothing 360 and drivable area identification, as disclosed in connection with FIG. 3B , for example.
[0075] In some implementations, additional neural networks, such as an emergency vehicle (EV) and construction detector 610, can be deployed to identify driving situations where drivability is likely to be restricted. The EV / construction detector 610 can be a separately trained model that processes one or more modalities of input data 301, including, but not limited to, any, some, or all of the camera imagery 202, radar imagery 204, lidar imagery 206, audio data 208, etc., and generates an identification of a target scenario or multiple target scenarios. A target scenario may include, by way of example and not limitation, the presence of emergency vehicles, warning tape, fire hoses, smoke bombs, lights, barricades, emergency responders, construction workers, etc. If indicators of a target scenario are not detected, the EV / construction detector 610 can generate a null output. The null output can indicate that no processing by the RDA detection model 132 will occur. If one or more of the target scenarios are detected, the EV / construction detector 610 can generate an indication of the corresponding scenario that can be used to trigger operation of the RDA detection model 132. As a result, computational and memory resources of the vehicle's perception system are not used unless there is an actual need for such processing.
[0076] Although not explicitly shown, the EV / construction detector 610 can be used in a similar manner to the E2E perception model shown in FIGS. 3A-3B and 5.
[0077] The various networks of the E2E perception model (including the RDA detection model), e.g., camera network 220, radar network 222, lidar network 224, BEV backbone 230, and road classification head 232, whose architectures and operations are disclosed above in conjunction with Figures 3A-3B and 5-6, can be trained together using suitable ground truth data. Training of the E2E perception model can be performed using supervised learning and imitation learning techniques. For example, supervised learning can be used to train the road graph mapping 232-1 head and the RDA detection 232-2 head. More specifically, sensory data collected during a driving mission (autonomous or driver-controlled) can be annotated by a human developer. The developer can identify at least one RDA, e.g., an emergency site, construction zone, etc., and further identify portions of the sensory data with timestamps associated with past vehicles that have traveled near the RDA. The developer can then use a graphical user interface (GUI) and suitable selection tools (mouse, touchpad, etc.) to mark areas in the camera image (and / or lidar / camera image) of the driving environment. During training, the E2E perception model generates training outputs, e.g., predicted RDA probability maps P RDA (x,y;t) (and similarly, the road graph map P R (x,y;t)) can be generated using an appropriate loss function, e.g., the binary cross-entropy loss function, to generate the predicted RDA probability map P RDAThe difference between (x,y;t) and the RDA (binary) ground truth annotated by the developer can be quantified. The difference quantified in the loss function can be used with various techniques, such as backpropagation and gradient descent, to modify parameters (e.g., weights and biases) of various networks in the E2E perception model, such as the camera network 220, radar network 222, lidar network 224, BEV backbone 230, and road classification head 232. During such E2E training, the parameters of any of these networks can be changed while training of other networks is still in progress (although it is not necessary to change the parameters of all networks during each training epoch).
[0078] In some implementations, imitation (self-supervised) learning can be used to train the trajectory prediction model 580 (see FIG. 5 ). More specifically, the trajectory prediction model 580 can output a map of driving probabilities PD(x,y) that mimic actual driving maneuvers performed by expert human drivers during past missions, e.g., determined from logs of past missions. The trajectory prediction model 580 can be trained using a binary cross-entropy loss function (or other suitable loss function) applied to various positions x,y. In some implementations, the trajectory prediction model 580 can be trained after other networks of the RDA detection model 132 have been trained. For example, during a first training phase, any, some, or all of the camera network 220, radar network 222, lidar network 224, BEV backbone 230, and road classification head 232 can first be pre-trained to identify correct road graph mapping and / or RDA detection (e.g., using human-annotated ground truth). During the second training phase, the trajectory prediction model 580 may be trained to mimic the driving behavior of a human driver, for example, as known from the same past driving mission. In other implementations, the training of the trajectory prediction model 580 may be performed simultaneously with the training of any, some, or all of the camera network 220, the BEV backbone 230, and the sign classification head 226.
[0079] Training of any, some, or all of the camera network 220, radar network 222, lidar network 224, BEV backbone 230, road classification head 232, trajectory prediction model 580, and / or other networks of the E2E perception model may include dropout techniques. More specifically, during a portion of a training epoch, sensory data of one or more modalities may be dropped (replaced with zeros) to train an efficient RDA and driving path detection E2E perception model that does not overly rely on any single sensory modality. For example, during a first training epoch, the E2E perception model may be trained to process camera images 204, radar images 204, and lidar images 206 (and possibly audio data 208). During a second training epoch, the E2E perception model may be trained to process camera images 204 and radar images 204 (replacing lidar images 206 with zero input). During the third training epoch, the E2E perception model can be trained to process camera image 202 and lidar image 206 (replacing radar image 204 with zero inputs). During the fourth training epoch, the E2E perception model can be trained to process radar image 204 and lidar image 206 (replacing camera image 202 with zero inputs), and so on.
[0080] After dropout training, the trained E2E perception model can be deployed to vehicles with three (or more) sensing modalities: camera, radar, and lidar (e.g., audio). Dropout training enables the E2E perception model to operate even in situations where one or more sensing modalities are temporarily disabled, for example, when the lidar transceiver is impaired by dirt, fog, or other environmental factors. In some implementations, the dropout-trained E2E perception model can be deployed to vehicles that do not have one or more sensing modalities, for example, in one example, a vehicle that has camera and radar sensors but no lidar sensor.
[0081] FIG. 7 illustrates an example method 700 for deploying an end-to-end perception model for RDA detection and navigation in a driving environment according to some implementations of the present disclosure. A processing device having one or more processing units (CPUs), one or more graphics processing units (GPUs), one or more parallel processing units (PPUs), and a memory device communicatively coupled to the CPUs, GPUs, and / or PPUs may implement method 700 and / or each of its individual functions, routines, subroutines, or operations. Method 700 may be directed to systems and components of a vehicle. In some implementations, the vehicle may be an autonomous vehicle. In some implementations, the vehicle may be a driver-operated vehicle with a driver assistance system, e.g., a level 2 or level 3 driver assistance system, that provides limited assistance with certain vehicle systems (e.g., steering, braking, acceleration, etc.) or under limited driving conditions (e.g., highway driving). The processing device performing method 700 may execute instructions issued by the perception and planning system 130 of FIG. 1, and more specifically, the RDA detection model 132, during vehicle operation. In certain implementations, a single processing thread may perform method 700. Alternatively, two or more processing threads may perform method 700, with each thread performing one or more individual functions, routines, subroutines, or operations of the method. In an example embodiment, the processing threads performing method 700 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads performing method 700 may execute asynchronously with respect to each other. Some operations of method 700 may be performed in an order different from the order shown in FIG. 7. Some operations of method 700 may be performed simultaneously with other operations. Some operations may be optional.
[0082] At block 710, the method 700 may include obtaining, using a sensing system of the vehicle, a first set of images of the environment, a second set of images of the environment, and a third set of images of the environment. The first set of images may include one or more perspective camera images of the environment (e.g., camera image 202 of FIG. 3A). The second set of images may include one or more radar images of the environment (e.g., radar image 204 of FIG. 3A), and the third set of images may include one or more lidar images of the environment (e.g., lidar image 206 of FIG. 3A).
[0083] At block 720, method 700 may include generating one or more camera features (e.g., camera feature tensor 312) characterizing the first set of camera images using a first neural network (e.g., camera network 220). As illustrated in top callout block 722, generating the one or more camera features may include mapping the one or more camera features from a viewpoint coordinate system to a coordinate system associated with the Earth's surface (e.g., BEV grid 305 of FIG. 3A).
[0084] At block 730, the method 700 may include using a second neural network (e.g., radar network 222) to generate one or more radar features (e.g., radar feature tensor 314 of FIG. 3A) that characterize the second set of radar images.
[0085] At block 740, the method 700 may include using a third neural network (e.g., the lidar network 224) to generate one or more lidar features (e.g., the lidar feature tensor 316 of FIG. 3A) that characterize the third set of lidar images.
[0086] At block 750, the method 700 may include processing one or more camera features, one or more radar features, and one or more lidar features to obtain an indication of a reduced drivability area (RDA) in the environment. In some implementations, the indication of the RDA may include multiple elements, each element of the multiple elements being mapped to a corresponding area of a plurality of areas of the environment and associated with a likelihood that each area of the plurality of areas belongs to the RDA.
[0087] As illustrated in the callouts at the bottom of FIG. 7 , processing the one or more camera features, one or more radar features, and one or more lidar features may include using a backbone neural network and one or more classification heads in block 752. In some implementations, the backbone neural network may include a BEV neural network (e.g., BEV backbone 230 of FIGS. 3A and 3B ) and one or more classification neural networks (e.g., road classification head 232 of FIG. 3B ). In some implementations, the first neural network, the second neural network, the third neural network, the backbone neural network, and the one or more classification heads are trained together, e.g., end-to-end. In some implementations, the one or more classification heads may include a driving trajectory classification head that outputs a target trajectory for the vehicle, where the target trajectory avoids RDA.
[0088] In some implementations, the driving trajectory classification head, the backbone neural network, and one or more of the first neural network, the second neural network, or the second neural network are trained using ground truth comprising one or more trajectories of human-operated vehicles navigating one or more respective past driving missions, each encountering at least one RDA.
[0089] In some implementations, the first neural network, the second neural network, the third neural network, the backbone neural network, and the one or more classification heads may be trained using one or more dropout training epochs, each dropout training epoch having the output of at least one of the first neural network, the second neural network, or the third neural network replaced with a null output.
[0090] In some implementations, the one or more camera features to be processed are a time series of camera features (e.g., FT C (t1), FT C (t2), F.T. C (t3), etc. One or more radar features may also include a time series of camera features (e.g., FT R (t1), FT R (t2), F.T. R The one or more lidar features may also include a time series of lidar features (e.g., FT L (t1), FT L (t2), F.T. L (t3, etc.). In some implementations, processing the time series of camera features, the time series of radar features, and the time series of lidar features may be performed simultaneously.
[0091] In some implementations, obtaining an indication of RDA in block 754 may include temporally smoothing a plurality of candidate indications of RDA using multiple sets of sensed data. More specifically, the sensing system may acquire second sensed data associated with a second time set. The second time set may include times t4, t5, t6, etc. that are all distinct from times in the first time set t1, t2, t3. In some implementations, the second time set may include times t3, t4, t5, etc. that partially overlap with the first time set t1, t2, t3. The second sensed data may include a fourth set of camera images of the environment, a fifth set of radar images of the environment, and a sixth set of lidar images of the environment. Method 700 may include generating one or more additional camera features (using a first neural network) that characterize the fourth set of camera images, generating one or more additional radar features (using a second neural network) that characterize the fifth set of radar images, and generating one or more additional lidar features (using a third neural network) that characterize the sixth set of lidar images. The operations of block 754 may include processing the one or more additional camera features, the one or more additional radar features, and the one or more additional lidar features to obtain an additional representation of the RDA. The representation of the RDA and the additional representation of the RDA may then be processed (e.g., averaged, aggregated, etc.) to obtain a temporally smoothed representation of the RDA. For example, the probability P RDA (x,y;t), P RDA Two (or more) maps, such as (x,y;t'), may be averaged to obtain an aggregated probability, for example, after mapping two or more probabilities to the same point x,y relative to the ground (e.g., of the BEV grid) (this may include performing a coordinate transformation from the reference frame of the moving vehicle to coordinates associated with the stationary ground). In some implementations, to obtain the representation of the RDA, method 700 may include, at block 756, obtaining a set of predicted representations of the RDA and eliminating one or more duplicate representations of the RDA from the set of predicted representations of the RDA.
[0092] In some implementations, the method 700 also associates the temporally smoothed representation of the RDA with a roadmap (e.g., a static, high-definition map of drivable roads, streets, intersections, lanes, etc.) to identify one or more lanes of travel that are blocked to traffic and are mapped to the roadmap.
[0093] In some implementations, the method 700 continues at block 760 with having a driving control system of the autonomous vehicle select a driving path for the autonomous vehicle taking into account the RDA (and / or the driving lanes in the roadmap), e.g., an indication of a driving path that avoids the RDA.
[0094] FIG. 8 illustrates a block diagram of an example computing device 800 capable of training and / or deploying E2E perception models, including one or more RDA detection models, that use a combination of camera, radar, and / or lidar imagery for accurate identification and navigation of restricted areas in a driving environment, according to some implementations of the present disclosure. The example computing device 800 may be connected to other computing devices within a LAN, an intranet, an extranet, and / or the Internet. The computing device 800 may operate in the capacity of a server in a client-server network environment. The computing device 800 may be a personal computer (PC), a set-top box (STB), a server, a network router, a switch, or a bridge, or any device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the device. Furthermore, while only a single example computing device is shown, the term “computer” shall also be considered to include any group of computers that, individually or jointly, execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.
[0095] The exemplary computing device 800 may include a processing device 802 (also referred to as a processor or CPU), a main memory 804 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), a static memory 806 (e.g., flash memory, static random access memory (SRAM), etc.), and a secondary memory (e.g., a data storage device 818), which may communicate with each other via a bus 830.
[0096] Processing device 802 (which may include logic processing 803) represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, processing device 802 may be a complex instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor executing other instruction sets, or a processor executing a combination of instruction sets. Processing device 802 may also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. According to one or more aspects of the present disclosure, processing device 802 may be configured to execute instructions implementing method 700 of deploying an E2E perception model for detection and navigation of RDA in a driving environment.
[0097] The exemplary computing device 800 may further comprise a network interface device 808, which may be communicatively coupled to a network 820. The exemplary computing device 800 may further comprise a video display 810 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 812 (e.g., a keyboard), a cursor control device 814 (e.g., a mouse), and an audio signal generating device 816 (e.g., a speaker).
[0098] The data storage device 818 may include a computer-readable storage medium (or more specifically, a non-transitory computer-readable storage medium) 828 having stored thereon one or more sets of executable instructions 822. According to one or more aspects of the present disclosure, the executable instructions 822 may include executable instructions that implement the method 700 of deploying an E2E perception model for detection and navigation of RDA in a driving environment.
[0099] The executable instructions 822 may also reside, completely or at least partially, within the main memory 804 and / or within the processing device 802 during execution thereof by the exemplary computing device 800, with the main memory 804 and the processing device 802 also constituting computer-readable storage media. The executable instructions 822 may further be transmitted or received over a network via the network interface device 808.
[0100] While computer-readable storage medium 828 is illustrated in FIG. 8 as a single medium, the term "computer-readable storage medium" should be considered to include a single medium or multiple media (e.g., centralized or distributed databases, and / or associated caches and servers) that store one or more sets of operating instructions. The term "computer-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for execution by a machine, causing the machine to perform any one or more of the methods described herein. Thus, the term "computer-readable storage medium" should be considered to include, but not be limited to, solid-state memory, and optical and magnetic media.
[0101] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, understood to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0102] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, and as will be apparent from the discussion that follows, throughout the description, discussions using terms such as "identify," "determine," "store," "adjust," "produce," "return," "compare," "generate," "stop," "load," "copy," "inject," "replace," "perform," or the like, will be understood to refer to the actions and processes of a computer system, or similar electronic computing device, that manipulate and transform data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display device.
[0103] Examples of the present disclosure also relate to an apparatus for performing the methods described herein. This apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer system selectively programmed by a computer program stored within the computer system. Such a computer program may be stored on a computer-readable storage medium, such as, but not limited to, any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random-access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0104] The methods and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems appears as set forth in the description below. Additionally, the scope of the present disclosure is not limited to any particular programming language. It will be understood that a variety of programming languages can be used to implement the teachings of the present disclosure.
[0105] It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other example embodiments will become apparent to those skilled in the art upon reading and understanding the above description. While the present disclosure describes particular examples, it will be recognized that the disclosed systems and methods are not limited to the examples described herein, but may be modified and practiced within the scope of the appended claims. Accordingly, the specification and drawings should be considered in an illustrative, and not a restrictive, sense. The scope of the present disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
Claims
1. 1. A system comprising:
1. A sensing system for a vehicle, the sensing system configured to obtain first sensory data associated with a first time set, the first sensory data comprising: a first set of camera images of the environment; a first set of radar images of the environment; and a sensing system including a first set of lidar images of the environment; A data processing system for the vehicle, the data processing system comprising: generating one or more camera features characterizing the first set of camera images using a first neural network (NN); generating one or more radar features characterizing said first set of radar images using a second neural network; generating one or more lidar features characterizing the first set of lidar images using a third neural network; and a data processing system further configured to process the one or more camera signatures, the one or more radar signatures, and the one or more lidar signatures to obtain an indication of reduced traf?c areas (RDAs) within the environment.
2. 2. The system of claim 1, wherein the representation of the RDA includes a plurality of elements, each element of the plurality of elements being mapped to a corresponding area of a plurality of areas of the environment, and each area of the plurality of areas being associated with a likelihood of belonging to the RDA.
3. To generate the one or more camera features, the data processing system: The system of claim 1 , further configured to map the one or more camera features from a viewpoint coordinate system to a coordinate system associated with the Earth's surface.
4. To process the one or more camera features, the one or more radar features, and the one or more lidar features, the data processing system: The system of claim 1 , configured to process the one or more camera features, the one or more radar features, and the one or more lidar features using a backbone NN and one or more classification heads.
5. the one or more camera features comprising a time series of camera features, the one or more radar features comprising a time series of radar features, and the one or more lidar features comprising a time series of lidar features, and for processing the one or more camera features, the one or more radar features, and the one or more lidar features, the data processing system: The system of claim 4 , configured to use the backbone NN to simultaneously process the camera feature time series, the radar feature time series, and the lidar feature time series.
6. The system of claim 4 , wherein the first NN, the second NN, the third NN, the backbone NN, and the one or more classification heads are trained jointly.
7. The system of claim 4 , wherein the one or more classification heads comprise a travel trajectory classification head that outputs a target trajectory for the vehicle, the target trajectory avoiding the RDA.
8. 8. The system of claim 7, wherein the driving trajectory classification head, the backbone neural network, and one or more of the first neural network, the second neural network, or the second neural network are trained using ground truth comprising one or more trajectories of a human-operated vehicle navigating a respective one or more past driving missions, each including at least one RDA.
9. 5. The system of claim 4, wherein the first NN, the second NN, the third NN, the backbone NN, and the one or more classification heads are trained using one or more dropout training epochs, each dropout training epoch replacing an output of at least one of the first NN, the second NN, or the third NN with a null output.
10. The sensing system is further configured to obtain second sensory data associated with a second time set, the second sensory data comprising: a fourth set of camera images of the environment; a fifth set of radar images of the environment; and a sixth set of lidar images of the environment; the data processing system, generating one or more additional camera features characterizing the fourth set of camera images using the first neural network; using the second neural network to generate one or more additional radar features characterizing the fifth set of radar images; using the third neural network to generate one or more additional lidar features that characterize the sixth set of lidar images; processing the one or more additional camera features, the one or more additional radar features, and the one or more additional lidar features to obtain an additional representation of the RDA; The system of claim 1 , further configured to process the representation of the RDA and the additional representation of the RDA to obtain a temporally smoothed representation of the RDA.
11. To obtain the representation of the RDA, the data processing system: obtaining a set of predicted representations of said RDA; The system of claim 1 , configured to eliminate one or more duplicate indications of the RDA from the set of predicted indications of the RDA to obtain the indication of the RDA.
12. the vehicle is an autonomous vehicle, and the data processing system comprises:
2. The system of claim 1, further configured to cause a driving control system of the autonomous vehicle to select a driving path for the autonomous vehicle taking into account the indication of the RDA.
13. 1. A method comprising: obtaining first sensory data associated with a first time set using a sensory system of the vehicle, the first sensory data comprising: a first set of camera images of the environment; a second set of radar images of the environment; and a third set of lidar images of the environment; and generating one or more camera features characterizing the first set of camera images using a first neural network (NN); generating one or more radar features characterizing said second set of radar images using a second neural network; generating one or more lidar features characterizing the third set of lidar images using a third neural network; and processing the one or more camera signatures, the one or more radar signatures, and the one or more lidar signatures to obtain an indication of reduced traf?c areas (RDA) within the environment.
14. 14. The method of claim 13, wherein the representation of the RDA includes a plurality of elements, each element of the plurality of elements being mapped to a corresponding area of a plurality of areas of the environment, and each area of the plurality of areas being associated with a likelihood of belonging to the RDA.
15. processing the one or more camera features, the one or more radar features, and the one or more lidar features; 14. The method of claim 13, comprising processing the one or more camera features, the one or more radar features, and the one or more lidar features using a backbone neural network (NN) and one or more classification heads, wherein the first NN, the second NN, the third NN, the backbone neural network, and the one or more classification heads are trained together.
16. 16. The method of claim 15, wherein the one or more classification heads comprise a trajectory classification head that outputs a target trajectory for the vehicle, the target trajectory avoiding the RDA, and wherein the trajectory classification head, the backbone neural network, and one or more of the first neural network, the second neural network, or the second neural network are trained using ground truth including one or more trajectories of a human-operated vehicle navigating a respective one or more past driving missions, each including at least one RDA.
17. 16. The method of claim 15, wherein the first NN, the second NN, the third NN, the backbone NN, and the one or more classification heads are trained using one or more dropout training epochs, each dropout training epoch replacing an output of at least one of the first NN, the second NN, or the third NN with a null output.
18. obtaining second sensory data associated with a second time set using the sensing system of the vehicle, the second sensory data comprising: a fourth set of camera images of the environment, a fifth set of radar images of the environment; and a sixth set of lidar images of the environment; and generating one or more additional camera features characterizing the fourth set of camera images using the first neural network; generating one or more additional radar features characterizing the fifth set of radar images using the second neural network; and generating one or more additional lidar features characterizing the sixth set of lidar images using the third neural network; and processing the one or more additional camera features, the one or more additional radar features, and the one or more additional lidar features to obtain an additional representation of the RDA; processing the representation of the RDA and the additional representation of the RDA to obtain a temporally smoothed representation of the RDA; 14. The method of claim 13, further comprising: correlating the temporally smoothed representation of the RDA with a roadmap to identify one or more blocked lanes of travel mapped to the roadmap.
19. the vehicle is an autonomous vehicle, and the method comprises:
14. The method of claim 13, further comprising causing a driving control system of the autonomous vehicle to select a driving path for the autonomous vehicle taking into account the indication of the RDA.
20. 1. An autonomous vehicle, comprising: a sensing system configured to acquire sensory data of a plurality of sensing modalities, wherein the plurality of sensing modalities are selected from at least a camera sensing modality, a radar sensing modality, or a radar sensing modality; 1. A data processing system comprising: generating one or more first features characterizing the sensory data of the first sensing modality using a first neural network (NN); generating one or more second features characterizing the sensory data of the second sensing modality using a second neural network; and a data processing system configured to process the one or more first features and the one or more second features using a third neural network to obtain an indication of reduced drivability areas (RDAs) in an environment of the autonomous vehicle, wherein the first NN, the second NN, and the third NN are jointly trained using training data for each sensing modality of the plurality of sensing modalities; and An operation control system, a driving control system configured to select a driving path for the autonomous vehicle taking into account the indication of the RDA.