Multiframe Temporal Aggregation and Dense Motion Estimation for Autonomous Vehicles

Multi-frame temporal aggregation and dense motion estimation using neural networks address the challenges of object detection and tracking in autonomous vehicles, enhancing accuracy and reducing computational demands for efficient object classification and tracking.

JP2026502808APending Publication Date: 2026-01-27WAYMO LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025530521
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-30
Publication Date
2026-01-27

AI Technical Summary

Technical Problem

Existing autonomous vehicle systems face challenges in accurately and efficiently detecting and tracking objects in their environment, particularly due to the limitations of individual sensing modalities like lidar and cameras, which require significant computational resources and maintenance, and struggle with mapping objects to a bird's-eye view for precise distance estimation.

Method used

Implementing multi-frame temporal aggregation and dense motion estimation using a set of neural networks trained on sensor data from autonomous vehicles, enabling robust object classification and tracking by converting sensor data into bird's-eye view feature sets, even in sensor dropout scenarios.

Benefits of technology

Enables rapid and accurate detection, identification, and tracking of objects with reduced computational overhead, allowing for efficient deployment on various platforms and improved safety and reliability in autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502808000001_ABST
    Figure 2026502808000001_ABST
Patent Text Reader

Abstract

The method includes acquiring, by a processing device, input data derived from a set of sensors associated with an autonomous vehicle (AV). The input data includes camera data and radar data. The method further includes extracting, by the processing device, a plurality of bird's eye view (BEV) feature sets from the input data. Each BEV feature set corresponds to a respective time step. The method further includes generating, by the processing device, an object flow of at least one object from the plurality of BEV feature sets. Generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation. The method further includes causing the processing device to modify a driving path of the AV in consideration of the object flow.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification relates generally to systems and applications for detecting and classifying objects, and particularly to autonomous vehicles and vehicles incorporating driver assistance technologies. More specifically, this specification relates to multi-frame temporal aggregation and dense motion estimation for autonomous vehicles (AVs) for faster and more resource-efficient detection of objects, including, but not limited to, vehicles, pedestrians, bicyclists, animals, etc. [Background technology]

[0002] Autonomous (fully or partially self-driving) vehicles (AVs) operate by sensing the external environment with a variety of electromagnetic (e.g., radar and optical) and non-electromagnetic (e.g., sound and humidity) sensors. Some autonomous vehicles chart a driving path through the environment based on the sensed data. The driving path may be determined based on Global Navigation Satellite System (GNSS) data and roadmap data. The GNSS and roadmap data may provide information about static aspects of the environment (such as buildings, street layouts, road closures, etc.), while dynamic information (such as information about other vehicles, pedestrians, street lights, etc.) is obtained from simultaneously collected sensory data. The accuracy and safety of the driving path, as well as the accuracy and safety of the speed regime selected by the autonomous vehicle, depend on the timely and accurate identification of various objects present in the driving environment and the ability of the driving algorithms to process information about the environment and provide correct instructions to the vehicle controls and drivetrain.

[0003] The present disclosure is presented by way of example, and not by way of limitation, and may be more fully understood by reference to the following detailed description when considered in conjunction with the figures in which: [Brief explanation of the drawings]

[0004] [Figure 1] FIG. 1 illustrates components of an exemplary autonomous vehicle (AV), according to some implementations of the present disclosure. [Figure 2] FIG. 1 illustrates an example architecture of a portion of a perception system capable of efficient object detection and classification, according to some embodiments of the present disclosure. [Figure 3] FIG. 1 illustrates an example flow according to some implementations of the present disclosure. [Figure 4A] FIG. 1 illustrates an example multi-frame temporal aggregation architecture in accordance with some implementations of the present disclosure. [Figure 4B] FIG. 1 illustrates an example multi-frame temporal aggregation architecture in accordance with some implementations of the present disclosure. [Figure 4C] FIG. 1 illustrates an example multi-frame temporal aggregation architecture in accordance with some implementations of the present disclosure. [Figure 5A] FIG. 1 illustrates an example multi-frame dense motion estimation architecture in accordance with some implementations of the present disclosure. [Figure 5B] FIG. 1 illustrates an example multi-frame dense motion estimation architecture in accordance with some implementations of the present disclosure. [Figure 5C] FIG. 1 illustrates an example multi-frame dense motion estimation architecture in accordance with some implementations of the present disclosure. [Figure 6] FIG. 1 illustrates an example method for implementing multi-frame temporal aggregation and dense motion estimation for an autonomous vehicle (AV), in accordance with some implementations of the present disclosure. [Figure 7] FIG. 1 illustrates an example method for implementing multi-frame temporal aggregation for an autonomous vehicle (AV) in accordance with some implementations of the present disclosure. [Figure 8A] FIG. 1 illustrates an example method for implementing dense motion estimation for an autonomous vehicle (AV), according to some implementations of the present disclosure. [Figure 8B] FIG. 1 illustrates an example method for implementing dense motion estimation for an autonomous vehicle (AV), according to some implementations of the present disclosure. [Figure 9]FIG. 1 is a block diagram of an example computing device capable of implementing multi-frame temporal aggregation and dense motion estimation for an autonomous vehicle (AV) in accordance with some implementations of the present disclosure. Summary of the Invention

[0005] In one implementation, a method is disclosed that includes acquiring, by a processing device, input data derived from a set of sensors associated with an autonomous vehicle (AV). The input data includes camera data and radar data. The method further includes extracting, by the processing device, a plurality of bird's eye view (BEV) feature sets from the input data. Each BEV feature set corresponds to a respective time step. The method further includes generating, by the processing device, an object flow of at least one object from the plurality of BEV feature sets. Generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation. The method further includes causing the processing device to modify a driving path of the AV in consideration of the object flow.

[0006] In another implementation, a system is disclosed that includes a memory and a processing device operably coupled to the memory, the processing device configured to acquire input data derived from a set of sensors associated with an autonomous vehicle (AV). The input data includes camera data and radar data. The processing device is further configured to extract a plurality of bird's eye view (BEV) feature sets from the input data. Each BEV feature set corresponds to a respective time step. The processing device is further configured to generate an object flow of at least one object from the plurality of BEV feature sets. Generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation. The processing device is further configured to modify a driving path of the AV in consideration of the object flow.

[0007] In yet another implementation, a non-transitory computer-readable storage medium is disclosed having stored thereon instructions that, when executed by a processing device, cause the processing device to perform operations including obtaining input data derived from a set of sensors associated with an autonomous vehicle (AV). The input data includes camera data and radar data. The operations further include extracting a plurality of bird's eye view (BEV) feature sets from the input data. Each BEV feature set corresponds to a respective time step. The operations further include generating an object flow of at least one object from the plurality of BEV feature sets. Generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation. The operations further include modifying a driving path of the AV in consideration of the object flow. DETAILED DESCRIPTION OF THE INVENTION

[0008] While various embodiments may be described below, using autonomous driving systems and driver assistance systems as examples for illustration, it should be understood that the techniques and systems described herein may be used to track objects in a wide range of applications, including aerology, marine applications, traffic control, animal control, industrial and academic research, public or personal safety, or in any other application where automatic detection of objects is advantageous.

[0009] In one example, for the safety of automated driving operations, it may be desirable to develop and deploy technology for rapid and accurate detection, classification, and tracking of various road users and other objects encountered on or near roadways, such as roadway obstacles, construction equipment, roadside structures, etc. Autonomous vehicles (as well as various driver assistance systems) can utilize several sensors to facilitate the detection of objects in the driving environment and determine the movement of such objects. Sensors typically include radio detection and ranging sensors (radar), light detection and ranging sensors (lidar), various types of digital cameras, sonar, and position sensors. Different types of sensors offer different, and often complementary, benefits. For example, radar and lidar emit electromagnetic signals (radio or optical signals) that reflect off objects and convey information that allows the distance to the object (e.g., from the signal's time of flight) and the object's velocity (e.g., from the Doppler shift in the signal's frequency) to be determined. Radar and lidar can cover a full 360-degree view, for example, by using a scanning transmitter of the sensing beam. The sensing beam may make numerous reflections that cover the driving environment in a dense grid of return points, each of which may be associated with the distance to the corresponding reflecting object and the radial velocity (the component of velocity along the line of sight) of the reflecting object.

[0010] Object identification and tracking methods use various sensing modalities, e.g., lidar, radar, camera, etc., to acquire images of an environment. The images can then be processed by a trained machine learning model to identify the location of various objects within the image (e.g., in the form of bounding boxes), the state of object motion (e.g., speed as detected by lidar or radar Doppler effect-based sensors), the type of object (e.g., vehicle or pedestrian), etc. Object motion (or any other evolution, such as splitting a single object into multiple objects) can be implemented by creating and maintaining tracks associated with specific objects.

[0011] Using multiple sensing modalities (e.g., lidar, radar, cameras) to acquire complementary data often improves the accuracy of object detection, identification, and tracking, but at the expense of significant costs for sensing hardware and processing software. For example, lidar sensors can provide valuable information about the distance to various reflective surfaces in the external environment. However, lidar sensors operate by actively probing the external environment using optical signals and are expensive optical and electronic devices that require significant maintenance and periodic calibration. Lidar returns (point clouds) must be processed, segmented into groups associated with distinct hypothesized objects, and matched with objects detected using other sensing modalities (e.g., cameras), which require additional processing and memory resources. On the other hand, cameras operate by passively collecting light (and / or infrared electromagnetic waves) emitted (or reflected) by objects in the environment and are significantly simpler and cheaper to design, install, and operate. As a result, various driver assistance systems that do not deploy lidar (for cost and maintenance reasons) typically employ one or more cameras. Cameras can also be more easily installed in a variety of fixed locations and used for traffic monitoring and control, public and private safety applications, etc. Cameras based on optical or infrared imaging technology have certain advantages over radar in that, while they allow for the detection of the distance to an object (and the object's speed), they operate in a range of wavelengths that inherently have lower resolution compared to cameras. Therefore, the ability to detect and identify objects based on camera images alone is beneficial.

[0012] However, cameras project the three-dimensional (3D) external environment onto a two-dimensional imaging surface (e.g., the camera's photodetector array), which can be flat or curved. This creates two related challenges. On the one hand, the distance to an object (often referred to as the object's depth in the image) is not immediately known (but can often be determined from the context of the imaged object). On the other hand, because camera images are subject to perspective distortion, the distance between objects varies depending on their depth, even if the number of pixels separating their images is the same. Furthermore, objects whose representations are close to each other may nevertheless be separated by a significant distance (e.g., a car and a pedestrian visible behind it). Some object detection machine learning techniques may attempt to map objects from a perspective view to a top-down view, also known as a bird's-eye view (BEV), in which objects are represented on a convenient manifold, such as a plane viewed from above, and characterized by a simple set of Cartesian coordinates. Object identification and tracking can then be performed directly within the BEV representation. The success of such techniques depends on accurately mapping objects to the BEV representation. This in turn requires accurate estimation of the distance to various objects, as misplacement of objects within the BEV can not only result in errors in ascertaining the distance to road users, but can also lead to the loss of important contextual information.

[0013] Motion data during AV operation may include flow data. For example, flow may be aligned to the latest time step (i.e., the latest frame) and represented by a two-dimensional (2D) displacement vector (u, v) for each cell calculated from successive three-dimensional (3D) bounding box labels. Flow may be captured using color coding, for example, where the color of an object identified in an image may specify the direction of the object's movement and the color saturation may specify the speed. For example, a green object may be an object determined to be moving north, a purple object may be an object determined to be moving south, an orange object may be an object determined to be moving west, and a teal object may be an object determined to be moving east.

[0014] One example of flow is base flow, which represents the movement of an object relative to the movement of an AV. Another example of flow is ego flow, which represents the movement of an AV (i.e., the movement of an object relative to an AV, assuming the object is stationary). Another example of flow is object flow, which represents the movement of an object. For example, object flow can represent the absolute displacement of an object relative to the ground. Object flow can be determined as the difference between base flow and ego flow. The state of movement of an object can be represented by the translational velocity of the object.

[0015] The object flow of an object may not be determined in an object-agnostic manner. This may make it difficult to classify the object into a known object class or an "unknown object" class. Therefore, it may be difficult for an AV to track the movement of an unknown object until the unknown object is safely outside the AV's operating environment.

[0016] Aspects and implementations of the present disclosure address these and other challenges of existing technologies by enabling methods and systems capable of implementing multi-frame temporal aggregation and dense motion estimation for AVs. As described in further detail herein, multi-frame temporal aggregation and dense motion estimation can be implemented by using multiple bird's eye view (BEV) feature sets, where each BEV feature set is extracted from data corresponding to a respective time step associated with each frame. For example, the time steps may include a current time step (time step t) and at least one previous time step (e.g., time step t-1). In particular, the disclosed technology provides an end-to-end perception model (EEPM) that may include a set of neural networks (NNs) trained to process input data from a set of sensors of an AV. For example, the input data may include camera data. Dense motion estimation can be used to determine (e.g., estimate) object flow (e.g., object motion) in an object-agnostic manner. This may enable classification of objects into any known object class and an "unknown object" class based on the object's motion patterns. The movement of an unknown object identified by the AV may be tracked by the AV until the object is safely outside the AV's operating environment.

[0017] The EEPM can be trained using sensor dropout scenarios in which some sensors are removed or not operational (e.g., at least one camera and / or at least one radar). For example, the right-facing camera can be removed, and information about objects in the portion of the space covered by the right-facing camera can be provided by other sensing modalities (e.g., lidar and / or radar sensors). Training scenarios can also include the complete dropout of certain sensing modalities, e.g., the dropout of the lidar data feed, so that all information about the environment is provided by the camera and radar. This trains the EEPM's output to be robust against failures of individual sensors and entire sensing modalities. Depending on the computational complexity and training sophistication, the EEPM can be used at various levels of driving automation, including level 2 driver assistance systems, level 3 contextual autonomous driving, level 4 primarily autonomous driving, level 5 fully autonomous driving, and other implementations.

[0018] Advantages of the described embodiments include (but are not limited to) rapid and accurate detection, identification, and tracking of objects in a manner that avoids the significant computational overhead of processing data from multiple sensing modalities. Because machine learning models trained and deployed as disclosed herein are capable of efficient object detection based on input data (e.g., camera data and radar data), EEPM models can be deployed on a variety of platforms (e.g., AVs), including systems with modest computational resources.

[0019] 1 is a diagram illustrating components of an exemplary autonomous vehicle (AV) 100, according to some implementations of the present disclosure. An autonomous vehicle may include a motor vehicle (such as a car, truck, bus, motorcycle, all-terrain vehicle, recreational vehicle, any specialized agricultural or construction vehicle, etc.), an aircraft (such as an airplane, helicopter, drone, etc.), a watercraft (such as a ship, boat, yacht, submarine, etc.), a spacecraft (a controllable object operating outside the Earth's atmosphere), or any other self-propelled vehicle (e.g., a robot, a factory or warehouse robotic vehicle, a sidewalk delivery robotic vehicle, etc.) capable of operating in an autonomous driving mode (without or reduced human input).

[0020] Vehicles such as those described herein may be configured to operate in one or more different driving modes. For example, in a manual driving mode, a driver may directly control acceleration, deceleration, and steering via inputs such as an accelerator pedal, brake pedal, steering wheel, etc. Vehicles may also operate in one or more autonomous driving modes, including, for example, a semi-autonomous or partially autonomous driving mode in which a human exercises some amount of direct or remote control over driving operations, or a fully autonomous driving mode in which the vehicle handles driving operations without direct or remote control by a human. These vehicles may be known by different names, including, for example, autonomous vehicles, automated vehicles, etc.

[0021] As described herein, in a semi-autonomous or partially autonomous driving mode, the vehicle assists with one or more driving maneuvers (e.g., steering, braking, and / or accelerating to perform lane centering, adaptive cruise control, advanced driver assistance systems (ADAS), or emergency braking), but the human driver is expected to maintain situational awareness of the vehicle's surroundings and supervise the assisted driving maneuvers. Here, the vehicle may perform all driving tasks in certain situations, but the human driver is expected to be responsible for assuming control as needed.

[0022] For simplicity and brevity, various systems and methods are described below in conjunction with autonomous vehicles, although similar technologies may be used in various driver assistance systems that fall short of fully autonomous driving systems. In the United States, the Society of Automotive Engineers (SAE) defines different levels of automated driving operation to indicate how much or how little control a vehicle has over the driving; however, different organizations in the United States, or elsewhere, may categorize the levels differently. More specifically, the disclosed systems and methods may be used in SAE Level 2 driver assistance systems, which implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. The disclosed systems and methods may be used in SAE Level 3 driver assistance systems, which are capable of autonomous driving under limited (e.g., highway) conditions. Similarly, the disclosed systems and methods may be used in vehicles using SAE Level 4 automated driving systems, which operate autonomously under most normal driving conditions and require only occasional attention from a human operator. In all such driver assistance systems, accurate lane estimation can be performed automatically without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, and overall safety of autonomous, semi-autonomous, and other driver assistance systems. As noted above, in addition to the way SAE classifies levels of autonomous driving operation, other organizations in the United States or other countries may classify levels of autonomous driving operation differently. Without limitation, the systems and methods disclosed herein may be used in driver assistance systems defined by the levels of autonomous driving operation of these other organizations.

[0023] The driving environment 101 may include any objects (moving or non-moving) located outside the AV, such as roads, buildings, trees, bushes, sidewalks, bridges, mountains, other vehicles, pedestrians, bridge piers, embankments, landing strips, animals, birds, etc. The driving environment 101 may be urban, suburban, rural, etc. In some implementations, the driving environment 101 may be an off-road environment (e.g., cultivated or other agricultural land). In some implementations, the driving environment 101 may be an indoor environment, such as an industrial plant, a shipping warehouse, a hazardous area of ​​a building, etc. In some implementations, the driving environment 101 may be substantially flat, with various objects moving parallel to the surface (e.g., parallel to the surface of the Earth). In other implementations, the driving environment 101 may be three-dimensional and include objects capable of moving along all three directions (e.g., balloons, fallen leaves, etc.). Hereinafter, the term "driving environment" should be understood to include all environments in which autonomous movement (e.g., SAE Level 5 and SAE Level 4 systems), conditional autonomous movement (e.g., SAE Level 3 systems), and / or movement of a vehicle equipped with driver assistance technology (e.g., SAE Level 2 systems) may occur. Furthermore, "driving environment" may include any possible flight environment of an aircraft (or spacecraft) or marine environment of a naval vessel. Objects in the driving environment 101 may be located at any distance from the AV, from a close distance of a few feet (or less) to several miles (or more).

[0024] Example AV 100 may include sensing system 110. Sensing system 110 may include various electromagnetic (e.g., optical, infrared, radio, etc.) and non-electromagnetic (e.g., acoustic) sensing subsystems and / or devices. Sensing system 110 may include one or more lidars 112, which may be laser-based units capable of determining the distance to and velocity of objects within driving environment 101. Sensing system 110 may include one or more radars 114, which may be any system that utilizes radio or microwave frequency signals to detect objects within AV 100's driving environment 101. LIDAR(s) 112 and / or RADAR(s) 114 may be configured to sense both the spatial location of objects (including their spatial dimensions) and their velocity (e.g., using Doppler shift techniques). Hereinafter, "velocity" refers to both how fast an object is moving (object speed) as well as the direction of the object's motion. Each of the lidar(s) 112 and radar(s) 114 may include a coherent sensor, such as a frequency-modulated continuous wave (FMCW) lidar or radar sensor. For example, the lidar(s) 112 and / or radar(s) 114 may use heterodyne detection for velocity determination. In some implementations, ToF and coherent lidar (or radar) functionality is combined into a lidar (or radar) unit that can simultaneously determine both the distance to a reflecting object and its radial velocity. Such a unit may be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection), or both modes simultaneously. In some embodiments, multiple lidars 112 and / or radars 114 may be mounted on the AV 100.

[0025] The lidar 112 (and / or radar 114) may include one or more light sources (and / or radio / microwave sources) that generate and emit signals and one or more detectors of signals reflected from objects. In some embodiments, the lidar 112 and / or radar 114 may perform a 360-degree scan in the horizontal direction. In some embodiments, the lidar 112 and / or radar 114 may be capable of spatial scanning along both the horizontal and vertical directions. In some implementations, the field of view may be up to 60 degrees vertically (e.g., at least a portion of the area above the horizon is scanned by the lidar or radar signal). In some embodiments (e.g., aerospace applications), the field of view may be spherical (consisting of two hemispheres).

[0026] The sensing system 110 may further include one or more cameras 118 to capture images of the driving environment 101. The cameras 118 may operate in the visible portion of the electromagnetic spectrum, for example, in the 300-800 nm wavelength range (also referred to herein as the optical range for simplicity). Some of the optical range cameras 118 may use a global shutter, while other cameras 118 may use a rolling shutter. The images may be two-dimensional projections of the driving environment 101 (or portions of the driving environment 101) onto the camera's projection surface (planar or non-planar). Some of the cameras 118 of the sensing system 110 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101. The sensing system 110 may also include one or more sonars 116, such as ultrasonic sonars, for active sound probing of the driving environment 101 and one or more microphones for passively listening to sounds in the driving environment 101. The sensing system 110 may also include one or more infrared (IR) sensors 119. For example, the IR sensor(s) 119 may include an IR camera. The IR camera(s) 119 may use light collection optics (e.g., made of germanium-based materials, silicon-based materials, etc.) configured to operate in the wavelength range of a few microns to tens of microns or more. The IR camera(s) 119 may include a phased array of IR detector elements. The pixels of the IR image produced by the IR sensor(s) 119 may represent the total amount of IR radiation collected by each detector element (associated with the pixel), the temperature of the physical object from which IR radiation is collected by each detector element, or any other suitable physical quantity.

[0027] The sensory data acquired by the sensing system 110 may be processed by the data processing system 120 of the AV 100. For example, the data processing system 120 may include a perception system 130. The perception system 130 may be configured to detect and track objects in the driving environment 101 and recognize the detected objects. For example, the perception system 130 may analyze images captured by the camera(s) 118 and may be capable of detecting traffic signals, road signs, road layouts (e.g., lane boundaries, intersection topology, parking designations, etc.), the presence of obstacles, etc. The perception system 130 may also receive radar sensory data (Doppler data and ToF data) to determine the distances to various objects in the environment 101 and the velocities of such objects (line of sight and, in some embodiments, lateral, as described below). In some implementations, the perception system 130 may use radar data in combination with data captured by the camera(s) 118, as described in more detail below.

[0028] The perception system 130 may include one or more components to facilitate object detection, classification, and tracking, including an end-to-end perception model (EEPM) 132 that may be used to process data provided by the sensing system 110. More specifically, in some implementations, the EEPM 132 may receive data from sensors of different sensing modalities. For example, the EEPM 132 may receive images from at least some of the lidar(s) 112, radar(s) 114, and (optical range) camera(s) 118, IR sensor(s) 119, sonar(s) 116, etc. In particular, the EEPM 132 may include one or more trained machine learning models (MLMs) that are used to process some or all of the above data to detect, classify, and track the movement of various objects within the driving environment 101. The EEPM 132 can use multiple classifier heads to determine various characteristics of the external environment, including, but not limited to, the occupancy of space by various objects, object types, object motion, the identification of objects that may be occluded, and the relationship of objects to roads, other objects, and traffic flow. The various models of the EEPM 132 can be trained using multiple sets of images / data and annotated to identify specific features in each piece of sensed data. In some implementations, the perception system 130 can include a behavior prediction module (BPM) 134 that predicts future movements of detected objects (e.g., objects detected by the EEPM 132).

[0029] The perception system 130 may further receive information from a Global Navigation Satellite System (GNSS) positioning subsystem (not shown in FIG. 1 ), which may include a GNSS transceiver (not shown) configured to obtain information about the position of the AV relative to the Earth and its surroundings. The positioning subsystem may use positioning data, e.g., GNSS and inertial measurement unit (IMU) data, in conjunction with the sensed data to help accurately determine the location of the AV with respect to fixed objects (e.g., roadways, lane boundaries, intersections, sidewalks, crosswalks, road signs, curbs, surrounding buildings, etc.) in the driving environment 101, whose locations may be provided by the map information 124. In some implementations, the data processing system 120 may receive non-electromagnetic data, such as audio data (e.g., ultrasonic sensor data from the sonar 116 or data from a microphone picking up an emergency vehicle siren), temperature sensor data, humidity sensor data, pressure sensor data, weather data (e.g., wind speed and direction, precipitation data), etc.

[0030] The data processing system 120 may further include an environment monitoring and prediction component 126, which can monitor how the driving environment 101 evolves over time, for example, by tracking the position (relative to the Earth) and velocity of moving objects. In some implementations, the environment monitoring and prediction component 126 can track the changing appearance of the environment due to the movement of the AV relative to the environment. In some implementations, the environment monitoring and prediction component 126 can make predictions about how various moving objects in the driving environment 101 will be positioned within a prediction time frame. The predictions may be based on the current state of the moving objects, including their current positions (coordinates) and velocities. Additionally, the predictions can be based on the history of the moving objects' motion (tracked dynamics) during a specific period preceding the current moment. For example, based on stored data regarding a first object indicating accelerated movement of the first object during the previous three seconds, the environment monitoring and prediction component 126 can conclude that the first object is resuming its motion from a stop sign or red light signal. Thus, the environmental monitoring and prediction component 126 can predict where a first object is likely to be within the next three or five seconds of movement, taking into account the layout of the roadway and the presence of other vehicles. As another example, based on stored data about a second object indicating the second object's slowed movement within the previous two seconds, the environmental monitoring and prediction component 126 can conclude that the second object is stopped at a stop sign or red light. Thus, the environmental monitoring and prediction component 126 can predict where the second object is likely to be within the next one or three seconds. The environmental monitoring and prediction component 126 can periodically check the accuracy of its prediction and revise the prediction based on new data obtained from the sensing system 110. The environmental monitoring and prediction component 126 can operate in conjunction with the EEPM 132. For example, the environmental monitoring and prediction component 126 can track the relative movement of the AV and various objects (e.g., reference objects that are stationary or moving relative to the Earth).

[0031] The data generated by the perception system 130, the GNSS processing module 122, and the environmental monitoring and prediction component 126 can be used by an autonomous driving system, such as the AV control system (AVCS) 140. The AVCS 140 may include one or more algorithms that control how the AV should behave in various driving situations and environments. For example, the AVCS 140 may include a navigation system for determining a global driving path to a destination. The AVCS 140 may also include a driving path selection system for selecting a particular path through the driving environment 101, which may include selecting a lane, navigating traffic jams, selecting a location to make a U-turn, selecting a trajectory for a parking maneuver, etc. The AVCS 140 may also include an obstacle avoidance system for safely avoiding various obstacles (e.g., rocks, stopped vehicles, etc.) in the AV's driving environment. The obstacle avoidance system may be configured to assess the size of an obstacle and the trajectory of the obstacle (if the obstacle is moving) and select an optimal driving strategy (e.g., braking, steering, accelerating, etc.) to avoid the obstacle.

[0032] The algorithms and modules of AVCS 140 may generate instructions for various systems and components of the vehicle, such as powertrain, braking, and steering 150, vehicle electronics 160, signaling 170, and other systems and components not explicitly shown in FIG. 1 . Powertrain, braking, and steering 150 may include an engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanism, and other systems. Vehicle electronics 160 may include an on-board computer, engine management, ignition, communication systems, car computers, telematics, in-car entertainment systems, and other systems and components. Signaling 170 may include high and low headlights, stop lights, turn signals and taillights, horns and alarms, interior lighting systems, dashboard notification systems, passenger notification systems, radio and wireless network transmission systems, etc. Some of the commands output by AVCS 140 may be delivered directly to powertrain, braking, and steering 150 (or signaling 170), while other commands output by AVCS 140 are first delivered to vehicle electronics 160, which generates commands to powertrain, braking, and steering 150 and / or signaling 170.

[0033] In one example, the EEPM 132 can determine that an image acquired by the camera(s) 118 includes a depiction of an object, and can further classify the object as a bicyclist. The environmental monitoring and prediction component 126 can track the bicyclist and determine that the bicyclist is traveling at a speed of 15 mph along an intersecting road perpendicular to the vehicle's direction of movement. In response to such a determination, the BPM 134 can determine that the vehicle needs to slow down to allow the bicyclist to clear the intersection. The AVCS 140 can output instructions to the powertrain, brakes, and steering 150 (either directly or via the vehicle electronics 160) to: (1) change the throttle setting to reduce fuel flow to the engine and lower engine speed; (2) downshift the drivetrain, via the automatic transmission, into a lower gear; and (3) activate the brake unit (in cooperation with the engine and transmission) to reduce vehicle speed. After the EEPM 132 and / or environmental monitoring and prediction component 126 determine that a bicyclist has crossed the intersection, the AVCS 140 can output commands to the powertrain, braking, and steering 150 to resume the vehicle's previous speed setting.

[0034] The output of the EEPM 132 may be used to track detected objects. In some implementations, the tracking may be reactive and include a history of the pose (position and orientation) and velocity of the tracked object. In some implementations, the tracking may be proactive and include a prediction of the tracked object's future pose and velocity. In some implementations, for example, future predictions may be generated by the BPM 134 based at least in part on the output of the EEPM 132. In some implementations, tracking by detection or instance segmentation may be used instead of building an explicit tracker. For example, the BPM 134 interface may include, for each object, a history of the object's recent location, range, heading, and velocity. In some implementations, flow information may be defined in terms of units of three-dimensional space (voxels). To add accuracy to the prediction, the flow information associated with individual voxels may include kinematic attributes such as curvature, yaw rate, and more, in addition to velocity. Based on this data, the BPM 134 can predict future trajectories in a manner that is more advantageous than more traditional tracking approaches. In some implementations, an alternative approach can be used that introduces a recurrent neural network (RNN) to smooth and interpolate location and velocity over time, which can be performed similarly to the operation of a Kalman filter.

[0035] The output of the EEPM 132 can be used to determine the vehicle's position. In some implementations, the BPM 134 can use lidar-based global mapping, which maps the entire area of ​​the 3D environment around the vehicle. In some implementations, the BPM 134 can employ a simpler system that uses acceleration measurements, odometry, GNNS data, and camera-based lane mapping to identify the vehicle's current position relative to map data.

[0036] In different implementations, the BPM 134 may have different levels of sophistication depending on the driving environment 101 (e.g., highway driving, urban driving, suburban driving, etc.). In L2 driver assistance implementations (“hands on the wheel”) where the driver is expected to take over control of the vehicle at any time, the BPM 134 may have minimal functionality and be able to predict the behavior of other road users within a short time horizon, such as a few seconds. For example, such predictions may include lane change interference by other vehicles (subjects). The BPM 134 may use various cues, such as turn signals, front wheel steering, or the driver turning their head in the direction of the turn. The BPM 134 may determine whether such an impending lane change requires the driver's attention. If the lane-changing subject is far enough away from the vehicle, the AVCS 140, acting on the BPM's 134 predictions, may alter the vehicle's trajectory (e.g., slow down the vehicle) without driver involvement. If the change requires immediate driver attention, the BPM 134 may output a signal to the driver indicating that the driver should assume control of the vehicle.

[0037] In an L3 driver assistance implementation (“hands off the wheel”), the objective may be to provide autonomous driving functionality for at least a certain time horizon (e.g., X seconds) so that if a condition requiring driver control arises, this condition is predicted at least X seconds before the condition occurs. The map data may further include camera and / or radar imagery of prominent landmarks (bridges, signs, roadside structures, etc.). In some implementations, the BPM 134 of the L3 system may output two trajectories, option A and backup option B, for X seconds at any given time. For example, when traveling on a street in the right-most lane of the street, the BPM 134 may calculate option A for the vehicle to stay in the right-most lane, and if a parked vehicle veers into the left-most lane, may further calculate option B for the vehicle to move to the left lane. The BPM 134 may predict that the left lane will remain available within the next X seconds and continue vehicle operation. At some point, BPM 134 may predict that in the left lane, a fast-moving entity will move close enough to the vehicle that the left lane (and thus option B) will no longer be available to the vehicle. Upon determining that option B is likely no longer available, BPM 134 may contact the driver to take control of the vehicle. Even in more sophisticated systems (e.g., autonomous L4 driving systems), if option B disappears, where no driver input is expected, AVCS 140 may park the vehicle on the side of the road until driving conditions change favorably.

[0038] To achieve reliable predictions, the BPM 134 can simulate multiple possible scenarios of how different road users might behave in different ways and estimate the probabilities and corresponding outcomes of various such scenarios. In some implementations, the BPM 134 can use a closed-loop approach to determine a distribution of probabilities that if a vehicle makes a certain driving path change (or maintains its current driving path), other vehicles will respond in a certain way—for example, by yielding to the vehicle, accelerating, or otherwise blocking the vehicle's driving path. The BPM 134 can evaluate multiple such scenarios and output probabilities for each or at least some of the scenarios. In some implementations, the BPM 134 can use an open-loop approach in which predictions are made based on the current movement state of an entity, and changes in the vehicle's movement do not affect the behavior of other entities. In some implementations, the predicted locations of various entities can be represented by a future occupancy heat map. Further details regarding the EEPM 132 are now described below with reference to FIG. 2.

[0039] 2 is a diagram illustrating an example network architecture of an end-to-end perception model (EEPM) 132 that may be deployed as part of a vehicle's perception system, according to some implementations of the present disclosure. Input data 201 may include data acquired by various components of sensing system 110 (as shown in FIG. 1), such as lidar(s) 112, radar(s) 114, optical (e.g., visible) range camera(s) 118, and IR sensor(s) 119. For example, as shown, input data 201 may include camera data 210 and radar data 220. Although not shown, input data 201 may further include, for example, lidar data.

[0040] The input data 201 may include images and / or any other data, such as metadata, e.g., voxel intensities, velocity data associated with voxels, and timestamps. The input data 201 may include, for example, directional data (e.g., angular coordinates of return points), distance data, and radial velocity data, as may be acquired by the lidar(s) 112 and / or radar(s) 114. Additionally, the input data 201 may further include road graph data stored (or accessible) by the perception system 130, e.g., as part of the map information 124. The road graph data may include any two-dimensional map of the road and a three-dimensional map of its surroundings (including any suitable mapping of stationary objects, e.g., identification of bounding boxes for such objects). It should be understood that this list of input data 201 is not exhaustive and that any suitable additional data may be used as part of the input data 201, e.g., IMU data, GNNS data, etc. Each modality of the input data 201 may be associated with a particular time at which the data was obtained. A set of available data (e.g., lidar, radar, camera, and / or IR camera images, etc.) associated with a particular time may be referred to as a sensing frame. In some implementations, images acquired by different sensors may be synchronized so that all images in a given sensing frame have the same timestamp (within the accuracy of the synchronization). In some implementations, some images in a given sensing frame may have a (controlled) time offset.

[0041] The image acquired by one of the sensors is converted into a corresponding intensity map I({x j}), where {x j} can be any set of coordinates, including three-dimensional (spherical, cylindrical, Cartesian, etc.) coordinates (e.g., for lidar and / or radar images) or two-dimensional coordinates (e.g., for camera data). The coordinates of various objects (or object surfaces) reflecting lidar and / or radar signals can be determined from direction data (e.g., polar angle θ and azimuth angle Φ of the lidar / radar transmission direction) and distance data (e.g., radial distance R determined from the ToF of the lidar / radar signal). The intensity map can identify the strength of the sensed signal detected by the corresponding sensor. Similarly, the lidar and / or radar sensor may generate a Doppler (frequency shift) map Δf({x ≠ 1 ) that identifies the radial velocity of the reflecting object based on a detected Doppler shift Δf in the frequency of the reflected radar signal. j}), where V=λΔf / 2, where λ is the lidar / radar wavelength, and positive values ​​Δf>0 are associated with objects (hence, vehicles) moving toward the lidar / radar, and negative values ​​Δf<0 are associated with objects moving away from the lidar / radar. In some implementations, for example, in a driving environment where objects are moving substantially within a particular plane (e.g., the Earth's surface), radar intensity maps and Doppler maps may be defined using two-dimensional coordinates such as radial distance and azimuth angle: I(R, φ), Δf(R, φ).

[0042] The camera feature network 212 receives the camera data 210 and can extract a camera data feature set from the camera data 210. For example, the camera data feature set can include a set of camera data feature vectors. More specifically, the camera data features can be two-dimensional (2D) camera data features. The camera data feature network 212 can obtain the camera data feature set using any suitable perspective backbone(s). Examples of suitable perspective backbones include ResNet, EfficientNet, etc. In some implementations, each camera sensor (e.g., front camera, rear camera, etc.) can use the same vision backbone (e.g., the same shared weights). Training the camera data feature network 212 using a common backbone can be advantageous because it prevents the network from learning to depend on a particular camera view. Training the camera data feature network 212 using a common backbone can promote the resilience of the EEPM 132 to various vehicle pose anomalies, such as vehicle yaw and roll. Each camera data feature can be associated with a particular pixel or cluster of pixels. Each pixel (or cluster of pixels) may be associated with a respective depth distribution and a respective depth feature. In some implementations, the processed camera data may be downsampled for computational efficiency. In some implementations, a pseudo camera may be used. The pseudo camera represents a crop of the image from the full-resolution image, providing finer details for long-range tasks. The pseudo camera may have a fixed crop or a crop that arises from the output of a coarse-resolution backbone. In some implementations, the crop may be trained directly. In some implementations, differentiable cropping may be used to train an attention mechanism end-to-end.

[0043] The camera data features may be provided to the camera data feature projection component 214. The camera data feature projection component 214 can utilize the camera data feature projection to transform the camera data feature set into a set of pixel points. For example, the set of pixel points may be a pixel point cloud. In some implementations, utilizing the camera data feature projection includes performing a lift transformation on the 2D camera data (e.g., from a 2D backbone, sensor intrinsics and extrinsics (or derived intrinsics and extrinsics for a pseudo camera)). To do so, the camera data feature transformation component 214 can project the 2D camera data into three-dimensional (3D) space. This projection may be performed using various depth distribution techniques. During training, depth ground truth may be available from other sensor data (e.g., lidar data) and may be used as a structured loss. The output of other sensors capable of providing 2D images (e.g., IR cameras) can be processed using the same (or similar) architecture. Therefore, the camera data feature projection component 214 can provide a lift-transformed camera "context" combined across the cameras of the AV. Further details regarding the generation of the pixel point set are described below with reference to FIG. 3.

[0044] More specifically, the lift transform can combine a depth distribution with a set of camera features (e.g., a feature vector). As an illustrative example, the lift transform can be used to transform a feature vector FV(c) w,h Each pixel w,h described by can be complemented with depth information from a depth distribution. For example, the lift transform can be used to find the corresponding depth distribution P(d) (of dimension D × 1) for the same pixel. w,h Using this, each feature vector FV(c) (of dimension C × 1) w,h The output of the lift transform is, for example,

[0045]

number

[0046] The characteristic may be represented by:

[0047] Then, the feature tensor FT(c,d) calculated for each pixel w,h Using the tensor function, we can obtain a combined feature tensor for the whole image by concatenating the feature tensors for different pixels: w,h}→CFT(c,d,w,h). The combined feature tensor CFT(c,d,w,h) has dimensions C×D×W×H. The combined feature tensor may then undergo 2D mapping. More specifically, the 2D mapping may yield a projected feature tensor that uses a convenient set of planar coordinates, e.g., Cartesian coordinates x and y, or polar coordinates r and θ, within the ground plane.

[0048] The 2D mapping can be a two-part transformation. During the first part, viewpoint coordinates d, w, h can be transformed into 3D Cartesian coordinates d, w, h → x, y, z (or 3D cylindrical coordinates, w, h → r, θ, z), where z is the vertical coordinate (the direction perpendicular to the ground). The transformation d, w, h → x, y, z can be a projective transformation parameterized by the camera's focal length, the direction of the camera's optical axis, and other similar parameters. If images are acquired by multiple cameras (or cameras with rotating optical axes), the transformation d, w, h → x, y, z can include multiple projective transformations, e.g., with separate transformations used for pixels w, h provided by different cameras.

[0049] During the second part, a 2D mapping can be performed by projecting the combined feature tensor, represented by new coordinates CFT(c,x,y,z), onto a horizontal plane to obtain a projected (BEV) feature tensor. For example, to obtain a C×W×H projected feature tensor PCT(c,x,y), the combined feature tensor can be mapped onto the elements associated with each vertical pillar of the pixel, e.g.,

[0050]

number

[0051] In some implementations, the sum (or average) can be calculated for coordinate z i The summation for different coordinates z i The different weights w assigned to i may be implemented with

[0052]

number

[0053] For example, the larger the weight w i is assigned to pixels that image objects within a certain elevation above the ground (e.g., up to a few meters), and lower weights are assigned to other elevations (e.g., to eliminate spurious objects such as tree branches or electrical cables that do not obstruct the vehicle's movement). The projected feature tensor can characterize objects and their locations within the BEV with perspective distortion reduced (e.g., eliminated).

[0054] The radar data feature network 222 can receive the radar data 220 and extract a set of radar data features from the radar data 220. For example, radar data features can be generated for each radar. The radar data feature network 222 can use any suitable radar backbone(s). Examples of suitable radar backbones include PointPillars, Range Sparse Net, etc. Each radar modality (e.g., intensity, second echo, Doppler shift, radar cross section) can have a different radar backbone and feature generation layer. In some implementations, a full cycle (spin) of the lidar / radar sensor can be used to obtain radar data features. In some implementations, a portion of the radar cycle can be used to obtain radar data features. Processing such a portion of the cycle allows the EEPM 132 to react more quickly to new entities (e.g., vehicles, pedestrians, etc.) or sudden movement of existing entities in some cases and can operate at the speed of the fastest sensor.

[0055] The radar data feature set can be provided to the radar data feature transform component 224. The radar data feature transform component 214 can utilize a radar data feature transform to transform the radar data feature set into a set of radar points. For example, the set of radar points can be a radar point cloud. Further details regarding generating a set of radar points are described below with reference to FIG. 3.

[0056] The set of pixel points generated by the camera data feature projection component 214 and the set of radar points generated by the radar data feature transformation component 224 may be provided to the BEV feature processing component 230. The BEV feature processing component 230 may generate a respective BEV feature set for each time step. Each BEV feature set may be integrated into a BEV grid (e.g., a BEV voxel grid). The radar data features of the radar data feature set may have a coordinate representation that is not suitable for integration into the BEV grid. Thus, in some implementations, performing the radar data feature transformation may include converting the coordinate representation of the radar data feature set into a coordinate representation appropriate for integration into the BEV grid. For example, the appropriate coordinate representation may be a Cartesian coordinate representation. Illustratively, the radar data feature network 222 may process the radar data 220 in a polar coordinate representation, and converting the coordinate representation includes converting from the polar coordinate representation to a Cartesian coordinate representation.

[0057] The BEV feature processing component 230 can perform multi-frame temporal aggregation to generate an aggregated BEV feature set from multiple BEV feature sets. More specifically, the multiple BEV feature sets can include a BEV feature set corresponding to a current time step (e.g., t) and one or more BEV features, each corresponding to a time step prior to the current time step (e.g., t-1 and t-2). For example, aggregating the BEV feature sets can include warping each BEV feature set corresponding to a time step prior to the current time step. The warping is performed to maintain the position of static objects across the BEV feature sets, thereby accounting for changes in the AV's position over time. This can enable improved prediction of object motion by excluding the AV's self-motion from consideration. Therefore, the warping can align the features that are combined to generate the aggregated BEV feature set.

[0058] In some implementations, the aggregated BEV feature set is a concatenation-based BEV feature set, as described below with reference to Figure 4A. In some implementations, the aggregated BEV feature set is an addition-based BEV feature set, as described below with reference to Figure 4B. In some implementations, the aggregated BEV feature set is a transformer-based BEV feature set, as described below with reference to Figure 4C.

[0059] The output of the BEV feature processing component 230 (e.g., each BEV feature set) can then be provided to the object flow processing component 240. The object flow processing component 240 can perform multi-frame dense motion estimation to generate object flow (e.g., the movement of one or more objects) using multiple BEV feature sets. More specifically, the object flow can be an object flow prediction.

[0060] In some implementations, the object flow processing component 240 includes a correlation network. For example, the object flow processing component can include multiple BEV feature networks (e.g., BEV backbones), each receiving a respective BEV feature set for each time step. For example, a first BEV feature network can receive a BEV feature set corresponding to the current time step (e.g., a BEV feature set corresponding to time step t), and a second BEV feature network can receive a BEV feature set corresponding to a previous time step (e.g., a warped BEV feature set corresponding to time step t-1). The outputs of the BEV feature networks (i.e., features extracted from each BEV feature set) can be correlated to generate a correlation output. For example, the correlation output can include a cost volume. Furthermore, the object flow processing component 240 can further include a context network for receiving the BEV feature set corresponding to the current time step and generating a context network output. The object flow processing component 240 can further include a flow estimation network that receives a dataset including the correlation output, the context network output, and an initial flow and can obtain a final flow from the dataset. In some implementations, the initial flow corresponds to a zero flow field. More specifically, the flow estimation network can determine a sequence of flows where the initial flow is the initial flow in the sequence and the final flow is the final flow in the sequence. Further details regarding these implementations are described below with reference to Figures 5A-5B.

[0061] In some implementations, the object flow processing component 240 includes a BEV feature network for extracting features from a BEV feature set (e.g., a set of aggregated BEV features) and a flow classification head ("flow head") for generating object flows from the extracted features. For example, the BEV feature network may include a BEV backbone. The flow head may output any suitable representation of object flows corresponding to various voxels in space (e.g., using motion vectors, etc.). For example, in the case of a 2D BEV grid, the flow head may output a 2D vector for each grid point. This 2D vector may represent displacement relative to the previous frame along the height (H) and width (W) axes with respect to the 2D BEV grid. In the case of a self-motion-compensated version of object flow, the displacement may be from the perspective of a grid fixed to a reference time step (so that static parts of the world have zero displacement vectors). This may also be extended to 3D. Further details regarding these implementations are described in more detail below with reference to FIG. 5C.

[0062] The EEPM 132 may include one or more additional feature networks (not shown). For example, the EEPM 132 may include a road graph feature network that can process road graph data and output road graph features that may include lanes and lane markings, road edges and medians, traffic lights and stop signs, crosswalks and speed bumps, driveways, parking lots and curb limits, railroad crossings, school zones, and vehicle-unaccessible areas. The road graph features may be voxelized into a coordinate frame. The road graph data may further include an elevation map. Such previous data may be treated as a separate modality. Such a framework may facilitate incorporating new location-based data, such as a heat map of object occurrences observed on previous trips. The road graph data may be accumulated during previous driving missions for a particular route. In some cases, if previous data is unavailable, the road graph data may be limited by the map information 124 available for a particular route. As with other modalities, road graph data may be missing, and during training, the EEPR 132 may be trained to incrementally incorporate road graph data rather than relying on such data.

[0063] 3 is a diagram 300 illustrating an example flow according to some implementations of the present disclosure. Diagram 300 shows base flow 310, which represents AV motion 312 and object motion of objects 314-1 and 314-2 observed in a driving environment. Diagram 300 further shows ego flow 320, which represents AV motion 312 (assuming all objects are stationary). Diagram 300 further shows object flow 330, which represents the opposing motion of objects 314-1 and 314-2. For example, object flow 330 may be determined as the difference between base flow 310 and ego flow 320.

[0064] 4A is a diagram illustrating an example temporal aggregation model architecture (“architecture”) 400A in accordance with some implementations of the present disclosure. More specifically, architecture 400A depicts a consolidation-based temporal aggregation model architecture. Architecture 400A can utilize constant weights per grid cell.

[0065] As shown, a BEV feature set 410-1 corresponding to time step t-2, a BEV feature set 410-2 corresponding to time step t-1, and a BEV feature set 410-3 corresponding to time step t (i.e., the current time step) can be obtained. More specifically, each BEV feature set 410-1 through 410-3 can be generated from a corresponding combination of camera data features and radar data features extracted from the camera data and radar data, respectively, as described above with reference to FIG. 2. Although only three BEV feature sets corresponding to three time steps are shown in FIG. 4A, architecture 400A can be used with any suitable number of BEV feature sets in accordance with implementations described herein.

[0066] The BEV feature set 410-1 may be provided as input to the warp component 420-1 to warp the camera data feature set and generate a warped BEV feature set 430A-1. More specifically, the warping is performed relative to a common reference time step, which may be the current time step t. The warping may be a mathematical transformation that eliminates the (independently known) motion of the sensing system (e.g., the self-motion of an autonomous vehicle). In some implementations, the warped BEV feature set 430A-1 includes warped feature vectors. As a result of the warping, an object that is stationary relative to the ground may be represented by elements of the warped BEV feature set 430A-1 associated with the same x, y coordinates, whereas a moving object will be represented by elements that span the direction of the object's motion.

[0067] The warped BEV feature set 430A-1 may be concatenated with the BEV feature set 410-2 to generate a first concatenation-based BEV feature set. The first concatenation-based BEV feature set may be provided as an input to the warp component 420-2 to generate the warped BEV feature set 430A-2. In some implementations, the warped BEV feature set 430A-2 includes a warped feature vector.

[0068] The warped BEV feature set 430A-2 can be concatenated with the BEV feature set 410-3 to generate a second concatenation-based BEV feature set, which can then be provided to the BEV feature processing component 240 for further processing, as described above with reference to FIG.

[0069] 4B is a diagram illustrating an example temporal aggregation model architecture (“architecture”) 400B in accordance with some implementations of the present disclosure. More specifically, architecture 400B depicts an addition-based temporal aggregation model architecture.

[0070] As shown, BEV feature set 410-1, BEV feature set 410-2, and BEV feature set 410-3 can be obtained as described above with reference to Figure 4A. Although only three time steps are shown in Figure 4B, architecture 400B can be used with any suitable number of time steps in accordance with implementations described herein.

[0071] Similar to the architecture 400A described above with reference to FIG. 4A, the BEV feature set 410-1 may be provided as an input to the warp component 420-1. In contrast to the architecture 400A described above with reference to FIG. 4A, the warped BEV feature set output by the warp component 420-1 may be combined with the BEV feature set 410-2 at adder 440-1 to generate the summation-based BEV feature set 430B-1. More specifically, the summation may be a weighted sum. The weights may be constant weights per grid cell. The summation-based BEV feature set 430B-1 may be provided as an input to the warp component 420-2. The warped BEV feature set output by the warp component 420-2 may be combined with the BEV feature set 410-3 at adder 440-2 to generate the summation-based BEV feature set 430B-2. The summation-based BEV feature set 430B-2 can then be provided to the BEV feature processing component 240 for further processing, as described above with reference to FIG.

[0072] 4C illustrates an example temporal aggregation model architecture ("architecture") 400C in accordance with some implementations of the present disclosure. More specifically, architecture 400C depicts an additive-based temporal aggregation model architecture with a transformer. Architecture 400C can be used to determine (e.g., predict) weights for each grid cell instead of using a fixed weight for each grid cell, as described above with respect to architectures 400A and 400B.

[0073] As shown, BEV feature set 410-1, BEV feature set 410-2, and BEV feature set 410-3 can be obtained as described above with reference to Figure 4A. Although only three time steps are shown in Figure 4C, architecture 400B can be used with any suitable number of time steps in accordance with implementations described herein.

[0074] Similar to architecture 400A described above with reference to FIG. 4A and architecture 400B described above with reference to FIG. 4B, BEV feature set 410-1 may be provided as an input to warp component 420-1. In contrast to architecture 400A described above with reference to FIG. 4A and architecture 400B described above with reference to FIG. 4B, the warped BEV feature set output by warp component 420-1 may be provided to transformer 450-1. Transformer 450-1 may be used to determine (e.g., predict) per-grid cell weights to use during subsequent summation, as opposed to the fixed per-grid cell weights used to generate the summation-based BEV feature set described above with reference to FIG. 4B. More specifically, transformer 450-1 may predict how the warped BEV feature set should be weighted relative to BEV feature set 410-2 and / or how they should be aggregated.

[0075] The output (e.g., weights) of the transformer 450-1 can then be used by the adder 440-1 to generate the transformer-based BEV feature set 430C-1. The transformer-based BEV feature set 430C-1 can be provided as an input to the warp component 420-2. The warped BEV feature set output by the warp component 420-2 can be provided to the transformer 450-2. The output (e.g., weights) of the transformer 450-2 can then be used by the adder 440-2 to generate the transformer-based BEV feature set 430C-2. The transformer-based BEV feature set 430C-2 can then be provided to the BEV feature processing component 240 for further processing, as described above with reference to FIG. 2.

[0076] FIG. 5A illustrates an exemplary multi-frame flow model architecture (“architecture”) 500A in accordance with some implementations of the present disclosure. Architecture 500A can be used to perform dense motion estimation and determine (e.g., estimate) the speed of an object. For example, the output of architecture 500A can be used to determine motion vectors (e.g., velocity vectors) and determine object flow. Dense motion estimation can enable objects observed in a driving environment to be classified among any known classes of objects and an “unknown object” class (e.g., based on the object’s motion pattern). The movement of the unknown object can be tracked until the object safely exits the driving environment. As will be described in further detail, architecture 500A is a specialized architecture capable of computing feature correlations between multiple frames.

[0077] As shown, BEV feature set 510-1 and BEV feature set 510-2 may be obtained. More specifically, BEV feature set 510-1 may correspond to time step t−1, and BEV feature set 510-2 may correspond to time step t. BEV feature set 510-1 may be provided as an input to warp component 520-1 to generate warped BEV feature set 530. In some implementations, warped BEV feature set 530 includes a warped feature vector.

[0078] As further shown, architecture 500A can further include object flow processing component 240. More specifically, object flow processing component 240 can include feature networks 540-1 and 540-2, a correlator 545, a context network 550, and a flow estimation network 570.

[0079] Warped BEV feature set 530 can be provided to feature network 540-1, and BEV feature set 510-1 can be provided to feature network 540-2. The outputs (i.e., extracted features) of feature networks 540-1 and 540-2 can be provided to correlator 545 to generate a feature correlation that is used to generate correlation output 560. Correlation output 560 measures the similarity between the features extracted by feature networks 540-1 and 540-2.

[0080] In some implementations, as shown in FIG. 5A, the correlation output 560 is a cost volume generated from the outputs of the feature networks 540-1 and 540-2. Generally, the cost volume can represent a full-pair correlation volume between a pair of objects of a given dimension. In some embodiments, the cost volume is a four-dimensional (4D) cost volume. The 4D cost volume can represent a full-pair correlation volume between a pair of 2D objects. For example, assume that a pair of BEV grids, including a first BEV grid and a second BEV grid, each have dimensions (H×W). For each grid point of the first BEV grid, correlations with every other point of the second BEV grid can be calculated and stored in a 4D cost volume. Thus, the 4D cost volume can have dimensions (H×W)×(H×W).

[0081] Additionally, the BEV feature set 510-2 can be provided to a context network 550 to generate a context network output. The context network 550 can include a convolutional neural network (CNN) for feature extraction that can maintain context information. The context information can include context details about the driving environment. For example, the context information can indicate which portions of the BEV features (e.g., grid cells of the BEV grid) belong to moving objects in the driving environment because they may share similar flow and / or displacement.

[0082] 5B, the correlation output 560, the context network output, and the initial flow 580 may be provided to a flow estimation network 570 to generate a final flow 590. In some implementations, the initial flow 580 corresponds to a zero flow field. More specifically, the final flow 590 may be generated by obtaining a sequence of flow estimates from the initial flow 560.

[0083] 5B illustrates an example flow estimation network 570 of architecture 500 according to some implementations of the present disclosure. More specifically, flow estimation network 570 may implement a recurrent neural network architecture. For example, flow estimation network 570 may include a set of gated recurrent units (GRUs) 582-1 through 582-N, a set of lookup components 584-1 through 584-(N-1), and a set of adders 586-1 through 586-N. Lookup refers to features (e.g., 4D cost volumes) sampled from correlation output 560.

[0084] As shown, a context network output 592 (e.g., the output of the context network 550) may be provided as an input to a GRU 582-1. The output of the GRU 582-1 may be combined with an initial flow 580 in a summer 586-1 to generate a first flow estimate of a sequence of flow estimates 588 (i.e., updating the initial flow 580 with the output of the GRU 582-1). To generate a second flow estimate of the sequence of flow estimates 588, the cost volume 560 and the first flow estimate output by the summer 586-1 may be provided as input to a lookup component 584-1. The output of the lookup component 584-1 and the output of the GRU 582-1 may be provided as input to a GRU 582-2. The output of the GRU 582-2 may be combined with the output of the summer 586-1 in a summer 586-2 to generate a second flow estimate. A similar process may be performed to generate the remaining flow estimates of the sequence of flow estimates until a final flow 590 is generated.

[0085] 5C illustrates an exemplary multi-frame flow model architecture (architecture) 500C according to some implementations of the present disclosure. Architecture 500C can be used to perform dense motion estimation and determine (e.g., estimate) the speed of an object. For example, the output of architecture 500C can be used to determine motion vectors (e.g., velocity vectors) and determine object flow. Dense motion estimation can enable objects observed in a driving environment to be classified among any known classes of objects and an "unknown object" class (e.g., based on the object's motion pattern). The movement of the unknown object can be tracked until the object safely exits the driving environment.

[0086] As shown, the BEV feature set 510-2 and the warped BEV feature set 530, described above with reference to FIG. 5A, can be combined to form an aggregated BEV feature set. In this illustrative example, the BEV feature set 510-2 and the warped BEV feature set 530 are concatenated to generate a concatenation-based BEV feature set. In an alternative implementation, the BEV feature set 510-2 and the warped BEV feature set 530 are combined (e.g., as described above with reference to FIG. 4B) to generate an addition-based BEV feature set. In an alternative implementation, the BEV feature set 510-2 and the warped BEV feature set 530 are combined (e.g., as described above with reference to FIG. 4C) to generate a transformer-based BEV feature set.

[0087] As further shown, architecture 500C can further include an object flow processing component 240. More specifically, object flow processing component 240 can include a BEV feature network 532 and a flow head 534. BEV feature network 532 can receive the aggregated BEV feature set and extract features from the aggregated BEV feature set. The features extracted by BEV feature network 532 can be provided to flow head 534 to generate object flow 536.

[0088] FIG. 6 is a flow diagram illustrating an example method for implementing multi-frame temporal aggregation and dense motion estimation for an autonomous vehicle (AV) according to some implementations of the present disclosure. A processing device having one or more processing units (CPUs) and a memory device communicatively coupled to the CPU(s) may perform method 600 and / or each of their individual functions, routines, subroutines, or operations. A processing device executing method 600 may implement instructions issued by various components of sensing system 110 or data processing system 120 of FIG. 1 , such as EEPM 132. In some implementations, method 600 may be directed to systems and components of an autonomous vehicle, such as autonomous vehicle 100 of FIG. 1 . In some implementations, method 600 may be performed by EEPM 132 or any other similar model that may be part of the perception system of an autonomous vehicle, a vehicle employing driver assistance technology, or any other application platform that uses object detection and classification.

[0089] Method 600 can be used to improve the performance of data processing system 120 and / or AVCS 140. In certain implementations, a single processing thread may perform method 600. Alternatively, two or more processing threads may perform method 600, with each thread executing one or more individual functions, routines, subroutines, or operations of method 600. In an illustrative example, the processing threads performing method 600 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads performing method 600 may execute asynchronously with respect to each other. The various operations of method 600 can be performed in an order different from that shown in FIG. 6. Some operations of method 600 may be performed concurrently with other operations. Some operations may be optional.

[0090] At operation 610, processing logic acquires input data associated with a plurality of time steps. For example, the plurality of time steps may include a current time step (t) and at least one time step prior to the current time step (e.g., t-1). The input data may include camera data and radar data. More specifically, the camera data may be acquired from one or more cameras associated with an autonomous vehicle (AV). The input data may include a plurality of images of the AV's driving environment, each image of the plurality of images corresponding to a respective time step of the plurality of time steps. Further details regarding acquiring the input data are described above with reference to FIG. 2.

[0091] At operation 620, processing logic extracts a plurality of BEV feature sets from the input data. Each BEV feature set of the plurality of feature sets corresponds to a respective time step of the plurality of time steps. In some implementations, each BEV feature set of the plurality of BEV feature sets includes a corresponding feature vector. Further details regarding the extraction of BEV feature sets are described above with reference to FIG. 2.

[0092] At operation 630, processing logic generates an object flow for at least one object using the multiple feature sets. In some implementations, generating the object flow includes performing multi-frame temporal aggregation. More specifically, performing multi-frame temporal aggregation can include generating an aggregated BEV feature set from the multiple BEV features. For example, performing multi-frame temporal aggregation can include warping a first feature set corresponding to a first time step to generate a warped BEV feature set, and combining the warped BEV feature set with a second feature set corresponding to a second time step subsequent to the first time step.

[0093] In some implementations, performing multi-frame temporal aggregation includes generating an aggregated BEV feature set as a concatenation-based BEV feature set by concatenating the warped BEV feature set and the second feature set. Further details regarding generating an aggregated BEV feature set as a concatenation-based BEV feature set are described above with reference to FIG. 4A and in further detail below with reference to FIG. 7.

[0094] In some implementations, performing multi-frame temporal aggregation includes generating the aggregated BEV feature set as an additive-based BEV feature set. Further details regarding generating the aggregated BEV feature set as an additive-based BEV feature set are described above with reference to FIG. 4B and in further detail below with reference to FIG. 7.

[0095] In some implementations, performing multi-frame temporal aggregation includes generating the aggregated BEV feature set as a Transformer-based BEV feature set. Further details regarding generating the aggregated BEV feature set as a Transformer-based BEV feature set are described above with reference to FIG. 4C and in further detail below with reference to FIG. 7.

[0096] Additionally or alternatively, in some implementations, generating the object flow includes performing multi-frame dense motion estimation using multiple feature sets. More specifically, performing multi-frame dense motion estimation can include generating at least one object flow for at least one object identified from the camera data. Further details regarding performing multi-frame dense motion estimation are described above with reference to Figures 5A-5B and in further detail below with reference to Figures 8A-8B.

[0097] At operation 640, the processing logic may modify the driving path in consideration of the at least one output. More specifically, the at least one output may include at least one of an aggregated BEV feature set or at least one object flow. For example, the at least one output may be processed by a data processing system of the AV (e.g., data processing system 120), and results of the processing by the data processing system may be provided to an AVCS of the AV (e.g., AVCS 140) to control the driving path of the AV. Further details regarding modifying the driving path in consideration of the at least one output are described above with reference to FIG. 1.

[0098] FIG. 7 illustrates an example method 700 for performing multi-frame temporal aggregation according to some implementations of the present disclosure. A processing device having one or more processing units (CPUs) and a memory device communicatively coupled to the CPU(s) may perform method 700 and / or each of their individual functions, routines, subroutines, or operations. A processing device executing method 700 may implement instructions issued by various components of sensing system 110 or data processing system 120 of FIG. 1 , such as EEPM 132. In some implementations, method 700 may be directed to systems and components of an autonomous vehicle, such as autonomous vehicle 100 of FIG. 1 . In some implementations, method 700 may be performed by EEPM 132 or any other similar model that may be part of the perception system of an autonomous vehicle, a vehicle employing driver assistance technology, or any other application platform that uses object detection and classification.

[0099] Method 700 can be used to improve the performance of data processing system 120 and / or AVCS 140. In certain implementations, a single processing thread may perform method 700. Alternatively, two or more processing threads may perform method 700, with each thread executing one or more individual functions, routines, subroutines, or operations of method 700. In an illustrative example, the processing threads performing method 700 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads performing method 700 may execute asynchronously with respect to one another. The various operations of method 700 can be performed in an order different from that shown in FIG. 7. Some operations of method 700 may be performed concurrently with other operations. Some operations may be optional.

[0100] At operation 710, processing logic obtains a BEV feature set associated with an initial time step. More specifically, the initial time step may be an initial time step in a sequence of time steps extending from the initial time step to the current time step. Illustratively, if the current time step is t, then the initial time step may be defined as tm, where m is a positive integer.

[0101] At operation 720, processing logic warps the BEV feature set to generate a warped BEV feature set. Further details regarding warping the BEV feature set are provided above.

[0102] At operation 730, processing logic obtains a BEV feature set associated with a subsequent time step. For example, the subsequent time step may be the time step immediately following an initial time step in the sequence of time steps. Illustratively, if the initial time step is defined as t, the subsequent time step may be defined as (t)+1.

[0103] At operation 740A, processing logic generates an aggregated BEV feature set. The aggregated BEV feature set may be generated by combining a BEV feature set associated with a subsequent time step with a warped BEV feature set generated from a BEV feature set associated with an initial time step.

[0104] In some embodiments, generating the aggregated BEV feature set includes generating a concatenation-based BEV feature set. More specifically, generating the concatenation-based BEV feature set includes concatenating a BEV feature set associated with a subsequent time step with the warped BEV feature set. Further details regarding generating a concatenation-based BEV feature set are described above with reference to FIG. 4A.

[0105] In some embodiments, generating the aggregated BEV feature set includes generating an addition-based BEV feature set. More specifically, generating the addition-based BEV feature set includes adding a BEV feature set associated with a subsequent time step and the warped BEV feature set. Further details regarding generating the addition-based BEV feature set are described above with reference to FIG. 4B.

[0106] In some embodiments, generating the aggregated BEV feature set includes generating a Transformer-based BEV feature set. More specifically, generating the Transformer-based BEV feature set includes transforming the warped BEV feature set to obtain a transformed warped BEV feature set, transforming the second BEV feature set to obtain a transformed second BEV feature set, and adding the transformed warped BEV feature set and the transformed second BEV feature set. Further details regarding generating a Transformer-based BEV feature set are described above with reference to FIG. 4C .

[0107] At operation 750, processing logic determines whether the subsequent time step is the same as the current time step. More specifically, determining whether the subsequent time step is the same as the current time step may include determining whether (tm)+1=t (i.e., the initial time step is t-1). If the subsequent time step is not the same as the current time step, this means that the subsequent time step is the time step that precedes the current time step in the sequence of time steps. The process can return to operation 720A to generate a second warped BEV feature set by warping the aggregated BEV feature set (e.g., a concatenation-based BEV feature set, an addition-based BEV feature set, or a transformer-based BEV feature set).

[0108] If the subsequent time step is the same as the current time step, this means that no additional time steps remain in the sequence. At operation 760A, processing logic may output an aggregated BEV feature set. The aggregated BEV feature set can be processed and used to modify the driving path (e.g., operation 650 of FIG. 6 ). For example, the aggregated BEV feature set can be used to cause an AVCS of the AV (e.g., AVCS 140 of FIG. 1 ) to control the operation of the AV.

[0109] FIG. 8A shows an example method 800A for performing multi-frame dense motion estimation according to some implementations of the present disclosure. A processing device having one or more processing units (CPUs) and a memory device communicatively coupled to the CPU(s) may perform method 800A and / or each of their individual functions, routines, subroutines, or operations. The processing device executing method 800A may implement instructions issued by various components of sensing system 110 or data processing system 120 of FIG. 1 , such as EEPM 132. In some implementations, method 800A may be directed to systems and components of an autonomous vehicle, such as autonomous vehicle 100 of FIG. 1 . In some implementations, method 800A may be performed by EEPM 132 or any other similar model that may be part of the perception system of an autonomous vehicle, a vehicle employing driver assistance technology, or any other application platform that uses object detection and classification.

[0110] Method 800A can be used to improve the performance of data processing system 120 and / or AVCS 140. In certain implementations, a single processing thread can perform method 800A. Alternatively, two or more processing threads can perform method 800A, with each thread executing one or more individual functions, routines, subroutines, or operations of method 800A. In an illustrative example, the processing threads performing method 800A can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads performing method 800A can be executed asynchronously with respect to each other. The various operations of method 800A can be performed in an order different from that shown in FIG. 8. Some operations of method 800A can be performed concurrently with other operations. Some operations may be optional.

[0111] At operation 810A, processing logic obtains a first BEV feature set associated with a first time step, a second BEV feature set associated with a second time step, and an initial flow. The first time step and the second time step may each be a time step in a sequence of time steps spanning from the initial time step to the current time step. More specifically, the second time step is the time step immediately following the first time step in the sequence of time steps. Illustratively, if the current time step is t and the first time step is t, the second time step may be defined as (t)+1, where m is a positive integer.

[0112] At operation 820A, processing logic warps the first BEV feature set to generate a warped BEV feature set.

[0113] At operation 830A, processing logic generates a correlation output by performing a correlation based on the second BEV feature set and the warped BEV feature set. For example, the correlation output may be a cost volume (e.g., a 4D cost volume).

[0114] At operation 840A, processing logic generates a context network output from the second BEV feature set.

[0115] At operation 850A, processing logic generates a final flow using the initial flow, the correlation output, and the context network output. For example, the context network output may be provided to a GRU. The output of the GRU may be combined with (e.g., summed) the initial flow to generate a modified flow. The correlation output (e.g., cost volume) may be provided to a lookup component to generate a lookup output. The output of the GRU and the lookup output may be provided as input to a next GRU. The output of the next GRU may be combined with the modified flow to obtain a second modified flow. A similar process may be performed until a final GRU is reached. The output of the final GRU may be combined with the penultimate modified flow to obtain a final flow (i.e., object flow). The final flow may be processed and used to modify the driving path (e.g., operation 650 of FIG. 6 ). For example, the final flow may be used to cause an AVCS of the AV (e.g., AVCS 140 of FIG. 1 ) to control the operation of the AV. Further details regarding operations 810A-850A are described with reference to Figures 5A-5B.

[0116] FIG. 8B shows an example method 800B for performing multi-frame dense motion estimation according to some implementations of the present disclosure. A processing device having one or more processing units (CPUs) and a memory device communicatively coupled to the CPU(s) may perform method 800B and / or each of their individual functions, routines, subroutines, or operations. A processing device executing method 800B may implement instructions issued by various components of sensing system 110 or data processing system 120 of FIG. 1 , such as EEPM 132. In some implementations, method 800B may be directed to systems and components of an autonomous vehicle, such as autonomous vehicle 100 of FIG. 1 . In some implementations, method 800B may be performed by EEPM 132 or any other similar model that may be part of the perception system of an autonomous vehicle, a vehicle employing driver assistance technology, or any other application platform that uses object detection and classification.

[0117] Method 800B can be used to improve the performance of data processing system 120 and / or AVCS 140. In certain implementations, a single processing thread can perform method 800B. Alternatively, two or more processing threads can perform method 800B, with each thread executing one or more individual functions, routines, subroutines, or operations of method 800B. In an illustrative example, the processing threads performing method 800B can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads performing method 800B can be executed asynchronously with respect to each other. The various operations of method 800B can be performed in an order different from that shown in FIG. 8. Some operations of method 800B can be performed concurrently with other operations. Some operations may be optional.

[0118] At operation 810B, processing logic obtains an aggregated BEV feature set. For example, the aggregated BEV feature set may be generated by combining a first BEV feature set associated with a first time step with a second BEV feature set associated with a second time step. More specifically, the first BEV feature set may be a warped BEV feature set generated by warping the BEV feature set. For example, the first time step may be a time step immediately before the current time step, and the second time step may be the current time step. Illustratively, if the current time step is t, the first time step may be t-1. However, such an example should not be considered limiting.

[0119] In some implementations, the aggregated BEV feature set includes a concatenation-based BEV feature set. Further details regarding concatenation-based BEV feature sets are described above with reference to FIGS. 4A and 7. In some embodiments, the aggregated BEV feature set includes an addition-based BEV feature set. Further details regarding addition-based BEV feature sets are described above with reference to FIGS. 4B and 7. In some embodiments, the aggregated BEV feature set includes a transformer-based BEV feature set. Further details regarding transformer-based BEV feature sets are described above with reference to FIGS. 4C and 7.

[0120] At operation 820B, processing logic extracts features from the aggregated BEV feature set to obtain extracted features. For example, the features can be extracted using a BEV feature network (e.g., a BEV backbone).

[0121] At operation 830B, processing logic generates an object flow based on the extracted features. The object flow can be processed and used to modify the driving path (e.g., operation 650 in FIG. 6). For example, the object flow can be used to cause an AVCS of the AV (e.g., AVCS 140 in FIG. 1) to control operation of the AV. Further details regarding operations 810B-830B are described above with reference to FIG. 5C.

[0122] FIG. 9 illustrates a block diagram of an exemplary computing device 900 capable of implementing multi-frame temporal aggregation and dense motion estimation for an autonomous vehicle (AV) according to some implementations of the present disclosure. The exemplary computing device 900 may be connected to other computing devices within a LAN, an intranet, an extranet, and / or the Internet. The computing device 900 may operate in the capacity of a server in a client-server network environment. The computing device 900 may be a personal computer (PC), a set-top box (STB), a server, a network router, a switch or bridge, or any device capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that device. Furthermore, while only a single example computing device is shown, the term “computer” shall also be considered to include any group of computers that, individually or jointly, execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.

[0123] An embodiment of a computing device 900 may include a processing device 902 (also referred to as a processor or CPU), a main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), etc.), a static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), and a secondary memory (e.g., a data storage device 918), which may communicate with each other via a bus 930.

[0124] Processing device 902 (which may include logic processing 903) represents one or more general-purpose processing devices, such as a microprocessor, a central processing device, or the like. More specifically, processing device 902 may be a complex instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that executes other instruction sets, or a processor that executes a combination of instruction sets. Processing device 902 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. According to one or more aspects of the present disclosure, processing device 902 may be configured to execute instructions that implement one or more of methods 600-800 of FIGS. 6-8.

[0125] The exemplary computing device 900 may further include a network interface device 608, which may be communicatively coupled to a network 920. An embodiment of the computing device 900 may further include a video display 910 (e.g., a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 912 (e.g., a keyboard), a cursor control device 914 (e.g., a mouse), and an audio signal generating device 916 (e.g., a speaker).

[0126] The data storage device 918 may include a computer-readable storage medium (or more specifically, a non-transitory computer-readable storage medium) 928 having stored thereon one or more sets of executable instructions 922. According to one or more aspects of the present disclosure, the executable instructions 922 may include executable instructions that implement one or more of the methods 600-800 of FIGS.

[0127] The executable instructions 922 may also reside, completely or at least partially, within the main memory 904 and / or within the processing device 902 during execution thereof by the exemplary computing device 900, with the main memory 904 and the processing device 902 also constituting computer-readable storage media. The executable instructions 922 may also be transmitted or received over a network via the network interface device 908.

[0128] While computer-readable storage medium 928 is shown in FIG. 9 as a single medium, the term "computer-readable storage medium" should be considered to include a single medium or multiple media (e.g., centralized or distributed databases, and / or associated caches and servers) that store one or more sets of operating instructions. The term "computer-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for execution by a machine, causing the machine to perform any one or more of the methods described herein. Thus, the term "computer-readable storage medium" should be considered to include, but not limited to, solid-state memory, and optical and magnetic media.

[0129] Some portions of the above detailed descriptions are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, understood to be a self-consistent sequence of steps leading to a desired result. The steps require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0130] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. Unless otherwise indicated, and as will be apparent from the discussion that follows, throughout the description, discussions utilizing terms such as "obtaining," "generating," "providing," "bringing about," "converting," "merging," "selecting," "performing," and the like will be understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's memory or registers, or other such information storage, transmission, or display device.

[0131] Examples of the present disclosure also relate to an apparatus for performing the methods described herein. This apparatus may be specially constructed for the required purposes, or it may be a general-purpose computer system selectively programmed by a computer program stored within the computer system. Such a computer program may be stored on a computer-readable storage medium, such as, but not limited to, any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random-access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0132] The methods and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems appears as set forth in the description below. Additionally, the scope of the present disclosure is not limited to any particular programming language. It will be understood that a variety of programming languages ​​can be used to implement the teachings of the present disclosure.

[0133] It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other example embodiments will become apparent to those skilled in the art upon reading and understanding the above description. While the present disclosure describes particular examples, it will be recognized that the systems and methods of the present disclosure are not limited to the examples described herein, but may be modified and practiced within the scope of the appended claims. Accordingly, the specification and drawings should be considered in an illustrative, and not a restrictive, sense. The scope of the present disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

1. obtaining, by a processing device, input data derived from a set of sensors associated with an autonomous vehicle (AV), the input data including camera data and radar data; extracting, by the processing device, a plurality of bird's eye view (BEV) feature sets from the input data, each BEV feature set corresponding to a respective time step; generating, by the processing device, an object flow for at least one object from the plurality of BEV feature sets, wherein generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation; and modifying, by the processing device, a driving path of the AV taking into account the object flow.

2. performing the multi-frame temporal aggregation obtaining a first set of BEV features associated with a first time step and a second set of BEV features associated with a second time step subsequent to the first time step; warping the first BEV feature set to obtain a warped BEV feature set; generating an aggregated BEV feature set; determining whether the second time step is the current time step; 2. The method of claim 1, further comprising: in response to determining that the second time step is the current time step, outputting the aggregated BEV feature set for generating the object flow.

3. The method of claim 2 , wherein generating the aggregated BEV feature set further comprises concatenating the warped BEV feature set and the second BEV feature set.

4. 3. The method of claim 2, wherein generating the aggregated BEV feature set further comprises adding the warped BEV feature set and the second BEV feature set based on a constant weight per grid cell.

5. 3. The method of claim 2, wherein generating the aggregated BEV feature set further comprises adding the warped BEV feature set and the second BEV feature set based on weights determined by a transformer.

6. performing multi-frame dense motion estimation obtaining a first set of BEV features associated with a first time step and a second set of BEV features associated with a second time step subsequent to the first time step; warping the first BEV feature set to generate a warped BEV feature set; generating a correlation output by performing a correlation based on the second BEV feature set and the warped BEV feature set; generating a context network output from the second BEV feature set; The method of claim 1 , further comprising: generating a final flow corresponding to the object flow using the correlation output, the context network output, and an initial flow.

7. performing multi-frame dense motion estimation Obtaining an aggregated BEV feature set; extracting features from the aggregated BEV feature set to obtain extracted features; The method of claim 1 , further comprising: generating the object flow based on the extracted features.

8. Memory and a processing device communicatively coupled to the memory, acquiring input data derived from a set of sensors associated with an autonomous vehicle (AV), the input data including camera data and radar data; extracting a plurality of bird's eye view (BEV) feature sets from the input data, each BEV feature set corresponding to a respective time step; generating an object flow for at least one object from the plurality of BEV feature sets, wherein generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation; and modifying a driving path of the AV in consideration of the object flow.

9. To perform the multi-frame temporal aggregation, the processing device: obtaining a first set of BEV features associated with a first time step and a second set of BEV features associated with a second time step subsequent to the first time step; warping the first BEV feature set to obtain a warped BEV feature set; generating an aggregated BEV feature set; determining whether the second time step is the current time step; 9. The system of claim 8, further configured to: in response to determining that the second time step is the current time step, output the aggregated BEV feature set for generating the object flow.

10. The system of claim 9 , wherein the aggregated BEV feature set is generated by concatenating the warped BEV feature set and the second BEV feature set.

11. 10. The system of claim 9, wherein the aggregated BEV feature set is generated by adding the warped BEV feature set and the second BEV feature set based on a constant weight per grid cell.

12. 10. The system of claim 9, wherein the aggregated BEV feature set is generated by adding the warped BEV feature set and the second BEV feature set based on weights determined by a transformer.

13. To perform the multi-frame dense motion estimation, the processing device: obtaining a first set of BEV features associated with a first time step and a second set of BEV features associated with a second time step subsequent to the first time step; warping the first BEV feature set to generate a warped BEV feature set; generating a correlation output by performing a correlation based on the second BEV feature set and the warped BEV feature set; generating a context network output from the second BEV feature set; The system of claim 8 , further configured to: generate a final flow corresponding to the object flow using the correlation output, the context network output, and an initial flow.

14. To perform the multi-frame dense motion estimation, the processing device: Obtaining an aggregated BEV feature set; extracting features from the aggregated BEV feature set to obtain extracted features; The system of claim 12 , further configured to: generate the object flow based on the extracted features.

15. A non-transitory computer-readable storage medium that, when executed by a processing device, causes the processing device to: acquiring input data derived from a set of sensors associated with an autonomous vehicle (AV), the input data including camera data and radar data; extracting a plurality of bird's eye view (BEV) feature sets from the input data, each BEV feature set corresponding to a respective time step; generating an object flow for at least one object from the plurality of BEV feature sets, wherein generating the object flow includes performing at least one of multi-frame temporal aggregation or multi-frame dense motion estimation; A non-transitory computer-readable storage medium having stored thereon instructions for performing an action including modifying a driving path of the AV in consideration of the object flow.

16. performing the multi-frame temporal aggregation obtaining a first set of BEV features associated with a first time step and a second set of BEV features associated with a second time step subsequent to the first time step; warping the first BEV feature set to obtain a warped BEV feature set; generating an aggregated BEV feature set; determining whether the second time step is the current time step; 16. The non-transitory computer-readable storage medium of claim 15, further comprising: in response to determining that the second time step is the current time step, outputting the aggregated BEV feature set for generating the object flow.

17. 17. The non-transitory computer-readable storage medium of claim 16, wherein generating the aggregated BEV feature set further comprises concatenating the warped BEV feature set and the second BEV feature set.

18. 17. The non-transitory computer-readable storage medium of claim 16, wherein generating the aggregated BEV feature set further comprises summing the warped BEV feature set and the second BEV feature set based on a constant weight per grid cell.

19. 17. The non-transitory computer-readable storage medium of claim 16, wherein generating the aggregated BEV feature set further comprises adding the warped BEV feature set and the second BEV feature set based on weights determined by a Transformer.

20. performing the multi-frame temporal aggregation obtaining a first set of BEV features associated with a first time step and a second set of BEV features associated with a second time step subsequent to the first time step; warping the first BEV feature set to generate a warped BEV feature set; generating a correlation output by performing a correlation based on the second BEV feature set and the warped BEV feature set; generating a context network output from the second BEV feature set; and generating a final flow corresponding to the object flow using the correlation output, the context network output, and an initial flow.

Citation Information

Patent Citations

  • Movement vector calculation method and movement vector calculation device

    JP2019020289A

  • Vehicle and control method thereof

    JP2019188932A