Object identification in a bird's-eye view reference system with explicit depth estimation cotraining.

By training neural networks with depth ground truth data to project camera images onto a bird's-eye view, the system addresses the challenges of costly multi-sensor object detection, achieving efficient and accurate object identification in autonomous vehicles.

JP7846827B2Active Publication Date: 2026-04-15WAYMO LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-08-07
Publication Date
2026-04-15

AI Technical Summary

Technical Problem

Existing object detection systems in autonomous vehicles rely heavily on multiple sensing modalities like LiDAR and cameras, which are costly and require significant computational resources, while camera images suffer from perspective distortion and lack immediate depth information, leading to inaccurate object mapping and loss of contextual information.

Method used

A system utilizing neural networks trained with depth ground truth data to estimate pixel depth and project images onto a bird's-eye view, enabling efficient object detection and classification using camera images alone, reducing the need for multiple sensing modalities.

Benefits of technology

Enables rapid and accurate object detection and classification with reduced computational overhead, suitable for autonomous vehicles and driver assistance systems, improving safety and efficiency by minimizing hardware and processing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007846827000002
    Figure 0007846827000002
  • Figure 0007846827000003
    Figure 0007846827000003
  • Figure 0007846827000004
    Figure 0007846827000004
Patent Text Reader

Abstract

The described aspects and embodiments deploy a bird's-eye view representation to enable efficient object detection and classification using a machine learning model trained using depth ground truth data. In one embodiment, a system and technique are disclosed that includes acquiring an image and generating a feature vector (FV) and depth distribution pixels of the image using a first neural network (NN), where the first NN is trained using training images and depth ground truth data for the training images. The technique further includes acquiring a feature tensor (FT) taking into account the FV and depth distribution, and processing the acquired FT using a second NN to identify one or more objects depicted in the image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification generally relates to systems and applications for detecting and classifying objects, and more particularly to autonomous vehicles and vehicles incorporating driver assistance technologies. More specifically, this specification relates to processing of through-camera images using machine learning techniques for faster and more resource-efficient detection and classification of objects, including, but not limited to, vehicles, pedestrians, bicycles, animals, and the like. [Background technology]

[0002] Autonomous (fully or partially autonomous) vehicles (AVs) operate by sensing the external environment using a variety of electromagnetic (e.g., radar and optical) and non-electromagnetic (e.g., voice and humidity) sensors. Some autonomous vehicles chart a driving path within the environment based on the sensed data. The driving path may be determined based on Global Navigation Satellite System (GNSS) data and roadmap data. GNSS and roadmap data can provide information about static aspects of the environment (buildings, street layout, road closures, etc.), while dynamic information (information about other vehicles, pedestrians, streetlights, etc.) is obtained from simultaneously collected sensing data. The accuracy and safety of the driving path, as well as the accuracy and safety of the speed regime selected by the autonomous vehicle, depend on the timely and accurate identification of various objects present in the driving environment and the ability of the driving algorithm to process information about the environment and provide correct instructions to vehicle control and the drivetrain. Brief explanation of the drawing

[0003] This disclosure is presented as an example, not an limitation, and may be more fully understood by referring to the following detailed description when considered in relation to the figures. [Brief explanation of the drawing]

[0004] [Figure 1]Figure 1 shows an exemplary autonomous vehicle (AV) component, according to some embodiments of the present disclosure, deploying a model that uses a bird's-eye view and is trained using depth ground truth data for efficient object detection and classification. [Figure 2] Figure 2 shows some exemplary architectures of perception systems capable of efficient object detection and classification, according to several embodiments of the present disclosure. [Figure 3A] Figure 3A is a schematic diagram illustrating exemplary operation of a model, according to several embodiments of the present disclosure, which uses a bird's-eye view and is trained using depth ground truth data for efficient object detection and classification. [Figure 3B] Figure 3B is a schematic diagram of one embodiment of a model that uses a bird's-eye view and is trained using depth ground truth data for efficient object detection and classification. [Figure 3C] Figure 3C shows a model in which a depth estimation network is pre-trained using depth ground truth data, according to some embodiments of the present disclosure. [Figure 4A] Figure 4A is a schematic diagram illustrating a distillation framework for training a bird's-eye view model using depth ground truth data, according to several embodiments of the present disclosure. [Figure 4B] Figure 4B is a schematic diagram illustrating another embodiment of the distillation framework for training a bird's-eye view model using depth ground truth data, according to some embodiments of the present disclosure. [Figure 5] Figure 5 is a schematic diagram illustrating the operation of a bird's-eye view model that provides instance segmentation and semantic segmentation according to some embodiments of the present disclosure. [Figure 6] Figure 6 is a schematic diagram illustrating the operation of a bird's-eye view model using temporal aggregation in some embodiments of the present disclosure. [Figure 7]Figure 7 shows an exemplary method of deploying a model trained using depth ground truth data for efficient detection and classification of objects using an aerial view representation, according to some embodiments of the present disclosure. [Figure 8] Figure 8 shows an exemplary method of using depth ground truth data to train a model that deploys an aerial view representation for efficient detection and classification of objects, according to some embodiments of the present disclosure. [Figure 9] Figure 9 shows a block diagram of an exemplary computer device that can operate and / or train a model trained using depth ground truth data for efficient detection and classification of objects using an aerial view, according to some embodiments of the present disclosure. **DETAILED DESCRIPTION**

[0005] <Summary> In one embodiment, a method is disclosed that includes obtaining one or more perspective camera images of an environment and, for each pixel of a set of pixels of the one or more perspective camera images, generating a feature vector (FV) and a depth distribution of a portion of the environment imaged by the corresponding pixel using a first neural network (NN). The first NN is trained using a plurality of training images and depth ground truth data for the plurality of training images. The method further includes obtaining a feature tensor (FT) for each pixel of the set of pixels, considering (i) the FV for each pixel and (ii) the depth distribution for each pixel. The method further includes using a second NN to process the obtained FT to identify one or more objects within the environment.

[0006] In another embodiment, a method for training a student model is disclosed. The method includes obtaining training images and using a first neural network (NN) of the student model to process the training images to generate a plurality of feature vectors (FV) and a plurality of depth distributions. Each FV of the plurality of FVs and each depth distribution of the plurality of depth distributions are associated with a respective pixel of a plurality of pixels of the training images. The method further includes obtaining a plurality of ground truth FVs generated by a first NN of an instructor model, where each ground truth FV of the plurality of ground truth FVs is associated with a respective pixel of a plurality of pixels of the training images. The method further includes obtaining a plurality of ground truth depth indicators, where each ground truth depth indicator of the plurality of ground truth depth indicators is associated with a respective pixel of at least a subset of the plurality of pixels of the training images. The method further includes adjusting parameters of the first NN of the student model. The adjustment is based on a comparison between the plurality of FVs and the plurality of ground truth FVs and is further based on a comparison between the plurality of depth distributions and the plurality of ground truth depth indicators.

[0007] In another embodiment, a system including a memory and a processing device is disclosed. The processing device is configured to obtain one or more perspective camera images of an environment and, using a first NN, generate for each pixel of a set of pixels of the one or more perspective camera images an FV and a depth distribution of a portion of the environment imaged by the respective pixel. The first NN is trained using a plurality of training images and depth ground truth data for the plurality of training images. The processing device is further configured to obtain a feature transform (FT) for each pixel of the set of pixels, considering (i) the FV for the respective pixel and (ii) the depth distribution for the respective pixel. The processing device is further configured to use a second NN to process the obtained FTs to identify one or more objects within the environment.

[0008] <Detailed Description> Various embodiments can be described below, but using autonomous driving systems and driver assistance systems as examples, it should be understood that the technologies and systems described herein can be used for tracking objects in a wide range of applications, including aerodynamics, marine applications, traffic control, animal control, industrial and academic research, public or personal safety, or any other application where automatic detection of objects is advantageous.

[0009] In one embodiment, for the safety of autonomous driving operations, it may be desirable to develop and deploy technologies for the rapid and accurate detection, classification, and tracking of various road users and other objects encountered on or near the road, such as road obstacles, construction equipment, and roadside structures. Autonomous vehicles (and various driver assistance systems) can utilize several sensors to facilitate the detection of objects in the driving environment and determine the movement of such objects. Sensors typically include radio wave detection and ranging sensors (radar), light detection and ranging sensors (lidar), various types of digital cameras, sonar, and position sensors. Different types of sensors offer various, and often complementary, benefits. For example, radar and lidar emit electromagnetic signals (radio signals and optical signals) reflected from objects, transmitting information that can determine the distance to the object (e.g., from the time of flight of the signal) and the speed of the object (e.g., from the Doppler shift of the signal frequency). Radar and lidar can cover a full 360-degree view, for example, by using a scanning transmitter of sensing beams. The sensing beam may produce numerous reflections, covering the operating environment within a high-density grid of return points. Each return point can be associated with the distance to the corresponding reflective object and the radial velocity of the reflective object (the component of the velocity along the line of sight).

[0010] Existing systems and methods for object identification and tracking use various sensing modalities, such as LiDAR, radar, and cameras, to acquire images of the environment. These images are then processed by trained machine learning models to identify the location of various objects within the image (e.g., bounding box shape), the state of their motion (e.g., velocity detected by LiDAR or radar Doppler effect-based sensors), the type of object (e.g., vehicle or pedestrian), and so on. Object motion (or any other evolution, such as splitting a single object into multiple objects) can be achieved by creating and maintaining tracks associated with specific objects.

[0011] Using multiple sensing modalities (e.g., LiDAR, radar, cameras) to acquire often complementary data improves the accuracy of object detection, identification, and tracking, but comes at a considerable cost for sensing hardware and processing software. For example, LiDAR sensors can provide valuable information about the distance to various reflective surfaces in the external environment. However, LiDAR sensors are expensive optical and electronic devices that operate by actively probing the external environment using optical signals and require considerable maintenance and periodic calibration. LiDAR returns (point clouds) must be processed, segmented into groups associated with distinct hypothetical objects, and matched with objects detected using other sensing modalities (e.g., cameras), which require additional processing and memory resources. Cameras, on the other hand, operate by passively collecting light (and / or infrared electromagnetic waves) emitted (or reflected) by objects in the environment and are significantly simpler and less expensive to design, install, and operate. As a result, various driver assistance systems that do not employ LiDAR (for cost and maintenance reasons) typically equip themselves with one or more cameras. Cameras can also be easily installed in various fixed locations and can be used for traffic monitoring and control, public and private safety applications, and more. Cameras based on optical or infrared imaging technology have certain advantages over radar, as they operate in a range of wavelengths that inherently have lower resolution compared to cameras, while enabling the detection of distance to objects (and their velocity). Therefore, the ability to detect and identify objects based solely on camera images is beneficial.

[0012] However, cameras project a three-dimensional (3D) external environment onto a two-dimensional imaging surface (e.g., the camera's photodetector array), which can be planar or curved. This presents two related challenges. On the one hand, the distance to an object (often referred to as the object's depth in the image) is not immediately known (but can often be determined from the context of the imaged object). On the other hand, camera images suffer from perspective distortion, so even if the number of pixels separating the object's image is the same, the distance between objects will differ depending on the object's depth. Furthermore, objects whose depictions are close to each other can nevertheless be separated by a considerable distance (e.g., a car and a pedestrian visible behind it). Existing machine learning techniques for object detection sometimes attempt to map objects to a top-down view, also known as a bird's-eye view (BEV), where objects are represented on a convenient manifold, such as a plane viewed from above, and characterized by a simple Cartesian coordinate set. Object identification and tracking can then be performed directly within the BEV representation. The success of such techniques depends on the accurate mapping of objects to the BEV. This then necessitates accurate estimation of distances to various objects, as misplacement of objects within the BEV can not only lead to errors in determining the distance to road users but also result in the loss of important contextual information.

[0013] Aspects and embodiments of this disclosure address these and other challenges of existing technologies by enabling methods and systems for using depth information training in machine learning object detection, classification, and tracking models. In particular, the disclosed technology includes a system of neural networks (NNs) trained to process fluoroscopy camera images. The first NN may be configured to estimate the probability that a pixel has several discretized depth values ​​so as to associate context with various pixels (e.g., in the form of feature vectors). As will be described in more detail below, the first neural network may be trained using ground truth-annotated training camera images, which include depth data for at least some pixels of the image. The depth data can be obtained by a suitable sensor capable of detecting the distance to an object, e.g., lidar, radar, sonar, or any other distance-aware sensor. During inference, the trained first NN may process a new set of fluoroscopy camera images and estimate depth for various pixels, for example, as the probability P(d) that a given pixel indicates an object (or part of an object) located at a distance d from the camera. The distance distribution, combined with the pixel feature vector FV(c) (where c is the context space index), can be used to obtain a feature tensor FT(c,d) that characterizes both the likely location (depth or distance d relative to the object) and the context of the object. Using feature tensors for various pixels with coordinates w, h in the image, a combined feature tensor {FT(c,d)} for the overall image can be obtained. w,h}→CFT(c,d,w,h) can be obtained. Next, a mapping transformation d,w,h→x,y,z can be performed from the perspective coordinates w,h,d (width, height, depth) to the Cartesian coordinate set x,y,z. Then, the combined feature tensor in the new coordinates, CFT(c,x,y,z), can be projected onto the horizontal plane by, for example, averaging or adding the elements of the combined feature tensor along the vertical column of pixels, to obtain the BEV projection context tensor, PCT(c,x,y)= Σi CFT(c,x,y,z i) can then be obtained. Next, the projected context tensor PCT(c,x,y) can be processed by a second trained NN that identifies objects, extracts semantic information (e.g., type of object), and performs precise localization of objects, etc.

[0014] Numerous variations of these techniques are described herein. In some embodiments, the first NN may include a first subnetwork trained to output a depth distribution P(d) and a separate second NN trained to output a feature vector FV(c). In some embodiments, the first and second NNs may be trained simultaneously. In such embodiments, the first and second NNs are subnetworks of an end-to-end model architecture trained together. In some embodiments, the first NN may be first trained using both training images and depth ground truth, and the second NN may be subsequently trained using training images but not depth ground truth. Any number of images taken simultaneously (or nearly simultaneously), for example, multiple images taken by a peripheral view camera (SVC) during a single camera cycle, can be processed simultaneously. In some implementations, images taken at different times can be processed simultaneously, for example, at a given time t j Using the images captured, each BEV context tensor PCT(c,x,y;t j) can be generated. Multiple BEV context tensors can then be processed at once by a second NN. Multiple depictions of the same object at the same or different locations at multiple times can increase the likelihood of correct segmentation and identification of objects in the image. In some embodiments, the second NN may have a common backbone and multiple classification heads. For example, one classification head may be trained to output the semantic segmentation of the input image. A second classification head may be trained to output the geometric centers of various objects. A third classification head may be trained to output the distances of various pixels in the BEV representation to the geometric centers of objects. The combined outputs of the classification heads can be used to provide object boundary identification along with object classification (type, class). In some embodiments, the training of the first NN and / or the second NN can be carried out using a teacher-student distillation framework. More specifically, the output of a teaching model trained using data acquired using multiple sensing modalities (camera, lidar, radar, sonar, etc.) can be used as ground truth for training the student model NN. A student model can be a distillation of the teaching model, for example, a model in which the number of neurons in the neuronal layer and / or within a particular layer is reduced. As a result, the student model can be more easily implemented on vehicles with less powerful processing and memory resources while maintaining the substantial functionality of the teaching model.

[0015] The embodiments described deviate from conventional object detection and classification paradigms that complement BEV segmentation techniques with efficient co-training of models using depth ground truth data. Benefits of the embodiments described include, but are not limited to, rapid and accurate detection, identification, and tracking of objects in a way that avoids the significant computational overhead of processing data from multiple sensing modalities. Since machine learning models trained and deployed as disclosed herein enable efficient object detection based on camera images, the models can be deployed on a variety of platforms, including systems with reasonable computational resources, such as autonomous vehicles and vehicles equipped with driver assistance technologies.

[0016] Figure 1 shows components of an exemplary autonomous vehicle (AV) 100, according to several embodiments of the present disclosure, deploying a model that uses a bird's-eye view and is trained using depth ground truth data for efficient detection and classification of objects. An autonomous vehicle may include motor vehicles (such as cars, trucks, buses, motorcycles, buggy vehicles, recreational vehicles, and any special agricultural or construction vehicles), aircraft (such as airplanes, helicopters, and drones), marine vessels (such as ships, boats, yachts, and submarines), spacecraft (controllable objects operating outside the Earth's atmosphere), or any other self-driving vehicle capable of operating in an autonomous driving mode (without or with reduced human input) (e.g., robots, factory and warehouse robotic vehicles, sidewalk delivery robotic vehicles, etc.).

[0017] Vehicles such as those described herein may be configured to operate in one or more different driving modes. For example, in manual driving mode, the driver may directly control acceleration, deceleration, and steering via inputs such as the accelerator pedal, brake pedal, and steering wheel. Vehicles may also operate in one or more autonomous driving modes, including, for example, semi-autonomous driving mode or partially autonomous driving mode, in which a person exercises some degree of direct or remote control over driving actions, or fully autonomous driving mode, in which the vehicle handles driving actions without direct or remote control by a person. These vehicles may be known by different names, such as autonomous driving vehicles, self-driving vehicles, etc.

[0018] As described herein, in semi-autonomous or partially autonomous driving modes, the vehicle assists with one or more driving operations (e.g., steering, braking, and / or acceleration for lane centering, adaptive cruise control, advanced driver-assistance systems (ADAS), or emergency braking), but the human driver is expected to situationally perceive the vehicle's surroundings and supervise the assisted driving operations. Here, the vehicle may perform all driving tasks in certain situations, but the human driver is expected to take control as needed.

[0019] For the sake of simplification and brevity, various systems and methods are described below in conjunction with autonomous vehicles, but similar technologies may be used in various driver assistance systems that do not reach the level of fully autonomous driving systems. In the United States, the Society of Automotive Engineers (SAE) defines different levels of autonomous driving operations to indicate how much or how little a vehicle controls the driving, although different organizations in the United States or other countries may classify the levels differently. More specifically, the disclosed systems and methods may be used in SAE Level 2 driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, and other driver support. The disclosed systems and methods may be used in SAE Level 3 driver assistance systems that enable autonomous driving under limited conditions (e.g., highways). Similarly, the disclosed systems and methods may be used in vehicles using SAE Level 4 autonomous driving systems that operate autonomously under most normal driving conditions and require only occasional attention from a human operator. In all such driver assistance systems, accurate lane estimation can be performed automatically without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, as well as overall safety for autonomous, semi-autonomous, and other driver assistance systems. As stated above, in addition to the way the SAE classifies levels of autonomous driving operations, other organizations in the United States or other countries may classify levels of autonomous driving operations differently. The systems and methods disclosed herein may be used in driver assistance systems defined by the levels of autonomous driving operations of these other organizations, but are not limited to those disclosed herein.

[0020] The driving environment 101 may include any (moving or stationary) objects located outside the AV, such as roads, buildings, trees, bushes, sidewalks, bridges, mountains, other vehicles, pedestrians, bridge piers, embankments, landing strips, animals, and birds. The driving environment 101 may be an urban area, suburban area, rural area, etc. In some embodiments, the driving environment 101 may be an off-road environment (e.g., agriculture or other farmland). In some embodiments, the driving environment may be an indoor environment, such as an industrial plant environment, a shipping warehouse, or a hazardous area of ​​a building. In some embodiments, the driving environment 101 may be substantially flat, with various objects moving parallel to the surface (e.g., parallel to the surface of the Earth). In other embodiments, the driving environment may be three-dimensional and may include objects (e.g., balloons, fallen leaves, etc.) that are capable of moving along all three directions. Hereinafter, the term “driving environment” should be understood to include all environments in which the motion of a vehicle equipped with autonomous motion (e.g., SAE Level 5 and SAE Level 4 systems), conditional autonomous motion (e.g., SAE Level 3 systems), and / or driver assistance technologies (e.g., SAE Level 2 systems) may occur. Furthermore, “driving environment” may include any possible flight environment of an aircraft (or spacecraft) or marine environment of a ship. Objects in the driving environment 101 may be located at any distance from the AV, from a few feet (or less) to several miles (or more).

[0021] The AV100 in this embodiment may include a sensing system 110. The sensing system 110 may include various electromagnetic (e.g., optical, infrared, radio, etc.) and non-electromagnetic (e.g., acoustic) sensing subsystems and / or devices. The sensing system 110 may include one or more lidars 112, which may be laser-based units capable of determining the distance to an object in the driving environment 101 and the velocity of the object. The sensing system 110 may include one or more radars 114, which may be any systems that utilize radio or microwave frequency signals to detect objects in the driving environment 101 of the AV100. The lidars 112 and / or radars 114 may be configured to sense both the spatial position of an object (including its spatial dimensions) and its velocity (e.g., using Doppler shift techniques). Hereinafter, “velocity” refers to both how fast an object is moving (object speed) and the direction of its motion. Each of the lidar(s) 112 and radar(s) 114 may include a coherent sensor, such as a frequency-modulated continuous-wave (FMCW) lidar or radar sensor. For example, the lidar(s) 112 and / or radar(s) 114 may use heterodyne detection for velocity determination. In some embodiments, the functionality of ToF and coherent lidar(or radar) is combined into a lidar(or radar) unit that can simultaneously determine both the distance to a reflective object and its radial velocity. Such a unit may be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode using heterodyne detection) or both modes simultaneously. In some embodiments, multiple lidar(s) 112 and / or radar(s) 114 can be mounted on the AV100.

[0022] Lidar 112 (and / or radar 114) may include one or more light sources (and / or radio / microwave sources) that generate and emit signals, and one or more detectors for signals reflected from objects. In some embodiments, Lidar 112 and / or radar 114 may be able to perform a 360-degree scan in the horizontal direction. In some embodiments, Lidar 112 and / or radar 114 may be able to scan space along both the horizontal and vertical directions. In some embodiments, the field of view may be up to 90 degrees vertically (e.g., at least a portion of the area above the horizon is scanned by the Lidar or radar signal). In some embodiments (e.g., aerospace applications), the field of view may be spherical (consisting of two hemispheres).

[0023] The sensing system 110 may further include one or more cameras 118 to capture images of the driving environment 101. The cameras 118 may operate in the visible portion of the electromagnetic spectrum, for example, in the wavelength range of 300–800 nm (also referred to here as the optical range for simplicity). Some of the optical range cameras 118 may use a global shutter, while others may use a rolling shutter. The image may be a two-dimensional projection of the driving environment 101 (or a portion of the driving environment 101) onto the projection surface (flat or non-flat) of the camera. Some of the cameras 118 of the sensing system 110 may be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101. The sensing system 110 may also include one or more sonars 116 for active audio probes of the driving environment 101, such as ultrasonic sonars, and one or more microphones for passively listening to sounds in the driving environment 101. The sensing system 110 may also include one or more infrared range cameras 119, also referred to herein as IR cameras 119. The IR cameras 119 may use focusing optics (e.g., made from germanium-based materials, silicon-based materials, etc.) configured to operate in the wavelength range of microns to tens of microns or more. The IR cameras 119 may include a phase array of IR detector elements. The pixels of the IR image produced by the camera(s) 119 may represent the total amount of IR radiation collected by each detector element (associated with the pixel), the temperature of a physical object from which the IR radiation is collected by each detector element, or any other suitable physical quantity.

[0024] The sensing data acquired by the sensing system 110 may be processed by the data processing system 120 of the AV 100. For example, the data processing system 120 may include a perception system 130. The perception system 130 may be configured to detect and track objects in the driving environment 101 and to recognize the detected objects. For example, the perception system 130 may be able to analyze images captured by the camera 118 and have the ability to detect traffic signals, road signs, road layout (e.g., lane boundaries, intersection topology, parking lot designations, etc.), the presence of obstacles, etc. The perception system 130 may also further receive radar sensing data (Doppler data and ToF data) to determine the distance to various objects in the environment 101 and the speed of such objects (in the line of sight and, in some embodiments, laterally, as described below). In some embodiments, the perception system 130 may use radar data in combination with data captured by the camera 118, as described in more detail below.

[0025] The perception system 130 may include one or more modules to facilitate efficient and reliable detection, identification, and tracking of objects, including an object detection and classification model (ODCM) 132 with depth cotraining, which can be used to process data provided by the sensing system 110. In some embodiments, during inference, the ODCM 132 may receive data from (optical range) cameras 118 and / or IR cameras 119. During training, as will be described in more detail in relation to Figures 2 and 4A-B, the ODCM 132 may process data from cameras 118 and / or IR cameras 119, while using range (distance) data acquired from at least some lidar 112, radar 114, sonar 116, etc. The ODCM 132 may include one or more trained machine learning models (MLMs) used to process the received images to detect objects depicted in the images and to classify the detected objects.

[0026] The perception system 130 may receive further information from a global navigation satellite system (GNSS) positioning subsystem (not shown in Figure 1), which may include a GNNS transceiver (not shown) configured to acquire information about the AV's position relative to the Earth and its surroundings. The positioning subsystem may use positioning data (e.g., GNNS and inertial measurement unit (IMU) data) in conjunction with sensing data to help accurately determine the AV's position relative to fixed objects in the driving environment 101 (e.g., roadways, lane boundaries, intersections, sidewalks, crosswalks, road signs, curbs, surrounding buildings, etc.) whose positions may be provided by map information 124. In some embodiments, the data processing system 120 may receive non-electromagnetic data such as voice data (e.g., ultrasonic sensor data from sonar 116, or data from a microphone picking up an emergency vehicle siren), temperature sensor data, humidity sensor data, pressure sensor data, and meteorological data (e.g., wind speed and direction, precipitation data).

[0027] The data processing system 120 may further include an environmental monitoring and prediction component 126, which can monitor how the driving environment 101 evolves over time, for example, by tracking the position and velocity (relative to the Earth) of moving objects. In some embodiments, the environmental monitoring and prediction component 126 can track the changing appearance of the environment due to the movement of AV relative to the environment. In some embodiments, the environmental monitoring and prediction component 126 can make predictions about how various moving objects in the driving environment 101 will be positioned within a predicted time frame. The predictions may be based on the current state of the moving object, including its current position (coordinates) and velocity. Furthermore, the predictions may be based on the history (tracked dynamics) of the moving object's movement over a specific period preceding the current moment. For example, based on stored data about a first object showing the movement of the first object accelerated over the previous three seconds, the environmental monitoring and prediction component 126 may conclude that the first object has resumed its movement from a stop sign or red light. Therefore, the environmental monitoring and prediction component 126 can predict where the first object is likely to be in the next 3 or 5 seconds of movement, taking into account the layout of the roadway and the presence of other vehicles. As another example, based on stored data about the second object showing deceleration movement of the second object in the previous 2 seconds, the environmental monitoring and prediction component 126 can conclude that the second object is stopped at a stop sign or red light. Therefore, the environmental monitoring and prediction component 126 can predict where the second object is likely to be within the next 1 or 3 seconds. The environmental monitoring and prediction component 126 can periodically check the accuracy of its predictions and modify them based on new data obtained from the sensing system 110. The environmental monitoring and prediction component 126 can operate in conjunction with the ODCM 132. For example, the environmental monitoring and prediction component 126 can track the relative movement of AV and various objects (e.g., reference objects that are stationary or moving relative to the Earth).

[0028] Data generated by the perception system 130, the GNSS processing module 122, and the environmental monitoring and prediction component 126 can be used by an autonomous driving system such as the AV control system (AVCS) 140. The AVCS 140 may include one or more algorithms that control how the AV should behave in various driving situations and environments. For example, the AVCS 140 may include a navigation system for determining a global driving route to a destination. The AVCS 140 may also include a driving route selection system for selecting a specific route through the immediate driving environment, which may include lane selection, avoiding traffic congestion, selecting a place to make a U-turn, and selecting a trajectory for parking maneuvers. The AVCS 140 may also include an obstacle avoidance system for safely avoiding various obstacles (rocks, stationary vehicles, etc.) in the AV's driving environment. The obstacle avoidance system may be configured to evaluate the size of the obstacle and its trajectory (if the obstacle is moving) and select the optimal driving strategy (e.g., braking, steering, acceleration, etc.) to avoid the obstacle.

[0029] The AVCS140 algorithms and modules can generate instructions for various systems and components of a vehicle, including the powertrain, brakes, and steering 150, vehicle electronics 160, signaling 170, and other systems and components not explicitly shown in Figure 1. The powertrain, brakes, and steering 150 may include the engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanism, and other systems. The vehicle electronics 160 may include the onboard computer, engine management, ignition, communication systems, car computer, telematics, in-car entertainment systems, and other systems and components. The signaling 170 may include high and low headlights, stop lights, turn signals and taillights, horns and alarms, interior lighting systems, dashboard notification systems, passenger notification systems, radio and wireless network transmission systems, etc. Some of the instructions output by AVCS140 can be delivered directly to the powertrain, brakes, and steering 150 (or signaling 170), while other instructions output by AVCS140 are first delivered to the vehicle electronics 160, which generates instructions for the powertrain, brakes, and steering 150, and / or signaling 170.

[0030] For example, the ODCM 132 may determine that images acquired by camera(s) 118 contain depictions of an object, and further classify the object as a bicycle. The environmental monitoring and prediction component 126 may track the bicycle and determine that it is moving at a speed of 15 mph along an intersecting road perpendicular to the direction of the vehicle's movement. In response to this determination, the data processing system 120 may determine that the vehicle needs to slow down to allow the bicycle to pass the intersection. The AVCS 140 may output instructions to the powertrain, brakes, and steering 150 (directly or via the vehicle electronics 160) such as: (1) reducing the engine speed by changing the throttle setting to reduce fuel flow to the engine; (2) shifting the drivetrain down to a lower gear via the automatic transmission; and (3) activating the brake unit (in cooperation with the engine and transmission) to reduce vehicle speed. After the ODCM132 and / or the environmental monitoring and prediction component 126 determine that the bicycle has crossed the intersection, the AVCS140 can output commands to the powertrain, brakes, and steering 150 to resume the vehicle's previous speed setting.

[0031] Figure 2 shows some exemplary architectures 200 of a perception system capable of efficient detection and classification of objects according to some embodiments of the present disclosure. Input to the perception system (e.g., perception system 130 in Figure 1) may include a plurality of camera images 202, which may be training images (during the training phase) or images taken onboard during runtime (during the estimation phase). The images 202 may be combined into a frame. A frame should be understood as any set of images showing the external environment along any direction to the sensing system (e.g., sensing system 110 in an autonomous vehicle). In particular, camera images 202 may refer to panoramic images taken by a peripheral view camera, directional cameras, e.g., a front view camera, a side view camera (SVC), a rear view camera, and similar, or images taken by any combination thereof. In some embodiments, images acquired by different cameras may be synchronized so that all images in a given frame have the same (to the highest degree of synchronization) timestamp. In some embodiments, some images in a given frame may have a time offset, e.g., a time offset associated with the scanning operation of an SVC (which can be controlled). The camera image 202 can be processed by the ODCM 132. Additional input to the ODCM 132 may include depth data 204, which may be data acquired using lidar sensors, radar sensors, sonar sensors, etc. In some implementations, depth data 204 may be used during training of the ODCM 132, but not during the inference phase. The image 202 and depth data 204 may include orientation indexing. More specifically, various pixels in the image 202 and return points in the depth data 204 may be associated with known orientations in space (e.g., from camera calibration). For example, the camera image may include an intensity map indexed by any suitable set of coordinates characterizing orientations in space, e.g., I(w,h) where w,h are Cartesian pixel coordinates (in the imaging plane), I(θ,φ) where θ and φ are polarity and azimuth angles, respectively, or any other set of coordinates.In some embodiments, a plurality of sets of coordinates can be used for different tasks and are facilitated by the stored mapping (transformation) between different sets. Depth data 204 can include, for example, the radial distance R(w, h) to an object associated with a particular pixel w, h in image 202, determined from the ToF of lidar / radar / sonar signals.

[0032] Each image 202 can have pixels of various intensities of one color I(w, h) (in the case of a black-and-white image) or multiple colors I c (w, h) (in the case of a color image). At least some of the images 202 can be infrared (IR) camera images obtained by an array of IR detectors (pixels) that can operate in a wavelength range from a few microns to tens of microns or more. The IR image can include an intensity IR(w, h) representing the total amount of IR radiation collected by each detector. In some embodiments, the IR image can include a pseudo-color map in which the presence of a particular pseudo-color can represent the total intensity IR(w, h) collected. In some implementations, the collected intensity can be used to determine a temperature map T(w, h) of the environment. Thus, in different embodiments, different representations (e.g., intensity map, pseudo-color map, temperature map, etc.) can be used to represent the IR camera data.

[0033] In some embodiments, the architecture 200 can include a normalization module (not shown in FIG. 2) that can change the size of each image 202 to match the size of the input to the ODCM132. In some implementations, the normalization module can further normalize the intensity of the pixels of the image 2 over a preset intensity range, such as [I min , I max , where I min is the minimum intensity and I max is the maximum intensity configured for the ODCM132 to process. In some embodiments, the minimum intensity can be zero, I min = 0. Further, the normalization module can perform other preprocessing of the image 202, including filtering, noise removal, etc.

[0034] The normalized and preprocessed images can be processed by various components of ODCM132 to detect the presence of objects 232 in the operating environment and classify the detected objects 232. ODCM132 may include any suitable machine learning model, such as lookup tables, geometric shape mapping, mathematical formulas, decision tree algorithms, support vector machines, deep neural networks, or any combination thereof. Deep neural networks may include convolutional neural networks, iterative neural networks (RNNs), fully connected neural networks, long-term short-term memory neural networks, Boltzmann machines, or any combination thereof.

[0035] The Depth Estimation Network (DEN) 210 can output predictions of the depth of an object imaged by various pixels in an image (or more) 202. The Context Feature Network (CFN) 220 can output feature vectors for various pixels. Feature vectors can be multi-element strings of data in feature space. Depth predictions and feature vectors can be combined into a feature tensor. The feature tensor can undergo one or more transformations, as will be described in more detail in relation to Figures 3A-C, mapping pixel data (intensity and depth) from a perspective view to a top-down BEV view. The feature tensor combined with the projected BEV context tensor can be processed by a BEV feature network 230 that outputs detected / classified objects 232, which can be classified into multiple classes, such as cars, trucks, buses, pedestrians, unknown objects, etc.

[0036] The detected / classified objects 232 may undergo post-processing 234, which may include object tracking to track the movement of the detected objects across multiple frames of the image 202. Each object may be assigned a detection track and may be characterized by some or all of the following: a bounding box for depicting a particular object across multiple frames, the type of object, the size of the object, the orientation of the object, and the motion of the object (e.g., velocity, acceleration, etc.). Post-processing 234 may further include generating any graphical, e.g., pixel-based (e.g., heatmap) or (curve-based) vectorized representation of the trajectory, including the trajectory, orientation, and velocity regime of various objects. In some embodiments, post-processing 234 may include processing the detected tracks using one or more models that predict the motion of the detected objects, e.g., models that track the velocity, acceleration, etc. of the detected objects. For example, a Kalman filter or any other suitable filter combining the predicted motion of a particular object with the detected motion of the object may be used to make a more accurate estimate of the object's position and motion.

[0037] The detected / classified objects 232 and the tracking data generated by post-processing 234 can be provided to the AVCS 140. The AVCS 140 evaluates the trajectories of the objects in various trajectories and, taking into account the position and velocity of the tracked objects, decides whether to modify the vehicle's current driving trajectory. For example, if a tracked pedestrian or cyclist is within a certain distance from the vehicle, the AVCS 140 may slow the vehicle down to a speed that ensures it can safely avoid the pedestrian or cyclist. Alternatively, the AVCS 140 may, for example, change lanes if there are no obstacles in the adjacent lane, or perform some other driving action.

[0038] Training may be performed by a training engine 242 hosted by a training server 240, which may be an external server deploying one or more processing units, such as a central processing unit (CPU) and a graphics processing unit (GPU). In some embodiments, the ODCM 132 may be trained by the training engine 242 and then downloaded to a vehicle deploying the perception system 130. The ODCM 132 can be trained using training data that includes training inputs 244 and corresponding target outputs 246 (correct matches for each training input), as shown in Figure 2. During training of the ODCM 132, the training engine 242 can find patterns in the training data that map each training input 244 to the corresponding target output 246.

[0039] In some implementations, the ODCM132 may be trained using images and other sensory data recorded and annotated with ground truth during a driving mission. For training the depth estimation network 210 of the ODCM132, the ground truth may include distances to various pixels in the image in the training input 244. Training the BEV feature network 230 may involve ground truth including correct identification of the locations (e.g., bounding boxes) of various objects in the training input 244, and semantic information (e.g., class, type, etc.) for the objects. In some embodiments, the ground truth for training any or all of the depth estimation network 210, the context feature network 220, and the BEV feature network 230 may include the outputs of one or more training models, as described later in conjunction with Figures 4A-B. The ground truth may include correct concatenation of the same object across batches of multiple images / frames, correct velocity of objects, etc.

[0040] The training engine 242 may have access to a data repository 250 that stores multiple camera / IR camera images 252 and lidar / radar (or sonar) images 254 acquired during driving conditions in various driving environments (e.g., urban driving missions, highway driving missions, rural driving missions, etc.). During training, the training engine 242 may select (e.g., randomly) a number of sets of camera / IR camera images 252 and a set of lidar / radar images 254 as training data. The training data can be annotated for correct object identification. In some embodiments, annotation can be performed by a developer before the annotated training data is placed in the data repository 250. The annotated training data acquired from the data repository 250 by the training server 240 may include one or more training inputs 244 and one or more target outputs 246. The training data may also include mapping data 248 that maps the training inputs 244 to the target outputs 246. For example, the mapping data 248 may identify the bounding box of a passenger's car in each of a batch of N consecutive frames acquired by the vehicle's forward-facing camera. The mapping data 248 may include the training data identifier, the position of the passenger car, the size and identification of the passenger car, the speed and direction of movement of the passenger car, and other suitable information.

[0041] During the training of ODCM132, the training engine 242 can evaluate the difference between the output of ODCM132 (or the various networks and subnetworks of ODCM132) and the target output 246 using an appropriate loss function 245. In some embodiments, different loss functions 245 may be used to train the depth estimation network 210, the context feature network 220, and / or the BEV feature network 230. During the training of ODCM132, the training engine 242 may modify the parameters (e.g., weights and biases) of the various networks and subnetworks of ODCM132 until the model minimizes the loss function 245 and successfully learns how to successfully identify and classify the target output 246, for example, various objects in the external environment. In some embodiments, the various networks and subnetworks of ODCM132 may be trained separately, for example, using the depth estimation network 210 which is trained before the context feature network 220 and / or the BEV feature network 230. In some embodiments, various networks and subnetworks of ODCM132 can be trained together (e.g., simultaneously). For example, the depth estimation network 210, the context feature network 220, and / or the BEV feature network 230 may be subnetworks of a single end-to-end neural network architecture that are trained together.

[0042] The data repository 250 can be persistent storage capable of storing camera / IR camera images, lidar / radar / sonar data, and data structures configured to facilitate detection and identification, according to embodiments of the present disclosure. The data repository 250 may be hosted by one or more storage devices such as main memory, magnetic or optical storage disks, tapes, or hard drives, network-attached storage (NAS), or storage area networks (SANs). Although shown separately from the training server 240, in embodiments the data repository 250 may be part of the training server 240. In some embodiments the data repository 250 may be a network-attached file server, while in other embodiments the data repository 250 may be some other type of persistent storage, such as an object-oriented database or a relational database (not shown in Figure 2), which may be hosted by a server machine or one or more different machines that can access the training server 240 via a network.

[0043] Figure 3A is a schematic diagram illustrating an example of the operation of Model 300, which uses a bird's-eye view and is trained using depth ground truth data for efficient object detection and classification, according to some embodiments of the present disclosure. The model shown in Figure 3A may be the ODCM 132 of the perceptual system 130 illustrated in Figure 1, or any other similar model. In some embodiments, the model shown in Figure 3A may include a set of neural networks (NNs), such as a depth estimation network (DEN) 310, a contextual feature network (CFN) 320, a BEF feature network (BEF FN) 350, etc. Although shown as separate blocks in Figure 3A, various exemplified NNs and subnetworks may be part of the same NN that are trained together. Neurons in a neural network are associated with learnable weights and biases. Neurons may be arranged in layers. Some of the layers may be hidden layers. Any of the NNs or subnetworks shown in Figure 3A may include multiple hidden neuron layers and may be configured to perform one or more functions that facilitate object detection and classification.

[0044] The input to the DEN310 and CFN320 may include one or more images 302, which may be perspective views acquired by one or more cameras. Any number of images 302 can be processed simultaneously. The input images 302 can depict any part of the external environment, up to a 360-degree panoramic periphery view. In some embodiments, the total number of pixels in all images 302 may be W × H, where W is the number of pixels along a first direction (e.g., horizontal) and H is the number of pixels along a second direction (e.g., vertical). In some embodiments, a “pixel” in image 302 may correspond to a single element of the camera’s charge-coupled device (CCD). In some embodiments, a “pixel” in image 302 may be a downsampled combination (superpixel) corresponding to multiple CCD elements of the camera. For example, one or more images from a camera may have a different number of pixels than W × H, e.g., the dimensions of the trained inputs to the DEN310 and CFN320. In these embodiments, the camera image can first be rescaled to a target input size W×H (e.g., using techniques such as interpolation, downsampling / upsampling). Similarly, the depth data 304 can be rescaled to a target input size W×H. In some embodiments, only a subset of W×H pixels of the depth data 304 may be available, for example, one depth value for every N pixels.

[0045] Image(s) 302 may be in any preferred digital format (JPEG, TIFF, GIG, BMP, CGM, SVG, etc.). Image(s) 302 may contain one or more intensity matrices I k It can be expressed via (w,h), where 0 ≤ w ≤ W and 0 ≤ h ≤ H, and the color index k has a single value (for a grayscale image), three values ​​(for an RGB image), four values ​​(for a CMYK image), etc. Intensity value I k This can be assumed to be a continuous or discrete value between 0 and 1 (or between any other selected limit, e.g., 0 and 255).

[0046] DEN310 calculates the probability P(d) that various pixels, for example, given pixels w and h, represent an object located at a distance d from the camera. w,h Any suitable neural network trained to identify a set of 312 depth distributions for can be any suitable neural network. In some implementations, the distance d is a number of intervals D, Δd1, Δd2, ..., Δd D It can be discretized between intervals Δd. i They do not need to be of equal size, but rather distance, for example, Δd1 ≤ Δd2 ≤ ... ≤ Δd D It can increase along with the last interval Δd. N This can extend from a specific distance (e.g., 100m, 200m, etc.) to an infinite distance.

[0047] The DEN310 can be trained using training images depicting objects in various environments, and using depth data 304 as ground truth. Depth data 304 can be provided by a lidar (radar, sonar, etc.) sensor(s) (multiple) and the actual distance d to the object (or part of the object) imaged by the corresponding camera pixels w, h. True (w, h) may be included. During the training phase, the training engine 242 in Figure 2 outputs the distribution P(d) from DEN310. w,h The center (e.g., mean, median, etc.) and the ground truth distance d True To minimize the loss function that characterizes the difference between (w,h), the parameters of the DEN310 neurons (e.g., weights and biases) can be modified. In some implementations, the distribution P(d) w,h This can be modeled as a categorical distribution, a Laplace distribution, a Gaussian distribution, or any other suitable distribution. The loss function minimized during training may include a focal loss function, a negative log-possibility loss function, or any other suitable loss function, such as a mean squares loss function or a cross-entropy loss function.

[0048] The CFN320 can be trained, for example, using the same training image to output a feature vector 322. The feature vector FV(c) is for pixels w,h. w,hrepresents a learned digital encoding that characterizes the appearance of the corresponding pixel and the graphical context provided by the other pixels of the image. The feature vector FV(c) can have any number of components (bits), c∈[1,C], selected considering the desired target accuracy for object detection and the specific computational resources on which the trained model (e.g., ODCM132) is deployed, such as C=32, 64, 128, 256, etc. Higher values ​​of C prefer higher accuracy for object detection and classification, while lower values ​​of C facilitate faster processing and / or deployment of the NN on systems with limited computational resources.

[0049] In some embodiments, DEN310 and CFN320 may be deep convolutional neural networks having a U-net architecture, for example, an encoder stage and a decoder stage. Each stage may include multiple convolutional neuron layers and one or more fully connected layers. The convolutions performed by each of DEN310 and CFN320 may include any number of convolutional kernels of different dimensions trained to capture both the local and global contexts of the input image. In some embodiments, DEN310 and CFN320 may be completely independent, for example, with no edges connecting the neurons of the two networks. In some embodiments, DEN310 and CFN320 may share a common backbone and may have separate heads that output depth distributions 312 and feature vectors 322. Each head may have any number of neuron layers and, in some embodiments, may have its own encoder-decoder architecture. In some implementations, at least a portion of the heads may include one or more fully connected layers.

[0050] The depth distribution 312 and the feature vector 322 can be combined into the feature tensor 332 by the lift transform 330. More specifically, the lift transform 330 is used to combine the feature vector FV(c) w,hEach pixel w,h described by can be interpolated with depth information from depth distribution 312. For example, the lift transform 330 can interpolate the corresponding depth distribution P(d) for the same pixel. w,h Using (dimension D × 1), each feature vector FV(c) w,h The cross product (of dimension C × 1) can be calculated. This allows us to calculate the feature tensor 332, for example, for pixels w, h, as schematically shown with the insert in Figure 3A.

number

[0051] Next, the feature tensor FT(c,d) calculated for each individual pixel w,h Using this method, for example, by concatenating feature tensor 332 for different pixels, a combined feature tensor 334 can be obtained for the entire image. {FT(c,d) w,h} → CFT(c,d,w,h). The combined feature tensor CFT(c,d,w,h) has dimensions C×D×W×H. The combined feature tensor 334 can then undergo a 2D mapping 340. More specifically, the 2D mapping 340 can generate a projected feature tensor 342 that uses a convenient set of planar coordinates, e.g., Cartesian coordinates x and y, or polar coordinates r and θ, within the ground plane.

[0052] 2D mapping 340 can be a bipartite transformation. During the first part, the viewpoint coordinates d, w, h can be transformed into 3D Cartesian coordinates d, w, h → x, y, z (or 3D cylindrical coordinates, w, h → r, θ, z), where z is the vertical coordinate (perpendicular to the ground). The transformation d, w, h → x, y, z can be a projection transformation, parameterized by the camera's focal length, the direction of the camera's optical axis, and other similar parameters. In an example where image 302 is acquired by multiple cameras (or cameras with a rotating optical axis), the transformation d, w, h → x, y, z can include multiple projection transformations, for example, using separate transformations used for pixels w, h provided by different cameras.

[0053] During the second part, the 2D mapping 340 can project the combined feature tensor, represented by the new coordinates CFT(c,x,y,z), onto the horizontal plane to obtain the projected (BEV) feature tensor 342. For example, to obtain the C×W×H projected feature tensor PCT(c,x,y), the combined feature tensor is assigned elements associated with each vertical pillar of the pixel, e.g., PCT(c,x,y) = Σi CFT(c,x,y,z i ) can be summed (or averaged). In some implementations, the coordinate z i The sum for each is different coordinate z i Different weights assigned to w i This can be done, and PCT(c,x,y)= Σi w i ×CFT(c,x,y,z i ) and for example, a larger weight w i However, pixels are assigned to image objects within a specific elevation from the ground (e.g., up to a few meters), and lower weights are assigned to other elevations (for example, to remove false objects such as tree branches or electrical cables that do not obstruct vehicle movement).

[0054] The projected feature tensor 342 characterizes the objects and their positions within the BEV with viewpoint distortion effectively eliminated. The BEV FN 350 can then process the projected feature tensor 342 to identify the objects shown in the image 302 and classify the identified objects. In some embodiments, the BEV FN 350 may be a network having both an encoder stage and a decoder stage. In some embodiments, the BEV FN 350 may be a network having a decoder stage, but the encoder stage is part of the CFN 320. In some embodiments, the BEV FN 350 may function as a backbone for one or more classification heads 360-n. Each classification head 360-n may output different kinds of information about the objects depicted in the image 302. For example, head A360-1 may output the bounding box of an object, head B360-2 may output the type and / or size of an object, head C360-3 may output the pose (position and orientation) of an object, and so on.

[0055] DEN 310, CFN 320, and BEV FN 350 are shown as separate blocks in Figure 3A, and each may have an encoder stage and a decoder stage. Other embodiments are also within the scope of this disclosure. More specifically, Figure 3B is a schematic diagram of one embodiment of Model 301, which uses a bird's-eye view and is trained using depth ground truth data for efficient detection and classification of objects. As shown in Figure 3B, DEN310 and CFN320 may be supported by a common backbone network, such as encoder network 305. Each of DEN310 and CFN320 may be a separate decoder trained to process the common output of encoder network 305. In some embodiments, CFN320 may be a decoder network, while DEN310 may be a head with one or more fully connected layers. In some embodiments, neither DEN310 nor CFN320 contains a decoder. In some embodiments, DEN310 may be a decoder network, while CFN320 may be a head with one or more fully connected layers. In some embodiments, one of DEN310 or CFN320 may be absent. For example, if CFN320 is absent, the feature vector 322 may be output directly by the encoder network 305, but DEN310 may include a decoder that further processes a copy of the output (feature vector 322) to generate a depth distribution 312.

[0056] In some embodiments, any or all of the encoder network 305, DEN310, CFN320, and BEV FN350 may be trained together. In some embodiments, some of the encoder network 305, DEN310, CFN320, or BEV FN350 may be trained stepwise. More specifically, Figure 3C shows Model 303 in which the depth estimation network 310 is pre-trained (together with the encoder network 305, if deployed) using depth ground truth data according to some embodiments of the present disclosure. The pre-training of DEN310 is performed at the center of the distribution P(d)w,h And the ground truth distance d provided by depth data 304 True The difference between (w,h) can be evaluated using an appropriate loss function. The difference can then be backpropagated through various layers of neurons in DEN310 (and encoder network 305, if deployed) until DEN310 (and encoder network 305) learns to predict the estimated depth of a pixel with the desired accuracy. After the pre-training 306 of DEN310 (and encoder network 305, if deployed) is complete, the output of DEN310 can be used as ground truth during the training of CFN320 and / or BEV FN350, which may be trained together or sequentially.

[0057] Figure 4A is a schematic diagram showing a distillation framework 400 for training a bird's-eye view model using depth ground truth data, according to some embodiments of the present disclosure. The distillation framework 400 may include training a student model 401-S using the output of a previously trained teaching model 401-T. The student model 401-S may include student DEN410-S, student CFN420-S, and student BEV FN450-S, which may operate substantially as described in relation to their respective networks in Figures 3A-C. In particular, student DEN410-S and student CFN420-S may process one or more images 402 and output depth distributions 412 and feature vectors 422 for various pixels of the image(s) 402, which can then undergo a lift-transform / 2D mapping 430, as well as a lift-transform 330 and 2D mapping 340, for example, to obtain a projected feature tensor 442. The Student BEV FN 450-S can process a projected feature tensor 442 to identify detected (and classified) objects 452.

[0058] Various outputs (including intermediate outputs) of the student ODCM401-S can be compared with ground truth obtained from the teaching ODCM401-T. In some embodiments, as shown in Figure 4A, the teaching ODCM401-T may have an architecture similar to that of the student ODCM401-S. In particular, the teaching DEN 410-T and teaching CFN 420-T can process one or more images 402, which may be the same images input to the student DEN 410-S and student CFN 420-S. Furthermore, the teaching DEN410-T can process depth data 404 (e.g., lidar range data) that associates the depth (distance to it) of the objects shown in the image 402. To accommodate additional depth data inputs 404, the teaching DEN410-T may have more input neurons compared to the student DEN-S 410-S. The teaching DEN 410-T and teaching CFN 420-T can output depth distributions 411 and feature vectors 421 for pixels of an image (or multiple image) 402, which undergo a lift transform / 2D mapping 430 to generate projected feature tensors 441, which are then processed by the teaching BEV FN 450-T to identify detected (and classified) objects 541.

[0059] To enable direct comparison between the various outputs (including intermediate outputs) of the student ODCM401-S and the outputs of the teaching ODCM401-T, at least some of the teaching DEN410-T, teaching CFN420-T, and teaching BEV FN450-T may have outputs (number of neurons in the output layer) of the same dimensions as the output dimensions of the student ODCM401-S, student DEN410-S, student CFN420-S, and student BEV FN450-S networks, respectively. On the other hand, the teaching networks may have greater complexity, including the number of neuron layers, the number of connections between layers, and the precision (number of bits) of the representation of the intermediate outputs. In some embodiments, the teaching DEN 410-T and teaching CFN 420-T may have more input neurons than their respective student networks and therefore may be configured to process image 402 at a higher resolution than the student networks. The Taught ODCM401-T can be used in off-board configurations, while the Student ODCM401-S can be used on vehicles, traffic monitoring stations, or any edge device with limited computing resources. As a result, the Taught ODCM401-T is not limited by complexity and / or computation time.

[0060] Student ODCM401-S may be a lightweight model that has substantially fewer neurons and neuronal layers than the teaching ODCM401-T, allowing for faster processing (given the same processing and memory resources). The distillation process to obtain Student ODCM401-S may include identifying and removing (culling, shearing) nodes and / or edges that have little or no effect on the model's output, and combining multiple neuron nodes and / or edges. Furthermore, the output of the teacher ODCM401-T can be used as ground truth during training of Student ODCM401-S, as schematically illustrated by the dashed arrows in Figure 4A. More specifically, the depth distribution 411 output by the teaching DEN410-T and the feature vector 421 output by the teaching CFN420-T can be used as ground truth during training of Student DEN410-S and Student CFN420-S. For example, the depth distribution 412 (feature vector 422) can be compared to the depth distribution 411 (feature vector 421) and the parameters of student DEN 410-S (student CFN 420-S) until the difference is minimized. Similarly, the projected feature tensor 441 and the detected object 451 can be used as ground truth values ​​compared to the projected feature tensor 442 and the detected object 452. In some embodiments, the difference between the detected object 452 and the ground truth detected object 451 can be backpropagated through student BEV FN 450-S but not through student DEN 410-S and / or student CFN 420-S. In these embodiments, the student DEN410-S is trained based on the difference between the depth distribution 412 and the ground truth depth distribution 444 (similarly, the student CFN420-S is trained based on the difference between the feature vector 422 and the ground truth feature vector 421).In other embodiments, the difference between the detected object 452 and the ground truth detected object 451 may be further backpropagated through the student DEN410-S and / or student CFN420-S.

[0061] Figure 4B is a schematic diagram showing another embodiment of the distillation framework 401 for training a bird's-eye view model using depth ground truth data, according to some embodiments of the present disclosure. Distillation framework 401 may differ from distillation framework 400 in Figure 4A in that there is no teacher DEN410-T, and instead of depth distribution 411, depth sensing data (e.g., lidar data or radar data) can be used as ground truth during training of student DEN410-S.

[0062] Figure 5 is a schematic diagram showing operation 500 of a bird's-eye view model providing instance segmentation and semantic segmentation according to some embodiments of the present disclosure. Operation 500 may include processing an image(s) 502 using a DEN 510 for outputting depth distributions 512, a CFN 520 for outputting feature vectors 522, and a lift transform / 2D mapping 530 for obtaining projected feature tensors 542. The DEN 510 may be trained using depth data, for example, as described in conjunction with Figures 3A-C and / or Figures 4A-B, or in any other similar manner. The BEV FN 550 may be a backbone network that processes the projected feature tensors 542 and then generates intermediate outputs provided to a plurality of classification heads 560-n. In some implementations, the classification head may include a semantic head 560-1 that outputs an instance segmentation map 562-1 that includes the classification of various locations of BEVs among several types, such as vehicles, vulnerable road users, roads, buildings, trees, and roadside structures. The classification head 560-n may include one or more instance segmentation heads. For example, an instance center head 560-2 may output the coordinates of the center 562-2 (e.g., the center of mass pixel) of various objects in the environment. An instance offset head 560-3 may output a map of offsets 560-3 that characterize the distance of various pixels of an object to its center. The various maps 562-n can be further processed by one or more layers of neurons (not shown in Figure 5) to obtain a detection and classification map (DCM) 564. The DCM 564 can combine the semantic and geometric (instance) segmentation generated by the classification heads 560-n to determine the location of an object and further identify the class of the object. For example, a “vehicle” object of the type identified by semantics head 560-1 may be further subdivided into classes such as “car,” “pickup truck,” “bus,” and “half-track,” taking geometric information into account.Similarly, the type "vulnerable road users" can be further subdivided into classes such as "pedestrians," "cyclists," "motorcyclists," and "skateboarders."

[0063] Figure 6 is a schematic diagram illustrating operation 600 of a bird's-eye view model using temporal aggregation according to some embodiments of the present disclosure. Operation 600 may include processing image(s) 602 using DEN610 for outputting depth distribution 612, CFN620 for outputting feature vectors 622, and lift transform / 2D mapping 630 for obtaining projected feature tensor 642. DEN610 may be trained using depth data, for example, as described in conjunction with Figures 3A-C and / or 4A-B, or in any other similar manner. Image 602 may be associated with frames taken at different times, t1, t2, t3... The operations of blocks 610-630 may be performed separately for different frames, for example, sequentially (using a single instance of DEN610 and CFN620) or in parallel (for example, using multiple instances of DEN610 and CFN620). Next, the projected feature tensor 642 can be warped using warping 644 to a common reference time, which may be one of the times t1, t2, t3..., e.g., the current time or the most recent available time (shown as time t3 in Figure 6). Warping 644 may be a mathematical transformation that eliminates the (independently known) motion of the sensing system (e.g., ego motion of an autonomous vehicle) and generates a warped feature tensor 645. As a result of warping 644, an object stationary relative to the ground is described by elements of the warped feature tensor 645 associated with the same coordinates x,y, while a moving object is described by elements spread along the direction of the object's motion. The warped feature tensor 645 can then be processed using an aggregation network 646, which may be a convolutional network with kernels extending across two or more temporal components of the warped feature tensor 645. The aggregation network 646 outputs an aggregation tensor 648 that can be processed by the BEV FN650 (and further processed by various classification heads 660-n) in the same way that the projected feature tensor is processed by the BEV feature network, for example, in the operation of Figures 3A-C or 4A-B.

[0064] Figures 7-8 illustrate exemplary methods 700-800 for deploying machine learning models that use a bird's-eye view and are trained using depth ground truth data for efficient object detection and classification. A processing unit having one or more processing units (CPUs) and memory devices communicably coupled to the CPU(s) can perform methods 700-800 and / or each of their individual functions, routines, subroutines, or operations. A processing unit performing methods 700-800 can execute instructions issued by various components of the sensing system 110 or data processing system 120 in Figure 1, such as the ODCM 132. In some embodiments, methods 700-800 may be directed to systems and components of an autonomous vehicle, such as the autonomous vehicle 100 in Figure 1. Methods 700-800 can be used to improve the performance of the processing system 120 and / or the autonomous vehicle control system 140. In a particular implementation, a single processing thread can perform methods 700-800. Alternatively, two or more processing threads may execute methods 700–800, each thread executing one or more individual functions, routines, subroutines, or operations of the method. In exemplary embodiments, the processing threads executing methods 700–800 may be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads executing methods 700–800 may execute asynchronously with respect to each other. The various operations of methods 700–800 may be executed in an order different from that shown in Figure 78. Some operations of methods 700–800 may be performed concurrently with others. Some operations may be optional.

[0065] Figure 7 shows an example of Method 700, which, according to some embodiments of the present disclosure, deploys a model that uses a bird's-eye view representation and is trained using depth ground truth data for efficient detection and classification of objects. Method 700 may use real-time images acquired by one or more cameras on a vehicle or by cameras mounted on any other suitable application platform. The cameras may be optical range cameras and / or IR cameras, including panoramic (peripheral view) cameras, partial panoramic cameras, high-definition (high-resolution) cameras, close-range cameras, cameras with a fixed field of view (relative to the vehicle), cameras with a dynamic (adjustable) field of view, cameras with a fixed or adjustable focal length, cameras with a fixed or adjustable numerical aperture, and any other suitable cameras. Optical range cameras may further include night vision cameras. The images acquired by the cameras may include various metadata that provides geometric associations between image pixels and the spatial location of objects, correspondences between pixels of different images, and so on. In some implementations, Method 700 may be implemented by any other similar model, which may be part of ODCM132, or part of a perception system in an autonomous vehicle, a vehicle implementing driver assistance technology, or any other application platform using object detection and classification.

[0066] In block 710, method 700 may include acquiring one or more fluoroscopy camera images of the environment (e.g., image 302 in Figures 3A-C). In block 720, method 700 may include using a first neural network (NN) to generate a feature vector (FV) and a depth distribution of the portion of the environment imaged by the corresponding pixel for each pixel in a set of pixels of one or more fluoroscopy camera images. In some embodiments, the set of pixels may include all pixels of one or more fluoroscopy camera images. In some embodiments, the set of pixels may include only a portion of all pixels of one or more fluoroscopy camera images.

[0067] The first NN may include one or more subnetworks, including encoder network 305, DEN310, CFN320, and / or other subnetworks. In some embodiments, the feature vector of each pixel in a set of pixels may be output by a first subnetwork of the first NN (e.g., CFN320). The depth distribution of each pixel in a set of pixels may be output by a second subnetwork of the first NN (e.g., DEN310). For example, the feature vector may be FV322 output by CFN320, and the depth distribution may be depth distribution 312 output by DEN310. In some embodiments, the first NN may be trained using a plurality of training images and depth ground truth data for the plurality of training images. In some embodiments, the depth ground truth data may include lidar determination distances for one or more objects in at least a subset of the plurality of training images. In some embodiments, a second subnetwork (e.g., DEN310) may be trained using depth ground truth data before training the first subnetwork (e.g., CFN320).

[0068] In some embodiments, the first NN may be part of a student model trained using distillation techniques, as illustrated in conjunction with Figures 4A-B, for example. More specifically, the depth ground truth data may include depth estimates for at least a subset of pixels in a plurality of training images. For example, depth estimates may be available for each Nth pixel in the training images. The depth estimates (e.g., depth distribution 411) may be output by the first NN (e.g., taught DEN410-S) of a teaching model (e.g., taught ODCM401-T).

[0069] In block 730, method 700 may include obtaining a feature tensor (FT) for each pixel in a set of pixels. The feature tensor (e.g., FT432) may be obtained considering the feature vectors of each pixel and the depth distribution of each pixel. For example, the feature tensor may be obtained using the cross product of the feature vectors and the depth distribution (for each pixel). Furthermore, obtaining the feature tensor may include performing a lift transform, for example, as described in relation to Figure 3A.

[0070] In block 740, method 700 may include processing the acquired feature tensor using a second NN to identify one or more objects in the environment. Processing the acquired feature tensor may include several operations, as shown in the callout portion of Figure 7. More specifically, in block 742, method 700 may include using each feature tensor of the set of pixels to obtain a combined feature tensor (e.g., combined FT334). Method 700 may then map the combined feature tensor to a ground plane (or any other reference plane) to obtain a projected feature tensor (e.g., projected FT342). For example, mapping a combined feature tensor to a ground plane may involve, in block 744, transforming the combined feature tensor to a set of coordinates associated with the ground plane (e.g., Cartesian coordinates x, y, z, cylindrical coordinates r, θ, z, or any other suitable set of coordinates), and in block 746, vertically aggregating the elements of the combined feature tensor to obtain a projected feature tensor. For example, to obtain a projected feature tensor, elements of the combined feature tensor associated with different values ​​of z (and the same values ​​of x, y or r, θ) can be added together.

[0071] In block 748, method 700 can continue to process the projected feature tensor using a second NN to identify one or more objects in the environment. In some implementations, the second NN may include a first classification head (e.g., semantic head 560-1 in Figure 5) configured to output semantic segmentation of one or more objects in the environment. The second NN may also include at least one second classification head (e.g., instance center head 560-2 and / or instance offset head 560-3) configured to output geometric information associated with the location of one or more objects in the environment. In some embodiments, for example, if the NN is trained using a distillation framework, the second NN (e.g., student BEV FN 450-S) may be trained using the output of the second NN (e.g., taught BEV FN 450-T) of a taught training model (e.g., taught ODCM 401-T).

[0072] In some embodiments, as illustrated in conjunction with Figure 6, for example, the model (e.g., ODCM132) can perform temporal aggregation of images acquired at different times. More specifically, one or more fluoroscopy camera images acquired in conjunction with block 710 may be associated with a first time (e.g., time t1). Temporal aggregation may include acquiring one or more additional fluoroscopy camera images associated at least a second time (e.g., times t2, t3, etc.). Temporal aggregation may then include generating additional projected feature tensors using the one or more additional fluoroscopy camera images, and performing concurrent processing of the projected feature tensors and the additional projected feature tensors (e.g., as exemplified by the processing of projected feature tensor 642). In some embodiments, concurrent processing may include the application of a warping transform (e.g., warping 644) and an aggregation NN (e.g., aggregation network 646). An aggregation neural network may include one or more convolutional kernels configured to aggregate elements of a projected feature tensor with elements of an additional projected feature tensor (or multiple additional feature tensors).

[0073] In embodiments where a perceptual system implementing method 700 is deployed on a vehicle, method 700 can continue in block 750 to cause the perceptual system to determine the vehicle's driving path, taking into account one or more identified objects in the vehicle's environment.

[0074] Figure 8 shows an exemplary method 800 of using depth ground truth data to train a model that unfolds a bird's-eye view representation for efficient detection and classification of objects, according to some embodiments of the present disclosure. Method 800 can be used to train ODCM132 or any other similar model. Method 800 can use previously recorded images and other sensing data acquired by scanning the environment of a vehicle (or any other relevant environment) using multiple sensors of the vehicle's sensing system, e.g., lidar, radar, sonar, etc. In block 810, Method 800 may include, for example, acquiring training images as part of a set of multiple training images. In block 815, Method 800 may include processing the training images to generate multiple feature vectors and multiple depth distributions using a first NN of the student model. For example, the first NN of the student model may include some or both of student DEN 410-S and student CFN 420-S, as shown, for example, in Figures 4A-B. Each feature vector of multiple feature vectors (e.g., FV422) and each depth distribution of multiple depth distributions (e.g., depth distribution 412) can be associated with each pixel of multiple pixels in the training image.

[0075] In block 820, method 800 may continue to acquire a plurality of ground truth feature vectors (e.g., feature vector 421) generated by a first NN of the teaching model. For example, the first NN of the teaching model may include some or both of the teaching DEN 410-T and teaching CFN 420-T. Each ground truth feature vector of the plurality of ground truth feature vectors may be associated with each pixel of a plurality of pixels in the training image. In block 825, method 800 may continue to acquire a plurality of ground truth depth indicators. Each ground truth depth indicator of the plurality of ground truth depth indicators may be associated with each pixel of at least a subset of a plurality of pixels in the training image. In some implementations, each of the plurality of ground truth depth indicators may include a depth distribution (e.g., depth distribution 411 as shown in Figure 4A) acquired by the first NN of the teaching model for the relevant pixel. In some embodiments, each of the multiple ground truth depth indicators may include a specific distance obtained by a distance sensing device to a portion of the environment captured by the associated pixel (e.g., depth data 404 as shown in Figure 4B).

[0076] In block 830, method 800 may include adjusting the parameters of a first neural network of the student model based on a comparison of multiple feature vectors (e.g., feature vector 422) with multiple ground truth feature vectors (e.g., feature vector 422) and a comparison of multiple depth distributions (e.g., depth distribution 412) with multiple ground truth depth indicators.

[0077] In block 835, method 800 may include obtaining multiple feature tensors. Each feature tensor of the multiple feature tensors may be obtained using the respective feature vectors of multiple feature vectors and the respective depth distributions of multiple depth distributions (for example, as illustrated in conjunction with Figure 3A). In block 840, method 800 may continue to obtain a combined feature tensor (e.g., combined feature tensor 334) using the multiple feature tensors. In block 845, method 800 may include mapping the combined feature tensor to a ground plane to obtain a projected feature tensor (e.g., using lift transform / 2D mapping 430). In block 850, method 800 may continue processing the projected feature tensor using a second NN of the Student model (e.g., Student BEV FN 450-S) to identify one or more objects in a training image (e.g., detected object 452). In block 855, method 800 may include acquiring one or more ground truth objects (e.g., detected object 451) identified by a second NN of a teaching model (e.g., teaching BEV FN 450-T) in a training image. In block 860, method 800 may continue to adjust the parameters of the second NN of the student model (e.g., student BEV FN 450-S) based on a comparison of one or more objects identified by the second NN of the student model with one or more objects identified by the second NN of the teaching model.

[0078] Figure 9 shows a block diagram of an exemplary computer device 900, according to several embodiments of the present disclosure, capable of operating and / or training a model that uses a bird's-eye view and is trained using depth ground truth data for efficient detection and classification of objects. Embodiments of computer device 900 may be connected to other computer devices in a LAN, intranet, extranet, and / or the Internet. Computer device 900 may operate as a server in a client-server network environment. Computer device 900 may be a personal computer (PC), set-top box (STB), server, network router, switch or bridge, or any device capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by such device. Furthermore, although only a single embodiment of a computer device is shown, the term “computer” shall also be considered to include any group of computers that individually or collectively execute a set of instructions (or sets of instructions) to carry out any one or more of the methods discussed herein.

[0079] Embodiments of the computer device 900 may include a processing unit 902 (also called a processor or CPU), main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM), static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), and secondary memory (e.g., data storage device 918), which can communicate with each other via bus 930.

[0080] The processing unit 902 (which may include the logic processing unit 903) represents one or more general-purpose processing units, such as a microprocessor, a central processing unit, or similar. More specifically, the processing unit 902 may be a composite instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that executes other instruction sets, or a processor that executes combinations of instruction sets. The processing unit 902 may also be one or more dedicated processing units, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or similar. According to one or more aspects of the present disclosure, the processing unit 902 may be configured to execute instructions for performing a method 700 of using depth ground truth data to train a model for deploying a bird's-eye view representation for efficient detection and classification of objects, and / or a method 800 of using depth ground truth data to train a model for deploying a bird's-eye view representation.

[0081] An embodiment of the computer device 900 may further include a network interface device 908 that can be communicatively coupled to a network 920. An embodiment of the computer device 900 may further include a video display 910 (e.g., a liquid crystal display (LCD), a touchscreen, or a cathode ray tube (CRT)), an alphanumeric input device 912 (e.g., a keyboard), a cursor control device 914 (e.g., a mouse), and an audio signal generator 916 (e.g., a speaker).

[0082] The data storage device 918 may include a computer-readable storage medium (or more specifically, a non-temporary computer-readable storage medium) 928 in which one or more sets of executable instructions 922 are stored. According to one or more aspects of the present disclosure, the executable instructions 922 may include executable instructions that perform a method 700 of using depth ground truth data to train a model that unfolds a bird's-eye view representation for efficient detection and classification of objects, and / or a method 800 of using depth ground truth data to train a model that unfolds a bird's-eye view representation.

[0083] The executable instruction 922 may also reside, all or at least partially, in the main memory 904 and / or the processing unit 902 while it is being executed by the exemplary computer device 900, and the main memory 904 and the processing unit 902 also constitute a computer-readable storage medium. The executable instruction 922 may also be further transmitted or received over a network via the network interface device 908.

[0084] Although the computer-readable storage medium 928 is shown as a single medium in Figure 9, the term “computer-readable storage medium” should be understood to include a single medium or multiple mediums (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of operational instructions. The term “computer-readable storage medium” should also be understood to include any medium capable of storing or encoding a set of instructions for machine execution that causes a machine to perform any one or more of the methods described herein. Thus, the term “computer-readable storage medium” should be understood to include, but not be limited to, solid memory, as well as optical and magnetic media.

[0085] Some parts of the detailed description above are presented with respect to algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic descriptions and representations are means used by those skilled in the field of data processing to most effectively communicate the content of the work to others skilled in the art. An algorithm is understood here, and generally, as a self-consistent set of steps that produce a desired result. The steps require the physical manipulation of physical quantities. Usually, but not always, these quantities take the form of electrical or magnetic signals that can be stored, moved, combined, compared, and otherwise manipulated. Referring to these signals as bits, values, elements, symbols, characters, terms, digits, or similar has sometimes proven convenient, primarily for reasons of general use.

[0086] However, it should be noted that all these terms and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specified, as will be evident from the following considerations, any considerations throughout the description using terms such as “specify,” “determine,” “store,” “adjust,” “produce,” “return,” “compare,” “generate,” “stop,” “load,” “copy,” “insert,” “replace,” “implement,” or similar terms are understood to refer to the actions and processes of a computer system or similar electronic computing device that manipulate and convert data represented as physical (electronic) quantities in the registers and memory of the computer system to other data similarly represented as physical quantities in the computer system memory or registers or other such information storage, transmission, or display devices.

[0087] Examples of this disclosure also relate to apparatus for carrying out the methods described herein. Such apparatus may be a general-purpose computer system that is specifically constructed for a required purpose or selectively programmed by computer programs stored within the computer system. Such computer programs may be stored on computer-readable storage media, each coupled to a computer system bus, including but not limited to disks of any type, such as optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of machine-accessible storage media, or any type of media suitable for storing electronic instructions.

[0088] The methods and displays presented herein are not inherently related to any particular computer or other device. Various general-purpose systems may be used with the programs taught herein, or it may be convenient to construct more specialized devices to carry out the necessary method steps. The necessary structures for various such systems will appear as described below. In addition, the scope of this disclosure is not limited to any particular programming language. It will be understood that various programming languages ​​may be used to carry out the teachings of this disclosure.

[0089] Naturally, the above description is intended to be illustrative and not restrictive. Many other embodiments will become apparent to those skilled in the art upon reading and understanding the above description. While this disclosure describes specific examples, it will be recognized that the systems and methods of this disclosure are not limited to the examples described herein and can be modified and implemented within the scope of the appended claims. Therefore, the specification and drawings should be considered illustrative, not restrictive. Accordingly, the scope of this disclosure should be determined by reference to the appended claims, along with the entire scope of equivalents to which such claims are entitled.

Claims

1. It is a method, Acquire images from one or more transparent cameras capturing the driving environment in which one or more vehicles are in motion, Using a first neural network (NN), for each pixel in the set of pixels of the one or more perspective camera images, Feature vectors (FVs) and A depth distribution for a portion of the driving environment captured by a corresponding pixel, wherein the first NN is trained using a plurality of training images and depth ground truth data for the plurality of training images, and generates a depth distribution. For each pixel in the set of pixels, (i) the FV for each pixel, and (ii) the depth distribution for each pixel are taken into consideration to obtain a feature tensor (FT), A method comprising processing the acquired FT using a second NN to identify one or more objects in the operating environment.

2. Processing the acquired FT is Using the FT for the aforementioned set of pixels, a combined FT is obtained, The above-mentioned combined FT is mapped to the ground surface to obtain the projected FT, The method according to claim 1, comprising processing the projected FT using the second NN.

3. Mapping the aforementioned combined FT to the ground surface is Convert the combined FT into a set of coordinates associated with the ground surface, The method according to claim 2, comprising aggregating the combined FT elements in a direction perpendicular to the ground surface to obtain the projected FT.

4. The one or more of the above-mentioned fluoroscopic camera images are associated with a first time, and the method is Obtain one or more additional fluoroscopic camera images associated with at least the second time point, Using the aforementioned one or more additional fluoroscopic camera images, an additional projected FT is generated. The method according to claim 2, further comprising performing simultaneous processing of the projected FT and the additional projected FT.

5. The method according to claim 4, wherein the simultaneous processing is performed by an aggregation NN comprising one or more convolution kernels configured to aggregate the elements of the projected FT with the additional elements of the projected FT.

6. The second NN is, A first classification head configured to output the semantic segmentation of one or more objects in the aforementioned operating environment, The method according to claim 1, further comprising: at least one second classification head configured to output geometric information associated with the positions of one or more objects in the operating environment.

7. The method according to claim 1, wherein the depth ground truth data includes depth estimates for at least a subset of pixels of the plurality of training images, and the depth estimates are output by a first NN of the teaching model.

8. The method according to claim 7, wherein the second NN is trained using the output of the second NN of the teaching model.

9. The method according to claim 1, wherein the FT for each pixel in the set of pixels is output by a first subnetwork of the first NN, the depth distribution for each pixel in the set of pixels is output by a second subnetwork of the first NN, and the second subnetwork is trained using the depth ground truth data before training the first subnetwork.

10. The method according to claim 1, wherein the depth ground truth data includes lidar determination distances for one or more objects in at least a subset of the plurality of training images.

11. A method for training student models, Acquiring training images, The training images are processed using the first neural network (NN) of the student model. Multiple feature vectors (FVs), To generate a plurality of depth distributions, wherein each of the plurality of FVs and each of the plurality of depth distributions is associated with each of the plurality of pixels in the training image, Obtaining a plurality of ground truth FVs generated by a first NN of the teaching model, wherein each of the plurality of ground truth FVs is associated with each of the plurality of pixels in the training image, Obtaining multiple ground truth depth indicators, wherein each of the multiple ground truth depth indicators is associated with each of the multiple pixels of the training image, The parameters of the first NN in the student model are as follows: Comparison of the plurality of FVs and the plurality of ground truth FVs, and A method comprising adjusting based on a comparison of the plurality of depth distributions and the plurality of ground truth depth indicators.

12. The method involves obtaining multiple feature tensors (FTs), wherein each of the multiple FTs is obtained using the respective FVs of the multiple FVs and the respective depth distributions of the multiple depth distributions. Using the aforementioned multiple FTs, a combined FT is obtained, The above-mentioned combined FT is mapped to the ground surface to obtain the projected FT, Using the second NN of the student model, the projected FT is processed to identify one or more objects in the training image, Obtaining one or more ground truth value objects identified by the second NN of the teaching model within the training image, Adjusting the parameters of the second NN of the student model based on a comparison between the one or more objects identified by the second NN of the student model and the one or more objects identified by the second NN of the teaching model, The method according to claim 11, further comprising:

13. Each of the aforementioned multiple ground truth depth indicators is (i) the depth distribution obtained by the first NN of the teaching model for the pixels associated with the ground truth depth indicator, and (ii) distance to a portion of the driving environment captured by a pixel associated with the ground truth depth indicator, which is acquired by a distance sensing device, wherein one or more vehicles are driving on the driving environment. The method according to claim 11, comprising at least one of the following.

14. It is a system, Memory and A processing unit that is communicatively coupled to the memory, Acquire images from one or more transparent cameras capturing the driving environment in which one or more vehicles are in motion, Using a first neural network (NN), for each pixel in the set of pixels of the one or more perspective camera images, Feature vectors (FVs) and A depth distribution for a portion of the driving environment captured by a corresponding pixel, wherein the first NN is trained using a plurality of training images and depth ground truth data for the plurality of training images, and generates a depth distribution. For each pixel in the set of pixels, (i) the FV for each pixel, and (ii) the depth distribution for each pixel are taken into consideration to obtain a feature tensor (FT), A processing apparatus is configured to use a second NN to process the acquired FT to identify one or more objects in the operating environment, A system equipped with these features.

15. In order to process the acquired FT, the processing device, Using the FT for the aforementioned set of pixels, a combined FT is obtained, The above-mentioned combined FT is mapped to the ground surface to obtain the projected FT, The system according to claim 14, wherein the second NN is used to process the projected FT.

16. In order to map the combined FT to the ground surface, the processing apparatus, Convert the combined FT into a set of coordinates associated with the ground surface, The system according to claim 15, wherein the combined FT elements are aggregated in a direction perpendicular to the ground surface to obtain the projected FT.

17. The one or more above-mentioned fluoroscopic camera images are associated with a first time, and the processing device, Obtain one or more additional fluoroscopic camera images associated with at least the second time point, Using the aforementioned one or more additional fluoroscopic camera images, an additional projected FT is generated. The system according to claim 15, further comprising performing simultaneous processing of the projected FT and the additional projected FT, wherein the simultaneous processing is performed by an aggregation NN comprising one or more convolutional kernels configured to aggregate elements of the projected FT with elements of the additional projected FT.

18. The second NN is, A first classification head configured to output the semantic segmentation of one or more objects in the aforementioned operating environment, The system according to claim 14, further comprising: at least one second classification head configured to output geometric information associated with the positions of one or more objects in the operating environment.

19. The system according to claim 14, wherein the FT for each pixel in the set of pixels is output by a first subnetwork of the first NN, the depth distribution for each pixel in the set of pixels is output by a second subnetwork of the first NN, and the second subnetwork is trained using the depth ground truth data before training the first subnetwork.

20. The system according to claim 14, wherein the depth ground truth data includes lidar determination distances for one or more objects in at least a subset of the plurality of training images.

Citation Information

Patent Citations

  • Vehicle recognition system and vehicle recognition method

    JP2020013480A

  • MULTI-VIEW DEEP NEURAL NETWORK FOR LiDAR PERCEPTION

    JP2021089723A

Cited By

  • Lane change prediction on highways

    US12673673B2

  • Lane change prediction on highways

    US20250222924A1