Preprocessed RADAR images for AI engine consumption

By using multi-signal classification algorithm and distributed hardware architecture to process FMCW RADAR data in autonomous robot systems, the problems of inaccurate azimuth estimation and excessive delay are solved, and the accuracy and efficiency of object detection are improved.

CN120303581APending Publication Date: 2025-07-11MOTIONAL AD LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380081091.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-28
Filing Date
2023-09-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When the existing FMCW RADAR is used in an autonomous robot system for object detection, the azimuth angle estimation is inaccurate and the cascading FFT preprocessing pipeline results in too long delays, affecting the accuracy and efficiency of object detection.

Method used

Multi-signal classification (MUSIC) algorithm is used to estimate azimuth in an independent preprocessing path, and combine distance and radial velocity to process using a distributed hardware architecture to reduce delays and improve accuracy.

Benefits of technology

It realizes more accurate azimuth estimation and lower delay processing, improving the accuracy and efficiency of object detection in autonomous robot systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303581A_ABST
    Figure CN120303581A_ABST
Patent Text Reader

Abstract

The disclosed embodiments provide a high precision, low latency preprocessing pipeline hardware architecture for Radar images. In an embodiment, a method comprises: receiving a time domain representation of a Radar image of a scene, the representation comprising at least one object; determining a distance to the at least one object based on a distance spectrum calculated from the time domain representation; determining a radial velocity of the at least one object based on a Doppler spectrum calculated from the time domain representation; determining an azimuth angle of the at least one object based on the Doppler spectrum; generating a data structure including the distance, the radial velocity and the azimuth angle; performing at least one machine learning process on the content of the data structure; and generating at least one control signal for controlling the vehicle based at least in part on a result of performing the machine learning process.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Autonomous robotic systems such as autonomous vehicles rely on a set of sensors to detect static or dynamic objects in a real-time operating environment. Detection of objects is typically performed by a perception subsystem of the autonomous robotic system, which includes a neural network backbone for real-time processing of large amounts of two-dimensional (2D) and / or three-dimensional (3D) sensor data and classification and localization of detected objects in the operating environment. The output of the perception subsystem is used by the planning system of the autonomous robotic system to plan a route through the operating environment.

[0002] Automotive sensors commonly used in advanced driver assistance systems (ADAS) and autonomous vehicles are frequency-modulated continuous wave (FMCW) RADARs. An FMCW RADAR transmits a sequence of frequency-modulated signals called chirps. The received signals reflected by objects in the vehicle's operating environment are recorded in a data structure (e.g., in a 3D tensor cube) for indicating chirp metrics, chirp sampling, and corresponding receiver antenna metrics in the time domain. In a preprocessing pipeline, this time-domain sensor data is transformed to the frequency domain, where the data is used to estimate the distance, radial velocity, and azimuth angle of the object relative to the RADAR using cascaded fast Fourier transforms (FFTs). These estimated parameters are then input into an artificial intelligence (AI) engine for object detection and localization.

[0003] The limitations of the above preprocessing method are that azimuth angle estimation is usually inaccurate, and when input into the AI engine, the azimuth angle estimation may produce overlapping bounding boxes for objects that are close to each other. Additionally, the cascaded FFT adds an undesirable delay to object detection / localization, which is not desirable for automotive sensing. BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1 is an example environment of a vehicle that can implement one or more components of an autonomous system;

[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;

[0006] Figure 3 is Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems;

[0007] Figure 4A is a diagram of certain components of an autonomous system;

[0008] Figure 4B is a diagram of an implementation of a neural network;

[0009] Figure 4C and Figure 4D is a diagram illustrating an exemplary operation of a CNN;

[0010] Figure 5 is a flowchart of a cascaded preprocessing pipeline for an ADC beat signal;

[0011] Figure 6 is a block diagram of a hardware architecture for high-precision, low-latency preprocessing of an ADC beat signal according to one or more embodiments;

[0012] Figure 7 is a flowchart of a method for high-precision, low-latency preprocessing of an ADC beat signal according to one or more embodiments; and

[0013] Figure 8 is for implementing according to one or more embodiments Figure 6 a block diagram of a chip layout of a computing unit of a preprocessing pipeline. DETAILED DESCRIPTION

[0014] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that the embodiments described herein may be practiced without these specific details. In some instances, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.

[0015] In the drawings, for ease of description, a specific arrangement or order of schematic elements (such as those representing systems, devices, modules, instruction blocks, and / or data elements, etc.) is illustrated. However, those skilled in the art will understand that unless explicitly described, the specific order or arrangement of schematic elements in the drawings is not intended to imply a required processing order or sequence, or a separation of processes. Additionally, unless explicitly described, including schematic elements in the drawings is not intended to imply that such elements are required in all embodiments, nor that the features represented by such elements cannot be included in some embodiments or combined with other elements in some embodiments.

[0016] In addition, in the drawings, connecting elements (such as solid lines, dashed lines, or arrows) are used to illustrate a connection, relationship, or association between or among two or more other schematic elements. The absence of any such connecting element is not intended to mean that a connection, relationship, or association cannot exist. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings so as not to obscure the present disclosure. Further, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of a signal, data, or instruction (e.g., "software instruction"), those skilled in the art will understand that such an element may represent one or more signal paths (e.g., a bus) that may be required to affect the communication.

[0017] Although terms such as "first," "second," and / or "third" etc. are used to describe various elements, these elements should not be limited by these terms. The terms "first," "second," and / or "third" are only used to distinguish one element from another. For example, without departing from the scope of the described embodiments, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.

[0018] The terms used in the description of the various embodiments herein are included only for the purpose of describing particular embodiments and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a," "an," and "the" are also intended to include the plural forms and may be used interchangeably with "one or more" or "at least one" unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein refers to and includes any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprises," "comprising," "includes," and / or "including" are used in this specification, they specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0019] As used herein, the terms "communicate" and "communicating" refer to at least one of receiving, receiving, transmitting, conveying, and / or providing information (or information represented by, for example, data, signals, messages, instructions, and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof, etc.) that is to communicate with another unit, this means that the unit can directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is inherently wired and / or wireless. Additionally, two units can communicate with each other even if the information transmitted between the first unit and the second unit can be modified, processed, relayed, and / or routed. For example, even if the first unit receives information passively and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message can refer to a network packet (e.g., a data packet, etc.) that includes data.

[0020] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "at the time of", "in response to determining as", and / or "in response to detecting", etc. Similarly, depending on the context, the phrase "if it has been determined" or "if [the stated condition or event] is detected" is optionally interpreted to mean "at the time of determining...", "in response to determining as", or "at the time of detecting [the stated condition or event]" and / or "in response to detecting [the stated condition or event]", etc. Additionally, as used herein, the terms "have", "having", or "possessing", etc. are intended to be open-ended terms. Further, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on".

[0021] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to those of ordinary skill in the art that the various described embodiments can be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0022] General Overview

[0023] The disclosed embodiments provide a high-precision, low-latency preprocessing pipeline hardware architecture for processing camera RADAR data. In the preprocessing pipeline, a first FFT (hereinafter referred to as "range-FFT") is used to estimate the distance of an object relative to the RADAR from the camera RADAR data. A second 2D FFT (hereinafter referred to as "Doppler-FFT") is used to estimate the radial velocity of the object relative to the RADAR from the camera RADAR data. Compared with existing camera RADAR preprocessing techniques, the azimuth angle relative to the RADAR is estimated in a preprocessing path independent of range and radial velocity by applying the Multiple Signal Classification (MUSIC) algorithm to the 2D Doppler spectrum output by the 2D Doppler-FFT. The estimated range, radial velocity, and azimuth angle are input into a data structure suitable for consumption by an AI engine. The AI engine (e.g., a CNN as described in Figures 4B to 4D ), as part of a perception system 402 as described in Figure 4A , can use these estimated parameters alone or in fusion with other sensor data to detect and classify the object and its location.

[0024] In some embodiments, a method includes: receiving, by at least one processor, a time-domain representation of a Radar image of a scene, the representation including at least one object; determining, by the at least one processor, a distance to the at least one object based on a range spectrum calculated from the time-domain representation; determining, by the at least one processor, a radial velocity of the at least one object based on a Doppler spectrum calculated from the time-domain representation; determining, by the at least one processor, an azimuth angle of the at least one object based on the Doppler spectrum; generating, by the at least one processor, a data structure including the distance, the radial velocity, and the azimuth angle; performing, by the at least one processor, at least one machine learning process on the content of the data structure; and generating, by the at least one processor, at least one control signal for controlling a vehicle at least in part based on a result of performing the machine learning process.

[0025] In some embodiments, the Multiple Signal Classification algorithm, i.e., the MUSIC algorithm, is used to generate the azimuth angle.

[0026] In some embodiments, the data structure is a three-dimensional range-angle-Doppler tensor, i.e., a 3D RAD tensor.

[0027] In some embodiments, the data structure is generated based on the type of the at least one processor.

[0028] In some embodiments, a system includes: a bus; a frame buffer; a frame buffer controller; and at least one processor, wherein: the bus transmits a time-domain representation of a Radar image, the representation including at least one object; the frame buffer controller stores an output in the frame buffer and retrieves the output from the buffer for further processing by the at least one processor; wherein the at least one processor: calculates a distance of the at least one object relative to the camera Radar based on a distance spectrum calculated from the time-domain representation; calculates a radial velocity of the at least one object based on a Doppler spectrum calculated from the time-domain representation; calculates an azimuth angle of the at least one object relative to the camera Radar based on the Doppler spectrum; generates a data structure storing the distance, the radial velocity, and the azimuth angle; performs at least one machine learning process on the content of the data structure; and generates at least one control signal for controlling a vehicle at least in part based on a result of performing the machine learning process.

[0029] In some embodiments, the system is a distributed hardware architecture, the at least one processor includes two or more processors, and the two or more processors include at least one of a neural processing unit (NPU), a graphics processing unit (GPU), a tensor processing unit (TPU), and an accelerator chip that share data through a computing fabric.

[0030] In some embodiments, the data structure is generated based on the type of the at least one processor.

[0031] By implementation of the systems and methods described herein, the disclosed high-precision, low-latency preprocessing hardware architecture for ADC signals output by a camera RADAR provides at least the following advantages: more accurate azimuth angle estimation and lower latency compared to existing cascaded FFT preprocessing pipelines for ADC signals.

[0032] Now refer to Figure 1, an exemplary environment 100 is illustrated, in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, environment 100 includes vehicles 102a - 102n, objects 104a - 104n, routes 106a - 106n, area 108, vehicle - to - infrastructure (V2I) devices 110, network 112, remote autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118. Vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 are interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections (e.g., establishing a connection for communication, etc.). In some embodiments, objects 104a - 104n are interconnected with at least one of vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, network 112, autonomous vehicle (AV) system 114, queue management system 116, and V2I system 118 via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0033] Vehicles 102a - 102n (individually referred to as vehicle 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or passengers. In some embodiments, vehicle 102 is configured to communicate with V2I device 110, remote AV system 114, queue management system 116, and / or V2I system 118 via network 112. In some embodiments, vehicle 102 includes cars, buses, trucks, and / or trains, etc. In some embodiments, vehicle 102 is the same as or similar to vehicle 200 described herein (see Figure 2 ). In some embodiments, vehicles 200 in the set of vehicles 200 are associated with an autonomous queue manager. In some embodiments, as described herein, vehicle 102 travels along corresponding routes 106a - 106n (individually referred to as route 106 and collectively referred to as routes 106). In some embodiments, one or more than one vehicle 102 includes an autonomous system (e.g., an autonomous system same as or similar to autonomous system 202).

[0034] The objects 104a - 104n (individually referred to as object 104 and collectively as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., building, sign, fire hydrant, etc.). Each object 104 (e.g., located at a fixed location and over a period of time) is stationary or (e.g., having a speed and associated with at least one trajectory) moving. In some embodiments, the object 104 is associated with a corresponding location in the region 108.

[0035] The routes 106a - 106n (individually referred to as route 106 and collectively as routes 106) are each associated with (e.g., defining) a sequence of actions (also referred to as a trajectory) that a connected AV can navigate along. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatio - temporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatio - temporal location different from the first spatio - temporal location) or a target zone (e.g., a subspace of acceptable states (e.g., termination states)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, the route 106 includes multiple acceptable sequences of states (e.g., multiple sequences of spatio - temporal locations), which are associated with (e.g., defining) multiple trajectories. In an example, the route 106 includes only high - level actions or imprecise state locations, such as a series of connected roads indicating a direction change at a roadway intersection, etc. Additionally or alternatively, the route 106 can include more precise actions or states, such as, for example, a specific target lane or precise location within a lane region and a target rate at those locations. In an example, the route 106 includes multiple precise state sequences along at least one high - level action with a finite look - ahead horizon to reach an intermediate target, where the combination of successive iterations of the finite - horizon state sequences cumulatively corresponds to multiple trajectories that together form a high - level route terminating at the final target state or zone.

[0036] Region 108 includes a physical region (e.g., a geographical region) that the vehicle 102 can navigate. In an example, region 108 includes at least one state (e.g., a country, a province, an individual state among multiple states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, region 108 includes at least one named arterial road (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, a city street, etc. Additionally or alternatively, in some examples, region 108 includes at least one unnamed road, such as a driveway, a section of a parking lot, a section of a vacant and / or undeveloped area, a dirt road, etc. In some embodiments, a road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, a road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.

[0037] A vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the platoon management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), a lane marking, a streetlight, a parking meter, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the platoon management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.

[0038] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a Long-Term Evolution (LTE) network, a third-generation (3G) network, a fourth-generation (4G) network, a fifth-generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.

[0039] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the queue management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the queue management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.

[0040] The queue management system 116 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ride-sharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems), etc.).

[0041] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the queue management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection different from the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipal authority or a private institution (e.g., a private institution for maintaining the V2I device 110, etc.).

[0042] Provide Figure 1 The number and arrangement of the illustrated elements are provided as examples. Compared with Figure 1 the illustrated elements, there may be additional elements, fewer elements, different elements, and / or elements with different arrangements. Additionally or alternatively, at least one element of the environment 100 may perform one or more functions described as being performed by Figure 1 at least one different element. Additionally or alternatively, at least one set of elements of the environment 100 may perform one or more functions described as being performed by at least one different set of elements of the environment 100.

[0043] Now refer to Figure 2 , the vehicle 200 (which may be the same as or similar to the vehicle 102 of Figure 1 ) includes an autonomous system 202, a powertrain control system 204, a steering control system 206, and a braking system 208, or is associated with the autonomous system 202, the powertrain control system 204, the steering control system 206, and the braking system 208. In some embodiments, the vehicle 200 is the same as or similar to the vehicle 102 (see Figure 1 ). In some embodiments, the autonomous system 202 is configured to endow the vehicle 200 with autonomous driving capabilities (e.g., implement at least one of the following driving functions, features, and / or devices that are automatic or based on maneuvering actions, and the at least one driving function, feature, and / or device that is automatic or based on maneuvering actions enables the vehicle 200 to operate partially or completely without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon dependence on human intervention, such as level 5 ADS-operated vehicles, etc.), highly autonomous vehicles (e.g., vehicles that abandon dependence on human intervention in certain situations, such as level 4 ADS-operated vehicles, etc.), and / or conditionally autonomous vehicles (e.g., vehicles that abandon dependence on human intervention in limited situations, such as level 3 ADS-operated vehicles, etc.), etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functionality required to enable the vehicle 200 to operate in road traffic and continuously perform a part or all of the dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) that includes driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., level 0) to full driving automation (e.g., level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International Standard J3016: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems, the entire content of which is incorporated by reference. In some embodiments, the vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.

[0044] The autonomous system 202 includes a sensor suite that includes one or more devices such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, the autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance the vehicle 200 has traveled, etc.). In some embodiments, the autonomous system 202 uses one or more devices included in the autonomous system 202 to generate data associated with the environment 100 described herein. The data generated by one or more devices of the autonomous system 202 can be used by one or more systems described herein to observe the environment (e.g., environment 100) in which the vehicle 200 is located. In some embodiments, the autonomous system 202 includes a communication device 202e, an autonomous vehicle computing 202f, a drive-by-wire (DBW) system 202h, and a safety controller 202g.

[0045] The camera 202a includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus 302 that is the same as or similar to Figure 3 the bus). The camera 202a includes at least one camera (e.g., a digital camera using an optical sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) for capturing images including physical objects (e.g., cars, buses, curbs, and / or people, etc.). In some embodiments, the camera 202a generates camera data as an output. In some examples, the camera 202a generates camera data that includes image data associated with the image. In this example, the image data can specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., and / or an image timestamp, etc.). In such an example, the image can be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured (e.g., positioned) on the vehicle for capturing images for the purpose of stereovision (stereo vision). In some examples, the camera 202a includes generating image data and transmitting the image data to the autonomous vehicle computing 202f and / or a queue management system (e.g., the same as Figure 1Multiple cameras of the same or similar queue management system as queue management system 116. In such an example, the autonomous vehicle computing 202f determines the depth to one or more objects in the fields of view of at least two of the multiple cameras based on image data from at least two cameras. In some embodiments, camera 202a is configured to capture images of objects within a distance relative to camera 202a (e.g., up to 100 meters and / or up to 1 kilometer, etc.). Thus, camera 202a includes features such as sensors and lenses that are optimized for sensing objects at one or more distances relative to camera 202a.

[0046] In an embodiment, camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, camera 202a generates traffic light data associated with one or more images. In some examples, camera 202a generates TLD (Traffic Light Detection) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, camera 202a that generates TLD data is different from other systems incorporating cameras described herein in that camera 202a may include one or more cameras having a wide field of view (e.g., a wide-angle lens, a fish-eye lens, and / or a lens having a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.

[0047] The Light Detection and Ranging (LiDAR) sensor 202b includes being configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., with Figure 3at least one device configured to communicate via a bus (e.g., a bus identical or similar to bus 302). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object it encounters. The LiDAR sensor 202b also includes at least one light detector that detects the light after the light emitted from the light emitter encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing the objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundary of the physical object in the field of view of the LiDAR sensor 202b.

[0048] A Radio Detection and Ranging (Radar) sensor 202c includes at least one device configured to communicate with a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g via a bus (e.g., a bus identical or similar to Figure 3 bus 302). The Radar sensor 202c includes a system configured to emit (pulsed or continuous) radio waves. The radio waves emitted by the Radar sensor 202c include radio waves within a predetermined spectrum. In some embodiments, during operation, the radio waves emitted by the Radar sensor 202c encounter a physical object and are reflected back to the Radar sensor 202c. In some embodiments, the radio waves emitted by the Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the Radar sensor 202c generates a signal representing the objects included in the field of view of the Radar sensor 202c. For example, at least one data processing system associated with the Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of the physical object in the field of view of the Radar sensor 202c.

[0049] The microphone 202d includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus the same or similar to the bus 302 of Figure 3 ). The microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, the microphone 202d includes a transducer device and / or a similar device. In some embodiments, one or more of the systems described herein may receive the data generated by the microphone 202d and determine the position (e.g., distance, etc.) of an object relative to the vehicle 200 based on the audio signal associated with the data.

[0050] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the DBW (drive-by-wire) system 202h. For example, the communication device 202e may include a device the same or similar to the communication interface 314 of Figure 3 . In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (e.g., a device for enabling wireless communication of data between vehicles).

[0051] The autonomous vehicle computing 202f includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the safety controller 202g, and / or the DBW system 202h. In some examples, the autonomous vehicle computing 202f includes devices such as a client device, a mobile device (e.g., a cellular phone and / or a tablet, etc.), and / or a server (e.g., a computing device including one or more central processing units and / or graphics processing units, etc.). In some embodiments, the autonomous vehicle computing 202f is configured to implement the autonomous vehicle software 400 described herein. In some embodiments, the autonomous vehicle computing 202f is the same or similar to the distributed computing architecture 500 described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114 of Figure 1 ), a queue management system (e.g., a queue management system the same or similar to the queue management system 116 of Figure 1 ), a V2I device (e.g., a V2I device the same or similar to the Figure 1the same or similar V2I devices as the V2I device 110) and / or a V2I system (e.g., a V2I system the same or similar to the V2I system 118) Figure 1 communicate with the same or similar V2I systems as the V2I system 118).

[0052] The safety controller 202g includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the autonomous vehicle computing 202f, and / or the DBW system 202h. In some examples, the safety controller 202g includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). In some embodiments, the safety controller 202g is configured to generate control signals that take precedence over (e.g., override) the control signals generated and / or transmitted by the autonomous vehicle computing 202f.

[0053] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (such as turn signals, headlights, door locks, and / or windshield wipers, etc.).

[0054] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to perform longitudinal vehicle movements (such as starting to move forward, stopping moving forward, starting to move backward, stopping moving backward, accelerating in a certain direction, decelerating in a certain direction, etc.) or perform lateral vehicle movements (such as making a left turn and / or making a right turn, etc.). In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (such as fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.

[0055] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates two front wheels and / or two rear wheels of the vehicle 200 left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activities required to regulate the y-axis component of the vehicle's movement.

[0056] The braking system 208 includes at least one device configured to actuate one or more brakes to decelerate the vehicle 200 and / or keep it stationary. In some examples, the braking system 208 includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle 200 on the corresponding rotors of the vehicle 200. Additionally or alternatively, in some examples, the braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.

[0057] In some embodiments, the vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring the nature of the state or condition of the vehicle 200. In some examples, the vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel braking pressure sensor, a wheel torque sensor, an engine torque sensor, and / or a steering angle sensor, etc. Although the braking system 208 is illustrated as being located Figure 2 proximal to the vehicle 200 within, the braking system 208 can be located anywhere within the vehicle 200.

[0058] Now refer to Figure 3 , a schematic diagram of the exemplary device 300. As illustrated, the device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, the device 300 corresponds to: at least one device of the vehicle 102 (e.g., at least one device of the system of the vehicle 102); and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112). In some embodiments, one or more devices of the vehicle 102 (e.g., one or more devices of the system of the vehicle 102), and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112) include at least one device 300 and / or at least one component of the device 300. As Figure 3As shown, device 300 includes bus 302, processor 304, memory 306, storage component 308, input interface 310, output interface 312, and communication interface 314.

[0059] Bus 302 includes components that permit communication among the components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), and / or a neural processing unit (NPU), etc.), a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic and / or static storage device that stores data and / or instructions for use by processor 304 (e.g., flash memory, magnetic memory, optical memory, and / or dynamic RAM (DRAM), etc.).

[0060] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, a CD-ROM, RAM, PROM, EPROM, FLASH-EPROM, NV-RAM, and / or another type of computer-readable medium, and corresponding drives.

[0061] Input interface 310 includes components that permit device 300 to receive information such as via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, input interface 310 includes sensors for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). Output interface 312 includes components for providing output information from device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).

[0062] In some embodiments, communication interface 314 includes transceiver-like components (e.g., transceivers and / or separate receivers and transmitters, etc.) that permit device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, an interface, and / or a cellular network interface, etc.

[0063] In some embodiments, device 300 performs one or more processes described herein. Device 300 performs these processes based on software instructions stored by a computer-readable medium such as memory 306 and / or storage component 308, executed by processor 304. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes storage space located within a single physical storage device or storage space distributed across multiple physical storage devices.

[0064] In some embodiments, software instructions are read into memory 306 and / or storage component 308 from another computer-readable medium or from another device via communication interface 314. When executed, the software instructions stored in memory 306 and / or storage component 308 cause processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with software instructions to perform one or more processes described herein. Thus, unless otherwise explicitly stated, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0065] Memory 306 and / or storage component 308 includes a data store or at least one data structure (e.g., a database, etc.). Device 300 is capable of receiving information from, storing information in, communicating information to, or searching for information stored in the data store or at least one data structure in memory 306 or storage component 308. In some examples, the information includes network data, input data, output data, or any combination thereof.

[0066] In some embodiments, device 300 is configured to execute software instructions stored in memory 306 and / or the memory of another device (e.g., another device that is the same as or similar to device 300). As used herein, the term "module" refers to at least one instruction stored in memory 306 and / or the memory of another device, which, when executed by processor 304 and / or the processor of another device (e.g., another device that is the same as or similar to device 300), causes device 300 (e.g., at least one component of device 300) to perform one or more processes described herein. In some embodiments, the module is implemented in software, firmware, and / or hardware, etc.

[0067] Provide Figure 3 The number and arrangement of the illustrated components are provided as examples. In some embodiments, compared to Figure 3 the illustrated components, device 300 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of device 300 (e.g., one or more components) may perform one or more functions described as being performed by another component or another set of components of device 300.

[0068] Referring now to FIG. 4, an example block diagram of an autonomous vehicle software 400 (sometimes referred to as an “AV stack”) is illustrated. As illustrated, the autonomous vehicle software 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a localization system 406 (sometimes referred to as a localization module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the localization system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of the vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the localization system 406, the control system 408, and the database 410 are included in one or more separate systems (e.g., one or more systems that are the same as or similar to the autonomous vehicle software 400, etc.). In some examples, the perception system 402, the planning system 404, the localization system 406, the control system 408, and the database 410 are included in one or more separate systems located in the vehicle and / or in at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle software 400 are implemented in software (e.g., software instructions stored in a memory), by computer hardware (e.g., by a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA), etc.), a die, or a distributed computing architecture. It will also be understood that, in some embodiments, the autonomous vehicle software 400 is configured to communicate with remote systems (e.g., an autonomous vehicle system that is the same as or similar to the remote AV system 114, a queue management system that is the same as or similar to the queue management system 116, and / or a V2I system that is the same as or similar to the V2I system 118, etc.).

[0069] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect the at least one physical object), and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with one or more physical objects within the field of view of the at least one camera (e.g., representing the one or more physical objects). In such examples, the perception system 402 classifies the at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical objects by the perception system 402, the perception system 402 transmits data associated with the classification of the physical objects to the planning system 404.

[0070] In some embodiments, the planning system 404 receives data associated with a destination and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data from the perception system 402 (e.g., the data associated with the classification of the physical objects described above), and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 102 in road traffic. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate speed, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with the updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.

[0071] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) the location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In certain examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on the respective point clouds. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with two-dimensional (2D) and / or three-dimensional (3D) maps of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of the roadway geometry, a map describing the connection nature of the road network, a map describing the physical properties of the roadways (such as traffic rate, traffic flow, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs or various other types of driving signal lights, etc.). In some embodiments, the map is generated in real time based on the data received by the perception system.

[0072] In another example, the positioning system 406 receives Global Navigation Satellite System (GNSS) data generated by a Global Positioning System (GPS) receiver. In some examples, the positioning system 406 receives GNSS data associated with the location of a vehicle in an area, and the positioning system 406 determines the latitude and longitude of the vehicle in the area. In such examples, the positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, the positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the positioning system 406 determining the position of the vehicle, the positioning system 406 generates data associated with the position of the vehicle. In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic properties corresponding to the position of the vehicle.

[0073] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to cause the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204, etc.), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208) to operate. For example, the control system 408 is configured to perform operational functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes activities required to regulate the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes activities required to regulate the x-axis component of the vehicle motion. In an example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.

[0074] In some embodiments, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model alone or in combination with one or more of the above systems. In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in the environment, etc.). The following are examples regarding Figures 4B to 4D the implementation of machine learning models.

[0075] The database 410 stores data transmitted to, received from, and / or updated by the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408. In some examples, the database 410 includes a storage component of at least one system for storing operation-related data and / or software and using the autonomous vehicle software 400 (e.g., associated with Figure 3the same or similar storage components as the storage component 308). In some embodiments, the database 410 stores data associated with 2D and / or 3D maps of at least one area. In some examples, the database 410 stores data associated with 2D and / or 3D maps of a part of a city, multiple parts of multiple cities, multiple cities, counties, states, and / or countries (e.g., nations), etc. In such examples, a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200) can drive along one or more drivable areas (e.g., single-lane roads, multi-lane roads, highways, backroads, and / or off-road paths, etc.), and cause at least one LiDAR sensor (e.g., a LiDAR sensor the same or similar to the LiDAR sensor 202b) to generate data associated with an image representing the objects included in the field of view of the at least one LiDAR sensor.

[0076] In some embodiments, the database 410 can be implemented across multiple devices. In some examples, the database 410 is included in a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114), a queue management system (e.g., a queue management system the same or similar to Figure 1 the queue management system 116), and / or a V2I system (e.g., a V2I system the same or similar to Figure 1 the V2I system 118), etc.

[0077] Now refer to Figure 4B , a diagram illustrating the implementation of a machine learning model. More specifically, a diagram illustrating the implementation of a convolutional neural network (CNN) 420. For illustrative purposes, the following description of the CNN 420 will be with respect to implementing the CNN 420 via the perception system 402. However, it will be understood that in some examples, the CNN 420 (e.g., one or more components of the CNN 420) is implemented by other systems different from or in addition to the perception system 402, such as the planning system 404, the positioning system 406, and / or the control system 408, etc. Although the CNN 420 includes certain features as described herein, these features are provided for illustrative purposes and are not intended to limit the present disclosure.

[0078] The CNN 420 includes a plurality of convolutional layers including a first convolutional layer 422, a second convolutional layer 424, and a convolutional layer 426. In some embodiments, the CNN 420 includes a subsampling layer 428 (sometimes referred to as a pooling layer). In some embodiments, the subsampling layer 428 and / or other subsampling layers have dimensions that are smaller than the dimensions of the upstream system (i.e., the number of nodes). By means of the subsampling layer 428 having dimensions that are smaller than the dimensions of the upstream layer, the CNN 420 combines the amount of data associated with the initial input and / or output of the upstream layer, thereby reducing the amount of computation required for the CNN 420 to perform downstream convolutional operations. Additionally or alternatively, by means of the subsampling layer 428 being associated with at least one subsampling function (e.g., being configured to perform at least one subsampling function) (as described below with respect to Figure 4C and Figure 4D ), the CNN 420 combines the amount of data associated with the initial input.

[0079] Based on the perception system 402 providing corresponding inputs and / or outputs associated with the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426 respectively to generate corresponding outputs, the perception system 402 performs convolutional operations. In some examples, based on the perception system 402 providing data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426, the perception system 402 implements the CNN 420. In such examples, based on the perception system 402 receiving data from one or more different systems (e.g., one or more systems of a vehicle identical or similar to the vehicle 102, a remote AV system identical or similar to the remote AV system 114, a queue management system identical or similar to the queue management system 116, and / or a V2I system identical or similar to the V2I system 118, etc.), the perception system 402 provides the data as inputs to the first convolutional layer 422, the second convolutional layer 424, and the convolutional layer 426. The following is a detailed description of Figure 4C including convolutional operations.

[0080] In some embodiments, the perception system 402 provides data associated with an input (referred to as an initial input) to the first convolutional layer 422, and the perception system 402 uses the first convolutional layer 422 to generate data associated with an output. In some embodiments, the perception system 402 provides the output generated by the convolutional layer as an input to a different convolutional layer. For example, the perception system 402 provides the output of the first convolutional layer 422 as an input to the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426. In such an example, the first convolutional layer 422 is referred to as an upstream layer, and the subsampling layer 428, the second convolutional layer 424, and / or the convolutional layer 426 are referred to as downstream layers. Similarly, in some embodiments, the perception system 402 provides the output of the subsampling layer 428 to the second convolutional layer 424 and / or the convolutional layer 426, and in this example, the subsampling layer 428 will be referred to as an upstream layer, and the second convolutional layer 424 and / or the convolutional layer 426 will be referred to as downstream layers.

[0081] In some embodiments, before the perception system 402 provides an input to the CNN 420, the perception system 402 processes data associated with the input provided to the CNN 420. For example, based on the perception system 402 normalizing sensor data (such as image data, LiDAR data, and / or Radar data, etc.), the perception system 402 processes data associated with the input provided to the CNN 420.

[0082] In some embodiments, based on the perception system 402 performing convolution operations associated with each convolutional layer, the CNN 420 generates an output. In some examples, based on the perception system 402 performing convolution operations associated with each convolutional layer and the initial input, the CNN 420 generates an output. In some embodiments, the perception system 402 generates an output and provides the output to the fully connected layer 430. In some examples, the perception system 402 provides the output of the convolutional layer 426 to the fully connected layer 430, where the fully connected layer 430 includes data associated with a plurality of eigenvalues referred to as F1, F2,..., FN. In this example, the output of the convolutional layer 426 includes data associated with a plurality of output eigenvalues representing predictions.

[0083] In some embodiments, the perception system 402 identifies an eigenvalue associated with the highest likelihood of being the correct prediction among multiple predictions, and the perception system 402 identifies a prediction from among the multiple predictions. For example, in a case where the fully connected layer 430 includes eigenvalues F1, F2, ..., FN and F1 is the largest eigenvalue, the perception system 402 identifies the prediction associated with F1 as the correct prediction among the multiple predictions. In some embodiments, the perception system 402 trains the CNN 420 to generate predictions. In some examples, the perception system 402 trains the CNN 420 to generate predictions based on providing training data associated with the prediction to the CNN 420.

[0084] Now refer to Figure 4C and Figure 4D , a diagram illustrating an example operation of the CNN 440 that utilizes the perception system 402. In some embodiments, the CNN 440 (e.g., one or more components of the CNN 440) is the same as or similar to the CNN 420 (e.g., one or more components of the CNN 420) (see Figure 4B ).

[0085] In step 450, the perception system 402 provides data associated with an image as an input to the CNN 440 (step 450). For example, as illustrated, the perception system 402 provides data associated with an image to the CNN 440, where the image is a grayscale image represented as values stored in a two-dimensional (2D) array. In some embodiments, the data associated with the image may include data associated with a color image, which is represented as values stored in a three-dimensional (3D) array. Additionally or alternatively, the data associated with the image may include data associated with an infrared image and / or a Radar image, etc.

[0086] In step 455, the CNN 440 performs a first convolution function. For example, based on the CNN 440 providing the values representing the image as inputs to one or more neurons (not explicitly illustrated) included in the first convolutional layer 442, the CNN 440 performs the first convolution function. In this example, the values representing the image may correspond to the values of a region (sometimes referred to as a receptive field) representing the image. In some embodiments, each neuron is associated with a filter (not explicitly illustrated). The filter (sometimes referred to as a kernel) can be represented as an array of values corresponding in size to the values provided as inputs to the neuron. In one example, the filter can be configured to identify edges (e.g., horizontal lines, vertical lines, and / or straight lines, etc.). In successive convolutional layers, the filters associated with the neurons can be configured to successively identify more complex patterns (e.g., arcs and / or objects, etc.).

[0087] In some embodiments, based on the CNN 440, the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 are multiplied by the values of the filters corresponding to each neuron among the same one or more neurons, and the CNN 440 performs a first convolutional function. For example, the CNN 440 may multiply the values provided as input to each neuron among one or more neurons included in the first convolutional layer 442 by the values of the filters corresponding to each neuron among the one or more neurons to generate a single value or an array of values as output. In some embodiments, the collective output of the neurons of the first convolutional layer 442 is referred to as the convolutional output. In some embodiments, when each neuron has the same filter, the convolutional output is referred to as a feature map.

[0088] In some embodiments, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the neurons of a downstream layer. For clarity, an upstream layer may be a layer that transmits data to a different layer (referred to as a downstream layer). For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of a subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444. In some embodiments, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregated set of all values provided to each neuron of the first subsampling layer 444. In such an example, the CNN 440 determines the final value to be provided to each neuron of the first subsampling layer 444 based on the aggregated set of all values provided to each neuron and the activation function associated with each neuron of the first subsampling layer 444.

[0089] In step 460, the CNN 440 performs a first subsampling function. For example, based on the CNN 440 providing the values output by the first convolutional layer 442 to the corresponding neurons of the first subsampling layer 444, the CNN 440 may perform a first subsampling function. In some embodiments, the CNN 440 performs the first subsampling function based on an aggregation function. In an example, based on the CNN 440 determining the maximum input among the values provided to a given neuron (referred to as the max pooling function), the CNN 440 performs the first subsampling function. In another example, based on the CNN 440 determining the average input among the values provided to a given neuron (referred to as the average pooling function), the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to each neuron of the first subsampling layer 444, the CNN 440 generates an output, which is sometimes referred to as the subsampled convolutional output.

[0090] In step 465, the CNN 440 performs a second convolution function. In some embodiments, the CNN 440 performs the second convolution function in a manner similar to how the CNN 440 performs the first convolution function as described above. In some embodiments, based on the values output by the first subsampling layer 444 being provided as inputs to one or more neurons (not explicitly illustrated) included in the second convolutional layer 446, the CNN 440 performs the second convolution function. In some embodiments, as described above, each neuron of the second convolutional layer 446 is associated with a filter. As described above, the (one or more) filters associated with the second convolutional layer 446 may be configured to identify more complex patterns compared to the filters associated with the first convolutional layer 442.

[0091] In some embodiments, based on the CNN 440 multiplying the values provided as inputs to each of the one or more neurons included in the second convolutional layer 446 by the values of the filters corresponding to each of the one or more neurons, the CNN 440 performs the second convolution function. For example, the CNN 440 may multiply the values provided as inputs to each of the one or more neurons included in the second convolutional layer 446 by the values of the filters corresponding to each of the one or more neurons to generate a single value or an array of values as output.

[0092] In some embodiments, the CNN 440 provides the output of each neuron of the second convolutional layer 446 to the neurons of a downstream layer. For example, the CNN 440 may provide the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the subsampling layer. In an example, the CNN 440 provides the output of each neuron of the first convolutional layer 442 to the corresponding neurons of the second subsampling layer 448. In some embodiments, the CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the downstream layer. For example, the CNN 440 adds a bias value to the aggregate set of all values provided to each neuron of the second subsampling layer 448. In such an example, the CNN 440 determines the final value provided to each neuron of the second subsampling layer 448 based on the aggregate set of all values provided to each neuron and the activation function associated with each neuron of the second subsampling layer 448.

[0093] At step 470, the CNN 440 performs a second subsampling function. For example, based on the values output by the second convolutional layer 446 being provided to the respective neurons of the second subsampling layer 448 by the CNN 440, the CNN 440 may perform the second subsampling function. In some embodiments, based on the CNN 440 using an aggregation function, the CNN 440 performs the second subsampling function. In an example, as described above, based on the CNN 440 determining the maximum input or average input among the values provided to a given neuron, the CNN 440 performs the first subsampling function. In some embodiments, based on the CNN 440 providing values to the respective neurons of the second subsampling layer 448, the CNN 440 generates an output.

[0094] At step 475, the CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449. For example, the CNN 440 provides the output of each neuron of the second subsampling layer 448 to the fully connected layer 449 such that the fully connected layer 449 generates an output. In some embodiments, the fully connected layer 449 is configured to generate an output associated with a prediction (sometimes referred to as classification). The prediction may include an indication that the object(s) included in the image provided as input to the CNN 440 include an object and / or a set of objects, etc. In some embodiments, the perception system 402 performs one or more operations and / or provides data associated with the prediction to different systems described herein.

[0095] Example Camera RADAR Preprocessing Pipeline

[0096] In some embodiments, an AV utilizes a camera RADAR to capture images of the AV's operating environment and any objects or agents approaching the AV. In some embodiments, the camera RADAR is an FMCW Radar that transmits a chirp sequence. The received signal reflected by an object in the AV operating environment is recorded in a 3D tensor cube (in the time domain) for indicating chirp metrics, chirp sampling, and corresponding receiver antenna metrics. This 3D tensor cube is hereinafter referred to as the ADC signal. Existing preprocessing techniques use Figure 5 the preprocessing pipeline 500 shown in to process the ADC signal to estimate the distance, radial velocity, and azimuth angle of the object relative to the camera RADAR.

[0097] Figure 5Illustrate a preprocessing pipeline 500 for an ADC signal 501 according to one or more embodiments. Use a first fast Fourier transform 502 (hereinafter "range-FFT") along a chirp sequence to extract the object distance from the ADC beat signal 501. Then apply a second FFT 503 (hereinafter "Doppler-FFT") along the chirp sampling axis to estimate the phase difference and derive the radial velocity of the reflecting surface. A third FFT 504 (hereinafter "angle-FFT") processes the ADC signal through an antenna pair to estimate the azimuth angle to the object. More detailed examples of the above processing can be found in the following literature: Kim, Bong-seok & Kim, Sangdong & Jin, Youngseok & Lee, Jonghun. (2020), Low-Complexity Joint Range and Doppler FMCW Radar Algorithm Based on Number of Targets, Sensors. 20.10.3390 / s20010051.

[0098] The cascaded sequence of FFTs 502, 503, 504 produces a range-angle-Doppler (RAD) tensor 505 that includes complex numbers, where each axis of the tensor includes discretized values of the corresponding physical measurement. The RAD tensor 505 is input into an AI engine 506, which can, for example, predict 2D or 3D bounding boxes and labels of the object and its location. However, the preprocessing pipeline 500 has several limitations, including inaccurate azimuth angle estimation and additional latency caused by the three cascaded FFTs 502, 503, 504. Figure 6 Illustrate an improved preprocessing pipeline for an ADC signal below.

[0099] High-precision, low-latency processing pipeline

[0100] Figure 6 is a block diagram of the hardware architecture of a high-precision, low-latency preprocessing pipeline 600 according to one or more embodiments. The number and arrangement of the components illustrated in Figure 6 are provided as an example. In some embodiments, the preprocessing pipeline 600 may include additional components, fewer components, different components, or components with a different arrangement compared to those illustrated in Figure 6 . Additionally or alternatively, a set of components (e.g., one or more components) of the preprocessing pipeline 600 may perform one or more functions described as being performed by another component or another set of components of the preprocessing pipeline 600.

[0101] Refer to Figure 6In the example embodiment, the preprocessing pipeline 600 includes a camera RADAR interface 601, a frame buffer controller 602, an ADC buffer 603, a 2D FFT calculation 604, a data structure calculation 608, and an AI engine calculation 609. The 2D FFT calculation 604 further includes a range-FFT 605 for estimating the distance of an object relative to the RADAR, and a Doppler-FFT 606 for estimating the radial velocity of the object relative to the RADAR based on the Doppler spectrum output by the 2D FFT calculation 604.

[0102] In a separate processing path, the output (spectrum) of the Doppler-FFT 606 is fed to a MUSIC calculation 607, which uses the MUSIC algorithm to estimate the azimuth angle of the object relative to the RADAR. The MUSIC algorithm provides a more accurate azimuth angle estimate than the angle-FFT 504 in the preprocessing pipeline 500. A more detailed description of the MUSIC algorithm applied to Radar ADC signals can be found in: Manokhin, Gleb & Erdyneev, Zhargal & Geltser, Andrey & Monastyrev, Evgeny. (2015), MUSIC-based algorithm for range-azimuth FMCW radar data processing without estimating number of targets, 1-4. 10.1109 / MMS.2015.7375471.

[0103] In operation, the ADC signal is output by the camera RADAR interface 601, which can be, for example, a serial interface. The ADC signal (e.g., a 3D tensor cube) is stored in the ADC buffer 603 by the frame buffer controller 602. The frame buffer controller 602 retrieves the ADC signal so that the ADC signal can be processed by the 2D FFT calculation 604, which calculates the estimated distance and radial velocity of at least one object relative to the camera RADAR. In a separate processing path, the output (2D Doppler spectrum) of the Doppler-FFT 606 is fed to the MUSIC calculation 607 to estimate the azimuth angle using the MUSIC algorithm. The high accuracy of the MUSIC algorithm helps to avoid problems of inaccurate azimuth angles that may result in overlapping bounding boxes output, for example, by the AI engine calculation 609. Additionally, only two 2D FFTs are performed in the preprocessing pipeline 500 instead of three 2D FFTs, which results in reduced latency.

[0104] The azimuth, radial velocity, and distance estimates are then input into a data structure computation 608 that modifies the data to make it suitable for consumption by a particular architecture (e.g., a particular type of processor) (e.g., scalar architecture, vector architecture, matrix architecture, spatial architecture, etc.) used by the AI engine computation 609. In some embodiments, as previously referenced Figure 5 and described, the data structure computation 608 generates a RAD tensor. In other embodiments, the data structure computation 608 generates scalar values, vectors, matrices, or any other suitable data structure according to the input requirements of the AI engine 609.

[0105] In some embodiments, the system 600 is implemented in a distributed processing hardware architecture 800 (see Figure 8 ), where various computations in the system 600 can be implemented on different hardware processors. For example, the 2D FFT computation 604 and the MUSIC computation 607 can be implemented on one hardware processor, and the AI engine computation 609 can be implemented on a different hardware processor with a computational fabric (such as intermediate results of arithmetic computations, etc.) that allows for shared memory. For example, the AI engine computation 609 can be implemented in a distributed hardware architecture using two or more AI accelerator chips, GPUs, neural processing units (NPUs), vector processing units (VPUs), or tensor processing units (TPUs). In some embodiments, the various computations performed by the system 600 can be performed in parallel on different processors, further reducing the overall latency of the object detection / classification / localization awareness tasks.

[0106] Figure 7 is a flowchart of a high-precision, low-latency processing 700 of an ADC beat signal according to one or more embodiments. The processing 700 can be implemented by a single processor or in a distributed architecture (such as the architecture shown in Figure 8 etc.).

[0107] In some embodiments, the processing 700 includes: receiving a time-domain representation of a Radar image of a scene having at least one object (701); determining a distance to the at least one object based on a Doppler spectrum computed from the time-domain representation (702); determining a radial velocity of the at least one object based on a distance spectrum computed from the time-domain representation (703); determining an azimuth of the at least one object based on the Doppler spectrum (704); generating a data structure containing the distance, the radial velocity, and the azimuth (705); performing at least one machine learning process on the contents of the data structure (706); and generating at least one control signal for controlling a vehicle based at least in part on the results of the machine learning process (707). The steps of the processing 700 were previously referenced Figure 6It is described above and can be implemented by at least one processor.

[0108] Figure 8 According to one or more embodiments, Figure 6 The block diagram of the chip layout of the computing unit of the pre-processing pipeline 600 of FIG. The computing unit 800 can be implemented in, for example, AV computing (e.g., AV computing 202f). The computing unit 800 includes a sensor multiplexer (MUX) 801, a main computing cluster 802-1 to 802-5, a failover computing cluster 802-6, and an Ethernet switch 802. The Ethernet switch 802 includes a plurality of Ethernet transceivers for sending commands 815 to a vehicle 803, wherein the commands 815 are transmitted by, for example, Figure 2 One or more of a drive-by-wire (DBW) system 202h, a safety controller 202g, a braking system 208, a powertrain control system 204, and / or a steering control system 206 are shown receiving.

[0109] The computing unit 800 can be used to implement Figure 6 , where different SoCs can be assigned to different pre-processing tasks, such as implementing 2D FFT on a first SoC, implementing a MUSIC algorithm on a second SoC, and implementing an AI engine on a third SoC, etc. Each of these SoCs can share memory / data through a computing fabric or a network on chip (NoC).

[0110] refer to Figure 8, the first main computing cluster 802-1 includes a System-on-Chip (SoC) 803-1, volatile memories 805-1, 805-2, a Power Management Integrated Circuit (PMIC) 804-1, and a flash boot 811-1. The second main computing cluster 802-2 includes an SoC 803-2, volatile memories 806-1, 806-2 (e.g., DRAM), a PMIC 804-2, and a flash operating system (OS) 812-2. The third main computing cluster 802-3 includes an SoC 803-3, volatile memories 807-1, 807-2, a PMIC 804-3, and a flash OS memory 812-1. The fourth main computing cluster 802-4 includes an SoC 803-5, volatile memories 808-1, 808-2, a PMIC 804-5, and a flash boot memory 811-2. The fifth main computing cluster 802-4 includes an SoC 803-4, volatile memories 809-1, 809-2, a PMIC 804-4, and a flash boot memory 811-3. The failover computing cluster 802-6 includes an SoC 803-6, volatile memories 810-1, 810-2, a PMIC 804-6, and a flash OS memory 812-3.

[0111] Each of the SoCs 803-1 to 803-6 can be a multi-processor System-on-Chip (MPSoC). The SoCs 803-1 to 803-6 can share memory through a cache coherence fabric (such as, for example, a cache coherence interconnect (CCIX) for accelerators, etc.).

[0112] In an embodiment, the PMICs 804-1 to 804-6 monitor relevant signals on a bus (e.g., a PCIe bus) and communicate with corresponding memory controllers (e.g., memory controllers in DRAM chips) to notify the memory controllers of a power mode change, such as a change from a normal mode to a low power mode or a change from a low power mode to a normal mode, etc. In an embodiment, the PMICs 804-1 to 804-6 also receive communication signals from their respective memory controllers that are monitoring the bus and operate to prepare the memory for a lower power mode. When the memory chip is ready to enter the low power mode, the memory controller communicates with its corresponding slave PMIC to indicate that the slave PMIC initiates the lower power mode.

[0113] In an embodiment, the sensor MUX 801 receives and multiplexes sensor data (e.g., video data, LiDAR point cloud, Radar data) from a sensor bus through a sensor interface 813 (which is a Low-Voltage Differential Signaling (LVDS) interface in some embodiments). For example, RADAR can be in a camera RADAR, as referenced Figure 6As described above, the camera RADAR generates ADS signals. In an embodiment, the sensor MUX 801 steers copies of video and / or camera RADAR data channels (e.g., Mobile Industry Processor Interface Camera Serial Interface (CSI) channels) that are sent to the failover clusters 802-6. The failover clusters 802-6 use the video or Radar data to provide backup to the primary computing clusters during failover 814 to operate the AV in the event of, for example, one or more of the primary computing clusters 802-1 failing. In some such cases, the failover clusters 802-6 may issue commands 816 to the vehicle 803.

[0114] The computing unit 800 is an example of a high-performance computing unit (such as for AV computing etc.) used in an autonomous robotic system, and other embodiments may include more or fewer clusters, and each cluster may have more or fewer SoCs, volatile memory chips, non-volatile memory chips, NPUs, GPUs, TPUs, VPUs, AI accelerators, and Ethernet switch / transceivers.

[0115] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to numerous specific details, which may vary depending on the implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than in a limiting sense. The sole and exclusive indication of the scope of the invention, and what the applicant desires to be the scope of the invention, is the literal and equivalent scope of the claims as issued from this application in the specific form of the issued claims, including any subsequent amendments. Any definition of terms expressly set forth herein for inclusion in such claims shall be construed in the sense such terms are used in the claims. Additionally, when the term "further comprises" is used in the foregoing specification or the appended claims, the text following such phrase may be an additional step or entity, or a sub-step / sub-entity of a previously described step or entity.

Claims

1. A method, comprising: Receiving, by at least one processor, a time-domain representation of a Radar image of a scene, the representation including at least one object; Determining, by the at least one processor, a distance to the at least one object based on a distance spectrum calculated from the time-domain representation; Determining, by the at least one processor, a radial velocity of the at least one object based on a Doppler spectrum calculated from the time-domain representation; Determining, by the at least one processor, an azimuth angle of the at least one object based on the Doppler spectrum; Generating, by the at least one processor, a data structure containing the distance, the radial velocity, and the azimuth angle; Performing, by the at least one processor, at least one machine learning process on the content of the data structure; And Generating, by the at least one processor, at least one control signal for controlling a vehicle, at least in part based on a result of performing the machine learning process.

2. The method according to claim 1, wherein Using a multiple signal classification algorithm, namely the MUSIC algorithm, to generate the azimuth angle.

3. The method according to claim 1, wherein The data structure is a three-dimensional range-angle-Doppler tensor, namely a 3DRAD tensor.

4. The method according to claim 1, wherein Generating the data structure based on the type of the at least one processor.

5. A system, comprising: A bus; A frame buffer; A frame buffer controller; And At least one processor, wherein: The bus transmits a time-domain representation of a Radar image, the representation including at least one object; The frame buffer controller stores an output in the frame buffer and retrieves the output from the buffer for further processing by the at least one processor; Wherein, the at least one processor: Calculates a distance of the at least one object relative to a camera Radar based on a distance spectrum calculated from the time-domain representation; Calculates a radial velocity of the at least one object based on a Doppler spectrum calculated from the time-domain representation; Calculates an azimuth angle of the at least one object relative to the camera Radar based on the Doppler spectrum; Generates a data structure storing the distance, the radial velocity, and the azimuth angle; Performs at least one machine learning process on the content of the data structure; and Generates at least one control signal for controlling a vehicle, at least in part based on a result of performing the machine learning process.

6. The system according to claim 5, wherein, Using a multiple signal classification algorithm, namely the MUSIC algorithm, to generate the azimuth angle.

7. The system according to claim 6, wherein, The data structure is a three-dimensional range-angle-Doppler tensor, namely a 3DRAD tensor.

8. The system according to claim 5, wherein, The system is a distributed hardware architecture, the at least one processor includes two or more processors, and the two or more processors include at least one of a neural processing unit (NPU), a graphics processing unit (GPU), a tensor processing unit (TPU), and an accelerator chip that share data through a computing fabric.

9. The system according to claim 5, wherein Generating the data structure based on the type of the at least one processor.

10. A system, comprising: A Radar; A Radar interface configured to receive a three-dimensional tensor cube, namely a 3D tensor cube, from the Radar; A frame buffer controller configured to store the 3D tensor cube and retrieve the 3D tensor cube from at least one buffer; A two-dimensional fast Fourier transform calculation, i.e., 2D FFT calculation, configured to receive the 3D tensor from the frame buffer controller and estimate the distance of the object relative to the Radar and the radial velocity of the object relative to the Radar based on the Doppler spectrum; An azimuth calculation configured to estimate the azimuth based on the Doppler spectrum; And An artificial intelligence calculation, i.e., AI calculation, configured to detect and classify the object and its location based on the estimated distance, the estimated radial velocity, and the estimated azimuth.

11. The system according to claim 10, wherein, The Radar is a frequency-modulated continuous-wave Radar, i.e., FMCW RADAR.

12. The system according to claim 10, wherein The azimuth calculation uses a multiple signal classification calculation, i.e., MUSIC calculation, to calculate the azimuth.

13. The system according to claim 10, further comprising a data structure calculation configured to generate a three-dimensional range-angle-Doppler tensor, i.e., 3D RAD tensor.

14. The system according to claim 10, wherein, The 2D FFT calculation and the MUSIC calculation are implemented on a first system-on-chip, i.e., the first SoC, and the AI engine is implemented on a second SoC, wherein each of the first SoC and the second SoC shares data through a compute fabric or a network-on-chip, i.e., NoC.

15. The system according to claim 10, wherein, The 2D FFT calculation is implemented on a first SoC, the MUSIC calculation is implemented on a second system-on-chip, i.e., the second SoC, and the AI engine is implemented on a third SoC, wherein each of the first SoC, the second SoC, and the third SoC shares data through a compute fabric or a network-on-chip, i.e., NoC.