Target detection using radio detection and ranging sensors
By FFT processing of Radar sensor data to generate multi-dimensional tensors and inputting them into machine learning models, the information loss and delay problems in Radar sensor signal processing are solved, and efficient object detection is achieved.
Patent Information
- Application Number
- CN202380085224.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-02
- Filing Date
- 2023-10-02
- Publication Date
- 2025-07-22
AI Technical Summary
The signal processing of existing Radar sensors overprocesses the original data, resulting in the loss of detailed information, and traditional methods require buffering and waiting for the entire data frame to be ready for FFT processing, increasing processing delay.
The data of the Radar sensor is processed by distance FFT, Doppler FFT and azimuth FFT to generate 1D distance thermogram tensor, 2D RD thermogram tensor, 2D RA thermogram tensor or 3D RAD matrix tensor, and input it into the machine learning model for object detection to reduce internal processing delay.
Through rich representations of early Radar output data, machine learning models can efficiently detect objects, reduce processing delays and retain detailed information, avoiding over-processing problems of traditional methods.
Smart Images

Figure CN120359432A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 416,459, filed on October 14, 2022, and U.S. Patent Application No. 18 / 105,183, filed on February 2, 2023, the entire contents of which are incorporated herein by reference. Background of the Invention
[0003] A Radio Detection and Ranging (Radar) sensor transmits an electromagnetic wave signal, which is reflected by an object in the environment. The Radar sensor captures the reflected signal and processes the reflected signal to determine various properties of the environment. Brief Description of the Drawings
[0004] Figure 1 is an example environment of a vehicle that can implement one or more components of an autonomous system;
[0005] Figure 2 is a diagram of one or more systems of a vehicle including an autonomous system;
[0006] Figure 3 is Figure 1 and Figure 2 a diagram of one or more devices and / or components of one or more systems of;
[0007] Figure 4 is a diagram of certain components of an autonomous system;
[0008] Figure 5 is a diagram of an implementation of a process for detecting an object using a sensor suite;
[0009] Figures 6(a) to 6(c) is different example pipelines of a Radar sensor;
[0010] Figures 7(a) to 7(d) is different example representations of Radar data;
[0011] Figure 8 is a diagram illustrating the generation of an example Range - Azimuth - Doppler (RAD) matrix tensor representing Radar data;
[0012] Figures 9(a) to 9(c) is different example pipelines of sensor data fusion; and
[0013] Figure 10 is an example flowchart of a process for generating a representation of Radar data. Detailed Description
[0014] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the embodiments described herein may be practiced without these specific details. In some instances, well-known structures and devices are illustrated in block diagram form to avoid unnecessarily obscuring aspects of the present disclosure.
[0015] In the drawings, for ease of description, the specific arrangements or orderings of illustrative elements (such as those representing systems, devices, modules, instruction blocks, and / or data elements, etc.) are illustrated. However, those skilled in the art will understand that unless explicitly described, the specific orderings or arrangements of the illustrative elements in the drawings are not intended to imply a required order or sequence of processing, or a separation of processing. Additionally, unless explicitly described, the inclusion of illustrative elements in the drawings is not intended to imply that such elements are required in all embodiments, nor that the features represented by such elements cannot be included in some embodiments or combined with other elements in some embodiments.
[0016] Furthermore, in the drawings, connecting elements (such as solid lines, dashed lines, or arrows, etc.) are used to illustrate connections, relationships, or associations between or among two or more other illustrative elements. The absence of any such connecting element is not intended to imply that no connection, relationship, or association can exist. In other words, some connections, relationships, or associations between elements are not illustrated in the drawings so as not to obscure the present disclosure. Additionally, for ease of illustration, a single connecting element may be used to represent multiple connections, relationships, or associations between elements. For example, if a connecting element represents the communication of a signal, data, or instruction (e.g., “software instruction”), those skilled in the art will understand that such an element may represent one or more than one signal path (e.g., a bus) that may be required to affect the communication.
[0017] Although terms such as “first,” “second,” and / or “third,” etc. are used to describe various elements, these elements should not be limited by these terms. The terms “first,” “second,” and / or “third” are only used to distinguish one element from another. For example, without departing from the scope of the described embodiments, a first contact may be referred to as a second contact, and similarly, a second contact may be referred to as a first contact. Both the first contact and the second contact are contacts, but they are not the same contact.
[0018] The terms used in the description of the various embodiments described herein are included only for the purpose of describing particular embodiments and are not intended to be limiting. As used in the description of the various embodiments and the appended claims, the singular forms "a", "an", and "the" are also intended to include the plural forms and may be used interchangeably with "one or more than one" or "at least one", unless the context clearly dictates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that when the terms "comprises", "comprising", "includes", and / or "including" are used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0019] As used herein, the terms "communicate" and "communicating" refer to at least one of receiving, receipt, transmission, conveyance, and / or provision of information (or information represented by, for example, data, signals, messages, instructions, and / or commands, etc.). For a unit (e.g., a device, a system, a component of a device or system, and / or a combination thereof, etc.) that is to communicate with another unit, this means that the unit can directly or indirectly receive information from the other unit and / or send (e.g., transmit) information to the other unit. This can refer to a direct or indirect connection that is inherently wired and / or wireless. Additionally, two units can communicate with each other even if the information transmitted therebetween is modified, processed, relayed, and / or routed. For example, even if the first unit receives information passively and does not actively transmit information to the second unit, the first unit can communicate with the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transmits the processed information to the second unit, the first unit can communicate with the second unit. In some embodiments, a message can refer to a network packet (e.g., a data packet, etc.) that includes data.
[0020] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "at the time of", "in response to determining that", and / or "in response to detecting", etc. Similarly, depending on the context, the phrase "if it has been determined" or "if [the stated condition or event] is detected" is optionally interpreted to mean "at the time of determining...", "in response to determining that" or "at the time of detecting [the stated condition or event]" and / or "in response to detecting [the stated condition or event]", etc. Further, as used herein, terms such as "have", "having", or "possessing" are intended to be open-ended terms. Additionally, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on".
[0021] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the various described embodiments. However, it will be apparent to those of ordinary skill in the art that the various described embodiments may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.
[0022] General Overview
[0023] The present disclosure provides a plurality of configurable Radar data representations from a Radar sensor (e.g., one-dimensional (1D) range heatmap tensor, two-dimensional (2D) range-Doppler (RD) heatmap tensor, 2D range-azimuth (RA) heatmap tensor, three-dimensional (3D) range-azimuth-Doppler (RAD) matrix tensor). At least one of the Radar data representations is input into a machine learning model to detect an object. In some embodiments, the Radar data representation is fused with a camera image, and the fused data is input into the machine learning model to detect an object.
[0024] In some embodiments, distance Fast Fourier Transform (FFT), Doppler FFT, and azimuth FFT are performed on the raw data output by the Radar sensor. The raw data is output by the analog-to-digital converter (ADC) of the Radar sensor. A 1D distance heatmap tensor for representing the Radar output is generated based on the distance FFT data. A 2D RD heatmap tensor for representing the Radar output is generated based on a combination of the distance FFT data and the Doppler FFT data. A 2D RA heatmap tensor for representing the Radar output is generated based on a combination of the distance FFT data and the azimuth FFT data. A 3D RAD matrix tensor for representing the Radar output is generated based on a combination of the distance FFT data, the azimuth FFT data, and the Doppler FFT data. The 1D distance heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, or the 3D RAD matrix tensor is provided to a machine learning model for obstacle detection. In some embodiments, the 1D distance heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, or the 3D RAD matrix tensor is fused with a camera image, and the fused data is provided to a machine learning model for obstacle detection. In some embodiments, the size of the 1D distance heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, or the 3D RAD matrix tensor is configured based on Radar specifications (such as sensing distance, distance resolution, velocity resolution, angle resolution, bandwidth of chirp, period of chirp, number of samples per chirp or sampling rate, number of ADC channels, etc.).
[0025] With the implementation of the systems, methods, and computer program products described herein, some advantages of these techniques include providing early Radar output data (Radar output with much less signal processing represented by the 1D distance heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, or the 3D RAD matrix tensor) to a machine learning model instead of sparse point clouds. The machine learning model detects objects based on the rich information in the early Radar data. In contrast, traditional Radar sensors use signal filters and clustering beamforming for signal processing of ADC raw data. This signal processing "over-processes" the ADC raw data, resulting in the loss of detailed information carried by the ADC raw data. Some advantages of these techniques also include configuring Radar output data based on Radar specifications (such as sensing distance, distance resolution, velocity resolution, angle resolution, bandwidth of chirp, period of chirp, number of samples per chirp or sampling rate, number of ADC channels, etc.).
[0026] Additionally, a 1D range heatmap tensor is generated based only on range FFT data. Range FFT can be performed in real time on the input raw ADC data (e.g., range FFT is performed for each chirp). Different from the 2D RD heatmap tensor, 2D RA heatmap tensor, or 3D RAD matrix tensor, the 1D range heatmap tensor does not need to buffer and wait until the entire data frame (e.g., one data frame includes 256 chirps) is ready for FFT processing. Thus, the 1D range heatmap tensor can reduce internal Radar sources (e.g., only perform range FFT inside the Radar sensor), and reduce the processing delay.
[0027] Now refer to Figure 1 , an exemplary environment 100 is illustrated, in which vehicles including autonomous systems and vehicles not including autonomous systems operate. As illustrated, the environment 100 includes vehicles 102a - 102n, objects 104a - 104n, routes 106a - 106n, regions 108, vehicle - to - infrastructure (V2I) devices 110, networks 112, remote autonomous vehicle (AV) systems 114, queue management systems 116, and V2I systems 118. The vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, networks 112, autonomous vehicle (AV) systems 114, queue management systems 116, and V2I systems 118 are interconnected via wired connections, wireless connections, or a combination of wired and wireless connections (e.g., establish connections for communication, etc.). In some embodiments, the objects 104a - 104n are interconnected with at least one of the vehicles 102a - 102n, vehicle - to - infrastructure (V2I) devices 110, networks 112, autonomous vehicle (AV) systems 114, queue management systems 116, and V2I systems 118 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0028] The vehicles 102a - 102n (individually referred to as vehicle 102 and collectively referred to as vehicles 102) include at least one device configured to transport goods and / or people. In some embodiments, the vehicle 102 is configured to communicate with the V2I device 110, remote AV system 114, queue management system 116, and / or V2I system 118 via the network 112. In some embodiments, the vehicle 102 includes cars, buses, trucks, and / or trains, etc. In some embodiments, the vehicle 102 is related to the vehicle 200 described herein (see Figure 2) are the same or similar. In some embodiments, the vehicle 200 in the set of vehicles 200 is associated with an autonomous queue manager. In some embodiments, as described herein, the vehicle 102 travels along corresponding routes 106a-106n (individually referred to as route 106 and collectively referred to as routes 106). In some embodiments, one or more vehicles 102 include an autonomous system (e.g., an autonomous system that is the same or similar to the autonomous system 202).
[0029] The objects 104a-104n (individually referred to as object 104 and collectively referred to as objects 104) include, for example, at least one vehicle, at least one pedestrian, at least one cyclist, and / or at least one structure (e.g., a building, a sign, a fire hydrant, etc.). Each object 104 (e.g., located at a fixed location and over a period of time) is stationary or (e.g., having a speed and associated with at least one trajectory) moving. In some embodiments, the object 104 is associated with a corresponding location in the region 108.
[0030] The routes 106a-106n (individually referred to as route 106 and collectively referred to as routes 106) are each associated with (e.g., define) a sequence of actions (also referred to as a trajectory) that connect the states along which the AV can navigate. Each route 106 begins at an initial state (e.g., a state corresponding to a first spatio-temporal location and / or speed, etc.) and ends at a final target state (e.g., a state corresponding to a second spatio-temporal location different from the first spatio-temporal location) or a target zone (e.g., a subspace of acceptable states (e.g., a termination state)). In some embodiments, the first state includes a location where one or more individuals will board the AV, and the second state or zone includes one or more locations where one or more individuals boarding the AV will disembark. In some embodiments, the route 106 includes multiple acceptable sequences of states (e.g., multiple sequences of spatio-temporal locations), which are associated with (e.g., define) multiple trajectories. In an example, the route 106 includes only high-level actions or imprecise state locations, such as a series of connected roads indicating a direction change at a roadway intersection. Additionally or alternatively, the route 106 can include more precise actions or states, such as, for example, a specific target lane or precise location within a lane area and the target rate at these locations. In an example, the route 106 includes multiple precise state sequences along at least one high-level action with a finite look-ahead horizon to reach an intermediate target, where the combination of consecutive iterations of the finite horizon state sequences cumulatively corresponds to multiple trajectories that together form a high-level route terminating at the final target state or zone.
[0031] Region 108 includes a physical region (e.g., a geographical region) that the vehicle 102 can navigate. In an example, region 108 includes at least one state (e.g., a country, a province, an individual state among a plurality of states included in a country, etc.), at least a portion of a state, at least one city, at least a portion of a city, etc. In some embodiments, region 108 includes at least one named arterial road (referred to herein as a "road"), such as a highway, an interstate highway, a parkway, an urban street, etc. Additionally or alternatively, in some examples, region 108 includes at least one unnamed road, such as a lane, a section of a parking lot, a section of a vacant and / or undeveloped area, a dirt road, etc. In some embodiments, a road includes at least one lane (e.g., a portion of the road that the vehicle 102 can traverse). In an example, a road includes at least one lane associated with (e.g., identified based on) at least one lane marking line.
[0032] A vehicle-to-infrastructure (V2I) device 110 (sometimes referred to as a vehicle-to-infrastructure or vehicle-to-everything (V2X) device) includes at least one device configured to communicate with the vehicle 102 and / or the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, the queue management system 116, and / or the V2I system 118 via the network 112. In some embodiments, the V2I device 110 includes a radio frequency identification (RFID) device, a sign, a camera (e.g., a two-dimensional (2D) and / or three-dimensional (3D) camera), a lane marking, a streetlight, a parking meter, etc. In some embodiments, the V2I device 110 is configured to communicate directly with the vehicle 102. Additionally or alternatively, in some embodiments, the V2I device 110 is configured to communicate with the vehicle 102, the remote AV system 114, and / or the queue management system 116 via the V2I system 118. In some embodiments, the V2I device 110 is configured to communicate with the V2I system 118 via the network 112.
[0033] The network 112 includes one or more wired and / or wireless networks. In an example, the network 112 includes a cellular network (e.g., a Long Term Evolution (LTE) network, a third-generation (3G) network, a fourth-generation (4G) network, a fifth-generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, a cloud computing network, etc., and / or a combination of some or all of these networks.
[0034] The remote AV system 114 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the network 112, the queue management system 116, and / or the V2I system 118 via the network 112. In an example, the remote AV system 114 includes a server, a server group, and / or other similar devices. In some embodiments, the remote AV system 114 is co-located with the queue management system 116. In some embodiments, the remote AV system 114 participates in the installation of some or all of the components of the vehicle (including autonomous systems, autonomous vehicle computing, and / or software implemented by autonomous vehicle computing, etc.). In some embodiments, the remote AV system 114 maintains (e.g., updates and / or replaces) these components and / or software during the life of the vehicle.
[0035] The queue management system 116 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the V2I system 118. In an example, the queue management system 116 includes a server, a server group, and / or other similar devices. In some embodiments, the queue management system 116 is associated with a ridesharing company (e.g., an organization for controlling the operation of multiple vehicles (e.g., vehicles including autonomous systems and / or vehicles not including autonomous systems), etc.).
[0036] In some embodiments, the V2I system 118 includes at least one device configured to communicate with the vehicle 102, the V2I device 110, the remote AV system 114, and / or the queue management system 116 via the network 112. In some examples, the V2I system 118 is configured to communicate with the V2I device 110 via a connection different from the network 112. In some embodiments, the V2I system 118 includes a server, a server group, and / or other similar devices. In some embodiments, the V2I system 118 is associated with a municipal authority or a private institution (e.g., a private institution for maintaining the V2I device 110, etc.).
[0037] Provide Figure 1 The number and arrangement of the illustrated elements are provided as examples. Compared with Figure 1 the illustrated elements, there may be additional elements, fewer elements, different elements, and / or elements with different arrangements. Additionally or alternatively, at least one element of the environment 100 may perform one or more functions described as being performed by Figure 1 at least one different element. Additionally or alternatively, at least one set of elements of the environment 100 may perform one or more functions described as being performed by at least one different set of elements of the environment 100.
[0038] Now refer to Figure 2 , vehicle 200 (which may be the same as or similar to Figure 1 vehicle 102) includes an autonomous system 202, a powertrain control system 204, a steering control system 206, and a braking system 208, or is associated with the autonomous system 202, the powertrain control system 204, the steering control system 206, and the braking system 208. In some embodiments, vehicle 200 is the same as or similar to vehicle 102 (see Figure 1 ). In some embodiments, the autonomous system 202 is configured to endow vehicle 200 with autonomous driving capabilities (e.g., implement at least one of the following driving functions, features, and / or devices that are automatic or based on maneuvers, where the at least one driving function, feature, and / or device that is automatic or based on maneuvers enables vehicle 200 to operate partially or fully without human intervention, including but not limited to fully autonomous vehicles (e.g., vehicles that abandon reliance on human intervention, such as 5 - level ADS - operated vehicles), highly autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in certain situations, such as 4 - level ADS - operated vehicles), and / or conditionally autonomous vehicles (e.g., vehicles that abandon reliance on human intervention in limited situations, such as 3 - level ADS - operated vehicles), etc.). In one embodiment, the autonomous system 202 includes the operational or tactical functionality required to operate vehicle 200 in road traffic and continuously perform part or all of the dynamic driving task (DDT). In another embodiment, the autonomous system 202 includes an advanced driver assistance system (ADAS) that includes driver support features. The autonomous system 202 supports various levels of driving automation ranging from no driving automation (e.g., level 0) to full driving automation (e.g., level 5). For a detailed description of fully autonomous vehicles and highly autonomous vehicles, reference can be made to SAE International standard J3016: Taxonomy and Definitions for Terms Related to On - Road Motor Vehicle Automated Driving Systems, the entire content of which is incorporated by reference. In some embodiments, vehicle 200 is associated with an autonomous queue manager and / or a ridesharing company.
[0039] The autonomous system 202 includes a sensor suite that includes one or more devices such as a camera 202a, a LiDAR sensor 202b, a Radar sensor 202c, and a microphone 202d. In some embodiments, the autonomous system 202 may include more or fewer devices and / or different devices (e.g., ultrasonic sensors, inertial sensors, GPS receivers (discussed below), and / or odometer sensors for generating data associated with an indication of the distance the vehicle 200 has traveled, etc.). In some embodiments, the autonomous system 202 uses one or more of the devices included in the autonomous system 202 to generate data associated with the environment 100 described herein. The data generated by one or more of the devices of the autonomous system 202 can be used by one or more of the systems described herein to observe the environment (e.g., environment 100) in which the vehicle 200 is located. In some embodiments, the autonomous system 202 includes a communication device 202e, an autonomous vehicle computing 202f, a drive-by-wire (DBW) system 202h, and a safety controller 202g.
[0040] The camera 202a includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus 302 that is the same as or similar to Figure 3 the bus). The camera 202a includes at least one camera (e.g., a digital camera using an optical sensor such as a charge-coupled device (CCD), a thermal camera, an infrared (IR) camera, and / or an event camera, etc.) for capturing images including physical objects (e.g., cars, buses, curbs, and / or people, etc.). In some embodiments, the camera 202a generates camera data as an output. In some examples, the camera 202a generates camera data that includes image data associated with the image. In such an example, the image data may specify at least one parameter corresponding to the image (e.g., image characteristics such as exposure, brightness, etc., and / or an image timestamp, etc.). In such an example, the image may be in a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a includes a plurality of independent cameras configured (e.g., positioned) on the vehicle for capturing images for the purpose of stereovision (stereo vision). In some examples, the camera 202a includes generating image data and transmitting the image data to the autonomous vehicle computing 202f and / or a queue management system (e.g., the same as Figure 1a plurality of cameras of the queue management system 116 or a similar queue management system). In such an example, the autonomous vehicle computing 202f determines the depth to one or more objects in the fields of view of at least two of the plurality of cameras based on image data from at least two cameras. In some embodiments, the camera 202a is configured to capture images of objects within a distance (e.g., up to 100 meters and / or up to 1 kilometer, etc.) relative to the camera 202a. Thus, the camera 202a includes features such as sensors and lenses optimized for sensing objects at one or more distances relative to the camera 202a.
[0041] In an embodiment, the camera 202a includes at least one camera configured to capture one or more images associated with one or more traffic lights, street signs, and / or other physical objects that provide visual navigation information. In some embodiments, the camera 202a generates traffic light data associated with one or more images. In some examples, the camera 202a generates TLD (Traffic Light Detection) data associated with one or more images including a format (e.g., RAW, JPEG, and / or PNG, etc.). In some embodiments, the camera 202a that generates TLD data is different from other systems incorporating cameras described herein in that the camera 202a may include one or more cameras having a wide field of view (e.g., a wide-angle lens, a fish-eye lens, and / or a lens having a viewing angle of about 120 degrees or greater, etc.) to generate images related to as many physical objects as possible.
[0042] The Light Detection and Ranging (LiDAR) sensor 202b includes being configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., with Figure 3At least one device that communicates via a bus (e.g., a bus identical or similar to bus 302). The LiDAR sensor 202b includes a system configured to emit light from a light emitter (e.g., a laser emitter). The light emitted by the LiDAR sensor 202b includes light outside the visible spectrum (e.g., infrared light, etc.). In some embodiments, during operation, the light emitted by the LiDAR sensor 202b encounters a physical object (e.g., a vehicle) and is reflected back to the LiDAR sensor 202b. In some embodiments, the light emitted by the LiDAR sensor 202b does not penetrate the physical object it encounters. The LiDAR sensor 202b also includes at least one light detector that detects the light after the light emitted from the light emitter encounters a physical object. In some embodiments, at least one data processing system associated with the LiDAR sensor 202b generates an image (e.g., a point cloud and / or a combined point cloud, etc.) representing the objects included in the field of view of the LiDAR sensor 202b. In some examples, at least one data processing system associated with the LiDAR sensor 202b generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In such examples, the image is used to determine the boundary of the physical object in the field of view of the LiDAR sensor 202b.
[0043] A radio detection and ranging (Radar) sensor 202c includes at least one device configured to communicate with a communication device 202e, an autonomous vehicle computer 202f, and / or a safety controller 202g via a bus (e.g., a bus identical or similar to Figure 3 bus 302). The Radar sensor 202c includes a system configured to emit (pulsed or continuous) radio waves. The radio waves emitted by the Radar sensor 202c include radio waves within a specific spectrum. In some embodiments, during operation, the radio waves emitted by the Radar sensor 202c encounter a physical object and are reflected back to the Radar sensor 202c. In some embodiments, the radio waves emitted by the Radar sensor 202c are not reflected by some objects. In some embodiments, at least one data processing system associated with the Radar sensor 202c generates a signal representing the objects included in the field of view of the Radar sensor 202c. For example, at least one data processing system associated with the Radar sensor 202c generates an image representing the boundary of a physical object and / or the surface of a physical object (e.g., the topology of the surface), etc. In some examples, the image is used to determine the boundary of the physical object in the field of view of the Radar sensor 202c.
[0044] The microphone 202d includes at least one device configured to communicate with the communication device 202e, the autonomous vehicle computing 202f, and / or the safety controller 202g via a bus (e.g., a bus the same as or similar to the bus 302 of Figure 3 . The microphone 202d includes one or more microphones (e.g., an array microphone and / or an external microphone, etc.) that capture an audio signal and generate data associated with (e.g., representing) the audio signal. In some examples, the microphone 202d includes a transducer device and / or a similar device. In some embodiments, one or more of the systems described herein may receive the data generated by the microphone 202d and determine the position (e.g., distance, etc.) of an object relative to the vehicle 200 based on the audio signal associated with the data.
[0045] The communication device 202e includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the autonomous vehicle computing 202f, the safety controller 202g, and / or the DBW (drive-by-wire) system 202h. For example, the communication device 202e may include a device the same as or similar to the communication interface 314 of Figure 3 . In some embodiments, the communication device 202e includes a vehicle-to-vehicle (V2V) communication device (e.g., a device for enabling wireless communication of data between vehicles).
[0046] The autonomous vehicle computing 202f includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the safety controller 202g, and / or the DBW system 202h. In some examples, the autonomous vehicle computing 202f includes devices such as a client device, a mobile device (e.g., a cellular phone and / or a tablet computer, etc.), and / or a server (e.g., a computing device including one or more central processing units and / or graphics processing units, etc.). In some embodiments, the autonomous vehicle computing 202f is the same as or similar to the autonomous vehicle (AV) computing 400 described herein. Additionally or alternatively, in some embodiments, the autonomous vehicle computing 202f is configured to communicate with an autonomous vehicle system (e.g., an autonomous vehicle system the same as or similar to the remote AV system 114 of Figure 1 , a queue management system (e.g., a queue management system the same as or similar to the queue management system 116 of Figure 1 ), a V2I device (e.g., a V2I device the same as or similar to the V2I device 110 of Figure 1 ), and / or a V2I system (e.g., a V2I system the same as or similar to the Figure 1communicate with the V2I system 118 or a V2I system that is the same as or similar to it.
[0047] The safety controller 202g includes at least one device configured to communicate with the camera 202a, the LiDAR sensor 202b, the Radar sensor 202c, the microphone 202d, the communication device 202e, the autonomous vehicle computing 202f, and / or the DBW system 202h. In some examples, the safety controller 202g includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). In some embodiments, the safety controller 202g is configured to generate control signals that take precedence over (e.g., override) the control signals generated and / or transmitted by the autonomous vehicle computing 202f.
[0048] The DBW system 202h includes at least one device configured to communicate with the communication device 202e and / or the autonomous vehicle computing 202f. In some examples, the DBW system 202h includes one or more controllers (such as an electrical controller and / or an electromechanical controller, etc.) configured to generate and / or transmit control signals to operate one or more devices of the vehicle 200 (such as the powertrain control system 204, the steering control system 206, and / or the braking system 208, etc.). Additionally or alternatively, one or more controllers of the DBW system 202h are configured to generate and / or transmit control signals to operate at least one different device of the vehicle 200 (such as turn signals, headlights, door locks, and / or windshield wipers, etc.).
[0049] The powertrain control system 204 includes at least one device configured to communicate with the DBW system 202h. In some examples, the powertrain control system 204 includes at least one controller and / or actuator, etc. In some embodiments, the powertrain control system 204 receives control signals from the DBW system 202h, and the powertrain control system 204 causes the vehicle 200 to perform longitudinal vehicle movements (such as starting to move forward, stopping moving forward, starting to move backward, stopping moving backward, accelerating in a certain direction, decelerating in a certain direction, etc.) or lateral vehicle movements (such as making a left turn and / or making a right turn, etc.). In an example, the powertrain control system 204 increases, maintains the same, or decreases the energy (such as fuel and / or electricity, etc.) provided to the motor of the vehicle, thereby causing at least one wheel of the vehicle 200 to rotate or not rotate.
[0050] The steering control system 206 includes at least one device configured to rotate one or more wheels of the vehicle 200. In some examples, the steering control system 206 includes at least one controller and / or actuator, etc. In some embodiments, the steering control system 206 rotates two front wheels and / or two rear wheels of the vehicle 200 left or right to turn the vehicle 200 left or right. In other words, the steering control system 206 causes the activities required to regulate the y-axis component of the vehicle's movement.
[0051] The braking system 208 includes at least one device configured to actuate one or more brakes to decelerate the vehicle 200 and / or keep it stationary. In some examples, the braking system 208 includes at least one controller and / or actuator configured to close one or more calipers associated with one or more wheels of the vehicle 200 on the corresponding rotors of the vehicle 200. Additionally or alternatively, in some examples, the braking system 208 includes an automatic emergency braking (AEB) system and / or a regenerative braking system, etc.
[0052] In some embodiments, the vehicle 200 includes at least one platform sensor (not explicitly illustrated) for measuring or inferring the nature of the state or condition of the vehicle 200. In some examples, the vehicle 200 includes platform sensors such as a global positioning system (GPS) receiver, an inertial measurement unit (IMU), a wheel speed sensor, a wheel braking pressure sensor, a wheel torque sensor, an engine torque sensor, and / or a steering angle sensor. Although the braking system 208 is illustrated as being located Figure 2 proximal to the vehicle 200 in, the braking system 208 can be located anywhere in the vehicle 200.
[0053] Now refer to Figure 3 , a schematic diagram of the illustrative device 300. As illustrated, the device 300 includes a processor 304, a memory 306, a storage component 308, an input interface 310, an output interface 312, a communication interface 314, and a bus 302. In some embodiments, the device 300 corresponds to: at least one device of the vehicle 102 (e.g., at least one device of the system of the vehicle 102); and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112). In some embodiments, one or more devices of the vehicle 102 (e.g., one or more devices of the system of the vehicle 102), and / or one or more devices of the network 112 (e.g., one or more devices of the system of the network 112) include at least one device 300 and / or at least one component of the device 300. As Figure 3As shown, device 300 includes bus 302, processor 304, memory 306, storage component 308, input interface 310, output interface 312, and communication interface 314.
[0054] Bus 302 includes components that permit communication between the components of device 300. In some cases, processor 304 includes a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or an accelerated processing unit (APU), etc.), a microphone, a digital signal processor (DSP), and / or any processing component that can be programmed to perform at least one function (e.g., a field programmable gate array (FPGA) and / or an application specific integrated circuit (ASIC), etc.). Memory 306 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic and / or static storage device that stores data and / or instructions for use by processor 304 (e.g., flash memory, magnetic memory, and / or optical memory, etc.).
[0055] Storage component 308 stores data and / or software related to the operation and use of device 300. In some examples, storage component 308 includes a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk, etc.), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cassette tape, a magnetic tape, a CD-ROM, RAM, PROM, EPROM, FLASH-EPROM, NV-RAM, and / or another type of computer-readable medium, and corresponding drives.
[0056] Input interface 310 includes components that permit device 300 to receive information such as via a user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, and / or a camera, etc.). Additionally or alternatively, in some embodiments, input interface 310 includes sensors for sensing information (e.g., a global positioning system (GPS) receiver, an accelerometer, a gyroscope, and / or an actuator, etc.). Output interface 312 includes components for providing output information from device 300 (e.g., a display, a speaker, and / or one or more light emitting diodes (LEDs), etc.).
[0057] In some embodiments, communication interface 314 includes transceiver-like components (e.g., transceivers and / or separate receivers and transmitters, etc.) that permit device 300 to communicate with other devices via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. In some examples, communication interface 314 permits device 300 to receive information from and / or provide information to another device. In some examples, communication interface 314 includes an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, and / or a cellular network interface, etc.
[0058] In some embodiments, device 300 performs one or more processes described herein. Device 300 performs these processes based on software instructions stored by a computer-readable medium such as memory 306 and / or storage component 308, executed by processor 304. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a storage space located within a single physical storage device or a storage space distributed across multiple physical storage devices.
[0059] In some embodiments, software instructions are read into memory 306 and / or storage component 308 from another computer-readable medium or from another device via communication interface 314. When executed, the software instructions stored in memory 306 and / or storage component 308 cause processor 304 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry is used in place of or in combination with software instructions to perform one or more processes described herein. Thus, unless otherwise explicitly stated, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.
[0060] Memory 306 and / or storage component 308 includes a data store or at least one data structure (e.g., a database, etc.). Device 300 is capable of receiving information from, storing information in, communicating information to, or searching for information stored in the data store or at least one data structure in memory 306 or storage component 308. In some examples, the information includes network data, input data, output data, or any combination thereof.
[0061] In some embodiments, apparatus 300 is configured to execute software instructions stored in memory 306 and / or in the memory of another apparatus (e.g., another apparatus that is the same as or similar to apparatus 300). As used herein, the term "module" refers to at least one instruction stored in memory 306 and / or in the memory of another apparatus, which, when executed by processor 304 and / or by a processor of another apparatus (e.g., another apparatus that is the same as or similar to apparatus 300), causes apparatus 300 (e.g., at least one component of apparatus 300) to perform one or more processes described herein. In some embodiments, a module is implemented in software, firmware, and / or hardware, etc.
[0062] Provide Figure 3 The number and arrangement of the illustrated components are provided as an example. In some embodiments, compared to Figure 3 the illustrated components, apparatus 300 may include additional components, fewer components, different components, or components arranged differently. Additionally or alternatively, a set of components of apparatus 300 (e.g., one or more components) may perform one or more functions described as being performed by another component or another set of components of apparatus 300.
[0063] Now refer to Figure 4, an example block diagram of an autonomous vehicle computing 400 (sometimes referred to as an "AV stack") is illustrated. As illustrated, the autonomous vehicle computing 400 includes a perception system 402 (sometimes referred to as a perception module), a planning system 404 (sometimes referred to as a planning module), a positioning system 406 (sometimes referred to as a positioning module), a control system 408 (sometimes referred to as a control module), and a database 410. In some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in and / or implemented in an automatic navigation system of the vehicle (e.g., the autonomous vehicle computing 202f of the vehicle 200). Additionally or alternatively, in some embodiments, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems (e.g., one or more systems identical or similar to the autonomous vehicle computing 400, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, the control system 408, and the database 410 are included in one or more independent systems located in the vehicle and / or at least one remote system as described herein. In some embodiments, any and / or all of the systems included in the autonomous vehicle computing 400 are implemented in software (e.g., software instructions stored in a memory), computer hardware (e.g., via a microprocessor, a microcontroller, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA), etc.), or a combination of computer software and computer hardware. It will also be understood that in some embodiments, the autonomous vehicle computing 400 is configured to communicate with remote systems (e.g., an autonomous vehicle system identical or similar to the remote AV system 114, a queue management system identical or similar to the queue management system 116, and / or a V2I system identical or similar to the V2I system 118, etc.).
[0064] In some embodiments, the perception system 402 receives data associated with at least one physical object in the environment (e.g., data used by the perception system 402 to detect at least one physical object), and classifies the at least one physical object. In some examples, the perception system 402 receives image data captured by at least one camera (e.g., camera 202a), the image being associated with one or more physical objects within the field of view of the at least one camera (e.g., representing the one or more physical objects). In such examples, the perception system 402 classifies the at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, based on the classification of the physical object by the perception system 402, the perception system 402 transmits data associated with the classification of the physical object to the planning system 404.
[0065] In some embodiments, the planning system 404 receives data associated with a destination, and generates data associated with at least one route (e.g., route 106) along which a vehicle (e.g., vehicle 102) can travel toward the destination. In some embodiments, the planning system 404 periodically or continuously receives data from the perception system 402 (e.g., the data associated with the classification of the physical object described above), and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the perception system 402. In other words, the planning system 404 can perform tasks related to the tactical functions required to operate the vehicle 102 in road traffic. Tactical efforts involve maneuvering the vehicle in traffic during the journey, which includes but is not limited to deciding whether and when to overtake another vehicle, change lanes, or select an appropriate speed, acceleration, deceleration, etc. In some embodiments, the planning system 404 receives data associated with an updated position of the vehicle (e.g., vehicle 102) from the positioning system 406, and the planning system 404 updates at least one trajectory or generates at least one different trajectory based on the data generated by the positioning system 406.
[0066] In some embodiments, the positioning system 406 receives data associated with (e.g., representing) the location of a vehicle (e.g., vehicle 102) in an area. In some examples, the positioning system 406 receives LiDAR data associated with at least one point cloud generated by at least one LiDAR sensor (e.g., LiDAR sensor 202b). In certain examples, the positioning system 406 receives data associated with at least one point cloud from multiple LiDAR sensors, and the positioning system 406 generates a combined point cloud based on the respective point clouds. In these examples, the positioning system 406 compares the at least one point cloud or the combined point cloud with a two-dimensional (2D) and / or three-dimensional (3D) map of the area stored in the database 410. Then, based on the positioning system 406 comparing the at least one point cloud or the combined point cloud with the map, the positioning system 406 determines the position of the vehicle in the area. In some embodiments, the map includes a combined point cloud of the area generated prior to the navigation of the vehicle. In some embodiments, the map includes, but is not limited to, a high-precision map of the roadway geometry, a map describing the connectivity of the road network, a map describing the physical properties of the roadways (such as traffic speed, traffic flow, the number of vehicle and bicycle traffic lanes, lane width, lane traffic direction, or the type and location of lane markings, or a combination thereof, etc.), and a map describing the spatial location of road features (such as crosswalks, traffic signs, or various other types of driving signal lights, etc.). In some embodiments, the map is generated in real time based on the data received by the perception system.
[0067] In another example, the positioning system 406 receives Global Navigation Satellite System (GNSS) data generated by a Global Positioning System (GPS) receiver. In some examples, the positioning system 406 receives GNSS data associated with the location of a vehicle in an area, and the positioning system 406 determines the latitude and longitude of the vehicle in the area. In such examples, the positioning system 406 determines the position of the vehicle in the area based on the latitude and longitude of the vehicle. In some embodiments, the positioning system 406 generates data associated with the position of the vehicle. In some examples, based on the positioning system 406 determining the position of the vehicle, the positioning system 406 generates data associated with the position of the vehicle. In such examples, the data associated with the position of the vehicle includes data associated with one or more semantic properties corresponding to the position of the vehicle.
[0068] In some embodiments, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle. In some examples, the control system 408 receives data associated with at least one trajectory from the planning system 404, and the control system 408 controls the operation of the vehicle by generating and transmitting control signals to cause the powertrain control system (e.g., the DBW system 202h and / or the powertrain control system 204, etc.), the steering control system (e.g., the steering control system 206), and / or the braking system (e.g., the braking system 208) to operate. For example, the control system 408 is configured to perform operating functions such as lateral vehicle motion control or longitudinal vehicle motion control. Lateral vehicle motion control causes activities required to regulate the y-axis component of the vehicle motion. Longitudinal vehicle motion control causes activities required to regulate the x-axis component of the vehicle motion. In an example, in the case where the trajectory includes a left turn, the control system 408 transmits a control signal to cause the steering control system 206 to adjust the steering angle of the vehicle 200, thereby causing the vehicle 200 to turn left. Additionally or alternatively, the control system 408 generates and transmits control signals to cause other devices of the vehicle 200 (e.g., headlights, turn signals, door locks, and / or windshield wipers, etc.) to change states.
[0069] In some embodiments, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model (e.g., at least one multi-layer perceptron (MLP), at least one convolutional neural network (CNN), at least one recurrent neural network (RNN), at least one autoencoder, and / or at least one transformer, etc.). In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model individually or in combination with one or more of the above systems. In some examples, the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408 implement at least one machine learning model as part of a pipeline (e.g., a pipeline for identifying one or more objects located in the environment, etc.).
[0070] The database 410 stores data transmitted to, received from, and / or updated by the perception system 402, the planning system 404, the positioning system 406, and / or the control system 408. In some examples, the database 410 includes a storage component for storing operation-related data and / or software and using the autonomous vehicle computing 400 of at least one system (e.g., associated with Figure 3the same or similar storage components as the storage component 308). In some embodiments, the database 410 stores data associated with 2D and / or 3D maps of at least one area. In some examples, the database 410 stores data associated with 2D and / or 3D maps of a part of a city, multiple parts of multiple cities, multiple cities, counties, states, and / or countries (e.g., a country), etc. In such examples, a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200) can drive along one or more drivable areas (e.g., a single-lane road, a multi-lane road, a highway, a back road, and / or an off-road path, etc.), and cause at least one LiDAR sensor (e.g., a LiDAR sensor the same or similar to the LiDAR sensor 202b) to generate data associated with an image representing the objects included in the field of view of the at least one LiDAR sensor.
[0071] In some embodiments, the database 410 can be implemented across multiple devices. In some examples, the database 410 is included in a vehicle (e.g., a vehicle the same or similar to the vehicle 102 and / or the vehicle 200), an autonomous vehicle system (e.g., an autonomous vehicle system the same or similar to the remote AV system 114), a queue management system (e.g., a queue management system the same or similar to Figure 1 the queue management system 116) and / or a V2I system (e.g., a V2I system the same or similar to Figure 1 the V2I system 118), etc.
[0072] Now referring to Figure 5 , a diagram of an implementation 500 of a process for using a sensor suite to detect objects is illustrated. In some embodiments, the implementation 500 includes an autonomous system (e.g., Figure 2 the autonomous system 202), which includes a camera 504a, a Radar sensor 504c (e.g., Figure 2 the camera 202a, the Radar sensor 202c) and an AV computing 502 (e.g., the AV computing 202f). In some embodiments, data generated by the camera 504a and the Radar sensor 504c is obtained by Figure 3 the device 300 to detect objects near the AV.
[0073] In the implementation 500, the AV computing 502 (e.g., Figure 2 the AV computing 202f) includes a perception system 506 (e.g., Figure 4 the perception system 402), a planning system 510 (e.g., Figure 4 the planning system 404) and a control system 514 (e.g., Figure 4The control system 408). The perception system 506 obtains Radar data output by at least one Radar sensor 504c, and the Radar data is associated with one or more physical objects within the field of view of at least one Radar sensor 504c (e.g., represents one or more physical objects). In an example, the Radar sensor is a four-dimensional (4D) Radar sensor having range, azimuth, elevation, and Doppler dimensions. The Radar sensor includes, for example, a millimeter-wave (mmWave) sensor for operating in a frequency range such as 60 - 64 GHz and 76 - 81 GHz frequencies. In an example, the 4D Radar sensor uses a multiple-input multiple-output antenna array system to capture high-resolution data corresponding to the environment.
[0074] The perception system 506 classifies at least one physical object based on one or more groupings of physical objects (e.g., bicycles, vehicles, traffic signs, and / or pedestrians, etc.). In some embodiments, the perception system 506 transmits data associated with the classification of a physical object (e.g., the classified object 508) to the planning system 510. In some examples, the perception system 506 also receives image data captured by at least one camera 504a, and the image is associated with one or more physical objects within the field of view of at least one camera 504a (e.g., represents the one or more physical objects). The image data captured by at least one camera 504a is fused with the Radar data output by at least one Radar sensor 504c, and the fused data associated with one or more physical objects (e.g., represents the one or more physical objects) is provided to the perception system 506 for object classification.
[0075] The planning system 510 determines a trajectory (512) to be used for AV navigation. For example, the planning system 510 periodically or continuously receives data from the perception system 506 (e.g., Figure 4 the perception system 402), and the data includes classified objects 508 in the environment (e.g., Figure 1 the environment 100). The planning system 510 determines at least one trajectory 512 based on the classified objects 508 generated by the perception system 402. The trajectory 512 is transmitted to the control system 514 for controlling the operation of the AV.
[0076] In some embodiments, the Radar data is represented as a 1D range heatmap tensor, a 2D range-Doppler (RD) heatmap tensor (also referred to as an "RD spectrum"), a 2D range-azimuth (RA) heatmap tensor (also referred to as an "RA spectrum"), or a three-dimensional (3D) range-azimuth-Doppler (RAD) matrix tensor (also referred to as an "RAD spectrum"). At least one of the Radar data representations is input into a machine learning model to detect objects, and the machine learning model outputs a classification of the detected objects. In an example, the machine learning model is implemented by the Radar sensor 504c, the perception system 506, or any combination thereof. In an example, the machine learning model is implemented by a device separate from the Radar sensor 504c or the perception system 506. In an example, the machine learning model is implemented on a controller (e.g., a domain controller) or a device of the AV (e.g., Figure 1 at least one device in the system of the vehicle 102) or a device of the autonomous vehicle computing 502. Additionally, in an example, the device separate from the Radar sensor 504c and the perception system 506 is custom hardwired logic, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), or firmware and / or program logic that, in combination with a computer system (e.g., the device 300), causes the computer system (e.g., the device 300) to become a special purpose machine or programs the computer system to be a special purpose machine.
[0077] In some embodiments, one of the Radar data representations is fused with a camera image from the camera 504a, and the fused data is input into a machine learning model for object detection and classification. In an example, the machine learning model outputs a classification of the detected objects. The machine learning model can also be implemented by a device separate from the Radar sensor 504c and the perception system 506. For example, the machine learning model is implemented by a controller (e.g., a domain controller), a device of the AV (e.g., Figure 1 at least one device in the system of the vehicle 102) or a device of the autonomous vehicle computing 502.
[0078] FIG. 6(a) is an example pipeline 600 implemented by a Radar sensor. In some embodiments, one or more of the steps of the pipeline 600 are performed (e.g., fully and / or partially, etc.) by a device or system (or group of devices and / or systems) separate from or including the autonomous system. For example, one or more of the steps of the pipeline 600 are performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102 or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of the vehicle 102 or 200),Figure 3 device 300 and / or Figure 5 Radar sensor 504c (e.g., completely and / or partially, etc.) to perform. In some embodiments, the steps of conduit 600 are performed between any of the above systems that cooperate with each other.
[0079] The transmitter (Tx) antenna 602 is an antenna for radiating the pulsed wave generated by the transmitter of the Radar sensor as a beam in a predetermined direction. The receiver (Rx) antenna 604 is an antenna for receiving the radio wave reflected by the object detected by the Radar sensor. The Radar sensor includes a radio frequency (RF) front end 606 that receives and demodulates the radio wave received from the Rx antenna 604 and generates a baseband signal. The baseband signal is further processed by a baseband signal processing chain 608, which includes one or more filters for removing signals in unwanted sidebands and image frequencies, and one or more amplifiers. The analog signal output from the baseband signal processing chain 608 is obtained by an analog-to-digital converter (ADC) 610. The digital signal output by the ADC 610 is raw data sampled at a high data rate (e.g., data recorded directly from the reflected wave captured by the Radar sensor). A range FFT 612 is performed on the ADC output data (digital signal) to extract range data from the ADC output data. A Doppler FFT 614 is performed on the ADC output data to extract velocity data from the ADC output data. An azimuth FFT 616 is performed on the ADC output data to extract angle data from the ADC output data. An elevation FFT 618 is performed on the ADC output data to extract elevation data from the ADC output data. The ADC output data may include noise and clutter, which may cause false detection of objects. To remove noise and clutter, signal processing is performed on the ADC output data. In some embodiments, a constant false alarm rate (CFAR) detection algorithm 620 is applied to achieve a false alarm probability below a predetermined threshold. In an example, the CFAR detection algorithm 620 is an adaptive algorithm used in the Radar system to detect target echoes in the presence of noise, clutter, and interference. Object association 622 is performed to track an object and its movement. Object association 622 associates a set of points with an object based on the position and velocity of each point, and then tracks the movement of the object as a whole. A tracking filter 624 is applied to improve the estimation of the tracking position of the object and correct errors in previous predictions. The tracking filter 624 includes, for example, an Alpha-Beta tracker and a Kalman filter. The filtered data is obtained by an electronic control unit (ECU) 626 for processing Radar sensor data to trigger critical advanced driver assistance system (ADAS) features. The output from the ECU 626 is 4D Radar data, which includes a 3D point cloud and velocity of surrounding objects near the AV.
[0080] Figure 6(b) is another example pipeline 630 implemented by a Radar sensor. In some embodiments, one or more steps in the pipeline 630 are performed (e.g., fully and / or partially, etc.) by a device or system (or group of devices and / or systems) separate from or including the autonomous system. For example, one or more steps of the pipeline 630 are performed by Figure 1 the remote AV system 114 of Figure 1 the vehicle 102 of Figure 2 the vehicle 200 of (e.g., the autonomous system 202 of vehicle 102 or 200), Figure 3 the device 300 of Figure 5 the Radar sensor 504c of (e.g., fully and / or partially, etc.). In some embodiments, the steps of the pipeline 630 are performed between any of the above systems that cooperate with each other.
[0081] Similar to FIG. 6(a), FIG. 6(b) shows a Tx antenna 602, an Rx antenna 604, an RF front end 606, a baseband signal processing chain 608, an ADC 610, a range FFT block 612, a Doppler FFT block 614, and an azimuth FFT block 616. Instead of the 4D Radar data output in FIG. 6(a), the output of FIG. 6(b) is one of four tensors (1D range heatmap tensor, 2D RD heatmap tensor, 2D RA heatmap tensor, 3D RAD matrix tensor). In some embodiments, the range data output from the range FFT block 612 may form a 1D range heatmap tensor, which is input to a machine learning model 632. In these embodiments, the machine learning model 632 for object detection replaces all the blocks from the Doppler FFT 614 to the ECU 626 in FIG. 6(a). In some embodiments, the range data output from the range FFT block 612 is combined with the velocity data output from the Doppler FFT block 614 to form a 2D RD heatmap tensor, which is input to the machine learning model 632. In these embodiments, the machine learning model 632 for object detection replaces all the blocks from the azimuth FFT 616 to the ECU 626 in FIG. 6(a). In some embodiments, the range data output from the range FFT block 612 is combined with the angle data output from the azimuth FFT block 616 to form a 2D RA heatmap tensor, which is input to the machine learning model 632. In these embodiments, the machine learning model 632 for object detection replaces the Doppler FFT 614 and all the blocks from the elevation FFT 618 to the ECU 626 in FIG. 6(a). In some embodiments, the range data output from the range FFT block 612, the velocity data output from the Doppler FFT block 614, and the angle data output from the azimuth FFT block 616 are combined to form a 3D RAD matrix tensor, which is input to the machine learning model 632. In these embodiments, the machine learning model 632 for object detection replaces all the blocks from the elevation FFT 618 to the ECU 626 in FIG. 6(a). The machine learning model 632 is configured to detect objects and is different (e.g., includes different layers) for each of the four tensors (1D range heatmap tensor, 2D RD heatmap tensor, 2D RA heatmap tensor, 3D RAD matrix tensor).
[0082] In some embodiments, the exemplary pipeline 630 may further include an elevation FFT block 618 as the pipeline 600 of FIG. 6(a). The machine learning model 632 for detecting objects replaces all the blocks from the CFAR detection 620 to the ECU 626 in FIG. 6(a). The elevation FFT block 618 may be placed between the azimuth FFT block 616 and the machine learning model 632. Compared with the angle data from the azimuth FFT block 616, the elevation data from the elevation FFT block 618 carries much less information because the measurement distance and resolution of the elevation data are much lower compared to the angle data.
[0083] Provide one of the 1D distance heatmap tensor, 2D RD heatmap tensor, 2D RA heatmap tensor, and 3D RAD matrix tensor to the machine learning model 632 (the machine learning model 632 is a different model for each tensor). The 1D distance heatmap tensor, 2D RD heatmap tensor, 2D RA heatmap tensor, or 3D RAD matrix tensor is a different representation of the Radar data. Without performing signal processing including the CFAR detection 620, object association 622, and tracking filter 624, the distance data, velocity data, and angle data are directly extracted from the raw data of the ADC 610. Even though signal processing can remove noise and clutter, the signal processing (e.g., the CFAR detection algorithm 620) may incorrectly remove useful data that is not noise or clutter. Compared with the 4D Radar data output from the ECU 626 of FIG. 6(a), the 1D distance heatmap tensor, 2D RD heatmap tensor, 2D RA heatmap tensor, or 3D RAD matrix tensor includes richer data.
[0084] In some embodiments, the machine learning model 632 for detecting and / or classifying objects is implemented by a Radar sensor (e.g., Figure 2 the Radar sensor 202c of Figure 5 or the Radar sensor 504c of Figure 4 ). In some embodiments, the machine learning model 632 is implemented by a perception system (e.g., Figure 5 the perception system 402 of Figure 3 or the perception system 506 of Figure 2 ). In some embodiments, the machine learning model 632 is implemented by a device separate from the Radar sensor and the perception system 506. For example, the machine learning model 632 is in a controller (e.g., a domain controller) or a computer in the AV (e.g., Figure 4 the device 300 of Figure 5It is implemented on the AVC 502). In some embodiments, the machine learning model 632 is any deep learning model for detecting objects (e.g., convolutional neural network (CNN), feature pyramid network (FPN), etc.). One of the tensors representing the Radar data is input into the machine learning model 632, and the machine learning model outputs the detected objects. In an example, the detected objects are identified by bounding boxes and classification labels around each object of interest in the Radar data frame.
[0085] FIG. 6(c) is another example pipeline 650 of the Radar sensor. In some embodiments, one or more of the steps in the pipeline 650 are performed (e.g., fully and / or partially, etc.) by a device or system (or a group of devices and / or systems) separate from or including the autonomous system. For example, one or more of the steps of the pipeline 650 are performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102 or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of the vehicle 102 or 200), Figure 3 the device 300 and / or Figure 5 the Radar sensor 504c (e.g., fully and / or partially, etc.). In some embodiments, the steps of the pipeline 650 are performed among any of the above systems that cooperate with each other.
[0086] The example pipeline 650 of the example Radar sensor does not include the azimuth FFT block 616 in FIGS. 6(a) and 6(b). Thus, the output of the example Radar sensor is a 2D RD heat map. However, in order to implement as Figures 9(a) to 9(c)The feature fusion shown requires angular data. The 2D RD heatmap tensor can be provided to the first machine learning model 652 to obtain a 2D RA heatmap including angular data. In the example of FIG. 6(c), the 2D RD heatmap tensor is generated as a combination of range data from the range FFT 612 and velocity data from the Doppler FFT 614. In the example, as shown in FIG. 6(c), the first Radar data representation (e.g., 2D RD heatmap tensor) is input to the first machine learning model 652, and the first machine learning model 652 outputs a second Radar data representation (e.g., 2D RA heatmap). The second Radar data representation (e.g., 2D RA heatmap) is input to the second machine learning model 654, and the second machine learning model outputs the location (e.g., bounding box) and classification of the object detected in the Radar data. In the example of FIG. 6(c), the 2D RD heatmap tensor is provided to the first machine learning model 652 to obtain a 2D RA heatmap tensor, which is provided to the second machine learning model 654 to detect objects near the AV.
[0087] In the example, as shown in FIG. 6(c), the first Radar data representation (e.g., 1D range heatmap tensor) is input to the first machine learning model 652, and the first machine learning model 652 outputs a second Radar data representation (e.g., 2D DA heatmap). The first machine learning model 652 varies according to the input (1D range heatmap tensor or 2D RD heatmap tensor) (e.g., includes different layers).
[0088] In some embodiments, the first machine learning model 652 includes an encoder / decoder neural network architecture. For example, the first machine learning model 652 includes a multi-input multi-output (MIMO) precoder 656, an FPN encoder 658, and a range-angle decoder 660. The MIMO precoder 656 reorganizes and compresses the 2D RD heatmap tensor to facilitate the subsequent utilization of MIMO information (to recover the angle) while keeping the data volume under control. The MIMO precoder 656 learns how to combine the input channels (number of receiver antennas) and compress the Radar data.
[0089] The FPN encoder 658 uses a pyramid structure to learn multi-scale features. In an embodiment, the FPN encoder 658 includes four Resnet blocks (RNBs) composed of 3, 6, 6, and 3 residual layers respectively. The feature maps of these residual layers form a feature pyramid. The channel dimension is selected to encode the azimuth angle within the entire distance range (i.e., at long distances, high resolution and narrow field of view, at short distances, low resolution and wider field of view). To prevent the loss of Radar data for small objects (usually several pixels in the RD heatmap tensor), the FPN encoder 658 downsamples each Resnet block by, for example, 2×2, so that the tensor size is reduced by a total of 16 times in height and width. The FPN encoder 658 uses, for example, a 3×3 convolutional kernel.
[0090] The distance-angle decoder 660 expands the input FPN feature map into a higher-resolution representation. The dimensions of the tensor provided to the distance-angle decoder 660 correspond to distance, Doppler, and azimuth angle respectively, while the feature map corresponds to a distance-azimuth representation. Therefore, the Doppler and azimuth axes are swapped to match the final axis ordering, and then the feature map is upscaled. A basic block (BB) of two Conv-BatchNorm-ReLU layers is also applied in the distance-angle decoder 660 to generate a distance-azimuth heatmap tensor.
[0091] In some embodiments, the second machine learning model 654 includes a detection head 662 for locating the vehicle in distance-azimuth coordinates and a segmentation head 664 for predicting the free driving space. The detection head 662 processes the input RA heatmap tensor using a first common sequence of four Conv-BatchNorm blocks (CB) each having, for example, 144, 96, 96, and 96 filters. In an example, the terms "backbone" and "head" refer to the structure of the second machine learning model 654. In an example, the backbone extracts features from the data, and one or more heads use these features to perform a predetermined task. In some embodiments, the segmentation head outputs a mask for each pixel indicating whether an object is present. Additionally, in some embodiments, the detection head includes a classification head and a bounding box regression head. The detection head outputs the classification and bounding box for each object in the Radar data.
[0092] The classification head includes a convolutional layer with sigmoid activation for predicting the probability map. The output of the classification head is a binary classification of whether each "pixel" is occupied by an object or not. The regression head finely predicts the distance and azimuth values corresponding to the detected object. The regression head applies a 3×3 convolutional layer to output two feature maps corresponding to the final distance and azimuth values.
[0093] The segmentation head 664 is formulated as a pixel-level binary classification. The segmentation mask has a resolution of, for example, 0.4 m in distance and 0.2° in azimuth. This is equivalent to half of the native distance and azimuth resolution while only considering half of the entire azimuth field of view (FoV) (within [-45°, 45°]). The RA heatmap tensor is processed by two sets of two consecutive Conv-BatchNorm-ReLu blocks (BBs), resulting in 128 and 64 feature maps respectively. A final 1×1 convolutional block is applied to output a 2D feature map, followed by sigmoid activation to estimate the probability of drivability at each location.
[0094] Figures 7(a) to 7(d) are different example representations of Radar data. The 3D RAD matrix tensor 702 includes the combination of distance data from the Figures 6(a) to 6(c) distance FFT block 612, velocity data from the Figures 6(a) to 6(c) Doppler FFT block 614, and angle data from the Figures 6(a) to 6(b) azimuth FFT block 616. The 2D RD heatmap tensor 704 includes the combination of distance data from the Figures 6(a) to 6(c) distance FFT block 612 and velocity data from the Figures 6(a) to 6(c) Doppler FFT block 614. The 2D RA heatmap tensor 706 includes the combination of distance data from the Figures 6(a) to 6(c) distance FFT block 612 and angle data from the Figures 6(a) to 6(b) azimuth FFT block 616. The 1D distance heatmap tensor 708 includes the distance data from the Figures 6(a) to 6(c) distance FFT block 612.
[0095] Figure 8 is a diagram illustrating the generation of an example 3D RAD matrix tensor representing Radar data. In the example, the Radar sensor transmits chirps, and each chirp is a frequency-swept signal. The frequency of the Radar signal is swept (or modulated) from low to high or from high to low over time. In the example, the chirp signal is a continuous sine wave whose frequency changes by several hundred megahertz from a specific frequency (e.g., 77 Ghz) over the entire chirp. As Figure 8As shown, the number of samples per chirp is m, the number of chirps (swept signals) per frame is n, and the number of ADC channels (the same as the number of Radar receivers) is h. An example 3D RAD matrix tensor (Radar data cube) is generated to represent this data. The size of the Radar data cube is m×n×h. In some embodiments, the values of m, n, and h are configurable such that the size of the Radar data cube is changed. The number of samples per chirp (the number of samples obtained during the period of the chirp) m is represented by the range data from the Figures 6(a) to 6(c) range FFT block 612; the number of chirps per frame n is represented by the velocity data from the Figures 6(a) to 6(c) Doppler FFT block 614; and the number of ADC channels h is represented by the angle data from the Figures 6(a) to 6(b) azimuth FFT block 616. In some embodiments, the size of the Radar data cube can be configured based on measurement parameters (e.g., sensing range (maximum range), range resolution, maximum velocity, velocity resolution, and angle resolution). The maximum velocity is related to the period of the chirp and can be represented by Equation (1): V max =λ / (4T c ), where λ is the wavelength of the electromagnetic wave signal, and T c is the period of the chirp. The velocity resolution is related to the period of the frame and can be represented by Equation (2): V res =λ / (2T f ), where λ is the wavelength of the electromagnetic wave signal, and T f is the period of the frame. The range resolution is related to the bandwidth of the chirp and can be represented by Equation (3): d res =C / (2B), where C is the speed of light, and B is the bandwidth of the chirp. The maximum range is related to the bandwidth of the chirp and the period of the chirp and can be represented by Equation (4): d max =CF IFmax / (2S), where C is the speed of light, F IFmax is the maximum intermediate frequency, S = B / T c (where B is the bandwidth of the chirp, and T c is the period of the chirp). The period T c of the chirp and the sampling rate determine the value of m. The period T c of the chirp and the period T f of the frame determine the value of n. The angle resolution is related to the number of ADC channels or the number of receivers and determines the value of h.
[0096] Similarly, the size of the 2D RD heatmap tensor is m×n, and the size of the 2D RA heatmap tensor is m×h, and the size of the 1D distance heatmap tensor is m. The size of the Radar data cube, 2D RD heatmap tensor, 2D RA heatmap tensor, or 1D distance heatmap tensor can be configured according to different Radar sensors. For example, for short-range Radar, fewer ADC channels are required, so the value of h is smaller. For another example, for long-range Radar, a lower range resolution (i.e., a higher range resolution value d res ), which results in a lower bandwidth (according to Equation (3)) and a higher range (according to Equation (4)). In an embodiment, the period T of the chirp c is slightly increased in combination with the lower bandwidth to obtain an even higher range (according to Equation (4)). Therefore, due to the increase in the period T of the chirp c , the value of m is larger.
[0097] FIG. 9(a) is an example pipeline 900 for sensor data fusion. In some embodiments, one or more steps in the steps of pipeline 900 are performed (e.g., completely and / or partially, etc.) by a device or system (or a group of devices and / or systems) separate from or including the autonomous system. For example, one or more steps of pipeline 900 are performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102 or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of vehicle 102 or 200), Figure 3 the device 300 and / or Figure 5 the Radar sensor 504c (e.g., completely and / or partially, etc.). In some embodiments, the steps of pipeline 900 are performed among any of the above systems that cooperate with each other.
[0098] Referring to FIGS. 9(a) and 6(b), range FFT 612, Doppler FFT 614, and azimuth FFT 616 are performed on the ADC raw data 610 of the Radar sensor to obtain a 3D RAD matrix tensor 902 representing the Radar data. The Radar data is fused with the camera image captured by a camera (e.g., Figure 2 the camera 202a or Figure 5 the camera 504a) for improved object detection.
[0099] The RAD matrix tensor 902 is provided to a feature extraction layer 904 that extracts features from the Radar data. A spatial transformation such as a polar-to-cartesian coordinate transformation is performed on the extracted Radar features by the spatial transformer 906.
[0100] Similarly, a camera image is provided to a feature extraction layer 904 that extracts features from the image. A spatial transformation, such as a homography transformation, is performed by a spatial transformer 906 to transform the camera image into Cartesian space. To calculate this projection mapping, it is assumed that a camera (e.g., Figure 2 camera 202a of Figure 5 or camera 504a of
[0101] is taking pictures of a planar scene (i.e., a Radar plane approximately parallel to the road plane). Then, a set of points in the Cartesian Radar plane is projected onto image coordinates using intrinsic and extrinsic calibration information. Then, a planar homography transformation is performed using, for example, a standard 4-point algorithm. If the calibration information is not available, multiple tie points can also be manually assigned, and finally, the best homography is solved using the least squares method. After the homography transformation, if the plane assumption is correct and the camera has not moved relative to the Radar sensor, the image coordinates match the Cartesian Radar image coordinates.
[0102] FIG. 9(b) is another example pipeline 930 for sensor data fusion. In some embodiments, one or more steps in the pipeline 930 are performed by a device or system (or a group of devices and / or systems) that is separate from or includes the autonomous system (e.g., fully and / or partially, etc.). For example, one or more steps in the pipeline 930 are performed by Figure 1 the remote AV system 114 of Figure 1 the vehicle 102 of Figure 2 or the vehicle 200 of Figure 3 e.g., the autonomous system 202 of vehicle 102 or 200), Figure 5 the device 300 of
[0103] Compared with the example pipeline 900 of FIG. 9(a), the example pipeline 930 does not include an azimuth FFT block 616. A 2D RD heat map tensor 932 is generated as a combination of range data from the range FFT 612 and velocity data from the Doppler FFT 614. The 2D RD heat map tensor 932 is provided to a first machine learning model 652 (e.g., the first machine learning model 652 of FIG. 6(c)) to obtain a 2D RA heat map tensor 934. Similar to the example pipeline 900 of FIG. 9(a), the camera image and the 2D RA heat map tensor 934 are then fused for object detection.
[0104] In some embodiments, a 1D range heat map tensor 931 is provided to the first machine learning model 652 to obtain a 2D RA heat map tensor 934. The first machine learning model 652 differs (e.g., includes different layers) depending on the input (the 1D range heat map tensor 931 or the 2D RD heat map tensor 932).
[0105] FIG. 9(c) is another example pipeline 950 for sensor data fusion. In some embodiments, one or more of the steps in the pipeline 950 are performed (e.g., fully and / or partially, etc.) by a device or system (or a group of devices and / or systems) separate from or including the autonomous system. For example, one or more of the steps in the pipeline 950 are performed by Figure 1 the remote AV system 114, Figure 1 the vehicle 102, or Figure 2 the vehicle 200 (e.g., the autonomous system 202 of the vehicle 102 or 200), Figure 3 the device 300, and / or Figure 5 the Radar sensor 504c (e.g., fully and / or partially, etc.). In some embodiments, the steps of the pipeline 950 are performed among any of the above systems that cooperate with each other.
[0106] As shown in FIG. 9(c), a 2D RA heat map tensor 934 is generated as a combination of range data from the range FFT 612 and angle data from the azimuth FFT 616. Similar to the example pipeline 900 of FIG. 9(a) and the example pipeline 930 of FIG. 9(b), the camera image and the 2D RA heat map tensor 934 are then fused for object detection.
[0107] Figure 10FIG. 1000 is an example flowchart of a process for generating a representation of Radar data. In some embodiments, one or more of the steps in pipeline 1000 are performed (e.g., fully and / or partially, etc.) by a device or system (or group of devices and / or systems) separate from or including the autonomous system. For example, one or more of the steps in pipeline 1000 are performed by Figure 1 remote AV system 114 of Figure 1 vehicle 102 of Figure 2 vehicle 200 of Figure 3 device 300 of Figure 5 Radar sensor 504c of (e.g., fully and / or partially, etc.). In some embodiments, the steps of pipeline 1000 are performed between any of the above systems that cooperate with each other.
[0108] In some embodiments, at block 1002, a processor (e.g., Figure 3 processor 304 of Figure 2 Radar sensor 202c of Figure 5 processor of Radar sensor 504c) receives ADC raw data of a Radar sensor of a vehicle ( Figure 2 Radar sensor 202c of Figure 5 Radar sensor 504c). Figures 6(a) to 6(c) ADC raw data 610).
[0109] At block 1004, the processor performs a range FFT ( Figures 6(a) to 6(c) range FFT 612), a Doppler FFT ( Figures 6(a) to 6(c) Doppler FFT 614), and an azimuth FFT ( Figures 6(a) to 6(b) azimuth FFT 616) on the ADC raw data. By performing the range FFT, range data is extracted from the ADC raw data. By performing the Doppler FFT, velocity data is extracted from the ADC raw data. By performing the azimuth FFT, angle data is extracted from the ADC raw data.
[0110] At block 1006, the processor generates a one-dimensional (1D) range heatmap tensor representing the range FFT, a two-dimensional range-Doppler (RD) heatmap tensor representing the combination of the range FFT and the Doppler FFT, a two-dimensional range-azimuth (RA) heatmap tensor representing the combination of the range FFT and the azimuth FFT, or a three-dimensional RAD matrix tensor representing the combination of the range FFT, the Doppler FFT, and the azimuth FFT. The processor generates at least one of four Radar data representations.
[0111] At block 1008, the processor provides a 1D distance heatmap tensor, a 2D RD heatmap tensor, a 2D RA heatmap tensor, or a 3D RAD matrix tensor to a machine learning model for detecting objects on a road network around a vehicle. At least one of the four Radar data representations is provided to the machine learning model for object detection.
[0112] The techniques of the present disclosure can provide Radar data representations to a machine learning model in place of sparse point clouds. The machine learning model detects objects based on the rich information in the Radar data representations as early Radar data. The Radar data representations are not subject to the signal processing of the ADC raw data typically performed by conventional Radar sensors. Without "overprocessing" the ADC raw data, the Radar data representations contain rich information and are provided in a format compatible with the machine learning model. The Radar data representations are used to train the machine learning model for object detection.
[0113] According to some non-limiting embodiments or examples, a method is provided that includes receiving analog-to-digital converter (ADC) raw data from a Radar sensor of a vehicle. The method includes performing a distance fast Fourier transform (FFT), a Doppler FFT, and an azimuth FFT on the ADC raw data. The method includes generating: a 2D range-Doppler (RD) heatmap tensor representing a combination of the range FFT and the Doppler FFT, a 2D range-azimuth (RA) heatmap tensor representing a combination of the range FFT and the azimuth FFT, or a 3D range-azimuth-Doppler (RAD) matrix tensor representing a combination of the range FFT, the Doppler FFT, and the azimuth FFT. The method includes providing at least one of a 1D distance heatmap tensor, a 2D RD heatmap tensor, a 2D RA heatmap tensor, and a 3D RAD matrix tensor to a machine learning model for detecting objects on a road network around the vehicle. A system and a computer program product are also provided.
[0114] According to some non-limiting embodiments or examples, a system is provided that includes: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including receiving analog-to-digital converter (ADC) raw data from a Radar sensor of a vehicle. The operations include performing a distance fast Fourier transform (FFT) and a Doppler FFT on the ADC raw data. The operations include generating a two-dimensional (2D) range-Doppler (RD) heatmap tensor representing a combination of the range FFT and the Doppler FFT. The operations include providing the 2D RD heatmap tensor to a machine learning model for detecting objects on a road network around the vehicle.
[0115] According to some non - limiting embodiments or examples, at least one non - transitory computer - readable medium is provided, which includes one or more than one instructions that, when executed by at least one processor, cause the at least one processor to perform operations. The operations include receiving the raw data of an analog - to - digital converter (ADC) of a Radar sensor of a vehicle. The operations include performing a range fast Fourier transform (FFT), a Doppler FFT, and an azimuth FFT on the ADC raw data. The operations include generating a three - dimensional (3D) range - azimuth - Doppler (RAD) matrix tensor representing a combination of the range FFT, the Doppler FFT, and the azimuth FFT. The operations include providing the 3D RAD matrix tensor to a machine - learning model for detecting objects on a road network around the vehicle.
[0116] Clause 1: A method, comprising: using at least one processor, receiving the raw data of an analog - to - digital converter (ADC) of a Radar sensor of a vehicle; using the at least one processor, performing at least one of a range fast Fourier transform (FFT), a Doppler FFT, and an azimuth FFT on the ADC raw data; using the at least one processor, generating a one - dimensional (1D) range heat - map tensor representing the range FFT, a two - dimensional (2D) range - Doppler (RD) heat - map tensor representing a combination of the range FFT and the Doppler FFT, a 2D range - azimuth (RA) heat - map tensor representing a combination of the range FFT and the azimuth FFT, or a three - dimensional (3D) range - azimuth - Doppler (RAD) matrix tensor representing a combination of the range FFT, the Doppler FFT, and the azimuth FFT; and using the at least one processor, providing at least one of the 1D range heat - map tensor, the 2D RD heat - map tensor, the 2D RA heat - map tensor, and the 3D RAD matrix tensor to a machine - learning model for detecting objects on a road network around the vehicle.
[0117] Clause 2: The method according to Clause 1, wherein providing the 1D range heat - map tensor, the 2D RD heat - map tensor, the 2D RA heat - map tensor, or the 3D RAD matrix tensor to the machine - learning model further includes: receiving a camera image from a camera of the vehicle; fusing the camera image with the 1D range heat - map tensor, the 2D RA heat - map tensor, or the 3D RAD matrix tensor; and providing the fused data to the machine - learning model.
[0118] Clause 3: The method according to Clause 1 or 2, wherein the 3D RAD matrix tensor includes the azimuth FFT representing the number of ADC channels, the Doppler FFT representing the number of chirps in each ADC channel, and the range FFT representing the number of samples per chirp.
[0119] Clause 4: The method according to Clause 3, wherein the size of the 3D RAD matrix tensor is configured based on one or more of the sensed distance, range resolution, velocity resolution, angular resolution, bandwidth of the chirp, period of the chirp, number of samples per chirp or sampling rate, and the number of ADC channels.
[0120] Clause 5: The method according to any one of the preceding clauses, wherein the 2D RD heatmap tensor includes the Doppler FFT representing the number of chirps in each ADC channel and the range FFT representing the number of samples per chirp.
[0121] Clause 6: The method according to Clause 5, wherein the size of the 2D RD heatmap tensor is configured based on one or more of the sensed distance, range resolution, velocity resolution, bandwidth of the chirp, period of the chirp, number of samples per chirp, and sampling rate.
[0122] Clause 7: The method according to any one of the preceding clauses, wherein the 2D RA heatmap tensor includes the azimuth FFT representing the number of ADC channels and the range FFT representing the number of samples per chirp.
[0123] Clause 8: The method according to Clause 7, wherein the size of the 2D RA heatmap tensor is configured based on one or more of the sensed distance, range resolution, angular resolution, bandwidth of the chirp, and the number of ADC channels.
[0124] Clause 9: The method according to any one of the preceding clauses, wherein the machine learning model includes a detection head and a segmentation head.
[0125] Clause 10: A system, comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: receiving analog-to-digital converter (ADC) raw data from a Radar sensor of a vehicle; performing range fast Fourier transform (FFT) and Doppler FFT on the ADC raw data; generating a two-dimensional (2D) range-Doppler (RD) heatmap tensor representing a combination of the range FFT and the Doppler FFT; and providing the 2D RD heatmap tensor to a machine learning model for detecting objects on a road network around the vehicle.
[0126] Clause 11: The system according to clause 10, wherein providing the 2D RD heatmap tensor to the machine learning model includes: providing the 2D RD heatmap tensor to a first machine learning model to obtain a 2D range-azimuth (RA) heatmap tensor; and providing the 2D RA heatmap tensor to the machine learning model for detecting objects on the road network around the vehicle.
[0127] Clause 12: The system according to clause 11, wherein the first machine learning model includes a pre-encoder, a shared feature pyramid network (FPN) encoder, and a range-angle decoder.
[0128] Clause 13: The system according to clause 11 or 12, wherein the 2D RD heatmap tensor includes the Doppler FFT representing the number of chirps in each ADC channel and the range FFT representing the number of samples per chirp, and wherein the size of the 2D RD heatmap tensor is configured based on one or more of the sensed range, range resolution, velocity resolution, bandwidth of the chirp, period of the chirp, number of samples per chirp, and sampling rate.
[0129] Clause 14: The system according to any one of clauses 11 to 13, wherein the 2D RA heatmap tensor includes the azimuth FFT representing the number of ADC channels and the range FFT representing the number of samples per chirp, and wherein the size of the 2D RA heatmap tensor is configured based on one or more of the sensed range, range resolution, angle resolution, bandwidth of the chirp, and number of ADC channels.
[0130] Clause 15: The system according to any one of clauses 11 to 14, wherein providing the 2D RA heatmap tensor to the machine learning model includes: receiving a camera image from a camera of the vehicle; fusing the camera image with the 2D RA heatmap tensor; and providing the fused data to the machine learning model.
[0131] Clause 16: The system according to clause 15, wherein the machine learning model includes a feature extraction layer, a spatial transformer, and a feature fusion layer.
[0132] Clause 17: A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations including: receiving analog-to-digital converter (ADC) raw data from a Radar sensor of a vehicle; performing range fast Fourier transform (FFT), Doppler FFT, and azimuth FFT on the ADC raw data; generating a three-dimensional (3D) range-azimuth-Doppler (RAD) matrix tensor representing a combination of the range FFT, the Doppler FFT, and the azimuth FFT; and providing the 3D RAD matrix tensor to a machine learning model for detecting objects on a road network around the vehicle.
[0133] Clause 18: The computer-readable storage medium according to Clause 17, wherein providing the 3D RAD matrix tensor to the machine learning model further includes: receiving a camera image from a camera of the vehicle; fusing the camera image with the 3D RAD matrix tensor; and providing the fused data to the machine learning model.
[0134] Clause 19: The computer-readable storage medium according to Clause 17 or 18, wherein the 3D RAD matrix tensor includes the azimuth FFT representing the number of ADC channels, the Doppler FFT representing the number of chirps in each ADC channel, and the range FFT representing the number of samples per chirp.
[0135] Clause 20: The computer-readable storage medium according to Clause 19, wherein the size of the 3D RAD matrix tensor is configured based on one or more of sensing distance, range resolution, velocity resolution, angular resolution, chirp bandwidth, chirp period, number of samples per chirp or sampling rate, and number of ADC channels.
[0136] Clause 21: A system including: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations including: receiving analog-to-digital converter (ADC) raw data from a Radar sensor of a vehicle; performing range fast Fourier transform (FFT) on the ADC raw data; generating a one-dimensional (1D) range heatmap tensor representing the range FFT; and inputting the 1D range heatmap tensor into a machine learning model for detecting objects on a road network around the vehicle.
[0137] In the foregoing description, aspects and embodiments of the present disclosure have been described with reference to numerous specific details, which may vary depending on the implementation. Accordingly, the specification and drawings are to be regarded as illustrative rather than in a limiting sense. The sole and exclusive indication of the scope of the invention, and what the applicant desires to be the scope of the invention, is the literal and equivalent scope of the claims that are issued from this application in the specific form of the issued claims, including any subsequent amendments. Any definition of terms expressly set forth herein for inclusion in such claims shall govern the meaning of such terms as used in the claims. Additionally, when the term "further comprises" is used in the foregoing specification or the appended claims, the text following such phrase may be additional steps or entities, or sub-steps / sub-entities of the previously described steps or entities.
Claims
1. A method, comprising: Receiving, by using at least one processor, raw analog-to-digital converter data (ADC raw data) of a radio detection and ranging sensor (Radar sensor) of a vehicle; Performing, by using the at least one processor, at least one of a range fast Fourier transform (range FFT), a Doppler FFT, and an azimuth FFT on the ADC raw data; Generating, by using the at least one processor: (a) A one-dimensional range heatmap tensor (1D range heatmap tensor) representing the range FFT, (b) A two-dimensional range-Doppler heatmap tensor (2D RD heatmap tensor) representing a combination of the range FFT and the Doppler FFT, (c) A two-dimensional range-azimuth heatmap tensor (2D RA heatmap tensor) representing a combination of the range FFT and the azimuth FFT, or (d) A three-dimensional range-azimuth-Doppler matrix tensor (3D RAD matrix tensor) representing a combination of the range FFT, the Doppler FFT, and the azimuth FFT; And Inputting, by using the at least one processor, at least one of the 1D range heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, and the 3D RAD matrix tensor into a machine learning model for detecting objects on a road network around the vehicle.
2. The method according to claim 1, wherein, Inputting the 1D range heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, or the 3D RAD matrix tensor into the machine learning model further comprises: Receiving a camera image from a camera of the vehicle; Fusing the camera image with the 1D range heatmap tensor, the 2D RD heatmap tensor, the 2D RA heatmap tensor, or the 3D RAD matrix tensor; and Inputting the fused data into the machine learning model.
3. The method according to claim 1 or 2, wherein The 3D RAD matrix tensor includes the azimuth FFT representing the number of ADC channels, the Doppler FFT representing the number of chirps in each ADC channel, and the range FFT representing the number of samples per chirp.
4. The method according to claim 3, wherein The size of the 3D RAD matrix tensor is configured based on one or more of a sensing range, a range resolution, a velocity resolution, an angle resolution, a chirp bandwidth, a chirp period, the number of samples per chirp or a sampling rate, and the number of ADC channels.
5. The method according to any one of the preceding claims, wherein, The 2D RD heatmap tensor includes the Doppler FFT representing the number of chirps in each ADC channel and the range FFT representing the number of samples per chirp.
6. The method according to claim 5, wherein The size of the 2D RD heatmap tensor is configured based on one or more of a sensing range, a range resolution, a velocity resolution, a chirp bandwidth, a chirp period, the number of samples per chirp, and a sampling rate.
7. The method according to any one of the preceding claims, wherein, The 2D RA heatmap tensor includes the azimuth FFT representing the number of ADC channels and the range FFT representing the number of samples per chirp.
8. The method according to claim 7, wherein The size of the 2D RA heatmap tensor is configured based on one or more of the sensing distance, distance resolution, angular resolution, bandwidth of the chirp, and the number of ADC channels.
9. The method according to any one of the preceding claims, wherein, The machine learning model includes a detection head and a segmentation head.
10. A system, comprising: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: receiving analog-to-digital converter raw data (ADC raw data) of a radio detection and ranging sensor (Radar sensor) of a vehicle; performing a distance fast Fourier transform (distance FFT) and a Doppler FFT on the ADC raw data; generating a two-dimensional range-Doppler heatmap tensor (2D RD heatmap tensor) representing a combination of the distance FFT and the Doppler FFT; and inputting the 2D RD heatmap tensor into a machine learning model for detecting objects on a road network around the vehicle.
11. The system according to claim 10, wherein, Inputting the 2D RD heatmap tensor into the machine learning model includes: inputting the 2D RD heatmap tensor into a first machine learning model to obtain a two-dimensional range-azimuth heatmap tensor (2D RA heatmap tensor); and inputting the 2D RA heatmap tensor into the machine learning model for detecting objects on a road network around the vehicle.
12. The system according to claim 11, wherein, The first machine learning model includes a pre-encoder, a shared feature pyramid network encoder (FPN encoder), and a range-angle decoder.
13. The system according to claim 11 or 12, wherein, The 2D RD heatmap tensor includes the Doppler FFT representing the number of chirps in each ADC channel and the distance FFT representing the number of samples per chirp, wherein the size of the 2D RD heatmap tensor is configured based on one or more of the sensing distance, distance resolution, velocity resolution, bandwidth of the chirp, period of the chirp, number of samples per chirp, and sampling rate.
14. The system according to any one of claims 11 to 13, wherein, The 2D RA heatmap tensor includes the azimuth FFT representing the number of ADC channels and the distance FFT representing the number of samples per chirp, wherein the size of the 2D RA heatmap tensor is configured based on one or more of the sensing distance, distance resolution, angular resolution, bandwidth of the chirp, and the number of ADC channels.
15. The system according to any one of claims 11 to 14, wherein, Inputting the 2D RA heatmap tensor into the machine learning model includes: receiving a camera image from a camera of the vehicle; fusing the camera image with the 2D RA heatmap tensor; and inputting the fused data into the machine learning model.
16. The system according to claim 15, wherein, The machine learning model includes a feature extraction layer, a spatial transformer, and a feature fusion layer.
17. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations including: receiving analog-to-digital converter raw data (ADC raw data) of a radio detection and ranging sensor (Radar sensor) of a vehicle; Perform range fast Fourier transform (range FFT), Doppler FFT, and azimuth FFT on the ADC raw data; Generate a three-dimensional range-azimuth-Doppler matrix tensor (3D RAD matrix tensor) representing the combination of the range FFT, the Doppler FFT, and the azimuth FFT; And Input the 3D RAD matrix tensor into a machine learning model for detecting objects on the road network around the vehicle.
18. The computer-readable storage medium according to claim 17, wherein, Inputting the 3D RAD matrix tensor into the machine learning model further includes: Receiving a camera image from a camera of the vehicle; Fusing the camera image with the 3D RAD matrix tensor; and Inputting the fused data into the machine learning model.
19. The computer-readable storage medium according to claim 17 or 18, wherein, The 3D RAD matrix tensor includes the azimuth FFT representing the number of ADC channels, the Doppler FFT representing the number of chirps in each ADC channel, and the range FFT representing the number of samples per chirp.
20. The computer-readable storage medium according to claim 19, wherein, The size of the 3D RAD matrix tensor is configured based on one or more of the sensing distance, range resolution, velocity resolution, angular resolution, chirp bandwidth, chirp period, number of samples per chirp or sampling rate, and number of ADC channels.
21. A system, comprising: At least one processor; And A memory storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations, the operations including: Receiving analog-to-digital converter raw data (ADC raw data) from a radio detection and ranging sensor (Radar sensor) of a vehicle; Performing range fast Fourier transform (range FFT) on the ADC raw data; Generating a one-dimensional range heatmap tensor (1D range heatmap tensor) representing the range FFT; and Inputting the 1D range heatmap tensor into a machine learning model for detecting objects on the road network around the vehicle.
Citation Information
Cited By
Semantic understanding method and device based on deep learning and medium
CN122414361A