Emergency Response Vehicle Detection for Autonomous Driving Applications
The system uses microphones and deep learning to accurately detect emergency response vehicles, overcoming occlusions and ensuring safe vehicle interactions by processing audio signals and employing acoustic triangulation.
Patent Information
- Application Number
- JP2021102113
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-18
- Filing Date
- 2021-06-21
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2041-06-21
AI Technical Summary
Conventional systems struggle to accurately and timely identify emergency response vehicles due to environmental occlusions, which can lead to unsafe interactions and violations of local rules.
The system uses microphones to capture audio signals, processes them with background noise suppression and beamforming, and employs a deep neural network (DNN) to identify emergency response vehicles by analyzing Mel frequency coefficients and audio patterns, enhancing detection accuracy through acoustic triangulation and deep learning.
Enables early identification of emergency response vehicles despite occlusions, allowing proactive and safe maneuvering to comply with local rules, thereby improving safety and compliance.
Smart Images

Figure 0007740911000001 
Figure 0007740911000002 
Figure 0007740911000003
Abstract
Description
[Background technology]
[0001] Designing a system that can autonomously and safely operate a vehicle without supervision is extremely challenging. An autonomous vehicle should at least be capable of performing as the functional equivalent of an attentive driver—using perception and action systems with a remarkable ability to identify and react to moving and static obstacles in complex environments—to avoid collisions with other objects or structures along the vehicle's path. Additionally, to operate fully autonomously, a vehicle should be capable of obeying the rules or conventions of the road, including obeying rules associated with emergency response vehicles. For example, depending on geographic location, different rules or conventions may be in place, such as pulling over and / or stopping when an emergency response vehicle is detected.
[0002] Some conventional systems attempt to identify emergency response vehicles using the ego vehicle's perception. For example, various sensor types—e.g., LiDAR, RADAR, cameras, etc.—may be used to detect emergency response vehicles when they are perceptible by the ego vehicle. However, these systems may have difficulty identifying emergency response vehicles—or at least identifying them with sufficient time to react—due to numerous occlusions in the environment. For example, if an emergency response vehicle is approaching an intersection blocked by a building or other structure, the ego vehicle may not notice the emergency response vehicle until it also enters the intersection. At this point, it may be too late for the ego vehicle to perform a maneuver that complies with local rules, thereby reducing the overall safety of the situation and / or impeding the movement of the emergency response vehicle. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] U.S. Patent Application No. 16 / 366,506 [Patent Document 2] U.S. Patent Application No. 16 / 101,232 Summary of the Invention [Means for solving the problem]
[0004] Embodiments of the present disclosure relate to emergency response vehicle detection for autonomous driving applications. For example, and without limitation, systems and methods are disclosed that detect and classify the following alerts from audio captured using microphones on an autonomous or semi-autonomous vehicle to identify the direction, location, and / or type of emergency response vehicle in the environment: sounds produced by sirens, alarms, horns, and other patterns emitted by emergency response vehicles. For example, multiple microphones—or a microphone array—can be placed on the vehicle and used to generate audio signals corresponding to sounds in the environment. These audio signals can be processed, for example, using background noise suppression, beamforming, and / or other preprocessing operations, to determine the location and / or direction of the emergency response vehicle (e.g., using triangulation). To further enhance the quality of the audio data, microphones can be mounted in various locations around the vehicle (e.g., front, rear, left, right, interior, etc.) and / or can include physical structures that aid in audio quality—e.g., windscreens to avoid wind forces on the microphones.
[0005] To identify siren types—and consequently, corresponding emergency response vehicle types—the audio signal can be transformed into the frequency domain by extracting Mel frequency coefficients to generate a Mel spectrogram. The Mel spectrogram can be processed using a deep neural network (DNN)—e.g., a convolutional recurrent neural network (CRNN)—to output confidence or probability that the audio data represents various alert types used by emergency response vehicles. Due to the spatial and temporal nature of alerts, the use of convolutional layers for feature extraction and recurrent layers for temporal feature identification enhances the accuracy of the DNN. For example, the DNN can operate on a window of audio data while preserving state—e.g., using gated recurrent units (GRUs)—to enable a more lightweight and accurate DNN that accounts for the spatial and temporal audio profile of the siren. Additionally, attention can be applied to the output of the temporal feature extraction layer to determine the probability of each type of emergency personnel alert within the window as a whole—e.g., to give higher weight to detected alert patterns regardless of their temporal location within the window. To train the DNN for accuracy across varying driving conditions—e.g., rain, wind, traffic, inside a tunnel, outside air, etc.—audio data from varying alert types can be augmented using time stretching, time shifting, pitch shifting, dynamic range compression, noise expansion at different signal-to-noise ratios (SNRs), etc.—to generate a more robust training dataset that accounts for echo, reverberation, attenuation, Doppler effect, and / or other effects of environmental physics. As such, once deployed, the DNN can accurately predict emergency personnel alert types across a variety of operating conditions.
[0006] Finally, the use of location, heading, and alert type may enable a vehicle to identify emergency response vehicles and make planning and / or control decisions in response according to local rules or practices. Additionally, by using audio (and, in embodiments, perception) rather than perception alone, emergency response vehicles may be identified earlier despite occlusions, thereby enabling the vehicle to make proactive planning decisions that contribute to the overall safety of the situation.
[0007] The present systems and methods for emergency response vehicle detection for autonomous driving applications are described in detail below with reference to the accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1A] 1 illustrates a data flow diagram of a process for emergency response vehicle detection according to an embodiment of the present disclosure. [Figure 1B] FIG. 1 is a diagram of an example architecture of a deep neural network, according to an embodiment of the present disclosure. [Figure 2] FIG. 1 illustrates an exemplary placement of a microphone array in a vehicle, according to an embodiment of the present disclosure. [Figure 3] 1 is a flow diagram of a method for emergency response vehicle detection according to an embodiment of the present disclosure. [Figure 4] 1 is a flow diagram of a method for emergency response vehicle detection according to an embodiment of the present disclosure. [Figure 5A] 1 is an illustration of an exemplary autonomous vehicle, according to some embodiments of the present disclosure. [Figure 5B] 5B is an illustration of camera positions and fields of view for the example autonomous vehicle of FIG. 5A, according to some embodiments of the present disclosure. [Figure 5C] FIG. 5B is a block diagram of an example system architecture of the example autonomous vehicle of FIG. 5A, in accordance with some embodiments of the present disclosure. [Figure 5D] FIG. 5B is a system diagram of communication between a cloud-based server and the example autonomous vehicle of FIG. 5A, according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a block diagram of an exemplary computing device suitable for use in implementing some embodiments of the present disclosure. [Figure 7] FIG. 1 is a block diagram of an exemplary data center suitable for use in implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] Systems and methods related to emergency response vehicle detection for autonomous driving applications are disclosed. The present disclosure may be described with reference to an exemplary autonomous vehicle 500 (an example of which is described herein with reference to FIGS. 5A-5D and alternatively referred to herein as “vehicle 500” or “ego vehicle 500”), but this is not intended to be limiting. For example, the systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., one or more advanced driver assistance systems (ADAS)), robots, warehouse vehicles, off-road vehicles, airships, boats, and / or other vehicle types. Additionally, the present disclosure may be described with reference to autonomous driving, but this is not intended to be limiting. For example, the systems and methods described herein may be used in robotics, aviation systems, marine systems (e.g., emergency vessel identification), simulation environments (e.g., emergency response vehicle detection of a virtual vehicle in a virtual simulation environment), and / or other technical fields.
[0010] Referring to FIG. 1A, FIG. 1A is a data flow diagram of a process 100 for emergency response vehicle detection, according to an embodiment of the present disclosure. It should be understood that this and other configurations described herein are provided by way of example only. Other configurations and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those illustrated, and some elements may be omitted entirely. Furthermore, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in memory.
[0011] The process 100 includes one or more microphones 102 (which may be similar to microphone 596 in FIGS. 5A and 5C ) that generate audio data 104. For example, any number of microphones 102 may be used and mounted or otherwise arranged in a vehicle 500 in any arrangement. By way of non-limiting example, and with reference to FIG. 2 , various microphone arrays 202A-202D may be mounted in a vehicle 200 (which may, in an embodiment, be similar to vehicle 500 in FIGS. 5A-5D ). The microphone arrays 202A-202D may each include multiple microphones—e.g., two, three, four, etc.—and each microphone may include a unidirectional, omnidirectional, or other type of microphone. Different configurations of microphone arrays 202 may be used. For example, a rectangular arrangement of four microphones in each array may be used, including a microphone located at each vertex of a square and soldered onto a printed circuit board (PCB). As another example, a circular array of seven microphones may be used, with six microphones in the circle and one in the center.
[0012] Various locations for the microphone array 202 may be used, depending on the embodiment. With the vehicle 200 acting as a sound barrier, the microphone array 202 may be positioned so that the vehicle 200 has minimal impact on occlusion sirens. As such, by including a microphone array 202 on each side (front, rear, left, right) of the vehicle 200, the vehicle 200 may not act as a barrier to at least one of the microphone arrays 202—thereby resulting in a higher quality audio signal and therefore more accurate predictions. For example, four microphone arrays 202A-202D may be used, one on the front, rear, left, and right of the vehicle 200. As another example, an additional, fifth microphone array 202E may be located on the top of the vehicle. In one embodiment, only a single microphone array—e.g., microphone array 202E—may be used on the top of the vehicle 200, which may result in more exposure to wind, rain, dust, etc. In other instances, there may be more or fewer microphone arrays 202 used in different locations of vehicle 200. For example, in some embodiments, microphones and / or microphone arrays 202 may be located inside vehicle 500 such that the resulting audio signal generated from within the cabin of vehicle 500 may be compared to known or learned audio signals of emergency response vehicle warnings or sirens within the cabin of vehicle 500.
[0013] The microphone array 202 may be positioned in different locations on the vehicle 200 (e.g., corresponding to different portions of the vehicle 200). For example, on the left or right side of the vehicle 200, the microphone array 202 may be positioned under a side mirror or under a door handle, and at the front or rear of the vehicle 200; the microphone array 202 may be positioned near the rear license plate, near a rearview camera, near the front license plate, above or below the windshield, and / or in another location. In some embodiments, positioning the microphones in these locations may reduce exposure of the microphone array 202 to dust, water, and / or wind during operation. For example, the location of a rearview camera is generally located within the cavity where the license plate is located, which may prevent water and dust from accumulating on the microphone array 202. Similarly, being under a rearview mirror or door handle may provide protection against dust, wind, and / or water.
[0014] Each microphone array 202 may be disposed within a housing, which may be disposed in vehicle 200. As such, the housing will be disposed external to vehicle 200 and therefore may be designed to resist dust and moisture ingress to reduce or eliminate corrosion or degradation of the microphones in each microphone array 202. In some embodiments, to resist dust and moisture, the thin film or fabric covering of each housing may be constructed of a material that does not significantly attenuate audio signals while simultaneously providing adequate protection against rain, snow, dirt, and dust, such as a hydrophobic spray foam, a moisture-resistant and dust-resistant composite synthetic fabric, etc. In some embodiments, to avoid wind forces on microphone array 202, microphone array 202 may each include one or more windscreens.
[0015] In some embodiments, the microphone may include an automotive-grade digital micro-electro-mechanical system (MEMS) microphone connected to an automotive audio bus (A2B) transceiver device using pulse density modulation (PDM)—e.g., to simplify the microphone topology in vehicle 200. For example, this configuration may enable capture of audio signals from one or more microphones mounted on the outside of vehicle 200 using low-cost and lightweight twisted pair (TP) cabling. In one such example, the siren frequency from an emergency response vehicle at any point within the radius of vehicle 200 may be represented by an oversampled 1-bit PDM audio stream connected to each A2B transceiver device in a daisy-chained configuration. The audio data 104 may be converted to multi-channel time-division multiplexed (TDM) data by a secondary node A2B transceiver before being passed or forwarded upstream to its nearest A2B node in the daisy chain (e.g., in the direction of the electronic control unit (ECU) containing the A2B primary node device). For example, the A2B secondary node transceiver closest to the system may collect TDM data from each microphone (e.g., each microphone from each microphone array 202) before forwarding the TDM data to the A2B primary node. The A2B primary node may forward the audio data 104 to another component of the system, such as a system on chip (SoC)—e.g., SoC 504 in FIG. 5C—via a TDM audio port. Control channel commands may be sent to the A2B node via an inter-integrated circuit (I2C) interface between the SoC and the A2B transceiver using a TP cable.Each secondary node may include an A2B transceiver and one or more PDM microphones that may be remotely powered in such a way that sufficient current can be supplied to all connected nodes using the same TP cable.
[0016] In an embodiment, the audio data 104 may undergo pre-processing using the pre-processor 106. For example, one or more background noise suppression algorithms may be used to suppress ambient noise such as wind, vehicle noise, road noise, and / or the like. For example, one or more beamforming algorithms may be implemented to perform background noise suppression of the audio data 104 generated by each microphone array 202. As used herein, audio data 104 may refer to raw audio data and / or pre-processed audio data.
[0017] In some embodiments, the audio data 104—e.g., after preprocessing—may be used by the position-determining device 108 to determine the output 110. For example, the position-determining device 108 may use one or more (e.g., passive) acoustic location algorithms—e.g., acoustic triangulation—to determine the location 112 and / or heading 114 of the emergency response vehicle. The position-determining device 108 may use acoustic triangulation to analyze the audio data 104 from each of multiple—e.g., three or more—microphone arrays 202 to determine the distance (e.g., from the vehicle 200) and / or direction (e.g., a range of angles defining an area of the environment in which the emergency response vehicle is located) to determine the location 112 of the emergency response vehicle. For example, acoustic triangulation may be used to determine the estimated distance of the emergency response vehicle from the vehicle 200 and the estimated source direction of an emergency response vehicle alert. The more microphones or microphone arrays 202 used, the more accurate the triangulation may be. However, the use of four microphone arrays 202 (e.g., 202A-202D in FIG. 2 ) may allow for accuracy within 30 degrees, which may be suitable for locating and reacting to an emergency response vehicle. In other embodiments, more or fewer microphone arrays 202 may be used, and the accuracy may vary as a result (e.g., 40 degrees, 25 degrees, up to 10 degrees, etc.). As one non-limiting example, depending on the geometry and general layout of the road network, accuracy within 30 degrees—in addition to a distance measure—may allow the vehicle 500 to identify which road the emergency response vehicle is on with sufficient accuracy to respond in accordance with local rules and customs. As such, when approaching a four-way intersection, the estimated location and distance may allow the vehicle 500 to determine that the emergency response vehicle is on a road entering the intersection from the left, and the heading 114—described in more detail herein—may be used to determine the emergency response vehicle's direction of travel on the road.As such, if the emergency response vehicle is traveling away from the intersection, vehicle 500 may decide to proceed through the intersection without accounting for the emergency response vehicle, but if the emergency response vehicle is traveling toward the intersection, vehicle 500 may decide to pull over and stop until the emergency response vehicle has cleared the intersection (in instances where pulling over and stopping is a local rule or practice). As a result, in some embodiments, knowledge of the road layout (e.g., determined using GNSS maps, high definition (HD) maps, vehicle perception, etc.) may additionally be used to determine position 112 and / or heading 114.
[0018] The direction of travel 114 may be determined by the position-determining device 108 by tracking the emergency response vehicle's position 112 over time, e.g., over multiple frames or time steps. For example, changes (e.g., increases or decreases) in the sound pressure, particle velocity, sound or audio frequency, and / or other physical quantities of the sound field as represented by the audio data 104 may indicate that the emergency response vehicle is approaching (e.g., an increase in sound pressure) or moving away (e.g., a decrease in sound pressure).
[0019] In addition to, or instead of, the location 112 and the heading 114, the process 100 may include determining the emergency vehicle warning type 122 using deep learning. For example, the audio data 104—e.g., before and / or after preprocessing by the preprocessor 106—may be analyzed by a spectrogram generator 116 to generate a spectrogram 118. In some embodiments, a spectrogram 118 may be generated for each microphone array 202, and N instances of the DNN 120 (where N corresponds to the number of microphone arrays 202) may be used to calculate N different outputs corresponding to the siren type 122. In other embodiments, the N input channels corresponding to the spectrograms 118 from each microphone array 202 may be input to the same DNN 120 to calculate the siren type 122. In further embodiments, the N signals from each microphone array 202 may be preprocessed and / or augmented to generate a single spectrogram 118, which may be used as input for the DNN 120.
[0020] To generate spectrograms 118, audio data 104 from each microphone array 202 may be converted—by spectrogram generator 116—to the frequency domain by extracting Mel frequency coefficients. The Mel frequency coefficients may be used to generate spectrograms 118, which may be used as input to one or more deep neural networks (DNNs) 120 to calculate outputs (e.g., confidences, probabilities, etc.) indicative of siren types 122. For example, audio data 104—e.g., a digital representation of air pressure samples over time—may be sampled in windows of some predetermined size (e.g., 1024, 2048, etc.), each time making a hop of the predetermined size (e.g., 256, 512, etc.) to sample the next window. A fast Fourier transform (FFT) may be calculated for each window to convert the data from the time domain to the frequency domain. The entire frequency spectrum may then be divided or partitioned into bins—e.g., 120 bins—and each bin may be converted to a corresponding Mel bin on the Mel scale. For each window, the magnitude of the signal may be decomposed into its components, which correspond to frequency on the Mel scale. As such, the y-axis, which corresponds to frequency, may be converted to a logarithmic scale, the color dimension, which corresponds to amplitude, may be converted to decibels to form a spectrogram, and in an embodiment, the y-axis, which corresponds to frequency, may be mapped to the Mel scale to form a Mel spectrogram. As such, spectrogram 118 may correspond to a spectrogram and / or a Mel spectrogram.
[0021] The spectrogram 118—e.g., conventional or Mel—may be provided as input to the DNN 120. In an embodiment, the DNN 120 may include, but is not limited to, a convolutional recurrent neural network (CRNN). For example, and without limitation, the DNN 120 may include linear regression, logistic regression, decision trees, support vector machines (SVMs), naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, long / short-term memory / LSTM, Hopfield, Boltzmann, deep belief, deconvolution, generative adversarial, liquid state machines, etc.), area-of-interest detection algorithms, machine learning models using computer vision algorithms, and / or other types of machine learning models.
[0022] A CRNN can be used as a DNN 120 to account for both spatial and temporal characteristics of emergency alerts—e.g., siren, alarm, or vehicle horn patterns. For example, each alert may be generated by an instrument or audio-emitting device on an emergency response vehicle to have a fixed pattern. Each emergency response vehicle may similarly have different alert types depending on the current state or situation (e.g., an ambulance may have a first alert when en route to a scene without knowledge of the severity, but a second alert when driving to a hospital simply transporting a patient with minor injuries). In addition, two different alerts may have the same signal, but one signal may extend over a longer period than the other; therefore, identifying this difference is important for accurate classification of siren type 122. As such, a CRNN can account for these fixed patterns by not only looking at frequency, amplitude, and / or other sound representations, but also the amount of time each of these sound representations is detected over a time window.
[0023] Referring to FIG. 1B, an exemplary architecture of a DNN 120A (e.g., a CRNN) is shown. T may correspond to time, which may be the time window to which the spectrogram corresponds. In some embodiments, T may be 250 milliseconds (ms), 500 ms, 800 ms, or another time window. F may correspond to frequency. The DNN 120A may include a series of convolutions followed by a series of RNNs. Convolutional layers (or feature detector layers) may generally be used to learn unique spatial features of the spectrogram 118 that are useful for classification. In the architecture of FIG. 1B, the convolutions are replaced with gated linear units (GLUs) 126—e.g., GLUs 126A and 126B—to increase accuracy and enable faster convergence during training with smaller training datasets. The outputs of GLUs 126A and 126B may be augmented and applied to max pooling layer 128, convolutional layer 130, and another max pooling layer 132. The output of the max pooling layer 132 may be applied to one or more RNNs or stateful layers of the DNN 120A. For example, a stateful RNN can help users identify longer warning sequences. The DNN 120A can use gated RNNs, such as gated recurrent units 134A and 134B, to increase accuracy with respect to sequences (e.g., warnings generally follow a predetermined pattern of sounds) while remaining lightweight and fast to converge during training using smaller training datasets. Additionally, to optimize latency and accuracy, the DNN 120A can operate on a window of audio data (as represented by the Mel spectrogram) for each inference, preserving the state of the GRU 134 so that subsequent inferences can continue on the audio sequence. The output of the RNN, such as the GRU 134, can be augmented and applied to one or more attention layers 140. For example, the output of GRU 134 may be applied to a dense layer with a sigmoid activation function and a dense layer with a SoftMax activation function, and the outputs of both dense layers may undergo a weighted average operation 142 to produce a final output of confidence or probability indicative of siren type 122.The attention layer 140 can help determine where within the time window defined by T the siren will begin. Thus, the attention layer 140 is used to determine the probability of each type of alert within the time window as a whole, giving a higher weight to a siren, alarm, horn, or other produced sound pattern regardless of its temporal location within the time window. The output of independent probabilities of different alert types within the time window allows multiple alert types 122 to be detected at any given time. For example, as opposed to using a confidence level for some alerts equal to 1, the probability of each alert type may be calculated without regard to other alert type probabilities or confidence levels. However, in some embodiments, confidence levels may be used.
[0024] As a result, the GLU 126 in the CRNN 100A can help the DNN learn the best features of different types of alerts quickly and without requiring a large training dataset. Additionally, the stateful GRU 134, coupled with the attention layer 140, helps detect alerts within a time window while preserving long-term temporal dependencies. This allows the DNN 120 to run faster, as opposed to sequentially detecting sirens for each time frame—for example, without a time component. Additionally, by dividing the inference into smaller time windows, the Doppler effect estimation can be more accurate. Furthermore, the use of the attention layer 140 helps the DNN 120 detect alerts faster, with no significant contribution to latency.
[0025] To train the DNN 120—e.g., the CRNN 100A—a training dataset may be generated. However, due to the difficulty of generating a real-world training dataset that includes sufficient variation in alert types and their transformations due to the physical properties of the environment, data augmentation may be used to train the DNN 120. For example, capturing real-world audio in which an alert is present is difficult, but capturing this same audio outdoors, in a tunnel, under different weather conditions, and / or under other different circumstances is even more difficult. Alert patterns may undergo an extensive set of transformations, such as Doppler, attenuation, echo, reverberation, and / or the like, before being captured by the microphone 102. In order for the DNN 120 to accurately predict the alert type 122, the DNN 120 may be trained to identify the alert type 122 after these transformations have occurred.
[0026] As such, to generate a robust training set, a dataset of audio data including warnings (such as, for example and without limitation, sirens, horns, alarms, or other produced sound patterns emitted by emergency response vehicles) may be generated—e.g., from a real-world gathering using a data collection vehicle, from audio tracks of the warnings, from videos containing audio associated with the warnings, and / or the like. Instances of training audio data may be tagged with semantic or class labels corresponding to different warning types 122 represented therein. Instances of training audio data may then be subjected to one or more transformations, such as time stretching, time shifting, pitch shifting, dynamic range compression at different signal-to-noise (SNR) ratios, noise compression at different SNR ratios, and / or other transformation types. For example, a single instance may be used to generate any number of additional instances using different combinations of transforms. As such, the training dataset may be expanded to generate an updated training set that includes some multiple—e.g., 25x, 50x, etc.—of the original training dataset size.
[0027] 1A , alert type 122 may include a type of emergency response vehicle—e.g., two or more alert types 122 may include the same name, such as, for example, “fire engine,” “police car,” or “ambulance.” In other examples, alert type 122 may include different alert types that do not have an association to a type of emergency response vehicle—e.g., “wail,” “yelp,” “warble,” “air horn,” “piercer,” “whoop,” “howler,” “priority,” “two-tone,” “rumbler,” etc. In some examples, a combination of the two may be used as alert type 122—e.g., “police car:piercer” or “ambulance:wail.” As such, depending on the output that DNN 120 is trained to detect, the probability or confidence output by DNN 120 may correspond to the alert type 122 of the emergency response vehicle name, the alert name, or a combination thereof.
[0028] In some embodiments, in addition to identifying the alert type 122, heading 114, and / or location 112, some or all of this information may be fused with additional information types. For example, in the case of perception from other sensors of the vehicle 500 (e.g., LiDAR, RADAR, cameras, etc., as described with respect to FIGS. 5A-5C ), the results from the process 100 may be combined with the sensory output of the vehicle 500 for redundancy or fusion to increase the robustness and accuracy of the results. As such, if an emergency response vehicle is identified and classified using object detection, for example, this sensory output may be combined with the location 112, heading 114, alert type 122 (or emergency response vehicle type as determined therefrom) to update or verify the prediction. Furthermore, in some embodiments, vehicle-to-vehicle communication may include additional sources of information to increase the robustness or accuracy of the results. For example, other vehicles in the environment may share information with the vehicle 500 regarding the detection of emergency response vehicles (e.g., location, heading, type, etc.).
[0029] The determined alert type 122—which may include one or more in any instance of the DNN 120—the emergency response vehicle's heading, the emergency response vehicle's location, the vehicle's 500 perception, and / or vehicle-to-vehicle communication may be used by the autonomous driving software stack (e.g., driving stack) 124 of the vehicle 500. For example, the perception layer of the driving stack 124 may use the process 100 for identifying and locating the emergency response vehicle to update the world model (e.g., using a world model manager) to localize the emergency response vehicle to the world model. The planning layer of the driving stack 124 may use the location 112, heading 114, and / or alert type to determine a route or path plan that describes the emergency response vehicle—e.g., slow down, pull over, stop, and / or perform another action. The control layer of the driving stack 124 may then use the route or path plan to control the vehicle 500 along the path. In some examples, the identification of the emergency response vehicle may trigger a remote operation request. As one non-limiting example, remote operations may be requested and performed similarly to those described in U.S. Patent Application Publication No. 2013 / 0129994, the entire contents of which are incorporated herein by reference. As such, the driving stack 124 may use the alert type 122 (and / or the corresponding emergency response vehicle type or emergency type as indicated by the alert type 122), the location 112, and / or the heading 114 to perform one or more actions to account for the presence of the emergency response vehicle in the environment.
[0030] Referring now to FIGS. 3-4 , each block of methods 300 and 400 described herein includes computational processes that can be implemented using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory. Methods 300 and 400 may also be implemented as computer-usable instructions stored on a computer storage medium. Methods 300 and 400 may be provided by a standalone application, a service, or a hosted service (standalone or in combination with another hosted service), or a plug-in to another product, to name a few. Additionally, methods 300 and 400 are described with reference to process 100 of FIG. 1A and autonomous vehicle 500 of FIGS. 5A-5D, by way of example. However, these methods may additionally or alternatively be performed by any one process and / or any one system, or any combination of processes and systems, including, but not limited to, those described herein.
[0031] 3, which is a flow diagram illustrating a method 300 for emergency response vehicle detection according to some embodiments of the present disclosure. The method 300 includes, at block B302, receiving audio data generated using multiple microphones of an autonomous machine. For example, the audio data 104 may be generated using a microphone 102, such as a microphone 102 of a microphone array 202.
[0032] The method 300 includes, at block B304, executing an acoustic triangulation algorithm to determine at least one of the location or heading of the emergency response vehicle. For example, the audio data 104 (before or after preprocessing by the preprocessor 106) may be used by the position-determining device 108 to determine the location 112 and / or heading 114.
[0033] Process 300 includes generating a Mel spectrogram at block B306. For example, spectrogram generator 116 can use audio data 104 to generate (Mel) spectrogram 118 (or another representation of a spectrum of one or more frequencies corresponding to one or more audio signals from the audio data).
[0034] The process 300 includes applying the first data representing the Mel spectrogram to a CRNN at block B308. For example, the data representing the (Mel) spectrogram 118 may be applied to a DNN 120 (e.g., CRNN 110A).
[0035] The process 300 includes calculating, at block B310, second data representing probabilities of a plurality of alert types using a CRNN. For example, the CRNN may calculate outputs corresponding to probabilities of a plurality of alert types 122.
[0036] The method 300 includes, at block B312, determining a type of emergency response vehicle based at least in part on the probability. For example, the alert type 122 with the highest probability (or a probability above a threshold) may be determined to be present, and the type of emergency response vehicle associated with the alert type 122 may be determined. In some embodiments, the threshold may be adjusted based on the location of the vehicle 500 and / or other information. For example, in addition to, or in lieu of, the proximity of the vehicle 500's location (as determined using GNSS, HD maps, vehicle perception, etc.) to a hospital, fire station, police station, etc., vehicle-to-vehicle communications indicating a nearby accident, SigAlert, and / or structure, forest, roadside, or other fire type may be used to lower the confidence or probability threshold due to an increased likelihood that emergency response may be in the area.
[0037] The method 300 includes, at block B314, performing one or more actions by the autonomous machine based at least in part on the type, location, and / or heading of the emergency response vehicle. For example, the type (and / or alert type 122), location 112, and / or heading 114 of the emergency response vehicle may be used to perform one or more actions, such as to comply with local rules or practices regarding emergency response vehicles.
[0038] 4, which is a flow diagram illustrating a method 400 for emergency response vehicle detection according to some embodiments of the present disclosure. The method 400 includes, at block B402, receiving audio data generated using multiple microphones. For example, the audio data 104 may be generated using a microphone 102, such as a microphone 102 of the microphone array 202.
[0039] Process 400 includes generating a spectrogram at block B406. For example, spectrogram generator 116 may use audio data 104 to generate spectrogram 118 (or another representation of a spectrum of one or more frequencies corresponding to one or more audio signals from the audio data).
[0040] The process 400 includes applying the first data representing the spectrogram to a DNN at block B 408. For example, the data representing the spectrogram 118 may be applied to a DNN 120 (e.g., the CRNN 110A).
[0041] At block B410, the process 400 includes computing second data using one or more feature extraction layers of the DNN and based at least in part on the first data. For example, the GLU 126 may be used to compute a feature map or feature vector using data representing the spectrogram 118.
[0042] Process 400 includes, at block B412, computing third data using one or more stateful layers of the DNN and based at least in part on the second data. For example, GRU 134 may compute the output from the output of GLU 126—e.g., before or after processing by one or more additional layers, e.g., layers 128, 130, and / or 132.
[0043] The process 400 includes, at block B414, calculating fourth data representing probabilities of the plurality of alert types using one or more attention layers of the DNN and based at least in part on the third data. For example, the dense layer 140 of the DNN 120 may be used to calculate outputs indicating probabilities of the plurality of alert types 122.
[0044] The method 400 includes performing one or more actions based at least in part on the probabilities at block B416. For example, the type of emergency response vehicle (and / or alert type 122), location 112, and / or heading 114 may be used to perform one or more actions, such as to comply with local rules or practices regarding emergency response vehicles.
[0045] Exemplary Autonomous Vehicle 5A is a diagram of an example autonomous vehicle 500 according to some embodiments of the present disclosure. Autonomous vehicle 500 (alternatively referred to herein as “vehicle 500”) may include, but is not limited to, a passenger vehicle, such as a car, a truck, a bus, a first responder vehicle, a shuttle, an electric or moped, a motorcycle, a fire engine, a police vehicle, an ambulance, a boat, a construction vehicle, a submarine, a drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described in terms of levels of automation as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (Standard No. J3016-201806 published June 15, 2018, Standard No. J3016-201609 published September 30, 2016, and previous and future versions of this standard). Vehicle 500 may be capable of functioning according to one or more of levels 3 through 5 of autonomous driving. For example, vehicle 500 may be capable of conditional automation (Level 3), highly automated (Level 4), and / or fully automated (Level 5), depending on the embodiment.
[0046] The vehicle 500 may include components such as a vehicle chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components. The vehicle 500 may include a propulsion system 550, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another propulsion system type. The propulsion system 550 may be connected to a drive train of the vehicle 500, which may include a transmission, to enable propulsion of the vehicle 500. The propulsion system 550 may be controlled in response to receiving a signal from a throttle / accelerator 552.
[0047] A steering system 554, which may include a steering wheel, may be used to steer the vehicle 500 (e.g., along a desired course or route) when the propulsion system 550 is operating (e.g., when the vehicle is moving). The steering system 554 may receive signals from a steering actuator 556. A steering wheel may be optional for fully automated (Level 5) functionality.
[0048] Brake sensor system 546 may be used to operate vehicle brakes in response to receiving signals from brake actuators 548 and / or brake sensors.
[0049] A controller 536, which may include one or more system on chip (SoC) 504 (FIG. 5C) and / or a GPU, can provide signals (e.g., representations of commands) to one or more components and / or systems of the vehicle 500. For example, the controller can send signals to operate vehicle brakes via one or more brake actuators 548, to operate a steering system 554 via one or more steering actuators 556, and to operate a propulsion system 550 via one or more throttle / acceleration devices 552. The controller 536 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable rhythmic driving and / or assist a driver in operating the vehicle 500. The controllers 536 may include a first controller 536 for autonomous driving functions, a second controller 536 for functional safety functions, a third controller 536 for artificial intelligence functions (e.g., computer vision), a fourth controller 536 for infotainment functions, a fifth controller 536 for redundancy in emergency situations, and / or other controllers. In some instances, a single controller 536 may handle two or more of the foregoing functions, and two or more controllers 536 may handle a single function and / or any combination thereof.
[0050] The controller 536 may provide signals to control one or more components and / or systems of the vehicle 500 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data may be received from, for example, and without limitation, global navigation satellite system sensors 558 (e.g., global positioning system sensors), RADAR sensors 560, ultrasonic sensors 562, LIDAR sensors 564, inertial measurement unit (IMU) sensors 566 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 596, stereo cameras 568, wide-view cameras 570 (e.g., fisheye cameras), infrared cameras 572, surround cameras 574 (e.g., 360-degree cameras), long-range and / or medium-range cameras 598, speed sensors 544 (e.g., for measuring the speed of the vehicle 500), vibration sensors 542, steering sensors 540, brake sensors (e.g., as part of a brake sensor system 546), and / or other sensor types.
[0051] One or more of the controllers 536 may receive input (e.g., represented by input data) from the instrument cluster 532 of the vehicle 500 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 534, an audible annunciator, a loudspeaker, and / or other components of the vehicle 500. The output may include information such as vehicle velocity, speed, time, map data (e.g., HD map 522 of FIG. 5C ), position data (e.g., the position of the vehicle 500, such as on a map), direction, the positions of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as known by the controller 536, etc. For example, the HMI display 534 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or a driving maneuver the vehicle has performed, is performing, or will perform (e.g., changing lanes now, taking exit 34B in 3.22 km (2 miles), etc.).
[0052] Vehicle 500 further includes a network interface 524 that can communicate over one or more networks using one or more wireless antennas 526 and / or a modem. For example, network interface 524 may be capable of communication over LTE, WCDMA, UMTS, GSM, CDMA2000, etc. Wireless antenna 526 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc., and / or low power wide-area networks (LPWANs) such as LoRaWAN, SigFox, etc.
[0053] 5B is an illustration of camera positions and fields of view for the example autonomous vehicle 500 of FIG. 5A, according to some embodiments of the present disclosure. The cameras and their respective fields of view are one illustrative example and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located in different positions on the vehicle 500.
[0054] The camera type may include, but is not limited to, a digital camera adapted for use with components and / or systems of vehicle 500. The camera may be capable of operating at automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may be capable of any image capture rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some instances, the color filter array may include a red clear clear clear (RCCC) color filter array, a red clear clear blue (RCCB) color filter array, a red blue green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in an effort to increase light sensitivity.
[0055] In some instances, one or more of the cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).
[0056] One or more of the cameras may be mounted in a mounting part, such as a custom-designed (e.g., 3D printed) part, to filter out stray light and reflections from within the vehicle (e.g., reflections from the dashboard reflected in the windshield mirror) that may interfere with the camera's image data capture ability. Referring to a side mirror mounting part, the side mirror part may be custom 3D printed so that the camera mounting plate fits the shape of the side mirror. In some instances, the camera may be integrated into the side mirror. For side view cameras, the camera may also be integrated into four posts at each corner of the cabin.
[0057] A camera (e.g., a forward-facing camera) with a field of view that includes a portion of the environment in front of the vehicle 500 may be used for surround view to help identify the forward path and obstacles and, with the assistance of one or more controllers 536 and / or control SoCs, provide information essential for generating an occupancy grid and / or determining a preferred vehicle path. Forward-facing cameras may be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. Forward-facing cameras may also be used for ADAS features and systems, including other functions such as lane departure warning (LDW), autonomous cruise control (ACC), and / or traffic sign recognition.
[0058] Various cameras may be used in a forward-facing configuration, including, for example, a monocular camera platform including a complementary metal oxide semiconductor (CMOS) color imager. Another example may be a wide-view camera 570 that may be used to understand objects entering the view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). While only one wide-view camera is shown in FIG. 5B, any number of wide-view cameras 570 may be present in the vehicle 500. Additionally, a long-range camera 598 (e.g., a long-view stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. The long-range camera 598 may also be used for object detection and classification, as well as basic object tracking.
[0059] One or more stereo cameras 568 may also be included in the forward-facing configuration. The stereo camera 568 may include an integrated control unit with an extensible processing unit, which may provide programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet interface integrated on a single chip. Such a unit may be used to generate a 3D map of the vehicle's environment, including distance estimates for all points in the image. An alternative stereo camera 568 may include a compact stereo vision sensor, which may include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to objects and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning features. Other types of stereo cameras 568 may be used in addition to or instead of those described herein.
[0060] Cameras having a field of view that includes portions of the environment to the sides of the vehicle 500 (e.g., side-view cameras) may be used for surround view, providing information used to create and update the occupancy grid and generate side-impact collision warnings. For example, surround cameras 574 (e.g., four surround cameras 574 as shown in FIG. 5B ) may be positioned on the vehicle 500. The surround cameras 574 may include wide-view cameras 570, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras may be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 574 (e.g., left, right, and rear) and utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround-view camera.
[0061] A camera having a field of view that includes the portion of the environment behind the vehicle 500 (e.g., a rearview camera) may be used for parking assistance, surround view, rear collision warning, and creating and updating an occupancy grid. As described herein, a wide variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range and / or mid-range camera 598, stereo camera 568, infrared camera 572, etc.).
[0062] FIG. 5C is a block diagram of an example system architecture for the example autonomous vehicle 500 of FIG. 5A , in accordance with some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are merely illustrative. Other arrangements and elements (e.g., machines, interfaces, functions, sequences, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted altogether. Furthermore, many of the elements described herein are functional entities that may be implemented as separate or distributed components or in conjunction with other components, and in any suitable combination and location. Various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in a memory.
[0063] Each of the components, features, and systems of the vehicle 500 in FIG. 5C is shown connected via a bus 502. The bus 502 may include a controller area network (CAN) data interface (alternatively referred to as a "CAN bus"). The CAN may be a network within the vehicle 500 used to help control various features and functions of the vehicle 500, such as braking, acceleration, braking, steering, windshield wiper operation, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus may be read to determine steering angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.
[0064] Although the bus 502 is described herein as being a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet may be used in addition to or as an alternative to a CAN bus. Additionally, although a single line is used to represent the bus 502, this is not intended to be limiting. There may be any number of buses 502, which may include, for example, one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some instances, two or more buses 502 may be used to perform different functions and / or for redundancy. For example, a first bus 502 may be used for collision avoidance functions, and a second bus 502 may be used for operational control. In any instance, each bus 502 may communicate with any of the components of the vehicle 500, and two or more buses 502 may communicate with the same component. In some instances, each SoC 504, each controller 536, and / or each computer in the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 500) and may be connected to a common bus, such as a CAN bus.
[0065] Vehicle 500 may include one or more controllers 536, such as those described herein with respect to FIG. 5A. Controller 536 may be used for a variety of functions. Controller 536 may be coupled to any of a variety of other components and systems of vehicle 500 and may be used for control of vehicle 500, artificial intelligence of vehicle 500, infotainment for vehicle 500, and / or the like.
[0066] The vehicle 500 may include a system-on-chip (SoC) 504. The SoC 504 may include a CPU 506, a GPU 508, a processor 510, a cache 512, an accelerator 514, a data store 516, and / or other components and features not shown. The SoC 504 may be used to control the vehicle 500 in a variety of platforms and systems. For example, the SoC 504 may be coupled in a system (e.g., that of the vehicle 500) with an HD map 522 that can obtain map refreshes and / or updates via a network interface 524 from one or more servers (e.g., server 578 of FIG. 5D ).
[0067] The CPU 506 may include a CPU cluster or CPU complex (alternatively referred to as a "CCPLEX"). The CPU 506 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 506 may include eight cores in a coherent multiprocessor configuration. In some embodiments, the CPU 506 may include four dual-core clusters, each with its own dedicated L2 cache (e.g., a 2M L2 cache). The CPU 506 (e.g., a CCPLEX) may be configured to support simultaneous cluster operation, allowing any combination of clusters of CPUs 506 to be active at any given time.
[0068] The CPU 506 may implement power management capabilities including one or more of the following features: individual hardware blocks may be automatically clock gated when idle to conserve dynamic power; each core clock may be gated when the core is not actively executing instructions by executing a WFI / WFE instruction; each core may be independently power gated; each core cluster may be independently clock gated when all cores are clock gated or power gated; and / or each core cluster may be independently power gated when all cores are power gated. The CPU 506 may further implement an enhanced algorithm for managing power states, where allowable power states and expected wake-up times are specified and hardware / microcode determines the best power state for entering the cores, clusters, and CCPLEX. The processing cores may support simplified power state entry sequences in software with work offloaded to microcode.
[0069] The GPU 508 may include an integrated GPU (alternatively referred to herein as an "iGPU"). The GPU 508 may be programmable and efficient for parallel workloads. In some instances, the GPU 508 may use an enhanced tensor instruction set. The GPU 508 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having at least 96 KB of storage capacity) and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, the GPU 508 may include at least eight streaming microprocessors. The GPU 508 may use a compute application programming interface (API). Additionally, the GPU 508 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).
[0070] The GPU 508 may be power-optimized for best performance in automotive and embedded use cases. For example, the GPU 508 may be fabricated on FinFET (Fin field-effect transistor) chips. However, this is not intended to be limiting, and the GPU 508 may be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor may incorporate several mixed-precision processing cores partitioned into multiple blocks. For example, and without limitation, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In such an example, each processing block may be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads with a mix of computational and addressing operations. Streaming microprocessors may include independent thread scheduling capabilities to allow finer-grained synchronization and coordination among concurrent threads. Streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0071] The GPU 508 may, in some instances, include a high bandwidth memory (HBM) and / or 16GB HBM2 memory subsystem to provide up to 900GB / s of peak memory bandwidth. In some instances, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used in addition to or in place of the HBM memory.
[0072] The GPU 508 may include unified memory technology that includes access counters to enable more accurate movement of memory pages to the processors that access them most frequently, thereby improving the efficiency of storage areas shared between processors. In some instances, address translation service (ATS) support may be used to enable the GPU 508 to directly access the CPU 506 page tables. In such instances, when the GPU 508 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 506. In response, the CPU 506 may consult its page table for a virtual-to-real mapping of addresses and send the translation back to the GPU 508. As such, unified memory technology may enable a single unified virtual address space for both the CPU 506 and GPU 508 memories, thereby simplifying GPU 508 programming and porting of applications to the GPU 508.
[0073] Additionally, GPU 508 may include access counters that can record the frequency of GPU 508's accesses to the memory of other processors. The access counters can help ensure that memory pages are moved to the physical memory of the processors that are accessing the pages most frequently.
[0074] The SoC 504 may include any number of caches 512, including those described herein. For example, the cache 512 may include an L3 cache available to both the CPU 506 and the GPU 508 (e.g., connected to both the CPU 506 and the GPU 508). The cache 512 may include a write-back cache that can record line state, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include 4 MB or more, depending on the implementation, although smaller cache sizes may also be used.
[0075] The SoC 504 may include an arithmetic logic unit (ALU) that may be utilized in performing processing for any of various tasks or operations (e.g., processing DNNs) of the vehicle 500. Additionally, the SoC 504 may include a floating point unit (FPU) (or other math co-processor or math co-processor type) for performing mathematical operations within the system. For example, the SoC 104 may include one or more FPUs integrated as execution units within the CPU 506 and / or GPU 508.
[0076] The SoC 504 may include one or more accelerators 514 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the SoC 504 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 508 and to offload some of the GPU 508's tasks (e.g., to free up more cycles for the GPU 508 to perform other tasks). As an example, the accelerator 514 may be used for target workloads that are sufficiently stable to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs), etc.). As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and Faster RCNNs (e.g., as used for object detection).
[0077] The accelerator 514 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured and optimized to perform image processing functions (e.g., CNN, RCNN, etc.). The DLA may also be optimized for a specific set of neural network types and floating-point operations, as well as inference. The DLA design can provide more performance per millimeter than a general-purpose GPU, significantly exceeding the performance of a CPU. The TPU can perform several functions, including, for example, single-instance convolution functions, supporting INT8, INT16, and FP16 data types for both features and weights, and post-processor functions.
[0078] The DLA can quickly and efficiently run neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, but not limited to, CNNs for object identification and detection using data from camera sensors, CNNs for distance estimation using data from camera sensors, CNNs for emergency response vehicle detection and identification using data from microphones, CNNs for face recognition and vehicle owner identification using data from camera sensors, and / or CNNs for security and / or safety related events.
[0079] The DLA can perform any function of the GPU 508, and by using an inference accelerator, for example, a designer can target either the DLA or the GPU 508 for any function. For example, a designer can focus on processing CNNs and floating-point operations on the DLA and offload other functions to the GPU 508 and / or other accelerators 514.
[0080] The accelerator 514 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA may provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0081] The RISC cores may interact with an image sensor (e.g., an image sensor in any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC cores may use any of several protocols, depending on the embodiment. In some instances, the RISC cores may execute a real-time operating system (RTOS). The RISC cores may be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores may include an instruction cache and / or tightly coupled RAM.
[0082] The DMA may enable components of the PVA to access system memory independent of the CPU 506. The DMA may support any number of features used to provide optimizations to the PVA, including, but not limited to, supporting multi-dimensional addressing and / or circular addressing. In some instances, the DMA may support up to six or more dimensions of addressing, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0083] A vector processor may be a programmable processor that can be designed to efficiently and flexibly execute computer vision algorithm programming and provide signal processing capabilities. In some instances, a PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem may act as the PVA's primary processing engine and may include a vector processing unit (VPU), an instruction cache, and / or a vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), or very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.
[0084] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some instances, each vector processor may be configured to execute independently of other vector processors. In other instances, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other instances, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. In particular, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system security.
[0085] The accelerator 514 (e.g., a hardware acceleration cluster) may include a computer vision network-on-chip and SRAM to provide high-bandwidth, low-latency SRAM for the accelerator 514. In some instances, the on-chip memory may include, for example, and without limitation, at least 4 MB of SRAM consisting of eight field-configurable memory blocks that may be accessible by both the PVA and DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and DLA can access the memory through a backbone that provides the PVA and DLA with high-speed access to the memory. The backbone may include a computer vision network-on-chip that interconnects the PVA and DLA to the memory (e.g., using the APB).
[0086] The computer vision network-on-chip may include an interface that determines, prior to the transmission of any control signals, addresses, or data, that both the PVA and DLA provide ready and valid signals. Such an interface may provide separate phases and separate channels for transmitting control signals, addresses, and data, as well as burst-type communication for continuous data transfer. This type of interface may conform to the ISO 26262 or IEC 61508 standards, although other standards and protocols may also be used.
[0087] In some instances, the SoC 504 may include a real-time ray tracing hardware accelerator, such as that described in U.S. Patent Application Publication No. 2007 / 0229994. The real-time ray tracing hardware accelerator may be used to quickly and efficiently determine the location and scale of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison to LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.
[0088] The accelerator 514 (e.g., a hardware accelerator cluster) has diverse applications for autonomous driving. The PVA may be a programmable vision accelerator that can be used for critical processing stages in ADAS and autonomous vehicles. The capabilities of the PVA make it well suited to algorithmic domains that require predictable processing at low power and low latency. In other words, the PVA works well for semi-dense or dense regular computations on small data sets that require predictable execution times with low latency and low power. Therefore, because the PVA is efficient at object detection and integer computation, in the context of a platform for autonomous vehicles, the PVA is designed to run classic computer vision algorithms.
[0089] For example, according to one embodiment of the present technology, PVA is used to perform computer stereo vision. A semi-global matching-based algorithm may be used in some instances, but this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., structure from motion, pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions with input from two monocular cameras.
[0090] In some instances, PVA may be used to perform dense optical flow by processing raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR. In other instances, PVA is used in time of flight depth processing, for example, by processing raw time of flight data to provide processed time of flight data.
[0091] DLA can be used to implement any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence measure for each object detection. Such a confidence value can be interpreted as a probability or as providing a relative "weight" for each detection compared to other detections. This confidence value allows the system to make further decisions regarding which detections should be considered true positives rather than false positives. For example, the system can set a confidence threshold and consider only detections above the threshold as true positives. In an automatic emergency braking (AEB) system, a false positive detection would cause the vehicle to automatically apply emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered to trigger AEB. DLA can implement a neural network that regresses the confidence value. The neural network may receive as its input at least some subset of parameters, such as bounding box dimensions, ground plane estimates obtained (e.g., from another subsystem), inertial measurement unit (IMU) sensor 566 outputs that correlate with vehicle 500 orientation, range, and 3D position estimates of objects obtained from the neural network and / or other sensors (e.g., LIDAR sensor 564 or RADAR sensor 560), and others.
[0092] The SoC 504 may include a data store 516 (e.g., memory). The data store 516 may be on-chip memory of the SoC 504 and may store neural networks to be executed by the GPU and / or DLA. In some instances, the data store 516 may have a capacity large enough to store multiple instances of the neural network for redundancy and safety. The data store 516 may comprise an L2 or L3 cache 512. References to the data store 516 may include references to memory associated with the GPU, DLA, and / or other accelerators 514, as described herein.
[0093] The SoC 504 may include one or more processors 510 (e.g., embedded processors). The processors 510 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC 504 boot sequence and may provide run-time power management services. The boot power and management processor may provide clock and voltage programming, assist with system low-power state transitions, manage the SoC 504 thermal and temperature sensors, and / or manage the SoC 504 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 504 may use the ring oscillator to detect the temperature of the CPU 506, GPU 508, and / or accelerator 514. If the temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine, place the SoC 504 in a lower power state, and / or place the vehicle 500 in a Chauffeur safe shutdown mode (e.g., bring the vehicle 500 to a safe shutdown).
[0094] The processor 510 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that allows full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In some instances, the audio processing engine is a dedicated processor core that includes a digital signal processor with dedicated RAM.
[0095] The processor 510 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0096] The processor 510 may further include a safety cluster engine that includes a processor subsystem dedicated to handling safety management for automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, the two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences between their operations.
[0097] The processor 510 may further include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.
[0098] The processor 510 may further include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.
[0099] The processor 510 may include a video image compositor, which may be a processing block (e.g., implemented in a microprocessor) that implements video post-processing functions required by the video playback application to produce the final image for the player window. The video image compositor may perform lens distortion correction on the wide-view camera 570, the surround camera 574, and / or the in-cabin surveillance camera sensor. The in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the advanced SoC, configured to identify in-cabin events and respond appropriately. The in-cabin system may perform lip reading to activate cellular service and make phone calls, dictate emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain features are available to the driver only when operating in autonomous mode and are disabled otherwise.
[0100] The video image combiner may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in the video, the noise reduction reduces the weight of information provided by adjacent frames and appropriately weights spatial information. When an image or portion of an image does not contain motion, the temporal noise reduction performed by the video image combiner can use information from previous images to reduce noise in the current image.
[0101] The video image compositor may also be configured to perform stereo rectification on the input stereo lens frames. The video image compositor may further be used for user interface compositing when the operating system desktop is in use, and the GPU 508 is not required to continuously render new surfaces. Even when the GPU 508 is powered on and actively performing 3D rendering, the video image compositor may be used to offload the GPU 508 to improve performance and responsiveness.
[0102] The SoC 504 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface for receiving video and input from a camera, and / or a video input block that may be used for camera and related pixel input functions. The SoC 504 may further include an input / output controller that may be controlled by software and that may be used to receive I / O signals that are not committed to a specific role.
[0103] The SoC 504 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management, and / or other devices. The SoC 504 may be used to process data from cameras (e.g., connected via gigabit multimedia serial links and Ethernet), sensors (e.g., LIDAR sensors 564, RADAR sensors 560, etc., which may be connected via Ethernet), data from the bus 502 (e.g., vehicle 500 speed, steering wheel position, etc.), and GNSS sensors 558 (e.g., connected via Ethernet or CAN bus). The SoC 504 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to offload routine data management tasks from the CPU 506.
[0104] The SoC 504 may be an end-to-end platform with a flexible architecture that spans levels 3-5 of automation, thereby providing a comprehensive functional safety architecture that leverages and efficiently uses computer vision and ADAS techniques for diversity and redundancy, and provides a platform for a flexible, reliable driving software stack along with deep learning tools. The SoC 504 may be faster, more reliable, and more energy- and space-efficient than conventional systems. For example, when the accelerator 514 is combined with the CPU 506, the GPU 508, and the data store 516, it can provide a fast and efficient platform for levels 3-5 of autonomous vehicles.
[0105] This technology therefore offers capabilities and functionality not achievable by conventional systems. For example, computer vision algorithms can be implemented on a central processing unit (CPU), which can be configured using a high-level programming language, such as the C programming language, to execute a wide variety of processing algorithms across a wide variety of visual data. However, CPUs often cannot meet the performance requirements of many computer vision applications, including those related to execution time and power consumption. Specifically, many CPUs cannot execute complex object detection algorithms in real time, a requirement for in-vehicle ADAS applications and practical Level 3-5 autonomous vehicles.
[0106] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein allows multiple neural networks to run simultaneously and / or serially and the results to be combined to enable Level 3-5 autonomous driving capabilities. For example, a CNN running on the DLA or dGPU (e.g., GPU520) can include text and word recognition, enabling the supercomputer to read and understand traffic signs, including signs for which the neural network was not specifically trained. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module running on the CPU complex.
[0107] As another example, multiple neural networks may be run simultaneously, as required for Level 3, 4, or 5 driving. For example, a warning sign consisting of "Caution: Flashing lights indicate icy conditions" along with a lightning flash may be interpreted independently or collectively by several neural networks. The sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" may be interpreted by a second deployed neural network that notifies the vehicle's route planning software (preferably running on a CPU complex) that icy conditions exist when the flashing light is detected. The flashing light may be identified by running a third deployed neural network over multiple frames, informing the vehicle's route planning software of the presence (or absence) of the flashing light. All three neural networks may run simultaneously, such as within the DLA and / or on the GPU 508.
[0108] In some instances, a CNN for facial recognition and vehicle owner identification can use data from the camera sensors to identify the presence of a legitimate driver and / or owner of the vehicle 500. An always-on sensor processing engine can be used to unlock the vehicle and turn on the lights when the owner approaches the driver's side door, and in security mode, to disable vehicle operation when the owner leaves the vehicle. In this way, the SoC 504 provides security against theft and / or carjacking.
[0109] In another example, a CNN for emergency response vehicle detection and identification can detect and identify emergency response vehicle sirens using data from microphone 596. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 504 uses CNNs for environmental and urban sound classification, as well as visual data classification. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative terminal velocity of emergency response vehicles (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency response vehicles specific to the local area in which the vehicle is operating, as identified by GNSS sensor 558. Thus, for example, when operating in Europe, the CNN would attempt to detect European sirens, and when in the United States, the CNN would attempt to identify only North American sirens. Once an emergency response vehicle is detected, the control program may be used to execute emergency response vehicle safety routines, such as slowing the vehicle, stopping the vehicle at the side of the road, parking the vehicle, and / or idling the vehicle, with the assistance of ultrasonic sensor 562, until the emergency response vehicle has passed.
[0110] The vehicle may include a CPU 518 (e.g., a discrete CPU or dCPU) that may be coupled to the SoC 504 via a high-speed interconnect (e.g., PCIe). The CPU 518 may include, for example, an X86 processor. The CPU 518 may be used to perform any of a variety of functions, including, for example, reconciling potentially inconsistent results between the ADAS sensors and the SoC 504 and / or monitoring the status and health of the controller 536 and / or infotainment SoC 530.
[0111] Vehicle 500 may include GPU 520 (e.g., a discrete GPU or dGPU) that may be coupled to SoC 504 via a high-speed interconnect (e.g., NVIDIA's NVLINK). GPU 520 may provide additional artificial intelligence functionality, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based on input (e.g., sensor data) from sensors of vehicle 500.
[0112] The vehicle 500 may further include a network interface 524, which may include one or more wireless antennas 526 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). The network interface 524 may be used to enable wireless connections with the cloud over the Internet (e.g., with a server 578 and / or other network devices), with other vehicles, and / or with computing devices (e.g., passenger client devices). To communicate with other vehicles, a direct link may be established between the two vehicles and / or an indirect link may be established (e.g., through a network and via the Internet). A direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide the vehicle 500 with information about vehicles in proximity to the vehicle 500 (e.g., vehicles in front of, beside, and / or behind the vehicle 500). This functionality may be part of a cooperative adaptive cruise control function of the vehicle 500.
[0113] The network interface 524 may include an SoC that provides modulation and demodulation functions and enables the controller 536 to communicate over a wireless network. The network interface 524 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion may be performed through well-known processes and / or may be performed using a superheterodyne process. In some instances, the radio frequency front end functionality may be provided by a separate chip. The network interface may include wireless functionality for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0114] The vehicle 500 may further include a data store 528, which may include off-chip (e.g., off-SoC 504) storage. The data store 528 may include one or more memory elements, including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0115] The vehicle 500 may further include a GNSS sensor 558. The GNSS sensor 558 (e.g., a GPS, an aided GPS sensor, a differential GPS (DGPS) sensor, etc.) aids in mapping, perception, occupancy grid generation, and / or route planning functions. Any number of GNSS sensors 558 may be used, including, for example, but not limited to, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.
[0116] The vehicle 500 may further include a RADAR sensor 560. The RADAR sensor 560 may be used by the vehicle 500 for long-range vehicle detection, even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some instances, the RADAR sensor 560 may use the CAN and / or bus 502 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 560), with access to Ethernet for accessing raw data. A wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor 560 may be suitable for front, rear, and side RADAR use. In some instances, a pulse-Doppler RADAR sensor is used.
[0117] The RADAR sensor 560 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, and short-range side coverage. In some instances, long-range RADAR may be used for adaptive cruise control functions. Long-range RADAR systems may provide a wide field of view achieved by two or more independent scans, such as within a 250-meter range. The RADAR sensor 560 may help distinguish between static and moving objects and may be used by ADAS systems for emergency brake assist and forward collision warning. Long-range RADAR sensors may include monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN, Ethernet, and / or FlexRay interfaces. In one example with six antennas, the center four antennas may create a focused beam pattern designed to record the vehicle 500's surroundings at high speeds with minimal interference from traffic in adjacent lanes. The other two antennas may widen the field of view, allowing for rapid detection of vehicles entering or leaving the vehicle's 500's lane.
[0118] As an example, a medium-range RADAR system may include a range of up to 560 meters (front) or 80 meters (rear) and a field of view of up to 42 degrees (front) or 550 degrees (rear). A short-range RADAR system may include, but is not limited to, a RADAR sensor designed to be mounted on either end of the rear bumper. When mounted on either end of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the vehicle.
[0119] Short-range RADAR systems may be used in ADAS systems for blind spot detection and / or lane change assist.
[0120] Vehicle 500 may further include ultrasonic sensors 562. The ultrasonic sensors 562, which may be positioned on the front, rear, and / or sides of vehicle 500, may be used for parking assistance and / or for creating and updating an occupancy grid. A variety of ultrasonic sensors 562 may be used, and different ultrasonic sensors 562 may be used for different ranges of detection (e.g., 2.5 m, 4 m). Ultrasonic sensors 562 may operate at an ASIL B functional safety level.
[0121] The vehicle 500 may include a LIDAR sensor 564. The LIDAR sensor 564 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 564 may be functional safety level ASIL B. In some instances, the vehicle 500 may include multiple (e.g., two, four, six, etc.) LIDAR sensors 564 that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0122] In some instances, the LIDAR sensor 564 may be capable of providing a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 564 may have an advertised range of approximately 500 m, for example, with an accuracy of 2 cm to 3 cm and support for a 500 Mbps Ethernet connection. In some instances, one or more non-protruding LIDAR sensors 564 may be used. In such instances, the LIDAR sensor 564 may be implemented as a small device that may be integrated into the front, rear, sides, and / or corners of the vehicle 500. In such instances, the LIDAR sensor 564 may have a range of 200 m, even for low-reflecting objects, and may provide up to a 120-degree horizontal and 35-degree vertical field of view. A front-mounted LIDAR sensor 564 may be configured for a horizontal field of view between 45 and 135 degrees.
[0123] In some instances, LIDAR technology such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the vehicle's surroundings up to approximately 200 meters. The flash LIDAR unit includes a receptor that records the laser pulse transit time and the reflected light at each pixel, which in turn corresponds to the range from the vehicle to the object. Flash LIDAR may enable a highly accurate and distortion-free image of the surroundings to be generated with every laser flash. In some instances, four flash LIDAR sensors may be deployed, one on each side of the vehicle 500. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) with no moving parts other than the blower. Flash LIDAR devices may use 5 nanosecond Class I (eye-safe) laser pulses per frame and may capture reflected laser light in the form of a 3D range point cloud and coregistered intensity data. By using flash LIDAR, and because flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 564 may be less susceptible to motion blur, vibration, and / or shock.
[0124] The vehicle may further include an IMU sensor 566. In some instances, the IMU sensor 566 may be positioned at the center of the rear axle of the vehicle 500. The IMU sensor 566 may include, for example, but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some instances, such as in a six-axis application, the IMU sensor 566 may include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 566 may include an accelerometer, a gyroscope, and a magnetometer.
[0125] In some embodiments, the IMU sensor 566 may be implemented as a miniature, high-performance GPS-Aided Inertial Navigation System (GPS / INS) that combines micro-electro-mechanical system (MEMS) inertial sensors, a highly sensitive GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. As such, in some instances, the IMU sensor 566 may enable the vehicle 500 to estimate heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 566. In some instances, the IMU sensor 566 and the GNSS sensor 558 may be combined in a single integrated unit.
[0126] The vehicle may include microphones 596 placed in and / or around the vehicle 500. The microphones 596 may be used for emergency response vehicle detection and identification, among other things.
[0127] The vehicle may further include any number of camera types, including stereo cameras 568, wide-view cameras 570, infrared cameras 572, surround cameras 574, long-range and / or mid-range cameras 598, and / or other camera types. The cameras may be used to capture image data around the entire exterior of the vehicle 500. The types of cameras used depend on the implementation and requirements of the vehicle 500, and any combination of camera types may be used to achieve the required coverage around the vehicle 500. Additionally, the number of cameras may vary depending on the implementation. For example, a vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras may support, by way of example only, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each camera is described in further detail herein with reference to FIGS. 5A and 5B.
[0128] The vehicle 500 may further include a vibration sensor 542. The vibration sensor 542 may measure vibrations of vehicle components, such as an axle. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 542 are used, the difference in vibration may be used to determine friction or slippage of the road surface (e.g., when the difference in vibration is between a powered axle and a free-spinning axle).
[0129] The vehicle 500 may include an ADAS system 538. In some instances, the ADAS system 538 may include an SoC. The ADAS system 538 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.
[0130] The ACC system may use a RADAR sensor 560, a LIDAR sensor 564, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle directly ahead of the vehicle 500 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and advises the vehicle 500 to change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.
[0131] CACC uses information from other vehicles, which may be received from other vehicles via a wireless link via network interface 524 and / or wireless antenna 526, or indirectly via a network connection (e.g., via the Internet). A direct link may be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link may be an infrastructure-to-vehicle (I2V) communication link. Generally, V2V communication concepts provide information about immediately preceding vehicles (e.g., vehicles directly ahead of vehicle 500 that are in the same lane as vehicle 500), while I2V communication concepts provide information about traffic further ahead. A CACC system may include either or both I2V and V2V information sources. Given information about vehicles ahead of vehicle 500, CACC may be more reliable, potentially allowing for smoother traffic flow and reducing road congestion.
[0132] FCW systems are designed to warn the driver of hazards so that the driver can take corrective action. FCW systems use forward-facing cameras and / or RADAR sensors 560 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs, electrically coupled to driver feedback such as displays, speakers, and / or vibration components. FCW systems can provide warnings in the form of audio, visual alerts, vibrations, and / or quick brake pulses.
[0133] An AEB system can detect an imminent forward collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. The AEB system can use a forward-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first alerts the driver to take corrective action to avoid the collision; if the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent, or at least mitigate, the effects of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or collision imminent braking.
[0134] The LDW system provides visual, audible, and / or tactile warnings, such as vibration of the steering wheel or seat, to alert the driver when the vehicle 500 crosses a lane marking. The LDW system does not activate when the driver indicates an intentional lane departure by activating a turn signal. The LDW system may use a forward-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration components.
[0135] The LKA system is a modification of the LDW system, which provides steering input or braking to correct the vehicle 500 if it begins to drift out of its lane.
[0136] The BSW system detects and warns the vehicle driver in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile warnings to indicate that merging or changing lanes is unsafe. The system can provide additional warnings when the driver uses a turn signal. The BSW system can use a rear-facing camera and / or RADAR sensor 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, e.g., a display, speaker, and / or vibration component.
[0137] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera when the vehicle 500 is backing up. Some RCTW systems include AEB to ensure vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 560 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, e.g., a display, speaker, and / or vibration components.
[0138] Because conventional ADAS systems alert the driver and allow the driver to determine whether a safety condition truly exists and act accordingly, conventional ADAS systems can be prone to producing false positives that, while not usually catastrophic, can be annoying and distracting to the driver. However, in an autonomous vehicle 500, when results conflict, the vehicle 500 itself must decide whether to listen to results from a primary computer or a secondary computer (e.g., the first controller 536 or the second controller 536). For example, in some embodiments, the ADAS system 538 may be a backup and / or secondary computer that provides perception information to a backup computer rationality module. The backup computer rationality monitor can run redundant software on hardware components to detect failures in perception and dynamic driving tasks. Output from the ADAS system 538 may be provided to a supervisory MCU. When the outputs from the primary and secondary computers conflict, the supervisory MCU must decide how to reconcile the conflict to ensure safe operation.
[0139] In some instances, the primary computer may be configured to provide a reliability score to the supervising MCU indicating the reliability of the primary computer in a selected outcome. If the reliability score exceeds a threshold, the supervising MCU may follow the primary computer's instructions regardless of whether the secondary computers provide conflicting or inconsistent results. If the reliability score does not meet the threshold, and the primary and secondary computers provide different (e.g., conflicting) results, the supervising MCU may arbitrate between the computers to determine the appropriate outcome.
[0140] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on outputs from the primary and secondary computers, conditions under which the secondary computer will provide a false alarm. Thus, the neural network in the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, when the secondary computer is a RADAR-based FCW system, the neural network in the supervisory MCU can learn when the FCW identifies a metal object that is not actually dangerous, such as a sewer grate or manhole cover, which triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network in the supervisory MCU can learn to ignore the LDW when a bicyclist or pedestrian is present and lane departure is, in fact, the safest maneuver. In embodiments including a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervising MCU may comprise and / or be included as a component of the SoC 504 .
[0141] In other instances, the ADAS system 538 may include a secondary computer that performs ADAS functions using traditional rules of computer vision. As such, the secondary computer may use classical computer vision rules (if-then), and the presence of a neural network in the supervisory MCU may improve reliability, safety, and performance. For example, diverse implementations and intentional non-identity may make the overall system more fault-tolerant, particularly to failures caused by software (or software-hardware interface) functions. For example, if a software bug or error exists in software running on the primary computer and non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that a bug in the software or hardware on the primary computer has not caused a critical error.
[0142] In some instances, the output of the ADAS system 538 can be fed to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 538 indicates a forward collision warning due to an object directly ahead, the perception block can use this information when identifying the object. In other instances, the secondary computer can have its own neural network that is trained as described herein, thus reducing the risk of false positives.
[0143] The vehicle 500 may further include an infotainment SoC 530 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system need not be an SoC and may include two or more separate components. The infotainment SoC 530 may include a combination of hardware and software that may be used to provide audio (e.g., music, personal digital assistants, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, reverse parking assist, wireless data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door opening / closing, air filter information, etc.) to the vehicle 500. For example, the infotainment SoC 530 may be a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, a car computer, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, a heads-up display (HUD), an HMI display 534, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 530 may further be used to provide information (e.g., visual and / or audible) to a user of the vehicle, such as information from an ADAS system 538, autonomous driving information such as planned vehicle maneuvers, trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0144] The infotainment SoC 530 may include GPU functionality. The infotainment SoC 530 may communicate with other devices, systems, and / or components of the vehicle 500 via the bus 502 (e.g., a CAN bus, Ethernet, etc.). In some instances, the infotainment SoC 530 may be coupled to a supervisory MCU so that the infotainment system's GPU can perform some self-drive functions in the event of a failure of the primary controller 536 (e.g., the vehicle's 500 primary and / or backup computer). In such instances, the infotainment SoC 530 may place the vehicle 500 in a Chauffeur safe shutdown mode, as described herein.
[0145] The vehicle 500 may further include an instrument cluster 532 (e.g., a digital dash, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 532 may include a controller and / or a supercomputer (e.g., a separate controller or supercomputer). The instrument cluster 532 may include a set of instruments such as a speedometer, fuel level, oil pressure, a tachometer, an odometer, turn signals, a gear shift position indicator, a seat belt warning light, a parking brake warning light, an engine malfunction light, airbag (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some instances, information may be displayed and / or shared between the infotainment SoC 530 and the instrument cluster 532. In other words, the instrument cluster 532 may be included as part of the infotainment SoC 530, or vice versa.
[0146] 5D is a system diagram of communication between the cloud-based server and the example autonomous vehicle 500 of FIG. 5A in accordance with some embodiments of the present disclosure. System 576 may include a server 578, a network 590, and a vehicle including vehicle 500. Server 578 may include multiple GPUs 584(A)-584(H) (collectively referred to herein as GPUs 584), PCIe switches 582(A)-582(H) (collectively referred to herein as PCIe switches 582), and / or CPUs 580(A)-580(B) (collectively referred to herein as CPUs 580). GPUs 584, CPUs 580, and PCIe switches may be interconnected with a high-speed interconnect, such as, but not limited to, an NVLink interface 588 developed by NVIDIA and / or a PCIe connection 586. In some instances, the GPUs 584 are connected via NVLink and / or NVSwitch SoCs, and the GPUs 584 and PCIe switches 582 are connected via PCIe interconnects. While eight GPUs 584, two CPUs 580, and two PCIe switches are illustrated, this is not intended to be limiting. Depending on the embodiment, each server 578 may include any number of GPUs 584, CPUs 580, and / or PCIe switches. For example, the servers 578 may each include 8, 16, 32, and / or more GPUs 584.
[0147] Server 578 may receive image data from vehicles over network 590 representing images showing unexpected or changed road conditions, such as recently started road construction. Server 578 may transmit neural network 592, updated neural network 592, and / or map information 594, including information about traffic and road conditions, to vehicles over network 590. Updates to map information 594 may include updates to HD map 522, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In some instances, neural network 592, updated neural network 592, and / or map information 594 may result from new training and / or experience represented in data received from any number of vehicles in the environment and / or based on training performed at a data center (e.g., using server 578 and / or other servers).
[0148] The server 578 may be used to train a machine learning model (e.g., a neural network) based on training data. The training data may be generated by a vehicle and / or generated in a simulation (e.g., using a game engine). In some instances, the training data is tagged (e.g., if the neural network benefits from supervised learning) and / or undergoes other pre-processing, while in other instances, the training data is not tagged and / or pre-processed (e.g., if the neural network does not require supervised learning). The training may be performed according to any one or more classes of machine learning techniques, including, but not limited to, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, federated learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including preliminary dictionary learning), rule-based machine learning, anomaly detection, and variations or combinations thereof. After the machine-learned model has been traced, it may be used by the vehicle (e.g., transmitted to the vehicle via network 590) and / or it may be used by server 578 to remotely monitor the vehicle.
[0149] In some instances, server 578 can receive data from vehicles and apply the data to state-of-the-art real-time neural networks for real-time intelligent inference. Server 578 can include deep learning supercomputers and / or dedicated AI computers powered by GPUs 584, such as the DGX and DGX Station machines developed by NVIDIA. However, in some instances, server 578 can include deep learning infrastructure that uses only CPU-powered data centers.
[0150] The deep learning infrastructure of server 578 may be capable of rapid real-time inference, which it may use to evaluate and verify the health of the processor, software, and / or associated hardware within vehicle 500. For example, the deep learning infrastructure may receive periodic updates from vehicle 500 (e.g., via computer vision and / or other machine learning object classification techniques), such as a sequence of images and / or objects where vehicle 500 was located within the sequence of images. The deep learning infrastructure may run its own neural network to identify objects and compare them to objects identified by vehicle 500; if the results are inconsistent and the infrastructure concludes that the AI within vehicle 500 is not functioning properly, server 578 may send a signal to vehicle 500 that infers control, notifies passengers, and commands the vehicle's 500 failsafe computer to complete a safe parking maneuver.
[0151] For inference, the server 578 may include a GPU 584 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other instances, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors may be used for inference.
[0152] Exemplary Computing Device 6 is a block diagram of an example computing device 600 suitable for use in implementing some embodiments of the present disclosure. Computing device 600 may include an interconnection system 602 that indirectly or directly couples the following devices: memory 604, one or more central processing units (CPUs) 606, one or more graphics processing units (GPUs) 608, a communication interface 610, input / output (I / O) ports 612, input / output components 614, a power supply 616, one or more presentation components 618 (e.g., displays), and one or more logic units 620. In at least one embodiment, computing device 600 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). As non-limiting examples, one or more of GPUs 608 may include one or more vGPUs, one or more of CPUs 606 may include one or more vCPUs, and / or one or more of logical units 620 may include one or more virtual logical units. As such, computing device 600 may include discrete components (e.g., an entire GPU dedicated to computing device 600), virtual components (e.g., a portion of a GPU dedicated to computing device 600), or a combination thereof.
[0153] While the various blocks in FIG. 6 are depicted as connected by lines via interconnection system 602, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 618, such as a display device, may be considered an I / O component 614 (e.g., if the display is a touch screen). As another example, CPU 606 and / or GPU 608 may include memory (e.g., memory 604 may represent a storage device in addition to the memory of GPU 608, CPU 606, and / or other components). In other words, the computing devices in FIG. 6 are merely exemplary. Categories such as “workstation,” “server,” “laptop,” “desktop,” “tablet,” “client device,” “mobile device,” “handheld device,” “gaming console,” “electronic control unit (ECU),” “virtual reality system,” and / or other device or system types are all intended to be within the scope of the computing devices in FIG. 6 and therefore will not be distinguished from one another.
[0154] Interconnect system 602 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. Interconnect system 602 may include one or more bus or link types, such as an industry standard architecture (ISA) bus, an extended industry standard architecture (EISA) bus, a video electronics standards association (VESA) bus, a peripheral component interconnect (PCI) bus, a peripheral component interconnect express (PCIe) bus, and / or another type of bus or link. In some embodiments, direct connections exist between components. As an example, CPU 606 may be directly connected to memory 604. Further, CPU 606 may be directly connected to GPU 608. When direct or point-to-point connections exist between components, interconnect system 602 may include a PCIe link to implement the connections. In these examples, a PCI bus need not be included in computing device 600.
[0155] Memory 604 may include any of a variety of computer-readable media. Computer-readable media may be any available media that can be accessed by computing device 600. Computer-readable media may include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media may comprise computer storage media and communication media.
[0156] Computer storage media may include both volatile and nonvolatile media, and / or removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 604 may store computer-readable instructions (e.g., representing programs and / or program elements), such as an operating system. Computer storage media may include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 600. As used herein, computer storage media does not include the signals themselves.
[0157] Computer storage media may embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism and include any information delivery media. The term "modulated data signal" may refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, computer storage media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Combinations of any of the above should also be included within the scope of computer-readable media.
[0158] The CPU 606 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 600 to perform one or more of the methods and / or processes described herein. The CPU 606 may include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores, each capable of simultaneously processing multiple software threads. The CPU 606 may include any type of processor, and may include different types of processors depending on the type of computing device 600 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 600, the processor may be an Advanced RISC Machines (ARM) processor implemented using Reduced Instruction Set Computing (RISC) or an x86 processor implemented using Complex Instruction Set Computing (CISC). Computing device 600 may include one or more CPUs 606 within one or more microprocessors or auxiliary coprocessors, such as computational coprocessors.
[0159] In addition to or instead of CPU 606, GPU 608 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 600 to perform one or more of the methods and / or processes described herein. One or more of GPUs 608 may be integrated GPUs (e.g., with one or more of CPUs 606 and / or one or more of GPUs 608 may be discrete GPUs. In an embodiment, one or more of GPUs 608 may be coprocessors of one or more of CPUs 606. GPU 608 may be used by computing device 600 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 608 may be used with GPGPU (General-Purpose Computing on a GPU) The GPU 608 may be used for graphics processing (GPU). The GPU 608 may include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 608 may generate pixel data for an output image in response to rendering commands (e.g., rendering commands from the CPU 606 received via a host interface). The GPU 608 may include graphics memory, e.g., display memory, for storing pixel data or any other suitable data, e.g., GPGPU data. The display memory may be included as part of the memory 604. GPU 608 may include two or more GPUs operating in parallel (e.g., via links). The links may connect the GPUs directly (e.g., using NVLINK) or may connect the GPUs via a switch (e.g., using NVSwitch). When coupled together, each GPU 608 may generate pixel data or GPGPU data for a different portion of the output or for a different output (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU may include its own memory or may share memory with other GPUs.
[0160] In addition to or instead of CPU 606 and / or GPU 608, logic unit 620 may be configured to execute at least some of the computer-readable instructions to control one or more of computing devices 600 to perform one or more of the methods and / or processes described herein. In an embodiment, CPU 606, GPU 608, and / or logic unit 620 may discretely or jointly execute any combination of methods, processes, and / or portions thereof. One or more of logic units 620 may be part of and / or integrated with one or more of CPU 606 and / or GPU 608, and / or one or more of logic units 620 may be discrete components to or otherwise external to CPU 606 and / or GPU 608. In an embodiment, one or more of logic units 620 may be a coprocessor of one or more of CPU 606 and / or GPU 608.
[0161] Examples of logic unit 620 include one or more processing cores and / or components thereof, such as a tensor core (TC), a tensor processing unit (TPU), a pixel visual core (PVC), a vision processing unit (VPU), a graphics processing cluster (GPC), a texture processing cluster (TPC), a streaming multiprocessor (SM), a tree traversal unit (TTU), an artificial intelligence accelerator (AIA), a deep learning accelerator (DLA), an arithmetic logic unit (ALU), an application specific integrated circuit (ASIC), a floating point unit (FPU), an input / output (I / O) element, a peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) element, and / or the like.
[0162] The communications interface 610 may include one or more receivers, transmitters, and / or transceivers that enable the computing device 600 to communicate with other computing devices over electronic communications networks, including wired and / or wireless communications. The communications interface 610 may include components and functionality to enable communication over any of several different networks, such as a wireless network (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), a wired network (e.g., communicating over Ethernet or InfiniBand), a low-power wide area network (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.
[0163] The I / O ports 612 may enable the computing device 600 to be logically coupled to other devices, including I / O components 614, presentation components 618, and / or other components, some of which may be built into (e.g., integrated with) the computing device 600. Exemplary I / O components 614 include a microphone, mouse, keyboard, joystick, game pad, game controller, satellite dish, scanner, printer, wireless device, etc. The I / O components 614 may provide a natural user interface (NUI) that processes air gestures, voice, or other physiological input generated by the user. In some cases, the input may be sent to an appropriate network element for further processing. The NUI may implement any combination of voice recognition, stylus recognition, facial recognition, biometric recognition, on-screen and adjacent-screen gesture recognition, air gestures, head and eye tracking, and touch recognition in connection with the display of the computing device 600 (as described in more detail below). Computing device 600 may include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, computing device 600 may include an accelerometer or gyroscope (e.g., as part of an inertia measurement unit (IMU)) to enable detection of movement. In some instances, the output of the accelerometer or gyroscope may be used by computing device 600 to render immersive augmented or virtual reality.
[0164] The power supply 616 may include a hardwired power supply, a battery power supply, or a combination thereof. The power supply 616 may provide power to the computing device 600 to enable the components of the computing device 600 to operate.
[0165] The presentation component 618 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), speakers, and / or other presentation components. The presentation component 618 can receive data from other components (e.g., GPU 608, CPU 606, etc.) and output data (e.g., as images, video, sound, etc.).
[0166] Exemplary Data Center 7 illustrates an example data center 700 that may be used in at least one embodiment of the present disclosure. The data center 700 may include a data center infrastructure layer 710, a framework layer 720, a software layer 730, and / or an application layer 740.
[0167] 7, the data center infrastructure layer 710 may include a resource orchestrator 712, grouped computational resources 714, and node computational resources (“node CRs”) 716(1) through 716(N), where “N” represents any integer, natural number. In at least one embodiment, the node CRs 716(1) through 716(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and / or cooling modules. In some embodiments, one or more of the nodes CR716(1)-716(N) may correspond to a server having one or more of the aforementioned computing resources. Additionally, in some embodiments, the nodes CR716(1)-716(N) may include one or more virtual components, such as a vGPU, a vCPU, and / or the like, and / or one or more of the nodes CR716(1)-716(N) may correspond to a virtual machine (VM).
[0168] In at least one embodiment, the grouped computing resources 714 may include separate groups of nodes CR716 housed within one or more racks (not shown), or multiple racks housed in data centers in various geographic locations (also not shown). The separate groups of nodes CR716 within the grouped computing resources 714 may include grouped computing, network, memory, or storage resources that can be configured or assigned to support one or more workloads. In at least one embodiment, several nodes CR716 including CPUs, GPUs, and / or other processors may be grouped within one or more racks to provide computing resources to support one or more workloads. The one or more racks may also include any number of power modules, cooling modules, and / or network switches, in any combination.
[0169] The resource orchestrator 722 can configure or otherwise control one or more nodes CR 716(1)-716(N) and / or grouped computational resources 714. In at least one embodiment, the resource orchestrator 722 can include a software design infrastructure (“SDI”) management entity of the data center 700. The resource orchestrator 722 can include hardware, software, or some combination thereof.
[0170] In at least one embodiment, as shown in FIG. 7 , framework layer 720 may include a job scheduler 732, a configuration manager 734, a resource manager 736, and / or a distributed file system 738. Framework layer 720 may include a framework to support software 732 in software layer 730 and / or one or more applications 742 in application layer 740. Software 732 or applications 742 may include web-based service software or applications, such as those offered by Amazon Web Services, Google Cloud, and Microsoft Azure, respectively. Framework layer 720 may be a type of free and open source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter “Spark”), which may use distributed file system 738 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 732 may include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 700. The configuration manager 734 may be capable of configuring different layers, for example, the software layer 730 and the framework layer 720, which includes Spark and a distributed file system 738 to support large-scale data processing. The resource manager 736 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support the distributed file system 738 and the job scheduler 732. In at least one embodiment, the clustered or grouped computing resources may include the computing resources 714 grouped in the data center infrastructure layer 710. The resource manager 1036 may coordinate with the resource orchestrator 712 to manage these mapped or allocated computing resources.
[0171] In at least one embodiment, software 732 included in software layer 730 may include software used by at least a portion of nodes CR 716(1)-716(N), grouped computational resources 714, and / or distributed file system 738 of framework layer 720. The one or more types of software may include, but are not limited to, internet web page searching software, email virus scanning software, database software, and streaming video content software.
[0172] In at least one embodiment, the applications 742 included in the application layer 740 may include one or more types of applications used by at least a portion of the nodes CR 716(1)-716(N), the grouped computational resources 714, and / or the distributed file system 738 of the framework layer 720. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.
[0173] In at least one embodiment, any of configuration manager 734, resource manager 736, and resource orchestrator 712 can implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically possible manner. The self-modifying actions can free data center operators of data center 700 from making potentially poor configuration decisions and possibly avoiding underutilized and / or underperforming portions of the data center.
[0174] Data center 700 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and / or computing resources described above with respect to data center 700. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 700, for example, by using weight parameters calculated via one or more training techniques, including but not limited to those described herein.
[0175] In at least one embodiment, data center 700 may use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, and / or other hardware (or corresponding virtual computing resources) for training and / or performing inference using such resources. Additionally, one or more of such software and / or hardware resources may be configured as services, such as image recognition, speech recognition, or other artificial intelligence services, to enable users to train or perform inference on information.
[0176] Example Network Environment A network environment suitable for use in implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other back-end devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented with one or more instances of computing device 600 of FIG. 6, e.g., each device may include similar components, features, and / or functionality of computing device 600. Additionally, if a back-end device (e.g., server, NAS, etc.) is implemented, the back-end device may be included as part of data center 700, examples of which are further detailed herein with respect to FIG. 7.
[0177] Components of a network environment may communicate with each other via a network, which may be wired, wireless, or both. A network may include multiple networks or a network of networks. Illustratively, a network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the Public Switched Telephone Network (PSTN), and / or one or more private networks. When a network includes a wireless telecommunications network, components such as base stations, communication towers, or access points (as well as other components) may provide wireless connectivity.
[0178] Compatible network environments may include one or more peer-to-peer network environments (wherein a server may not be included in the network environment) and one or more client-server network environments (wherein a server or servers may be included in the network environment). In a peer-to-peer network environment, functionality described herein with respect to a server may be implemented in any number of client devices.
[0179] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework to support software in the software layer and / or one or more applications in the application layer. The software or applications may each include web-based service software or applications. In an embodiment, one or more of the client devices may use the web-based service software or applications (e.g., by accessing the service software and / or applications via one or more application programming interfaces (APIs)). The framework layer may be, but is not limited to, a type of free and open source software web application framework that may use a distributed file system for large-scale data processing (e.g., “big data”).
[0180] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed across multiple locations from a central or core server (e.g., one or more data centers that may be distributed across a state, region, country, or the world). When a user (e.g., a client device) is connected relatively close to an edge server, the core server may delegate at least a portion of its functionality to the edge server. A cloud-based network environment may be private (e.g., limited to a single organization), public (e.g., available to multiple organizations), and / or a combination thereof (e.g., a hybrid cloud environment).
[0181] A client device may include at least some of the components, features, and functionality of the exemplary computing device 600 described herein with respect to Figure 6. By way of illustration, and not limitation, a client device may be embodied as a personal computer (PC), a laptop computer, a mobile device, a smartphone, a tablet computer, a smart watch, a wearable computer, a personal digital assistant (PDA), an MP3 player, a virtual reality headset, a global positioning system (GPS) or device, a video player, a video camera, a surveillance device or system, a vehicle, a boat, an airship, a virtual machine, a drone, a robot, a handheld communications device, a hospital device, a gaming device or system, an entertainment system, a vehicle computing system, an embedded system controller, a remote control, an appliance, a consumer electronic device, a workstation, an edge device, any combination of these depicted devices, or any other suitable device.
[0182] The present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a personal digital assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in distributed computing environments where tasks are performed by remote processing devices linked through a communications network.
[0183] As used herein, the term "and / or" in reference to two or more elements should be interpreted to mean one element only or a combination of elements. For example, "element A, element B, and / or element C" may include element A only, element B only, element C only, elements A and B, elements A and C, elements B and C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Furthermore, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0184] The subject matter of the present disclosure has been described with specificity to meet statutory requirements. However, that description itself is not intended to limit the scope of the disclosure. Rather, the inventors contemplate that the claimed subject matter may be implemented in other ways, including different steps or combinations of steps similar to those described herein, in conjunction with other current or future technologies. Furthermore, although the terms "step" and / or "block" may be used herein to connote different elements of the method used, these terms should not be construed as implying any particular order among the various steps disclosed herein unless and when the order of individual steps is explicitly described.
Claims
1. receiving audio data generated using a plurality of microphones of an autonomous or semi-autonomous machine; performing an acoustic triangulation algorithm based at least in part on the audio data to determine at least one of a location of an emergency response vehicle or a heading of the emergency response vehicle; generating, based at least in part on the audio data, a representation of a spectrum of one or more frequencies corresponding to one or more audio signals from the audio data; applying first data representing the representation to a neural network; calculating second data representing probabilities of at least a first alert type and a second alert type using the neural network and based at least in part on the first data; determining a type of the emergency response vehicle based at least in part on the probability; performing, by the autonomous or semi-autonomous machine, one or more actions based at least in part on the type of the emergency response vehicle and one or more of the location of the emergency response vehicle or the direction of travel of the emergency response vehicle; A method comprising:
2. The method of claim 1 , wherein the representation of a spectrum of one or more frequencies comprises a Mel spectrogram.
3. The method of claim 1 , wherein the neural network comprises a convolutional recurrent neural network (CRNN).
4. The first alert type and the second alert type are siren, one or more siren sequences or patterns; alarm, a sequence or pattern of one or more alarms; One or more emissions from a vehicle horn; or A sequence or pattern of emissions from at least one vehicle horn The method of claim 1 , wherein the at least one of
5. 10. The method of claim 1, wherein the plurality of microphones are arranged in a plurality of microphone arrays of the autonomous or semi-autonomous machine, each microphone array of the plurality of microphone arrays including a plurality of microphones.
6. 6. The method of claim 5, wherein the plurality of microphone arrays includes two or more of: a first microphone array located at a front of the autonomous or semi-autonomous machine; a second microphone array located at a rear of the autonomous or semi-autonomous machine; a third microphone array located at a left side of the autonomous or semi-autonomous machine; a fourth microphone array located at a right side of the autonomous or semi-autonomous machine; or a fifth microphone array located at a top of the autonomous or semi-autonomous machine.
7. The method of claim 5 , wherein each microphone array of the plurality of microphone arrays includes a windscreen disposed thereon.
8. The method of claim 7, further comprising the step of removing background noise from the audio data to generate processed audio data. further comprising The method of claim 1 , wherein at least one of the executing the acoustic triangulation algorithm or the generating the representation is based at least in part on the processed audio data.
9. The method of claim 8 , wherein the step of removing the background noise comprises the step of performing a beamforming algorithm.
10. The step of calculating the second data comprises: processing the first data representing the representation using one or more gated linear units (GLUs) of the neural network; processing third data generated based at least in part on the output of the one or more GLUs using one or more gated recursive units (GRUs); The method of claim 1 , comprising:
11. The method of claim 10 , wherein attention is applied to fourth data generated based at least in part on the output of the GRU.
12. The method of claim 11 , wherein the attention is applied by one or more dense layers using at least one of a softmax function or a sigmoid function.
13. 2. The method of claim 1, wherein the neural network is trained using recorded audio data and extended audio data, and the extended audio data is generated using the recorded audio data and one or more extension techniques, the one or more extension techniques comprising at least one of time stretching, time shifting, pitch shifting, dynamic range compression, and noise extension at different signal-to-noise ratios (SNRs).
14. receiving audio data generated using a plurality of microphones; generating a spectrogram based at least in part on the audio data; applying first data representing the spectrogram to a deep neural network (DNN); calculating second data using one or more feature extraction layers of the DNN and based at least in part on the first data; computing third data using one or more stateful layers of the DNN and based at least in part on the second data; calculating fourth data representing probabilities of a plurality of alert types using one or more attention layers of the DNN and based at least in part on the third data; performing one or more actions based at least in part on said probabilities; A method comprising:
15. The method of claim 14 , wherein the spectrogram is a Mel spectrogram.
16. pre-processing the audio data using an ambient noise suppression algorithm to generate processed audio data; further comprising generating the spectrogram is based at least in part on the processed audio data; 15. The method of claim 14.
17. The method of claim 14 , wherein the one or more feature extraction layers comprise one or more gated linear units (GLUs).
18. The method of claim 14 , wherein the one or more stateful layers include one or more gated recursive units (GRUs).
19. 15. The method of claim 14, wherein the one or more attention layers include one or more dense layers that use at least one of a softmax function or a sigmoid function.
20. calculating a weighted average of a first output using the softmax function and a second output using the sigmoid function; further comprising the probability corresponds to the weighted average; 20. The method of claim 19.
21. a plurality of microphone arrays, each of the microphone arrays including a plurality of microphones; one or more processing units; one or more memory devices that store instructions that, when executed by the one or more processing units, receiving audio data generated using the plurality of microphone arrays; performing an acoustic triangulation algorithm based at least in part on the audio data to determine at least one of a location of an emergency response vehicle or a heading of the emergency response vehicle; generating, based at least in part on said audio data, a representation of the frequency spectrum of one or more audio signals of said audio data; applying first data representing said representation to a neural network; calculating second data representing probabilities of at least a first alert type and a second alert type using the neural network and based at least in part on the first data; determining a type of the emergency response vehicle based at least in part on the probability; and performing one or more autonomous or semi-autonomous machine actions based at least in part on the type of the emergency response vehicle and one or more of the location of the emergency response vehicle or the direction of travel of the emergency response vehicle. one or more memory devices that cause the one or more processing units to perform operations including A system comprising:
22. 22. The system of claim 21, wherein the plurality of microphone arrays comprises two or more of: a first microphone array located at a front of the autonomous or semi-autonomous machine, a second microphone array located at a rear of the autonomous or semi-autonomous machine, a third microphone array located at a left side of the autonomous or semi-autonomous machine, a fourth microphone array located at a right side of the autonomous or semi-autonomous machine, or a fifth microphone array located at a top of the autonomous or semi-autonomous machine.
23. 22. The system of claim 21, wherein each microphone array of the plurality of microphone arrays includes a windscreen disposed thereon.
Citation Information
Patent Citations
Detection, recognition and position specification for siren signal source
JP2016057295A
Information processor, information processing method and program
JP2016180791A
Sound source detection system and sound source detection method
JP2019121887A
Device, method, and program for controlling mobile body
JP2020044930A
Street light system
JP2020087883A