Autonomous vehicle driving path selection in presence of emergency vehicle sounds

By employing a voice-based emergency vehicle detection and localization system, autonomous vehicles can accurately estimate the position and trajectory of emergency vehicles, thereby modifying their routes to ensure safety and efficiency.

JP2025093302APending Publication Date: 2025-06-23WAYMO LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024206050
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-11-27
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Existing autonomous vehicle systems struggle to accurately detect and respond to emergency vehicles, often leading to confusion among human drivers and potential traffic congestion or safety hazards.

Method used

The implementation of a voice-based emergency vehicle detection and localization system that uses audio localization models to estimate the position and trajectory of emergency vehicles, allowing the autonomous vehicle to modify its driving route accordingly.

Benefits of technology

This solution enables autonomous vehicles to timely and accurately estimate the trajectory of emergency vehicles, selecting efficient and safe driving routes that avoid potential hazards and reduce the risk of traffic congestion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093302000001_ABST
    Figure 2025093302000001_ABST
Patent Text Reader

Abstract

To support sound-based emergency vehicle detection, localization, and tracking for an autonomous vehicle and a driver-assist system.SOLUTION: Techniques of the present invention include obtaining, using one or more audio detectors of a vehicle, a sound recording that includes a sound emitted by an emergency vehicle (EV). The techniques further include applying a sound localization (SL) model to the sound recording to obtain an SL output, which includes a first map of possible locations of the EV in a driving environment of the vehicle and can further include a second map of possible velocities of the EV. The techniques further include simulating, using the SL output, trajectories of simulated EV(s) in the driving environment of the vehicle, and causing, responsive to proximity of the simulated trajectories to a driving path of the vehicle, modification of the driving path of the vehicle.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification generally relates to autonomous vehicles and driving assistance systems. More specifically, this specification relates to the selection of an autonomous vehicle driving route in the presence of emergency vehicle sounds.

Background Art

[0002] Autonomous (fully and partially self-driving) vehicles (AVs) operate by sensing the external environment with various electromagnetic (e.g., radar and optical) as well as non-electromagnetic (e.g., voice and humidity) sensors. Some autonomous vehicles chart a driving route within the environment based on the sensed data. The driving route can be determined based on global positioning system (GPS) data and road map data. GPS and road map data can provide information about the static aspects of the environment (buildings, street layouts, road closures, etc.), while dynamic information (information about other vehicles, pedestrians, street lights, etc.) is obtained from simultaneously collected sensed data. The accuracy and safety of the driving route, as well as the accuracy and safety of the speed regime selected by the autonomous vehicle, depend on the timely and accurate identification of various objects present in the external environment, as well as the ability of the driving algorithm to process information about the environment and provide correct instructions to the vehicle control and drive train.

Brief Description of the Drawings

[0003] This disclosure is shown by way of example and not limitation and can be more fully understood by referring to the following detailed description when considered in connection with the figures.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0004] In one implementation, a method is disclosed that includes obtaining an audio recording including audio emitted by an emergency vehicle (EV) using one or more audio detectors of a vehicle. The method further includes applying an audio localization (SL) model to the audio recording using a processing device to generate an SL output. The SL output includes a first map of possible locations of the EV within the driving environment of the vehicle. The method further includes simulating, by the processing device, one or more simulated EV trajectories within the driving environment of the vehicle using the SL output, and modifying the driving route of the vehicle in response to proximity of one or more of the simulated trajectories to the driving route of the vehicle.

[0005] In another implementation, a method is disclosed that includes a vehicle's sensing system and a vehicle's perception system. The sensing system includes one or more voice detectors configured to obtain a voice recording that includes the voice emitted by an EV. The perception system is configured to apply an SL model to the voice recording to obtain an SL output that includes a first map of possible positions of the EV within the vehicle's driving environment. The perception system is further configured to use the SL output to simulate one or more simulated EV trajectories within the vehicle's driving environment. The perception system is further configured to cause the vehicle's driving route to be modified in response to the proximity of one or more of the simulated trajectories to the vehicle's driving route.

[0006] In another implementation, a non-transitory computer-readable storage medium is disclosed that stores instructions that, when executed by a processing device, cause the processing device to perform operations including obtaining a voice recording that includes the voice emitted by an EV using one or more voice detectors of a vehicle. The operations further include applying an SL model to the voice recording to generate an SL output that includes a first map of possible positions of the EV within the vehicle's driving environment. The operations further include using the SL output to simulate one or more simulated EV trajectories within the vehicle's driving environment and causing the vehicle's driving route to be modified in response to the proximity of one or more of the simulated trajectories to the vehicle's driving route.

DETAILED DESCRIPTION OF THE INVENTION

[0007] Autonomous vehicles (AVs) or vehicles incorporating various driver assistance technologies can use multiple sensor modalities to facilitate the detection of objects in the external environment and determine the trajectories of movement of such objects. The sensors can include radio detection and ranging (radar) sensors, light detection and ranging (Lidar) sensors, multiple digital cameras, voice sensors (microphones), position sensors, and the like. Different types of sensors can provide different complementary benefits. For example, radar and Lidar emit electromagnetic signals (radio signals or optical signals) that are reflected from objects and convey information about the distance to the object (e.g., from the time of flight of the signal) and the velocity of the object (e.g., from the Doppler shift of the frequency of the reflected signal). Radar and Lidar can scan the entire 360-degree view by using a series of continuous sensing frames. The sensing frames can include a large number of reflections that cover the external environment within a high-density grid of return points. Each return point can be associated with the distance to the corresponding reflecting object and the radial velocity of the reflecting object (the component of the velocity along the line of sight). Cameras (e.g., still or video cameras) can acquire high-resolution images at short and long distances and can complement Lidar and radar data. Microphones can detect significant sounds such as sirens, crossing bell klaxons, train klaxons, and / or the like.

[0008] Lidar, radar, and cameras (including infrared cameras) operate using electromagnetic waves with relatively small wavelengths (having radar with the longest wavelength within the centimeter range or less). As a result, the sensing data obtained by electromagnetic sensors is mainly limited to direct line-of-sight detection. On the other hand, human drivers have sensory capabilities beyond line-of-sight perception. In particular, human drivers can hear the sirens of emergency vehicles approaching even when the emergency vehicle is approaching along a different (e.g., perpendicular) road and / or is blocked by other vehicles or buildings, including situations like this. An emergency vehicle (EV) may have a recognizable shape and appearance (e.g., a fire truck, ambulance, etc.) and be equipped with emergency lighting, but timely detection of an emergency vehicle on a life-saving mission based solely on the detection of emergency lighting and / or the vehicle's exterior appearance is difficult and may be insufficient in many situations. However, the sound waves of an emergency siren typically have wavelengths in the range of 20 - 90 centimeters and are thus very efficient at carrying sound around obstacles. Human hearing can extract a lot of useful information from sound waves, such as the distance and direction to the sound source (even when the sound source is not directly visible), the state of movement of the sound source (e.g., stationary, approaching, departing, passing, and / or the like), an estimate (to some extent) of the speed of this movement, and / or the like. In particular, existing computer systems and autonomous driving systems still do not have capabilities comparable to or close to those of human hearing and perception. Therefore, an AV that detects emergency sounds (e.g., sirens) may have difficulty in selecting an optimal driving route. For example, with due care, an AV can pull to the side of the road or stop before entering an intersection even when the EV has already passed through the intersection or is moving away from the AV. This can be confusing for other road participants (e.g., human drivers) and may lead to traffic congestion and / or the creation of dangerous driving conditions.

[0009] Aspects and implementations of the present disclosure address these and other challenges of existing autonomous driving technologies by disclosing a method and system for determining the likelihood of the position of an emergency vehicle on a drivable road within a vehicle, e.g., in an AV environment. Determining such a position, referred to herein as localization, can involve a combination of machine learning techniques that estimate the possible direction, and distance, to the sound source of an audible EV signal, as well as drivable road layout data. More specifically, a digital representation of sound in an AV environment, e.g., a spectrogram of audio data (e.g., collected by a microphone), can be used as an input to an audio localization model that outputs a probability map in which the sound source of the EV sound is positioned at some location relative to the AV. Next, an AV planner module can overlay the probability map, considered as a map of initial hypotheses of the EV position, onto a drivable road layout and exclude at least a portion of the probability map. For example, a portion of the probability map can correspond to non-drivable areas of the environment, e.g., areas occupied by buildings or structures, fenced-off areas, impassable areas, and / or the like, and can thus be excluded from the probability map. Another portion of the probability map can be identified as the visible (to the AV's perception system) portion of the environment (e.g., based on the drivable road layout and on-board perception data). Such portions can also be excluded from the probability map under the assumption that, if the EV is located in the visible area of the environment, the presence of a flashing light (associated with the EV sound) would be detected by the perception system. Yet another portion of the probability map can be excluded based on the corresponding probability being too low, e.g., based on falling below an empirically set threshold.

[0010] The resulting area (referred to herein as the candidate area) can be used by the AV's planner module for EV motion simulation. More specifically, the planner module can place the simulated (hypothesized) EV within the candidate area. The simulated EV can obtain both a position and a velocity (understood as both the speed of movement and its direction), and the planner module (which can be part of the vehicle's perception and planning system) can predict the movement of the simulated EV to determine whether the EV will pass close to the AV (e.g., closer than an empirically set minimum distance) and / or whether the EV's path can be blocked by the movement or position of the AV. In some implementations, the output of the voice localization model can further include the likelihoods of various speeds

Number

[0011] In some cases, for example, due to high levels of ambient audio noise, the audio localization model may output a probability map (of the possible positions and / or velocities of EVs) with low reliability, e.g., a value below a threshold reliability. In such cases, the planner may place simulated EVs throughout the occluded regions not visible to the perception system and run a simulation that is not constrained by the velocities of the simulated EVs up to, or even above, a sufficiently high set maximum speed, e.g., twice the maximum legal speed for the region. The movement of the simulated EVs may, in some cases, not need to be restricted to a particular side of the road (e.g., the right side) as the EVs may be able to move with respect to traffic. Numerous other implementations and uses of the systems and techniques of the present disclosure are shown below.

[0012] Advantages of the systems and techniques of the present disclosure include, but are not limited to, obtaining a timely and accurate estimate of the trajectory of an EV and selecting an efficient and safe driving route for an AV. This ensures that, on the one hand, the AV avoids entering the active regions where an EV may be present and, on the other hand, prevents the AV from making unnecessary stops and / or decelerations that could confuse other road users and create the potential for traffic congestion or vehicle collisions.

[0013] In these examples where the implementation description refers to autonomous vehicles, it should be understood that similar techniques can be used in various driver assistance systems that do not scale up to the level of a fully autonomous driving system. More specifically, the disclosed techniques can be used in level 2 driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, etc., and other driver support. Similarly, the disclosed techniques can be used in level 3 driver assistance systems that enable autonomous driving under limited (e.g., highway) conditions. In such systems, timely and accurate detection of approaching EVs can be used to notify the driver (e.g., in a level 2 system) that a change in the driving route may be recommended, or to make specific driving decisions such as reducing speed, pulling over to the side of the road (e.g., in a level 3 system) without requiring driver feedback.

[0014] Figure 1 shows the components of an exemplary autonomous vehicle (AV) 100 that can estimate the location of an EV using voice detection and processing techniques and determine the driving route of the autonomous vehicle, according to some implementations of the present disclosure. Autonomous vehicles can include motor vehicles (cars, trucks, buses, motorcycles, buggy vehicles, recreational vehicles, any special agricultural or construction vehicles, etc.), aircraft (airplanes, helicopters, drones, etc.), marine vessels (ships, boats, yachts, submarines, etc.), or any other self-propelled vehicle capable of operating in an autonomous driving mode (e.g., without human input or with reduced human input) (e.g., robots, robotic vehicles in factories and warehouses, sidewalk delivery robotic vehicles, etc.).

[0015] As described herein, in a semi-autonomous driving mode or a partial autonomous driving mode, the vehicle assists with one or more driving operations (e.g., steering, braking, and / or accelerating to perform lane centering, adaptive cruise control, an advanced driver assistance system (ADAS), or emergency braking), but the human driver is expected to situationally perceive the surroundings of the vehicle and monitor the assisted driving operations. Here, the vehicle may perform all driving tasks in certain situations, but the human driver is expected to take responsibility for control as needed.

[0016] For simplicity and conciseness, various systems and methods are described below in conjunction with autonomous vehicles, but the same techniques can be used in various driver assistance systems that do not reach the level of a fully autonomous driving system. In the United States, the Society of Automotive Engineers (SAE) has defined different levels of automated driving operation to indicate how much or how little a vehicle controls the driving, but different organizations in the United States or other countries may classify the levels differently. More specifically, the disclosed systems and methods can be used in SAE Level 2 (L2) driver assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, etc., and other driver support. The disclosed systems and methods can be used in SAE Level 3 (L3) driver assistance systems that can drive autonomously under limited (e.g., highway) conditions. Similarly, the disclosed systems and methods can be used in vehicles that use SAE Level 4 (L4) automated driving systems that operate autonomously in most normal driving situations and require only occasional attention from a human operator. In all such systems, accurate lane estimation is automatically performed without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, as well as overall safety of autonomous driving, semi-autonomous driving, and other driver assistance systems. As noted above, in addition to the way SAE classifies the levels of automated driving operation, other organizations in the United States or other countries may classify the levels of automated driving operation differently. Without limitation, the disclosed systems and methods in this specification can be used in driver assistance systems defined by the levels of automated driving operation of these other organizations.

[0017] The driving environment 101 can include any (moving or stationary) object located outside the AV, such as roads, buildings, trees, bushes, sidewalks, bridges, mountains, other vehicles, pedestrians, etc. The driving environment 101 can be an urban area, a suburban area, a rural area, etc. In some implementations, the driving environment 101 can be an off-road environment (e.g., agriculture, or farmland). In some implementations, the driving environment can be an indoor environment, such as the environment of an industrial plant, a shipping warehouse, a hazardous area of a building, etc. In some implementations, the driving environment 101 can be substantially flat with various objects moving parallel to the surface (e.g., parallel to the surface of the earth). In other implementations, the driving environment can be three-dimensional, and the vehicle can include three directions (e.g., balloon, leaf, etc.) located on the hill side and / or objects that can move along all roads such that the visibility of other vehicles and objects can be limited when the direct line of sight is blocked by the terrain of the ground. Hereinafter, the term "driving environment" should be understood to include all environments in which autonomous movement of a self-driving vehicle can occur. For example, the "driving environment" can include any possible flight environment of an aircraft or marine environment of a marine vessel. The objects in the driving environment 101 can be located at any distance from the AV, from a short distance of several feet (or less) to several miles (or more).

[0018] For simplicity and conciseness, various systems and methods are described below in conjunction with autonomous vehicles, but the same technologies can be used in various driving assistance systems that do not reach the level of a fully autonomous driving system. In the United States, the Society of Automotive Engineers (SAE) has defined different levels of automated driving operation to indicate how much or how little a vehicle controls the driving, but different organizations in the United States or other countries may classify the levels differently. More specifically, the disclosed systems and methods can be used in SAE Level 2 (L2) driving assistance systems that implement steering, braking, acceleration, lane centering, adaptive cruise control, etc., and other driver support. The disclosed systems and methods can be used in SAE Level 3 (L3) driving assistance systems that are capable of autonomous driving under limited (e.g., highway) conditions. Similarly, the disclosed systems and methods can be used in vehicles that use SAE Level 4 (L4) automated driving systems that operate autonomously in most normal driving situations and require only occasional attention from a human operator. In all such driving assistance systems, accurate lane estimation is automatically performed without driver input or control (e.g., while the vehicle is moving), resulting in improved reliability of vehicle positioning and navigation, as well as the overall safety of autonomous driving, semi-autonomous driving, and other driving assistance systems. As noted above, in addition to the way SAE classifies the levels of automated driving operation, other organizations in the United States or other countries may classify the levels of automated driving operation differently. Without limitation, the disclosed systems and methods in this specification can be used in driving assistance systems defined by the levels of automated driving operation of these other organizations.

[0019] The AV100 of the embodiment may include a sensing system 110. The sensing system 110 may include various electromagnetic (e.g., optical) and non-electromagnetic (e.g., audio) sensing subsystems and / or devices. The sensing system 110 can include radar 114 (or multiple radars 114), which can be any system that utilizes radio or microwave frequency signals to sense objects within the operating environment 101 of the AV100. The radar(s) 114 can be configured to sense both the spatial position of objects (including their spatial dimensions) and their velocity (e.g., using Doppler shift techniques). Hereinafter, "velocity" refers to both how fast an object is moving (the speed of the object) and the direction of the object's movement. The sensing system 110 may include Lidar 112, which can be a laser-based unit that can determine the distance to objects within the operating environment 101 and the velocity of the objects. Each of Lidar 112 and radar 114 can include a coherent sensor, such as a frequency-modulated continuous-wave (FMCW) Lidar or radar sensor. For example, radar 114 can use heterodyne detection for velocity determination. In some implementations, the functionality of ToF and coherent radar is combined in a radar unit that can simultaneously determine both the distance to a reflecting object and the radial velocity. Such units can be configured to operate in a non-coherent sensing mode (ToF mode) and / or a coherent sensing mode (e.g., a mode that uses heterodyne detection) or both modes simultaneously. In some implementations, multiple Lidar 112 or radar 114 can be mounted on the AV100.

[0020] Lidar 112 may include one or more light sources that generate and emit signals, and one or more detectors for signals reflected back from objects. In some implementations, Lidar 112 can perform a 360-degree scan in the horizontal direction. In some implementations, Lidar 112 may be capable of spatial scanning along both the horizontal and vertical directions. In some implementations, the field of view may be up to 90 degrees in the vertical direction (e.g., at least a portion of the region above the horizon is scanned by the radar signal). In some implementations, the field of view may be global (consisting of two hemispheres).

[0021] The perception system 110 can further include one or more cameras 118 (which may include one or more infrared sensors) to capture images of the driving environment 101. The image can be a two-dimensional projection of the driving environment 101 (or a part of the driving environment 101) onto the projection plane (flat or non-flat) of the camera. Some of the cameras 118 of the perception system 110 can be video cameras configured to capture a continuous (or quasi-continuous) stream of images of the driving environment 101. The perception system 110 can further include one or more ultrasonic sensors 116 that can be used, in some implementations, to identify objects located near the AV100 and / or to assist in parking and other AV operations at low speeds. The perception system 110 can also include one or more microphones 119 that can be positioned around the AV100. In some implementations, each microphone 119 can be arranged within a microphone array of two or more microphones. The AV100 can have a plurality of such microphone arrays, for example, a four-microphone array, an eight-microphone array, or some other number of microphone arrays. In one embodiment, two microphone arrays can be deployed near the front left corner and the front right corner of the AV100, and two microphone arrays can be deployed near the rear left corner and the rear right corner of the AV100. In some implementations, different microphones of a given array can be positioned at a distance of 1 to 5 centimeters from each other. In some implementations, the microphones can be positioned at a distance of, for example, more than 10 cm from each other, up to a maximum. In some implementations, the microphones within a given array can be time-synchronized. In some implementations, different arrays of microphones are not synchronized. In some implementations, the microphones can also be synchronized across different arrays.

[0022] The perception data acquired by the perception system 110 can be processed by the data processing system 120 of the AV100. For example, the data processing system 120 may include a perception and planning system 130. The perception system 130 may be configured to detect and track objects in the driving environment 101 and recognize the detected objects. For example, the perception and planning system 130 can analyze the images captured by the camera 118 and further has the ability to detect traffic signals, road signs, road layouts (e.g., lane boundaries, intersection topologies, parking lot designations, etc.), the presence of obstacles, etc. The perception and planning system 130 may also receive radar perception data (Doppler data and ToF data) to determine the distances to various objects in the environment 101 and the speeds of such objects (in the line-of-sight direction and, in some implementations, laterally as described below). In some implementations, the perception and planning system 130 can use radar data in combination with the data captured by the camera(s) 118, as described in more detail below.

[0023] The perception and planning system 130 may include some components and / or modules to facilitate the detection and localization of EVs, as disclosed herein. In some implementations, the perception system 130 uses the audio data collected by the microphone 119 to determine the distance D to the sound source, the direction θ (azimuthal angle or bearing) to the sound source, and the speed

Number

[0024] The perception and planning system 130 may also include a behavior prediction module (not explicitly shown) that can monitor how the driving environment 101 evolves over time, for example, by continuously tracking the position and velocity of moving objects (e.g., relative to the Earth). In some implementations, the behavior prediction module can track the changing appearance of the environment due to the movement of the AV through the environment. In some implementations, the behavior prediction module can predict how the various tracked objects in the driving environment 101 will be positioned within a prediction time frame. The prediction may be based on the current position and velocity of the tracked objects, including the EV whose position is determined using the output of the ESLM 132. In some implementations, the output of the ESLM 132 can be combined with the output of Lidar / radar / camera-based object tracking.

[0025] In some implementations, after the ESLM 138 determines a probability map predicting the approximate distance, direction, and / or velocity of an occluded EV (e.g., a fire truck) with its siren on, the planner module 134 executes a simulation to predict a set of future events at specific time intervals, e.g., 0.5 seconds, or some other interval, t1, t2,...t N of one or more, at which the position (trajectory) of the EV

Number

Number

[0026] The perception and planning system 130 can also receive information from a positioning subsystem 122 that can include a GPS transceiver and / or an inertial measurement unit (IMU) configured to obtain information about the position of the AV with respect to the Earth and its surroundings. The positioning subsystem 122 can use positioning data, e.g., GPS data and IMU data, in combination with the perception data to assist in accurately determining the position of the AV100 relative to fixed objects in the driving environment 101, such as roadways, lane boundaries, intersections, sidewalks, crosswalks, road signs, surrounding buildings, etc., whose positions are provided by the road layout information 124. In some implementations, the data processing system 120 can receive non-electromagnetic data such as audio data (e.g., ultrasonic sensor data, or data from a microphone that picked up the siren of an emergency vehicle), temperature sensor data, humidity sensor data, pressure sensor data, weather data (e.g., wind speed and direction, precipitation data), etc.

[0027] The AVCS 140 may include one or more algorithms that control how the AV should behave in various driving situations and environments. For example, the AVCS 140 can include a navigation system for determining a global driving route to a destination. The AVCS 140 can also include a driving route selection system for selecting a specific route through the current driving environment, which may include lane selection, getting around traffic jams, choosing where to make a U-turn, selecting a trajectory for parking operations, and the like. The AVCS 140 can also include an obstacle avoidance system for safely avoiding various obstacles (such as rocks, stalled vehicles, pedestrians crossing against traffic rules and signals, etc.) within the driving environment of the AV. The obstacle avoidance system can be configured to evaluate the size of the obstacle and the trajectory of the obstacle (if the obstacle is moving), and select an optimal driving strategy (such as braking, steering, accelerating, etc.) for avoiding the obstacle.

[0028] The algorithms and modules of the AVCS 140 can generate instructions for various vehicle systems and components, such as the powertrain, brakes, and steering 150, vehicle electronics 160, signaling 170, and other systems and components not explicitly shown in FIG. 1. The powertrain, brakes, and steering 150 can include an engine (internal combustion engine, electric engine, etc.), transmission, differential, axles, wheels, steering mechanism, and other systems. Vehicle electronics 160 can include an on-board computer, engine management, ignition, communication systems, car computer, telematics, in-vehicle entertainment systems, and other systems and components. Signaling 170 can include high and low beam headlights, stoplights, turn signals and backup lights, sirens and alarms, interior lighting systems, dashboard notification systems, passenger notification systems, wireless and wireless network transmission systems, and the like. Some of the instructions output by the AVCS 140 can be delivered directly to the powertrain, brakes, and steering 150 (or signaling 170), while other instructions output by the AVCS 140 are first delivered to the vehicle electronics 160, which generates instructions for the powertrain, brakes, and steering 150, and / or signaling 170.

[0029] In one example, the AVCS 140 may determine that an obstacle identified by the data processing system 120 should be avoided by decelerating the vehicle until a safe speed is reached and then steering the vehicle around the obstacle. The AVCS 140 outputs commands to the power train, brakes, and steering 150 (either directly or via the vehicle electronics 160) to (1) reduce the flow of fuel to the engine and decrease the engine rpm by changing the throttle setting, (2) downshift the drive train to a lower gear via the automatic transmission, (3) engage the brake unit to decrease the speed of the vehicle until a safe speed is reached (while operating in cooperation with the engine and transmission), and (4) perform a steering operation using the power steering mechanism until the obstacle is safely bypassed. Thereafter, the AVCS 140 may output commands to the power train, brakes, and steering 150 to resume the previous speed setting of the vehicle.

[0030] "Autonomous vehicle" can include motor vehicles (such as cars, trucks, buses, motorcycles, buggy vehicles, recreational vehicles, any special agricultural or construction vehicles, etc.), aircraft (such as airplanes, helicopters, drones, etc.), marine vessels (such as ships, boats, yachts, submarines, etc.), robotic vehicles (e.g., factory, warehouse, sidewalk delivery robots) or any other self-propelled vehicle capable of operating in an autonomous driving mode (without human input or with reduced human input). "Object" can include any vehicle body, item, device, vehicle body, or article (moving or stationary) located outside the autonomous vehicle, such as roads, buildings, trees, grasslands, sidewalks, bridges, mountains, other vehicles, piers, dikes, runways, animals, birds, or other things.

[0031] FIG. 2A is a diagram showing an exemplary voice-based emergency vehicle detection, localization, and tracking pipeline 200 that can be used as part of a perception and planning system for an autonomous vehicle according to some implementations of the present disclosure. The voice separation and processing pipeline 200 may include a voice sensor 202 such as the microphone 119 of FIG. 1, which may be arranged in one or more time-synchronized arrays and may be located, for example, around the AV. The microphones within the microphone array may be located at a distance of 1 to 5 cm from each other. In some implementations, the microphones within the array may be located at a distance less than 1 cm or greater than 5 cm (e.g., 10 cm or more). The microphones may be omnidirectional (cardioid) microphones, bidirectional microphones, omnidirectional microphones, dynamic microphones, multi-pattern microphones, and / or some combination thereof. In some implementations, the microphone array may include directional microphones with different orientations of the maximum sensitivity axis.

[0032] The voice collected by the voice sensor 202 may be in any suitable raw voice format, spectrogram format, or some other digital format. More specifically, the voice sensor 202 converts the air pressure fluctuations caused by the arriving sound waves into analog electromagnetic signals, and in some implementations, these analog signals may be digitized. The digitized signals may be provided to a spectrum analyzer 204 that calculates a Fourier transform (e.g., using a fast Fourier transform) at various time intervals over a predetermined period to obtain voice frames. Each individual voice frame may represent the voice content at each time interval. In some implementations, the amplitude of the signal in the frequency domain may be represented using a logarithmic (decibel) scale. In some implementations, the spectrogram may be a mel spectrogram in which the frequency f (measured in Hz) is converted to the mel domain according to f→m = 1607·ln(1 + f / 700) or a similar transformation.

[0033] The audio frame can be input into an Emergency Sound Location Model (ESLM) 132 trained to generate a prediction 220 regarding the position and movement state of the EV in the external environment. In some implementations, the input to the ESLM 132 can include multiple types of frames (or spectrograms). For example, the F-frame 206 can have a high frequency resolution for efficiently capturing the frequency content of the sound wave detected by the audio sensor 202. The high frequency resolution can be achieved by a long sampling time. With the high frequency resolution, the ESLM 132 can identify the signature of the changes in the pitch and amplitude of the sound wave and can ensure good accuracy in detecting the distance D to the emergency sound source. The T-frame 208 can have a high time resolution for efficiently capturing the difference in the arrival time of the sound at different microphones of the time synchronization array. The high time resolution can be achieved by a shortened interval of points (high sampling rate) used in the Fourier transform. With the high time resolution, the ESLM 132 can identify the orientation of the front of the wave of the arriving sound wave and thus ensure good accuracy in detecting the direction θ to the emergency sound source. Both the F-frame 206 and the T-frame 208 can include information that improves the accuracy of speed detection because the speed can be related to both the radial speed V D =ΔD / Δt (related to the change in the distance ΔD that is efficiently tracked through the progression of the F-frame 206) and the azimuthal speed V θ =DΔθ / Δt (related to the change in the azimuthal angle (bearing) Δθ that is efficiently tracked through the progression of the T-frame 208).

[0034] More specifically, the complex signal S(t) detected by any given audio sensor 202 can be sampled N times over the window T, for example, using a suitable analog-to-digital (ADC) converter of the spectrum analyzer 204, to generate a set of (time) points t j =jT / N, where j = 0, 1, 2,..., N - 1. Then, the spectrum analyzer 204 applies an N-point Fourier transform to the signal S(t j ) to obtain, for example, N complex Fourier coefficients S(f k) can generate a set of: [Number] In the formula, the frequency is fk = k / T, where k = 0, 1, 2, ..., N - 1.

[0035] The F frame 206 can be generated by selecting a longer time interval, for example, T = 2 seconds, 3 seconds, etc. Accordingly, in order to ensure that typical EV sound frequencies f = 650 - 1000 Hz are properly detected, the maximum frequency f MAX ≒ N / T should be greater than the frequency f, for example, some coefficient f MAX > 2, 3, etc., should exceed the frequency f. For example, when T = 3 seconds and f = 1000 Hz, by selecting N = 104, f MAX / f ≒ 3.3 is ensured.

[0036] The T frame 208 can be generated by selecting a shorter time interval so that the individual points of the Fourier transform are separated for a time short enough to be sensitive to the difference in the arrival time of the sound at different voice sensors 202 (to capture the direction dependence of the arriving sound wave). For example, in a microphone array where the microphones are separated by 10 cm, the sound can reach different microphones with an order of time difference (Equation 2) (for example, in this embodiment, the time it takes for the sound to cover half the distance between the microphones, for example, 5 cm). Accordingly, in order to ensure that such time is fully resolved, the separation of the individual points of the Fourier transform T / N should be substantially smaller (for example, by a factor of 2, 3,...) than Δt. For the same number of points, for N = 104, this means that the duration of the sampling window can be T = 0.5 seconds or less (for example, 0.3 seconds, 0.2 seconds, etc.) to ensure that the interval T / N = 0.5×10 -4 seconds is at least three times less than the target time difference Δt. [Number]

[0037] The durations of the sampling window T and the sampling rate N / T for the F-frame 206 and the T-frame 208 should be understood as examples and are not limiting. In some implementations, for higher resolution, the sampling window duration T can be more than 3 seconds for the F-frame 206 and less than 0.5 seconds for the T-frame 208. Higher resolution (in both types of frames) can be achieved by using a Fourier transform with more points N than 104. In the implementation shown above, the number of points N used to acquire the F-frame 206 and the T-frame 208 is the same, but in other implementations, the number of points used to acquire the F-frame 206 and the T-frame 208 can be different. For simplicity and clarity, the conversion from the audio signal S(t j ) to the Fourier spectrogram S(f k ) is described, but in other implementations, the audio signal S(t j ) can be represented in mel spectrogram format.

[0038] In some implementations, the spectrogram S(f k ) can include separate spectrograms of the real part S'(t j ) and the imaginary part S''(t j ) of the complex audio signal, S(t j ) = S'(t j ) + iS''(t j ). In some implementations, the spectrogram S(fk) can include separate spectrograms for the amplitude A(t j ) and the phase φ(t j ) of the complex audio signal.

Number

[0039] The F-frame 206 and the T-frame 208 (also referred to herein as the speech frames) can be processed using the ESLM 132. In some implementations, the speech frames can be converted from a spectrogram representation to a speech embedding representation using, for example, the Wav2vec algorithm or some other similar waveform-embedding converter (not explicitly shown in FIG. 2A). The embedding (feature vector) should be understood as a vector (string) of any number M of components that can have, as any suitable digital representation of the input data, for example, integer values or floating-point values. The embedding can be regarded as a vector or a point within an M-dimensional embedding space. The dimension M of the embedding space (defined as part of the ESLM 132 architecture) can be smaller than the size of the input data (speech frames). During training, the ESLM 132 learns to associate a similar set of training speech frames with similar embeddings represented by points located close to each other within the embedding space, and further learns to associate a different set of training speech frames with points located further apart within that space. In some implementations, a distinct speech embedding (or a distinct set of speech embeddings) can represent a given speech frame.

[0040] In some implementations, the ESLM 132 can be (or include) a neural network, for example, a deep neural network having a plurality of layers of neurons. The first (input) neuron layer of the ESLM 132 can receive the generated embedding, process the embedding, and pass the output of the processing to the next layer, and so on until the last neuron layer of the ESLM 132 generates a final output, for example, a prediction 220. In some implementations, the ESLM 132 can process a combination of multiple voices simultaneously or using batch processing.

[0041] In some implementations, the ESLM 132 may include a convolutional neural network. In some implementations, the convolution may be performed over the time domain (over different frames) and over different embeddings corresponding to a given frame. In some implementations, some of the convolutional layers of the ESLM 132 may be transposed convolutional layers. Some of the layers of the ESLM 132 may be fully connected (dense) layers. The ESLM 132 may further include one or more skipped connections and one or more batch normalization layers.

[0042] The prediction 220 may include any suitable representation of the estimated position of the sound source of an emergency signal (e.g., a fire truck, an ambulance, a police vehicle, etc.). Due to the relatively large wavelength of the sound and the resulting diffraction and multiple scattering of the sound from various objects in the environment, the accuracy of the prediction 220 may be lower than the accuracy of direct line-of-sight optical (camera) detection or Lidar / radar detection. In some implementations, the estimated position of the sound source is a probability map (heat map) P(X i , Y j ; τ) for a set of audio frames (F frame 206 and T frame 208) associated with a particular time τ. The probability map P(X i , Y j ; τ) can characterize the probability that a particular region (pixel) of space with coordinates X i , Y j (some suitable dimensions ΔX, ΔY) is occupied by an EV emitting an emergency sound. The prediction 220 may further include the estimated speed of the EV. In some implementations, the estimated speed can likewise be output via another probability map P(X i , Y j ; τ) within the speed space VX, VY. Some implementations may unfold the polar coordinates of the coordinate probability map P(D, θ; τ) and / or the speed probability map P(V D , V θ ; τ).

[0043] In some implementations, the prediction 220 may also include one or more confidence scores, e.g., C D , C θ , CV can include C D is a reliability score that describes the confidence of the ESLM132 at the predicted distance to the EV, C θ is a reliability score that describes the confidence of the ESLM132 at the predicted direction to the EV, C V is a reliability score that describes the confidence of the ESLM132 at the predicted speed of the EV. In some implementations, the set of output reliability scores can have fewer than three values (e.g., the ESLM132 can output a single aggregated reliability score that characterizes the overall confidence in the coordinate speed prediction) or more than three values.

[0044] Prediction 220 can be provided to the planner module 134, which, as disclosed in more detail below in conjunction with FIG. 4, combines the coordinate and velocity probability maps with the road layout data to identify the likely positions of the EVs and performs EV motion simulations.

[0045] In some implementations, an additional emergency speech identification model (ESIM) 205 can be deployed in front of the ESLM132, e.g., as a gateway to the ESLM132. The ESIM 205 can be a lightweight model with fewer neuron layers / neurons compared to the ESLM132 and can be trained using a smaller set of training data. The ESIM 205 can quickly identify whether the emergency sound is audible. If no emergency sound is detected, there is no need to deploy the ESLM132. This prevents the AV perception and planning system from performing unnecessary processing under normal driving conditions. In some implementations, the ESIM 205 can process the F frame 206 and / or the T frame 208. In some implementations, the ESIM 205 can process a set of frames different from the F frame 206 and / or the T frame 208, e.g., low-resolution audio frames, with a smaller number of Fourier components N.

[0046] ESLM132 and (when used) ESIM205 can be trained by the training server 240. The model(s) can be trained using voice data 254, e.g., emergency sounds recorded in various driving environments including urban driving environments, highway driving environments, rural driving environments, off-road driving environments, and / or the like. In the case of supervised training, the training data can be annotated with ground truth 256, which can include the position and speed of the EVs emitting the emergency sounds. In some implementations, additional sensor data, e.g., one or more of Lidar data 212, radar data 214, or camera data 216, can be used to perform the annotation. Since the most relevant voice data 254 can record the sounds emitted by EVs that are not within the line of sight of the vehicle recording the voice data 254, the ground truth 256 can be generated using sensor data collected by a fleet of multiple vehicles that operate in cooperation with the vehicle recording the voice data 254. For example, a fleet of vehicles A1, A2,...A M can be deployed. For example, when an emergency sound is detected by one of the vehicles such as vehicle A1, at least some of the other vehicles in the fleet such as vehicle A2 may have a direct line of sight to the EV emitting the sound and can accurately determine the position of the EV (as part of the ground truth 256). In some implementations, for example, when vehicle A2 is an autonomous vehicle, the sensor data (Lidar data 212, radar data 214, and / or camera data 216) can be processed by the trained object detection / classification model 230 to identify and locate the EV emitting the emergency sound. In some implementations, the additional input to the object detection / classification model 230 can include road layout information 124 for the accurate placement of the EV with respect to the road. In some implementations, for example, when vehicle A2 is a human driver-operated vehicle, the (timestamped) sensor data can be recorded together with the voice data 254 collected by vehicle A1 and then processed offline. The output of the object detection / classification model 230 can be saved as the ground truth 256 (annotation) of the voice data 254 in the data repository 250.

[0047] The training of ESLM132 and ESIM205 can be carried out by a training engine 242 hosted by a training server 240, which can be an external server introducing one or more processing devices, such as a central processing unit (CPU), a graphics processing unit (GPU), and / or the like. In some implementations, one or both models can be trained by the training engine 242 and then downloaded onto the perception and planning system of the autonomous vehicle. The various models illustrated in FIG. 2A can be trained using training data including training inputs 244 and corresponding target outputs 246 (correct matches for each training input). During the training of the model, the training engine 242 can find patterns in the training data that map each training input 244 to its respective target output 246.

[0048] The training engine 242 can have access to a data repository 250 that stores teacher voice data 254 and ground truth 256 for actual driving situations in various environments. The training data stored in the data repository 250 can include, for example, a large dataset having hundreds or thousands of voice recordings. During training, the training server 240 can retrieve training data from the data repository 250, generate one or more training inputs 244 and one or more target outputs 246, and use the training inputs 244 and target outputs 246 to train ESLM132 and / or ESIM205.

[0049] During model training, the training engine 242 can change the model's parameters (e.g., weights and biases) until the model successfully learns how to perform each task, such as estimating the position and movement of an acoustic generation EV (for ESLM132) or identifying the presence of EV sounds within acoustic data (for ESIM205). In some implementations, the different models of FIG. 2A can be trained separately. In some implementations, the models can be trained together (e.g., simultaneously). Different models can have different architectures (e.g., different numbers of neuron layers and neural connections of different topologies), can have different settings (e.g., activation functions, etc.), and can be trained using different hyperparameters.

[0050] The data repository 250 can be persistent storage capable of accommodating acoustic data 254, the ground truth of the acoustic data 256, and any additional data such as Lidar data 212, radar data 214, camera data (images) 216, and / or road layout information 124 that can be used to train a model operating according to various implementations of the present disclosure. The data repository 250 can be hosted by one or more storage devices such as main memory, magnetic or optical storage disks, tapes, or hard drives, network-connected storage devices (NAS), storage area networks (SAN), etc. Although shown separately from the training server 240, in an implementation, the data repository 250 can be part of the training server 240. In some implementations, the data repository 250 can be a network-connected file server, while in other implementations, the data repository 250 can be some other type of persistent storage, such as an object-oriented database, a relational database, etc., that can be hosted by one or more different machines accessible to the training server 240 via a server machine or a network (not shown in FIG. 2A).

[0051] In some implementations, in addition to the F-frame 206 and the T-frame, the input to the ESLM 132 can include (in training and inference) an embedding obtained by processing any of the auxiliary inputs 209, such as any of Lidar data 212, radar data 214, camera data (images) 216, and / or road layout information 124, which are sensed data collected by the AV's sensing system. This auxiliary input 209 can provide additional context for processing the voice data input.

[0052] FIG. 2B is a diagram showing an exemplary architecture of an emergency voice localization model 132 that can be deployed as part of the detection, localization, and tracking pipeline 200 of FIG. 2A, according to some implementations of the present disclosure. As shown in FIG. 2B, the ESLM 132 can include separate inputs for receiving the F-frame 206 and the T-frame 208. The F-network 260 can process the F-frame 206 or some suitable representation (e.g., an embedding) of the F-frame 206. Similarly, the T-network 264 can process the T-frame 208 or some representations of the F-frame 206. Each of the F-network 260 and the T-network 264 can process the corresponding input and generate intermediate embeddings 262 and 266. The intermediate embeddings 262 and 266 can be, for example, concatenated and translated into a fused embedding 270. In those implementations where the auxiliary input 209 is used, the embedding 268 associated with the auxiliary input 209 can also be combined into the fused embedding 270. Next, the fused embedding 270 can be processed by a fusion network 280 that outputs a prediction 220.

[0053] In some implementations, the F network 260 and the T network 264 can be or include convolutional neural networks. The fusion network 280 can be or include one or more fully connected layers and a final classifier layer. The inputs to the F network 260 and the T network 264 include frames obtained by processing audio data collected from one, some, or all of the AV's audio sensors (e.g., the audio sensor 202 of FIG. 2A). The training of the ESLM 132 can include dropout techniques. More specifically, the input from at least some of the audio sensors can be turned off (e.g., can be replaced with zeros). The audio sensors can be selected to dropout based on any given schedule or randomly. For example, if the AV deploys two arrays of three microphones each, one or two microphones can be turned off for a given training epoch. During some (e.g., randomly selected) epochs, the audio data from all microphones can be used as input to the ESLM 132. In some implementations, dropout can be implemented as part of the application of a suitable loss function, such as the cross-entropy loss function, by weighing the contributions from one or more microphones having small weights that can be, for example, zero. Dropout techniques can have multiple advantages. In particular, such techniques train the ESLM 132 for situations where one or more microphones and / or microphone arrays stop functioning. Further, dropout techniques tune the ESLM 132 to be more efficient in the use of the input data of individual microphones / arrays.

[0054] In some implementations, the microphone array used to generate the audio frames can have three or more synchronous microphones per array, and the microphones in the array are positioned at the vertices of a triangle rather than along a single line. Such a microphone array can completely clarify the audio arriving from different directions. As a result, a single time-synchronized array with three or more microphones can be used to generate the audio frames used as input to the ESLM132. In some implementations, deploying a two-microphone array may be more cost-effective. The two-microphone array may not be able to clarify the sound waves arriving along the opposite direction of the center line perpendicular to the line connecting the two microphones. In such implementations, even if not time-synchronized between different arrays, two or more two-microphone arrays positioned along different (e.g., perpendicular) lines can provide the ESLM132 with sufficient audio information to clarify the arrival direction of the sound waves.

[0055] Figure 3 shows an exemplary data flow 300 for audio and perception data processing for behavior prediction using an emergency voice localization model for driving route selection in an autonomous driving application, according to some implementations of the present disclosure. Predictions generated by the ESLM132 (e.g., disclosed in conjunction with FIGS. 2A-2B above), e.g., the position probability map P(D,θ;τ), and / or the velocity probability map P(V D ,V θ; Different timestamps τ (such as those corresponding to the start, center, or end of a sliding window used to generate a given set of audio frames) can be provided to the planner module 134. Further, the planner module 134 can receive live (electromagnetic) sensing data collected by the AV sensing system. For example, the sensing data can include one, some, or all of Lidar data 302 (generated by Lidar 112 in FIG. 1), radar data 304 (generated by radar 114 in FIG. 1), and / or camera data 306 (generated by camera(s) 118 in FIG. 1). The sensing data can identify the visible portion of the environment where the presence of an EV can be excluded, based on, for example, the absence of blinking lights on a vehicle located within the visible portion, the visual appearance of a vehicle that does not match the appearance of a known EV, and / or the like.

[0056] FIG. 4 shows an exemplary environment 400 of an autonomous vehicle 402 that can deploy an emergency voice location identification model for driving route selection in an autonomous driving application, according to some implementations of the present disclosure. As shown in FIG. 4, the AV 402 detects the presence of an emergency sound emitted by an emergency vehicle within the environment 400. The planner module 134 (see FIG. 3) can identify blocked regions of the environment, such as regions 404, 406, and 408 blocked by various objects 410 (such as buildings, structures, trees and plants, road signs, etc.). In response, the planner module 134 can determine (based on an inspection of the visible region of the environment 400) that the EV is currently located within the blocked regions 404, 406, and 408. The planner module 134 can apply a position probability map P(D,θ;t j ) to the blocked regions 404, 406, and 408. Thereby, the planner module 134 can limit the possible positions of the EV to candidate regions 420 that represent the intersection of the blocked regions with a portion of the environment 400 where the position probability is some threshold probability or greater. P(D,θ;t j )≧P T . The threshold probability P Tcan be empirically set to, for example, 30%, 50%, 65%, or some other value. The intersection area 420 is illustrated with a uniform shadow. For example, based on the probability map P(D,θ;t j ), the planner module 134 can determine that the emergency sound is coming from the right side of the AV and not from the left side of the AV, and thus can exclude the occlusion area 404. Accordingly, the planner module 134 can limit the possible positions of the EV to the front candidate area 420-F and the rear candidate area 420-B.

[0057] The BP module can further exclude at least a part of the occluded area portion based on the road layout information 124 by excluding non-drivable areas of the occlusion area, such as areas occupied by buildings, or structures, enclosures, impassable areas, and / or land of the same kind. For example, only the portions of the candidate areas 420-F and 420-B that overlap with the road (obtained using the road layout information 124) can be regarded as candidate positions where the EV is possible. As shown in FIG. 4, next, the simulated EV 430 can be placed within the candidate areas 420-A and 420-B, but other positions can be excluded. For example, the simulated EV 434 can be excluded as a low-probability position, and the simulated EV 438 can be excluded as an off-road position.

[0058] The target space can include a region that is (i) enclosed, (ii) drivable, and (iii) characterized by at least a minimum (threshold) probability P(D,θ;τ) of having an EV thereon. This target space represents a hypothesis space of possible EV positions (for a given time τ). Then, a series of simulated EVs can be placed at various positions in the target space, such as simulated EV 430, simulated EV 432, etc., and several simulated speeds can be further given to each simulated EV. The simulated speeds can be subject to one or more constraints. One constraint can be due to a specific road layout (obtained as part of road layout information 124) and can limit the movement of the simulated EV into the drivable space. The drivable space can include drivable off-road surfaces such as roads, lanes, parking lots, and sidewalks. The drivable space can exclude driving routes that cross physical obstacles such as buildings, structures (e.g., bus stops), utility poles, trees, road signs, patches of grass, curves exceeding a certain height (e.g., 30 cm), and various other obstacles that the EV cannot physically overcome. The road layout constraint does not need to include the condition that the simulated EV is on the correct side of the road since the EV can move against traffic. Another constraint can be due to the calculated probability map P(V D ,V θ ;τ) and can have various values V D ,V θ that are predicted with a probability below some (e.g., empirically set) threshold and are excluded from the simulation. For example, if the speed probability map P(V D ,V θ ;τ) overlaid on the position probability P(D,θ;τ) indicates that the EV is moving away from the AV in a particular region (e.g., an enclosed region), such regions can be excluded from further simulation. On the other hand, the speed of the simulated EV does not need to be constrained by the maximum legal speed of the environment (or the average speed of other vehicles in that environment).

[0059] In some implementations, the simulation can be performed using one or more Monte Carlo techniques. For example, a plurality of simulated EVs can be selected (sampled, created) and placed within the target space, for example, subject to the above conditions (i)-(iii). Sampling can be based on the calculated position probability P(D,θ;τ), and the simulated EVs can be assigned velocities based on the calculated velocity probability P(V D ,V θ ;τ). The simulated EVs can then follow some simulated trajectories for a specific target time, for example, 1 second, 2 seconds, 3 seconds, or some other time. In one exemplary implementation, the trajectories can be simulated based on the assumption that the simulated EVs maintain the assigned velocities during the target time. Next, the planner module 134 can identify a set of positions, called the EV active region, where the simulated EVs can be at the target time or later (or before the expiration of the target time). In one embodiment, the EV active region can then pass through the AVCS (see, for example, FIG. 3, AVCS 140), and the AVCS can control the driving path of the AV to avoid the AV entering (or approaching) the EV active region.

[0060] Controlling the driving route of the AV includes maintaining the current driving route of the AV if the current driving route is maintained (e.g., if it is determined that the AV will not enter the EV active area within the target time and the current driving route is maintained), or modifying the current driving route of the AV by braking (e.g., moving to the side of the road), moving away from the EV active area (e.g., turning the AV onto a side path or changing to another route), and / or similar things (if the AV is determined to enter the EV active area during the target time or be within a certain (empirically set) distance from the EV active area, e.g., 20m). For example, the simulation performed by the BP module (e.g., the planner module 134 in FIG. 3) can determine that the simulated EV432 is behind the AV402 and the AV402 is moving away from the rear EV active area 440 (associated with the possible trajectory of the simulated EV432). Next, the planner module can determine that the EV traveling on the road behind the AV402 will not interfere with the trajectory of the AV402. Next, the AVCS can re-evaluate the EV active area dynamically if additional audio data and / or other sensing data indicates that the EV is making a correct turn and capturing the AV402, and the modified EV active area 440 can extend to the new position of the AV402), and maintain the current driving route of the AV402 until at least a later time. At the same time, the simulation performed by the planner module can determine that the simulated EV430 is in front of the AV402 and both the simulated EV430 and the AV402 are moving towards the same intersection that overlaps with the front EV active area 442 (associated with the possible trajectory of the simulated EV430), and which AV402 will enter first. Next, the planner module can provide an instruction to the AVCS to decelerate the AV402 to ensure that the AV402 does not reach the intersection before the simulated EV430 passes.Accordingly, the Planner Module 134 can provide an instruction to the AVCS to stop AV402 (or otherwise, the AVCS can decelerate / stop AV402).

[0061] In some implementations, to respond to the potential presence of an EV as soon as possible and yield the road to the EV located behind AV402 in a timely manner (e.g., by pulling over), the stop / pull-over decision may be based on a comparison of the estimated distance to the EV with a predetermined threshold distance D and / or a threshold time T. For example, if the distance to the simulated EV (e.g., based on position probability) is less than D (e.g., 75m), or if the simulated EV (e.g., based on both position probability and speed probability) will reach AV402 within time T (e.g., 5 seconds), the AVCS of AV402 will pull over, stop, and / or perform similar actions. This can be implemented even in situations where the simulated EV is not directly behind AV402 but is still separated from AV402 by one or more turns, provided that there is a perceived possibility that the EV can capture the AV within time T.

[0062] The above operations can be repeated at one or more subsequent times τ, e.g., every second, 2 seconds, half a second, or any other set time interval. Slide window techniques can be used to obtain audio frames for subsequent times τ. Then, a new set of prediction 220 (as well as a new set of Lidar / radar / camera data and updated road layout information) can be used by the BP module to perform a new set of simulations, update the expected EV active area, and accordingly adjust the driving route of the AV.

[0063] In some implementations, the road layout information 124 is not used (e.g., may not be available). In such cases, the planner module 134 does not know the exact road layout in the blocked areas of the driving environment. Accordingly, the planner module 134 can assume that the entire blocked area is drivable and that simulated EVs can be placed anywhere within the blocked area, which may further be subject to the constraints provided by the probability map generated by the ESLM 132.

[0064] In some examples, for instance, due to the presence of noise in the environment or as a result of multiple reflections of sound from buildings and other objects, as part of the prediction 220 and characterizing the reliability of the ESLM 132 within the predicted probability map, the reliability score C output by the ESLM 132 may fall below a certain minimum reliability CMIN. In some implementations, the reliability score C is the overall reliability score calculated as a combination of individual reliability scores C D 、C θ 、C V output separately by the ESLM 132 for individual predictions. For example, the reliability score C can be the average (e.g., arithmetic mean, geometric mean, and / or the like) or weighted average of the individual reliability scores. In low-reliability cases, C < C MIN 、the planner module 134 can perform simulations with EVs placed throughout the blocked areas not visible to the perception system (which may be subject to the constraints of the road layout if available). For example, additional simulated EVs 434, 436, and 438 can then be sampled along with the simulated EVs 430 and 432.

[0065] In some implementations, other techniques for operating the AV in the presence of the EV's audio can be used in addition to the techniques disclosed above. More specifically, such additional techniques can include stopping the AV (or decelerating the AV to some set speed, such as 5 mph or the like) if the most likely distance from the EV, as estimated based on the position probability map output by the ESLM132, is less than a minimum distance, such as 75 m, 100 m, and / or the like. The additional techniques can include, for example, stopping or decelerating the AV each time the EV sound source approaches, based on position probability maps generated two or more times in succession. The additional techniques can include stopping the AV or stopping the AV if the emergency sound source cannot be detected using the direct line-of-sight electromagnetic sensing data.

[0066] Figure 5 shows a method 500 for controlling a driving route of a vehicle in the presence of the voice of an emergency vehicle, according to some implementations of the present disclosure. A processing device having one or more processing units (CPUs) and / or one or more graphics processing units (GPUs), and a memory device communicatively coupled to the CPU(s) and / or GPU can implement each of the method 500 and / or its individual functions, routines, subroutines, or operations. The method 500 can be directed to vehicle systems and components. In some implementations, the vehicle can be an autonomous vehicle (AV) such as the AV100 of FIG. 1. In some implementations, the vehicle can be a driver-operated vehicle equipped with a driving assistance system, such as a level 2 or level 3 driving assistance system, that provides limited assistance to a particular vehicle system (e.g., systems such as steering, braking, acceleration, etc.) or operates under limited driving conditions (e.g., highway driving). The processing device executing the method 500 can execute instructions issued by various components of the perception and planning system 130 of FIG. 1, such as the ESLM 132, the planner module 134, and / or the like. The method 500 can be used to improve the performance of the autonomous vehicle control system 140. In a particular implementation, a single processing thread can execute each of the method 500. Alternatively, two or more processing threads can execute each of the method 500, and each thread can execute one or more individual functions, routines, subroutines, or operations of the method. In an exemplary embodiment, the processing threads executing the method 500 can be synchronized (e.g., using semaphores, critical sections, and / or other thread synchronization mechanisms). Alternatively, the processing threads executing the method 500 can execute asynchronously with respect to each other. The various operations of the method 500 can be performed in a different order (e.g., reversed) compared to the order shown in FIG. 5. Some operations of the method 500 can be performed simultaneously with other operations. Some operations can be optional.

[0067] In block 510, method 500 may include obtaining an audio recording that includes audio emitted by an emergency vehicle (EV) using one or more audio detectors of the vehicle. The audio recording can include a plurality of files, for example, separate files recorded by each audio detector (microphone). In block 520, method 500 can continue by applying an audio localization (SL) model (e.g., ESLM132) to the audio recording to obtain an SL output (e.g., prediction 220). The SL output can include a first map of the possible locations of the EV within the driving environment of the vehicle. For example, the first map can include the probabilistic occupancy of a plurality of locations in the driving environment of the vehicle by the EV. In some implementations, the probabilistic occupancy can include the probability that the EV occupies various locations

Number

Number

[0068] In some implementations, the SL output can include a second map of the possible speeds of the EV. For example, the second map can include a plurality of probabilities, each of which characterizes the likelihood that the EV is moving at each of a plurality of speeds

Number

[0069] In some implementations, the simulated trajectories of one or more simulated EVs include the operations indicated by the arrow portion at the top of FIG. 5. More specifically, applying the SL model to the voice recording may include, at block 522, obtaining a first spectrogram for the voice recording and a second spectrogram for the voice recording. The first spectrogram (e.g., F frame 206) may be obtained using a first sampling window size and a first sampling rate. The second spectrogram (e.g., T frame 208) may be obtained using a second sampling window size and a second sampling rate. The first sampling window size may be larger than the second sampling window size, and the first sampling rate may be smaller than the second sampling rate.

[0070] At block 524, the first spectrogram and the second spectrogram may be processed using the SL model. For example, the first spectrogram may be processed using a first neural network (e.g., F network 260 of FIG. 2B) that outputs a first embedding (e.g., intermediate embedding 262 of FIG. 2B). Similarly, the second spectrogram may be processed using a second neural network (e.g., T network 208 of FIG. 2B) that generates a second embedding (e.g., intermediate embedding 266 of FIG. 2B). At block 526, method 500 may continue to obtain a fused embedding (e.g., fused embedding 270 of FIG. 2B) by fusing the first embedding and the second embedding. At block 528, method 500 may process the fused embedding using a third neural network (e.g., fused network 280 of FIG. 2B) to obtain an SL output.

[0071] In block 530, method 500 may include simulating the trajectories of one or more simulated EVs in the vehicle's operating environment using the SL output. In some implementations, the simulated trajectories of the one or more simulated EVs include the operations shown in the arrow portion at the bottom of FIG. 5. More specifically, in block 532, method 500 may identify one or more occluded regions of the vehicle's operating environment using electromagnetic sensor data collected by the vehicle's sensing system. The electromagnetic sensor data may be or include Lidar data, radar data, camera data, and / or the like. In block 534, method 500 may include identifying drivable regions within the one or more occluded regions using road layout information.

[0072] In block 536, method 500 may continue to simulate the positions of the one or more simulated EVs. In some implementations, the selected positions of the one or more simulated EVs may be selected within the identified drivable regions and / or within the one or more occluded regions. In some implementations, the SL output may further include a confidence score. In some implementations, in response to the confidence score being less than a threshold confidence score, method 500 may include ignoring a first map when selecting the positions of the one or more simulated EVs. In some implementations, the selection of the positions of the one or more simulated EVs may be performed using the first map. In block 538, method 500 may include selecting the speeds of the one or more simulated EVs using a second map. In block 539, method 500 may continue to calculate the simulated trajectories using the selected positions and the selected speeds.

[0073] In block 540, method 500 may continue to modify the driving route of the vehicle in response to proximity to the driving routes of one or more vehicles on the simulated track. Bringing one or more of the simulated tracks close to the driving route of the vehicle may mean that the distance from a vehicle planned to follow the driving route over a given time to one or more of the simulated tracks is less than a given distance (e.g., less than 30 m, less than 50 m, etc.).

[0074] In some implementations, the SL model may be trained using a plurality of training audio recordings obtained by a plurality of audio sensors. During at least one training epoch, one or more of the plurality of training audio recordings can be replaced with a null input to the SL model.

[0075] FIG. 6 shows a block diagram of an exemplary computer device 600 that can support an audio-based emergency vehicle detection, localization, and tracking pipeline that can be used as part of a perception and planning system for an autonomous vehicle, according to some implementations of the present disclosure. Examples of computer device 600 can be connected to other computer devices within a LAN, intranet, extranet, and / or the Internet. Computer device 600 can operate with the capabilities of a server in a client-server network environment. Computer device 600 can be a personal computer (PC), a set-top box (STB), a server, a network router, a switch or bridge, or any device capable of executing a set (sequentially or otherwise) of instructions that specify actions to be taken by that device. Further, although only a single example of a computer device is shown, the term "computer" shall also be considered to include any group of computers that execute, individually or jointly, a set (or sets) of instructions to perform any one or more of the methods discussed herein.

[0076] An embodiment of the computer device 600 can include a processing device 602 (also referred to as a processor or CPU), a main memory 604 (such as a dynamic random access memory (DRAM) like a read-only memory (ROM), flash memory, synchronous DRAM (SDRAM), etc.), a static memory 606 (such as flash memory, static random access memory (SRAM), etc.), and a secondary memory (such as a data storage device 618), which can communicate with each other via a bus 630.

[0077] The processing device 602 (which may include logical processing 603) represents one or more general-purpose processing devices such as a microprocessor, a central processing unit, or the like. More specifically, the processing device 602 can be a complex instruction set computer (CISC) microprocessor, a reduced instruction set computer (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor that executes other instruction sets, or a processor that executes a combination of instruction sets. The processing device 602 can also be one or more dedicated processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. According to one or more aspects of the present disclosure, the processing device 602 can be configured to execute instructions for executing a method 500 for controlling a driving route of a vehicle in the presence of the voice of an emergency vehicle.

[0078] An embodiment of the computer device 600 can further include a network interface device 608 that can be communicatively coupled to a network 620. An embodiment of the computer device 600 can further include a video display 610 (such as a liquid crystal display (LCD), a touch screen, or a cathode ray tube (CRT)), an alphanumeric input device 612 (such as a keyboard), a cursor control device 614 (such as a mouse), and an audio signal generating device 616 (such as a speaker).

[0079] The data storage device 618 can include a computer-readable storage medium (or more specifically, a non-transitory computer-readable storage medium) 628 in which one or more sets of executable instructions 622 are stored. According to one or more aspects of the present disclosure, the executable instructions 622 can include executable instructions that execute a method 500 for controlling a driving route of a vehicle in the presence of the voice of an emergency vehicle.

[0080] The executable instructions 622 can also be present, in whole or at least in part, within the main memory 604 and / or within the processing device 602 during its execution in the computer device 600, for example. The main memory 604 and the processing device 602 also constitute a computer-readable storage medium. The executable instructions 622 can further be further transmitted or received over a network via the network interface device 608.

[0081] The computer-readable storage medium 628 is shown in FIG. 6 as a single medium, but the term "computer-readable storage medium" should be considered to include a single medium or a plurality of media (e.g., a centralized or distributed database, and / or associated caches and servers) that store one or more sets of operating instructions. The term "computer-readable storage medium" should also be considered to include any medium that has the ability to store or encode a set of instructions for machine execution that cause a machine to execute any one or more of the methods described herein. Thus, the term "computer-readable storage medium" should be considered to include, but not be limited to, solid-state memory, as well as optical and magnetic media.

[0082] Some portions of the detailed descriptions above are presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing art to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, understood to be a self-consistent sequence of steps that bring about a desired result. The steps are those requiring physical operations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated. Referring to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like is primarily for reasons of common usage and has sometimes proven convenient.

[0083] However, it should be noted that all of these terms and similar terms are associated with appropriate physical quantities and are nothing more than convenient labels applied to these quantities. Unless otherwise specified, as will be apparent from the following discussion, throughout the description, discussions using terms such as "specify", "determine", "store", "adjust", "cause", "return", "compare", "generate", "stop", "load", "copy", "input", "replace", "perform", or the like are understood to refer to the actions and processes of a computer system, or similar electronic computing device, that operate on and transform data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the computer system memory or registers or other such information storage, transmission, or display devices.

[0084] Examples of the present disclosure also relate to an apparatus for implementing the methods described herein. This apparatus can be specially constructed for the required purpose or can be a general-purpose computer system selectively programmed by a computer program stored within the computer system. Such computer programs can be stored on a computer-readable storage medium, which can include any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic disk storage media, optical storage media, flash memory devices, other types of storage media accessible by machines, or any type of media suitable for storing electronic instructions, but are not limited thereto.

[0085] The methods and displays presented herein are not inherently related to any particular computer or other device. A variety of general-purpose systems can be used with the programs according to the teachings herein, or it may prove convenient to construct more specialized devices to perform the required method steps. The required structure for these various systems will appear as described in the following description. Additionally, the scope of the present disclosure is not limited to any particular programming language. It will be understood that various programming languages can be used to implement the teachings of the present disclosure.

[0086] Of course, the above description is intended to be illustrative and not restrictive. Many other implementation examples will be apparent to those skilled in the art upon reading and understanding the above description. Although the present disclosure describes specific examples, it is recognized that the systems and methods of the present disclosure are not limited to the examples described herein and can be modified and implemented within the scope of the appended claims. Therefore, the specification and drawings should be considered in an illustrative sense rather than a restrictive sense. Accordingly, the scope of the present disclosure should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

1. 1. A method comprising: obtaining an audio recording that includes audio emitted by an emergency vehicle (EV) using one or more audio detectors in the vehicle; applying a voice localization (SL) model to the voice recording using a processing device to generate a SL output, the SL output including a first map of likely locations of the EV within a driving environment of the vehicle; simulating, by the processing device, trajectories of one or more simulated EVs within the driving environment of the vehicle using the SL output; modifying the driving path of the vehicle in response to proximity of one or more of the simulated trajectories to a driving path of the vehicle; A method comprising:

2. The method of claim 1 , wherein the first map includes probabilistic occupancy by the EV of a plurality of locations of the driving environment of the vehicle.

3. The method of claim 2 , wherein the SL output includes a second map of possible speeds of the EV.

4. the second map includes a plurality of probabilities, each of the plurality of probabilities characterizing a likelihood that the electric vehicle is traveling at a respective one of a plurality of speeds; and simulating a trajectory of the one or more simulated electric vehicles. selecting a location for the one or more simulated EVs using the first map; and selecting a speed of the one or more simulated EVs using the second map; and calculating the simulated trajectory using the selected positions and the selected velocities; The method of claim 3 , comprising:

5. applying the SL model to the audio recording; obtaining a first spectrogram for the audio recording, the first spectrogram being obtained using a first sampling window size and a first sampling rate; acquiring a second spectrogram for the audio recording, the second spectrogram being acquired using a second sampling window size and a second sampling rate, the first sampling window size being greater than the second sampling window size and the first sampling rate being less than the second sampling rate; processing the first spectrogram and the second spectrogram using the SL model; and The method of claim 1 , comprising:

6. Processing the first spectrogram and the second spectrogram includes: processing the first spectrogram using a first neural network to obtain a first embedding; processing the second spectrogram using a second neural network to obtain a second embedding; fusing the first embedding and the second embedding to obtain a fused embedding; processing the fused embedding using a third neural network to obtain the SL output; The method of claim 5 , comprising:

7. Simulating trajectories of one or more simulated EVs includes: identifying one or more occluded regions of the driving environment of the vehicle using electromagnetic sensor data collected by a sensing system of the vehicle, the electromagnetic sensor data comprising: Lidar data, Radar data, or including one or more of the camera data; selecting a location for the one or more simulated EVs within the one or more occlusion regions; The method of claim 1 , comprising:

8. using road layout information to identify a drivable area within the one or more occluded areas; The method of claim 7 , wherein the selected locations of the one or more simulated EVs are selected within the identified drivable area.

9. The SL output further comprises a confidence score, and the method further comprises:

8. The method of claim 7, further comprising: responsive to the reliability score being less than a threshold reliability score, ignoring the first map in selecting the location for the one or more simulated EVs.

10. 2. The method of claim 1 , wherein the proximity of the one or more of the simulated trajectories to the driving path of the vehicle comprises a distance from the vehicle planned to follow the driving path for a predetermined time to the one or more of the simulated trajectories that is less than a predetermined distance.

11. 2. The method of claim 1, wherein the SL model is trained using a plurality of training speech recordings acquired by a plurality of speech sensors, and during at least one training epoch, one or more training speech recordings of the plurality of training speech recordings are replaced with a null input to the SL model.

12. 1. A system comprising:

1. A sensing system for a vehicle, comprising: a sensing system including one or more audio detectors configured to obtain an audio recording including audio emitted by an emergency vehicle (EV); A perception system for the vehicle, applying a voice localization (SL) model to the voice recording to obtain a SL output, the SL output including a first map of likely locations of the EV within a driving environment of the vehicle; simulating trajectories of one or more simulated EVs within the driving environment of the vehicle using the SL output; modifying the driving path of the vehicle in response to proximity of one or more of the simulated trajectories to a driving path of the vehicle; A perception system configured to: A system equipped with

13. 13. The system of claim 12, wherein the first map includes probabilistic occupancy by the EV of a plurality of locations of the driving environment of the vehicle, and the SL output includes a second map of possible speeds of the EV.

14. the second map includes a plurality of probabilities, each of the plurality of probabilities characterizing a likelihood that the electric vehicle is traveling at a respective one of a plurality of velocities; and to simulate a trajectory of the one or more simulated electric vehicles, the perception system: selecting a location for the one or more simulated EVs using the first map; and selecting a speed of the one or more simulated EVs using the second map; and calculating the simulated trajectory using the selected positions and the selected velocities; The system of claim 13 configured to:

15. To apply the SL model to the audio recording, the perception system comprises: obtaining a first spectrogram for the audio recording, the first spectrogram being obtained using a first sampling window size and a first sampling rate; acquiring a second spectrogram for the audio recording, the second spectrogram being acquired using a second sampling window size and a second sampling rate, the first sampling window size being greater than the second sampling window size and the first sampling rate being less than the second sampling rate; processing the first spectrogram and the second spectrogram using the SL model; and The system of claim 12 configured to:

16. To simulate a trajectory of one or more simulated electric vehicles, the perception system includes: identifying one or more occluded regions of the driving environment of the vehicle using electromagnetic sensor data collected by a sensing system of the vehicle, the electromagnetic sensor data comprising: Lidar data, Radar data, or including one or more of the camera data; selecting a location for the one or more simulated EVs within the one or more occlusion regions; The system of claim 12 configured to:

17. the perception system further comprising: configured to identify a drivable area within the one or more occluded areas using road layout information; The system of claim 16 , wherein the selected locations of the one or more simulated EVs are selected within the identified drivable area.

18. The SL output further comprises a confidence score, and the perception system further comprises:

17. The system of claim 16, configured to ignore the first map in selecting the locations of the one or more simulated EVs in response to the reliability score being less than a threshold reliability score.

19. 13. The system of claim 12, wherein the SL model is trained using a plurality of training speech recordings acquired by a plurality of speech sensors, and during at least one training epoch, one or more training speech recordings of the plurality of training speech recordings are replaced with a null input to the SL model.

20. A non-transitory computer readable storage medium that, when executed by a processing device, causes the processing device to perform operations including: obtaining an audio recording that includes audio emitted by an emergency vehicle (EV) using one or more audio detectors in the vehicle; applying a voice localization (SL) model to the voice recording to generate a SL output, the SL output including a first map of likely locations of the EV within a driving environment of the vehicle; simulating trajectories of one or more simulated EVs within the driving environment of the vehicle using the SL output; modifying the driving path of the vehicle in response to proximity of one or more of the simulated trajectories to a driving path of the vehicle; A non-transitory computer-readable storage medium having stored thereon instructions for causing a