Systems, methods, and computer-readable media for sound source localization
By installing a microphone array on top of the autonomous vehicle and synchronizing it with an internal clock, the difference in timestamps is used to locate emergency vehicles, solving the problem that autonomous vehicles have difficulty detecting the location of emergency vehicles and achieving accurate positioning and safe operation in sensor blind spots.
Patent Information
- Application Number
- CN202180051347.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-21
- Filing Date
- 2021-07-21
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-07-21
AI Technical Summary
Autonomous vehicles struggle to effectively detect and respond to the location of emergency vehicles, especially when the emergency vehicle's flashing lights are not visible, which affects their maneuverability.
A microphone array is installed on top of the autonomous vehicle. Multiple microphones are used to generate timestamped audio frames. The internal clock is synchronized by a computing system, and the timestamp differences are analyzed to locate the sound source and generate control commands.
It can accurately locate emergency vehicles even when they are undetectable by sensors, improving the responsiveness of autonomous vehicles and ensuring safe operation.
Smart Images

Figure CN115943643B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit of U.S. Patent Application No. 16 / 999,830, filed August 21, 2020, the contents of which are incorporated herein by reference. Background Technology
[0003] Autonomous vehicles, or vehicles operating in autonomous mode, may encounter scenarios where they may need to take rapid maneuvers based on unexpected changes in the surrounding environment. As a non-limiting example, if an emergency vehicle activates its siren, an autonomous vehicle may responsively veer to the side of the road and come to a stop.
[0004] Typically, autonomous vehicles use sensors to determine their surroundings. For example, autonomous vehicles may use optical detection and ranging (LIDAR) devices, radio detection and ranging (RADAR) devices, and / or cameras to capture environmental data around them. However, certain aspects of objects in the environment surrounding an autonomous vehicle may not be easily detected by such sensors. As a non-limiting example, if the flashing lights of an approaching emergency vehicle are not visible, the autonomous vehicle may not be able to easily detect the location of the emergency vehicle. As a result, the autonomous vehicle's ability to perform maneuvers in response to an approaching emergency vehicle may be adversely affected. Summary of the Invention
[0005] This disclosure generally relates to using a microphone array positioned on the roof of an autonomous vehicle to locate the source of a detected sound, such as a siren. In an example embodiment, the microphones of the microphone array generate timestamped audio frames of the detected sound. A computing system analyzes the timestamped audio frames from each microphone array to locate the source of the detected sound. As a non-limiting example, using timestamped audio frames from a first microphone array, the computing system may perform a first operation to determine the time difference of arrival of the detected sound and calculate the azimuth of the source based on that difference. As another non-limiting example, using timestamped audio frames from a second microphone array, the computing system may perform a second operation to determine the time difference of arrival of the detected sound and calculate the azimuth of the source based on that difference. For additional microphone arrays positioned on the roof, timestamped audio frames can be used to perform similar operations. The computing system can further locate or verify the position of a specific source by comparing or analyzing the results of the first and second operations.
[0006] It should be understood that the techniques described herein can be implemented using various numbers of microphone arrays. For example, the techniques described herein can be implemented using two microphone arrays, three microphone arrays, four microphone arrays, etc. Furthermore, the techniques described herein can be implemented using various microphone array configurations. For example, the techniques described herein can be implemented using microphone arrays with two microphones, microphone arrays with three microphones, microphone arrays with four microphones, microphone arrays with microphone rings, etc. As used herein, each microphone array is associated with a microphone unit or module having an internal clock for generating timestamps.
[0007] In a first aspect, the system includes a first microphone array disposed at a first location on an autonomous vehicle. The first microphone array includes a first plurality of microphones, and each of the first plurality of microphones is capable of capturing a specific sound from a specific source to generate a corresponding audio frame with a timestamp based on a first internal clock associated with the first microphone array. The timestamp indicates the time at which the corresponding microphone among the first plurality of microphones captured the specific sound. The system also includes a second microphone array disposed at a second location on the autonomous vehicle. The second microphone array includes a second plurality of microphones, and each of the second plurality of microphones is capable of capturing a specific sound from a specific source to generate a corresponding audio frame with a timestamp based on a second internal clock associated with the second microphone array. The timestamp indicates the time at which the corresponding microphone among the second plurality of microphones captured the specific sound. The system further includes a processor configured to synchronize the first internal clock with the second internal clock. The processor is also configured to perform a first operation to locate a specific source relative to the autonomous vehicle based on the timestamp of the audio frame generated by the first plurality of microphones. The processor is also configured to perform a second operation to locate a specific source relative to the autonomous vehicle based on the timestamp of the audio frame generated by the second plurality of microphones. The processor is also configured to determine the position of a specific source relative to the autonomous vehicle based on the first and second operations. The processor is further configured to generate commands to manipulate the autonomous vehicle based on the position of the specific source relative to the autonomous vehicle.
[0008] In a second aspect, the method includes synchronizing at a processor a first internal clock associated with a first microphone array positioned at a first location on the autonomous vehicle and a second internal clock associated with a second microphone array positioned at a second location on the autonomous vehicle. The method also includes receiving audio frames and corresponding timestamps from each microphone in the first microphone array. Each audio frame is generated in response to a corresponding microphone in the first microphone array capturing a specific sound from a specific source, and each timestamp is generated using the first internal clock and indicates the time at which the corresponding microphone in the first microphone array captured the specific sound. The method further includes receiving audio frames and corresponding timestamps from each microphone in the second microphone array. Each audio frame is generated in response to a corresponding microphone in the second microphone array capturing a specific sound from a specific source, and each timestamp is generated using the second internal clock and indicates the time at which the corresponding microphone in the second microphone array captured the specific sound. The method also includes performing a first operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the microphones in the first microphone array. The method further includes performing a second operation to locate the specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the microphones in the second microphone array. The method also includes determining the position of the specific source relative to the autonomous vehicle based on the first and second operations. The method also includes generating commands to manipulate the autonomous vehicle based on the position of a specific source relative to the autonomous vehicle.
[0009] In a third aspect, the non-transitory computer-readable medium stores instructions executable by a computing device to cause the computing device to perform a function. This function includes synchronizing a first internal clock associated with a first microphone array positioned at a first location on the autonomous vehicle and a second internal clock associated with a second microphone array positioned at a second location on the autonomous vehicle. The function also includes receiving audio frames and corresponding timestamps from each microphone in the first microphone array. Each audio frame is generated in response to a corresponding microphone in the first microphone array capturing a specific sound from a specific source, and each timestamp is generated using the first internal clock and indicates the time when the corresponding microphone in the first microphone array captured the specific sound. The function also includes receiving audio frames and corresponding timestamps from each microphone in the second microphone array. Each audio frame is generated in response to a corresponding microphone in the second microphone array capturing a specific sound from a specific source, and each timestamp is generated using the second internal clock and indicates the time when the corresponding microphone in the second microphone array captured the specific sound. The function further includes performing a first operation to locate a specific source relative to the autonomous vehicle based on the timestamps of the audio frames generated by the microphones in the first microphone array. The function also includes performing a second operation to locate a specific source relative to the autonomous vehicle based on the timestamps of the audio frames generated by the microphones in the second microphone array. This functionality also includes determining the position of a specific source relative to the autonomous vehicle based on the first and second operations. Furthermore, it includes generating commands to manipulate the autonomous vehicle based on the position of the specific source relative to the autonomous vehicle.
[0010] Other aspects, embodiments, and implementation methods will become clear to those skilled in the art upon reading the following detailed description and with appropriate reference to the accompanying drawings. Attached Figure Description
[0011] Figure 1 This is a functional diagram illustrating the components of an autonomous vehicle according to an example embodiment.
[0012] Figure 2 This is a diagram of a microphone array according to an example embodiment.
[0013] Figure 3 A diagram depicts a microphone coupled to a microphone board according to an example embodiment.
[0014] Figure 4 This is another functional diagram illustrating the components of an autonomous vehicle according to an example embodiment.
[0015] Figure 5 Diagrams depicting different top positions of the coupled microphone unit according to an example embodiment.
[0016] Figure 6 This is a flowchart of a method according to an example embodiment.
[0017] Figure 7 This is a flowchart of another method according to an example embodiment. Detailed Implementation
[0018] This document describes example methods, devices, and systems. It should be understood that the terms "example" and "exemplary" as used herein mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as an "example" or "exemplary" is not necessarily to be construed as being more preferred or advantageous than other embodiments or features. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.
[0019] Therefore, the exemplary embodiments described herein are not intended to be limiting. As generally described herein and shown in the accompanying drawings, aspects of this disclosure can be arranged, replaced, combined, separated, and designed in a variety of different configurations, all of which are contemplated herein.
[0020] Furthermore, unless the context otherwise suggests, the features shown in each figure can be used in combination with each other. Therefore, the figures should generally be considered as aspects of one or more overall embodiments, and it should be understood that not all features shown are necessary for every embodiment.
[0021] I. Overview
[0022] This disclosure generally relates to using microphone arrays mounted on an autonomous vehicle (e.g., mounted on top of the autonomous vehicle) to locate the source of detected sound. For example, multiple microphone arrays may be coupled to the top of the autonomous vehicle at different locations. To illustrate, a left microphone array may be coupled to the left side of the top to capture sound from the left side of the autonomous vehicle, a right microphone array may be coupled to the right side of the top to capture sound from the right side of the autonomous vehicle, and a rear microphone array may be coupled to the rear of the top to capture sound from the rear of the autonomous vehicle. Each microphone array may be associated with a separate internal clock for generating a timestamp indicating the time when the corresponding microphone array captured sound. A computing system within the autonomous vehicle sends a synchronization signal to each internal clock, synchronizing the internal clocks associated with each microphone array coupled to the top. According to one embodiment, synchronization may be performed using a Precise Time Protocol (PTP) operation.
[0023] The microphone array is used in conjunction with a computing system to detect the position of a sound source relative to the autonomous vehicle. As a non-limiting example, if an ambulance is approaching the autonomous vehicle from the left rear, the left and rear microphone arrays can detect the ambulance siren sound before the right microphone array detects it. Partly, the left and rear microphone arrays can detect the ambulance siren sound before the right microphone array detects it because the left and rear microphone arrays are closer to the ambulance and slightly angled towards it. As a result, the left and rear microphone arrays can generate an audio frame representing the captured ambulance siren sound before the right microphone array generates one. Therefore, if the internal clocks are synchronized, the timestamp on the audio frame generated by the left and rear microphone arrays (for the captured ambulance siren sound) should indicate an earlier time than the timestamp on the audio frame generated by the right microphone array (for the captured ambulance siren sound). The computing system can read timestamps to determine which microphone arrays are closer to the ambulance, thus determining the ambulance's approximate position relative to the autonomous vehicle. For example, the computing system can determine that the ambulance is located to the left rear of the autonomous vehicle.
[0024] In response to determining the location of the ambulance, the computing system can generate commands to manipulate the autonomous vehicle. As a non-limiting example, the computing system could generate commands to steer the autonomous vehicle to the right side of the road, commands to reduce the speed of the autonomous vehicle, etc.
[0025] Therefore, a microphone array coupled to the top of the autonomous vehicle can be used to locate the ambulance siren sound and determine the ambulance's position relative to the autonomous vehicle. Furthermore, the microphone array can be used to determine the ambulance's location in scenarios where it is not detected by other sensors (e.g., cameras, lidar, or radar).
[0026] It should be understood that the sound source localization technology described in this article can be used to determine the location of various sound types (e.g., horn sounds, siren sounds, etc.) and sound sources, including emergency vehicles, pedestrians, school buses, other motor vehicles, etc.
[0027] II. Example Implementation
[0028] Figure 1This is a functional diagram illustrating the components of an autonomous vehicle 100 according to an example embodiment. For example, the autonomous vehicle 100 may take the form of a car, truck, motorcycle, bus, boat, airplane, helicopter, lawnmower, bulldozer, snowmobile, aircraft, recreational vehicle, amusement park vehicle, farm equipment, construction equipment, tram, golf cart, train, and trolley. Other vehicles are also possible. The autonomous vehicle 100 can be configured to operate fully or partially in an autonomous mode. For example, the autonomous vehicle 100 may control itself in an autonomous mode and may be operable to determine the current state of the autonomous vehicle 100 and its environment, determine the predicted behavior of at least one other vehicle in the environment, determine the confidence level of the probability that the predicted behavior can be performed corresponding to at least one other vehicle, and control the autonomous vehicle 100 based on the determined information. When in autonomous mode, the autonomous vehicle 100 may be configured to operate without human interaction.
[0029] exist Figure 1 The top 102 of an autonomous vehicle 100 is shown in the diagram. Three microphone units 150, 160, and 170 are coupled to the top 102 of the autonomous vehicle 100 at different locations. For example, a first microphone unit 150 is positioned at a first location on the top 102, a second microphone unit 160 is positioned at a second location on the top 102, and a third microphone unit 170 is positioned at a third location on the top 102. In the example embodiment, the first location associated with the first microphone unit 150 corresponds to the left side of the top 102, the second location associated with the second microphone unit 160 corresponds to the right side of the top 102, and the third location associated with the third microphone unit 170 corresponds to the rear of the top 102. It should be understood that in other embodiments, additional (or fewer) microphone units may be coupled to the top 102 of the autonomous vehicle 100 at different locations. For example, in one embodiment, four microphone units may be coupled to the top 102 of the autonomous vehicle 100 at different locations. In another embodiment, two microphone units may be coupled to the top 102 of the autonomous vehicle 100 at different locations. In yet another embodiment, multiple microphone units may be coupled to the top 102 in a circular pattern. Therefore, regarding Figure 1 The microphone units 150, 160, 170 and their corresponding positions shown and described are for illustrative purposes only and should not be construed as limiting.
[0030] The first microphone unit 150 includes a first microphone array 151 and a first internal clock 153. The first microphone array 151 includes microphones 151A, 151B, and 151C. Each microphone 151A-151C is coupled to a microphone board 157 and positioned above an opening in the microphone board 157. According to the above embodiment, the first microphone array 151 can be positioned facing the left side of the autonomous vehicle 100 to capture sound originating from the left side of the autonomous vehicle 100 with relatively high accuracy.
[0031] The second microphone unit 160 includes a second microphone array 161 and a second internal clock 163. The second microphone array 161 includes microphones 161A, 161B, and 161C. Each microphone 161A-161C is coupled to a microphone board 167 and positioned above an opening in the microphone board 167. According to the above embodiment, the second microphone array 161 can be positioned facing the right side of the autonomous vehicle 100 to capture sound originating from the right side of the autonomous vehicle 100 with relatively high accuracy.
[0032] The third microphone unit 170 includes a third microphone array 171 and a third internal clock 173. The third microphone array 171 includes microphones 171A, 171B, and 171C. Each microphone 171A-171C is coupled to a microphone board 177 and positioned above an opening in the microphone board 177. According to the above embodiment, the third microphone 171 can be positioned towards the rear of the autonomous vehicle 100 to capture sound originating from the rear of the autonomous vehicle 100 with relatively high accuracy.
[0033] It should be understood that although three microphones are shown in each microphone array 151, 161, 171, in some embodiments, one or more of microphone arrays 151, 161, 171 may include additional (or fewer) microphones. As a non-limiting example, one or more of microphone arrays 151, 161, 171 may include two microphones. For illustration, microphone array 151 may include a first microphone facing a first direction and a second microphone facing a second direction different from the first direction. As another non-limiting example, one or more of microphone arrays 151, 161, 171 may include a microphone ring. For illustration, microphone array 151 may include a ring of microphones.
[0034] According to some implementations, the microphones in a microphone array can be oriented in different directions to improve sound source localization, as described below. For example, refer to... Figure 2 Non-limiting examples of microphones 151A-151C coupled to microphone board 157 are shown. Figure 2In this configuration, microphone 151A faces a first direction, microphone 151B faces a second direction at a 120-degree angle to the first direction, and microphone 151C faces a third direction at a 120-degree angle to both the first and second directions. For example, in... Figure 2 In this configuration, microphone 151A is oriented at 0 degrees, microphone 151B at 120 degrees, and microphone 151C at 240 degrees. Furthermore, microphones 151A-151C can be separated by a specific distance "x". According to one embodiment, the specific distance "x" is equal to 6 centimeters. It should be understood that... Figure 2 The examples shown are for illustrative purposes only and should not be construed as limiting.
[0035] Return to reference Figure 1 Each microphone unit 150, 160, 170 is coupled to the computing system 110 via bus 104. Although in Figure 1 The bus is shown as a physical bus 104 (e.g., a wired connection), but in some embodiments, bus 104 may be a wireless communication medium for communicating messages and signals between microphone units 150, 160, 170 and computing system 110. According to one embodiment, bus 104 is a two-wire interface bus connecting a single master device (e.g., computing system 110) and multiple slave devices (e.g., microphone units 150, 160, 170). For example, bus 104 may be an automotive audio bus. Figure 1 As shown, the computing system 110 can be integrated into the cab 103 of the autonomous vehicle 100. For example, the computing system 110 can be integrated into the front console or the central console of the autonomous vehicle 100.
[0036] The computing system 110 includes a processor 112 coupled to a memory 114. The memory 114 may be a non-transitory computer-readable medium storing instructions 124 executable by the processor 112. The processor 112 includes a clock synchronization module 116, a sound classification module 118, a location determination module 120, and a command generation module 122. According to some embodiments, one or more of modules 116, 118, 120, and 122 may correspond to software (e.g., instructions 124) executable by the processor 112. According to other embodiments, one or more modules 116, 118, 120, and 122 may correspond to special-purpose circuitry (e.g., application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs)) integrated into the processor 112.
[0037] Clock synchronization module 116 is configured to synchronize a first internal clock 153 of the first microphone unit 150 with a second internal clock 163 of the second microphone unit 160 and a third internal clock 173 of the third microphone unit 170. For example, clock synchronization module 116 may generate a synchronization signal 140 based on the clock (not shown) of computing system 110. Processor 112 may send synchronization signal 140 to the first microphone unit 150, second microphone unit 160, and third microphone unit 170 to synchronize the internal clocks 153, 163, and 173 of each microphone unit 150, 160, and 170. According to some embodiments, processor 112 may periodically send synchronization signal 140 to each microphone unit 150, 160, and 170 to ensure synchronization of the internal clocks 153, 163, and 173 of the microphone units 150, 160, and 170. As a non-limiting example, processor 112 may send a synchronization signal 140 to each microphone unit 150, 160, 170 every M clock cycles, where M is any integer value.
[0038] At least one microphone 151A-151C in the first microphone array 151 is configured to capture a specific sound 190 to generate a first audio frame 152. For example, microphone 151A is configured to capture a specific sound 190 to generate the first audio frame 152. The first audio frame 152 may include a first timestamp 154 and a first audio attribute 156. For example, the first microphone unit 150 may use a first internal clock 153 associated with microphone 151A to generate the first timestamp 154. The first timestamp 154 indicates the time when microphone 151A captures the specific sound 190. The first audio attribute 156 may correspond to at least one of the frequency characteristics of the specific sound 190 when captured by microphone 151A, the pitch characteristics of the specific sound 190 when captured by microphone 151A, the reverberation characteristics of the specific sound 190 when captured by microphone 151A, etc. The first microphone unit 150 is configured to send the first audio frame 152 to the computing system 110 via bus 104 in response to generating the first audio frame 152. It should be understood that when a specific sound 190 is detected, microphones 151B and 151C can capture the specific sound 190 and generate an audio frame with a corresponding timestamp, in a manner similar to that of microphone 151A.
[0039] At least one microphone 161A-161C in the second microphone array 161 is configured to capture a specific sound 190 to generate a second audio frame 162. For example, microphone 161A is configured to capture a specific sound 190 to generate the second audio frame 162. The second audio frame 162 may include a second timestamp 164 and a second audio attribute 166. For example, the second microphone unit 160 may use a second internal clock 163 associated with microphone 161A to generate the second timestamp 164. The second timestamp 164 indicates the time when microphone 161A captures the specific sound 190. The second audio attribute 166 may correspond to at least one of the frequency characteristics of the specific sound 190 when captured by microphone 161A, the pitch characteristics of the specific sound 190 when captured by microphone 161A, the reverberation characteristics of the specific sound 190 when captured by microphone 161A, etc. The second microphone unit 160 is configured to send the second audio frame 162 to the computing system 110 via bus 104 in response to generating the second audio frame 162. It should be understood that when a specific sound 190 is detected, microphones 161B and 161C can capture the specific sound 190 and generate an audio frame with a corresponding timestamp, in a manner similar to that of microphone 161A.
[0040] At least one microphone 171A-171C in the third microphone array 171 is configured to capture a specific sound 190 to generate a third audio frame 172. For example, microphone 171A is configured to capture a specific sound 190 to generate the third audio frame 172. The third audio frame 172 may include a third timestamp 174 and a third audio attribute 176. For example, the third microphone unit 170 may use a third internal clock 173 associated with microphone 171A to generate the third timestamp 174. The third timestamp 174 indicates the time when microphone 171A captures the specific sound 190. The third audio attribute 176 may correspond to at least one of the frequency characteristics of the specific sound 190 when captured by microphone 171A, the pitch characteristics of the specific sound 190 when captured by microphone 171A, the reverberation characteristics of the specific sound 190 when captured by microphone 171A, etc. The third microphone unit 170 is configured to send the third audio frame 172 to the computing system 110 via bus 104 in response to generating the third audio frame 172. It should be understood that when a specific sound 190 is detected, microphones 171B and 171C can capture the specific sound 190 and generate an audio frame with a corresponding timestamp, in a manner similar to that of microphone 171A.
[0041] Processor 112 is configured to receive audio frames 152, 162, and 172 from microphone units 150, 160, and 170 respectively via bus 104. Upon receiving audio frames 152, 162, and 172, sound classification module 118 is configured to determine (or classify) a specific source 192 of a specific sound 190. For example, sound classification module 118 can determine whether the specific source 192 is an emergency vehicle, a pedestrian, an ice cream truck, another motor vehicle, etc.
[0042] For ease of explanation, the sound classification module 118 is described below as determining whether a specific source 192 of a particular sound 190 is an emergency vehicle. As a non-limiting example, the sound classification module 118 is described as determining whether the specific source 192 is a police car siren, ambulance siren, fire truck siren, etc. As described below, detecting a sound associated with an emergency vehicle may necessitate maneuvering the autonomous vehicle 100 (e.g., pulling the autonomous vehicle 100 to the side of the road, reducing the speed of the autonomous vehicle 100, etc.).
[0043] Although the description focuses on determining whether a specific source 192 is an emergency vehicle, it should be understood that in scenarios where specific source 192 is not an emergency vehicle, the classification techniques described herein can be implemented to classify specific source 192. As a non-limiting example, sound classification module 118 can determine whether specific source 192 is a train, and if so, may force autonomous vehicle 100 to stop. As another non-limiting example, sound classification module 118 can determine whether specific source 192 is a child, and if so, may also force autonomous vehicle 100 to stop. However, for ease of description, the following example focuses on determining whether specific source 192 is an emergency vehicle.
[0044] To determine whether a specific source 192 is an emergency vehicle, the sound classification module 118 can compare the first audio attribute 156 of the first audio frame 152 with multiple sound models associated with emergency vehicle sounds. For illustration, sound model data 126 is stored in memory 114 and is accessible by processor 112. Sound model data 126 may include datasets of sound models of different emergency vehicle sounds. As a non-limiting example, sound model data 126 may include datasets of sound models of police car sirens, ambulance sirens, fire truck sirens, etc. Sound model data 126 can be continuously updated using machine learning or through remote transmission of sound model datasets from manufacturers.
[0045] The sound classification module 118 can compare the first audio attribute 156 of the first audio frame 152 with each sound model dataset to determine a similarity score. The similarity score can be based on at least one of the following: (i) the similarity between the frequency characteristics indicated in the first audio attribute 156 and the frequency characteristics of the selected sound model dataset; (ii) the similarity between the pitch characteristics indicated in the first audio attribute 156 and the pitch characteristics of the selected sound model dataset; (iii) the similarity between the reverberation characteristics indicated in the first audio attribute 156 and the reverberation characteristics of the selected sound model dataset, etc. If the similarity score of the selected sound model dataset is higher than a certain threshold, the sound classification module 118 can determine that the specific source 192 is the corresponding emergency vehicle.
[0046] To illustrate, if the similarity score is higher than a certain threshold when comparing the first audio attribute 156 with the sound model dataset associated with an ambulance siren, the sound classification module 118 can determine that the specific source 192 is an emergency vehicle (e.g., an ambulance). However, if the sound classification module 118 compares the first audio attribute 156 with each sound model dataset associated with an emergency vehicle (or with any other sound source that might force the operation of the autonomous vehicle 100), and no result similarity score is higher than the certain threshold, the sound classification module 118 can determine that the specific source 192 is not an emergency vehicle and can ignore the specific sound 190. Similarly, the sound classification module 118 can compare the second audio attribute 166 of the second audio frame 162 and the third audio attribute 176 of the third audio frame 172 with multiple sound models associated with emergency vehicle sounds.
[0047] In some implementations, before comparing audio attributes 156, 166, 176 with multiple sound models, the first audio attribute 156 of the first audio frame 152 is compared with the second and third audio attributes 166, 176 of the second and third audio frames 162, 172 to ensure that audio frames corresponding to the same sound (e.g., a specific sound 190) are compared with multiple sound models. For example, the processor 112 may be configured to determine whether both the first audio frame 152 and the second audio frame 162 capture a specific sound 190 from a specific source 192. If the deviation between audio attributes 156, 166 is greater than a threshold deviation, the sound classification module 118 may avoid using audio frames 152, 162 to classify the specific source 192 of the specific sound 190 and use subsequent audio frames from microphone units 150, 160 with a smaller deviation.
[0048] It should be understood that comparing the audio attributes 166, 176 of other audio frames 162, 172 with multiple sound models can reduce the likelihood of false positives. For example, if a comparison based on the first audio attribute 156 of the first audio frame 152 leads to the determination that a specific source 192 is an emergency vehicle, and a comparison based on the other audio attributes 166, 176 of other audio frames 162, 172 leads to the determination that a specific source 192 is not an emergency vehicle, then the sound classification module 118 can determine that a false positive may exist. As a result of a potential false positive, a comparison based on the audio attributes of additional (e.g., subsequent) audio frames generated by the microphone units 150, 160, 170 can be performed to classify the specific source 192 of a specific sound 190.
[0049] According to one implementation, in response to classifying a specific source 192 of a specific sound 190, the location determination module 120 is configured to determine the location of the specific source 192 relative to the autonomous vehicle 100 based on timestamps 154, 164, and 174 of audio frames 152, 162, and 172. The location of the specific source 192 relative to the autonomous vehicle 100 can be determined based on determining which of the microphone arrays 151, 161, and 171 first detected the specific sound 190.
[0050] For illustration, as described above, each microphone array 151, 161, 171 is positioned at a different location on the top 102, and each microphone 151A-151C, 161A-161C, 171A-171C can face different directions. Using the above non-limiting example, microphone 151A is located on the left side of the top 102 and faces left, microphone 161A is located on the right side of the top 102 and faces right, and microphone 171A is located at the rear of the top 102 and faces the rear of the autonomous vehicle 100. The microphone closest to and facing the specific source 192 will likely detect the specific sound 190 before the microphone further away from and facing away from the specific source 192. For ease of description and illustration, it is assumed that the specific source 192 is approaching the autonomous vehicle from the left front side, such that microphone 151A is closer to the specific source 192 than microphone 161A, and microphone 161A is closer to the specific source 192 than microphone unit 171A. In this scenario, microphone 151A may detect the specific sound 190 and generate the first audio frame 152 before microphone 161A detects the specific sound 190 and generates the second audio frame 162. As a result, the first timestamp 154 of the first audio frame 152 will indicate a time earlier than the second timestamp 164 of the second audio frame 162. Furthermore, in the above scenario, microphone 161A may detect the specific sound 190 and generate the second audio frame 162 before microphone 171A detects the specific sound 190 and generates the third audio frame 172. As a result, the second timestamp 164 of the second audio frame 162 will indicate a time earlier than the third timestamp 174 of the third audio frame 172.
[0051] The location determination module 120 determines the position of a specific source 192 relative to the autonomous vehicle 100 based on a comparison of timestamps 154, 164, and 174. For example, in response to determining that the first timestamp 154 indicates an earlier time than other timestamps 164 and 174, the location determination module 120 can determine that the specific source 192 is approaching from the side of the autonomous vehicle 100 at a first position proximate to the first microphone unit 150, which is the opposite of approaching from the side at a position proximate to the other microphone units 160 and 170.
[0052] For reference Figure 4 As described, in some implementations, using the timestamps of individual microphone units, processor 112 can perform operations to determine the position of a specific source 192 relative to the autonomous vehicle 100. When determining the position of a specific source 192 based on the timestamps of each individual microphone unit, processor 112 can determine the position of a specific source 192 based on an individualized determination.
[0053] Command generation module 122 is configured to generate command 142 to manipulate autonomous vehicle 100 based on the location of specific source 192. As a non-limiting example, if location determination module 120 indicates that the location of specific source 192 is to the left front of autonomous vehicle 100, command generation module 122 may generate command 142 to navigate autonomous vehicle 100 to the right side of the road. Additionally, or alternatively, command generation module 122 may generate command 142 to reduce the speed of autonomous vehicle 100.
[0054] Command generation module 122 can send command 142 to autonomous vehicle control unit 106 via bus 108. Autonomous vehicle control unit 106 can be coupled to control different components of autonomous vehicle 100, such as steering wheel, brakes, accelerator, turning signals, etc. Based on command 142, autonomous vehicle control unit 106 can send signals to different components of autonomous vehicle 100. For example, autonomous vehicle control unit 106 can send a signal to enable steering wheel to maneuver autonomous vehicle 100 to the side of the road, autonomous vehicle control unit 106 can send a signal to activate brakes to reduce the speed (or stop) of autonomous vehicle 100, etc.
[0055] According to one embodiment, in response to determining the location of a specific source 192, the command generation module 122 can generate a command 142 to change the mode of the autonomous vehicle 100 to a user-assisted mode. In this embodiment, in response to receiving the command 142, the autonomous vehicle control unit 106 can send a signal to the components of the autonomous vehicle 100 to disable the autonomous operation mode, allowing the driver to control the operation of the autonomous vehicle 100.
[0056] about Figure 1 The described technique uses microphone arrays 151, 161, and 171 positioned at different locations on the top 102 of the autonomous vehicle 100 to enable sound source localization. For example, based on timestamps 154, 164, and 174 indicating when the respective microphone arrays 151, 161, and 171 captured a specific sound 190, the computing system 110 can determine the location of a specific source 192 and generate commands to manipulate the autonomous vehicle 100 based on the position of the specific source 192 relative to the autonomous vehicle 100. Therefore, the position and orientation of the microphone arrays 151, 161, and 171 are used to locate the specific source 192.
[0057] Figure 3 A schematic diagram of a microphone coupled to a microphone board according to an example embodiment is depicted. For example, Figure 3 A first example 300 of a microphone coupled to a microphone board and a second example 350 of a microphone coupled to a microphone board are depicted.
[0058] According to the first example 300, microphone 151A is directly coupled to microphone board 157 and positioned above an opening in microphone board 157. It should be understood that the configuration in the first example 300 can be applied to the reference... Figure 1 Any microphone described. In the first example 300, there is a strong flush to support improved acoustic effects.
[0059] According to the second example 350, microphone 151A is coupled to object 352 (e.g., metal, plastic, etc.), and object 352 is coupled to microphone plate 157. In the second example 350, the height (h) of object 352 is greater than the distance (d) of the opening in microphone plate 157. As a result, in the second example 350, there is a stronger acoustic scouring than in the first example 300, leading to even better acoustic performance. It should be understood that the configuration in the second example 350 can be applied to the reference... Figure 1 Any microphone described.
[0060] Figure 4 This is another functional diagram illustrating the components of an autonomous vehicle 100 according to an example embodiment. Specifically, Figure 4 Microphone unit 150, microphone unit 160, computing system 110, and autonomous vehicle control unit 106 are shown. It should be understood that, in reference... Figure 4 The described technology may include additional microphone units, such as microphone unit 170. It should also be understood that the number of microphones in each microphone unit 150, 160 may vary depending on the implementation.
[0061] exist Figure 4 In this configuration, each microphone of the first microphone array 151 is configured to capture a specific sound 190 from a specific source 192 to generate a corresponding audio frame with a timestamp based on a first internal clock 153. For example, microphone 151A is configured to capture a specific sound 190 from a specific source 192 to generate an audio frame 152A with a timestamp 154A based on the first internal clock 153, microphone 151B is configured to capture a specific sound 190 from a specific source 192 to generate an audio frame 152B with a timestamp 154B based on the first internal clock 153, and microphone 151C is configured to capture a specific sound 190 from a specific source 192 to generate an audio frame 152C with a timestamp 154C based on the first internal clock 153. Because microphones 151A-151C face different directions and have different positions, as... Figure 2As shown, microphones 151A-151C capture a specific sound 190 at slightly different times, resulting in slightly different times indicated by timestamps 154A-154C. Audio frames 152A-152C include audio attributes 156A-156C for the specific sound 190 captured by microphones 151A-151C, respectively. Among other attributes, audio attributes 156A-156C include the sound level of the specific sound 190.
[0062] Similarly, each microphone in the second microphone array 161 is configured to capture a specific sound 190 from a specific source 192 to generate a corresponding audio frame with a timestamp based on a second internal clock 163. For example, microphone 161A is configured to capture the specific sound 190 from the specific source 192 to generate an audio frame 162A with a timestamp 164A based on the second internal clock 163, microphone 161B is configured to capture the specific sound 190 from the specific source 192 to generate an audio frame 162B with a timestamp 164B based on the second internal clock 163, and microphone 161C is configured to capture the specific sound 190 from the specific source 192 to generate an audio frame 162C with a timestamp 164C based on the second internal clock 163. Because microphones 161A-161C face different directions and have different positions, they capture the specific sound 190 at slightly different times, resulting in slightly different times indicated by timestamps 164A-164C. Audio frames 162A-162C include audio attributes 166A-166C for a specific sound 190 captured by microphones 161A-161C, respectively. Among other attributes, audio attributes 166A-166C include the sound level of the specific sound 190.
[0063] Processor 112 is configured to perform a first operation to locate a specific source 192 relative to autonomous vehicle 100 based on timestamps 154A-154C associated with the first microphone unit 150. To perform the first operation, processor 112 is configured to determine whether each audio frame 152A-152C captures a specific sound 190 from the specific source 192. For example, processor 112 may compare the audio attributes 156A-156C of audio frames 152A-152C to each other and to sound model data 126 to determine that each audio frame 152A-152C captures a specific sound 190, as per [the relevant information]. Figure 1As described, when it is determined that each audio frame 152A-152C captures a specific sound 190, the processor 112 is configured to compare the timestamps 154A-154C of each audio frame 152A-152C to determine the specific microphone in the first microphone array 151 that first captures the specific sound 190. For example, the processor 112 may determine which timestamp indicates the earliest time to determine the specific microphone that first captures the specific sound 190.
[0064] For descriptive purposes, it is assumed that timestamp 154A indicates a time earlier than timestamps 154B and 154C. Based on this assumption, processor 112 is configured to locate a specific source 192 relative to autonomous vehicle 100 based on the attributes of microphone 151A compared with the attributes of other microphones 151B and 151C. According to one embodiment, the attributes of microphone 151A may correspond to the location of microphone 151A. Therefore, if microphones 151A-151C are as follows... Figure 2 As shown in the arrangement, the processor 112 can position the specific source 192 closer to the microphone 151A than the positions of microphones 151B and 151C. According to another embodiment, the properties of microphone 151A can correspond to the orientation of microphone 151A. Therefore, if microphones 151A-151C are as shown... Figure 2 Given the orientation shown, the processor 112 can locate the angle of arrival of a specific sound 190 to approximately 0 degrees. According to some embodiments, using timestamps 154A-154C, the processor 112 can determine the time difference of arrival of a specific sound 190 and calculate the azimuth angle of a specific source 192 based on that difference.
[0065] Processor 112 is also configured to perform a second operation to locate a specific source 192 relative to the autonomous vehicle based on timestamps 164A-164C associated with the second microphone unit 160. To perform the second operation, processor 112 is configured to determine whether each audio frame 162A-162C captures a specific sound 190 from the specific source 192. For example, processor 112 may compare the audio attributes 166A-166C of audio frames 162A-162C with each other and with sound model data 126 to determine whether each audio frame 162A-162C captures a specific sound 190, as per [the relevant information]. Figure 1 As described, when it is determined that each audio frame 162A-162C captures a specific sound 190, the processor 112 is configured to compare the timestamps 164A-164C of each audio frame 162A-162C to determine the specific microphone in the second microphone array 161 that first captures the specific sound 190. For example, the processor 112 may determine which timestamp indicates the earliest time to determine the specific microphone that first captures the specific sound 190.
[0066] For descriptive purposes, it is assumed that timestamp 164B indicates an earlier time than timestamps 164A and 164C. Based on this assumption, processor 112 is configured to locate a specific source 192 relative to autonomous vehicle 100 based on the attributes of microphone 161B compared with the attributes of other microphones 161A and 161C. According to one embodiment, the attributes of microphone 161B may correspond to the position of microphone 161B. According to another embodiment, the attributes of microphone 161B may correspond to the orientation of microphone 161B. According to some embodiments, using timestamps 164A-164C, processor 112 can determine the time difference of arrival of a specific sound 190 and calculate the azimuth angle of the specific source 192 based on this difference.
[0067] Processor 112 is also configured to determine the position of a specific source 192 relative to autonomous vehicle 100 based on a first operation and a second operation. For example, in response to determining that the earliest timestamp 154A from the first microphone array 151 is associated with microphone 151A and the earliest timestamp 164B from the second microphone array 161 is associated with microphone 161B, processor 112 can also locate the specific source 192 relative to autonomous vehicle 100 based on the attributes of microphones 151A and 161B. For example, processor 112 can use the orientation of microphones 151A and 161B to determine the angle of arrival of a specific sound 190. To illustrate, if the orientation of microphone 151A indicates that the angle of arrival of the specific sound 190 is 0 degrees and the orientation of microphone 161B indicates that the angle of arrival of the specific sound 190 is 10 degrees, then processor 112 can locate the specific source 192 at 5 degrees. In a similar manner, processor 112 can locate the specific source 192 based on the position of microphones 151A and 161B.
[0068] According to one implementation, processor 112 can determine the distance of a specific source 192 from autonomous vehicle 100 based on the sound levels (e.g., audio attributes 156A, 166B) of audio frames 152A, 162B. For example, processor 112 can determine a first operating distance based on the sound level of audio frame 152A and a second operating distance based on the sound level of audio frame 162B. In response to individually determining the distance based on the sound levels of audio frames 152A, 162B, processor 112 can input the first and second operating distances into an algorithm to determine the distance of the specific source 192 from autonomous vehicle 100. In some implementations, this distance can be equal to the mean (e.g., average) distance of the individual distances.
[0069] refer to Figure 4The described technique uses microphone arrays 151, 161 positioned at different locations on the top 102 of the autonomous vehicle 100 to enable sound source localization. For example, based on timestamps 154A-154C, 164A-164C indicating when the respective microphones 151A-151C, 161A-161C captured a specific sound 190, the computing system 110 can determine the location of a specific source 192. For example, the computing system 110 can determine which microphones 151A, 161B in each microphone unit 150, 160 first captured the specific sound 190, and determine the location of the specific source 192 based on the attributes of microphones 151A, 161B (e.g., orientation and position) compared to the attributes of other microphones.
[0070] Figure 5 Schematic diagrams depicting different top positions of the coupling microphone unit according to an example embodiment are provided. Figure 5 In the diagram, different locations on the top 102 (e.g., locations A-E) are depicted as potential locations for coupling microphone units (such as microphone units 150, 160, 170). It should be understood that... Figure 5 The locations depicted are for illustrative purposes only and should not be construed as limiting. According to some embodiments, the microphone unit is placed in a ring surrounding the top 102 to improve sound source localization.
[0071] The position of the microphone unit can be determined based on the detected wind speed. For example, in a scenario where a limited number of microphone units are available, the microphone unit can be coupled to the top 102 at a location with relatively low wind speeds. Simulated data can be generated to detect wind speeds at different locations. For example, during a simulation, a sensor can be placed on the top 102 of the autonomous vehicle 100 to detect various wind speeds at different locations. Figure 5 In the non-limiting illustrative example, location A has a wind speed of 35 m / sec, location B has a wind speed of 30 m / sec, location C has a wind speed of 10 m / sec, location D has a wind speed of 5 m / sec, and location E has a wind speed of 3 m / sec. Therefore, according to Figure 5 In the non-limiting illustrative example, position E is a relatively good location for the coupled microphone unit, position D is the second best location for the coupled microphone unit, position C is the third best location for the coupled microphone unit, position B is the next best location for the coupled microphone unit, and position A is the worst location for the coupled microphone unit.
[0072] It should be understood that the location of the microphone unit can vary based on the structure of the autonomous vehicle. Therefore, different models of autonomous vehicles may have different preferred locations for coupling the microphone unit to the top.
[0073] III. Example Method
[0074] Figure 6 This is a flowchart of method 600 according to an example embodiment. Method 600 may be performed by microphone unit 150, 160, 170, computing system 110, or a combination thereof.
[0075] Method 600 includes, at 602, synchronizing at a processor a first internal clock associated with a first microphone array positioned at a first location on top of the autonomous vehicle and a second internal clock associated with a second microphone array positioned at a second location on top of the autonomous vehicle. For example, refer to... Figure 1 The clock synchronization module 116 sends a synchronization signal 140 to the internal clocks 153 and 163 to synchronize them. It should be understood that synchronizing the internal clocks 153 and 163 enables the position determination module 120 to determine relatively accurately which audio frames 152 and 162 were generated first based on timestamps 154 and 164. If the internal clocks 153 and 163 are not synchronized, it may introduce a certain degree of error, causing later-generated audio frames to have earlier timestamps than earlier-generated audio frames.
[0076] Method 600 further includes receiving a first audio frame and a corresponding first timestamp from at least one microphone in the first microphone array at 604. For example, refer to Figure 1 The processor 112 receives a first audio frame 152 (including a first timestamp 154) from a microphone 151A located at a first position on the top 102 of the autonomous vehicle 100. The microphone 151A generates the first audio frame 152 in response to the microphone 151A capturing a specific sound 190 from a specific source 192. The first timestamp 154 is generated by the microphone 151A (e.g., the first microphone unit 150) using a first internal clock 153. The first timestamp 154 indicates the time at which the microphone 151A captured the specific sound 190.
[0077] Method 600 further includes receiving a second audio frame and a corresponding second timestamp from at least one microphone in the second microphone array at 606. For example, refer to Figure 1 The processor 112 receives a second audio frame 162 (including a second timestamp 164) from a microphone 161A located at a second position on the top 102 of the autonomous vehicle 100. The microphone 161A generates the second audio frame 162 in response to the microphone 161A capturing a specific sound 190 from a specific source 192. The second timestamp 164 is generated by the microphone 161A (e.g., the second microphone unit 160) using a second internal clock 163. The second timestamp 164 indicates the time at which the second microphone 161A captured the specific sound 190.
[0078] Method 600 further includes determining at 608 that both the first audio frame and the second audio frame capture a specific sound from a specific source. For example, refer to Figure 1 Before comparing audio attributes 156, 166, and 176 with multiple sound models, the first audio attribute 156 of the first audio frame 152 is compared with the second and third audio attributes 166 and 176 of the second and third audio frames 162 and 172 to ensure that audio frames corresponding to the same sound (e.g., a specific sound 190) are compared with multiple sound models. The processor 112 uses this comparison to determine that the first audio frame 152 and the second audio frame 162 captured a specific sound 190 from a specific source 192. If the deviation between audio attributes 156 and 166 is greater than a threshold deviation, the sound classification module 118 can bypass using audio frames 152 and 162 to classify the specific source 192 of the specific sound 190 and use subsequent audio frames from microphone units 150 and 160 with a smaller deviation.
[0079] Method 600 further includes determining the position of a specific source relative to an autonomous vehicle at 610 based on a first timestamp and a second timestamp. For example, refer to Figure 1 In response to the sound classification module 118 determining that a specific sound 190 is associated with an emergency vehicle, the location determination module 120 determines the location of the specific source 190 based on a first timestamp 154 and a second timestamp 164.
[0080] For illustration, using the above non-limiting example, microphone 151A of the first microphone array 151 is located on the left side of the top 102 and faces left, microphone 161A of the second microphone array 161 is located on the right side of the top 102 and faces right, and microphone 171A of the third microphone array 171 is located at the rear of the top 102 and faces the rear of the autonomous vehicle 100. The microphone closest to and facing the specific source 192 will likely detect the specific sound 190 before microphones further away from the emergency vehicle. For ease of description and illustration, assume that the specific source 192 is approaching the autonomous vehicle from the left front side, such that microphone 151A is closer to the specific source 192 than microphone 161A, and microphone 161A is closer to the specific source 192 than microphone unit 171A. In this scenario, microphone 151A may detect the specific sound 190 and generate the first audio frame 152 before microphone 161A detects the specific sound 190 and generates the second audio frame 162. As a result, the first timestamp 154 of the first audio frame 152 will indicate a time earlier than the second timestamp 164 of the second audio frame 162. Additionally, in the above scenario, before microphone 171A detects the specific sound 190 and generates the second audio frame 162, microphone 161A may detect the specific sound 190 and generate the second audio frame 162. Consequently, the second timestamp 164 of the second audio frame 162 will indicate a time earlier than the third timestamp 174 of the third audio frame 172.
[0081] The location determination module 120 determines the position of a specific source 192 relative to the autonomous vehicle 100 based on a comparison of timestamps 154, 164, and 174. For example, in response to determining that the first timestamp 154 indicates a time earlier than the other timestamps 164 and 174, the location determination module 120 determines that the specific source 192 is approaching (or is located at a first position closer to the first microphone array 151 than at a position closer to the other microphone arrays 161 and 171) from a first position of the first microphone array 151 rather than at a position closer to the other microphone arrays 161 and 171.
[0082] Method 600 also includes generating commands for maneuvering the autonomous vehicle at 612 based on the position of a specific source relative to the autonomous vehicle. For example, refer to Figure 1Command generation module 122 generates a command 142 that is sent to autonomous vehicle control unit 106. Command 142 instructs autonomous vehicle control unit 106 to manipulate autonomous vehicle 100 based on the location of a specific source 192. According to one embodiment of method 600, command 142 for manipulating autonomous vehicle 100 includes a command to reduce the speed of autonomous vehicle 100. According to another embodiment of method 600, command 142 for manipulating autonomous vehicle 100 includes a command to navigate autonomous vehicle 100 to one side of a road. According to yet another embodiment of method 600, command 142 for manipulating autonomous vehicle 100 includes a command to change the mode of autonomous vehicle 100 to user-assisted mode.
[0083] According to one implementation, method 600 further includes determining whether a specific sound is associated with an emergency vehicle based on a first audio frame and a second audio frame. For example, see reference... Figure 1 The sound classification module 118 can determine whether a specific sound 190 is associated with an emergency vehicle based on a first audio frame 152 and a second audio frame 162. According to one embodiment of method 600, the sound classification module 118 compares a first audio attribute 156 of the first audio frame 152 and a second audio attribute 166 of the second audio frame 162 with a plurality of sound models (e.g., sound model data 126) associated with the emergency vehicle sound. In response to determining that the first audio attribute 156 and the second audio attribute 166 substantially match at least one of the plurality of sound models, the specific sound 190 is associated with an emergency vehicle.
[0084] In some embodiments of method 600, an additional microphone array may be coupled to the top of the autonomous vehicle to more accurately locate the sound source, such as an emergency vehicle. For illustration, method 600 may also include receiving a third audio frame and a corresponding third timestamp from at least one microphone of a third microphone array located at a third position on the top of the autonomous vehicle. For example, refer to... Figure 1 The processor 112 receives a third audio frame 172 (including a third timestamp 174) from a microphone 171A located at a third position on the top 102 of the autonomous vehicle 100. The third audio frame 172 is generated by the microphone 171A in response to the microphone 171A capturing a specific sound 190. The third timestamp 174 is generated by the microphone 171A (e.g., the third microphone unit 170) using a third internal clock 173. The third timestamp 174 indicates the time when the third microphone 171 captured the specific sound 190. According to the above embodiment of method 600, wherein the third microphone 171 is located at a third position on the top 102, the sound classification module 118 compares the third audio attribute 176 of the third audio frame 172 with a plurality of sound models (e.g., sound model data 126) associated with vehicle sound.
[0085] Method 600 uses microphone arrays 151, 161, and 171 positioned at different locations on the top 102 of the autonomous vehicle 100 to enable sound source localization. For example, based on timestamps 154, 164, and 174 indicating when the respective microphone arrays 151 and 161 capture a specific sound 190 from a specific source 192, the computing system 110 can determine the position of the specific source 192 relative to the autonomous vehicle 100 and generate commands to manipulate the autonomous vehicle 100 based on the position of the specific source 192. Therefore, the position and orientation of the microphone arrays 151 and 161 are used to locate the specific source 192.
[0086] Figure 7 This is a flowchart of method 700 according to an example embodiment. Method 700 may be performed by microphone unit 150, 160, 170, computing system 110, or a combination thereof.
[0087] Method 700 includes, at 702, synchronizing at a processor a first internal clock associated with a first microphone array positioned at a first location on top of the autonomous vehicle and a second internal clock associated with a second microphone array positioned at a second location on top of the autonomous vehicle. For example, refer to... Figure 1 and 4 The clock synchronization module 116 sends a synchronization signal 140 to the internal clocks 153 and 163 to synchronize the internal clocks 153 and 163.
[0088] Method 700 further includes receiving audio frames and corresponding timestamps from each microphone in the first microphone array at 704. For example, refer to Figure 4 The computing system 110 receives audio frames 152A-152C from each microphone 151A-151C in the first microphone array 151. Each audio frame 152A-152C is generated in response to the corresponding microphone 151A-151C in the first microphone array 151 capturing a specific sound 190 from a specific source 192, and each timestamp 154A-154C is generated using a first internal clock 153 and indicates the time when the corresponding microphone 151A-151C in the first microphone array 151 captured the specific sound 190.
[0089] Method 700 also includes receiving audio frames and corresponding timestamps from each microphone in the second microphone array at 706. For example, refer to Figure 4The computing system 110 receives audio frames 162A-162C from each microphone 161A-161C in the second microphone array 161. Each audio frame 162A-162C is generated in response to the corresponding microphone 161A-161C in the second microphone array 161 capturing a specific sound 190 from a specific source 192, and each timestamp 164A-164C is generated using a second internal clock 163 and indicates the time when the corresponding microphone 161A-161C in the second microphone array 161 captured the specific sound 190.
[0090] Method 700 further includes, at 708, performing a first operation to locate a specific source relative to the autonomous vehicle based on timestamps of audio frames generated by microphones in the first microphone array. For example, refer to Figure 4 The location determination module 120 performs a first operation to locate a specific source 192 relative to the autonomous vehicle 100 based on the timestamps 154A-154C of audio frames 152A-152C generated by microphones 151A-151C in the first microphone array 151.
[0091] Method 700 further includes, at 710, performing a second operation to locate a specific source relative to the autonomous vehicle based on timestamps of audio frames generated by microphones in the second microphone array. For example, refer to Figure 4 The location determination module 120 performs a second operation to locate a specific source 192 relative to the autonomous vehicle 100 based on the timestamps 164A-164C of audio frames 162A-162C generated by microphones 161A-161C in the second microphone array 161.
[0092] Method 700 further includes determining the position of a specific source relative to an autonomous vehicle at 712 based on the first and second operations.
[0093] Method 700 also includes generating, at 714, commands to manipulate the autonomous vehicle based on the position of a specific source relative to the autonomous vehicle.
[0094] Method 700 uses microphone arrays 151, 161, and 171 positioned at different locations on the top 102 of the autonomous vehicle 100 to enable sound source localization. For example, based on timestamps 154, 164, and 174 indicating when the respective microphone arrays 151 and 161 capture a specific sound 190 from a specific source 192, the computing system 110 can determine the position of the specific source 192 relative to the autonomous vehicle 100 and generate commands to manipulate the autonomous vehicle 100 based on the position of the specific source 192. Therefore, the position and orientation of the microphone arrays 151 and 161 are used to locate the specific source 192.
[0095] IV. Conclusion
[0096] The specific arrangements shown in the accompanying drawings should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in the given drawings. Furthermore, some of the shown elements may be combined or omitted. Additionally, illustrative embodiments may include elements not shown in the drawings.
[0097] A step or block representing information processing may correspond to a circuit that can be configured to perform a specific logical function of the method or technique described herein. Alternatively or additionally, a step or block representing information processing may correspond to a module, segment, or portion of program code (including associated data). The program code may include one or more instructions executable by a processor for implementing a specific logical function or action in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device including a disk, hard disk drive, or other storage medium.
[0098] Computer-readable media can also include non-transitory computer-readable media, such as short-term data storage computer-readable media like register memory, processor cache, and random access memory (RAM). Computer-readable media can also include long-term storage non-transitory computer-readable media for program code and / or data. Therefore, for example, computer-readable media can include secondary or permanent long-term storage, such as read-only memory (ROM), optical discs or magnetic disks, and optical disc read-only memory (CD-ROM). Computer-readable media can also be any other volatile or non-volatile storage system. For example, computer-readable media can be considered a computer-readable storage medium, or a tangible storage device.
[0099] While various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various disclosed examples and embodiments are for illustrative purposes and not for limitation; the true scope is indicated by the appended claims.
Claims
1. A system for sound source localization, comprising: A first microphone array, disposed at a first location on an autonomous vehicle, includes a first plurality of microphones, each of the first plurality of microphones being capable of capturing a specific sound from a specific source to generate a corresponding audio frame with a timestamp based on a first internal clock associated with the first microphone array, the timestamp indicating the time when the corresponding microphone among the first plurality of microphones captured the specific sound; A second microphone array, positioned at a second location on the autonomous vehicle, comprises a second plurality of microphones, each of which is capable of capturing a specific sound from a specific source to generate a corresponding audio frame with a timestamp based on a second internal clock associated with the second microphone array, the timestamp indicating the time when the corresponding microphone among the second plurality of microphones captured the specific sound; as well as The processor is configured as follows: Synchronize the first internal clock with the second internal clock; Perform the first operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the first plurality of microphones; Perform a second operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by a second plurality of microphones; The position of a specific source relative to the autonomous vehicle is determined based on the first and second operations; as well as Commands to manipulate the autonomous vehicle are generated based on the position of a specific source relative to the autonomous vehicle.
2. The system according to claim 1, wherein, In order to perform the first operation, the processor is configured to: It was determined that the audio frames from the first plurality of microphones captured a specific sound from a specific source; Compare the timestamps of each audio frame from the first plurality of microphones to determine the specific microphone that first captured a particular sound among the first plurality of microphones; and The attributes of a specific microphone among a plurality of microphones are compared with the attributes of other microphones in the plurality of microphones to locate a specific source relative to the autonomous vehicle.
3. The system according to claim 2, wherein, The attributes of a particular microphone among a first plurality of microphones correspond to the position of that particular microphone among the first plurality of microphones, and wherein the particular source is located based on the position of the particular microphone among the first plurality of microphones compared with the positions of the other microphones among the first plurality of microphones.
4. The system according to claim 2, wherein, The attribute of a particular microphone among the first plurality of microphones corresponds to the orientation of that particular microphone among the first plurality of microphones, and wherein the particular source is located based on the orientation of the particular microphone among the first plurality of microphones compared with the orientation of the other microphones among the first plurality of microphones.
5. The system according to claim 1, wherein, In order to perform the second operation, the processor is configured to: It was determined that each audio frame from a second set of microphones captured a specific sound from a specific source. Compare the timestamps of each audio frame from the second plurality of microphones to determine the specific microphone that first captures a particular sound among the second plurality of microphones; and The attributes of a specific microphone in a second plurality of microphones are compared with the attributes of other microphones in the second plurality of microphones to locate a specific source relative to the autonomous vehicle.
6. The system according to claim 5, wherein, The attributes of a specific microphone in the second plurality of microphones correspond to the position of that specific microphone in the second plurality of microphones, and wherein the specific source is located based on the position of the specific microphone in the second plurality of microphones compared with the positions of other microphones in the second plurality of microphones.
7. The system according to claim 5, wherein, The attribute of a specific microphone in the second plurality of microphones corresponds to the orientation of the specific microphone in the second plurality of microphones, and wherein the specific source is located based on the orientation of the specific microphone in the second plurality of microphones compared with the orientation of the other microphones in the second plurality of microphones.
8. The system according to claim 1, further comprising: A first microphone unit is coupled to the top of the autonomous vehicle at a first location. The first microphone unit includes a first microphone array and a first internal clock. as well as A second microphone unit, coupled to the top of the autonomous vehicle at a second location, includes a second microphone array and a second internal clock.
9. The system according to claim 1, wherein, The first microphone array includes: The first microphone is facing the first direction; The second microphone is oriented in a second direction, 120 degrees away from the first direction; and The third microphone is oriented in a third direction, 120 degrees from both the first and second directions.
10. The system according to claim 1, wherein, The first microphone array includes: The first microphone, facing the first direction; and The second microphone is oriented in a different direction than the first microphone.
11. The system according to claim 1, wherein, The first microphone array includes a microphone ring.
12. The system according to claim 1, wherein, The commands for maneuvering the autonomous vehicle include commands to reduce the speed of the autonomous vehicle.
13. The system according to claim 1, wherein, The commands for maneuvering the autonomous vehicle include commands to navigate the autonomous vehicle to one side of the road.
14. The system according to claim 1, wherein, The commands for manipulating the autonomous vehicle include commands to change the autonomous vehicle's mode to user-assisted mode.
15. A method for sound source localization, comprising: At the processor, a first internal clock associated with a first microphone array located at a first position on the autonomous vehicle and a second internal clock associated with a second microphone array located at a second position on the autonomous vehicle are synchronized. The first microphone array includes a first plurality of microphones, and the second microphone array includes a second plurality of microphones. Audio frames and corresponding timestamps are received from each microphone in the first microphone array. Each audio frame is generated in response to the capture of a specific sound from a specific source by the corresponding microphone in the first microphone array. Each timestamp is generated using a first internal clock and each timestamp indicates the time when the corresponding microphone in the first microphone array captured the specific sound. Audio frames and corresponding timestamps are received from each microphone in the second microphone array. Each audio frame is generated in response to the corresponding microphone in the second microphone array capturing a specific sound from a specific source. Each timestamp is generated using a second internal clock and each timestamp indicates the time when the corresponding microphone in the second microphone array captured the specific sound. Perform a first operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the microphones in the first microphone array; Perform a second operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the microphones in the second microphone array; The position of a specific source relative to the autonomous vehicle is determined based on the first and second operations; as well as Commands to manipulate the autonomous vehicle are generated based on the position of a specific source relative to the autonomous vehicle.
16. The method according to claim 15, wherein, Performing the first operation includes determining the azimuth of a specific source based on the timestamps of audio frames generated by microphones in the first microphone array.
17. The method according to claim 15, wherein, Performing the second operation includes determining the azimuth of a specific source based on the timestamps of audio frames generated by microphones in the second microphone array.
18. The method according to claim 15, wherein, Performing the first operation includes: Determine that each audio frame from the microphones in the first microphone array captures a specific sound from a specific source; Compare the timestamps of each audio frame from the microphones in the first microphone array to determine the specific microphone that first captured a particular sound; and The properties of a specific microphone are used to locate a specific source relative to the autonomous vehicle by comparing the properties of the microphones with those of other microphones in the first microphone array.
19. The method according to claim 15, wherein, Performing the second operation includes: Determine that each audio frame from the microphones in the second microphone array captures a specific sound from a specific source; Compare the timestamps of each audio frame from the microphones in the second microphone array to determine the specific microphone that first captured a particular sound; and The properties of a specific microphone are used to locate a specific source relative to the autonomous vehicle by comparing its properties with those of other microphones in the second microphone array.
20. A non-transitory computer-readable medium storing instructions executable by a computing device to cause the computing device to perform a function, the function including: At the processor, a first internal clock associated with a first microphone array located at a first position on the autonomous vehicle and a second internal clock associated with a second microphone array located at a second position on the autonomous vehicle are synchronized. The first microphone array includes a first plurality of microphones, and the second microphone array includes a second plurality of microphones. Audio frames and corresponding timestamps are received from each microphone in the first microphone array. Each audio frame is generated in response to the capture of a specific sound from a specific source by the corresponding microphone in the first microphone array. Each timestamp is generated using a first internal clock and each timestamp indicates the time when the corresponding microphone in the first microphone array captured the specific sound. Audio frames and corresponding timestamps are received from each microphone in the second microphone array. Each audio frame is generated in response to the corresponding microphone in the second microphone array capturing a specific sound from a specific source. Each timestamp is generated using a second internal clock and each timestamp indicates the time when the corresponding microphone in the second microphone array captured the specific sound. Perform a first operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the microphones in the first microphone array; Perform a second operation to locate a specific source relative to the autonomous vehicle based on the timestamps of audio frames generated by the microphones in the second microphone array; The position of a specific source relative to the autonomous vehicle is determined based on the first and second operations; as well as Commands to manipulate the autonomous vehicle are generated based on the position of a specific source relative to the autonomous vehicle.
Citation Information
Patent Citations
Audio processing for vehicle sensory systems
CN110573398A
System and method of identifying a vehicle and determining the location and the velocity of the vehicle by sound
US20170213459A1