Use of structured audio output to detect playback and / or adapt to inconsistent playback in wireless speakers

By determining the audio delay and adding a delayed segment to the audio stream, the solution addresses incomplete audio rendering in vehicle speakers, ensuring complete playback and optimizing noise cancellation.

JP7723716B2Active Publication Date: 2025-08-14GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023185621
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-01-29
Filing Date
2023-10-30
Publication Date
2025-08-14
Estimated Expiration
2039-02-12

AI Technical Summary

Technical Problem

Vehicle computing devices often fail to render initial portions of audio data transmitted by mobile devices through speakers due to delays or improper modes, leading to incomplete audio playback and wasted computational resources.

Method used

Determine the audio delay between a computing device and vehicle speakers, adding a delayed audio segment to the audio stream to ensure complete rendering, and adapt noise reduction techniques based on the delay to improve audio playback.

Benefits of technology

Ensures complete audio rendering and reduces computational waste by preventing incomplete audio playback and optimizing noise cancellation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007723716000001
    Figure 0007723716000001
  • Figure 0007723716000002
    Figure 0007723716000002
  • Figure 0007723716000003
    Figure 0007723716000003
Patent Text Reader

Abstract

To provide a method for determining audio latency of a computing device by causing the computing device to transmit an audio data stream over a wireless communication channel.SOLUTION: A method uses an audio data stream to render audio output generated via a speaker, captures the rendered audio output via a microphone, and determines the audio delay by comparing the captured audio output to the audio data stream. Then, the delayed audio segment is appended to an additional audio data stream to be sent to the computing device, length of the delayed audio segment is determined by using the audio delay, and a noise elimination filter is additionally or alternatively adapted based on the audio delay.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans may engage in human-computer dialogues using interactive software applications referred to herein as "automated assistants" (which may also be referred to as "digital agents," "chatbots," "conversational personal assistants," "intelligent personal assistants," "assistant applications," "conversational agents," etc.). For example, a human (who, when interacting with an automated assistant, may be referred to as a "user") may provide commands and / or requests to the automated assistant using oral natural language input (i.e., utterances), which may in some cases be converted to text and then processed, and / or by providing textual (e.g., typed) natural language input. The automated assistant complies with the request by providing a responsive user interface output, which may include audible and / or visual interface output.

[0002] As described above, many automated assistants are configured to be interacted with via verbal utterances. To protect user privacy and / or conserve resources, the user must often explicitly invoke the automated assistant before the automated assistant will fully process the verbal utterance. Explicit invocation of the automated assistant generally occurs in response to certain user interface input received at the client device. The client device includes an assistant interface that provides a user of the client device with an interface for interfacing with the automated assistant (e.g., receives verbal and / or typed input from the user and provides audible and / or graphical responses) and interfaces with one or more additional components that implement the automated assistant (e.g., a remote server device that processes the user input and generates appropriate responses).

[0003] Some user interface inputs that can invoke an automated assistant via a client device include hardware and / or virtual buttons on the client device for invoking the automated assistant (e.g., tapping a hardware button, selecting a graphical interface element displayed by the client device). Many automated assistants may additionally or alternatively be invoked in response to one or more verbal invocation phrases, also known as "hot words / phrases" or "trigger words / phrases." For example, an automated assistant can be invoked by uttering verbal invocation phrases such as "hey, assistant," "ok, assistant," and / or "assistant."

[0004] A user may wish to interact with an automated assistant while in a vehicle. For example, a user may call an automated assistant on a mobile smartphone to request driving directions. Additionally, a client device (such as a mobile smartphone) may be communicatively coupled to the vehicle such that audio data provided by the client device may be rendered through one or more vehicle speakers. For example, a mobile smartphone may be communicatively coupled to the vehicle via Bluetooth, and audio data from the mobile smartphone may be transmitted to the vehicle computing device via Bluetooth and rendered by the vehicle computing device via the vehicle speakers. This audio data may include natural language responses provided by the automated assistant to the mobile smartphone.

[0005] However, many vehicle computing devices will fail to render the entire audio data transmitted by a mobile smartphone through vehicle speakers under one or more conditions. For example, some vehicle computing devices may receive the entire audio data but will not render an initial portion of the audio data through the speaker, e.g., due to a delay in initiating the components for rendering the audio data. Thus, the initial portion of the audio data will fail to render through the vehicle speaker. For example, the audio data may include a synthesized voice saying, "Turn left onto Main Street," but the vehicle computing device will only render "on Main Street." This is problematic because the associated "turn left" is not rendered, forcing the user to activate the mobile smartphone display to verify the turn direction and / or causing the user to inadvertently turn "onto Main Street" in the wrong direction. Both of these scenarios result in wasted computational resources. Additionally or alternatively, many vehicle computing devices will fail to render any audio data transmitted by a mobile smartphone through vehicle speakers when the vehicle computing device is not in an appropriate mode (e.g., Bluetooth mode). This can result in wasted transmission of audio data because the audio data is not actually rendered through the vehicle speakers, requiring the user to retransmit the audio data after manually switching to the appropriate mode. Summary of the Invention [Means for solving the problem]

[0006] Implementations described herein relate to determining an audio delay between a computing device and one or more additional speakers (e.g., vehicle speakers) driven by an additional computing device (e.g., a vehicle computing device), where the computing device and the additional computing device are communicatively coupled via a wireless communication channel. In some versions of these implementations, a corresponding delayed audio segment having a duration determined using the audio delay is added to the additional audio stream transmitted to the additional computing device. By adding the delayed audio segment to the additional audio stream, at least a portion of the delayed audio segment is not rendered by the additional speaker, but this ensures that the additional audio stream is rendered by the additional speaker. In some versions of these implementations, the audio delay is additionally or alternatively utilized to adapt noise reduction of the computing device and / or the additional computing device. For example, the adapted noise reduction may be a noise cancellation filter that filters audio data provided for rendering through the additional speaker from captured audio data (captured via a microphone), and the audio delay may be utilized to accurately determine a predicted timing for actually rendering the audio data through the additional speaker.

[0007] Implementations described herein additionally and / or alternatively relate to determining whether an audio data stream transmitted to an additional computing device (e.g., a vehicle computing device) for rendering through one or more additional speakers (e.g., vehicle speakers) driven by the additional computing device is actually being rendered through the additional speakers. If so, the additional audio data may be transmitted to the additional computing device for rendering through the additional speakers (based on the assumption that the audio data will similarly be rendered through the additional speakers). If not, the additional audio data may instead be rendered using an alternative speaker not driven by the additional computing device. In these and other methods, when it is determined that previously transmitted audio data was actually rendered through the vehicle speakers, the additional audio data may be transmitted for rendering through the vehicle speakers. However, when it is determined that the previously transmitted audio data was not actually rendered through the vehicle speakers, the additional audio data may instead be rendered through the alternative speakers. This ensures that the audio data is actually being rendered and can be perceived by the user, and prevents the user from being forced to request a retransmission of the audio data and / or another rendering attempt. Additionally, this may optionally ensure that transmission of audio data to the additional computing device occurs only when it is determined that the data is to be audibly rendered by the additional computing device.

[0008] As an example, a smartphone can be communicatively coupled to a vehicle computing device via Bluetooth, and the smartphone can transmit audio data to the vehicle computing device for rendering through the vehicle's speakers. For example, the automated assistant client on the smartphone can generate audio data in response to a request from a user and cause the audio data to be transmitted to the vehicle computing device in response to the request.

[0009] In many implementations, audio data transmitted from a computing device for rendering by the vehicle computing device using a vehicle speaker may be delayed by up to several seconds before the audio data is rendered through the vehicle speaker. For example, when the vehicle computing device switches to Bluetooth mode, a portion (e.g., 1 second, 1.5 seconds, etc.) of the audio data transmitted to the vehicle computing device may fail to render on the vehicle speaker. In many implementations, this portion of the audio data is received by the vehicle computing device but discarded by the vehicle computing device (i.e., the vehicle computing device discards any received portions of the audio data stream until the vehicle computing device switches to Bluetooth mode). As a result, this portion of the audio data is not rendered through the vehicle speaker at all, even though it is intended to be rendered through the vehicle speaker. This can be problematic for various audio data, such as audio data capturing a natural language response generated using an automated assistant client. Such natural language responses are often short, and relevant portions (or even the entirety) of the natural language response may fail to render through the vehicle speaker.

[0010] In many implementations, the audio delay between the computing device and the vehicle computing device can be automatically determined, and a delayed audio segment, the size of which is determined based on the length of the delay, can be added to future audio data. For example, the client device can determine a 0.5-second delay in the audio data to be sent to the vehicle computing device and therefore add a delayed audio segment containing 0.5 seconds of delayed audio to the beginning of the future audio data stream to be sent to the vehicle computing device. In this way, adding the delayed audio segment to the future audio data stream ensures that the entire future audio stream will be rendered, even if the delayed audio segment will not be rendered (at least completely). The audio delay can occur at the beginning of the audio data stream, at the end of the audio data stream, or at both the beginning and the end of the audio data stream. Thus, the delayed audio segment can be added to the beginning, the end, or both the beginning and the end of the audio data stream.

[0011] To determine the vehicle audio delay, the computing device may transmit a known sequence of audio data to the vehicle computing device. The vehicle computing device may render the audio data using one or more vehicle speakers, and the computing device may capture the rendering of the audio data. For example, a mobile smartphone may transmit an audio data stream to the vehicle computing device and capture the audio output produced using the vehicle's speakers. By comparing the captured audio data to a known audio data stream, the computing device may determine the vehicle audio delay. This audio data sequence may be audible to the user, inaudible to the user (e.g., high-frequency audio), and / or a combination of audible and inaudible audio data. For example, the audio data stream may include a segment of audio data at a single frequency of a known length. The client device may compare the length of the captured audio data to the known length of the audio data stream to determine the delay. Additionally and / or alternatively, the audio data stream may include a sequence of frequency segments. The captured sequence of frequency segments may be compared to the transmitted sequence of frequency segments to determine the delay. In various implementations, background noise may interfere with the captured audio output rendered using the vehicle speakers (i.e., traffic outside the vehicle, people talking inside the vehicle, etc. may interfere with the captured audio output). The audio data stream may include a sequence of co-occurring frequency segments (e.g., a dual-tone frequency segment, a tri-tone frequency segment, a quad-tone frequency segment, etc.). In many cases, the computing device may still capture at least one frequency in the co-occurring frequency segment despite the background noise.In various implementations, the audio data used to determine the vehicle audio delay is a sequence of dual-tone multi-frequency (DTMF) audio.

[0012] In various implementations, the vehicle interface device may additionally be communicatively coupled to the computing device. The vehicle interface device may provide additional and / or alternative user interface inputs and / or outputs, such as additional microphones and / or additional speakers. For example, the vehicle interface device may be communicatively coupled to the computing device via Bluetooth and may include one or more additional microphones for capturing the audio output of the vehicle speakers as well as speech emitted by the user. In some implementations, the microphone of the vehicle interface device may be better positioned and / or more suitable for capturing audio output than the microphone of the computing device. For example, a user may have their mobile smartphone in a backpack while driving, and the backpack may prevent the microphone of the mobile smartphone from capturing audio output to the same extent as the microphone of the vehicle interface device (which is not blocked by the backpack). As another example, the vehicle interface device may include a far-field microphone that may be better equipped to capture various speeches within the vehicle, whereas the smartphone may not have a far-field microphone. Additionally or alternatively, the vehicle interface device may detect spoken invocation of an invocation phrase and, after detection of the invocation phrase, transmit audio data to the computing device.

[0013] In many implementations, the client device may send the audio data stream to the vehicle computing device to determine whether the audio data stream will actually be rendered using one or more vehicle speakers. If the client device (or a vehicle interface device communicatively coupled to the client device) does not capture (via a microphone) any audio data output corresponding to the audio data stream, the client device may render the future audio data stream using an alternative speaker (e.g., the client device speaker, the vehicle interface device speaker, etc.). On the other hand, if an audio data output corresponding to the audio data stream is captured, the client device may send the future audio data stream to the vehicle computing device for rendering through the vehicle speakers. In many implementations, the client device may send the audio data stream on a periodic basis (e.g., every minute, every two minutes, every five minutes, etc.) or at other regular or non-regular intervals. Additionally or alternatively, the client device may determine whether to send the audio data stream based on the amount of time elapsed since the client device last sent the audio data stream. For example, the client device will transmit the audio data stream if a threshold time has passed since the last transmission (e.g., the client device will transmit the audio data stream if 10 seconds have passed since the last transmission, 30 seconds have passed since the last transmission, 1 minute has passed since the last transmission, etc.). In many implementations, the client device may transmit the audio data stream in response to detecting a call to the automated assistant (e.g., in response to detecting a hotword or actuation of a call button) while the automated assistant client is "busy" determining a response to user-provided input.In these and other methods, the client device may dynamically update its determination as to whether the audio data stream should be provided for rendering through the vehicle speakers or, alternatively, through an alternative speaker.

[0014] In some implementations, the determined audio delay may additionally or alternatively be used with various noise reduction techniques. For example, a user may provide a verbal input "OK Assistant, what time?" to the automated assistant client, and the automated assistant client may respond with "It's 3:05 PM." The client device may transmit an audio data stream including a text-to-speech conversion of "It's 3:05 PM" to a vehicle computing device, which may render the audio data stream through a vehicle speaker. The client device and / or a separate vehicle interface device may utilize the transmitted audio data stream and the determined audio delay when erasing the text-to-speech conversion of "It's 3:05 PM" from the captured audio data (captured via a microphone). In other words, knowledge of the audio data stream being rendered may be used to erase the audio data stream from the captured audio data, thereby enabling better recognition of any co-occurring verbal utterances of the user. Knowledge of the audio data stream is utilized along with the determined audio delay to enable noise cancellation to eliminate the audio data stream at the appropriate time (e.g., to know that "it is 3:05 PM" will actually be rendered with a 1.2 second delay (or other delay)). For example, the client device may send the "it is 3:05 PM" audio stream to a vehicle interface device for use by the vehicle interface device in noise reduction. The vehicle interface device will use the audio stream and the vehicle audio delay to know when the audio output will be rendered using the vehicle speakers and may filter "it is 3:05 PM" at the appropriate time from any captured audio output sent to the client device.

[0015] It should be appreciated that all combinations of the foregoing concepts and additional concepts described in detail herein are contemplated as part of the presently disclosed subject matter, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as part of the presently disclosed subject matter. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 illustrates an exemplary environment in which various implementations disclosed herein may be implemented. [Figure 2] FIG. 1 illustrates another exemplary environment in which various implementations disclosed herein may be implemented. [Figure 3] FIG. 1 illustrates another exemplary environment in which various implementations disclosed herein may be implemented. [Figure 4] 1A-1C illustrate various examples of an exemplary audio data stream and captured audio data according to various implementations disclosed herein. [Figure 5] 1 is a flowchart illustrating an example process according to various implementations disclosed herein. [Figure 6] 10 is a flowchart illustrating another example process according to various implementations disclosed herein. [Figure 7] FIG. 1 is a block diagram illustrating an example environment in which various implementations disclosed herein may be implemented. [Figure 8] FIG. 1 illustrates an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION

[0017] Figures 1, 2, and 3 illustrate a computing device communicatively coupled to a vehicle computing device according to many implementations described herein. While Figures 1-3 illustrate the computing device and vehicle interface device (Figures 2 and 3) external to the vehicle for simplicity, it should be understood that the computing device and / or vehicle interface device would be located within the vehicle during performance of the various techniques described herein.

[0018] 1 illustrates a computing device 106 communicatively coupled to a vehicle computing device 102 via a wireless communication channel 104. The computing device 102 may be, for example, a laptop computing device, a tablet computing device, a mobile smartphone computing device, and / or a user-wearable device including a computing device (e.g., a watch with a computing device, glasses with a computing device, or a virtual or augmented reality computing device, etc.). Additional and / or alternative client devices may be provided. In various implementations, the computing device 106 includes various user interface input and / or output devices, such as a microphone, a speaker, and / or additional user interface devices. The computing device 106 may be mounted within the vehicle (e.g., on a car mount, suctioned to a window) and / or may be powered and / or charged by auxiliary power provided by the vehicle (e.g., a 12V vehicle receptacle, a USB port, or an auxiliary standard plug, such as a “Type A” plug in the United States). However, the computing device 106 is not integrated with the vehicle, can be easily removed from the vehicle and placed in another vehicle, and may be a smartphone or other device utilized by users in a variety of environments.

[0019] The vehicle computing device 102 of the vehicle may be, for example, an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system, etc. Additional and / or alternative vehicle computing devices may be provided. In various implementations, the vehicle computing device 102 is integrated with the vehicle and directly drives vehicle speakers that are also integrated with the vehicle. The vehicle computing device 102 may be original equipment for the vehicle or an aftermarket installed accessory. The vehicle computing device 102 is integrated such that it directly drives vehicle speakers and / or does not require the use of specialized tools and / or cannot be removed from the vehicle without requiring significant time and / or expertise. For example, the vehicle computing device 102 may be connected to the vehicle's controller area network (CAN) bus and / or may be powered via a vehicle-specific connector (e.g., rather than a 12V vehicle receptacle or an easily accessible auxiliary standard plug). In many implementations, the vehicle computing device 102 may include various user interfaces, including a microphone, a speaker, and / or additional user interface devices. For example, the audio input may be rendered through one or more vehicle speakers driven by the vehicle computing device.

[0020] The wireless communication channel 104 may include various wireless communication networks that may optionally utilize one or more standard communication technologies, protocols, and / or inter-process communication techniques. For example, the wireless communication channel 104 may be a Bluetooth channel 104, and the mobile smartphone computing device 106 may be communicatively coupled to the vehicle computing device 102 via the Bluetooth channel 104. As a further example, the client device 106 may transmit an audio data stream to the vehicle computing device 102 via the Bluetooth channel 104, which can cause the vehicle computing device 102 to render a corresponding audio output that may be captured by a microphone within the vehicle, and this captured data may be used to determine vehicle audio delay.

[0021] 2 shows computing device 206 communicatively coupled to vehicle computing device 202 via wireless communications network 204. Additionally, computing device 206 is communicatively coupled to vehicle interface device 210 via wireless communications network 208. As described above with respect to FIG. 1, computing device 206 may include various computing devices, vehicle computing device 202 may include various computing devices of a vehicle, and / or wireless communications channels 204 and 208 may include various communications channels.

[0022] In various implementations, the computing device 206 may additionally and / or alternatively be coupled to the vehicle interface device 210 via a wireless communication channel 208. The vehicle interface device 210 may provide additional and / or alternative user interface inputs and / or outputs, such as one or more additional microphones, one or more additional speakers, one or more additional buttons, etc. In various implementations, the vehicle interface device 210 may be powered using a 12V vehicle receptacle (also referred to herein as a cigarette lighter receptacle), a vehicle USB port, a battery, etc. For example, the vehicle interface device 210 may be powered by a 12V receptacle of the vehicle and may be located on or around the center console of the vehicle (i.e., located near the driver of the vehicle such that one or more microphones of the vehicle interface device 210 may capture verbal utterances provided by the driver and / or additional vehicle passengers). The computing device 206, such as a mobile smartphone, may be communicatively coupled to the vehicle interface device 210 via the wireless communication channel 210. As a further example, a mobile smartphone computing device 206 may be communicatively coupled to the vehicle interface device 202 via a first Bluetooth channel 204, and the computing device 206 may be communicatively coupled to the vehicle interface device 210 via a second Bluetooth channel 208.

[0023] 3 illustrates an alternative configuration of computing devices communicatively coupled to a vehicle computing device and a vehicle interface device. The computing device 304, the vehicle interface device 302, and / or the vehicle interface device 308 are described above with respect to FIGS. 1 and 2. In various implementations, the vehicle is not communicatively coupled to the computing device via a wireless communication channel (e.g., the vehicle may lack the ability to connect to a computing device via a wireless communication channel). In some such implementations, the computing device 304 may be communicatively coupled to the vehicle interface device 308 via a wireless communication channel 306. Additionally, the vehicle interface device 308 may be communicatively coupled to the vehicle computing device via a communication channel 310. For example, a mobile smartphone (i.e., the computing device 304) may be communicatively coupled to the vehicle interface device 308 via a Bluetooth channel (i.e., the wireless communication channel 306). The vehicle interface device 308 may additionally or alternatively be communicatively coupled to the vehicle computing device 302 via an auxiliary cable (i.e., the communication channel 310).

[0024] In various implementations, a computing device (e.g., 106 of FIG. 1 , 206 of FIG. 2 , and / or 304 of FIG. 3 ) may automatically determine the vehicle device delay by sending an audio data stream to a vehicle computing device (e.g., 102 of FIG. 1 , 202 of FIG. 2 , and / or 302 of FIG. 3 ) and comparing the captured audio output (rendered using one or more vehicle speakers) with the audio data stream. Audio data streams according to many implementations are described herein with reference to FIG. 4 . In many implementations, the captured audio output rendered by the vehicle speakers may be captured using one or more microphones of the computing device and / or one or more microphones of the vehicle interface device.

[0025] In various implementations, once the delay is determined, the delayed audio data may be added onto a future audio data stream, where the determined delay is used to determine the length of the delayed audio data. Additionally or alternatively, the determined delay may be utilized as part of a noise reduction process.

[0026] In many implementations, the audio data stream is transmitted to determine whether the rendered audio output can be captured through one or more vehicle speakers. That is, a test audio signal may be transmitted to the vehicle computing device, and if the computing device and / or vehicle interface device is unable to capture the rendered audio output through the vehicle speakers, future audio data streams may be rendered using the speakers of the computing device and / or the vehicle interface device.

[0027] While the implementations described herein relate to a computing device communicatively coupled to a vehicle computing device, it should be understood that additional or alternative computing devices may be coupled to the computing device. For example, the computing device may be communicatively coupled to a computing device of a stand-alone wireless speaker (e.g., a mobile smartphone communicatively coupled to a Bluetooth wireless speaker). The computing device may be coupled to additional and / or alternative computing devices.

[0028] 4 illustrates an exemplary audio data stream and various captured audio data according to various implementations. Audio data stream 402 includes a sequence of five frequency segments: frequency segment "1" 404, frequency segment "2" 406, frequency segment "3" 408, frequency segment "4" 410, and frequency segment "5" 412. In many implementations, a computing device transmits audio data stream 402 to a vehicle computing device for rendering using vehicle speakers. Corresponding audio output rendered using the vehicle speakers can then be captured and compared to audio data 402 to determine any vehicle audio delay.

[0029] For example, the vehicle audio delay may be shorter than the first frequency segment. The captured audio data 414 exhibits a delay approximately half the length of the first frequency segment 404 and captures the sequence frequency segment “1” 416, frequency segment “2” 418, frequency segment “3” 420, frequency segment “4” 422, and frequency segment “5” 424. Due to the audio device delay, frequency segment “1” 416 of the captured audio data 414 is shorter than frequency segment “1” 404 of the audio data stream 402. In many implementations, the delayed audio segment may be determined using the difference between the end of frequency segment “1” 416 and the end of frequency segment “1” 404. Additional frequency segments “2,” “3,” “4,” and / or “5” will have similar delays, and the computing device may additionally and / or alternatively determine the delays using the additional captured frequency segments. For example, the audio data stream may be 2.5 seconds long and include five 0.5-second-long frequency segments. The captured audio data may capture frequency segment "1" that is 0.3 seconds long (i.e., the captured audio data may capture 2.3-second frequency segments). The computing device may compare frequency segment "1" 404 with frequency segment "1" 416 to determine a 0.2-second delay. Similarly, the computing device may compare frequency segment "2" 406 with frequency segment "2" 418 to determine a 0.25-second delay, frequency segment "3" 408 with frequency segment "3" 420 to determine a 0.2-second delay, frequency segment "4" 410 with frequency segment "4" 422 to determine a 0.3-second delay, and frequency segment "5" 412 with frequency segment "5" 424 to determine a 0.2-second delay. The computing device may select 0.3 seconds as the delay (i.e., 0.3 seconds is the largest delay among the determined delays of 0.2 seconds, 0.25 seconds, 0.2 seconds, 0.3 seconds, and 0.2 seconds).

[0030] In many implementations, entire frequency segments may be missing in the captured audio data. The system may compare the frequency segments in the audio data stream 402 to captured audio data capturing the sequence of frequency segment “2” 428, frequency segment “3” 430, frequency segment “4” 432, and frequency segment “5” 434. In other words, frequency segment “1” 404 of the audio data stream 402 does not have a corresponding representation in the captured audio data stream 426. For example, the audio data stream 402 may be five seconds long and include five one-second frequency segments. The computing device may determine that the captured audio data stream 426 does not include any of the frequency segments “1” 404. To determine the one-second delay, the number of missing frequency segments may be multiplied by the one-second-long frequency segments in the audio data stream 402.

[0031] In many implementations, entire frequency segments may be missing, as well as portions of frequency segments. Captured audio data 436 shows captured audio in which frequency segment “1” and frequency segment “2” are missing in their entirety and a portion of frequency segment “3” is missing. In other words, captured audio data 436 includes frequency segment “3” 438, frequency segment “4” 440, and frequency segment “5” 442, where frequency segment “3” 438 of captured audio data 436 is shorter than frequency segment “3” 408 of audio data stream 402. Device delay may be determined using a combination of the missing frequency segments and the length of the missing portion of the first captured frequency segment, as described above. For example, audio data stream 402 may include five frequency segments that are 0.3 seconds long (i.e., audio data stream 402 is 1.5 seconds long). The captured audio data stream may capture only 0.7 seconds of audio data stream 402. The 0.7 second delay is determined by comparing the captured audio data stream 436 to the audio data stream 402 and determining that frequency segments corresponding to frequency segment "1" 404 and frequency segment "2" 406 are not captured in the captured audio data stream 436. Additionally, by comparing frequency segment "3" 408 to the captured frequency segment "3" 438, it may be determined that only the 0.1 second frequency segment "3" 438 is captured. The computing device may determine the delay by combining the delay of the missing frequency segment (0.3 seconds from missing frequency segment "1" + 0.3 seconds from missing frequency segment "2") with the delay of the first captured frequency segment "3" 438 (0.2 seconds) to determine a delay of 0.8 seconds (0.3 + 0.3 + 0.2).

[0032] Additionally or alternatively, the captured audio data may be missing portions from both the beginning and end of the audio data stream. For example, captured audio data 444 includes frequency segment “2” 446 and frequency segment “3” 448, where frequency segment “2” 446 is shorter than frequency segment “2” 406 of audio data stream 402. In other words, within captured audio data 444, frequency segments “1,” “4,” and “5” are completely missing, and a portion of frequency segment “2” is missing. A first vehicle delay may be determined based on the missing frequency segment “1” and the missing portion of frequency segment “2.” Additionally or alternatively, a second vehicle delay may be determined based on the missing frequency segments “4” and “5.” For example, audio data stream 402 may include five frequency segments, each one second long (i.e., the audio data stream is five seconds long). By comparing the audio data stream 402 with the captured audio data stream 444, it may be determined that the captured audio data stream 444 does not capture frequency segments corresponding to frequency segment “1” 404, frequency segment “4” 410, and frequency segment “5” 412. Additionally, an additional 0.4 second delay may be determined by comparing captured frequency segment “2” 446 with frequency segment “2” 406 and capturing frequency segment “3” 448 with frequency segment “3” 408. A first audio delay occurring at the beginning of the captured audio data stream may be determined to be 1.4 seconds by combining the delay of captured frequency segment “2” (0.4 seconds) with the length of missing frequency segment “1” (1 second). Additionally or alternatively, a second audio delay occurring at the end of the 2 second captured audio data stream may be determined by combining the length of missing frequency segment “4” (1 second) with the length of missing frequency segment “5” (1 second).

[0033] Although a particular sequence of frequency segments is described with respect to FIG. 4 , various audio data streams (and corresponding captured audio data) according to many implementations may be used. For example, the audio data stream may be a segment of a single frequency. For example, the audio data stream may be an 8-second long segment of a single frequency, and the captured audio data may capture only a single frequency for 6.5 seconds, and a 1.5-second vehicle audio delay may be determined based on a comparison of the expected duration of the segment (8 seconds) to its actual duration (6.5 seconds) in the captured audio data. As another example, each frequency segment may be several co-occurring frequencies (e.g., dual-tone co-occurring frequencies, tri-tone co-occurring frequencies, etc.). In many implementations, the sequence of frequency segments includes a non-repeating sequence of frequency segments. In many implementations, the sequence of frequency segments includes repeating frequency segments, where missing frequency segments are uniquely identifiable. For example, the sequence may be a frequency segment representation of “1”, “2”, “3”, “4”, “5”, “4”, “3”, “2”, “1”. The audio data stream may be of various lengths, such as 0.5 seconds, 1 second, 1.5 seconds, 2 seconds, and so on.

[0034] Referring to FIG. 5 , an example process 500 for determining vehicle audio delay is shown, according to implementations disclosed herein. For convenience, operations of some aspects of the flowchart of FIG. 5 are described with reference to a system that performs those operations. The system may include various components of various computer systems and / or one or more of a GPU, a CPU, and / or a TPU. For example, the system may include a smartphone or other computing device and / or a vehicle interface device. Furthermore, although the operations of process 500 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.

[0035] In block 502, the system causes the computing device to transmit an audio data stream to the vehicle computing device over a wireless communication channel. For example, a mobile smartphone may transmit the audio data stream to the vehicle computing device over Bluetooth. As another example, the mobile smartphone may transmit the audio data stream to the vehicle interface device over Bluetooth, and the vehicle interface device may transmit the audio data stream to the vehicle computing device. As yet a further example, the vehicle interface device may transmit the audio data stream to the vehicle computing device over Bluetooth and / or a wired communication channel.

[0036] At block 504, the system causes the vehicle computing device to render an audible output generated using the audio data stream through one or more speakers of the vehicle, the one or more speakers of the vehicle being driven by the vehicle computing device. For example, the vehicle computing device may drive a vehicle speaker integrated with the vehicle based on all or a portion of the audio data stream, thereby causing the vehicle speaker to render a corresponding audible output. As described herein, if the vehicle computing device has no delay, the corresponding audible output will include the entire audio data stream. However, if the vehicle computing device has a delay, the corresponding audible output may omit one or more portions of the audio data stream.

[0037] At block 506, the system receives captured audio data that captures the audible output rendered through one or more speakers of the vehicle. The captured audio data is captured by at least one microphone in the vehicle. In some implementations, the at least one microphone in the vehicle includes a microphone of a computing device, such as the computing device that transmitted the audio data stream in block 502. In some implementations, the at least one microphone in the vehicle additionally or alternatively includes a microphone of a vehicle interface device, which may be separate from the computing device that transmitted the audio data stream in block 502. Additionally or alternatively, the audible output may be captured by both the at least one microphone of the computing device as well as the at least one microphone of the vehicle interface device.

[0038] At block 508, the system determines the vehicle audio delay by comparing the captured audio data to the audio data stream. Several non-limiting examples of determining the vehicle audio delay are described herein (e.g., above with respect to FIG. 4).

[0039] In block 510, the system determines whether there is an additional audio data stream to send to the vehicle computing device. In many implementations, the automated assistant client of the computing device generates the additional audio data stream. In many implementations, the automated assistant client of the vehicle interface device generates the additional audio data stream. If so, the system proceeds to block 512, where the system adds a delayed audio segment to the additional audio data stream, the duration of the delayed audio segment being determined using the vehicle audio delay. In various implementations, the delayed audio segment may include a variety of audio, including white noise, high-frequency segments of sounds impossible to humans, and additional other sounds. The delayed audio segment may be a single length that is repeated as needed (i.e., a 0.2-second delayed audio segment may be added once for each 0.1-second delay and 0.2-second delay, a 0.2-second delayed audio data segment may be added twice for each 0.3-second delay and 0.4-second delay, etc.). Additionally or alternatively, the length of the delayed audio segment may be customized to the determined audio delay (i.e., when a 0.5 second delay is determined, a 0.5 second delayed audio segment may be added, when a 0.75 second delay is determined, a 0.75 second delayed audio segment may be added, etc.). Furthermore, a delayed audio segment slightly longer than the determined audio delay may be added (i.e., when a 0.25 second audio delay is determined, a 0.3 second delayed audio segment may be added, when a 0.5 second audio delay is determined, a 0.75 second delayed audio segment may be added, etc.).

[0040] At block 514, the system causes the computing device to transmit the additional audio stream with the added delayed audio segment to the vehicle computing device over the wireless communication channel. Once the system has transmitted the additional audio data stream, the process ends.

[0041] If, at block 510, the system determines that there are no additional audio data streams to send to the vehicle computing device, the system proceeds to block 516, where the system determines whether a noise cancellation filter exists. If the system determines that a noise cancellation filter does not exist, the process ends. If, at block 516, the system determines that a noise cancellation filter exists, the system proceeds to block 518, where the system causes the computing device to adapt the noise cancellation filter based on the vehicle audio delay before the process ends. In many implementations, the noise cancellation filter is stored locally on the computing device. In many implementations, the noise cancellation filter is stored in a separate computing device (e.g., a separate vehicle interface device). If the noise cancellation filter is stored in a separate computing device, block 512 may include sending data to the separate computing device based on the vehicle audio delay and causing the computing device to adapt its local noise cancellation filter based on the vehicle audio delay.

[0042] 5 shows a process that includes both adding a delayed audio segment based on a determined vehicle audio delay and adapting a noise cancellation filter based on the determined vehicle audio delay. However, as described herein, in various implementations, the delayed audio segment may be added without any adaptation of the noise cancellation filter, or adaptation of the noise cancellation filter may occur without any addition of the delayed audio segment.

[0043] Referring to FIG. 6 , an example process 600 for determining whether one or more speakers driven by a vehicle computing device render an audible output generated using an audio data stream is shown, according to implementations disclosed herein. For convenience, the operations of some aspects of the flowchart of FIG. 6 are described with reference to a system that performs those operations. The system may include various components of various computer systems and / or one or more of a GPU, a CPU, and / or a TPU. For example, the system may include a smartphone or other computing device and / or a vehicle interface device. Furthermore, although the operations of process 600 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.

[0044] In block 602, the system determines whether to transmit an audio data stream from the computing device to the vehicle computing device over a communication channel. In many implementations, the system determines whether the vehicle is in a communication channel mode (i.e., whether the vehicle is in Bluetooth mode, whether the vehicle supports automatic switching to Bluetooth mode, etc.). In many implementations, the system determines whether the volume of one or more speakers driven by the vehicle computing device is too low for the rendered audio output to be captured via one or more microphones in the vehicle. If the system determines that the vehicle is in a communication channel mode (or supports automatic switching to a communication channel mode) and the volume of the speakers driven by the vehicle computing device is not too low, the system proceeds to block 604. If the system determines that the vehicle is not in a communication channel mode or the system determines that the volume of the speakers driven by the vehicle computing device is too low, the system proceeds to block 612.

[0045] In block 604, the system causes the computing device to transmit the audio data stream to the vehicle computing device over a communication channel. In some implementations, the communication channel is a wireless communication channel (e.g., a Bluetooth channel). In other implementations, the communication channel is a wired communication channel (e.g., an auxiliary cable).

[0046] At block 606, the system causes the vehicle computing device to render an audible output generated based on the audio data stream through one or more speakers driven by the vehicle computing device.

[0047] In block 608, the system determines whether the audible output is captured by at least one microphone in the vehicle. If the system determines that the audible output is captured by at least one microphone, the system proceeds to block 610. If the system determines that the audible output is not captured by at least one microphone, the system proceeds to block 612. In many implementations, the audible output is captured by at least one microphone of the computing device. In many implementations, the audible output is captured by at least one microphone of the vehicle interface device. In many implementations, the audible output is captured by at least one microphone of the computing device and at least one microphone of the vehicle interface device.

[0048] At block 610, the system causes the computing device to transmit the additional audio data stream to the vehicle computing device for rendering through one or more speakers driven by the vehicle computing device.

[0049] In block 612, the system causes the additional audio data stream to be rendered on one or more alternative speakers in the vehicle, the one or more alternative speakers not driven by the vehicle computing device. In many implementations, the one or more alternative speakers are speakers of the computing device. In many implementations, the one or more alternative speakers are speakers of the vehicle interface device.

[0050] Referring to Figure 7, an example environment in which implementations disclosed herein may be implemented is shown. Figure 7 includes a client computing device 702 that executes an instantiation of an automated assistant client 704. One or more cloud-based automated assistant components 712 may be implemented on one or more computing systems (collectively "cloud" computing systems) communicatively coupled to the client device 702 via one or more local area and / or wide area networks (e.g., the Internet), generally indicated by 710.

[0051] An instance of automated assistant client 704, through its interaction with one or more cloud-based automated assistant components 712, may form what appears from a user's perspective as a logical instance of automated assistant 700 with which the user may engage in a human-computer dialogue. It should be understood, therefore, that in some implementations, a user engaging with an automated assistant client 704 running on a client device 702 may in effect engage with their own logical instance of automated assistant 700. For brevity and simplicity, the term "automated assistant," as used herein as "serving" a particular user, will often refer to the combination of the automated assistant client 704 running on the user-operated client device 702 and one or more cloud-based automated assistant components 712 (which may be shared among multiple automated assistant clients on multiple client computing devices). It should also be understood that in some implementations, automated assistant 700 may respond to requests from any user, regardless of whether the user is actually "served" by that particular instance of automated assistant 700.

[0052] The client computing device 702 may be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile smartphone computing device, a standalone interactive speaker, a smart appliance, and / or a user-wearable device including a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided. Additionally or alternatively, the operations of the client computing device 702 may be distributed among multiple computing devices. For example, one or more operations of the client computing device 702 may be distributed between a mobile smartphone and a vehicle computing device. Furthermore, the operations of the client computing device 702 may be replicated among multiple computing devices (which may in some cases be communicatively coupled). As a further example, a mobile smartphone and a vehicle interface device may each implement the operations of the automated assistant 700, such as a mobile smartphone and a vehicle interface device, both of which include a call engine (described below). In various implementations, client computing device 702 may optionally run one or more other applications, such as a messaging client (e.g., SMS, MMS, online chat), a browser, etc., in addition to automated assistant client 704. In some of these various implementations, one or more of the other applications may optionally interface with automated assistant 704 (e.g., via an application programming interface) or may include its own instance of an automated assistant application (which may also interface with cloud-based automated assistant component 712).

[0053] The automated assistant 700 engages in a human-to-computer dialog session with the user via the user interface input and output devices of the client device 702. To protect the user's privacy and / or to conserve resources, in many situations the user must often explicitly invoke the automated assistant 700 before the automated assistant will fully process the verbal utterance. Explicit invocation of the automated assistant 700 can occur in response to certain user interface input received at the client device 702. For example, user interface inputs that can invoke the automated assistant 700 via the client device 702 can optionally include activation of hardware and / or virtual buttons on the client device 702. Additionally, the automated assistant client can include one or more local engines 708, such as an invocation engine, operable to detect the presence of one or more verbal invocation phrases. The invocation engine can invoke the automated assistant 700 in response to detecting one or more of the verbal invocation phrases. For example, the invocation engine may invoke the automated assistant 700 in response to detecting a verbal invocation phrase, such as "hey, assistant," "OK assistant," and / or "assistant." The invocation engine may continuously process (e.g., when not in "inactive" mode) a stream of audio data frames based on output from one or more microphones of the client device 702 to monitor for the occurrence of the verbal invocation phrase. While monitoring for the occurrence of the verbal invocation phrase, the invocation engine discards (e.g., after temporary storage in a buffer) any audio data frames that do not contain the verbal invocation phrase. However, when the invocation engine detects the occurrence of the verbal invocation phrase within a processed audio data frame, the invocation engine may invoke the automated assistant 700. As used herein, "invoking" the automated assistant 700 may include activating one or more previously inactive features of the automated assistant 700.For example, invoking the automated assistant 700 may include having one or more local engines 708 and / or cloud-based automated assistant component 712 further process the audio data frame based on which the invocation phrase was detected and / or one or more subsequent audio data frames (whereas no further processing of the audio data frames occurred prior to the invocation).

[0054] The one or more local engines 708 of the automated assistant 704 are optional and may include, for example, the invocation engine described above, a local speech-to-text ("STT") engine (which converts captured audio to text), a local text-to-speech ("TTS") engine (which converts text to speech), a local natural language processor (which determines the semantic meaning of the audio and / or text converted from the audio), and / or other local components. Because the client device 702 is relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the local engines 108 may have limited functionality relative to any counterparts included in the cloud-based automated assistant component 712.

[0055] The automated assistant client 704 may additionally include a delay engine 706 and audio data 720. The delay engine 706 may be utilized by the automated assistant client 704 according to various implementations, including sending an audio data stream to a vehicle computing device, sending the audio data stream to a vehicle interface device, determining a vehicle device delay, adding an audio delay segment to the audio data stream, sending the vehicle device delay to the vehicle interface device, capturing audio data rendered using a vehicle speaker, etc. In many implementations, the delay engine 706 may select an audio data stream from an audio data database 720.

[0056] The cloud-based automated assistant component 712 leverages the virtually unlimited resources of the cloud to perform more robust and / or more accurate processing of audio data and / or other user interface input relative to any counterpart of the local engine 708. Again, in various implementations, the client device 702 may provide the audio data and / or other data to the cloud-based automated assistant component 712 in response to detection of a spoken invocation phrase by the invocation engine or some other explicit invocation of the automated assistant 700.

[0057] The illustrated cloud-based automated assistant component 712 includes a cloud-based TTS module 714, a cloud-based STT module 716, and a natural language processor 718. In some implementations, one or more of the engines and / or modules of the automated assistant 700 may be omitted, combined, and / or implemented in a component separate from the automated assistant 700. Additionally, in some implementations, the automated assistant 700 may include additional and / or alternative engines and / or modules.

[0058] The cloud-based STT module 716 can convert the audio data to text, which may then be provided to the natural language processor 718. In various implementations, the cloud-based STT module 716 may convert the audio data to text based at least in part on the speaker label indications and assignments provided by an assignment engine (not shown).

[0059] The cloud-based TTS module 714 may convert text data (e.g., natural language responses composed by the automated assistant 700) into computer-generated voice output. In some implementations, the TTS module 714 may provide the computer-generated voice output to the client device 702 for direct output, e.g., using one or more speakers. In other implementations, the text data (e.g., natural language responses) generated by the automated assistant 700 may be provided to one of the local engines 708, which may then convert the text data into computer-generated voice for local output.

[0060] The natural language processor 718 of the automated assistant 700 processes free-form natural language input and generates annotated output based on the natural language input for use by one or more other components of the automated assistant 700. For example, the natural language processor 718 may process natural language free-form input, which is text input that is a transformation by the STT module 716 of audio data provided by a user via the client device 702. The generated annotated output may include one or more annotations of the natural language input and, optionally, one or more (e.g., all) of the terms of the natural language input. In some implementations, the natural language processor 718 is configured to identify and annotate various types of grammatical information within the natural language input. For example, the natural language processor 718 may include a portion of a speech tagger (not shown) configured to annotate terms with their grammatical roles. Also, for example, in some implementations, the natural language processor 718 may additionally and / or alternatively include a dependency parser (not shown) configured to determine semantic relationships between terms of the natural language input.

[0061] In some implementations, the natural language processor 718 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references in one or more segments, such as references to people (e.g., including literary characters, famous people, public figures, etc.), organizations, locations (real and fictional), etc. The entity tagger of the natural language processor 718 may annotate references to entities at a high level of granularity (e.g., to enable identification of all references to an entire class, such as people) and / or at a low level of granularity (e.g., to enable identification of all references to a particular entity, such as a particular person). The entity tagger may rely on the content of the natural language input to parse particular entities and / or may optionally communicate with a knowledge graph or other entity database to parse particular entities.

[0062] In some implementations, the natural language processor 718 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, a coreference resolver may be utilized to parse the term "there" into "Hypothetical Cafe" in the natural language input "The last time I was there, I liked Hypothetical Cafe."

[0063] In some implementations, one or more components of the natural language processor 718 may rely on annotations from one or more other components of the natural language processor 718. For example, in some implementations, a named entity tagger may rely on annotations from a coreference resolver and / or a dependency parser when annotating all mentions to a particular entity. Also, for example, in some implementations, a coreference resolver may rely on annotations from a dependency parser when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 718 may use related previous input and / or other relevant data outside the particular natural language input to determine one or more annotations.

[0064] 8 is a block diagram of an example computing device 810 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client computing devices and / or other components may include one or more components of the example computing device 810.

[0065] Computing device 810 typically includes at least one processor 814 that communicates with several peripheral devices via a bus subsystem 812. These peripheral devices may include, for example, a storage subsystem 824 including a memory subsystem 825 and a file storage subsystem 826, a user interface output device 820, a user interface input device 822, and a network interface subsystem 816. The input and output devices enable user interaction with computing device 810. Network interface subsystem 816 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.

[0066] The user interface input devices 822 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphical tablet, a scanner, a touchscreen integrated into a display, a voice recognition system, an audio input device such as a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all conceivable types of devices and methods for inputting information into the computing device 810 or over a communications network.

[0067] The user interface output devices 820 may include non-visual devices such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all conceivable types of devices and methods for outputting information from the computing device 810 to a user, or to another machine or computing device.

[0068] The storage subsystem 824 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 824 may include logic for executing selected aspects of one or more of the processors of Figures 5 and / or 6, as well as for implementing the various components shown in Figure 7.

[0069] These software modules are generally executed by processor 814 alone or in combination with other processors. The memory 825 used in storage subsystem 824 may include several memories, including a main random access memory ("RAM") 830 for storing instructions and data during program execution and a read-only memory ("ROM") 832 in which fixed instructions are stored. The file storage subsystem 826 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored in storage subsystem 824 by file storage subsystem 826 or in other machines accessible by processor 814.

[0070] The bus subsystem 812 provides a mechanism for allowing the various components and subsystems of the computing device 810 to communicate with each other as intended. Although the bus subsystem 812 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0071] Computing device 810 may be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of computing device 810 shown in Figure 8 is intended to be merely a specific example to illustrate some implementations. Many other configurations of computing device 810 are possible, having more or fewer components than the computing device shown in Figure 8.

[0072] In some implementations, a method implemented by one or more processors is provided, the method comprising: causing a computing device to transmit an audio data stream to a vehicle computing device of a vehicle over a wireless communication channel, the transmitting step including causing the vehicle computing device to render an audible output through one or more vehicle speakers of the vehicle, the audible output being generated by the vehicle computing device based at least in part on the audio data stream. The method further includes receiving captured audio data captured by at least one microphone in the vehicle, the captured audio data capturing the audible output rendered by the at least one vehicle speaker. The method further includes determining a vehicle audio delay based on comparing the captured audio data to the audio data stream. In response to determining the vehicle audio delay, the method further includes causing the computing device to add a corresponding delayed audio segment to the additional audio data stream prior to transmitting the additional audio data stream to the vehicle computing device over the wireless communication channel, the duration of the delayed audio segment being determined using the vehicle audio delay.

[0073] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0074] In some implementations, determining the vehicle audio delay based on comparing the captured audio data with the audio data stream includes determining a temporal indication of the particular feature within the captured audio data. In some of those implementations, the method further includes determining an additional temporal indication of the particular feature within the audio data stream. In some of those implementations, the method further includes determining the vehicle audio delay based on a difference between the temporal indication of the particular feature within the captured audio data and the additional temporal indication of the particular feature within the audio data stream. In some versions of those implementations, the audio data stream includes a defined sequence of frequency segments, and the particular feature is a particular frequency segment of the defined sequence of frequency segments. In some versions of those implementations, each frequency segment of the sequence of frequency segments includes at least two corresponding co-occurring frequencies.

[0075] In some implementations, determining the temporal indication of the particular feature within the captured audio data includes determining a captured position of a particular frequency segment within the captured audio data, and determining the additional temporal indication of the particular feature within the audio data stream includes determining a stream position of the particular frequency segment within the audio data stream. In some versions of these versions, determining the vehicle audio delay based on the difference between the temporal indication of the particular feature within the captured audio data and the additional temporal indication of the particular feature within the audio data stream includes determining that the captured position of the particular frequency segment indicates that the frequency segment is a first-occurring frequency segment within the captured audio data and that the stream position of the particular frequency segment within the audio data stream indicates that the frequency segment is not a first-occurring frequency segment within the audio data stream, and determining the difference between the temporal indication of the particular feature within the captured audio data and the additional temporal indication of the particular feature within the audio data stream includes determining a position offset between the captured position and the stream position.

[0076] In some implementations, determining the vehicle audio delay based on comparing the captured audio data to the audio data stream includes, for each of a plurality of frequency segments in the sequence of frequency segments, determining a corresponding temporal offset between the frequency segment in the captured audio data and the frequency segment in the audio data stream. In some versions of these implementations, determining the vehicle audio delay based on comparing the captured audio data to the audio data stream includes determining the vehicle audio delay based on a maximum offset of the corresponding temporal offsets.

[0077] In some implementations, causing the computing device to append a corresponding delayed audio segment to the additional data stream prior to transmitting the additional data stream to the vehicle computing device over the wireless communication channel includes causing the computing device to append the corresponding delayed audio segment to a beginning of the additional data stream.

[0078] In some implementations, causing the computing device to append a corresponding delayed audio segment to the additional data stream prior to transmitting the additional data stream to the vehicle computing device over the wireless communication channel includes causing the computing device to append the corresponding delayed audio segment to an end of the additional data stream.

[0079] In some implementations, the wireless communication channel is a Bluetooth channel.

[0080] In some implementations, the computing device includes an automated assistant client. In some versions of these implementations, the additional audio data stream is transmitted to the vehicle computing device in response to receiving verbal input by the automated assistant client via one or more microphones, the additional audio data stream being an automated assistant response generated in response to the verbal input. In some versions of these implementations, the at least one microphone capturing the captured audio data includes at least one computing device microphone of the computing device. In some versions of these implementations, the at least one microphone capturing the captured audio data includes at least one interface microphone of a vehicle interface device communicating with the computing device via a second wireless communication channel, and receiving the captured audio data includes receiving the captured audio data from the vehicle interface device via the second communication channel.

[0081] In some implementations, the vehicle interface device is communicatively coupled to the vehicle computing device via an additional wireless communication channel.

[0082] In some implementations, the vehicle interface device is communicatively coupled to the vehicle computing device via a wired communication channel.

[0083] In some implementations, the method further includes causing the vehicle interface device to adapt the local noise cancellation filter based on the vehicle audio delay.

[0084] In some implementations, a method implemented by one or more processors includes causing a computing device to transmit an audio data stream to a vehicle computing device of a vehicle over a communication channel, wherein transmitting the audio data stream causes the vehicle computing device to render an audible output through one or more vehicle speakers driven by the vehicle computing device when the vehicle computing device is in a communication channel mode, the audible output being generated by the vehicle computing device based at least in part on the audio data stream. The method further includes determining whether the audible output is captured by at least one microphone in the vehicle. In response to determining that the audible output is captured by the at least one microphone in the vehicle, the method further includes causing the computing device to transmit an additional audio data stream to the vehicle computing device over the communication channel to render the additional audio data stream by the vehicle computing device through the one or more vehicle speakers. In response to determining that the audible output is not captured by the at least one microphone in the vehicle, the method further includes rendering the additional audio data stream on one or more alternative speakers instead, the one or more alternative speakers being in the vehicle but not the one or more vehicle speakers driven by the vehicle computing device.

[0085] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0086] In some implementations, the one or more alternative speakers are of the computing device. In some versions of these implementations, the method further includes detecting a call to the automated assistant client of the computing device, where the call causes the automated assistant client to transition from a first state to a second state, and causing the computing device to transmit the audio data stream to the vehicle computing device of the vehicle in response to the detection of the call.

[0087] In some implementations, detecting the call includes detecting an occurrence of the call phrase in audio data captured via at least one microphone of the computing device.

[0088] In some implementations, detecting the call includes detecting the call based on receiving an indication of the call from the vehicle interface device via the additional communication channel, wherein the vehicle interface device sends the indication of the call in response to user interaction with a hardware interface element or in response to detecting the occurrence of a call phrase in audio data captured via at least one microphone of the vehicle interface device.

[0089] In some implementations, causing the computing device to transmit the audio data stream to the vehicle computing device further includes, in response to a user interface input to the automated assistant client of the computing device, transmitting a request to the remote server device including the user interface input and / or additional data based on the user interface input. In some versions of these implementations, causing the computing device to transmit the audio data stream to the vehicle computing device further includes, in response to transmitting the request, receiving the additional audio data stream from the remote server device, and transmitting the audio data stream to the vehicle computing device occurs prior to receiving the entire additional audio data stream from the remote server device.

[0090] In some implementations, the at least one microphone in the vehicle includes at least one microphone of the computing device.

[0091] In some implementations, the method further includes determining a temporal indication indicating a time at which the automated assistant client caused the computing device to transmit the audio data stream to the vehicle computing device of the vehicle via the communication channel. In some versions of these implementations, the method further includes determining a current temporal indication indicating the current time. In some versions of these implementations, the method further includes determining a difference between the current temporal indication and the temporal indication. In response to determining that the difference between the current temporal indication and the temporal indication exceeds a threshold, some versions of these implementations further include causing the automated assistant client of the computing device to transmit a second audio data stream to the vehicle computing device of the vehicle via the communication channel, wherein transmitting the second audio data stream causes the vehicle computing device to render an additional audible output via one or more speakers of the vehicle computing device when the vehicle computing device is in the communication channel mode, the additional audible output being generated by the vehicle computing device based at least in part on the second audio data stream. In some versions of these implementations, the method further includes determining whether the additional audible output is captured by at least one microphone in the vehicle. In some versions of those implementations, in response to determining that the additional audible output is captured by at least one microphone in the vehicle, the method further includes causing the computing device to transmit a third audio data stream to the vehicle computing device over the communication channel. In response to determining that the additional audible output is not captured by the at least one microphone in the vehicle, in some versions of those implementations, the method further includes causing the third audible output to be rendered at one or more alternative speakers.

[0092] In some implementations, a method implemented by one or more processors includes causing a vehicle computing device of a vehicle to transmit an audio data stream via a wireless communication channel, the transmitting of the audio data stream causing the vehicle computing device to render an audible output through one or more vehicle speakers of the vehicle, the audible output being generated by the vehicle computing device based at least in part on the audio data stream. In some of these implementations, the method further includes receiving captured audio data captured by at least one microphone of the computing device in the vehicle, the captured audio data capturing the audible output rendered by the at least one vehicle speaker. In some of these implementations, the method further includes determining a vehicle audio delay based on comparing the captured audio data with the audio data stream. In some versions of these implementations, in response to determining the vehicle audio delay, the method further includes causing the computing device to adapt local noise cancellation based on the vehicle audio delay.

[0093] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0094] In some implementations, the local noise cancellation reduces a known source audio data stream transmitted over a wireless communication channel from the subsequently captured audio data for rendering by the vehicle computing device through one or more vehicle speakers, and adapting the local noise cancellation includes adapting an expected time to detect the known source audio data stream based on a vehicle audio delay.

[0095] In some implementations, the computing device is a vehicle interface device powered by a cigarette lighter receptacle of the vehicle. In some versions of those implementations, determining the vehicle audio delay is by the vehicle interface device. In some versions of those implementations, determining the vehicle audio delay is by a smartphone in communication with the vehicle interface device over a communications channel, and adapting the local noise cancellation based on the vehicle audio delay includes transmitting the vehicle audio delay and / or additional data determined based on the vehicle audio delay to the vehicle interface device.

[0096] In some implementations, a method implemented by one or more processors is provided, the method comprising: causing a computing device to transmit an audio data stream to an additional computing device over a wireless communication channel, the transmitting step including causing the additional computing device to render an audible output through one or more additional speakers driven by the additional computing device, the audible output being generated by the additional computing device based at least in part on the audio data stream. The method further includes receiving captured audio data captured by at least one microphone, the captured audio data capturing the audible output rendered by the at least one additional speaker. The method further includes determining an audio delay based on comparing the captured audio data with the audio data stream. In response to determining the audio delay, the method further includes causing the computing device to add a corresponding delayed audio segment to the additional audio data stream prior to transmitting the additional audio data stream to the additional computing device over the wireless communication channel, the duration of the delayed audio segment being determined using the audio delay, and / or adapting local noise cancellation based on the audio delay.

[0097] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0098] The additional computing device may be a vehicle computing device, and the one or more speakers may be one or more vehicle speakers.

[0099] The additional computing device may be a Bluetooth-enabled device that incorporates one or more additional speakers or is directly coupled to one or more additional speakers via an audio cable.

[0100] In some implementations, a method implemented by one or more processors is provided, the method including causing a computing device to transmit an audio data stream to an additional computing device via a communication channel, wherein transmitting the audio data stream, when the additional computing device is in a communication channel mode, causes the vehicle computing device to render an audible output via one or more additional speakers driven by the additional computing device, the audible output being generated by the additional computing device based at least in part on the audio data stream. The method further includes determining whether the audible output is captured by at least one microphone. In response to determining that the audible output is captured by the at least one microphone, the method further includes determining an audio delay based on comparing the captured audio data with the audio data stream. In response to determining that the audible output is captured by the at least one microphone, the method further includes causing the computing device to transmit the additional audio data stream via the communication channel to the additional computing device for rendering the additional audio data stream by the vehicle computing device via the one or more additional speakers. In response to determining that the audible output is not captured by the at least one microphone, the method further includes causing the additional audio data stream to be rendered instead at one or more alternative speakers, where the one or more alternative speakers are not the one or more vehicle speakers driven by the additional computing device.

[0101] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0102] The additional computing device may be a vehicle computing device, and the one or more speakers may be one or more vehicle speakers.

[0103] The additional computing device may be a Bluetooth-enabled device that incorporates one or more additional speakers or is directly coupled to one or more additional speakers via an audio cable.

[0104] Additionally, some implementations include one or more processors (e.g., central processing units (CPUs), graphical processing units (GPUs), and / or tensor processing units (TPUs) of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause any of the methods described herein to be performed. Some implementations also include one or more non-transitory computer-readable storage media having stored thereon computer instructions executable by the one or more processors to perform any of the methods described herein. [Explanation of symbols]

[0105] 102 Vehicle computing device, computing device 104 Wireless communication channels, Bluetooth channels 106 Computing Devices, Mobile Smartphone Computing Devices, Client Devices 202 Vehicle Computing Device 204 wireless communication network, first Bluetooth channel 206 Computing Devices 208 Wireless communication network, second Bluetooth channel 210 Vehicle Interface Device 302 Vehicle Interface Device 304 Computing Devices 306 Wireless Communication Channels 308 Vehicle Interface Device 310 Communication Channels 402 Audio Data Stream, Audio Data 404 frequency segment "1", the first frequency segment 406 Frequency segment "2" 408 Frequency segment "3" 410 frequency segment "4" 412 Frequency segment "5" 414 Audio Data 416 Sequence frequency segment "1", frequency segment "1" 418 Frequency Segment "2" 420 frequency segment "3" 422 frequency segment "4" 424 Frequency segment "5" 426 Audio Data Stream 428 Frequency Segment "2" 430 Frequency segment "3" 432 frequency segment "4" 434 Frequency segment "5" 436 Audio Data, Audio Data Stream 438 Frequency Segment "3" 440 frequency segment "4" 442 Frequency segment "5" 444 Audio Data, Audio Data Stream 446 Frequency segment "2" 448 Frequency Segment "3" 500 processes 600 processes 700 Automated Assistants 702 Client Computing Device, Client Device 704 Automated Assistant Client, Automated Assistant 706 Delay Engine 708 Local Engine 712 Cloud-Based Automated Assistant Components 714 Cloud-based TTS module, TTS module 716 Cloud-based STT module, STT module 718 Natural Language Processor 720 Audio Data, Audio Data Database 810 Computing Devices 812 Bus Subsystem 814 processor 816 Network Interface Subsystem 820 User Interface Output Device 822 User Interface Input Devices 824 Memory Subsystem 825 Memory Subsystem, Memory 826 File Storage Subsystem 830 Main Random Access Memory ("RAM") 832 Read Only Memory ("ROM")

Claims

1. A method performed by a first computing device located in a vehicle, comprising: transmitting the audio data stream to a vehicle computing device of the vehicle over a wireless communication channel; causing the vehicle computing device to render an audible output through one or more vehicle speakers of the vehicle; the audible output is generated by the vehicle computing device based at least in part on the audio data stream; capturing audio data corresponding to the audible output with at least one microphone of the first computing device; determining a vehicle audio delay based on comparing the captured audio data with the audio data stream; In response to determining the vehicle audio delay, adapting local noise cancellation based on the vehicle audio delay; the local noise cancellation reduces a known source audio data stream transmitted over the wireless communication channel from subsequently captured audio data for rendering by the vehicle computing device through the one or more vehicle speakers; adapting the local noise cancellation includes adapting an expected time to detect the known source audio data stream based on the vehicle audio delay; A method comprising:

2. The method of claim 1 , wherein the first computing device is a vehicle interface device powered by a cigarette lighter receptacle of the vehicle or a USB port of the vehicle.

3. The method of claim 2 , wherein the determining the vehicle audio delay is by the vehicle interface device.

Citation Information

Patent Citations

  • Hands-free speech communication method and equipment

    JP2003289373A

  • Echo canceler, video phone terminal and echo cancellation method

    JP2007214976A

  • Handsfree calling device

    JP2009033216A

  • Communication apparatus and communication method

    JP2012142910A

  • Hands-free call system

    JP2015076696A