Speaker diarization using speaker embedding and a trained generative model

By employing speaker embeddings and a generative model to process audio data, the system effectively addresses the challenges of accurately identifying speakers in multi-speaker environments, resulting in improved diarization accuracy and reduced errors in speech processing.

JP7690651B2Active Publication Date: 2025-06-10GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024097959
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-06-10
Estimated Expiration
2038-09-25

Smart Images

  • Figure 0007690651000005
    Figure 0007690651000005
  • Figure 0007690651000006
    Figure 0007690651000006
  • Figure 0007690651000007
    Figure 0007690651000007
Patent Text Reader

Abstract

To provide a speaker diarization technique for allowing processing of audio data to generate one or a plurality of refined versions of the audio data.SOLUTION: Each refined version of audio data separates one or a plurality of utterances of a single human speaker. Various implementations generate speaker embedding for the single human speaker, process the audio data using a trained generation model, and use speaker embedding when activation of a hidden layer of the trained generation model is determined during processing to generate the refined version of the audio data which separates the utterance of the single human speaker. Output is generated on the trained generation model on the basis of the processing and output is the refined version of the audio data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Speaker diarization is the process of splitting an input audio stream into homogeneous segments according to speaker identification. It answers the question "who spoke when" in a multi-speaker environment. For example, speaker diarization can identify that the first segment of an input audio stream is due to a first human speaker (without particularly identifying who the first human speaker is), identify that the second segment of the input audio stream is due to a different second human speaker (without particularly identifying who the first human speaker is), and identify that the third segment of the input audio stream is due to the first human speaker, etc. Speaker diarization has a variety of applications including multimedia information retrieval, speaker turn analysis, and audio processing.

[0002] A typical speaker diarization system usually consists of four steps: (1) audio segmentation where the input audio is segmented into short sections assumed to have a single speaker and non-audio sections are removed, (2) audio embedding extraction where specific features are extracted from the segmented sections, (3) clustering where the number of speakers is determined and the extracted audio embeddings are clustered into these speakers, and optionally (4) re-segmentation where the results of clustering are further refined to generate the final diarization result.

[0003] In such a typical speaker diarization system, diarization fails to accurately recognize the occurrence of multiple speakers speaking in a given segment. Rather, such a typical system fails to attribute a given segment to only one speaker or fails to attribute a given segment to any of the speakers. This leads to inaccurate diarization and may have an adverse impact on other applications that may depend on the diarization results.

[0004] Furthermore, in such a typical speaker diarization system, errors can be introduced at each step and can propagate to further steps, thereby causing incorrect diarization results and thereby adversely affecting further applications that may depend on the incorrect diarization results. For example, errors can be introduced in audio segmentation as a result of low resolution of long segments and / or as a result of short segments having insufficient audio to generate accurate audio embeddings. As another example, audio embeddings can be generated locally without using any global information, which can additionally or alternatively introduce errors. As yet another example, clustering of audio embeddings involves unsupervised learning which is said to be of low accuracy and can thus additionally or alternatively introduce errors. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0005] This specification describes speaker diarization techniques that enable the processing of a sequence of audio data to generate one or more refined versions of the audio data, where each of the refined versions of the audio data separates one or more utterances of a single respective human speaker, thereby enabling determination of which portions of the sequence of audio data correspond to each respective human speaker. For example, consider a sequence of audio data that includes a first utterance from a first human speaker, a second utterance from a second human speaker, and various occurrences of background noise. The implementations disclosed herein can be utilized to generate a first refined audio data that includes only the first utterance from the first human speaker and excludes the second utterance and the background noise. Additionally, a second refined audio data can be generated that includes only the second utterance from the second human speaker and excludes the first utterance and the background noise. Further, in those implementations, the first and second utterances can be separated even if one or more of the first and second utterances overlap in the sequence of audio data.

[0006] Various implementations generate a speaker embedding for a single human speaker, process the audio data using a trained generative model, and use the speaker embedding when determining the activation of hidden layers of the trained generative model during processing to generate a refined version of the audio data that separates the utterances of the single human speaker. The output is sequentially generated on the trained generative model based on the processing, and the output is a refined version of the audio data.

[0007] When generating speaker embeddings for a single human speaker, one or more instances of speaker audio data corresponding to the human speaker can be processed using a trained speaker embedding model to generate one or more respective instances of the output. The speaker embeddings can then be generated based on one or more respective instances of the output. The trained speaker embedding model can be a machine learning model such as a recurrent neural network (RNN) model that accepts as input a sequence of features of each audio data frame of any length and can be utilized to generate respective embeddings as output based on the input. Each of the features of the sequence of audio data frames processed using the trained speaker embedding model to generate each embedding can be based on each respective portion of each instance of the audio data, such as 25 milliseconds or other duration segments. The features of the audio data frames can be, for example, Mel-frequency cepstral coefficients (MFCCs) and / or other features of the audio data frames. When the trained speaker embedding model is an RNN model, the RNN model includes one or more memory layers each including one or more memory units to which the input can be sequentially applied, and in each iteration of the applied input, the memory units can be utilized to compute a new hidden state based on the input of that iteration and based on the current hidden state (which can be based on the input of the previous iteration). In some implementations, the memory units can be long short-term (LSTM) units. In some implementations, additional or alternative memory units such as gated recurrent units (“GRUs”) can be utilized.

[0008] As an example of generating a speaker embedding for a given speaker, the speaker embedding can be generated during a registration process in which the given speaker speaks multiple utterances. Each utterance can be of the same phase (text-dependent) or of a different phase (text-independent). The features of the audio data corresponding to each instance of the given speaker speaking each utterance can be processed on a speaker embedding model to generate respective outputs that are vectors of respective values. For example, the first audio data regarding the first utterance can be processed to generate a first vector of values, the second audio data regarding the second utterance can be processed to generate a second vector of values, and so on. The speaker embedding can then be generated based on the vectors of values. For example, the speaker embedding can itself be a vector of values such as the centroid or other function of the respective vectors of values.

[0009] In an implementation that utilizes a speaker embedding for a given speaker that was generated beforehand (e.g., during a registration process), the techniques for separating utterances of the given speaker described herein can utilize the speaker embedding that was generated beforehand when generating a refined version of the audio data, and the audio data is received from a user via a client device and / or digital system (e.g., an automated assistant) associated with the registration process. For example, if the audio data is received via a client device of a given user and / or after verification of the given user (e.g., using voice fingerprinting and / or other biometrics from previous utterances), the speaker embedding for the given user can be utilized to generate a refined version of the audio data in real time. Such a refined version can be utilized for various purposes such as conversion of the audio of the refined audio data to text, verification that segments of the audio data are from the user, and / or other purposes described herein.

[0010] In some additional or alternative implementations, the speaker embeddings utilized in generating a refined version of the audio data can be based on one or more instances of the audio data itself (to be refined). For example, a voice activity detector (VAD) can be utilized to determine a first instance of voice activity in the audio data, and a portion of the first instance can be utilized in generating a first speaker embedding for a first human speaker. For example, the first speaker embedding can be generated based on processing the first X (e.g., 0.5, 1.0, 1.5, 2.0) seconds of features of the first instance of voice activity (the first instance of voice activity can be considered to be from a single speaker) using a speaker embedding model. For example, a vector of values generated as an output based on the processing can be utilized as the first speaker embedding. The first speaker embedding can then be utilized to generate a first refined version of the audio data that separates the utterances of the first speaker, as described herein. In some of those implementations, the first refined version of the audio data can be utilized to determine those segments of the audio data corresponding to the utterances of the first speaker, and the VAD can be utilized to determine additional instances (if any) of voice activity in the audio data occurring outside of those segments. If additional instances are determined, a second speaker embedding can be generated for a second human speaker based on processing a portion of the additional instances using a speaker embedding model. The second speaker embedding can then be utilized to generate a second refined version of the audio data that separates the utterances of the second speaker, as described herein. This process can continue, for example, until no further utterances due to additional human speakers are identified in the audio data. Thus, in these implementations, the speaker embeddings utilized in generating a refined version of the audio data can be generated from the audio data itself.

[0011] Regardless of the technique used to generate the speaker embedding, the implementations disclosed herein use a trained generative model to process the audio data and the speaker embedding to generate a refined version of the audio data that separates the speaker's utterance (if any) corresponding to the speaker embedding. For example, the audio data can be processed sequentially using the trained generative model, and the speaker embedding can be used to determine the activation of the layers of the trained generative model during sequential processing. The trained generative model can be an inter-sequence model, and the refined version of the audio data can be generated sequentially as a direct output from the trained generative model. The layers of the trained generative model are hidden layers and can include, for example, a stack of dilated causal convolutional layers. The stack of dilated causal convolutional layers enables the receptive field of the convolutional layer to grow exponentially with depth, which can help to model long-term dependencies in the audio signal. In various implementations, the audio data processed using the trained generative model can be at the waveform level, and the refined audio data generated using the trained generative model can also be at the waveform level. In some implementations, the trained generative model has a WaveNet model architecture and is trained according to the techniques described herein.

[0012] In various implementations, the trained generative model is trained to model the conditional distribution p(x|h), where x represents the refined version of the audio data and h represents the speaker embedding. More formally, p(x|h) can be represented as

[0013]

Equation

[0014] and can be represented as such, where x 1 ...x t-1 represents a sequence of T refined audio sample predictions (which can be conditioned on the source audio data), and x trepresents the following refined audio sample prediction (i.e., the following audio sample prediction regarding the refined version of the audio data). As described above, h represents the speaker embedding and can be a real-valued vector with a fixed-size dimension.

[0015] Furthermore, in various implementations, the output filter in each of one or more layers of the trained generative model (e.g., in each causal convolutional layer) can be represented by the following equation

[0016] [Number]

[0017] where W represents a filter and V represents another filter. Thus, the combination of the speaker embedding h and the audio sample x is performed by transforming h with filter V, transforming the audio sample x with filter W, and summing the results of both operations. The result (z) of both operations becomes the input x to the next layer. The weights of the generative model learned during the training of the generative model can include the weights of filters W and Y. Additional descriptions of the generative model and its training are provided herein.

[0018] Utilizing the trained generative model to process given audio data considering a given speaker embedding of a given speaker results in refined audio data that is the same as the given audio data if the given audio data contains only utterances of the given speaker. Further, it results in refined audio data that is null / zero if the given audio data lacks any utterances from the given speaker. Further, it results in refined audio data with the utterances from the given speaker separated and the additional sounds excluded if the given audio data contains utterances from the given speaker and additional sounds (e.g., overlapping and / or non-overlapping utterances of other human speakers).

[0019] The refined version of the audio data can be utilized by various components and for various purposes. As an example, speech-to-text processing can be performed on the refined version of the audio data that separates utterances from a single human speaker. Performing speech-to-text processing on the refined version of the audio data can improve the accuracy of the speech-to-text processing, for example, because the refined version has no background noise, other users' utterances (e.g., overlapping utterances), etc., as compared to performing processing on the audio data (or an alternative pre-processed version of the audio data). Further, performing speech-to-text processing on the refined version of the audio data guarantees that the resulting text belongs to a single speaker. The improved accuracy and / or guaranteeing that the resulting text belongs to a single speaker can directly result in further technical advantages. For example, the improved accuracy of the text can increase the accuracy of one or more downstream components that depend on the resulting text (e.g., a natural language processor, a module that generates a response based on the intent and parameters determined based on natural language processing of the text). Also, for example, when implemented in combination with an automated assistant and / or other interactive systems, the improved accuracy of the text can reduce the likelihood that the interactive system will be unable to convert an utterance to text and / or reduce the likelihood that the interactive system will incorrectly convert an utterance to text, thereby leading to an incorrect response to the utterance by the interactive system. This can reduce the amount of dialogue turns that would otherwise be required for the user to provide the utterance and / or other explanations to the interactive system again.

[0020] Additionally or alternatively, a refined version of the audio data can be utilized to assign segments of the audio data to corresponding human speakers. The assignment of segments to human speakers can be semantically meaningful (e.g., identifying speaker attributes), or can simply indicate to which of one or more semantically meaningless speaker labels a segment belongs. The implementations disclosed herein may result in generating more robust and / or more accurate assignments compared to other speaker diarization techniques. For example, various errors introduced by other techniques, such as those mentioned in the background above, can be reduced through the utilization of the implementations disclosed herein. Additionally or alternatively, the use of the implementations disclosed herein can enable the determination of segments that include each utterance from each of two or more human speakers, which cannot be achieved by various other speaker diarization techniques. Further, additionally or alternatively, the implementations disclosed herein can enable speaker diarization to be performed in a more computationally efficient manner compared to previous speaker diarization techniques. For example, the computationally intensive clustering of previous techniques can be excluded in various implementations disclosed herein.

[0021] In various implementations, the techniques described herein are utilized to generate speaker diarization results, to perform automatic speech recognition (ASR) (e.g., speech-to-text processing), and / or for other processing of audio data submitted as part of a speech processing request (e.g., via an application programming interface (API)). In some of those implementations, the results of the processing of the audio data are generated in response to the speech processing request and sent back to the computing device that sent the speech processing request, or to an associated computing device.

[0022] In various implementations, the techniques described herein are utilized to generate speaker diarization results for audio data captured by a microphone of a client device that includes an automated assistant interface for an automated assistant. For example, the audio data can be a stream of audio data that captures utterances from one or more speakers, and the techniques described herein can be utilized to generate speaker diarization results, to perform automatic speech recognition (ASR), and / or for other processing of the stream of audio data.

[0023] The above description is provided as an overview of various implementations disclosed herein. Those various implementations, as well as additional implementations, are described in more detail herein.

[0024] In some implementations, a method is provided that includes generating a speaker embedding for a human speaker. The step of generating a speaker embedding for a human speaker can optionally include processing one or more instances of speaker audio data corresponding to the human speaker using a trained speaker embedding model, and generating a speaker embedding based on one or more instances of output each generated based on processing each of the one or more instances of speaker audio data using the trained speaker embedding model. The method further includes receiving audio data that captures one or more utterances of the human speaker and also captures one or more additional sounds not from the human speaker, generating a refined version of the audio data, where the refined version of the audio data separates the one or more utterances of the human speaker from the one or more additional sounds not from the human speaker, and performing further processing on the refined version of the audio data.

[0025] These and other implementations can include one or more of the following features.

[0026] In some implementations, the step of generating a refined version of the audio data includes sequentially processing the audio data using a trained generation model and, during the sequential processing, using a speaker embedding when determining the activation of the layers of the trained generation model, and sequentially generating a refined version of the audio data based on the sequential processing and as a direct output from the trained generation model.

[0027] In some implementations, the step of performing further processing includes performing a speech-to-text process on the refined version of the audio data to generate predicted text for one or more utterances of a human speaker and / or assigning a single given speaker label to one or more temporal portions of the audio data based on one or more temporal portions of the audio data corresponding to at least a threshold level of audio in the refined version of the audio data.

[0028] In some implementations, a method is provided that includes the step of invoking an automated assistant client on a client device, the step of invoking the automated assistant client responding to detecting one or more call queues in a received user interface input. The method further includes, in response to invoking the automated assistant client, executing specific processing of a first verbal input received via one or more microphones of the client device, generating a response action based on the specific processing of the first verbal input, causing execution of the response action, and determining that an ongoing listening mode is activated for the automated assistant client on the client device. The method further includes, in response to the ongoing listening mode being activated, automatically monitoring for additional verbal input after causing execution of at least a portion of the response action, receiving audio data while automatically monitoring, and determining whether the audio data includes any additional verbal input from the same human speaker who provided the first verbal input. The step of determining whether the audio data includes any additional verbal input from the same human speaker includes identifying a speaker embedding for the human speaker who provided the first verbal input, and generating a refined version of the audio data that separates any of the audio data from the human speaker, the step of generating the refined version of the audio data including processing the audio data using a trained generation model and, during processing, using the speaker embedding when determining activation of layers of the trained generation model, and generating a refined version of the audio data based on the processing, and determining whether the audio data includes any additional verbal input from the same human speaker based on whether any portion of the refined version of the audio data corresponds to at least a threshold level of audio.The method further includes suppressing one or both of the execution of at least some of the specific processing for the audio data and the generation of any additional response actions adjusted according to the audio data in response to determining that the audio data does not include any additional verbal input from the same human speaker.

[0029] In some implementations, receiving a stream of audio data captured via one or more microphones of a client device, obtaining a previously generated speaker embedding for a human user of the client device from local storage of the client device, and while receiving the stream of audio data, generating a refined version of the audio data, the refined version of the audio data separating one or more utterances of the human user from any additional sounds not from the human speaker, and the step of generating the refined version of the audio data includes processing the audio data using a trained generation model and using the speaker embedding (e.g., when determining the activation of layers of the trained generation model during processing), and generating a refined version of the audio data based on the processing and as a direct output from the trained generation model. The method further includes performing a local speech texturing process on the refined version of the audio data and / or transmitting the refined version of the audio data to a remote system to cause the remote system to perform a remote speech texturing process on the refined version of the audio data.

[0030] In some implementations, a method of training a machine learning model to generate a refined version of audio data that separates arbitrary utterances of a target human speaker is provided. The method includes identifying an instance of audio data that includes verbal input from only a first human speaker, generating a speaker embedding for the first human speaker, identifying an additional instance of audio data that includes no verbal input from the first human speaker and includes verbal input from at least one additional human speaker, generating a mixed instance of audio data that combines the instance of audio data and the additional instance of audio data, processing the mixed instance of audio data using the machine learning model and using the speaker embedding when determining the activation of a layer of the machine learning model during processing, generating a predicted refined version of the audio data based on the processing and as a direct output from the machine learning model, generating a loss based on comparing the predicted refined version of the audio data with the instance of audio data that includes verbal input from only the first human speaker, and updating one or more weights of the machine learning model based on the loss.

[0031] Additionally, some implementations include one or more processors of one or more computing devices, the one or more processors being operable to execute instructions stored in a related memory, the instructions being configured to cause execution of any of the methods described herein. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to execute any of the methods described herein.

[0032] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein.

Brief Description of the Drawings

[0033]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 8

Figure 9

Figure 10

Figure 11

[0034] First, turning to FIG. 1, an example of training a generative model 156 is shown. The generative model 156 can be a sequence-to-sequence model having a hidden layer that includes a stack of dilated causal convolutional layers. The stack of dilated causal convolutional layers allows the receptive field of the convolutional layer to grow exponentially with depth, which may help to model long-term dependencies in audio signals. In some implementations, the trained generative model has a WaveNet model architecture.

[0035] The generation model 156 is trained to be used to generate a refined version of the audio data that separates the utterance (if any) of the target human speaker based on the processing of the audio data and the speaker embedding of the target person. As described herein, the generation model 156 accepts, as input for each iteration, each sequence of the audio data for that iteration (the number of time steps of the audio data within the sequence depends on the size of the receptive field of the generation model), and can be trained to generate, as output for each iteration, the corresponding time steps of a sequence of the refined version of the audio data. The corresponding time steps of the sequence of the refined version of the audio data can be based on the corresponding time steps of the audio data and the preceding time steps of the audio data captured in the receptive field, but not on the future time steps of the audio data. As further described herein, the output is also generated using the speaker embedding of the target speaker. For example, the activation of the layers of the generation model 156 during the processing of a sequence of audio data can be based on the speaker embedding. Thus, once trained, the generation model 156 can be used to generate, as a direct output of the generation model 156, a refined version of the audio data, and is used to generate a refined version based on the processing of the audio data and the speaker embedding of the target speaker. The generation model 156 can accept the audio data in its raw waveform format and can likewise generate the refined audio data in the raw waveform format.

[0036] The trained generation model 156 can be trained to model the conditional distribution p(x|h), where x represents the refined version of the audio data and h represents the speaker embedding. More formally, p(x|h,s) is

[0037]

Number

[0038] can be represented as, where x 1 ...x t-1 represents a sequence of T refined audio sample predictions (which may be conditioned in the source audio data), and x t represents the next refined audio sample prediction (i.e., the next audio sample prediction regarding the refined version of the audio data), h represents the speaker embedding, and can be a real-valued vector having a fixed-size dimension. Further, in various implementations, the output filters in each of one or more layers of the trained generative model are

[0039]

Number

[0040] can be represented by, where W represents a filter and V represents another filter.

[0041] Turning to FIG. 10, an example of a process that can be performed by the trained generative model is shown. In FIG. 10, the current time step 1070 of the input audio data 1070 (source audio data) t is processed on the hidden layers 157A, 157B, and 157C of the trained generative model together with the preceding 15 time steps of the input audio data 1070 (1070 t is the only preceding one numbered in FIG. 10 for simplicity) to directly generate the current time step 1073 of the refined audio data t-15 . Due to the stacked causal convolutional nature of layers 157A - 157C, the processing of the current time step 1073 of the refined audio t is related to the current time step 1070 of the input audio 1070 tIt is also affected by the processing of the previous 15 time steps of the input audio 1070. The influence of the previous time steps is a result of the increased receptive field provided by the nature of the stacked causal convolutions of layers 157A - 157C, and the other solid arrows represent the current time step 1073 of the refined audio data t presents a visualization of the influence of the previous 15 time steps of the input audio data 1070 on generating t . The previous 15 time steps of the refined audio data 1073 generated in the previous iteration of the process (1073 t-15 which is the only previous one numbered in FIG. 10 for simplicity) are also shown in FIG. 10. Such previous time steps of the refined audio data 1073 can similarly be affected by the previous processing of the input audio 1070 that occurred before those time steps were generated (the dashed lines represent such previous influences). As described herein, the processing performed using each of the hidden layers 157A - 157C can be affected by the speaker embedding for which the refined audio data 1073 is being generated. For example, one or more of the filters modeled by such hidden layers 157A - 157C can be affected by the speaker embedding. Specific examples are shown in FIG. 10, but variations such as additional layers, variations with larger receptive fields, etc. are contemplated.

[0042] Returning to FIG. 1, a training instance engine 130 that generates a plurality of training instances 170A - 170N stored in a training instance database 170 for use in training the generative model 156 is also shown. Training instance 170A is shown in detail in FIG. 1. Training instance 170A includes a mixed instance 171A of audio data, an embedding 172A of a given speaker, and ground truth audio data 173A (also referred to herein as "reference audio data") which is an instance of audio data having only utterances from the given speaker corresponding to the embedding 172A.

[0043] The training instance engine 130 can generate training instances 170A to 170N based on instances of audio data from an instance of the audio data database 160 and through interaction with the speaker embedding engine 125. For example, the training instance engine 130 can obtain the ground truth audio data 173A from an instance of the audio data database 160 and use it as the ground truth audio data for the training instance 170A.

[0044] Furthermore, the training instance engine 130 can provide the ground truth audio data 173A to the speaker embedding engine 125 in order to receive an embedding 172A of a given speaker from the speaker embedding engine 125. The speaker embedding engine 125 can use the speaker embedding model 152 to process one or more segments of the ground truth audio data 173A in order to generate an embedding 172A of a given speaker. For example, the speaker embedding engine 125 can determine one or more segments of the ground truth audio data 173A that include voice activity and utilize VAD to determine the embedding 172A of a given speaker based on processing one or more of those segments using the speaker embedding model 152. For example, all of the segments can be processed using the speaker embedding model 152, and the final output of the result generated based on the processing can be used as the embedding 172A of a given speaker. Also, for example, the first segment can be processed using the speaker embedding model 152 to generate a first output, the second segment can be processed using the speaker embedding model 152 to generate a second output, etc., and the center of gravity of the outputs is utilized as the embedding 172A of a given speaker.

[0045] The training instance engine 130 generates a mixed instance 171A of audio data by combining the ground truth audio data 173A with additional instances of audio data from an instance of the audio data database 160. For example, the additional instances of audio data can include one or more other human speakers and / or background noise.

[0046] Referring to FIG. 2, the ground truth audio data 173A is schematically shown. The arrows indicate time, and the three vertical shaded regions within the ground truth audio data 173A represent segments of audio data where "Speaker A" is providing each utterance. In particular, the ground truth audio data 173A contains no (or minimal) additional sound. As described above, the ground truth audio data can be utilized in the training instance 170A. Further, the embedding 172A of a given speaker in the training instance 170A can be generated based on the processing of one or more segments of the ground truth audio data 173A by the speaker embedding engine 125.

[0047] Additional audio data A 164A that can be obtained from an instance of the audio data database 160 (FIG. 1) is also schematically shown in FIG. 2. The arrows indicate time, and the two diagonal shaded regions within the additional audio data A 164A represent segments of audio data where "Speaker C" is providing each utterance. In particular, "Speaker C" is different from "Speaker A" who provides utterances in the ground truth audio data 173A.

[0048] The mixed audio data A 171A generated by combining the ground truth audio data 163A and the additional audio data A 164A is also shown in FIG. 2. The mixed audio data A 171A includes a masked area (vertical masking) from the ground truth audio data 173A and a masked area (diagonal masking) from the additional audio data A 164A. Thus, in the mixed audio data A 171A, the utterances of both "Speaker A" and "Speaker C" are included, and a part of the utterance of "Speaker A" overlaps with a part of the utterance of "Speaker B". As described above, the mixed audio data A 171A can be used in the training instance 170A.

[0049] Furthermore, additional training instances of the training instances 170A to 170N can be generated, including the embedding 172A of a given speaker, the ground truth audio data 173A, and the mixed audio data B 171B. The mixed audio data B 171B is schematically shown in FIG. 2 and is generated by combining the ground truth audio data 163A and the additional audio data B 164B. The additional audio data B 164B includes the utterance (dotted masking) from "Speaker B" and includes background noise (hatched masking). The mixed audio data B 171B includes a masked area (vertical masking) from the ground truth audio data 173A and a masked area (dotted masking and hatched masking) from the additional audio data B 164B. Thus, in the mixed audio data B 171B, the utterances of both "Speaker A" and "Speaker B", as well as "background noise", are included. Furthermore, a part of the utterance of "Speaker A" overlaps with a part of the "background noise" and the utterance of "Speaker B".

[0050] Although only two training instances are described with respect to FIG. 2, it is understood that more training instances can be generated. For example, additional training instances that utilize ground truth audio data 173A can be generated (e.g., by generating additional mixed audio data based on additional audio data instances). Also, for example, additional training instances can be generated that each utilize alternative ground truth audio data (each including an alternative speaker), and each include an alternative speaker embedding (corresponding to the alternative speaker of the ground truth audio data).

[0051] When training the generation model 156 based on the training instance 170A, the refinement engine 120 utilizes the embedding 172A of a given speaker to determine the activation of the layers of the generation model 156 in order to sequentially generate time steps of the predicted audio data 175A, and sequentially applies the mixed instance 171A of the audio data as an input to the generation model 156. For example, the refinement engine 120 can apply the first time step of the mixed instance 171A of the audio data as an input to the generation model 156 in the first iteration, and generate the first time step of the predicted audio data based on processing the input on the model 156 (and based on the embedding 172A). Continuing with the example, the refinement engine 120 can then apply the first and second time steps of the mixed instance 171A of the audio data as an input to the generation model 156 in the second iteration, and generate the second time step of the predicted audio data based on processing the input on the model 156 (and based on the embedding 172A). This can continue until all time steps of the mixed instance 171A of the audio data are processed (later iterations include the current time step of the audio data and up to a maximum of N preceding time steps, where N depends on the receptive field of the model 156).

[0052] The loss module 132 generates a loss 174A as a function of the predicted audio data 175A and the ground truth audio data 173A. The loss 174A is provided to an update module 134 that updates the generation model 156 based on the loss. For example, the update module 134 can update the causal convolutional layer of the generation model 156 (modeling the filters in the above equation) based on the loss and using backpropagation (e.g., gradient descent).

[0053] Although FIG. 1 shows only a single training instance 170A in detail, it is understood that the training instance database 170 can include a large number of additional training instances. The additional training instances can include training instances of various lengths (e.g., based on various durations of audio data), training instances having various ground truth audio data and speaker embeddings, and training instances having various additional sounds within each mixed instance of the audio data. Further, it is understood that a large number of additional training instances are utilized to train the generation model 156.

[0054] Turning now to FIG. 3, an exemplary method 300 for generating training instances for training a generation model according to various implementations disclosed herein is shown. For convenience, the operations of the flowchart of FIG. 3 are described with reference to a system that performs the operations. This system can include various components of various computer systems, such as a training instance engine 130 and / or one or more GPUs, CPUs, and / or TPUs. Further, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.

[0055] In block 302, the system selects a ground truth instance of audio data that includes an oral input from a single human speaker.

[0056] In block 304, the system generates a speaker embedding for a single human speaker. For example, the speaker embedding can be generated by processing one or more segments of the ground truth instance of the audio data using a speaker embedding model.

[0057] In block 306, the system selects an additional instance of audio data without oral input from a single human speaker. For example, the additional instance of audio data can include oral input from other speakers and / or background noise (e.g., music, sirens, air conditioner noise, etc.).

[0058] In block 308, the system generates a mixed instance of audio data by combining the ground truth instance of the audio data and the additional instance of the audio data.

[0059] In block 310, the system generates and stores a training instance that includes the mixed instance of the audio data, the speaker embedding, and the ground truth instance of the audio data.

[0060] In block 312, the system determines whether to generate an additional training instance using the same ground truth instance of the audio data and the same speaker embedding but without using a different mixed instance of the audio data based on a different additional instance. If so, the system returns to block 306, selects a different additional instance, proceeds to block 308, generates a different mixed instance of the audio data by combining the same ground truth instance of the audio data and the different additional instance, and then proceeds to block 310 to generate and store the corresponding training instance.

[0061] In the iteration of block 312, if the system determines not to generate additional training instances using the same ground truth instance of the audio data and the same speaker embedding, the system proceeds to block 314 and determines whether to generate additional training instances using a different ground truth instance of the audio data. If so, the system utilizes different ground truth instances of the audio data by different human speakers, utilizes different speaker embeddings for different human speakers, and optionally utilizes different additional instances of the audio data to perform another iteration of blocks 302, 304, 306, 308, 310, and 312.

[0062] In the iteration of block 314, if the system determines not to generate another training instance using a different ground truth instance of the audio data, the system proceeds to block 316 and the generation of training instances ends.

[0063] Turning now to FIG. 4, an exemplary method 400 for training a generative model according to various implementations disclosed herein is shown. For convenience, the operations of the flowchart of FIG. 4 are described with reference to a system that performs the operations. This system may include various components of various computer systems, such as a refinement engine 120 and / or one or more GPUs, CPUs, and / or TPUs. Further, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be rearranged, omitted, or added.

[0064] In block 402, the system selects a training instance that includes a mixed instance of audio data, a speaker embedding, and ground truth audio data. For example, the system can select a training instance generated according to method 300 of FIG. 3.

[0065] In block 404, the system uses a machine learning model (e.g., generative model 156) and speaker embeddings to sequentially process mixed instances of audio data when determining the activation of layers of the machine learning model.

[0066] In block 406, the system sequentially generates a predicted refined version of the mixed instance of audio data as the direct output of the machine learning model based on the sequential processing in block 404.

[0067] In block 408, the system generates a loss based on comparing the predicted refined version of the mixed instance of audio data to the ground truth instance of the audio data.

[0068] In block 410, the system updates the weights of the machine learning model based on the generated loss.

[0069] In block 412, the system determines whether to perform further training of the machine learning model. If so, the system returns to block 402, selects additional training instances, then performs iterations of blocks 404, 406, 408, and 410 based on the additional training instances, and then performs additional iterations of block 412. In some implementations, the system can determine to perform further training if one or more additional unprocessed training instances exist and / or other criteria have not yet been met. Other criteria can include, for example, whether a threshold number of epochs have occurred and / or whether training of a threshold duration has occurred. Method 400 is described with respect to non-batch learning techniques, but batch learning can be additionally and / or alternatively utilized.

[0070] In the iteration of block 412, if the system determines not to perform further training, the system can proceed to block 416, where the system provides the machine learning model to be used, taking into account the trained machine learning model. For example, the system can provide a trained machine learning model for use in one or both of methods 600 (FIG. 6A) and 700 (FIG. 7A) described herein.

[0071] FIG. 5 shows an example of generating a refined version of audio data using the audio data, speaker embeddings, and a generative model. The generative model 156 can be the same as the generative model 156 of FIG. 1, but is trained (e.g., trained as described with respect to FIGS. 1 and / or 3).

[0072] In FIG. 5, the refinement engine 120 receives a sequence of audio data 570. The audio data 570 can be, for example, streaming audio data processed in an online manner (e.g., in real-time or near real-time), or non-streaming audio data previously recorded and provided to the refinement engine 120. The refinement engine 120 also receives the speaker embedding 126 from the speaker embedding engine 125. The speaker embedding 126 is an embedding for a given human speaker, and the speaker embedding engine 125 can generate the speaker embedding 126 based on processing one or more instances of audio data from the given speaker using the speaker embedding model 125. As described herein, in some implementations, the speaker embedding 126 was previously generated by the speaker embedding engine 125 based on previous instances of audio data from the given speaker. In some of those implementations, the speaker embedding 126 is associated with the account of the given speaker and / or the client device of the given speaker, and the speaker embedding 126 can be provided for use with the audio data 570 based on the audio data 570 coming from the client device and / or from a digital system approved by the account. As also described herein, in some implementations, the speaker embedding 126 is generated by the speaker embedding engine 125 based on the audio data 570 itself. For example, VAD can be performed on the audio data 570 to determine a first instance of voice activity within the audio data, and a portion of the first instance can be utilized by the speaker embedding engine 125 in generating the speaker embedding 126.

[0073] The refinement engine 120 sequentially applies the audio data 570 as input to the generation model 156 using the speaker embedding 126 when determining the activation of the layers of the generation model 156 to sequentially generate time steps of the refined audio data 573. The refined audio data 573 can be the same as the audio data 570 if the audio data 570 contains only utterances from the speaker corresponding to the speaker embedding 126, can be null / zero if the audio data 570 lacks any utterances from the speaker corresponding to the speaker embedding 126, or can exclude additional sounds while separating the utterances from the speaker corresponding to the speaker embedding 126 if the audio data 570 contains utterances from the speaker and additional sounds (e.g., overlapping utterances of other human speakers).

[0074] The refinement engine 120 then provides the refined audio data 573 to one or more additional components 135. FIG. 5 shows generating a single instance of the refined audio data 573 based on a single speaker embedding 126, but it is understood that in various implementations, multiple instances of the refined audio data can be generated, each instance being based on the audio data 570 and a unique speaker embedding for a unique human speaker.

[0075] As described above, the refinement engine 120 provides the refined audio data 573 to one or more additional components 135. In some implementations, the refinement engine 120 provides the refined audio data 573 in an online manner (e.g., a portion of the refined audio data 573 may be provided while the remainder is still being generated). In some implementations, the additional components 135 include a client device or other computing device (e.g., a server device), and the audio data 570 is received as part of an audio processing request submitted by the computing device (or a related computing device). In those implementations, the refined audio data 573 is generated in response to receiving the audio processing request and is transmitted to the computing device in response to receiving the audio processing request. Optionally, other (not shown) audio processing may additionally be performed in response to the audio processing request (e.g., speech-to-text processing, natural language understanding), and the results of such audio processing are transmitted additionally or alternatively in response to the request.

[0076] In some implementations, the additional components 135 include one or more components of an automated assistant, such as an automatic speech recognition (ASR) component (e.g., performing conversion from speech to text) and / or a natural language understanding component. For example, the audio data 570 can be streaming audio data based on the output from one or more microphones of a client device that includes an automated assistant interface for interacting with the automated assistant. The automated assistant can include (or be communicable with) the refinement engine 120, and transmitting the refined audio data 573 can include transmitting it to one or more other components of the automated assistant.

[0077] Turning now to FIGS. 6A, 6B, 7A, and 7B, additional explanations are provided for generating a refined version of the audio data and using such a refined version.

[0078] FIG. 6A shows an exemplary method 600 for generating a refined version of audio data using audio data, speaker embeddings, and a generation model according to various implementations disclosed herein. For convenience, the operations of certain aspects of the flowchart of FIG. 6A are described with reference to audio data 670 and refined audio data 675 schematically represented in FIG. 6B. Also for convenience, the operations of the flowchart of FIG. 6A are described with reference to a system that performs the operations. This system may include various components of various computer systems such as a speaker embedding engine 125, a refinement engine 120, and / or one or more GPUs, CPUs, and / or TPUs. In various implementations, one or more blocks of FIG. 6A may be performed by a client device using speaker embeddings and a machine learning model stored locally on the client device. Further, the operations of method 600 are shown in a particular order, but this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0079]

[0080] In block 602, the system receives audio data that captures utterances of a human speaker and additional sounds that are not from a human speaker. In some implementations, the audio data is streaming audio data. As an example, in block 602, the system can receive the audio data 670 of FIG. 6B, which includes an utterance from "Speaker A" (vertical hatching), as well as an utterance from "Speaker B" (dotted hatching) and "background noise" (hatched hatching).In block 604, the system selects a previously generated speaker embedding for the human speaker. For example, the system can select a previously generated speaker embedding for "Speaker A". For example, the speaker embedding may have been previously generated based on the most recent utterance from "Speaker A" received at the client device that generated the audio data, and may be selected based on "Speaker A" being the speaker of the most recent utterance. Also, for example, the speaker embedding may have been previously generated during an enrollment process performed by "Speaker A" for the automated assistant, the client device, and / or other digital systems. In such cases, the speaker embedding may be selected based on the audio data being generated by the client device and / or via the account of "Speaker A" for the digital system. As one specific example, the audio data received in block 602 may be determined to be from "Speaker A" based on "Speaker A" having been recently verified as an active user with respect to the digital system. For example, voice fingerprinting, image authentication, a passcode, and / or other verification may have been utilized to determine that "Speaker A" is currently active, and as a result, a speaker embedding for "Speaker A" may be selected.

[0081] In block 606, the system sequentially processes the audio data using a machine learning model (e.g., generative model 156) and using the speaker embedding when determining the activation of the layers of the machine learning model.

[0082] In block 608, the system sequentially generates a refined version of the audio data as the direct output of the machine learning model based on the sequential processing of block 606. For example, the system can generate refined audio data 675, shown schematically in FIG. 6, in which only the utterances of "Speaker A" remain.

[0083] The system then optionally executes optional blocks 610, 612, and 614, and / or optional blocks 616, 618, and / or 620.

[0084] In block 610, the system determines whether the audio data includes verbal input from a human speaker corresponding to the speaker embedding of block 604, based on the refined audio data generated in block 608. For example, if the refined audio data is null / zero (e.g., all audio data is below a threshold level), the system can determine that the audio data does not include any verbal input from a human speaker corresponding to the speaker embedding. On the other hand, if the refined audio data includes one or more non-null segments (e.g., above a threshold level), the system can determine that the audio data includes verbal input from a human speaker corresponding to the speaker embedding.

[0085] In an iteration of block 610, if the system determines that the audio data does not include any verbal input from a human speaker corresponding to the speaker embedding, the system proceeds to block 612 and then to block 614, and determines to suppress the ASR and / or other processing of the audio data and the refined audio data. Thus, in those implementations, resources such as the computational resources consumed in the execution of the ASR and / or other processing, and / or the network resources utilized to transmit the audio data (or the refined audio data) when the ASR and / or other processing is executed by a remote system, are not unnecessarily consumed when executing the ASR and / or other processing.

[0086] In some implementations, blocks 610, 612, and 614 can be performed in various situations to ensure that any verbal input in the received audio data is from a particular human speaker before performing various processes (e.g., ASR and / or intent determination) on such audio data and / or before transmitting such audio data from the device where it was first generated. In addition to conserving computational and network resources, this can also promote the privacy of audio data that is not from a particular human speaker and is not intended for further processing by an automated assistant and / or other digital systems. In some of those implementations, blocks 610, 612, and 614 can be performed at least when the client device that received the audio data is in a continuous listening mode when the audio data is received. For example, the device can be in a continuous listening mode after a particular human speaker has previously explicitly invoked the automated assistant via the client device, provided verbal input, and received response content from the automated assistant. For example, the automated assistant can continue to perform limited processing of the audio data for a certain duration after providing at least a portion of the response content. In those implementations, the limited processing can include (or be limited to) blocks 602, 604, 606, 608, 610, 612, and 614 and can be performed to ensure that further processing of the audio data is only performed in response to a "yes" decision in block 612. For example, the speaker embedding in block 614 can be an embedding for a particular human speaker based on the particular human speaker having previously explicitly invoked the automated assistant, thereby ensuring that further processing is only performed if the same particular human speaker provides further utterances.This can prevent waste of resources as described above, and can further prevent incorrect actions from being taken based on utterances from other human speakers who may have been captured by the audio data (e.g., other human speakers coexisting with a particular human speaker, other human speakers on a TV or radio captured within the audio data, etc.).

[0087] In block 616, the system generates text by performing speech-to-text processing on the refined audio data. As described above, block 616 can be executed following block 608 in some implementations, or can be executed only after the execution of block 610 and a determination of "yes" in block 612 in some implementations.

[0088] In block 618, the system performs natural language processing (NLP) on the text generated in block 616. For example, the system can perform NLP to determine the intent of the text (and as a result, the utterance of the human speaker), and optionally one or more parameters related to the intent.

[0089] In block 620, the system generates and provides a response based on the NLP of block 620. For example, if the text is "weather in Los Angeles", the intent from the NLP can be "weather forecast" with a location parameter of "Los Angeles, CA", and the system can provide a structured request for the weather forecast in Los Angeles to a remote system. The system can receive a response in response to the structured request and provide a response for rendering the response to the utterance of the human speaker audibly and / or visually.

[0090] Figure 7A shows an exemplary method 700 for generating multiple refined versions of audio data using audio data, speaker embeddings, and a generative model, according to various implementations disclosed herein. For convenience, the operations of certain aspects of the flowchart of FIG. 7A are described with reference to audio data 770, first refined audio data 775A, and second refined audio data 775B schematically represented in FIG. 7B. Also for convenience, the operations of the flowchart of FIG. 7A are described with reference to a system that performs the operations. This system may include various components of various computer systems, such as a speaker embedding engine 125, a refinement engine 120, and / or one or more GPUs, CPUs, and / or TPUs. In various implementations, one or more blocks of FIG. 7A may be performed by a client device using speaker embeddings and machine learning models stored locally on the client device. Further, the operations of method 700 are shown in a particular order, but this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0091] In block 702, the system receives audio data that captures the utterances of one or more human speakers. As an example, in block 702, the system may receive the audio data 770 of FIG. 7B, which includes utterances from "Speaker A" (vertical hatching) and "Speaker B" (dotted hatching), as well as "background noise" (hatched hatching).

[0092] In block 704, the system selects a portion of the audio data based on the portion being from the first occurrence of voice activity detection within the audio data. For example, the system may select portion 777A of the audio data 770 of FIG. 7B based on portion 777A being from the first segment of voice activity (i.e., the first utterance of "Speaker A").

[0093] In block 706, the system generates a first speaker embedding based on the portion of the audio data selected in block 704. For example, the system can generate the first speaker embedding based on processing portion 777A (FIG. 7B) using a trained speaker embedding model.

[0094] In block 708, the system sequentially processes the audio data using a machine learning model (e.g., generative model 156) and using the first speaker embedding when determining the activation of the layers of the machine learning model.

[0095] In block 710, the system sequentially generates a first refined version of the audio data as the direct output of the machine learning model based on the sequential processing of block 708. For example, the system can generate first refined audio data 775A, shown schematically in FIG. 7B, in which only the utterances of "Speaker A" remain.

[0096] In block 712, the system determines segments of the audio data in which a first human speaker is speaking based on the first refined version of the audio data. For example, the system can determine that the non-null segments of the first refined audio data 775A temporally correspond to the respective segments of the audio data in which the first human speaker is speaking.

[0097] In block 714, the system determines whether there is one or more occurrences of voice activity detection outside of those segments of the audio data in which the first human speaker is speaking. If not, the system can proceed to block 728 (described in more detail below). If so, the system can proceed to block 716. For example, for audio data 770 (FIG. 7B), the system can determine in block 714 that there is an occurrence of voice activity detection outside of the segments in which the first human speaker is speaking, such as portions of the utterance of "Speaker B" that do not overlap with the utterance of "Speaker A".

[0098] In block 716, the system generates a second speaker embedding based on an additional portion of the audio data from the occurrence of voice activity detection within the audio data outside the first segment. For example, the system can select portion 777B of the audio data 770 in FIG. 7B based on the fact that portion 777B is from the occurrence of voice activity and is outside the first segment (where the first segment is determined based on the first refined audio data 775A as described above). Further, for example, the system can generate a second speaker embedding based on portion 777B (e.g., using a speaker embedding model).

[0099] In block 718, the system sequentially processes the audio data using a machine learning model (e.g., generative model 156) and using the second speaker embedding when determining the activation of the layers of the machine learning model.

[0100] In block 720, the system sequentially generates a second refined version of the audio data as the direct output of the machine learning model based on the sequential processing in block 718. For example, the system can generate second refined audio data 775B as schematically shown in FIG. 7B where only the utterances of "Speaker B" remain.

[0101] In block 722, the system determines segments of the audio where the second human speaker is speaking based on the second refined version of the audio data. For example, the system can determine that the non-null segments of the second refined audio data 775B temporally correspond to the respective segments of the audio data where the second human speaker is speaking.

[0102] In block 724, the system determines whether there is one or more occurrences of voice activity detection outside of those segments where the first human speaker is speaking and outside of those segments where the second human speaker is speaking. If not, the system can proceed to block 728 (described in more detail below). If so, the system proceeds to block 726 and can repeat the variations of blocks 716, 718, 720, 722, and 724 (wherein, in the variation, a third speaker embedding and a third refined version of the audio data are generated). This can continue until a "no" decision is made at block 274. As an example of block 724, for audio data 770 (FIG. 7B), the system can determine in block 724 that there are no other utterances from other human speakers (only background noise), so there are no occurrences of voice activity detection outside of the segments where the first human speaker is speaking and where the second human speaker is speaking. Thus, in such an example, the system can proceed to block 728 in the first iteration of block 724.

[0103] In block 728, the system performs further processing based on one or more of the generated refined versions of the audio data (e.g., the first refined version, the second refined version, additional refined versions), and / or based on the identified speaker segments (i.e., the identification of which speaker corresponds to which temporal segment of the audio data). In some implementations, block 728 includes transmitting an indicator of the refined version of the audio data and / or the identified speaker segments to one or more remote systems. In some implementations, block 728 additionally or alternatively includes performing ASR, NLP, and / or other processing. In some versions of those implementations, ASR is performed, the audio data is from audio-visual content, and block 728 includes generating temporally synchronized closed captions for the audio-visual content.

[0104] Turning now to FIGS. 8 and 9, two exemplary environments in which various implementations may be performed are shown. FIG. 8 is first described and includes a client computing device 106 that runs an instance of an automated assistant client 107. One or more cloud-based automated assistant components 180 may be implemented on one or more computing devices (collectively referred to as a “cloud” computing system) communicatively coupled to the client device 106 via one or more local and / or wide area networks (e.g., the Internet), shown generally at 110.

[0105] An instance of the automated assistant client 107 can, through its interaction with one or more cloud-based automated assistant components 180, form, from the user's perspective, something that appears to be a logical instance of an automated assistant 140 with which the user can engage in a human-computer interaction. An example of such an automated assistant 140 is shown in FIG. 8. Thus, in some implementations, it should be understood that each user associated with the automated assistant client 107 running on the client device 106 can, in effect, interact with a logical instance of the automated assistant 140 that is the user's own. For brevity and simplicity, the term "automated assistant" as used herein to refer to what "serves" a particular user often refers to a combination of the automated assistant client 107 running on the client device 106 operated by the user and one or more cloud-based automated assistant components 180 (which can be shared among multiple automated assistant clients of multiple client computing devices). In some implementations, it should also be understood that the automated assistant 140 can respond to requests from any user, regardless of whether the user is actually "served" by that particular instance of the automated assistant 140.

[0106] The client computing device 106 can be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a cellular phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker, a smart home appliance such as a smart TV, and / or a wearable device of the user that includes a computing device (e.g., a user's wristwatch having a computing device, a user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices can be provided. In various implementations, the client computing device 106 can optionally operate one or more other applications, such as a messaging client (e.g., SMS, MMS, online chat), a browser, in addition to the automation assistant client 107. In some of those various implementations, one or more of the other applications can optionally interface with the automation assistant 140 (e.g., via an application programming interface) or can include their own instances of an automation assistant application (which can also interface with the cloud-based automation assistant component 180).

[0107] Automation assistant 104 is involved in a human-computer interaction session with the user via the user interface input and output devices of client device 106. In many situations, in order to protect the user's privacy and / or save resources, the user often has to explicitly invoke automation assistant 140 before the automation assistant fully processes the utterance. The explicit invocation of automation assistant 140 can occur in response to a specific user interface input received at client device 106. For example, the user interface input that can be used to invoke automation assistant 140 via client device 106 can optionally include the actuation of hardware and / or virtual buttons of client device 106. Additionally, the automation assistant client can include one or more local engines 108, such as a call engine operable to detect the presence of one or more verbal call phrases. The call engine can invoke automation assistant 140 in response to the detection of one of the verbal call phrases. For example, the call engine can invoke automation assistant 140 in response to detecting a verbal call phrase such as "Hey Assistant", "OK Assistant", and / or "Assistant". The call engine can continuously process a stream of audio data frames based on the output from one or more microphones of client device 106 to monitor for the occurrence of a verbal call phrase (e.g., when not in an "inactive" mode). While monitoring for the occurrence of a verbal call phrase, the call engine discards any audio data frames that do not contain a verbal call phrase (e.g., after temporarily storing them in a buffer). However, if the call engine detects the occurrence of a verbal call phrase in the processed audio data frame, the call engine can invoke automation assistant 140.As used herein, "invoking" the automation assistant 140 can include activating one or more previously inactive functions of the automation assistant 140. For example, invoking the automation assistant 140 can include causing one or more local engines 108 and / or cloud-based automation assistant components 180 to further process the audio data frame in which the invocation phrase was detected, and / or one or more subsequent audio data frames (whereas prior to the invocation, no further processing of the audio data frames occurred). For example, local and / or cloud-based components can generate a refined version of the audio data and / or perform other processing in response to an invocation of the automation assistant 140. In some implementations, the spoken invocation phrase can be processed to generate a speaker embedding that is used in generating a refined version of the audio data that follows the spoken invocation phrase. In some implementations, the spoken invocation phrase can be processed to identify the account associated with the speaker of the spoken invocation phrase and the stored speaker embedding associated with the account utilized in generating a refined version of the audio data that follows the spoken invocation phrase.

[0108] One or more local engines 108 of the automation assistant 140 are optional and can include, for example, the call engine described above, a local speech-to-text (“STT”) engine (which converts captured audio to text), a local text-to-speech (“TTS”) engine (which converts text to audio), a local natural language processor (which determines the semantic meaning of audio and / or text converted from audio), and / or other local components. Since the client device 106 is relatively constrained with respect to computing resources (e.g., processor cycles, memory, battery, etc.), the local engine 108 may have limited functionality compared to any components included within the cloud-based automation assistant components 180.

[0109] The cloud-based automation assistant components 180 utilize the substantially unlimited resources of the cloud to perform more robust and / or more accurate processing of audio data and / or other user interface inputs compared to any corresponding ones of the local engines 108. Again, in various implementations, the client device 106 can provide audio data and / or other data to the cloud-based automation assistant components 180 in response to the call engine detecting a spoken call phrase or another explicit call of the automation assistant 140.

[0110] The illustrated cloud-based automated assistant component 180 includes a cloud-based TTS module 181, a cloud-based STT module 182, a natural language processor 183, a dialogue state tracker 184, and a dialogue manager 185. The illustrated cloud-based automated assistant component 180 also includes a refinement engine 120 that utilizes a generation model 156 when generating a refined version of audio data and provides the refined version to one or more other cloud-based automated assistant components 180 (e.g., STT module 182, natural language processor 183, dialogue state tracker 184, and / or dialogue manager 185). Further, the cloud-based automated assistant component 180 includes a speaker embedding engine 125 that utilizes a speaker embedding model for various purposes described herein.

[0111] In some implementations, one or more of the engines and / or modules of the automated assistant 140 may be omitted, combined, and / or implemented within a component separate from the automated assistant 140. For example, in some implementations, the refinement engine 120, the generation model 156, the speaker embedding engine 125, and / or the speaker embedding model 152 may be implemented, in whole or in part, on the client device 106. Further, in some implementations, the automated assistant 140 may include additional and / or alternative engines and / or modules.

[0112] The cloud-based STT module 182 can convert audio data to text, which can then be provided to the natural language processor 183. In various implementations, the cloud-based STT module 182 can convert audio data to text based at least in part on a refined version of the audio data provided by the refinement engine 120.

[0113] The cloud-based TTS module 181 can convert text data (e.g., a natural language response formulated by the automation assistant 140) into computer-generated voice output. In some implementations, the TTS module 181 may provide the computer-generated voice output to the client device 106 for direct output by, for example, one or more speakers. In other implementations, the text data (e.g., natural language response) generated by the automation assistant 140 may be provided to one of the local engines 108, and one of the local engines 108 may then convert the text data into computer-generated voice output that is output locally.

[0114] The natural language processor 183 of the automation assistant 140 processes free-form natural language input and generates an annotated output for use by one or more other components of the automation assistant 140 based on the natural language input. For example, the natural language processor 183 can process free-form natural language input that is text input resulting from the conversion by the STT module 182 of audio data provided by the user via the client device 106. The generated annotated output may include one or more annotations of the natural language input and optionally one or more (e.g., all) of the terms of the natural language input.

[0115] In some implementations, the natural language processor 183 is configured to identify and annotate various types of grammatical information within the natural language input. In some implementations, the natural language processor 183 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references within one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), and the like. In some implementations, the natural language processor 183 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more contextual cues. For example, the coreference resolver may be utilized to resolve the term "there" in the natural language input "I liked Hypothetical Cafe last time we ate there" to "Hypothetical Cafe". In some implementations, one or more components of the natural language processor 183 may depend on annotations from one or more other components of the natural language processor 183. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 183 may use related pre-input and / or other related data outside of the particular natural language input to determine one or more annotations.

[0116] In some implementations, the dialogue state tracker 184 may be configured to track a "dialogue state" that includes, for example, the belief state of the goals (or "intentions") of one or more users over a human-computer dialogue session and / or over multiple dialogue sessions. When determining the dialogue state, some dialogue state trackers may attempt to determine the most likely values for the slots instantiated in the dialogue based on the utterances of the user and the system in the dialogue session. Some techniques utilize a fixed ontology that defines a set of slots and a set of values associated with those slots. Some techniques may additionally or alternatively be tailored to individual slots and / or domains. For example, some techniques may require training a model for each slot type in each domain.

[0117] The dialogue manager 185 can be configured to map, for example, the current dialogue state provided by the dialogue state tracker 184 to one or more "response actions" among a plurality of candidate response actions to be executed by the automated assistant 140. The response actions can appear in various forms depending on the current dialogue state. For example, the initial and intermediate dialogue states corresponding to the turns of a dialogue session that occur before the last turn (e.g., when the task desired by the final user is executed) can be mapped to various response actions including the automated assistant 140 outputting additional natural language dialogue. This response dialogue can include, for example, asking the user to provide (i.e., fill a slot) parameters for some action that the dialogue state tracker 184 believes the user intends to perform. In some implementations, the response actions can include actions such as "request" (e.g., looking for parameters to fill a slot), "offer" (e.g., proposing an action or course of action to the user), "select", "notify" (e.g., providing the requested information to the user), "mismatch" (e.g., notifying the user that the user's last input was not understood), commands to peripheral devices (e.g., turning off a light bulb), etc.

[0118] Turning now to FIG. 9, another exemplary environment in which the implementations disclosed herein can be executed is shown. The automated assistant is not included in FIG. 9. Rather, in FIG. 9, the client device 106 does not include an automated assistant client. Further, in FIG. 9, a remote audio processing system 190 is included instead of cloud-based automated assistant components.

[0119] In FIG. 9, client device 106 submits a request 970 via one or more local and / or wide area networks (e.g., the Internet) shown generally at 110. Request 970 is an audio processing request and can be submitted via an API defined for remote audio processing system 190. Request 970 can include audio data and can optionally define the type of audio processing to be performed on the audio data in response to request 970. Remote audio processing system 190 can process requests from client device 106 as well as requests from various other computing devices. Remote audio processing system 190 can be implemented, for example, on a cluster of one or more server devices.

[0120] In response to receiving request 970, remote audio processing system 190 performs one or more audio processing functions, such as those described herein, on the audio data included in the request and returns the audio processing results in the form of a response 971 sent back to client device 106 via network 110. For example, remote audio processing system 190 includes a refinement engine 120 that utilizes generation model 156 when generating a refined version of the audio data, and a speaker embedding engine 125 that utilizes speaker embedding model 152 when generating a speaker embedding. These engines can be utilized in cooperation to generate the speaker diarization results included in response 971 and / or the refined version of the audio data included in response 971. Additionally, remote audio processing system 190 also includes an STT module 182 and a natural language processor 183. Those components can have the same and / or similar functions as described above with respect to FIG. 8. In some implementations, the output generated using those components can be additionally included in response 971.

[0121] FIG. 11 is a block diagram of an exemplary computing device 1110 that may optionally be utilized to execute one or more aspects of the techniques described herein. For example, client device 106 can include one or more components of exemplary computing device 1110, and / or one or more server devices implementing cloud-based automated assistant component 180 and / or remote voice processing system 190 can include one or more components of exemplary computing device 1110.

[0122] Computing device 1110 typically includes at least one processor 1114 that communicates with several peripheral devices via bus subsystem 1112. These peripheral devices can include, for example, storage subsystem 1124, which includes memory subsystem 1125 and file storage subsystem 1126, user interface output device 1120, user interface input device 1122, and network interface subsystem 1116. The input and output devices enable user interaction with computing device 1110. Network interface subsystem 1116 provides an interface to external networks and is coupled to corresponding interface devices within other computing devices.

[0123] User interface input device 1122 can include input devices such as a keyboard, mouse, trackball, pointing device such as a touchpad or graphics tablet, scanner, touch screen incorporated in a display, audio input devices such as a voice recognition system, microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into computing device 1110 or a communication network.

[0124] The user interface output device 1120 may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or another mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 1110 to the user or another machine or computing device.

[0125] The memory subsystem 1124 stores programming and data constructs that provide some or all of the functionality of some of the modules described herein. For example, the memory subsystem 1124 may include logic for performing selected aspects of the methods described herein and / or for implementing the various components shown herein.

[0126] These software modules are generally executed by the processor 1114 alone or in combination with other processors. The memory 1125 used in the memory subsystem 1124 can include several memories, including a main random access memory (RAM) 1130 for storing instructions and data during program execution and a read-only memory (ROM) 1132 in which fixed instructions are stored. The file storage subsystem 1126 can provide permanent storage for program files and data files and can include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functions of a particular implementation can be stored by the file storage subsystem 1126 within the memory subsystem 1124 or within another machine accessible by the processor 1114.

[0127] The bus subsystem 1112 provides a mechanism for the various components and subsystems of the computing device 1110 to communicate with each other as intended. Although the bus subsystem 1112 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0128] The computing device 1110 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 1110 shown in FIG. 11 is intended only as a specific example for the purpose of showing some implementations. Many other configurations of the computing device 1110 with more or fewer components than the computing device shown in FIG. 11 are possible.

[0129] In situations where the specific implementations discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, as well as the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use a user's personal information only when they have received explicit permission from the relevant user to do so.

[0130] For example, the user is provided with control over whether the program or function collects user information about that particular user or other users related to the program or function. Each user from whom personal information is collected is presented with one or more options that enable control of information collection related to that user, in order to provide permission or approval regarding whether the information is to be collected and which parts of the information are to be collected. For example, the user may be provided with one or more such control options via a communication network. Additionally, certain data may be processed in one or more ways before being stored or used such that information that can identify an individual is removed. As an example, the user's identification information may be processed such that it cannot be used to determine information that can identify an individual. As another example, the user's geographical location may be generalized to a larger region such that the user's specific location cannot be identified.

Explanation of Signs

[0131] 106 Client computing device, client device 107 Automated assistant client 108 Local engine 110 Local and / or wide area network, network 120 Refinement engine 125 Speaker embedding engine 126 Speaker embedding 130 Training instance engine 132 Loss module 134 Update module 135 Additional component 140 Automated assistant 152 Speaker embedding model 156 Generation model 157A Hidden layer, layer 157B Hidden layer, layer 157C Hidden layer, layer 160 Audio data database 164A Additional Audio Data A 164B Additional Audio Data B 170 Training Instance Database 170A~170N Training Instances 171A Mixed Instance of Audio Data, Mixed Audio Data A 171B Mixed Audio Data B 172A Embedding 173A Ground Truth Audio Data 174A Loss 175A Predicted Audio Data 180 Automation Assistant Component 181 Cloud-based TTS Module, TTS Module 182 Cloud-based STT Module, STT Module 183 Natural Language Processor 184 Dialogue State Tracker 185 Dialogue Manager 190 Remote Audio Processing System 570 Audio Data 573 Refined Audio Data 670 Audio Data 675 Refined Audio Data 770 Audio Data 775A First Refined Audio Data 775B Second Refined Audio Data 777A Portion 777B Portion 970 Request 971 Response 1070 Input Audio Data, Input Audio 1070 t Time Step 1070 t-15 Time Step 1073 Refined Audio Data 1073 t Time Step 1073 t-15 Time Step 1110 Computing device 1112 Bus subsystem 1114 Processor 1116 Network interface subsystem 1120 User interface output device 1122 User interface input device 1124 Memory subsystem 1125 Memory subsystem, memory 1126 File memory subsystem 1130 Main random access memory (RAM) 1132 Read-only memory (ROM)

Claims

1. A method implemented by one or more processors of a client device, the method comprising: receiving detected audio data via one or more microphones of the client device; the audio data captures speech of a human speaker and also captures one or more additional sounds not from the human speaker; selecting a particular speaker embedding from a plurality of locally stored speaker embeddings for different human speakers, the selecting of the particular speaker embedding being based on the particular speaker embedding corresponding to the human speaker and based on determining that the human speaker is currently an active user of the client device; generating a refined version of the audio data, the refined version of the audio data separating the one or more utterances of the human speaker from the one or more additional sounds not from the human speaker; processing the audio data and the particular speaker embeddings using the neural network model to generate the refined audio data as a direct output of the neural network model; the particular speaker embedding is used in place of any other of the locally stored speaker embeddings in response to selecting the particular speaker embedding; and performing further processing on the refined version of the audio data, said further processing comprising: and performing a speech-to-text process on the refined version of the audio data to generate predictive text for the one or more utterances of the human speaker. method.

2. The method of claim 1, further comprising a step of generating the particular speaker embedding based on one or more enrollment utterances spoken by the human speaker during enrollment with a digital system before the audio data is captured.

3. The method of claim 1, further comprising a step of determining that the human speaker is currently the active user of the client device based on processing of additional audio data preceding the sequence of audio data.

4. The method of claim 3, wherein the additional audio data is an invocation phrase for invoking an automation assistant, and the audio data is received via an automation assistant interface of the automation assistant.

5. The method of claim 4, further comprising the step of generating an automated assistant response based on the predictive text and rendering the automated assistant response on the client device.

6. The method of claim 1, further comprising the step of causing the predictive text to be rendered via a display of the client device. The method of claim 1.

7. The method of claim 1, further comprising a step of determining that the human speaker is currently the active user of the client device based on performing image authentication.

8. One or more microphones; A memory for storing instructions; one or more processors for executing the instructions stored in the memory to perform the method of any one of claims 1 to 7; A client device comprising:

Citation Information

Patent Citations

  • Voice recognition and dialog device and voice recognition and dialog processing method

    JP2005122194A

  • Anchored speech detection and speech recognition

    US20170270919A1

  • Contextual hotwords

    WO2018125292A1